Dev.to AI πŸ€– Ai πŸ‘ 0 πŸ“– 3 min read

How I Built a Self-Improving AI Reply System for E-commerce Sellers with pgvector

Marketplace sellers in Turkey get a constant stream of customer questions: "Is this table waterproof?", "Will it fit a small balcony?", "Why are the reviews so bad?". Every unanswered question is a lost sale, but writing

Marketplace sellers in Turkey get a constant stream of customer questions: "Is this table waterproof?", "Will it fit a small balcony?", "Why are the reviews so bad?". Every unanswered question is a lost sale, but writing thoughtful replies all day doesn't scale.

While building Netaliz, a profit analytics tool for Trendyol sellers, I added an AI module that drafts replies to these questions. The interesting part isn't calling an LLM. It's making the system get better every time the seller approves an answer.

Here's how it works.

The problem with "just prompt it"

My first version was simple: send the product info and the question to an LLM, get an answer back. It worked, but:

  • Answers sounded generic, not like the seller's brand.
  • The model didn't know seller-specific rules ("always say shipping takes 1–3 business days").
  • It made the same stylistic mistakes over and over, because nothing was learned.

Sellers were editing almost every draft. That's not automation, that's extra work.

The architecture

The final system has four layers that get assembled into the prompt:

  1. Product context: material, dimensions, warranty, pulled from the marketplace API.
  2. Brand voice: the seller picks a tone (friendly, corporate, or sales-focused) and defines banned words.
  3. Template rules: hard instructions like "for shipping questions, say 1–3 business days".
  4. Few-shot examples: previously approved answers to similar questions, retrieved with vector search.

Layer 4 is what makes it self-improving.

Storing approved answers with pgvector

Every time a seller approves a draft (or edits and sends it), we store the question, the final answer and an embedding of the question:

CREATE EXTENSION IF NOT EXISTS vector;

CREATE TABLE approved_answers (
  id          BIGSERIAL PRIMARY KEY,
  store_id    BIGINT NOT NULL,
  question    TEXT NOT NULL,
  answer      TEXT NOT NULL,
  embedding   vector(1536),
  created_at  TIMESTAMPTZ DEFAULT now()
);

CREATE INDEX ON approved_answers
  USING hnsw (embedding vector_cosine_ops);

Keeping this in Postgres instead of a separate vector database was a deliberate choice. The data already lives there, it's scoped per store with a simple WHERE, and it's one less service to run.

Retrieving similar examples

When a new question comes in, we embed it and pull the closest approved answers from the same store:

SELECT question, answer
FROM approved_answers
WHERE store_id = $1
ORDER BY embedding <=> $2
LIMIT 3;

Those examples go into the prompt as few-shot demonstrations. If a seller always answers sizing questions in a particular way, the model sees that pattern and follows it, without any fine-tuning.

Structuring the answer

We also give the model a fixed three-step structure, which made answers noticeably more persuasive:

  1. Acknowledge the concern: show the customer they were heard.
  2. Give a concrete argument: a specific fact about the product (material, warranty, dimensions).
  3. Close with trust: offer further help.

A simplified version of the prompt assembly:

function buildPrompt(ctx: ReplyContext): string {
  return [
    `You are a customer support writer for a marketplace seller.`,
    `Tone: ${ctx.brandVoice}. Never use: ${ctx.bannedWords.join(", ")}.`,
    `Rules:\n${ctx.templateRules.map(r => `- ${r}`).join("\n")}`,
    `Product facts:\n${JSON.stringify(ctx.product)}`,
    `Structure: acknowledge the concern, give one concrete argument, close with trust.`,
    `Examples of approved answers:\n${ctx.examples
      .map(e => `Q: ${e.question}\nA: ${e.answer}`)
      .join("\n\n")}`,
    `Customer question: ${ctx.question}`,
  ].join("\n\n");
}

Keeping humans in control

By default, nothing is sent automatically. The AI drafts, the seller reviews and clicks send. Auto-send is opt-in, and even then the brand voice, banned words and template rules still apply.

This turned out to matter for two reasons: sellers trust the system more, and every human approval becomes a new training example. The review step is the learning loop.

What I learned

  • Retrieval beats fine-tuning for per-customer style. Each store gets its own "memory" instantly, with no training pipeline.
  • Postgres + pgvector is enough at this scale. Don't add infrastructure you don't need.
  • Human-in-the-loop is a feature, not a limitation. It builds trust and generates your best data.

If you're curious about the product side, Netaliz has a demo store with sample data, no signup needed: netaliz.com/demo.

I'd love to hear how others are handling per-user style in LLM apps. Are you using retrieval, fine-tuning, or something else?

πŸ“° Read the original article on Dev.to AI

Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes β€” full credit and traffic to the original publisher.