Building a Production RAG Assistant: 7 Lessons From a Real Product

RAG demos take a weekend. RAG products take discipline. Seven hard won lessons from shipping an LLM assistant that real users depend on every day.

Ijaz KhanSeptember 24, 2026 5 min read
On this page

The demo took a weekend. The product took four months. That gap is this post.

I shipped an LLM assistant into a live healthcare platform, the kind of deployment where users ask real questions about their own data and a wrong answer has consequences. Every tutorial made RAG look like three steps: embed your documents, retrieve the top chunks, stuff them into a prompt. All of that is true and none of it is the job. Here is what the job actually was, organised as the seven lessons I wish someone had handed me on day one.

The architecture, so the lessons have context

Our setup was deliberately boring. Documents flow through a preprocessing worker that cleans, chunks and embeds them into a vector database. At question time we run hybrid retrieval, assemble a grounded prompt, and stream the answer from the model. A FastAPI service owns the pipeline, Celery workers own ingestion, Postgres owns the metadata, and every answer is logged with the exact chunks that produced it.

That last sentence is the most important one in this post. If you cannot see what the model saw, you cannot debug anything.

Streams of data, the raw material every retrieval pipeline has to tame

1. Retrieval quality beats model quality

We spent our first week arguing about which LLM to use. Wrong conversation. When answers were bad, it was almost never the model. It was retrieval. Garbage chunks in, garbage answers out, no matter how clever the model at the end of the pipe.

Before touching model settings, instrument retrieval. Log the query, the retrieved chunks and their scores, then sit down and read fifty of them. It is tedious and it is the highest leverage hour you will spend. You will immediately see the real problems: chunks cut mid sentence, tables shredded into confetti, or the right document ranked sixth when you only pass five.

2. Chunking is a product decision, not a preprocessing step

Fixed 512 token chunks are where quality goes to die. What worked for us:

  • Chunk along document structure. Headings, sections and list boundaries, not token counts. A chunk should be a thought, not a measurement.
  • Prepend context to every chunk. Document title, section path, date. A paragraph that says "this option costs $49" is useless without knowing which option and whether the price sheet is from this year.
  • Store more than you embed. Embed a tight passage for sharp similarity, but store the surrounding context so the model sees enough to answer coherently.

We rebuilt our chunking three times. Each rebuild moved answer quality more than any model upgrade did.

3. Hybrid search saves you from embedding blind spots

Pure vector search fails on exact identifiers: SKUs, error codes, invoice numbers, people's names. A user asking about "ERR_4012" does not want the semantically similar error. They want that error.

Combine vector similarity with keyword search and merge the results. Most vector databases support this natively now, so there is no excuse to skip it. Our rule of thumb after measuring: keyword search catches roughly one in five queries that pure vectors fumble, and those tend to be the angriest users.

4. Ground the model, then make refusal a feature

Our system prompt evolved into a contract: answer only from the provided context, and if the answer is not there, say so and point to a human. Users trust an assistant that says "I don't have that information" far more than one that hallucinates confidently. Every hallucination costs you ten correct answers' worth of trust, and in regulated domains it costs more than trust.

We also added citations. Each answer links back to the source passages. It looks like a UX nicety. It is actually a trust machine, and it turned our support team from skeptics into the feature's loudest advocates.

5. Build the eval set before you think you need it

We collected around 100 real question and answer pairs from support logs and stakeholders, then scored every pipeline change against them. It turned "this feels better" into "this improved answer accuracy from 71 percent to 84 percent".

Without an eval set you are doing vibes driven engineering, and vibes do not survive a prompt tweak that silently breaks a use case you forgot existed. The eval set does not need to be fancy. A spreadsheet and a script that runs the pipeline against it is enough to start. What matters is that it exists before the first "quick prompt improvement" lands.

6. Latency budgets change your architecture

Cold retrieval plus a large model plus streaming still needs to feel instant. What helped, in order of impact:

  • Stream tokens from the first byte. Perceived latency matters more than total latency, and a stream that starts in 400 milliseconds feels faster than a complete answer in three seconds.
  • Cache embeddings for repeated queries. Users ask the same things, constantly.
  • Run retrieval and any classification steps in parallel, not sequentially.
  • Keep a smaller, faster model for query rewriting and routing. Save the big model for the final answer.

7. Cost is an engineering metric

Track tokens per answer from day one. We cut our per conversation cost by more than half without touching quality: deduplicating retrieved chunks, trimming boilerplate out of the system prompt, and capping history to what the conversation actually needed. None of that was glamorous. All of it compounded.

The mistakes I see most often

Three patterns show up in almost every struggling RAG project I get called into. Teams index everything instead of curating what deserves to be retrievable. Teams skip the eval set and then argue about anecdotes. And teams treat the system prompt as a magic incantation to be endlessly tuned, when the real fix is almost always one layer down, in retrieval.

The takeaway

RAG is not a library you install. It is a pipeline you own. Retrieval, chunking, grounding, evaluation, latency and cost: each one is unglamorous, and each one is the difference between a demo and a product people rely on.

Building an AI assistant for your product? This is exactly the kind of work I do. Let's talk.

#rag#llm#vector-databases#ai#production

Need this built?

These services can turn the ideas in this article into production software.

Keep reading

All articles
For Founders6 min read

Adding AI to Your SaaS: Scope, Risks and Launch Checklist

A founder’s guide to adding AI to an existing SaaS: choose one workflow, define permissions and acceptance criteria, then plan a controlled rollout.

Read article
AI Engineering4 min read

Serving a Custom ML Model in Production: The Plumbing Nobody Talks About

The data scientists hand you a model that works in a notebook. Making it survive real traffic, uploads, GPUs, queues and versioning is a different job entirely.

Read article