jellyace
← All posts
AI·3 min read

RAG that actually works: building an assistant over your documents

Why most retrieval-augmented generation demos disappoint in production, and the chunking, retrieval and evaluation practices that fix them.

Retrieval-augmented generation (RAG) lets a language model answer questions using your documents — help-centre articles, internal wikis, contracts, support tickets — without training a custom model. Before answering, the system finds the most relevant passages and gives them to the model as context.

A basic demo takes an afternoon. Getting reliable answers for real users takes more care. Here's what makes the difference.

The three stages

RAG pipeline: ingestion splits, embeds and stores chunks in an index when documents change; for every question, retrieval finds and reranks chunks, and generation answers from them with citations.

  1. Ingestion — split documents into chunks, create embeddings, and store them with metadata.
  2. Retrieval — for each question, find the most relevant chunks.
  3. Generation — ask the model to answer using only those chunks, and cite them.

Most failures happen at retrieval. If the right passage never reaches the model, even the best model can only guess — and it will guess confidently.

Chunk by structure, not by character count

Splitting every 1,000 characters cuts sentences, tables and sections in half. Instead:

  • Split along the document's own structure: headings, sections, list items.
  • Keep chunks big enough to make sense on their own, with a little overlap.
  • Store metadata with every chunk: source document, title, section, date, and who's allowed to see it.

Adding the document and section title to each chunk's text also helps, because a paragraph often doesn't mention what it's about.

Don't rely on vector search alone

Embeddings are good at meaning, but can miss exact terms — product codes, error messages, names. Two changes make a large difference:

  • Hybrid search: combine keyword search with vector search, so exact matches aren't lost.
  • Reranking: retrieve more candidates than you need, then use a reranking model to pick the best few.

Keyword search finds exact terms such as error codes and SKUs; vector search finds similar meaning; a reranker re-scores the merged results and keeps the best few for the model.

Use metadata filters too, for example to restrict answers to one product or to current documents only.

Enforce permissions at retrieval time

If some documents are restricted, filter them out before they reach the model. Telling the model "don't reveal confidential information" in a prompt is not a security boundary — if the text is in the context, assume it can come out.

Make it cite its sources

Instruct the model to answer only from the provided context, to cite which chunks it used, and to say clearly when the answer isn't there. Citations let users verify answers, and "I don't know" is far better than a plausible invention.

Evaluate before you tune

Build a test set of 50–100 real questions with the expected answer or source document. Then measure two things separately:

  • Retrieval: does the correct chunk appear in the top results?
  • Answers: is the response correct, grounded in the sources, and properly cited?

Rerun this every time you change chunking, embeddings, prompts or models. Without it, you're tuning by feel.

Measuring the two separately tells you where to look:

Retrieval Answers What it means Fix
Good Good Working Keep the test set growing
Good Bad The model ignores or misreads the chunks Prompt, model, or chunk formatting
Bad Bad The right passage never arrives Chunking, hybrid search, reranking, filters
Bad Good Lucky, or answering from its own knowledge Check citations; tighten "answer only from context"

Where to store embeddings on AWS

  • pgvector on Amazon RDS or Aurora PostgreSQL — a great default if you already use Postgres; your vectors live next to your data.
  • Amazon OpenSearch Service — strong when you need heavy hybrid search and filtering.
  • Amazon Bedrock Knowledge Bases — a managed pipeline that handles ingestion and retrieval for you, with less control over the details.

Update (October 2026): Amazon S3 Vectors (generally available since December 2025) adds low-cost vector storage that works with Bedrock Knowledge Bases and OpenSearch. Bedrock Knowledge Bases now also offers a fully managed option with agentic retrieval.

The database matters less than people expect. Chunking, retrieval quality and evaluation are what make users trust the answers.

Want this done for you?See our AI services

Have something you need built?

Tell us a bit about your product and what you’re trying to get done. You’ll hear back from an engineer, not a sales team — no obligation.

hello@jellyace.net