A retrieval-augmented generation (RAG) demo takes an afternoon: chunk some PDFs, embed them, put them in a vector store, and prompt an LLM. Then users arrive, and the answers turn out wrong, slow, or expensive.
Here is what we focus on to close the gap between a demo and a production system.
A production RAG architecture on AWS
- Ingestion pipeline: documents land in S3, and an event-driven pipeline (Lambda or Step Functions) parses, chunks, enriches with metadata, and embeds them.
- Retrieval store: Amazon OpenSearch Serverless, Aurora PostgreSQL with
pgvector, or Bedrock Knowledge Bases if you want a managed option. - Orchestration layer: an API (API Gateway + Lambda or containers) that handles query rewriting, retrieval, re-ranking, prompting and guardrails.
- Models: foundation models served by Amazon Bedrock, so data stays in your AWS account and you can switch models easily.
- Observability & evaluation: log every prompt, retrieved context, response, latency and token cost.
What separates demos from production
Retrieval quality is the product
Most bad answers are retrieval failures, not LLM failures. What helps:
- Chunk by structure (headings, sections, tables), not by a fixed character count
- Hybrid search: combine keyword (BM25) and vector similarity
- Re-ranking of the top results before they reach the prompt
- Metadata filters for tenant, document type, date and access level
Security belongs in retrieval
If a user can’t see a document in the source system, they must not see it in an answer. Enforce document-level permissions at query time through metadata filters. Don’t rely on the prompt to do it.
Evaluate continuously
Build a golden dataset of real questions with expected answers and sources. Score every change (chunking, prompts, models) on:
- Retrieval hit rate: was the right source retrieved?
- Faithfulness: is the answer grounded in that context?
- Answer relevance and completeness
Without evaluation, every change is a guess.
Control cost and latency
- Cache frequent queries and embeddings
- Route simple questions to smaller, cheaper models
- Keep prompts lean, since retrieved context is usually the largest token cost
- Stream responses to improve perceived latency
Add guardrails
Use Bedrock Guardrails (or similar) to filter harmful content, redact PII, and keep the assistant on topic. Also give users a clear way to flag bad answers, which then feed back into the evaluation set.
The takeaway
Production RAG is mostly a data engineering and evaluation problem wrapped around an LLM call. Teams that treat it that way ship assistants people actually trust.
Moving a GenAI prototype toward production? Talk to us about architecture reviews and delivery support.