RAG on AWS: Taking LLM Prototypes to Production

Retrieval-augmented generation demos are easy. Production RAG is not. Here is what changes between the two, and how to build it on Amazon Bedrock.

GenAILLMAmazon BedrockAWS

A retrieval-augmented generation (RAG) demo takes an afternoon: chunk some PDFs, embed them, put them in a vector store, and prompt an LLM. Then users arrive, and the answers turn out wrong, slow, or expensive.

Here is what we focus on to close the gap between a demo and a production system.

A production RAG architecture on AWS

  1. Ingestion pipeline: documents land in S3, and an event-driven pipeline (Lambda or Step Functions) parses, chunks, enriches with metadata, and embeds them.
  2. Retrieval store: Amazon OpenSearch Serverless, Aurora PostgreSQL with pgvector, or Bedrock Knowledge Bases if you want a managed option.
  3. Orchestration layer: an API (API Gateway + Lambda or containers) that handles query rewriting, retrieval, re-ranking, prompting and guardrails.
  4. Models: foundation models served by Amazon Bedrock, so data stays in your AWS account and you can switch models easily.
  5. Observability & evaluation: log every prompt, retrieved context, response, latency and token cost.

What separates demos from production

Retrieval quality is the product

Most bad answers are retrieval failures, not LLM failures. What helps:

  • Chunk by structure (headings, sections, tables), not by a fixed character count
  • Hybrid search: combine keyword (BM25) and vector similarity
  • Re-ranking of the top results before they reach the prompt
  • Metadata filters for tenant, document type, date and access level

Security belongs in retrieval

If a user can’t see a document in the source system, they must not see it in an answer. Enforce document-level permissions at query time through metadata filters. Don’t rely on the prompt to do it.

Evaluate continuously

Build a golden dataset of real questions with expected answers and sources. Score every change (chunking, prompts, models) on:

  • Retrieval hit rate: was the right source retrieved?
  • Faithfulness: is the answer grounded in that context?
  • Answer relevance and completeness

Without evaluation, every change is a guess.

Control cost and latency

  • Cache frequent queries and embeddings
  • Route simple questions to smaller, cheaper models
  • Keep prompts lean, since retrieved context is usually the largest token cost
  • Stream responses to improve perceived latency

Add guardrails

Use Bedrock Guardrails (or similar) to filter harmful content, redact PII, and keep the assistant on topic. Also give users a clear way to flag bad answers, which then feed back into the evaluation set.

The takeaway

Production RAG is mostly a data engineering and evaluation problem wrapped around an LLM call. Teams that treat it that way ship assistants people actually trust.

Moving a GenAI prototype toward production? Talk to us about architecture reviews and delivery support.

Working on something similar?

We help teams design and deliver data and AI systems like the one described here. Let's compare notes.