ModelRefs / Retrieval-Augmented Generation — Architecture Blueprint
Retrieval-Augmented Generation — Architecture Blueprint
Production architecture blueprint for Retrieval-Augmented Generation: components, deployment patterns, cost & latency, failure modes, evaluation and governance, with sources and review dates.
Overview
This is the implementation view of Retrieval-Augmented Generation: the components it requires, where it can run, what it costs in latency and spend, how it fails, and what you must measure before putting it in front of users.
4 components to assemble, 4 documented failure modes, high implementation complexity. Every statement below comes from the canonical workflow record with its sources and review date; where the evidence does not settle a question, the page says so rather than filling the gap.
What this workflow takes in and produces
Takes in
- documents
- knowledge-base records
- user queries
Produces
- grounded answers
- citations
- retrieval diagnostics
Applied to
- grounded question answering
- enterprise search
- knowledge assistants
Components you need to assemble
A working implementation needs 4 distinct components. Each is a build-or-buy decision in its own right.
- document ingestion
- embedding model
- retrieval index
- evaluation harness
Implementation complexity: high. This describes the integration and evaluation effort, not the difficulty of any single component.
Deployment patterns
Deployment options recorded for this workflow: managed-api, self-hosted, hybrid.
Topologies it has been recorded against: serverless-api, managed-container, hybrid-private-cloud. Each changes the data-residency, scaling and cost profile, so confirm the one you need against current provider documentation.
Cost and latency
- Retrieval, reranking, and generation each add latency and cost.
- Measure end-to-end latency on the target corpus and query mix.
How this workflow fails
Observed failure modes for this class of workflow. Design a check for each one before shipping, not after.
- retrieval miss
- irrelevant context
- unsupported answer
- stale index
Risk areas the evidence covers
- retrieval quality
- groundedness
- citation support
Proving it works before you ship
Evaluation readiness: Partial — Retrieval and answer-quality evaluation is defined conceptually; workload-specific acceptance thresholds remain required.
Worked evaluation case: Enterprise knowledge assistant
Answer employee questions over a permissioned internal corpus with citations and an abstention path.
What to measure
- retrieval recall and precision
- answer grounding and citation accuracy
- permission-filter behavior
- unsupported-answer and abstention rate
- end-to-end latency and cost
Governance and data handling
- Enforce source permissions during retrieval.
- Review retention and provider data handling for private corpora.
Implementation notes
- Evaluate retrieval quality separately from answer quality so a retrieval miss is not hidden by fluent generation.
- Use held-out questions that reflect the target corpus, user distribution, and permission model.
What this blueprint does not establish
- The canonical record does not prescribe a vector database, chunking policy, or production acceptance threshold.
- Compatibility links do not prove quality on a specific corpus.
Source coverage: Partial — The primary RAG paper supports retrieval-grounded generation, and official evaluation guidance supports task-specific held-out evaluation; production controls and acceptance thresholds remain workload-specific.
Registry relationships reviewed 2026-06-28; mutable provider and tool behavior requires current primary-documentation review.
Sources
- Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks Facebook AI Research and research collaborators · primary-research · accessed 2026-06-28
- Evaluation best practices OpenAI · provider-reported · accessed 2026-06-28
Candidate models and benchmarks
Candidate models with published references, the providers behind them, and the benchmarks whose task shape bears on this workflow are on the Retrieval-Augmented Generation workflow reference. This blueprint covers implementation; that page covers selection.
Continue your research
Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to Retrieval-Augmented Generation — Architecture Blueprint.