RAG vs. fine-tuning: choosing an enterprise LLM architecture
A practical decision guide for teams putting private data behind a language model.
Once a team decides to put its own data behind a language model, the first architectural question is almost always the same: retrieve or retrain?
Since million-token context windows arrived, there is a third answer: skip both and put the documents straight into the prompt. Each of the three wins somewhere, and many production systems combine two.
The short answer: retrieve facts, tune behavior
Use retrieval-augmented generation (RAG) to give a model knowledge, and fine-tuning to change how it behaves. Long context sits in between: it works well when the knowledge is small and stable enough to send with every request.
The research backs the first half of that rule. When Ovadia and colleagues compared the two methods for injecting knowledge, RAG consistently outperformed unsupervised fine-tuning, both for facts the model had seen in training and for entirely new ones. Models struggle to learn new facts from fine-tuning alone. Fine-tuning earns its place elsewhere: format, tone, classification and cost.
When retrieval (RAG) wins
- Your knowledge changes weekly or daily.
- Answers must cite a source document.
- Access rules differ by user or team.
RAG searches your documents at request time and gives the model only the passages it needs. That has three practical effects. Updating knowledge means re-indexing a document, not retraining a model. Every answer can point back to the passage it came from, which is what auditors, lawyers and analysts ask for first. And because retrieval runs per request, it can filter by the user’s permissions before the model sees anything.
Our Relm build shows the pattern. Analysts upload rent rolls, offering memorandums, P&Ls and leases; Relm reads them with OCR, indexes them for vector search, and links every finding in its reports back to the record it came from. That traceability is the reason to choose retrieval: a commercial real estate analyst will not act on a number nobody can check.
When long context is enough
If the material is small and stable, skip the retrieval pipeline and put it in the prompt. Anthropic’s guidance is specific: a knowledge base under about 200,000 tokens, roughly 500 pages, can simply be included in the prompt, and prompt caching makes that much faster and cheaper on repeated requests.
Long context has limits that matter in enterprise work:
- Cost per query: you pay for every token on every request, so a large context sent thousands of times a day gets expensive, even with caching.
- Attention: the “Lost in the Middle” study by Liu and colleagues found that models use information at the start or end of a long input much better than information buried in the middle.
- Permissions: one shared prompt cannot easily respect different access rules for different users.
- Citations: pointing to the exact source passage is harder when the model read everything at once.
Long context is a good fit for a policy assistant over one handbook, or a contract review tool that reads one agreement at a time. It is a poor fit for a company-wide assistant over thousands of permissioned documents.
When fine-tuning wins
- You need a consistent format, tone or classification behavior.
- Latency or cost demands a smaller model.
- The task is narrow and examples are plentiful.
Fine-tuning trains a model on examples of the inputs and outputs you want. It is the right tool when prompting cannot reliably produce a strict JSON schema, a house writing style, or a classification label across thousands of edge cases. It can also move a task that works on a large model to a smaller, cheaper and faster one, which matters for voice agents and high-volume pipelines.
OpenAI’s own guidance puts fine-tuning last in the sequence: start with evaluations and prompt engineering, which may be all you need, and fine-tune when prompting alone does not deliver the behavior or performance required. We follow the same order. Fine-tuning before you have an evaluation set means you cannot tell whether it helped.
A decision table for enterprise LLM architecture
| Situation | Best fit | Why |
|---|---|---|
| Large document set that changes often | RAG | Re-index instead of retraining; cost grows with the passages retrieved, not the whole corpus |
| Answers must cite their source | RAG | Each answer links to the passages it used |
| Different users may see different documents | RAG | Filter by permission before the model reads anything |
| One small, stable reference set | Long context | No retrieval pipeline to build; prompt caching keeps repeat queries cheap |
| Reasoning across a whole document at once | Long context | The model sees the full text, not fragments |
| Strict output format or house style | Fine-tuning | Teaches behavior that prompts do not hold reliably |
| High-volume task on a tight latency or cost budget | Fine-tuning | Moves the task to a smaller, faster model |
| Company assistant over permissioned knowledge | RAG plus prompting | Facts from retrieval, behavior from instructions; tune later only if evaluations show a gap |
Why production systems often combine them
“Retrieval for facts, instructions for behavior, and tuning only where evaluations show a gap.”
The approaches are not rivals. A typical enterprise assistant retrieves the relevant passages, sends them in a context window large enough to keep surrounding detail, and relies on careful instructions, or occasionally a tuned model, for format and tone. The architecture question is really about which part of the problem each method should own.
What each approach costs to run
The three approaches spend money in different places. RAG costs engineering up front, for ingestion, chunking, search and permissions, and then stays relatively cheap per query, because the model reads only a few passages. Long context costs almost nothing to build and more on every request, because the model reads the whole context each time; caching softens that for repeated content but does not remove it. Fine-tuning costs data preparation and training runs up front, then can be the cheapest per request if it lets you move to a smaller model.
Latency follows the same pattern. Retrieval adds a search step, usually small next to generation. Very long prompts slow the first token. A smaller tuned model is often the fastest option of all, which is why fine-tuning matters most in latency-sensitive products such as voice agents.
Retrieval quality decides how accurate a RAG system is
When a RAG system gives a wrong answer, the cause is usually that it retrieved the wrong passage, not that the model reasoned badly. Retrieval is worth engineering properly. Anthropic reported that adding context to each chunk before indexing, combined with keyword search, reduced top-20 retrieval failures by 49% in its tests, and by 67% with a reranking step. The exact numbers will differ on your data, but the lesson holds: chunking, hybrid search and reranking move accuracy more than swapping the model.
Build evaluations first
Whichever you choose, build evaluations first. A set of real questions with expected answers is what lets you swap models safely when a better or cheaper one ships.
Start with fifty to a hundred questions your users actually ask, the documents that answer them, and what a correct answer looks like. Score retrieval and answers separately, so you know which part to fix. Run the set on every change to prompts, chunking, models or tuning data. It is the cheapest part of the project and the one that makes every later decision measurable.
If you are choosing an architecture for private data, our enterprise LLM and RAG development work starts with exactly this: your documents, your access rules and an evaluation set, before any model is chosen, and RAG development cost explains what drives the price. For why so many AI projects stall after the demo, read why most AI pilots never reach production.
Sources
- Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks, arXiv (Lewis et al., 2020)
- Fine-Tuning or Retrieval? Comparing Knowledge Injection in LLMs, arXiv (Ovadia et al., 2023)
- Lost in the Middle: How Language Models Use Long Contexts, arXiv (Liu et al., 2023)
- Introducing Contextual Retrieval, Anthropic
- Model optimization, OpenAI API docs
Enterprise LLM & RAG
Private LLM architectures, retrieval (RAG) systems and fine-tuned models that answer from your data, cite their sources and never leak it.



