Engineering · · 6 min

RAG vs. fine-tuning: choosing an enterprise LLM architecture

A practical decision guide for teams putting private data behind a language model.

Once a team decides to put its own data behind a language model, the first architectural question is almost always the same: retrieve or retrain?

Since million-token context windows arrived, there is a third answer: skip both and put the documents straight into the prompt. Each of the three wins somewhere, and many production systems combine two.

The short answer: retrieve facts, tune behavior

Use retrieval-augmented generation (RAG) to give a model knowledge, and fine-tuning to change how it behaves. Long context sits in between: it works well when the knowledge is small and stable enough to send with every request.

The research backs the first half of that rule. When Ovadia and colleagues compared the two methods for injecting knowledge, RAG consistently outperformed unsupervised fine-tuning, both for facts the model had seen in training and for entirely new ones. Models struggle to learn new facts from fine-tuning alone. Fine-tuning earns its place elsewhere: format, tone, classification and cost.

When retrieval (RAG) wins

  • Your knowledge changes weekly or daily.
  • Answers must cite a source document.
  • Access rules differ by user or team.

RAG searches your documents at request time and gives the model only the passages it needs. That has three practical effects. Updating knowledge means re-indexing a document, not retraining a model. Every answer can point back to the passage it came from, which is what auditors, lawyers and analysts ask for first. And because retrieval runs per request, it can filter by the user’s permissions before the model sees anything.

Our Relm build shows the pattern. Analysts upload rent rolls, offering memorandums, P&Ls and leases; Relm reads them with OCR, indexes them for vector search, and links every finding in its reports back to the record it came from. That traceability is the reason to choose retrieval: a commercial real estate analyst will not act on a number nobody can check.

When long context is enough

If the material is small and stable, skip the retrieval pipeline and put it in the prompt. Anthropic’s guidance is specific: a knowledge base under about 200,000 tokens, roughly 500 pages, can simply be included in the prompt, and prompt caching makes that much faster and cheaper on repeated requests.

Long context has limits that matter in enterprise work:

  • Cost per query: you pay for every token on every request, so a large context sent thousands of times a day gets expensive, even with caching.
  • Attention: the “Lost in the Middle” study by Liu and colleagues found that models use information at the start or end of a long input much better than information buried in the middle.
  • Permissions: one shared prompt cannot easily respect different access rules for different users.
  • Citations: pointing to the exact source passage is harder when the model read everything at once.

Long context is a good fit for a policy assistant over one handbook, or a contract review tool that reads one agreement at a time. It is a poor fit for a company-wide assistant over thousands of permissioned documents.

When fine-tuning wins

  • You need a consistent format, tone or classification behavior.
  • Latency or cost demands a smaller model.
  • The task is narrow and examples are plentiful.

Fine-tuning trains a model on examples of the inputs and outputs you want. It is the right tool when prompting cannot reliably produce a strict JSON schema, a house writing style, or a classification label across thousands of edge cases. It can also move a task that works on a large model to a smaller, cheaper and faster one, which matters for voice agents and high-volume pipelines.

OpenAI’s own guidance puts fine-tuning last in the sequence: start with evaluations and prompt engineering, which may be all you need, and fine-tune when prompting alone does not deliver the behavior or performance required. We follow the same order. Fine-tuning before you have an evaluation set means you cannot tell whether it helped.

A decision table for enterprise LLM architecture

Which approach fits which situation
SituationBest fitWhy
Large document set that changes oftenRAGRe-index instead of retraining; cost grows with the passages retrieved, not the whole corpus
Answers must cite their sourceRAGEach answer links to the passages it used
Different users may see different documentsRAGFilter by permission before the model reads anything
One small, stable reference setLong contextNo retrieval pipeline to build; prompt caching keeps repeat queries cheap
Reasoning across a whole document at onceLong contextThe model sees the full text, not fragments
Strict output format or house styleFine-tuningTeaches behavior that prompts do not hold reliably
High-volume task on a tight latency or cost budgetFine-tuningMoves the task to a smaller, faster model
Company assistant over permissioned knowledgeRAG plus promptingFacts from retrieval, behavior from instructions; tune later only if evaluations show a gap

Why production systems often combine them

“Retrieval for facts, instructions for behavior, and tuning only where evaluations show a gap.”

The approaches are not rivals. A typical enterprise assistant retrieves the relevant passages, sends them in a context window large enough to keep surrounding detail, and relies on careful instructions, or occasionally a tuned model, for format and tone. The architecture question is really about which part of the problem each method should own.

What each approach costs to run

The three approaches spend money in different places. RAG costs engineering up front, for ingestion, chunking, search and permissions, and then stays relatively cheap per query, because the model reads only a few passages. Long context costs almost nothing to build and more on every request, because the model reads the whole context each time; caching softens that for repeated content but does not remove it. Fine-tuning costs data preparation and training runs up front, then can be the cheapest per request if it lets you move to a smaller model.

Latency follows the same pattern. Retrieval adds a search step, usually small next to generation. Very long prompts slow the first token. A smaller tuned model is often the fastest option of all, which is why fine-tuning matters most in latency-sensitive products such as voice agents.

Retrieval quality decides how accurate a RAG system is

When a RAG system gives a wrong answer, the cause is usually that it retrieved the wrong passage, not that the model reasoned badly. Retrieval is worth engineering properly. Anthropic reported that adding context to each chunk before indexing, combined with keyword search, reduced top-20 retrieval failures by 49% in its tests, and by 67% with a reranking step. The exact numbers will differ on your data, but the lesson holds: chunking, hybrid search and reranking move accuracy more than swapping the model.

Build evaluations first

Whichever you choose, build evaluations first. A set of real questions with expected answers is what lets you swap models safely when a better or cheaper one ships.

Start with fifty to a hundred questions your users actually ask, the documents that answer them, and what a correct answer looks like. Score retrieval and answers separately, so you know which part to fix. Run the set on every change to prompts, chunking, models or tuning data. It is the cheapest part of the project and the one that makes every later decision measurable.

If you are choosing an architecture for private data, our enterprise LLM and RAG development work starts with exactly this: your documents, your access rules and an evaluation set, before any model is chosen, and RAG development cost explains what drives the price. For why so many AI projects stall after the demo, read why most AI pilots never reach production.

Related service

Private LLM architectures, retrieval (RAG) systems and fine-tuned models that answer from your data, cite their sources and never leak it.

Want this applied to your business?
Free 45-min AI audit with a senior architect.
Book the audit
FAQ

Common questions.

Is RAG better than fine-tuning?

For adding knowledge, usually yes. In a 2023 comparison by Ovadia and colleagues, retrieval-augmented generation consistently outperformed unsupervised fine-tuning at injecting both existing and new facts. Fine-tuning is the better tool for changing behavior, such as output format, tone or classification.

When should you fine-tune an LLM instead of using RAG?

When the problem is how the model responds, not what it knows: a strict output format, a consistent tone, a narrow classification task with plenty of labeled examples, or moving a task to a smaller, cheaper and faster model. Try careful prompting and an evaluation set first.

Do long context windows make RAG obsolete?

No. If your material fits comfortably in the prompt and rarely changes, long context with prompt caching can be simpler than RAG. For large, changing or permissioned document sets, retrieval is still cheaper per query, easier to cite and easier to secure.

Can you combine RAG and fine-tuning?

Yes, and many production systems do. Retrieval supplies current facts and citations at request time, while prompting or light fine-tuning shapes format and behavior. An evaluation set of real questions lets you change either part safely.