Buying software · · 8 min

RAG development cost in 2026: what drives the price

Why data preparation, permissions and evaluation, not the vector database, decide what a retrieval-augmented generation system costs to build and run.

An internal assistant that answers 10,000 questions a month from 20,000 pages of company documents costs about $80 to $210 a month in tokens and storage. Indexing all 20,000 pages costs about 20 cents. The vector database and the model, the parts most RAG (retrieval-augmented generation) cost guides price, are the cheap parts.

The money goes into getting your documents into a state where retrieval works, enforcing who can see what, and proving the answers are right. So we price a RAG system the way we scope one: by what the data and the permissions demand.

How much does RAG development cost?

A production RAG system costs $25k–$120k to build with us, as one fixed price. We set it at the end of a two-week Discovery Sprint ($5,000) that tests retrieval on a sample of your real documents, so the quote reflects the data you have rather than a clean demo set.

Running it is the small number. The worked example below lands at about $80 to $210 a month at 10,000 queries, before hosting and monitoring.

What a RAG system is made of, and what each part costs

A RAG system has six parts. The first and the last take the most engineering; the middle ones are mostly configuration and usage fees.

RAG components: build effort and running cost
ComponentWhat it doesBuild effortRunning cost
Ingestion and cleanupPulls documents from each source, runs OCR, splits tables and text, removes duplicates and stale versionsHighLow, mostly compute during syncs
Chunking and embeddingsSplits text into passages and turns each into a vectorMediumVery low: text-embedding-3-small lists at $0.02 per million tokens
Vector store and searchStores vectors and finds the passages closest to a question, often with keyword search alongsideLow to medium$0 extra in Postgres with pgvector; Pinecone Standard has a $50 monthly minimum
PermissionsFilters results to what each user is allowed to seeMedium to highNegligible
Generation with citationsSends the question and passages to a model and returns an answer with links to sourcesMediumThe main usage cost: model tokens
Evaluation and monitoringTests answers against known questions and tracks quality after launchHighLow, scheduled test runs

RAG cost by scope

The scope of a RAG system is set by three questions: how many sources, how messy the documents are, and whether different users may see different things. Each answer pushes a project toward one row of the table below. The rows are our own placement of typical scopes inside the published $25k–$120k band, not quotes; the real number comes out of discovery once we have seen your documents.

Three RAG scopes and where each sits in our price band
ScopeTypical shapeWhere it usually lands
Single-source assistantOne document library, mostly digital text, everyone can see everything, answers with citationsBottom of the band, from $25k
Multi-source with permissionsSeveral systems such as a drive, a wiki and a ticketing tool, scheduled syncs, per-user access rulesMid-band
Document intelligence and agentic RAGScans and complex tables, structured extraction, agents that call tools, tenant isolationTop of the band, up to $120k, or split into phases

What drives the price of a RAG system?

Six things move a RAG quote. The vector database and the model are not on the list, because they barely change the build.

Data preparation is the largest cost

Retrieval can only find what ingestion made findable. Scanned PDFs need OCR and layout-aware extraction; tables need to be kept intact rather than split mid-row; duplicate and outdated versions need to be removed or they will be cited as if current. The messier the documents, the more of the build goes here, which is why discovery tests real files.

Number of sources and how fresh they must be

Each source needs a connector, authentication and a sync schedule. Nightly batch syncs are simple; near-real-time updates need change tracking and more careful re-indexing. Five sources cost meaningfully more than one, even if the total volume of text is the same.

Permissions

If everyone can see every document, permissions are free. If a salesperson must never see HR files, every passage needs access metadata copied from its source system, and every query must be filtered by it. Getting this wrong is a data leak, so it needs its own tests.

Retrieval quality

Basic vector search is quick to set up but misses exact terms such as product codes and names. Combining it with keyword search and adding a reranking step costs more to build and pays back in accuracy. In Anthropic’s published tests, contextual embeddings plus keyword search cut failed retrievals by 49%, and adding reranking took the cut to 67%. Preparing the contextualized chunks cost about $1.02 per million document tokens, once.

Evaluation

An evaluation set is a list of real questions with known correct answers and source passages. It is how you know a change made things better rather than worse. Building it takes time with your subject-matter experts, and it is the part most often skipped in pilots.

Hosting and compliance

Running in your own cloud, using zero-retention model endpoints or self-hosted models, keeping data in a region, and logging access for audits all add work. For regulated data, such as health or financial records, these are not optional.

Worked example: monthly RAG run cost at 10,000 queries

Here is the arithmetic for a mid-sized internal knowledge assistant, using list prices from each vendor’s pricing page.

Assumptions: 20,000 pages of documents at about 500 tokens a page, so 10 million tokens to index. 10,000 queries a month. Each query sends about 6,000 input tokens (the question, instructions and retrieved passages) and gets back about 400 output tokens.

RAG run cost, 10,000 queries a month (list prices)
Line itemCalculationCost
Index the documents (one-time)10M tokens × $0.02 per million (text-embedding-3-small)about $0.20
Contextual chunks (optional, one-time)10M tokens × about $1.02 per million (Anthropic’s figure)about $10
Embed the queries10,000 × about 50 tokens × $0.02 per millionabout $0.01 a month
Answers on Claude Haiku 4.510,000 × (6,000 × $1 + 400 × $5) per millionabout $80 a month
Answers on Claude Sonnet 5.510,000 × (6,000 × $2 + 400 × $10) per millionabout $160 a month
Vector storepgvector in existing Postgres, or Pinecone Standard minimum$0 or $50 a month

So tokens and storage come to about $80 to $210 a month. At this size the index itself is small: 25,000 passages of 1,536-dimension vectors take roughly 0.15 GB. Add the hosting for your application and sync jobs, logging and scheduled evaluation runs, which depend on your cloud and retention rules. If query volume grows tenfold, the token line grows tenfold; the build does not change.

How long does a RAG project take?

A typical RAG project takes two weeks of discovery and then 6–16 weeks to build, depending on scope. A single-source assistant with clean digital documents is at the short end. Several sources, per-user permissions and scanned documents push it to the long end, or into phases.

The timeline follows the same drivers as the price:

  • Week one and two: discovery on a sample of real documents, with retrieval tested against real questions and a fixed quote at the end.
  • Early build sprints: connectors, ingestion and cleanup for each source, then chunking, embeddings and search.
  • Middle sprints: permissions, citations, the answer interface and integration with where your team works.
  • Final sprints: the evaluation set run end to end, security review, load testing and launch.

What slows projects down is rarely engineering. It is waiting for access to source systems, agreeing who may see what, and finding subject-matter experts to write and check evaluation questions. Lining those up before the build starts is the single most effective way to keep the schedule.

RAG vs long context vs fine-tuning: the cost angle

RAG is not the only way to put your knowledge in front of a model, and the alternatives have very different cost curves.

Three ways to use your documents with a model
RAGLong contextFine-tuning
Cost per queryLow: only relevant passages are sentHigh: the whole set is sent every timeLow at inference, plus training runs
Updating knowledgeRe-index changed documentsEdit the documentsRetrain
Citations and permissionsYes, per passage and per userHard to enforce per userNo
Best forLarge or changing knowledgeSmall, stable setsConsistent behavior or format

Long context has become a real option: Anthropic includes a one-million-token context window at standard pricing on Claude 4.6 and later models, and its retrieval research puts the cut-off for skipping RAG at roughly 500 pages (about 200,000 tokens). Past that size, every query pays for every token. Filling that window on Claude Sonnet 5.5 costs about $2 per query at list price before caching, against roughly $0.016 for the RAG query in our example, a difference of more than 100 times. For a deeper comparison of the architectures, see RAG vs fine-tuning.

When an off-the-shelf tool is enough

You do not always need a custom RAG build. If your documents live in one workspace tool and its built-in AI search answers your team’s questions well enough, use it. The same applies if a document-chat product already connects to your sources, respects their permissions and keeps your data where your policies require.

Custom RAG is worth paying for when answers must come from several systems at once, when per-user permissions are strict, when documents are scans or complex tables, when the assistant must act through tools rather than only answer, or when it is part of a product you sell. If none of those apply, start with what you have. We set out a fuller decision method in build vs buy AI agents.

A real example: document intelligence in Relm

Relm, a commercial real estate platform we built, shows what the top end of the scope table involves. Its document intelligence runs OCR, extraction and vector search across rent rolls, offering memoranda, P&Ls and leases. Every finding is traceable to a source record, and tenant isolation covers the database, the vector store and the AI context, so one firm’s documents can never surface in another firm’s answers. The work that made it reliable was ingestion, citations and isolation, exactly the drivers above, not the choice of vector database.

How to keep RAG development cost down

The cheapest RAG project is a narrow one that grows on evidence.

  • Start with one source and one group of users, and add sources once answers are trusted.
  • Prototype on your messiest real documents in discovery, not a clean sample.
  • Write fifty real questions with known answers before the build starts; they become your evaluation set.
  • Use the database you already run for vectors unless scale says otherwise.
  • Route simple questions to a smaller model and cache stable instructions.
  • Keep regulated documents out of version one if the use case allows it.

To get a fixed price on your own documents, start with a free AI audit from our enterprise LLM and RAG team. Discovery then ends with retrieval tested on your data and a quote that does not move unless the scope does.

Related service

Private LLM architectures, retrieval (RAG) systems and fine-tuned models that answer from your data, cite their sources and never leak it.

Want this applied to your business?
Free 45-min AI audit with a senior architect.
Book the audit
FAQ

Common questions.

How much does it cost to build a RAG system?

With us, $25k–$120k as a fixed price, set after a $5,000 Discovery Sprint that tests retrieval on your own documents; the build then takes 6–16 weeks. One clean document library sits at the low end. Several sources with per-user permissions, scanned documents and agent tools sit at the top.

How much does a RAG system cost to run each month?

In our worked example of 10,000 queries a month, model tokens cost about $80 on Claude Haiku 4.5 or $160 on Claude Sonnet 5.5 at list prices, query embeddings cost about a cent, and the vector store costs nothing extra in Postgres with pgvector or $50 a month minimum on Pinecone Standard. Hosting and monitoring come on top.

What is the most expensive part of a RAG project?

Preparing the data. Scanned PDFs, tables, duplicates, outdated versions and documents with mixed permissions all have to be handled before retrieval works well. Evaluation is the next largest part, because you need a set of real questions with known answers to prove the system retrieves the right passages.

Is RAG cheaper than fine-tuning?

Usually, for knowledge that changes. RAG updates when you re-index a document, cites its sources and needs no training run. Fine-tuning suits consistent behavior or format rather than facts that change. Many production systems combine both, with RAG supplying the facts.

Can we skip RAG and put all our documents in a long context window?

For small, stable document sets, sometimes. But cost scales with every token sent on every query: filling a one-million-token context on Claude Sonnet 5.5 costs about $2 per query at list price before caching, against roughly $0.016 for a RAG query that sends 6,000 tokens. RAG also enforces per-user permissions, which a shared context cannot.

Which vector database should we use for RAG?

For most business systems, the one you already run. pgvector in an existing Postgres database adds no license cost and keeps data and permissions in one place. A managed vector service makes sense at large scale or when your team does not want to operate a database; Pinecone Standard, for example, has a $50 monthly minimum.