RAG development cost in 2026: what drives the price
Why data preparation, permissions and evaluation, not the vector database, decide what a retrieval-augmented generation system costs to build and run.
An internal assistant that answers 10,000 questions a month from 20,000 pages of company documents costs about $80 to $210 a month in tokens and storage. Indexing all 20,000 pages costs about 20 cents. The vector database and the model, the parts most RAG (retrieval-augmented generation) cost guides price, are the cheap parts.
The money goes into getting your documents into a state where retrieval works, enforcing who can see what, and proving the answers are right. So we price a RAG system the way we scope one: by what the data and the permissions demand.
How much does RAG development cost?
A production RAG system costs $25k–$120k to build with us, as one fixed price. We set it at the end of a two-week Discovery Sprint ($5,000) that tests retrieval on a sample of your real documents, so the quote reflects the data you have rather than a clean demo set.
Running it is the small number. The worked example below lands at about $80 to $210 a month at 10,000 queries, before hosting and monitoring.
What a RAG system is made of, and what each part costs
A RAG system has six parts. The first and the last take the most engineering; the middle ones are mostly configuration and usage fees.
| Component | What it does | Build effort | Running cost |
|---|---|---|---|
| Ingestion and cleanup | Pulls documents from each source, runs OCR, splits tables and text, removes duplicates and stale versions | High | Low, mostly compute during syncs |
| Chunking and embeddings | Splits text into passages and turns each into a vector | Medium | Very low: text-embedding-3-small lists at $0.02 per million tokens |
| Vector store and search | Stores vectors and finds the passages closest to a question, often with keyword search alongside | Low to medium | $0 extra in Postgres with pgvector; Pinecone Standard has a $50 monthly minimum |
| Permissions | Filters results to what each user is allowed to see | Medium to high | Negligible |
| Generation with citations | Sends the question and passages to a model and returns an answer with links to sources | Medium | The main usage cost: model tokens |
| Evaluation and monitoring | Tests answers against known questions and tracks quality after launch | High | Low, scheduled test runs |
RAG cost by scope
The scope of a RAG system is set by three questions: how many sources, how messy the documents are, and whether different users may see different things. Each answer pushes a project toward one row of the table below. The rows are our own placement of typical scopes inside the published $25k–$120k band, not quotes; the real number comes out of discovery once we have seen your documents.
| Scope | Typical shape | Where it usually lands |
|---|---|---|
| Single-source assistant | One document library, mostly digital text, everyone can see everything, answers with citations | Bottom of the band, from $25k |
| Multi-source with permissions | Several systems such as a drive, a wiki and a ticketing tool, scheduled syncs, per-user access rules | Mid-band |
| Document intelligence and agentic RAG | Scans and complex tables, structured extraction, agents that call tools, tenant isolation | Top of the band, up to $120k, or split into phases |
What drives the price of a RAG system?
Six things move a RAG quote. The vector database and the model are not on the list, because they barely change the build.
Data preparation is the largest cost
Retrieval can only find what ingestion made findable. Scanned PDFs need OCR and layout-aware extraction; tables need to be kept intact rather than split mid-row; duplicate and outdated versions need to be removed or they will be cited as if current. The messier the documents, the more of the build goes here, which is why discovery tests real files.
Number of sources and how fresh they must be
Each source needs a connector, authentication and a sync schedule. Nightly batch syncs are simple; near-real-time updates need change tracking and more careful re-indexing. Five sources cost meaningfully more than one, even if the total volume of text is the same.
Permissions
If everyone can see every document, permissions are free. If a salesperson must never see HR files, every passage needs access metadata copied from its source system, and every query must be filtered by it. Getting this wrong is a data leak, so it needs its own tests.
Retrieval quality
Basic vector search is quick to set up but misses exact terms such as product codes and names. Combining it with keyword search and adding a reranking step costs more to build and pays back in accuracy. In Anthropic’s published tests, contextual embeddings plus keyword search cut failed retrievals by 49%, and adding reranking took the cut to 67%. Preparing the contextualized chunks cost about $1.02 per million document tokens, once.
Evaluation
An evaluation set is a list of real questions with known correct answers and source passages. It is how you know a change made things better rather than worse. Building it takes time with your subject-matter experts, and it is the part most often skipped in pilots.
Hosting and compliance
Running in your own cloud, using zero-retention model endpoints or self-hosted models, keeping data in a region, and logging access for audits all add work. For regulated data, such as health or financial records, these are not optional.
Worked example: monthly RAG run cost at 10,000 queries
Here is the arithmetic for a mid-sized internal knowledge assistant, using list prices from each vendor’s pricing page.
Assumptions: 20,000 pages of documents at about 500 tokens a page, so 10 million tokens to index. 10,000 queries a month. Each query sends about 6,000 input tokens (the question, instructions and retrieved passages) and gets back about 400 output tokens.
| Line item | Calculation | Cost |
|---|---|---|
| Index the documents (one-time) | 10M tokens × $0.02 per million (text-embedding-3-small) | about $0.20 |
| Contextual chunks (optional, one-time) | 10M tokens × about $1.02 per million (Anthropic’s figure) | about $10 |
| Embed the queries | 10,000 × about 50 tokens × $0.02 per million | about $0.01 a month |
| Answers on Claude Haiku 4.5 | 10,000 × (6,000 × $1 + 400 × $5) per million | about $80 a month |
| Answers on Claude Sonnet 5.5 | 10,000 × (6,000 × $2 + 400 × $10) per million | about $160 a month |
| Vector store | pgvector in existing Postgres, or Pinecone Standard minimum | $0 or $50 a month |
So tokens and storage come to about $80 to $210 a month. At this size the index itself is small: 25,000 passages of 1,536-dimension vectors take roughly 0.15 GB. Add the hosting for your application and sync jobs, logging and scheduled evaluation runs, which depend on your cloud and retention rules. If query volume grows tenfold, the token line grows tenfold; the build does not change.
How long does a RAG project take?
A typical RAG project takes two weeks of discovery and then 6–16 weeks to build, depending on scope. A single-source assistant with clean digital documents is at the short end. Several sources, per-user permissions and scanned documents push it to the long end, or into phases.
The timeline follows the same drivers as the price:
- Week one and two: discovery on a sample of real documents, with retrieval tested against real questions and a fixed quote at the end.
- Early build sprints: connectors, ingestion and cleanup for each source, then chunking, embeddings and search.
- Middle sprints: permissions, citations, the answer interface and integration with where your team works.
- Final sprints: the evaluation set run end to end, security review, load testing and launch.
What slows projects down is rarely engineering. It is waiting for access to source systems, agreeing who may see what, and finding subject-matter experts to write and check evaluation questions. Lining those up before the build starts is the single most effective way to keep the schedule.
RAG vs long context vs fine-tuning: the cost angle
RAG is not the only way to put your knowledge in front of a model, and the alternatives have very different cost curves.
| RAG | Long context | Fine-tuning | |
|---|---|---|---|
| Cost per query | Low: only relevant passages are sent | High: the whole set is sent every time | Low at inference, plus training runs |
| Updating knowledge | Re-index changed documents | Edit the documents | Retrain |
| Citations and permissions | Yes, per passage and per user | Hard to enforce per user | No |
| Best for | Large or changing knowledge | Small, stable sets | Consistent behavior or format |
Long context has become a real option: Anthropic includes a one-million-token context window at standard pricing on Claude 4.6 and later models, and its retrieval research puts the cut-off for skipping RAG at roughly 500 pages (about 200,000 tokens). Past that size, every query pays for every token. Filling that window on Claude Sonnet 5.5 costs about $2 per query at list price before caching, against roughly $0.016 for the RAG query in our example, a difference of more than 100 times. For a deeper comparison of the architectures, see RAG vs fine-tuning.
When an off-the-shelf tool is enough
You do not always need a custom RAG build. If your documents live in one workspace tool and its built-in AI search answers your team’s questions well enough, use it. The same applies if a document-chat product already connects to your sources, respects their permissions and keeps your data where your policies require.
Custom RAG is worth paying for when answers must come from several systems at once, when per-user permissions are strict, when documents are scans or complex tables, when the assistant must act through tools rather than only answer, or when it is part of a product you sell. If none of those apply, start with what you have. We set out a fuller decision method in build vs buy AI agents.
A real example: document intelligence in Relm
Relm, a commercial real estate platform we built, shows what the top end of the scope table involves. Its document intelligence runs OCR, extraction and vector search across rent rolls, offering memoranda, P&Ls and leases. Every finding is traceable to a source record, and tenant isolation covers the database, the vector store and the AI context, so one firm’s documents can never surface in another firm’s answers. The work that made it reliable was ingestion, citations and isolation, exactly the drivers above, not the choice of vector database.
How to keep RAG development cost down
The cheapest RAG project is a narrow one that grows on evidence.
- Start with one source and one group of users, and add sources once answers are trusted.
- Prototype on your messiest real documents in discovery, not a clean sample.
- Write fifty real questions with known answers before the build starts; they become your evaluation set.
- Use the database you already run for vectors unless scale says otherwise.
- Route simple questions to a smaller model and cache stable instructions.
- Keep regulated documents out of version one if the use case allows it.
To get a fixed price on your own documents, start with a free AI audit from our enterprise LLM and RAG team. Discovery then ends with retrieval tested on your data and a quote that does not move unless the scope does.
Sources
- Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks, arXiv (Lewis et al.)
- Introducing Contextual Retrieval, Anthropic
- Claude API pricing, Anthropic
- OpenAI API pricing, OpenAI
- Pinecone pricing, Pinecone
Enterprise LLM & RAG
Private LLM architectures, retrieval (RAG) systems and fine-tuned models that answer from your data, cite their sources and never leak it.



