Web · API · MCPYour knowledge. Your model. Your cloud.
Private LLM architectures, retrieval (RAG) systems and fine-tuned models that answer from your data, cite their sources and never leak it.
Everything to go from idea to production.
RAG pipelines
Ingestion, chunking, embeddings and hybrid search over your documents.
Fine-tuning
Domain-tuned models for tone, format and classification tasks.
Access control
Answers respect per-user and per-team permissions.
Document intelligence
OCR and extraction for PDFs, scans, contracts and spreadsheets.
Private deployment
Run inside your AWS, Azure or GCP account, or on-prem.
Evaluation & monitoring
Accuracy, hallucination and latency tracked continuously.
How we deliver
- STEP 01
Audit
Data sources, access rules, use cases.
- STEP 02
Index
Ingest and embed your knowledge.
- STEP 03
Tune
Retrieval, prompts and evals.
- STEP 04
Deploy
Private, monitored, scaled.
A private ChatGPT for your company.
One chat assistant for your whole team. It answers from your own documents with citations and keeps every prompt and file inside infrastructure you control.
Managed API
Frontier models from Anthropic or OpenAI under business API terms that exclude training on your data, with zero retention where offered. Fastest to launch.
Your cloud’s model service
The same model families served inside your AWS Bedrock, Azure OpenAI or Google Cloud account, in the region you choose, under your existing cloud contract.
Self-hosted open model
Llama or another open-weight model on your own GPUs, in your cloud or on-prem. No external calls, at the cost of more operations work.
Data residency
Pick the region once. Documents, embeddings, chat history and logs stay in it, and sign-in goes through your identity provider, so people only get answers from files they can already open.
What drives the cost
Users and questions per day, the model tier, how many sources you connect, and whether you self-host. Self-hosting swaps per-question fees for a fixed GPU bill.
Our default: start on your cloud’s model service. Move to self-hosting only when policy forbids any external model call.
Stack by category
Retrieval
- Pinecone
- pgvector
- Weaviate
Models
- Llama
- Anthropic
- OpenAI
Cloud
- Azure
- AWS
Industries
RAG, fine-tuning or an off-the-shelf chatbot?
Most teams need retrieval first. Fine-tuning and off-the-shelf tools each have a narrower job.
| Your situation | Our recommendation |
|---|---|
| Staff want writing help and Q&A over a few uploaded files | Buy an enterprise chatbot plan |
| Answers must cite your systems and respect permissions | Build a private RAG system |
| Facts change weekly: policies, prices, inventory | RAG, not fine-tuning |
| You need a fixed format, tone or classification | Add fine-tuning on top of RAG |
Our recommendation: buy for general help; build RAG when answers carry your data and your liability.
What a private RAG system costs and how long it takes.
- 01 · 2 weeks
Discovery Sprint
$5,000fixed fee · 2 weeksData sources and access rules audited, plus a prototype answering questions over a sample of your documents.
- 02 · 6–16 weeks
Fixed-Price Build
$25k–$120kper project · 6–16 weeksProduction RAG with connectors, hybrid search, permissions, evaluations and private deployment.
- 03 · Ongoing
Dedicated Team
From $12kper month · cancel anytimeNew sources, use cases and model upgrades, with accuracy tracked every sprint.
What moves the price: the number of sources, how messy the documents are, how strict the access rules are, and whether models run in your own account. Compare engagement models
Shipped with this.
Web · API · MCPRead before you build.
- Buying software · 8 min readRAG development cost in 2026: what drives the price
- Engineering · 6 min readRAG vs. fine-tuning: choosing an enterprise LLM architecture
- AI strategy · 8 min readBuild vs buy AI agents: a CTO decision framework
- AI strategy · 5 min readWhy most AI pilots never reach production — and how to be the exception
Common questions.
RAG or fine-tuning?
RAG for facts that change and need citations; fine-tuning for consistent behaviour. Most systems use both.
Will our data train public models?
No. We use zero-retention endpoints or self-hosted models, and your data stays in your environment.
Can it handle scanned documents?
Yes. OCR and layout-aware extraction turn scans and PDFs into searchable, structured data.
How much does a RAG or enterprise LLM system cost?
Most engagements start with a 2-week Discovery Sprint ($5,000 fixed fee, credited to your build), followed by a fixed-price build ($25k–$120k per project, 6–16 weeks) or a dedicated team (from $12k per month). The build price is fixed after discovery and billed in milestones tied to demoed features. Scanned PDFs, tables and spreadsheets take more ingestion work than clean text. Embedding, vector storage and model usage are separate running costs we estimate in discovery.
How long does a RAG system take to build?
Discovery takes two weeks and ends with a prototype answering questions over a sample of your documents. A production system with connectors, access control, evaluations and monitoring is a fixed-price build of 6–16 weeks.
Where does our data live?
In the cloud account and region you choose (AWS, Azure or GCP) or on-prem. Documents, embeddings, chat history and logs stay there. Model calls go to API endpoints that do not train on your data (zero retention where the provider offers it) or to a model hosted in your own account.
Which vector database do you use?
Postgres with pgvector when you already run Postgres and want one less system; Pinecone or Weaviate when you need managed scale or advanced filtering. Retrieval combines keyword and vector search either way.
How do you measure answer quality?
We build an evaluation set of real questions with known answers and sources, then track retrieval accuracy, answer correctness, citation coverage and hallucination rate on every change, before and after launch.
Should we build a private LLM or buy an enterprise chatbot?
Buy an enterprise chatbot plan when general writing help over a few uploaded files is enough. Build when answers must respect per-user permissions, cite your systems of record, run in your own cloud, or feed other software through an API.


