Voice agents · · 9 min

AI voice agent cost in 2026: per-minute vs custom build

Per-minute rates from vendor price lists, the line items behind them, and the call volume at which a custom build starts paying for itself.

Vapi advertises $0.05 a minute. Add the speech-to-text, the language model and the voice that every call needs, and the same agent costs $0.082 to $0.129 a minute on Vapi’s own calculator. The advertised rate covers one layer of a five-layer stack; the other four arrive as separate line items.

Every figure below is a list price checked on October 7, 2026, and every calculation is shown, so you can rerun it with your own call volume.

How much does an AI voice agent cost per minute?

At list prices, a hosted voice platform costs about $0.08 to $0.15 per minute once every layer is included. A custom stack that pays each vendor directly costs about $0.04 per minute in usage, but you have to build and host it. The platform’s advertised rate, $0.05 on Vapi or $0.055 on Retell, is only the orchestration fee.

All-in cost per minute, list prices (Oct 2026)
OptionPer minuteWhat is included
Vapi$0.082–$0.129$0.05 hosting plus speech-to-text, model and voice passed through at cost, on Vapi telephony (no transport fee)
Retell, Claude 4.5 Haiku$0.110$0.055 infrastructure, $0.015 voice, $0.025 model, $0.015 telephony
Retell, recommended tier$0.149Same, with a $0.064 recommended-tier model such as Claude 5 Sonnet
Custom stackabout $0.044Twilio, Deepgram Nova-3, Deepgram Aura-2 voice, Claude Sonnet 5.5; hosting and build not included

The five line items inside every per-minute price

Every AI phone call pays for the same five things, whether a platform bundles them or you buy them separately. Knowing them is how you compare quotes.

  • Telephony carries the call. Twilio lists US inbound calls to a local number at $0.0085 a minute, and streaming the audio to your server at $0.0044 a minute.
  • Speech-to-text transcribes the caller. Deepgram lists streaming Nova-3 at a regular price of $0.0077 a minute, currently discounted to $0.0048; we use the regular price.
  • The language model decides what to say and which tools to call. It is billed per token, so its per-minute cost depends on prompt size and how often the caller speaks.
  • Text-to-speech speaks the reply. Deepgram lists Aura-2 at $0.030 per 1,000 characters; Retell lists voices from $0.015 to $0.10 a minute depending on the provider.
  • Orchestration handles turn-taking, interruptions and tool calls. Platforms charge for it per minute; on a custom stack you host it yourself.

How we turned token prices into a per-minute cost

Two of these layers are not priced per minute, so we converted them with stated assumptions. For the model, we assume five model calls per minute of conversation, each sending 3,000 input tokens (2,500 of them a cached system prompt) and receiving 80 output tokens. At Claude Sonnet 5.5 list prices of $2 per million input tokens, $0.20 cached and $10 output, that is $0.0023 per call and $0.0115 a minute. For the voice, we assume the agent talks for about 45% of the call at around 150 words a minute, roughly 400 characters, which is $0.012 a minute on Aura-2.

Platform pricing: what Vapi and Retell actually charge

Both platforms publish their prices, so a fair comparison is possible. They differ in what is bundled and what costs extra.

Platform price lists compared (Oct 2026)
VapiRetell
Platform fee$0.05/min hosting$0.055/min voice infrastructure
Models and voicesPassed through at cost; calculator examples $0.0095–$0.0099 transcriber, $0.0077–$0.0452 model, $0.0146–$0.0238 voicePriced per minute: voices $0.015–$0.10; models from under $0.01 to over $0.30, with Claude 4.5 Haiku at $0.025 and the recommended tier at $0.064
TelephonyVapi telephony and SIP free, one number on the free tier; Twilio numbers billed by Twilio$0.015/min, numbers $2 a month
Concurrent calls4 on the free tier, extra lines $10 each a monthFirst 20 free, then $8 each a month
HIPAA$2,000 a month add-onHIPAA and BAA on the enterprise plan, custom pricing

The choice of model and voice moves the platform price more than the choice of platform. On Retell, swapping a recommended-tier model for Claude 4.5 Haiku takes the all-in rate from $0.149 to $0.110 a minute, the most expensive models on Retell’s list cost over $0.30 a minute on their own, and moving to a premium ElevenLabs voice can add up to $0.085.

What a custom AI voice agent costs to build

With us, a custom voice agent falls in the same $25k–$120k band as any agent we build, with the price fixed after discovery (the pricing page has the model). “Custom” means you own the orchestration code, run it in your cloud and pay each vendor directly, typically on an open-source framework such as LiveKit or Pipecat instead of a hosted platform.

Five things set where a voice build lands in that band:

  • the systems the agent books into or updates, such as a scheduler, CRM or order system,
  • the number of call flows, from a single booking flow to triage across several departments,
  • warm transfer to a person with a summary of the call,
  • latency tuning, because slow replies make a conversation feel broken (see voice agent latency),
  • compliance, such as HIPAA, call recording consent and data retention.

A platform build is not free either. Someone still writes the prompts, connects the agent to your systems and tests the call flows. That work is similar on both routes, so the comparison below charges only the custom route with a build, which favors the platforms.

Total cost at 1,000, 10,000 and 50,000 minutes a month

At 1,000 minutes a month a platform costs roughly a tenth of a custom stack; at 50,000 the order flips. The table adds usage at list prices and, for the custom stack, two fixed monthly costs we state as assumptions: $500 a month for hosting and monitoring, and the minimum build of $25k spread over 36 months, about $694 a month.

Monthly cost by call volume (list prices, assumptions above)
Minutes a monthVapiRetellCustom usageCustom all-in
1,000$82–$129$110–$149$44$1,239
10,000$818–$1,289$1,100–$1,490$441$1,635
50,000$4,090–$6,445$5,500–$7,450$2,205$3,399

The platform columns exclude concurrency charges and HIPAA add-ons, and the custom column excludes the cost of your team’s time to operate it. Both omissions matter at the edges, so treat the table as a starting model, not a quote.

The break-even point for a custom voice agent

On these list prices, a custom stack breaks even at about 11,400 minutes a month against Retell’s recommended tier, 14,100 against Vapi’s higher configuration, 18,100 against Retell with Claude 4.5 Haiku, and 31,700 against Vapi’s lowest configuration. The arithmetic is the custom stack’s fixed monthly cost, about $1,194, divided by the per-minute saving.

Break-even minutes per month for a custom stack
Compared withPlatform per minuteSaving per minuteBreak-even
Retell, recommended tier$0.149$0.105about 11,400
Vapi, higher configuration$0.129$0.085about 14,100
Retell, Claude 4.5 Haiku$0.110$0.066about 18,100
Vapi, lowest configuration$0.082$0.038about 31,700

Two things move the break-even sharply, in opposite directions. A HIPAA add-on of $2,000 a month on a platform is larger than the custom stack’s entire fixed cost in this model, so it pulls the break-even down, although a custom stack has its own compliance costs, such as Twilio’s Security Edition for a BAA. A larger build, say for several call flows and integrations, pushes it up by roughly 6,600 to 18,400 minutes for every additional $25k spread over three years, depending on the platform you compare against.

Costs that per-minute quotes leave out

A per-minute rate is the biggest line on the bill, but not the only one. Budget for these as well, on any route:

  • Concurrent calls. Platforms include a set number of simultaneous calls and charge for more: Retell includes 20 and charges $8 a month for each extra, and Vapi includes 4 on its free tier and charges $10 a month per additional line. Size this for your busiest hour, not your average.
  • Phone numbers. Small, but they add up across locations: Twilio lists local numbers at $1.15 a month and toll-free at $2.15.
  • Transfers. When the agent hands a caller to a person, the outbound leg is billed too; Twilio lists outbound calls at $0.0140 a minute.
  • Recordings and transcripts. Storage and retention rules matter more than the storage price, especially for regulated calls.
  • Testing. Every prompt or model change should be replayed against recorded test calls before it goes live, and those test calls use real minutes.
  • Change. Models, voices and prices change; someone has to evaluate and apply those changes.

HIPAA and the BAA chain

A voice agent that hears patient information is only as compliant as the weakest vendor on the call. Under HIPAA, every vendor that receives protected health information must sign a business associate agreement, and on a voice call that is the carrier, speech-to-text, the language model, text-to-speech, the orchestration layer and wherever recordings and transcripts are stored.

Price this chain before you choose a stack. Vapi lists HIPAA-eligible data handling at $2,000 a month. Retell offers HIPAA and a BAA on its enterprise plan at custom pricing. Twilio signs a BAA only for customers on its Security or Enterprise Edition. On a custom stack you sign each agreement yourself and can choose vendors and regions accordingly, which is often the deciding factor for healthcare work. There is no such thing as being “HIPAA certified”, so ask each vendor for its BAA, not a badge.

When a platform is the right answer

For most first voice agents, a platform is the better choice. Choose one when:

  • volume is under about 10,000 minutes a month and not growing fast;
  • the call flow is standard, such as answering questions, booking into a common scheduler or taking messages;
  • you want to test whether callers accept an AI agent before you invest in a build;
  • or nobody will run voice infrastructure after launch.

Go custom when volume is high or rising, when compliance needs control over every vendor in the chain, when the agent must act inside your own systems in ways a platform’s integrations cannot, or when latency and voice quality are part of your product. Many teams start on a platform and move once the numbers above favor it; designing prompts and tools to be portable makes that move cheaper.

Is an AI voice agent worth it? Compare with missed calls

The right comparison is not the agent’s cost against a receptionist’s salary; it is the agent’s cost against the revenue lost to calls nobody answers. A business taking 60 calls a day at three minutes each uses about 5,400 minutes a month, which costs roughly $440 to $800 a month in platform usage at the list prices above.

Then estimate what missed calls cost you: calls a month, times the share you miss, times the share of those callers who would have bought, times the value of a customer. We walk through that model in what missed calls really cost, and you can run it with your own numbers in the AI voice agent ROI calculator. If the lost revenue is a multiple of the agent’s monthly cost, the case is easy at any of the prices in this guide.

How to lower the per-minute cost of a voice agent

The biggest savings come from the model and the voice, not from haggling over the platform fee.

  • Use the smallest model that passes your call tests, and keep the system prompt stable so it can be cached.
  • Pick a standard voice unless the voice is part of your brand; premium voices can cost several times more per minute.
  • Keep calls short with clear flows and fast handoff to a person when the agent cannot help.
  • Route by intent, so simple calls never reach the expensive model.
  • Recheck list prices before you commit. This guide uses prices checked on October 7, 2026, and vendors update them.

If you want these numbers worked out for your own call volume and systems, our AI voice agents team can model both routes in a free audit before you commit to either.

Related service
Want this applied to your business?
Free 45-min AI audit with a senior architect.
Book the audit
FAQ

Common questions.

How much does an AI voice agent cost per minute?

At October 2026 list prices, a hosted platform costs about $0.08 to $0.13 per minute on Vapi and $0.11 to $0.15 on Retell once you add speech-to-text, the language model, the voice and telephony. A custom stack built on Twilio, Deepgram and Claude costs about $0.04 per minute in usage, before hosting and the build.

How much does it cost to build a custom AI voice agent?

A custom voice agent from Apptycoons costs $25k–$120k as a fixed-price build over 6–16 weeks. The quote comes at the end of a $5,000 two-week Discovery Sprint and depends on the systems the agent books into or updates, how many call flows it handles, and compliance needs such as HIPAA.

Is a platform like Vapi or Retell cheaper than a custom voice agent?

Below roughly 11,000 minutes a month, yes, on list prices. A platform has no build to pay off and its per-minute fee covers infrastructure you would otherwise run yourself. Above roughly 32,000 minutes a month, a custom stack is cheaper against every platform configuration in our model. In between it depends on which platform and model you compare against.

What does an AI receptionist cost compared with a human?

At 5,400 minutes a month, about 60 calls a day at three minutes each, platform usage costs roughly $440 to $800 a month at list prices. Compare that with the revenue lost to missed calls rather than with a salary alone, because the agent’s job is to answer every call, including after hours.

Do HIPAA-compliant voice agents cost more?

Usually, yes. Every vendor that handles patient data on the call has to sign a business associate agreement. Vapi lists HIPAA-eligible data handling as a $2,000 a month add-on, Retell offers HIPAA and BAA on its enterprise plan, and Twilio requires its Security or Enterprise Edition before it signs a BAA.

What are the line items in a voice agent’s per-minute cost?

Five: telephony to carry the call, speech-to-text to transcribe the caller, the language model that decides what to say, text-to-speech to speak the reply, and the platform or orchestration fee that ties them together. Platforms bundle these into one rate; a custom stack pays each vendor directly.