Voice agents · · 5 min

Why sub-second latency makes or breaks a voice agent

Callers judge an AI agent in the pause before it speaks. Here is where those milliseconds go.

In conversation, people leave about 200 ms between turns. When an agent takes two seconds to respond, callers start talking over it — or hang up.

That pause is the first thing a caller judges, before they notice the voice or what the agent says, and most of it can be engineered away.

How fast should a voice agent respond?

A voice agent should start speaking within about one second of the caller finishing, and faster is noticeably better. That target comes from how people talk to each other, not from what the technology finds easy.

Conversation research is consistent here. Levinson and Torreira summarize it as gaps between turns “of the order of 200 ms”, while the time a person needs to plan even a short utterance is over 600 ms. People manage this by predicting how the other speaker will finish and preparing a reply while they are still listening. A comparison of ten languages by Stivers and colleagues found the same pattern everywhere: speakers avoid overlap and minimize silence, and the average gap varies by language only within about 250 ms of the cross-language mean.

The phone network has its own guideline. ITU-T G.114 sets a target of 150 ms one-way delay for most voice uses, before any AI is involved. Together, these numbers explain why a two-second silence feels broken: it is roughly ten times longer than the gap a caller expects.

Where the time goes

  • Speech-to-text finishing the caller’s sentence.
  • The language model deciding what to say or which tool to call.
  • Text-to-speech producing the first audio.
  • Network hops between each service.

Each item hides a further delay. Before transcription can finish, the system has to decide that the caller has actually stopped talking. The model’s first token arrives quickly, but a long system prompt, a large model or a tool call can push the first useful sentence much later. Speech synthesis needs a full phrase to sound natural. And on a phone call, audio travels from the carrier to your telephony provider, then to each AI service and back, so a service hosted far from the others adds a round trip on every turn.

End-of-turn detection is the hidden delay

The most overlooked delay is deciding when the caller has finished. A simple voice activity detector waits for a fixed stretch of silence before it ends the turn. OpenAI’s documentation for its realtime API describes the trade-off directly: with a shorter silence duration, turns are detected more quickly. Set it too short, and the agent interrupts people who pause mid-sentence; set it too long, and every reply starts late.

Newer approaches look at meaning as well as silence. LiveKit’s agent framework, for example, uses a turn detector model that predicts the end of a turn from both what was said and how it sounded, on top of voice activity detection, and can adapt its delay to the caller’s own pauses. OpenAI offers a semantic mode with an adjustable eagerness setting. Neither removes the trade-off, but both let the agent answer quickly when a sentence is clearly complete and wait when it is not.

How to cut latency without making the agent worse

We stream every stage, start speaking before the full response is generated, and run tool calls in parallel with filler acknowledgements — so the agent feels present even while it works.

That comes down to six choices:

  • Stream transcription, generation and speech, so each stage starts on partial output instead of waiting for the previous one to finish.
  • Send the first sentence to speech synthesis as soon as it is complete, and keep generating the rest while it plays.
  • Keep the speech, model and telephony services in the same region, and keep connections warm between calls.
  • Keep the system prompt short and move rarely needed instructions into tools, so the model reads less on every turn.
  • Use a smaller, faster model for routing and simple turns, and a larger one only where reasoning is needed.
  • Cache data that does not change during a call, such as opening hours or the caller’s account, after the first lookup.

None of these require a worse answer. The goal is to remove waiting, not thinking.

Interruptions matter as much as speed

A fast agent that cannot be interrupted still feels robotic. When a caller starts talking, the agent should stop speaking almost at once, keep what it heard, and respond to the interruption rather than finishing its script. Frameworks such as LiveKit pause agent speech when they detect the caller and can tell a real interruption from a backchannel like “uh-huh”, so the agent does not stop every time someone agrees with it.

Test this on purpose. Interrupt the agent mid-sentence, talk over it in a noisy room, and pause halfway through a long account number. These are the moments callers remember.

Platform or custom pipeline: who controls the latency?

Hosted voice platforms such as Vapi and Retell assemble the pipeline for you: telephony, transcription, model and voice behind one API. That is the fastest way to a working agent, and for many call flows their latency is good enough. The trade-off is control. You choose from the vendors and regions the platform supports, and when a turn is slow, you can see less of why.

A custom pipeline, built on frameworks such as LiveKit or Pipecat, puts every stage under your control: where each service runs, which models handle which turns, how end-of-turn detection is tuned, and how tool calls are scheduled. It takes more engineering and more operational work. It is worth it when calls involve heavy lookups in your own systems, when data must stay in a specific region, or when volume makes per-minute platform fees a large line item. Whichever route you take, measure it the same way.

How to measure voice agent latency

Measure what the caller hears, not what a dashboard reports. The number that matters is the time from the end of the caller’s speech to the first audio of the reply, recorded on a real phone call.

  • Place test calls over the phone network, not only in a browser demo, because the carrier and telephony hops add delay a demo never sees.
  • Record both sides of the call and mark the end of caller speech and the start of agent audio.
  • Report the median and the 95th percentile across many turns; a good median with a slow tail still produces awkward calls.
  • Log each stage separately, so a slow turn points to transcription, the model, a tool or speech synthesis.
  • Repeat the test with background noise, accents and the longest tool calls in your flow.

Run this test on your own call flows before launch, because a latency figure measured on a different script, network or vendor mix says little about your calls.

“Latency isn’t a technical metric. It’s the difference between a conversation and a phone tree.”

What to do next

If you are planning a voice agent, set a latency target alongside your business goals, and ask any vendor or platform how they measured theirs. Our AI voice agent development work starts with your call flows, integrations and a latency budget for each one. To put a value on every answered call, read what missed calls really cost or put your own call volume into the AI voice agent ROI calculator.

Related service
Want this applied to your business?
Free 45-min AI audit with a senior architect.
Book the audit
FAQ

Common questions.

What is a good latency for an AI voice agent?

Under about one second from the end of the caller’s speech to the first word of the reply, measured on a real phone call. Faster is better: people leave gaps of roughly 200 ms between turns in normal conversation, so every extra few hundred milliseconds is noticeable.

Why do AI voice agents feel slow?

Usually because each stage waits for the one before it. The agent waits for silence to decide the caller has finished, then for a full transcript, a full model response and a full audio file. Streaming every stage and speaking the first sentence as soon as it is ready removes most of that waiting.

Does a phone call add latency compared with a browser demo?

Yes. A phone call adds the carrier network and the hop between the telephony provider and your speech services, which a browser demo does not have. Always measure on a real phone line before launch, not only in a browser.

How do you measure voice agent latency?

Record test calls and measure the time from the end of the caller’s speech to the first audio of the reply, as the caller hears it. Report the median and the 95th percentile across many turns, and log each stage separately so you know which one to fix.

How can a voice agent look things up without long silences?

Start the lookup in parallel with a short spoken acknowledgement, keep the tool APIs fast and close to the agent, and cache data that does not change during the call. If a lookup will take several seconds, the agent should say so, as a person would.