Why sub-second latency makes or breaks a voice agent
Callers judge an AI agent in the pause before it speaks. Here is where those milliseconds go.
In conversation, people leave about 200 ms between turns. When an agent takes two seconds to respond, callers start talking over it — or hang up.
That pause is the first thing a caller judges, before they notice the voice or what the agent says, and most of it can be engineered away.
How fast should a voice agent respond?
A voice agent should start speaking within about one second of the caller finishing, and faster is noticeably better. That target comes from how people talk to each other, not from what the technology finds easy.
Conversation research is consistent here. Levinson and Torreira summarize it as gaps between turns “of the order of 200 ms”, while the time a person needs to plan even a short utterance is over 600 ms. People manage this by predicting how the other speaker will finish and preparing a reply while they are still listening. A comparison of ten languages by Stivers and colleagues found the same pattern everywhere: speakers avoid overlap and minimize silence, and the average gap varies by language only within about 250 ms of the cross-language mean.
The phone network has its own guideline. ITU-T G.114 sets a target of 150 ms one-way delay for most voice uses, before any AI is involved. Together, these numbers explain why a two-second silence feels broken: it is roughly ten times longer than the gap a caller expects.
Where the time goes
- Speech-to-text finishing the caller’s sentence.
- The language model deciding what to say or which tool to call.
- Text-to-speech producing the first audio.
- Network hops between each service.
Each item hides a further delay. Before transcription can finish, the system has to decide that the caller has actually stopped talking. The model’s first token arrives quickly, but a long system prompt, a large model or a tool call can push the first useful sentence much later. Speech synthesis needs a full phrase to sound natural. And on a phone call, audio travels from the carrier to your telephony provider, then to each AI service and back, so a service hosted far from the others adds a round trip on every turn.
End-of-turn detection is the hidden delay
The most overlooked delay is deciding when the caller has finished. A simple voice activity detector waits for a fixed stretch of silence before it ends the turn. OpenAI’s documentation for its realtime API describes the trade-off directly: with a shorter silence duration, turns are detected more quickly. Set it too short, and the agent interrupts people who pause mid-sentence; set it too long, and every reply starts late.
Newer approaches look at meaning as well as silence. LiveKit’s agent framework, for example, uses a turn detector model that predicts the end of a turn from both what was said and how it sounded, on top of voice activity detection, and can adapt its delay to the caller’s own pauses. OpenAI offers a semantic mode with an adjustable eagerness setting. Neither removes the trade-off, but both let the agent answer quickly when a sentence is clearly complete and wait when it is not.
How to cut latency without making the agent worse
We stream every stage, start speaking before the full response is generated, and run tool calls in parallel with filler acknowledgements — so the agent feels present even while it works.
That comes down to six choices:
- Stream transcription, generation and speech, so each stage starts on partial output instead of waiting for the previous one to finish.
- Send the first sentence to speech synthesis as soon as it is complete, and keep generating the rest while it plays.
- Keep the speech, model and telephony services in the same region, and keep connections warm between calls.
- Keep the system prompt short and move rarely needed instructions into tools, so the model reads less on every turn.
- Use a smaller, faster model for routing and simple turns, and a larger one only where reasoning is needed.
- Cache data that does not change during a call, such as opening hours or the caller’s account, after the first lookup.
None of these require a worse answer. The goal is to remove waiting, not thinking.
Interruptions matter as much as speed
A fast agent that cannot be interrupted still feels robotic. When a caller starts talking, the agent should stop speaking almost at once, keep what it heard, and respond to the interruption rather than finishing its script. Frameworks such as LiveKit pause agent speech when they detect the caller and can tell a real interruption from a backchannel like “uh-huh”, so the agent does not stop every time someone agrees with it.
Test this on purpose. Interrupt the agent mid-sentence, talk over it in a noisy room, and pause halfway through a long account number. These are the moments callers remember.
Platform or custom pipeline: who controls the latency?
Hosted voice platforms such as Vapi and Retell assemble the pipeline for you: telephony, transcription, model and voice behind one API. That is the fastest way to a working agent, and for many call flows their latency is good enough. The trade-off is control. You choose from the vendors and regions the platform supports, and when a turn is slow, you can see less of why.
A custom pipeline, built on frameworks such as LiveKit or Pipecat, puts every stage under your control: where each service runs, which models handle which turns, how end-of-turn detection is tuned, and how tool calls are scheduled. It takes more engineering and more operational work. It is worth it when calls involve heavy lookups in your own systems, when data must stay in a specific region, or when volume makes per-minute platform fees a large line item. Whichever route you take, measure it the same way.
How to measure voice agent latency
Measure what the caller hears, not what a dashboard reports. The number that matters is the time from the end of the caller’s speech to the first audio of the reply, recorded on a real phone call.
- Place test calls over the phone network, not only in a browser demo, because the carrier and telephony hops add delay a demo never sees.
- Record both sides of the call and mark the end of caller speech and the start of agent audio.
- Report the median and the 95th percentile across many turns; a good median with a slow tail still produces awkward calls.
- Log each stage separately, so a slow turn points to transcription, the model, a tool or speech synthesis.
- Repeat the test with background noise, accents and the longest tool calls in your flow.
Run this test on your own call flows before launch, because a latency figure measured on a different script, network or vendor mix says little about your calls.
“Latency isn’t a technical metric. It’s the difference between a conversation and a phone tree.”
What to do next
If you are planning a voice agent, set a latency target alongside your business goals, and ask any vendor or platform how they measured theirs. Our AI voice agent development work starts with your call flows, integrations and a latency budget for each one. To put a value on every answered call, read what missed calls really cost or put your own call volume into the AI voice agent ROI calculator.
Sources
- Timing in turn-taking and its implications for processing models of language, Frontiers in Psychology (Levinson & Torreira, 2015)
- Universals and cultural variation in turn-taking in conversation, PNAS (Stivers et al., 2009)
- What is latency? (ITU-T G.114 one-way delay target), Twilio
- Turn detection and interruptions, LiveKit Agents docs
- Voice activity detection (VAD), OpenAI API docs



