All posts
researchcomparisonai

How fast is real-time call translation, actually?

Vendors quote "sub-3-second" and "1.5 second" latency for four different things. Here are our own production measurements — first audio at 4 to 4.7 seconds, synthesis first-byte ranging from 975ms to 6.2 seconds, and one turn that took 14.6.

Prakash Vakhesa · September 27, 2026 · 6 min read

Disclosure: we make TellAcross, a real-time call translation tool. The numbers below are ours, including the unflattering ones. We are publishing them because we could not find anyone else publishing theirs, and because "real-time" has quietly come to mean four different things.

Every vendor in this category, us included, says "real-time". The published figures range from 300 milliseconds to 4 seconds, and they are not measuring the same event.

That is not usually dishonesty. It is that a translated call has at least four moments you could call "the latency", and each vendor reports the one their architecture is best at.

The four things called latency

What's measured When the clock stops Typical published figure
Final transcript latency The translated text stops changing 1.5s median
StreamLAAL (academic) Average lag of the text stream, quality-adjusted 2.94s at BLEU 31.96, IWSLT 2025
Time to first audio byte (TTFB) Synthesis emits its first byte Rarely published
Time to first audible word The listener actually hears something Rarely published
End-to-end turn The listener has heard the whole sentence Sub-3s claimed at scale

The gap between row one and row four is the entire problem. A 1.5-second text latency is a real achievement and a real number. It also tells you nothing about when a human hears a voice, because text is finished before synthesis has started.

If you are buying for a conversation rather than for subtitles, only the last two rows describe your experience.

Our numbers

Measured in production, across real calls on real networks, using ElevenLabs Scribe v2 for recognition, GPT-4o or DeepL for translation, and ElevenLabs for speech.

Time to first audible word: roughly 4 to 4.7 seconds from the moment the speaker stops talking.

That breaks down as:

  • ~2s — recognition settles and the translation returns
  • 1–1.7s — synthesis time-to-first-byte, on a good run
  • 1–2s — a lead buffer we add on purpose (see below)

Synthesis TTFB is the volatile part. Across observed calls it ranged from 975ms to 6,216ms for the same model on the same account. Not the same sentence twice — the same service, varying by a factor of six depending on the minute.

Worst single turn observed: 14.6 seconds. One turn, one user, on a day when synthesis stalled mid-sentence. We are including it because an average hides exactly the experience that makes someone close the tab.

Why we deliberately add delay

This is the part that surprises people, so it is worth stating plainly: we hold audio back before playing it.

Our speech model for Indic languages stalls mid-sentence. If we start playing the instant the first byte arrives, the listener gets three words, a gap, two more words, another gap. Choppy audio is read as broken. Late audio is read as slow.

So we buffer 1.5 to 3 seconds of speech before starting playback. That buys smoothness and costs latency, and it is the right trade when synthesis is stalling. It is the wrong trade when synthesis is prompt — which is most of the time.

We handle that by watching TTFB on each turn: past a 2-second threshold we cut the buffer to 600ms and get audio out rather than waiting for a smooth run that is not coming.

There is no neutral setting here. Any vendor quoting a single latency number is either not buffering, or not telling you they are.

Some languages are structurally slower

Our default speech model is fast. It also cannot speak Hindi or Gujarati at all.

For those we use a higher-quality, slower model, which means there is no faster fallback available for Indic languages. When it is slow, we wait. A caption on a demo video reading "under 2 seconds" was almost certainly recorded in Spanish, French or German.

This generalises beyond us. Ask any vendor for their latency figure in the specific pair you will use, not their best pair. The spread between a well-resourced pair and a poorly-resourced one is larger than the spread between vendors.

What benchmarks don't measure

Published benchmarks run on clean audio. The 2026 vendor survey notes it directly: real conference audio with echo and overlap is poorly represented in the standard test sets.

Two things happen in real rooms that no BLEU score captures.

Your own output comes back as input. Translated speech leaves a laptop speaker, enters the microphone a metre away, and gets transcribed as a person talking. We logged this verbatim:

15:43:26  [TTS] Start → "Can you tell me your name? My name is Prakash."
15:43:31  [STT] COMMITTED     "Can you tell me your name? My name is Prakash."

Five seconds later, attributed to a human, and translated a second time. On a clean benchmark this never occurs.

Recognition sessions expire. Our streaming recognition session dies after about 45 seconds of silence. In a two-person conversation, 45 seconds of not speaking is completely ordinary — the other person is talking. Reconnecting cost the next speaker about 1.2 seconds, while they were already mid-sentence. We now recycle the session pre-emptively at 36 seconds. That entire class of delay is invisible to a benchmark that feeds continuous audio.

What to ask a vendor

Five questions that produce comparable answers:

  1. Which moment does your number measure — text settled, first audio byte, or first audible word?
  2. What is the p95, not the median? The median describes a good call. The p95 describes the call someone complains about.
  3. What is the figure in my language pair, not your best one?
  4. Do you buffer before playback? If yes, how much, and is it adaptive?
  5. What happens when your speech provider is slow — do you wait, degrade, or fall back to another model?

A vendor who answers all five with specifics is worth more attention than one with a lower headline number.

The honest summary

For a well-resourced pair on a good connection, expect three to five seconds from the end of one person's sentence to the start of the translated audio. Expect worse on Indic and other lower-resource languages. Expect occasional turns far outside that range, because one of the components is an external service having a bad minute.

That is slower than any marketing page in this category, ours included, will tell you. It is also fast enough to hold a real conversation, which is the only test that matters — and it is considerably faster than waiting two days for an interpreter to be scheduled.

We will republish these numbers as they change. If they get worse, we will say so.

Found this useful? Share it.

Keep reading