There's a delay. Here's why, and what it costs you.
Every way of being understood in a language you don't speak has a lag. The question isn't whether Taiwa has one. It's how it compares to the alternatives, and what we do to keep it small.
Three ways to be translated
Three ways your customer gets understood. All three have a delay. Only one of them is usually described as instant, and it isn't.
A native speaker
The lag: none. The wait: everything else.
No translation delay, because there's no translation. But you have to find them, transfer to them, and hope they're on shift — which means hiring multilingual agents for every language you take calls in. That doesn't hold up for long-hours coverage or the languages you get a handful of calls a week in.
A telephone interpreter
The lag: two to six seconds, interpreting live.
A third person on the line. In simultaneous interpreting, published studies put professionals roughly two to six seconds behind the speaker, and past about four seconds accuracy drops, because they're holding more of the sentence in memory. Many phone calls are consecutive instead — each side speaks, stops, and waits — which avoids that lag but stretches the call.
Taiwa
The lag: several seconds behind, per turn.
In one controlled test, Taiwa's translated audio started about three to seven seconds after the customer began speaking — most languages within the same broad ear-voice span reported for interpreters, Korean and Chinese a little above it. A machine isn't holding the sentence in working memory the way a person is, but that doesn't make the delay vanish. In that test the translation was usually already playing before the customer had finished their sentence. No third person on the line, no transfer, no wait for someone to be available.
Interpreter figures from published research on ear–voice span in simultaneous interpreting. ↓
Why there's a floor
You can't translate a sentence you haven't heard.
Translation isn't transcription. To turn a sentence into another language you need enough of it to know what it means, and "enough" is often most of it.
Take "I went to the bank". Until the next few words arrive, that's either a riverbank or a financial institution, and in most languages those are different words. Guess early and you're fast and wrong. Wait, and you're slower and right.
Every simultaneous interpreter makes this trade continuously: run too close and you commit to a meaning before the speaker has finished making it.
Part of that floor is linguistic, not technical. Faster hardware can shave the processing, but it can't let any system translate meaning that hasn't arrived yet. It's the reason no real-time translation is instant, ours included.
The German problem
Some languages make it worse, and it isn't anyone's fault.
German often leaves the decisive part of the verb until the end of the clause. Japanese is even more consistently predicate-final. Until that piece arrives, you may know who and what were involved, but not what actually happened.
Ich habe den Vertrag gestern zusammen mit meinem Kollegen …
I have the contract yesterday together with my colleague …
gekündigt.
cancelled.
Everything before that last word is compatible with signing it, reading it, or losing it. The decisive action arrives last.
And it isn't only German, or only verbs. Mandarin, Japanese and Korean can mark a yes/no question at the very end — a final 吗, か, or Korean verb ending. There are often clues earlier, but the safe English shape can hang on that last piece: statement, or question?
Human interpreters hit this exact wall — the research describes them being forced to increase their lag when working from German, waiting for a predicate that arrives at the end of the clause. It's not a limitation of machine translation. It's a property of the language.
So we tested the simple version of that fear. In one controlled run, with synthetic voices, a German sentence whose meaning flips on its last two words started translating no slower than a plain one. Short Mandarin, Japanese and Korean questions came back as fast as their matching statements — and correctly as questions, with the "Did…" that English puts first. That doesn't prove every long sentence behaves the same way. It does show the engine doesn't simply sit silent, waiting for the final verb or particle: it starts on the opening and fills the rest in behind.
You can watch it happen.
One controlled sentence, spoken by a synthetic native voice in each of thirteen languages and translated to English through DeepL Voice at true real-time. The grey bar is the source speech; the black bar is the first translated audio. In this single run, eleven of the thirteen black bars start before the grey one ends — the translation already playing while the sentence is still being spoken.
Measured through the DeepL Voice engine — one sentence, one synthetic native voice per language, a single run, source loudness-normalised and silence-trimmed. This is what happened in that run, not a benchmark. Korean and Chinese began about 0.6–0.7 seconds after the source speech ended. DeepL keeps improving its engine, so we re-run this and update the page rather than leave a flattering snapshot up.
What we do about it
We don't wait for the full sentence.
DeepL streams a phrase as it forms, revising the text as more arrives, rather than sitting silent until the speaker stops. Often the translation starts before the sentence ends — though, as the chart above shows, not always.
The transcript arrives before the audio does.
Text usually appears before the spoken translation, because the speech has to be synthesised. So your agent can follow the customer's meaning as they read it, while the sentence finishes settling before anyone acts on it.
The transcript corrects itself in place.
A phrase appears as a tentative line and updates as the sentence resolves — the riverbank becomes the financial institution, in the transcript, without anyone repeating themselves. The spoken audio is the opposite: once it's played, it can't be taken back.
Reading ahead of listening
The delay is real. So are the costs of the alternatives.
A phone interpreter adds setup time and another person to the call. A native speaker has no lag — if one's on shift. Taiwa is for the calls where that trade beats waiting, transferring, or staffing every language yourself.
