When a live translation shows a short delay, the instinct is to think it is falling behind the conversation, almost broken. It is the opposite. That delay is not a flaw to fix. It is the direct result of a physical constraint that not even the best human interpreters have ever managed to avoid.
The delay is not a bug, it is a strategy
Conference interpreters call it the ear-voice span: the time between hearing a word and starting to say it in another language. For experienced professionals, this gap is usually two to four seconds. Not because they are slow to understand, but because they wait on purpose. They need enough material to get the meaning right.
Take a German sentence: the main verb often comes at the very end. Ich habe gestern meinen alten Freund aus der Schule zufällig im Supermarkt getroffen is, word for word, "I have yesterday my old friend from school by chance at the supermarket met." An interpreter who translates word by word, as each word arrives, would produce nonsense. They have to wait for the verb to know what actually happened. This is a constant trade-off between speed and accuracy, and no one, human or machine, fully escapes it.
A real-time translation system faces the same dilemma, just in a different form. It can start sooner, but the meaning may be less precise. Or it can wait for more context, at the cost of a bit more delay. This trade-off cannot be removed. It can only be tuned.
What understanding really means
There is a difference between recognizing words and understanding a sentence. Machines have been good at recognizing for a long time. Understanding means resolving ambiguities that a human sorts out without even noticing: the tone used, the implied meaning, a reference to something said three sentences earlier.
Here is a simple example. In English, "I saw her duck" can mean two things: I saw her pet duck, or I saw her bend down. No grammar rule can settle this, only the situation can. It is this kind of small, everyday case that makes automatic translation much harder than voice recognition alone.
Japanese pushes this even further. The language very often leaves out the subject of a sentence, because it is assumed to be clear from context. Translating into English, which almost always needs an explicit subject, means rebuilding information that is not in the original text. At that point, it is no longer translation in the strict sense. It is interpretation, in the true meaning of the word.

Not all languages face this problem equally
Some language pairs work together more easily than others. English and Spanish share a similar word order (subject-verb-object), a comparable grammar structure, and a syntax that allows fairly direct, step-by-step translation. English and Japanese, on the other hand, pile up the obstacles: reversed word order, no equivalent grammatical markers, missing subjects, and levels of politeness with no direct match.
It is no coincidence that translation systems, human or automated, perform very differently depending on the language pair. The difficulty is not about how rich a vocabulary is. It is about how far apart two languages are in the way they organize thought into words. Across the 60+ languages Glot supports, that distance varies enormously, and the engineering has to account for it.
The real goal is not to remove the delay
It is tempting to think that progress means shrinking this delay until it disappears. That would be solving the wrong problem. A system with zero delay would have to translate word by word, without waiting for context, and it would produce, just like our rushed German interpreter, a result that looks right but means something wrong.
So the real question is not how do we go faster. It is how do we decide, at each moment, whether we know enough to speak. This is exactly the problem that the best real-time translation systems are working on today, Glot included: not erasing the delay, but making it just long enough to stay accurate, and just short enough to stay natural.
Next time a live translation pauses for a second or two before answering, that is not dead time. That is the time it takes to get it right.
Want to see how this engineering plays out on a real stage?
Discover the technology