Simultaneous interpretation is now a given in major international institutions, conferences, and multilingual events. Yet behind this practice lies a rich history closely tied to the political and technological upheavals of the 20th century, from the booths of Nuremberg to contemporary neural algorithms.
What is easy to forget, listening to a smooth simultaneous rendering at a modern summit, is how recently this was considered nearly impossible. For centuries, high-stakes multilingual diplomacy relied on consecutive interpretation, a speaker pausing every few sentences so an interpreter could relay them, which could double or triple the length of any negotiation. The idea of translating a speech as it was being delivered, without that pause, was treated by many linguists of the era as a cognitive feat bordering on the implausible.
1. Origins: the shock of Nuremberg
Simultaneous interpretation officially emerged during the Nuremberg Trials in 1945. IBM developed a revolutionary system: soundproof booths where interpreters translated in real time, while listeners received the translation through multi-channel headsets. This innovation immediately became a reference model.
The stakes at Nuremberg made the experiment unavoidable rather than optional: the trials involved defendants, witnesses, judges, and prosecutors across English, French, Russian, and German, and consecutive interpretation would have stretched proceedings that already ran for months into something logistically and politically untenable. The interpreters who staffed those booths, several of them refugees who had fled the very regime on trial, effectively invented a new profession under enormous pressure, in real time, with no established training programs or techniques to draw on.
2. The golden age of human interpreters
From the 1950s onward, the United Nations and other international organisations widely adopted simultaneous interpretation. Highly trained and often multilingual, interpreters became indispensable facilitators of diplomatic dialogue. Their work requires exceptional skills: active listening, short-term memory, rapid reformulation, and mastery of cultural nuance.
Dedicated interpreting schools emerged to formalize what Nuremberg's pioneers had improvised, and by the 1970s simultaneous interpretation booths were a standard fixture at any major international gathering, from NATO summits to the Olympic Games. This period also cemented interpretation as a specialized, well-compensated profession precisely because the skill set was so demanding: a competent conference interpreter typically works in short shifts, often twenty to thirty minutes at a time, because sustained simultaneous translation is cognitively exhausting in a way few other jobs are.
3. The AI era: promises and limitations
Since around 2016, neural machine translation has shifted the landscape. The pipeline "speech → text → translation → speech" has become technically viable, with latency measured in just a few seconds. However, errors in rare languages, difficulty handling metaphors, and data privacy concerns remain real challenges.
This is the same technical foundation behind the live captioning and speech translation available on our product page today, extended to 60+ languages rather than the handful a single human interpreter can realistically cover. Unlike a human interpreter who needs advance notice and a booked slot, this kind of system can be spun up for an unplanned session with no lead time, which matters enormously for smaller organisations that could never justify booking a professional interpreting team for every meeting where a language barrier might come up.
4. Cost: a major barrier
A team of interpreters for a two-day conference with three languages can cost more than €15,000. In contrast, a scalable AI-based solution can reduce this cost by a factor of ten, opening the door to multilingual events for associations, SMEs, and local NGOs that previously could not afford them.
That cost gap compounds with scale. A single interpreting team is priced per language pair, per day, which means an event with five languages instead of two does not simply cost a bit more, it often requires an entirely separate team and booth setup for each additional pair. An AI-based system, by contrast, tends to scale far more gently: adding a language is a configuration choice rather than a new line item requiring its own logistics, travel, and per-diem costs.
None of this erases the value of a skilled human interpreter, particularly in settings where a mistranslation carries diplomatic or legal weight. What it does is widen who gets access to real-time understanding at all. A UN summit could always justify a full interpreting team; a regional business association's annual meeting, a community organisation's multilingual town hall, or a mid-sized company's international town hall meeting historically could not, and simply proceeded in whichever language the majority of attendees happened to share, leaving everyone else to follow along as best they could.
Conclusion
From Nuremberg to AI, simultaneous interpretation reflects humanity's constant need for dialogue. The future will be hybrid: AI for large-scale translation, humans for nuanced diplomacy and critical contexts. Far from disappearing, simultaneous interpretation is entering a new era.
Sources: Gaiba, Francesca. The Origins of Simultaneous Interpretation (1998). OpenAI, Introducing Whisper (2022).