</>Technical documentation

Technical details

TranscriLive is a real-time transcription, translation, diarization and voice engine, powered by AI models running locally on GPU. Everything runs on your own infrastructure — no data is ever sent to the cloud.

Architecture

The real-time pipeline

Input
Transcription
Translation
Diarization
Voice
Outputs

An in-process orchestrator captures audio, transcribes, translates, identifies speakers, then vocalizes it. A per-language pub/sub bus feeds pluggable sinks: several simultaneous outputs per language.

Inputs & outputs

What goes in, what comes out

Inputs

• Browser microphone (WebRTC)

• Audio device / sound card (CoreAudio, ASIO)

• Dante (Dante Virtual Soundcard)

• RTSP network stream

• HLS stream

• Audio/video files (batch mode)

Outputs (several per language)

• WebSocket captions (real-time JSON)

• WebVTT / SRT / TTML / EBU-TT Live

• Vocalized audio per language → device / Dante channel

• RTSP / SRT network broadcast (ffmpeg)

• HLS stream (remote viewers, QR code)

• Multilingual web caption pages (overlay, lower-third)

AI processing

What the engine does

Transcription

Real-time speech recognition, even in noisy environments and with multiple speakers.

Translation

Simultaneous translation of the transcribed text, live, into several languages at once.

Diarization

Speaker identification: who speaks, and when — for clear, attributed transcripts.

Voice

Voice synthesis of the translated text, natural voices, in real time — with optional voice cloning.

All processing runs locally on GPU. The AI models are interchangeable depending on the accuracy and speed required.

Languages & voices

Multilingual, from text to voice

Languages

• Transcription in 90+ languages

• Fast mode on 25 European languages

• Real-time multilingual translation

• Simultaneous outputs in several languages

Synthetic voices

• Library of male and female voices

• Per-language vocalization, in real time

• Natural prosody and punctuation

Voice cloning

• From a 5 to 15-second sample

• Automatic transcription of the sample

• Cloned voice reusable for vocalization

Real-time & broadcast

Latency and mass broadcasting

Stage
Typical latency
WebSocket captions
20 – 100 ms
Transcription (STT)
200 – 500 ms
Translation
~ 100 – 300 ms
Voice
real time (streaming)

Mass broadcasting: share an accessible link via QR code and the whole audience follows live on their own device — up to hundreds of thousands of viewers. Caption overlays for the control room and RTSP/SRT streams re-injectable into an existing broadcast chain.

Integration

API & control

REST / WebSocket API

• Versioned /api/v1 API, self-documented (Swagger)

• API-key authentication

/ws/subtitles?lang= · /ws/listen · /subtitles/<lang>.vtt · /voices/upload · /routing

Operations

• Savable input/output profiles

• Strict device validation (CoreAudio UID)

• Restart identically, no intervention

• Web console for setup and testing

Security & sovereignty

Everything stays with you

100% on-premise

No data sent to the cloud.

GDPR native

Local processing, no third-party.

EAA / RGAA compliant

Accessibility of live events.

Sovereignty

You keep control of the whole chain.

System requirements

Recommended hardware per environment

TranscriLive relies on GPU acceleration. The power required depends on the enabled stack (transcription only, + translation, + voice) and the number of languages broadcast simultaneously.

Windows (NVIDIA CUDA)

• GPU : NVIDIA RTX 4090 24 Go recommended (full stack)

• Transcription + translation: RTX 4070 / 4080 (12–16 GB)

• Multi-language / broadcast: RTX 5090 or RTX A6000

• CPU 8+ cores · 32 Go RAM · Windows 10 / 11

Linux (NVIDIA CUDA)

• Ideal for server / control room, sovereign deployment

• GPU : RTX 4090 24 Go or RTX A6000 / L40S workstation

• Multiple GPUs for many parallel languages

• CPU 8+ cores · 32 Go RAM · Ubuntu 22.04+

macOS (Apple Silicon)

• Native Apple Silicon MLX pipeline (M1 → M4)

• Recommended: Mac Studio M2 / M3 Max or Ultra

• 30+ GPU cores · 32–64 GB unified memory

• MacBook Pro M-series for transcription + translation

Processing runs on the GPU (CUDA on PC, Neural Engine / Metal on Mac). No cloud, no specific network card required beyond the local network.

Comparison

What you gain over the others

Our most direct competitor, KUDO, translates conferences in the cloud and bills per usage (meeting, minute, language); TranscriLive is unlimited once installed, and nothing leaves your servers. As for voice — vocalization, cloning — it almost only exists in the cloud, on US servers.

TranscriLive
on-premise
KUDO
cloud · usage billing
Event interpretation
Wordly · Interprefy…
Cloud transcription
Otter · Deepgram · Google · Azure…
Cloud voice AI
OpenAI · ElevenLabs…
100% on-premise, on your servers~
No data in the cloud
Sovereignty / GDPR (EU hosting)~~
Unlimited — no per-minute/hour billing
Real-time transcription~
Real-time multilingual translation~~
Diarization (who speaks, when)~~~
Vocalization / live synthetic voice
Voice cloning~
Mass broadcast (QR, captions, RTSP/SRT)~~
Works offline / local network
Yes~ Partial / optional NoIndicative comparison by category; offerings evolve.

Want to try TranscriLive?

On your own infrastructure: we install a free pilot on your environment, results measured on your own live events. Or right now: an online demo, available on our SaaS test platform.