Speech-to-speech translation for two
Two people, two computers, one shared link.
Each picks the language they speak. When one talks, the other hears it in their own.
No buttons. You speak, you pause, and the pause ends your turn. Video calling with the transcript beside the faces, and both languages written down, so a word lost to a bad line can be read instead of asked for again.
2.20 s
median from a speaker stopping to the listener's audio starting, everything local — wait 1.25, recognition 0.20, translation 0.50, synthesis 0.171
0.8 s
the pause that ends a turn, chosen by sweeping five values on 40.4 minutes of real conversational Telugu, not picked
+0.12 s
what video costs a turn against the same run with video off — 3%, smaller than the turn-to-turn spread
5 languages
Hindi, Telugu, Bengali, Tamil and Kannada, translated directly between any pair — never pivoted through English
The problem
A translator you have to operate is a translator you stop using.
Two people without a shared language can already get a translation out of a phone. What they cannot easily get is a conversation. Anything that has to be operated — pressed, aimed, held up, passed across a table — stops being a conversation and becomes two people taking turns at a device. Three things stand between a translator and a conversation, and all three are structural rather than cosmetic.
It has to know who is speaking
Two voices reaching one microphone leave the system guessing, and a wrong guess sends your sentence into the wrong language. Separating two speakers from a shared microphone measures 20.6% diarization error here — roughly one turn in five put on the wrong person — and that is the ceiling of the speaker model rather than something tuning reaches past.
A control is one more thing to hold
Press to start, release to stop, twice an exchange, while also trying to follow what someone is telling you. Workable for a minute and exhausting for ten — and the moment either person forgets, half a sentence is gone.
Where the audio goes is yours to decide
A translator that must reach a server is somebody else's copy of your conversation. That call belongs to the people talking, which means the machine in front of you has to be able to do the whole job on its own.
The product
Send a link. Pick your language. Talk.
Opening the site creates a room and hands you a link. Sending that link is the entire invitation — no account, no lobby, exactly two seats. Each person chooses their own language on their own screen, and from then on the pause in their speech is the only control either of them touches.
The pause is the turn
Voice activity detection per seat; a 0.8 s silence closes the turn, and a 300 ms breath inside a sentence does not split it.
Two devices delete the hard part
The speaker is whoever's socket the audio arrived on. Diarization is removed rather than improved.
A cough is not a turn
Minimum speech duration and minimum transcript length together — noise is rejected while a real one-word haan still gets through.
Interruption has a rule
Playback stops, the interrupting turn is captured from its true start with a 0.30 s pre-roll, and the other person is told.
Both languages, written down
Each side sees what was said and what it became. Nobody is played their own words back.
Faces beside the words
Video on the same link, transcript alongside, so you read what was said while looking at who said it.
Local by default
Recognition and translation run on the host machine. Which stages go to the cloud is a per-room choice, shown to both people.
Five languages, direct
Hindi, Telugu, Bengali, Tamil, Kannada — translated Indic↔Indic, never pivoted through English.
Two things it deliberately does not do.
- Pretend a room is private. The link is the key: anyone holding it can take the second seat. That is the right trade for a two-person tool, and it is said out loud rather than dressed up as security.
- Hold more than two people. Two seats by construction. A third socket is refused with a reason the page renders, never a bare close.
Architecture
One host, two browsers, and the video that never touches the host.
The one architectural decision worth arguing about is that the video does not pass through the host. The host is already running recognition, translation and synthesis for both people; two encoded video streams through it would spend exactly the budget the translation needs. It was verified rather than assumed: getStats() reported the nominated candidate pair as host → host with 3.2 MB received on the far side, so no media byte entered the server process.
Deployment and migration notes
Where it runs today, and what a second lane would cost.
How it runs
Recognition and translation never leave the host. Only the voice does, and only if you leave it there.
- One process under a launchd agent,
RunAtLoadandKeepAlive, bound to127.0.0.1:8000. - Models live on disk: IndicConformer for recognition, IndicTrans2 for translation, Silero for voice activity. Loaded once, about three gigabytes.
- Reachable from a second machine over HTTPS with a real certificate, because a browser will not grant a microphone on an insecure origin. The name resolves publicly, but the address behind it is only routable inside the tailnet — deliberate, for a two-person tool.
- Each stage is a knob, not a fork:
SANGAM_ASR,SANGAM_MT,SANGAM_TTS, eachlocaloropenrouter, resolved from the environment thenconfig.yaml. Choosing them per room in the interface is the same switch. - Nothing is persisted. A room holds its transcript while it is alive and is swept when it empties.
# the whole lane switch, no code change — every stage is local | openrouter SANGAM_ASR=local SANGAM_MT=local SANGAM_TTS=local \ .venv/bin/python3 -m uvicorn sangam.server:app --host 127.0.0.1 --port 8000
This betadoc. One port, 6789, registered beside the project's private build cockpit so nothing else can claim it. The site at / is public; /app is PIN-gated and, once unlocked, proxies the real application — including tunnelling the room websocket frame for frame without reading any of them.
Exposure. No tunnel, no static-host front, no subdomain: one port on this machine, deliberately.
Secrets. The cloud key lives in a gitignored .env and is never logged; every provider reads it through one helper, so none of them reaches for it its own way.
Privacy posture, stated. With the default settings nothing leaves the machine: recognition, translation and the voice all run on the host. A cloud voice is available for whoever prefers how it sounds, and is the only setting that sends any audio outward.
1 Every figure on this page comes from a run recorded in the project's build tracker on this hardware, most recently 10 Sep 2026, when the 107-test suite was re-run and the six-turn conversation in the video was recorded through this port.
Try it now.
The real application on this machine. Opening it creates an empty two-seat room and hands you the link — the same link you would send to the other person. Everything works from one browser except the second seat, which is the whole point: open the link on a second device to hear yourself translated.