Sangam documentation
The manual for the live demo, the walkthrough, the architecture, and the protocol two browsers speak to the host. Everything described here is what the running application does on 10 Sep 2026.
1 · Demo access
The PIN, and what it opens.
- Demo PIN
- 284617
- Where
- /app on this same address — one port serves both the betadoc and the demo
- Session
- twelve hours, in an
HttpOnlycookie. Restarting the server invalidates every session immediately - Rate limit
- eight wrong attempts from one address locks that address out for fifteen minutes
Entering the PIN does one thing: it asks the running application for a fresh room and sends you into it. There is no account, no profile, and nothing of yours is stored — the room is thrown away when both of you leave it.
localhost unless the page is served over HTTPS. Driving the demo from the machine that is hosting it works, because loopback counts as a secure context. From a second machine, use the HTTPS address (https://sangam-app.hayman.ink) instead of this port, or the second person will get no microphone.There is nothing synthetic behind the PIN. It is the same process on the same machine that the operator uses, with the same models loaded. What makes it safe to hand out is that the application holds no data: a room exists only while somebody is in it.
2 · What this is
Two people, two computers, one shared link.
One person opens the site, gets a link, and sends it to the other. Each of them picks the language they speak. From that point on, when one talks the other hears it in their own language — Hindi, Telugu, Bengali, Tamil or Kannada — and both of them can read what was said and what it became.
There is no button. You speak; when you stop, that pause is the end of your turn. The pause length is 0.8 seconds and it was chosen by sweeping five candidate values on forty minutes of real conversational Telugu, not by picking a number that sounded right.
Video is on the same link, and the transcript sits beside the faces so you can read what was just said while looking at the person who said it. The video itself goes browser to browser and never passes through the host.
3 · The manual
Everything on the screen, in the order you meet it.
The room and the link
Opening the demo creates a room and drops you straight into it. Along the top of the page is the room bar: what is happening right now on the left, and the share link on the right with two controls beside it.
- Copy puts the link on your clipboard. What it copies is the address the other machine can reach — never
127.0.0.1, which would be useless to them. - QR opens a code for pointing a phone at, so somebody across a table can join without typing anything.
The room id is cryptographically random and long enough that it cannot be walked. Rooms are collected once they are empty, so an old link stops working and shows a page saying so, with a way to start a new one. A third person is refused with a reason the page renders — never a silent close, which is indistinguishable from a crashed server.
If your laptop sleeps, your wifi changes, or the connection simply drops, reconnecting within the grace period reclaims your seat and your language rather than being treated as a third person. The other side sees reconnecting, not left.
Names and languages
You answer two questions for yourself: what to call you, and which language you speak. The other person answers the same two on their own device. Nobody sets anybody else's language — changing yours does not touch theirs, and it takes effect on your very next turn.
Below that are three dropdowns, one for each stage of the pipeline — recognition, translation and speech — each set to run on this device or in the cloud. That choice belongs to the room, not to a person, because both of you share one pipeline on the host; changing it is broadcast to both screens so they cannot disagree. The defaults are what measurement chose rather than what sounds best: every stage local. The cloud voice sounds better and is one dropdown away; it is not the default because it is the slowest stage in the turn and the only part that would leave the machine.
If you are alone in the room, it says so. Talking into a room with nobody in it and blaming the software is the failure this exists to prevent.
Talking, hands free
Press Start conversation once, allow the microphone once, and then never touch anything again. Your browser streams continuously and the host decides where your turns begin and end.
- A pause of 0.8 s ends your turn. A 300 ms breath in the middle of a sentence does not.
- Noise does not become a turn. A cough, a door, a keyboard: a turn needs both a minimum duration and a minimum transcript before it is sent onward. A real one-word reply still gets through, which is why the guard could not simply be "ignore short things".
- You can interrupt. Start talking while their translation is playing and the playback stops. Your turn is captured from its true beginning — a 0.30 s pre-roll is kept and re-attached, because the detector only reports a start once it is sure — and the other person is told you interrupted.
- You never hear yourself. The translated audio goes to the listener only. You get your own transcript and the translation for reference, with no audio.
Faces and the transcript
Turn camera on starts your video. It is entirely optional and never load-bearing: refusing the camera, losing it, or switching it off mid-call leaves the microphone, the endpointing and the whole turn pipeline untouched. Someone with the camera off appears as a named tile carrying their initial, and the other side is told the camera is off rather than left staring at a black rectangle wondering whether the call broke.
The tile of whoever is speaking is ringed as they speak, driven by the host's own speech events rather than a second detector guessing in the browser. While a translation is playing to you, your tile says so.
Hear their voice too is off by default, and deliberately. If the peer's raw voice plays alongside the translation you hear the same sentence twice in two languages and follow neither — and your loudspeaker gains a second source for your microphone to pick up, which is exactly the loop the echo work exists to prevent. The control is there for anyone who wants the original underneath.
At laptop width the faces and the words are both on screen at once. At phone width the layout stacks, with the faces above the words rather than shrinking both into uselessness.
What every state means
| On screen | What is happening |
|---|---|
| listening | Your microphone is open and the host is watching for speech. Nothing is being sent onward yet. |
| hearing | Speech has been detected and your turn is open. It closes 0.8 s after you stop. |
| working | Your turn closed and is being recognised, translated and spoken. The timings appear on the transcript row when it lands. |
| speaking | The other person's translated audio is playing to you. Start talking to interrupt it. |
| Nobody else is here yet | You are alone in the room. Send them the link. |
| reconnecting | The socket dropped and the page is trying to reclaim the same seat. After four failed retries it says the server is not answering rather than spinning forever. |
| a stage failed | Recognition, translation or speech raised an error. The other stages still deliver and the row states which part is missing and why — you are never left with silence and no reason. |
The rule behind that table: no state lasts more than a second without something visible changing. At any instant the screen should answer is it hearing me, has my turn ended, is it working, is the other person speaking, has something failed.
4 · A conversation, start to finish
Ravi speaks Telugu at a ticket counter. Priya answers in Hindi.
This is the exact conversation in the video on the betadoc, recorded through this port on 10 Sep 2026 with real audio.
Ravi opens the demo
He enters the PIN. A room is created and he lands in it. The bar says Nobody else is here yet — send them the link.
He says who he is
Types Ravi, taps తెలుగు. The three stage dropdowns already read: recognition, translation and speech all on this device.
He sends the link
One press on copy, or the QR if Priya is holding a phone. What lands in her message is the reachable HTTPS address, not a loopback one.
Priya opens it on her own computer
She types Priya, taps हिन्दी. Ravi's bar changes to Talking with Priya, in हिन्दी. Neither of them chose the other's language.
Both press Start once, and allow the microphone
That is the last button either of them touches.
Ravi asks how much a ticket to Hyderabad costs
He stops talking. 0.8 s later his turn closes. His tile shows working; Priya's transcript gets his Telugu, the Hindi translation, and about four seconds after he stopped, the Hindi audio plays on her speakers. Ravi hears nothing — nobody is played their own words back.
The first turn of a session is slower than the rest, because the models load on the first call. In the recorded run it was 7.1 s; the five turns after it ranged 4.5 to 6.5 s.
Priya answers — three hundred rupees, how many do you need?
Same in reverse. Her Hindi, his Telugu, audio on his side only. Each row carries its own timings: heard, translated, spoken, total.
They finish the exchange
Two tickets, four in the afternoon from platform two, and a canteen beside the platform. Six turns, no buttons, and both of them can scroll back and read either language.
They close their tabs
The room empties and is swept. The link stops working, and there is no transcript left anywhere.
5 · When something goes wrong
Every one of these says what happened, and what to do about it.
| What you see | What it means and what to do |
|---|---|
| This address cannot use the microphone — open it over https | You are on an origin that is not a secure context. Outside one, navigator.mediaDevices is not merely blocked, it is undefined, so this is checked rather than discovered: start is disabled and the HTTPS link is offered instead. Use that. |
| Microphone blocked | Permission was refused. Allow it in the browser's site settings and press Start again. |
| Microphone held by another app | Another call or recorder owns the device. Close it and press Start again. |
| This browser cannot run it | Checked before you type anything — websockets, AudioContext, AudioWorklet and mediaDevices. Start stays disabled with the reason, instead of failing after you have filled everything in. |
| This room already has two people in it | Both seats are taken. Ask for a new link. |
| Nobody else is here yet — send them the link | You are alone. Nothing is broken. |
| The server is not answering — it may be restarting | Shown after four failed reconnects, rather than an eternal reconnecting. |
| Video could not connect directly | The two browsers could not find a direct path between them. Audio translation is unaffected — the camera is never load-bearing, and the conversation continues without it. |
| Person A chose Telugu, but that sounded like Kannada | A warning, never a re-route. You chose your language; the system says what it heard and offers to switch, and does nothing unless you say so. |
6 · Technical architecture
One FastAPI process, two browsers, and one websocket each.
Each browser captures at 16 kHz through an AudioWorklet with the browser's own echo cancellation on, and streams 50 ms packets over a single websocket. That same socket carries everything else: the roster, the transcripts, the translated audio, and the WebRTC signalling.
On the host
- Room registry — two seats, cryptographically random ids, reconnect tokens, and a time-based sweep that discards empty rooms rather than waiting for a request to trigger it.
- Endpointer, one per seat — Silero VAD, 0.8 s of silence closes a turn, 0.30 s of pre-roll is kept and re-attached so the beginning of a turn is not lost to the moment the detector became sure.
- Turn queue, one per room — closed turns are processed in the order they closed. A slow turn is never overtaken, two turns never interleave their audio, and a turn whose synthesis fails does not block the one behind it.
- Recognition — IndicConformer-600M quantised to INT8, on ONNX Runtime CPU threads so it does not compete with anything on the GPU. Needs an explicit language code, which is why each person states their language rather than the machine guessing.
- Translation — IndicTrans2
indic-indic-dist-320M, on CPU deliberately: the GPU measured slower for this model size. Direct Indic↔Indic, never pivoted through English. - Speech — the local MMS voice by default, 0.62 s at RTF 0.08, covering all five languages. A cloud voice sounds better and costs 2.06 s, which was the single largest cost in a turn until this became the default.
What crosses which wire
- Audio in — browser to host, over the websocket.
- Translated audio out — host to the listener's socket only.
- Both transcripts — to both sockets.
- WebRTC signalling — over the same websocket. The host is a letterbox: it relays offer, answer and ICE to the other seat verbatim and never inspects, rewrites or stores them. A signalling message from a socket with no seat is ignored; a message with nobody to send it to is refused with a reason rather than dropped.
- Video — browser to browser, and nowhere near the host. Verified with
getStats(): the nominated candidate pair was host → host with 3.2 MB received on the far side.
Nothing is persisted
There is no database. A room holds its transcript in memory while it is alive and is discarded when it empties. There are no accounts, no roles and no recordings — anything that needs an audit trail of what was said needs a different product.
7 · What was measured
Every model in the pipeline had to earn its place.
These are runs on this machine, recorded in the project's build tracker. Character error rate rather than word error rate throughout, and chrF rather than BLEU, because word-level scoring punishes a correct stem with a different inflection in agglutinative languages.
| Stage | What runs | Score | What the comparison showed |
|---|---|---|---|
| Recognition | IndicConformer-600M INT8 · local ONNX CPU | CER 0.7–13.5% | By language, on FLEURS dev sets. The cloud recogniser scores 2.2% on the matched subset — more accurate, six times slower, and the audio leaves the machine. |
| Translation | IndicTrans2 indic-indic-dist-320M · local CPU | chrF 57.0 / 55.4 | hi→te and te→hi, against NLLB-600M's 53.9 / 52.5 on the identical 25 parallel sentences. Better in both directions, three times faster, half the size. NLLB was ruled out on the numbers before its licence was considered. |
| Speech, local | MMS-TTS · one VITS checkpoint per language | RTF 0.08 · CER 0.072 | Median round-trip character error over twelve real sentences per language, and it covers all five. Twelve times faster than realtime on CPU. Round-trip through our own recogniser, so it measures intelligibility, not whether the voice sounds natural. |
| Turn boundary | Silero VAD · 0.8 s end-of-turn pause | swept, not chosen | 0.4 / 0.6 / 0.8 / 1.0 / 1.5 s over 40.4 minutes of real conversational Telugu, 85 participant channels. 0.8 s is the knee: 1155 turns against 1256 expected with 61 clauses cut, where 0.4 s cut 155. |
| Echo | Browser WebRTC echo cancellation | ≥ 33.8 dB ERLE | Measured on this machine's own speakers and microphone playing a real synthesised translation, not a tone. Cancellation off: correlation 0.556 at a 52 ms lag. On: 0.041 with no consistent lag, so 33.8 dB is a floor rather than a point estimate. |
| Whole turn | everything local — the default | 2.20 s median | Twelve turns, alternating directions, from the speaker stopping to the listener's audio starting. Wait 1.25, recognition 0.20, translation 0.50, synthesis 0.17 — the synthesis figure fell by two thirds when Piper replaced MMS as the local voice. The endpoint wait is now the largest part: 1.25 s against a configured 0.80 s threshold, so about 0.45 s is detection and delivery — cheaper to attack than the threshold. |
| Whole turn | local recognition + local translation + cloud voice | 4.05 s median | The previous default, and the reason the default moved. Identical but for the voice: synthesis alone is 2.06 s of it, against 0.62 s locally. The cloud voice is the better-sounding one, so this row is a quality choice rather than a slower mistake. |
| Whole turn, video on | the cloud-voice pipeline, plus two 640×480 peer streams | 4.17 s median | +0.12 s, 3% of a turn and smaller than the run's own 3.37–5.38 s spread. Measured under encode and decode at about 1.1 Mbps, 10 fps. |
8 · Protocol reference
What the browser and the host actually say to each other.
Everything below is the application's own surface, reachable through this betadoc's proxy at the same paths once the PIN cookie is set.
HTTP routes
| Route | What it does |
|---|---|
| GET / | Creates a room and returns 303 to /r/<id>. There is no lobby. (On this betadoc's port, / is the betadoc site and /app is what asks for a room.) |
| GET /r/{id} | Serves the application for that room. An unknown id returns 404 with a readable page and a link to start a new room. |
| GET /qr/{id} | The join link as an SVG QR code, rendered server-side because the page cannot know the address another machine can reach. 404 for an unknown room. |
| GET /where | {"https": "<public base>"} or null — the address to hand to the other device. |
| GET /health | {"ok": true, "rate": 16000}. |
| GET /static/* | The page, the capture worklet and the loader engine. |
| GET /aec | The echo-return-loss measurement page. Needs a real microphone and real speakers, so it is not something a headless browser can run. |
| GET /vidload | A canvas-based video load generator, used to measure what video costs a turn without needing a camera. |
| WS /ws/room/{id} | A seat in a room. This is the whole application. |
| WS /ws/session | The single-device socket from the first arc, kept for the bench pages. |
Websocket, client to host
The first message on /ws/room/{id} is a JSON object carrying the reconnect token, or {} for a fresh seat. Everything after that is either a binary frame of 16-bit PCM audio, or a JSON object with a cmd field.
| cmd | Fields | Effect |
|---|---|---|
| hello | rate | Declares the sample rate the browser actually got. Sending the true rate matters: every duration measured downstream depends on it. |
| name | name | Sets your display name and broadcasts a new roster. |
| lang | lang | Sets your language. Never the other person's. Takes effect on your next turn. |
| providers | asr, mt, tts | Each local or openrouter. Applies to the room, and is broadcast so both screens agree. |
| listen | on | Opens or closes continuous capture for this seat. |
| camera | on | Records whether your tile is dark by choice or by fault; the roster carries it either way. |
| signal | payload | WebRTC offer, answer or ICE candidate. Relayed verbatim to the other seat. |
| interrupt | — | You started talking over the audio being played to you. Your page stops the playback; this tells the other person why their sentence went quiet. |
| turn_start / turn_stop | — | Explicit turn boundaries, from the first arc's push-to-talk mode. Not used in a hands-free room. |
| echo, record, stop, stats, example | various | Diagnostic commands used by the measurement pages. |
Websocket, host to client
| event | Meaning |
|---|---|
| seated | You have a seat. Carries the room id, your seat letter, your reconnect token, whether the seat was reclaimed, the roster, the room's provider choices and whether the room is ready. |
| refused | You did not get a seat, with a machine-readable reason — room_not_found or room_full. Always sent before the close, never a bare close frame. |
| room | The roster changed: somebody joined, changed language, or left. Carries every seat's state, name, language and camera flag. |
| providers | The room's three-stage choice changed. |
| listening / hearing | Capture is open; speech has been detected and a turn is open. |
| level | The other person's audio level, for the waveform under their tile. |
| turn_open / turn_done / turn_empty | A turn opened, completed, or produced nothing worth sending — the noise guard. |
| utterance / said | What was recognised, in the speaker's own language. |
| translation | The same turn in the listener's language. |
| speaking / spoke | Audio is about to play, and has finished. Also what marks the tile. |
| interrupted / interrupt_ok | The other person started talking over you; your own interruption was accepted. |
| stage_failed | Recognition, translation or speech failed, with which one and why. The rest of the turn still arrives. |
| lang_warning | What you said did not sound like the language you chose. A warning with an offer to switch, never an automatic re-route. |
| signal / signal_refused | Relayed WebRTC payload from the other seat, or a refusal with a reason: not_in_a_room, nobody_to_signal. |
| cannot_talk | You tried to start a turn before choosing a language. |
Binary frames from the host to a client are the translated audio, 16-bit PCM, and they are sent to the listener's socket only.
9 · Running it
Two launchd agents, two ports, and three environment knobs.
The application
launchctl print gui/$(id -u)/com.onec.sangam # the app, 127.0.0.1:8000 launchctl kickstart -k gui/$(id -u)/com.onec.sangam # restart it
A user launchd agent with RunAtLoad and KeepAlive, so it comes back after a crash and after a boot, and the models reload with no manual step.
This betadoc
launchctl print gui/$(id -u)/com.onec.betadoc.sangam # site + gate, :6789 launchctl kickstart -k gui/$(id -u)/com.onec.betadoc.sangam # after editing the site python3 serve.py --check # its one self-check
Port 6789 sits next to the project's private build cockpit on 6788 and is registered in the shared port registry with "kind": "betadoc", so the allocator skips it. The knobs live in betadoc.env: host, port, PIN, and the upstream to proxy.
Choosing where the work happens
| Variable | Values | Default and why |
|---|---|---|
| SANGAM_ASR | local | openrouter | local — 0.20 s a turn against roughly 1.6 s in the cloud, and the audio stays here. |
| SANGAM_MT | local | openrouter | local — better chrF than the 200-language generalist it was measured against, and three times faster. |
| SANGAM_TTS | local | openrouter | openrouter — the historical default. The local voice measures RTF 0.08 across all five languages and is the obvious next change. |
Resolution order is environment, then config.yaml, then the built-in default. The same three choices are exposed per room in the interface, so nothing has to be restarted to change them.
Reaching it from another machine
The application is served over HTTPS with a real certificate at sangam-app.hayman.ink, through an outbound tunnel — so nothing has to be reachable from outside and no port is opened. It answers only while the machine hosting it is running, which is inherent rather than incidental: the models, the microphone and the audio socket are all on that machine.
These documents are not. They are static snapshots and are served from a content network at sangam.hayman.ink, so a link to a page stays good whether or not the application behind it is up.