# CoreDeskAI Voice Bridge (Phase 1) — Python

> **⚠️ SUPERSEDED (2026-06-08).** This custom bridge (Telnyx Media Streaming ↔ Voxtral STT ↔
> Mistral ↔ Supertonic TTS) is slated for retirement in favor of **LiveKit Agents + a realtime
> speech-to-speech model** — see `../../LIVEKIT_VOICE_MIGRATION.md`. LiveKit provides VAD,
> turn detection, barge-in, and AEC as framework features, replacing the hand-coded equivalents
> here. Do not invest further in this bridge without checking that doc first.


Always-on service bridging **Telnyx Media Streaming** ↔ **Voxtral Realtime STT** for the
Vapi → Telnyx voice migration. Convex can't hold a raw audio WebSocket, so this does the
plumbing; Convex stays the orchestrator (wired in Phase 2). Python, because the Mistral
realtime SDK **and** Supertonic (Phase-2 TTS) both live there.

## Milestone 1 (this scaffold) — inbound STT proof
Telnyx forks the caller's audio (PCMU 8kHz) → bridge decodes μ-law + upsamples to PCM16/16kHz
(`audioop`) → streams to Voxtral Realtime (`client.audio.realtime.transcribe_stream`) →
**live transcript logged.** Inbound-only on purpose — prove real-time STT before the full loop.

## Files
- `bridge.py` — the Telnyx WS server → Voxtral bridge.
- `test_stt.py` — validate the Voxtral STT leg standalone (feed a WAV; no telephony needed).
- `requirements.txt` — `mistralai[realtime]`, `websockets`, `audioop-lts` (Py 3.13 dropped `audioop`).

## Run
```bash
cd services/voice-bridge
python -m venv .venv && ./.venv/Scripts/python -m pip install -r requirements.txt   # (.venv\Scripts on Windows)
MISTRAL_API_KEY=<key> ./.venv/Scripts/python bridge.py     # listens on :8080
ngrok http 8080                                            # -> wss://<id>.ngrok.app  (DEV ONLY)
```

> **Production deploy** — always-on EU host with co-located Supertonic (the real sub-1.5s
> latency fix): see **`DEPLOY.md`** (+ `Dockerfile`, `fly.toml`). ngrok/cloudflared are for
> local dev only — a tunnel reintroduces the RTT this migration is trying to remove.

## Validate the STT leg first (no phone call needed)
```bash
# make a 16k mono wav: ffmpeg -i any.mp3 -ac 1 -ar 16000 -sample_fmt s16 speech.wav
MISTRAL_API_KEY=<key> ./.venv/Scripts/python test_stt.py speech.wav
```
If you see the transcript print, Voxtral realtime works — the hardest leg is proven.

## Wire a Telnyx call to the bridge
The Telnyx webhook (`/webhooks/telnyx` in `packages/backend/convex/http.ts`) already issues
`streaming_start` on `call.answered` **when `VOICE_BRIDGE_WS_URL` is set** (otherwise it stays
in Phase-0 dial+speak mode). So:
```bash
# point the webhook at the running bridge, then redeploy convex
npx convex env set VOICE_BRIDGE_WS_URL wss://<bridge-host>/    # (run from packages/backend)
npx convex dev --once --typecheck=disable
```
Then place a call (the Phase-0 `initiateTelnyxCall`), answer, speak — watch `bridge.py` log the transcript.

## Env
| Var | Purpose | Default |
|---|---|---|
| `MISTRAL_API_KEY` | Voxtral STT auth (+ Mistral chat when not using Convex) | — (required) |
| `SUPERTONIC_URL` | Supertonic TTS server | `http://127.0.0.1:7788` |
| `VOXTRAL_MODEL` | STT model id | `voxtral-mini-transcribe-realtime-2602` |
| `VOXTRAL_DELAY_MS` | STT streaming delay (latency/accuracy) | `300` |
| `ENDPOINT_SILENCE_MS` | end-of-turn silence gap | `600` |
| `CONVEX_VOICE_TURN_URL` | if set, FULL app orchestration (`/voice/turn`: workflows/tools/KB/compliance + persistence). Highest precedence. | — (unset) |
| `CONVEX_LLM_URL` | LLM-reply layer only (`/webhooks/vapi/llm`: org LLM + naturalness, streamed); used if voice-turn unset | — (unset) |
| `VOICE_ORG_ID` | org this call belongs to (sent as `x-organization-id`) | — |
| `VAPI_WEBHOOK_SECRET` | `x-vapi-secret` if the Convex endpoint enforces it | — |
| `VOICE_GREETING` | spoken on stream start (bridge owns out-audio) | French CoreDeskAI greeting |
| `MISTRAL_CHAT_MODEL` | chat model for the direct-Mistral fallback | `mistral-small-latest` |
| `PORT` | listen port | `8080` |

## Phase 2 — DONE
- ✅ greeting on stream start (Supertonic; replaced the Telnyx `speak`)
- ✅ endpointing + streaming per-sentence reply (Mistral or Convex) → Supertonic → bidirectional WS back to Telnyx
- ✅ **barge-in** — STT activity during playback cancels the bot (`test_barge_sim.py`)
- ✅ **Convex orchestration (LLM layer)** — `CONVEX_LLM_URL` → org LLM + naturalness (streamed)
- ✅ **Deep orchestration** — `CONVEX_VOICE_TURN_URL` → `/voice/turn` runs the full `messageProcessor` pipeline (workflows/tools/KB/MCP/compliance) **and persists** to `conversations`/`messages`; one phone conversation per Telnyx `call_control_id`

## Not yet (Phase 3)
- latency: co-locate the bridge off the home tunnel (the real sub-1.5s fix) — heavy pipeline is ~7s on the dev rig.
  ✅ **Combined image DONE** → `Dockerfile` builds bridge + Supertonic (`supertonic[serve]`, weights baked at
  build via `supertonic download`) and runs both via `entrypoint.sh`. Ready to `fly deploy` — see `DEPLOY.md`.
- ✅ FR-ize tool-error messages — `messageProcessor` data-tool + MCP error replies now follow `context.language`.
- ✅ candidate-path campaign chaining — event-driven via Telnyx `client_state` + the `call.hangup` webhook.
- run under `voiceMigrationConfigs` (shadow → pilot → full) for measured rollout — **last gated step** (real calls).
- env note: org CData subscription must be active for data tools (`NOT_SUBSCRIBED` 403 otherwise)
