# Voice Bridge — Production Deploy (Phase 3 co-location)

Goal: get the bridge off the home cloudflared quick tunnel and onto an always-on EU host,
co-located with Supertonic. **This is the real sub-1.5s latency fix** — the dev-rig 2.7–7.6s
is dominated by tunnel RTT (Telnyx-EU ↔ home machine in Tunisia ↔ Mistral-EU).

Recommended host: **Fly.io, region `cdg` (Paris)** — close to Telnyx-EU + Mistral-EU, simple
Docker deploy, transparent WebSocket proxying, easy always-on. (Swappable to Render / a small
EU VM; only `fly.toml` + the deploy commands change.)

## Architecture

```
Telnyx (wss://coredeskai-voice-bridge.fly.dev/?org=&call=&from=)
   └─▶ Fly machine (cdg)
         ├─ bridge.py            (:8080, public)         ── Voxtral STT + Mistral  (api.mistral.ai)
         └─ supertonic serve     (:7788, localhost only)  ◀── bridge calls SUPERTONIC_URL
   bridge.py ──▶ Convex /voice/turn (CONVEX_VOICE_TURN_URL) for orchestration + persistence
```

Both processes run **in one container/machine** so `SUPERTONIC_URL=http://127.0.0.1:7788`
stays a localhost call (Fly process *groups* run on separate machines — localhost wouldn't
reach across them, so a single combined image is required for the co-location win).

## ✅ Combined image is complete

The `Dockerfile` now builds the **combined bridge + Supertonic image** (the earlier
"need the pip package + weights" blocker is resolved):

- **Package:** `supertonic[serve]` on PyPI (`>=1.3.1`) — the `[serve]` extra pulls FastAPI +
  `uvicorn[standard]`, giving the `supertonic serve` HTTP server the bridge calls on `/v1/tts`.
- **System libs:** `libsndfile1` (for `soundfile`) and `libgomp1` (for `onnxruntime`) — installed
  via `apt-get` in the image.
- **Model weights:** auto-downloaded from HuggingFace Hub (`Supertone/supertonic-3`, pinned SHA).
  We run `supertonic download` at **build** time so the weights are baked into the image
  (`/root/.cache/supertonic3`) and a running container never blocks a call on a cold fetch.
- **Process startup:** `entrypoint.sh` runs `supertonic serve --host 127.0.0.1 --port 7788`
  (loopback) **and** `python bridge.py` (public `:8080`); `wait -n` tears the container down if
  either exits so Fly restarts it cleanly.

The model repo is public, so **no HF token is required**. To use a different model, set
`SUPERTONIC_MODEL` (e.g. `supertonic-2`) — but note the build bakes `supertonic-3` (the default),
so a different runtime model would download on first call unless the build is changed to match.

## Env vars — which platform holds which

| Var | Where to set | Notes |
|---|---|---|
| `MISTRAL_API_KEY` | **Fly** (`fly secrets set`) | Voxtral STT auth (+ Mistral chat fallback). Rotate at deploy per the key-rotation note. |
| `CONVEX_VOICE_TURN_URL` | **Fly** | `https://<prod-deployment>.convex.site/voice/turn` — full orchestration + persistence. |
| `VAPI_WEBHOOK_SECRET` | **Fly** (+ already on Convex) | only if the Convex voice endpoints enforce `x-vapi-secret`. |
| `SUPERTONIC_URL` | Fly `[env]` | `http://127.0.0.1:7788` (co-located). Already in `fly.toml`. |
| `VOXTRAL_DELAY_MS`, `ENDPOINT_SILENCE_MS`, `TTS_LANG`, `PORT` | Fly `[env]` | tunables; already in `fly.toml`. |
| `VOICE_ORG_ID` | Fly `[env]` *(optional)* | fallback org only; in prod the org arrives per-call via the WS `?org=` query, so leave unset. |
| `VOICE_BRIDGE_WS_URL` | **Convex** (`npx convex env set`) | `wss://coredeskai-voice-bridge.fly.dev/` — tells `/webhooks/telnyx` to stream to the bridge. **This is the wire that points Telnyx at the deployed bridge.** |
| `NEXT_PUBLIC_CONVEX_URL` | **Vercel** | unchanged; frontend ↔ Convex. |

## Deploy steps

```bash
# 0. one-time
fly auth login
cd coredeskai/services/voice-bridge

# 1. create the app (no immediate deploy) — reads fly.toml
fly launch --no-deploy --copy-config --name coredeskai-voice-bridge --region cdg

# 2. set secrets (NOT in fly.toml)
fly secrets set MISTRAL_API_KEY=… CONVEX_VOICE_TURN_URL=https://<deployment>.convex.site/voice/turn
# (add VAPI_WEBHOOK_SECRET if the Convex endpoints enforce it)

# 3. deploy
fly deploy

# 4. point Convex/Telnyx at the deployed bridge, then redeploy the backend
cd ../../packages/backend
npx convex env set VOICE_BRIDGE_WS_URL wss://coredeskai-voice-bridge.fly.dev/
npx convex deploy        # (or `npx convex dev --once` for the dev deployment)
```

## Validate

1. `fly logs` → expect `[telnyx] media stream connected` on a real inbound/outbound call.
2. Place a gated test call (per the no-outbound-calls rule — confirm before dialing). Watch the
   `[lat]` line in `fly logs`; target end-of-speech→first-audio well under the 2.7s tunnel baseline.
3. Re-check `+216` delivery — that's a route/infra issue (premium route / local DID), independent
   of where the bridge runs.

## Then: the measured ramp (rest of Phase 3)

Once latency is confirmed in prod, flip a pilot org to `policyMode:pilot` + candidate under
`voiceMigrationConfigs` and run the shadow→pilot→full ramp with acceptance gates / auto-rollback.
**This is the last gated step** — it places real billed calls, so it needs explicit per-call
confirmation (the no-outbound-calls rule).

Already closed this round: candidate-path campaign chaining (now event-driven — the remaining
queue rides in Telnyx `client_state` and `/webhooks/telnyx` `call.hangup` dials the next contact;
capped at 20 in-flight for the `client_state` size guard) and FR-izing the pipeline tool-error
messages (`messageProcessor` data-tool + MCP error replies now follow `context.language`).
