# BeebBeeb on Kubernetes

Manifests for the telemetry-carrying half of the platform. They encode the decisions that make
the fleet scale, so read the comments in the YAML before changing a number.

## What is here

| File | What it does |
|---|---|
| `base/00-namespace.yaml` | the `beebbeeb` namespace |
| `base/01-config.yaml` | non-secret tuning shared by every service |
| `base/02-secrets.example.yaml` | template; never commit real values |
| `base/10-fleet-service.yaml` | telemetry consumer, Deployment + Service + HPA |
| `base/11-booking-service.yaml` | ride lifecycle and trip tracks, same shape |
| `base/12-iot-bridge-service.yaml` | MQTT bridge, deliberately single-replica |
| `base/20-kafka-topics.yaml` | declares partitions, because auto-created topics get 1 |

Infrastructure (Postgres, Redis, Kafka, EMQX) is intentionally **not** here. Run those as managed
services or with their own operators; a StatefulSet of Postgres written by hand is a database you
have to operate yourself at 3am.

## Apply

```bash
kubectl apply -f base/00-namespace.yaml
kubectl -n beebbeeb create secret generic beebbeeb-secrets \
  --from-literal=JWT_SECRET="$(openssl rand -base64 48)" \
  --from-literal=CONFIG_MASTER_KEY="$(openssl rand -base64 32)" \
  --from-literal=DB_PASSWORD='...'
kubectl apply -f base/01-config.yaml
kubectl apply -f base/20-kafka-topics.yaml      # partitions first
kubectl apply -f base/10-fleet-service.yaml -f base/11-booking-service.yaml \
               -f base/12-iot-bridge-service.yaml
```

## The one rule that catches people

**Consumer parallelism is capped by partitions, not by replicas.**

`replicas x KAFKA_CONCURRENCY` must stay at or below a topic's partition count. With 12 partitions
and `KAFKA_CONCURRENCY: 3`, four pods saturate it; a fifth pod holds no partition, consumes
nothing, and costs money. Raise partitions first if you need more throughput, then raise
`maxReplicas`. Partitions can be increased but never decreased, so start generously.

## Why iot-bridge is not autoscaled

Each pod opens its own MQTT subscription, so two pods receive two copies of every device frame and
the platform processes each report twice. Before running more than one:

1. switch its subscriptions to EMQX shared subscriptions (`$share/group/topic`), then
2. raise `replicas` and change the strategy back to `RollingUpdate`.

Until then it is one pod with `Recreate`, which is correct rather than lazy.

## Scaling checklist, in order of effect

1. **Device cadence.** Reporting every 3 s matters during a ride; parked vehicles are fine at
   30-60 s. This is a device setting and it is worth more than any amount of hardware.
2. **Partitions**, then replicas. See the rule above.
3. **`LIVE_FLUSH_SECONDS`.** Higher means fewer Postgres writes; the live answer is in Redis either
   way, so this only affects how stale the stored position is after a crash.
4. **`HISTORY_SAMPLE_SECONDS`.** The difference between a month of history being millions of rows
   or billions. Full-resolution ride paths are recorded separately and are unaffected.
5. **Postgres.** Only after the above. The telemetry path no longer writes per frame, so the
   database is no longer the constraint it was.
