# Observability (Grafana / Loki / Prometheus)

Optional compose profile that adds log search, resource metrics, and Telegram
alerts — without a new subdomain.

## URL

**https://topup.kedebah.com/monitoring/**

Edge proxies `/monitoring/` → Grafana (subpath mode). Host nginx that already
forwards the whole host to the edge does not need a new vhost.

## What runs

| Service | Role |
|---------|------|
| `grafana` | UI + alerting |
| `loki` | Log store (~7 day retention) |
| `promtail` | Ships Docker logs for project `kedebah-commerce` |
| `prometheus` | Metrics (~15 day retention) |
| `cadvisor` | Per-container CPU/RAM |
| `node-exporter` | Host memory/CPU/disk |

Approx RAM budget: **2–3 GB** (mem_limits set on each service).

## Enable

```bash
cd /var/www/html/KEDEBAH/kedebah-commerce-deploy   # or your deploy path

# 1) .env
# GRAFANA_ADMIN_USER=admin
# GRAFANA_ADMIN_PASSWORD=<strong password>
# TELEGRAM_BOT_TOKEN=<from @BotFather>
# TELEGRAM_CHAT_ID=<group or user id>

# 2) Rebuild edge so /monitoring/ exists, then start the profile
docker compose build edge
docker compose up -d edge
docker compose --profile observability up -d

# 3) Confirm
docker compose --profile observability ps
curl -sI https://topup.kedebah.com/monitoring/ | head -5
```

Login with `GRAFANA_ADMIN_*`. Open dashboard **Kedebah Commerce — Logs & Resources**,
or **Explore → Loki** with:

```
{project="kedebah-commerce"} |= "ERROR"
{service="auth"} |~ "(?i)ERROR|SQLSTATE"
```

## Telegram alerts

Provisioned contact point **telegram** + sample rules:

- ERROR / CRITICAL / SQLSTATE spike (Loki, >15 lines / 5m)
- Host MemAvailable &lt; 12% for 5m (Prometheus)

After setting tokens:

```bash
docker compose --profile observability up -d --force-recreate grafana
```

In Grafana: **Alerting → Contact points → telegram → Test**.

Get `TELEGRAM_CHAT_ID`: message the bot, then open  
`https://api.telegram.org/bot<TOKEN>/getUpdates` and read `chat.id`.

In `.env`, always quote the chat id as a **string** (group ids are often negative):

```bash
TELEGRAM_CHAT_ID="-1001234567890"
```

Telegram contact points are written at container start by
`docker/observability/grafana/entrypoint.sh` (quoted strings). Do **not** put a
static `contactpoints.yml` under provisioning — Grafana’s `$VAR` expansion
turns numeric chat ids into JSON numbers and crash-loops Grafana (502).


## Disable

```bash
docker compose --profile observability stop
# or remove:
docker compose --profile observability down
```

`/monitoring/` will 502 until the profile is up again (edge route stays).

## Notes

- Apps already log to stderr (`LOG_CHANNEL=stderr`) — Promtail picks them up.
- Do not expose Grafana without a strong admin password.
- Tune alert thresholds in the Grafana UI after a few days of real traffic.
