# AI Observability & Trace System

## Overview

Full observability and tracing system for the CoreDeskAI platform. Every AI request now generates a structured trace with step-level latency, token usage, cost estimation, and a replayable execution graph.

---

## Architecture

```
Message Request
    │
    ▼
┌─────────────────────────────────────────┐
│  createTracer({ sessionId, orgId })     │
│                                         │
│  ┌─ quota_check ─────────────────────┐  │
│  └───────────────────────────────────┘  │
│  ┌─ context_extraction ─────────────┐   │
│  │  (runAgent / intent + entities)  │   │
│  └──────────────────────────────────┘   │
│  ┌─ escalation (if triggered) ──────┐   │
│  └──────────────────────────────────┘   │
│  ┌─ workflow_selection ─────────────┐   │
│  │  (4-tier matching)              │   │
│  └──────────────────────────────────┘   │
│  ┌─ knowledge_base_search ─────────┐    │
│  │  (TF-IDF / Voyage + LLM)       │    │
│  └─────────────────────────────────┘    │
│  ┌─ response_composition ──────────┐    │
│  │  (formulateReply + LLM call)    │    │
│  └─────────────────────────────────┘    │
│  ┌─ workflow_execution ────────────┐    │
│  │  ┌─ tool_call ──────────────┐   │    │
│  │  │  (API/DB tool execution) │   │    │
│  │  └──────────────────────────┘   │    │
│  └─────────────────────────────────┘    │
│                                         │
│  tracer.finish() → TraceSummary         │
└─────────────────────────────────────────┘
    │
    ▼
  Response + { trace: { totalLatencyMs, totalTokens, totalCostUsd, breakdown } }
```

---

## Test Results

### Scenario 1: Simple Greeting (no LLM)

```
Trace ID:      tr_1782918484604_acn7gp
Status:        success
Total Latency: 44ms
Total Tokens:  0
Total Cost:    $0.000000
Step Count:    4

Steps:
  [quota_check] checkQuota — 6ms
  [context_extraction] runAgent — 22ms
  [workflow_selection] matchWorkflow — 14ms
  [response_composition] formulateReply:shortcut — 2ms
```

Conversational shortcut detected → no LLM call, zero cost.

---

### Scenario 2: KB Search (RAG + LLM)

```
Trace ID:      tr_1782918484648_7t8sde
Status:        success
Total Latency: 164ms
Total Tokens:  225
Total Cost:    $0.000013
Step Count:    6

Steps:
  [quota_check] checkQuota — 4ms
  [context_extraction] runAgent — 21ms
  [workflow_selection] matchWorkflow — 3ms
  [knowledge_base_search] kbSearch — 136ms
    └─ [rag_query] tfidfSearch — 3ms
    └─ [llm_call] kbInterpretation — 122ms [225 tokens, $0.000013]
```

RAG search (3ms) + LLM interpretation (122ms) — full DAG visible.

---

### Scenario 3: Workflow Trigger (Refund + Tools)

```
Trace ID:      tr_1782918484812_lfegxb
Status:        success
Total Latency: 292ms
Total Tokens:  510
Total Cost:    $0.001300
Step Count:    8

Steps:
  [quota_check] checkQuota — 9ms
  [context_extraction] runAgent — 25ms
  [workflow_selection] matchWorkflow — 23ms
  [workflow_execution] execute:Process Refund — 133ms
    └─ [tool_call] stripe_refund — 81ms
    └─ [tool_call] send_email — 52ms
  [response_composition] formulateReply — 102ms
    └─ [llm_call] formulateReply:llm — 102ms [255 tokens, $0.000650]
```

2 tool calls (Stripe + email) nested under workflow, LLM reply with Mistral pricing.

---

### Scenario 4: Escalation (Human Handoff)

```
Trace ID:      tr_1782918485104_v38ask
Status:        success
Total Latency: 125ms
Total Tokens:  165
Total Cost:    $0.000009
Step Count:    5

Steps:
  [quota_check] checkQuota — 15ms
  [context_extraction] runAgent — 20ms
  [escalation] escalateConversation — 90ms
    └─ [llm_call] escalationRuleCheck — 34ms [65 tokens]
    └─ [llm_call] escalationReply — 56ms [100 tokens]
```

Arabic language escalation on WhatsApp — 2 LLM calls nested under escalation step.

---

### Scenario 5: Complex Workflow (Error + Retry)

```
Trace ID:      tr_1782918485229_fpcas4
Status:        partial
Total Latency: 342ms
Total Tokens:  580
Total Cost:    $0.001480
Step Count:    9
Error Count:   1

Steps:
  [quota_check] checkQuota — 15ms
  [context_extraction] runAgent — 35ms
  [workflow_selection] matchWorkflow — 5ms
  [workflow_execution] execute:Book Appointment — 181ms
    └─ [tool_call] google_calendar_create — 65ms (ERROR)
    └─ [tool_call] google_calendar_create_retry — 73ms (success)
    └─ [tool_call] send_confirmation_email — 41ms
  [response_composition] formulateReply — 106ms
    └─ [llm_call] formulateReply:llm — 106ms [290 tokens, $0.000740]
```

Tool error detected (65ms failure) → retry succeeded (73ms) → status `partial` with `errorCount: 1`.

---

### Cross-Scenario Summary

| # | Scenario | Latency | Tokens | Cost | Steps | Status |
|---|----------|---------|--------|------|-------|--------|
| 1 | Greeting (shortcut) | 44ms | 0 | $0.000000 | 4 | success |
| 2 | KB Search (RAG+LLM) | 164ms | 225 | $0.000013 | 6 | success |
| 3 | Workflow (refund+tools) | 292ms | 510 | $0.001300 | 8 | success |
| 4 | Escalation (human handoff) | 125ms | 165 | $0.000009 | 5 | success |
| 5 | Complex (multi-step+retry) | 342ms | 580 | $0.001480 | 9 | partial |

---

## Files Created (10)

### Core Types & Engine

| File | Purpose |
|------|---------|
| `packages/backend/convex/observability/trace.ts` | Core types: `Trace`, `TraceStep`, `TraceSummary`, `TraceEvent`, pricing table, helpers |
| `packages/backend/convex/observability/tracer.ts` | Main `Tracer` engine — creates traces, manages step spans, computes summaries |
| `packages/backend/convex/observability/llmMetrics.ts` | `LLMMetricsTracker` — wraps LLM calls to capture tokens, cost, latency |
| `packages/backend/convex/observability/toolTracker.ts` | `ToolTracker` — wraps tool/API executions with timing and success tracking |

### Storage

| File | Purpose |
|------|---------|
| `packages/backend/convex/observability/storage/TraceStorage.ts` | Pluggable storage interface |
| `packages/backend/convex/observability/storage/InMemoryTraceStorage.ts` | In-memory implementation (dev/test) |

### Integration Wrappers

| File | Purpose |
|------|---------|
| `packages/backend/convex/observability/formulateReplyTraced.ts` | Trace-aware wrapper for `formulateReply` |
| `packages/backend/convex/observability/toolExecutorTraced.ts` | Trace-aware wrapper for `CustomApiToolExecutor` |

### Query & Export

| File | Purpose |
|------|---------|
| `packages/backend/convex/observability/traceQuery.ts` | Convex queries: `getTrace`, `listTraces`, `listTraceSummaries` |
| `packages/backend/convex/observability/index.ts` | Barrel export |

### Test

| File | Purpose |
|------|---------|
| `eval/test-observability.ts` | 5-scenario integration test suite |

---

## Files Modified (2)

| File | Changes |
|------|---------|
| `packages/backend/convex/system/messageProcessor.ts` | Added tracing around: quota check, context extraction, escalation, workflow matching, KB search, fallback reply, workflow execution. Returns `trace` summary with every response. |
| `packages/backend/convex/lib/actionExecutor/actionExecutor.ts` | Added tracing around each workflow step execution (`executeStep` call) |

---

## Trace Output Format

Every `processMessage` call now returns:

```json
{
  "type": "workflow_result",
  "message": "...",
  "trace": {
    "traceId": "tr_1234567890_abc123",
    "totalLatencyMs": 1234,
    "totalTokens": 450,
    "totalCostUsd": 0.00042,
    "stepCount": 6,
    "errorCount": 0,
    "breakdown": {
      "llmMs": 800,
      "toolMs": 200,
      "ragMs": 150,
      "orchestrationMs": 84,
      "intentDetectionMs": 0,
      "workflowSelectionMs": 0,
      "responseCompositionMs": 0
    },
    "status": "success",
    "metadata": {
      "channel": "widget",
      "workflowName": "Book Appointment",
      "language": "fr"
    }
  }
}
```

---

## Step Types Traced

| Step Type | What's Captured |
|-----------|-----------------|
| `quota_check` | Credit check latency |
| `context_extraction` | Intent detection, entity extraction, language detection |
| `workflow_selection` | Matching logic, matched workflow name |
| `knowledge_base_search` | Chunk count, hit/miss, LLM interpretation |
| `response_composition` | LLM model, tokens in/out, cost, latency |
| `workflow_execution` | Step execution, success/failure |
| `tool_call` | Tool name, input/output, timing, success |
| `llm_call` | Model, provider, tokens, cost, latency |
| `escalation` | Escalation detection, human handoff |
| `rag_query` | Vector/TF-IDF search, top-k results |

---

## Pricing Table

Default cost estimates per model (USD per 1K tokens):

| Model | Input | Output |
|-------|-------|--------|
| llama-3.1-8b-instant | $0.00005 | $0.00008 |
| llama-3.3-70b-versatile | $0.00059 | $0.00079 |
| gpt-4o | $0.0025 | $0.01 |
| gpt-4o-mini | $0.00015 | $0.0006 |
| claude-sonnet-4-20250514 | $0.003 | $0.015 |
| mistral-large-latest | $0.002 | $0.006 |
| gemini-2.0-flash | $0.0001 | $0.0004 |

---

## Design Principles

1. **Zero blind execution** — no LLM/tool call runs without being traced
2. **No implicit steps** — if it happens in code, it's a TraceStep
3. **No performance degradation** — async-safe, <5% overhead target
4. **Fail-safe behavior** — if tracing fails, system continues with minimal logging
5. **Deterministic structure** — structured JSON, machine-readable, replayable

---

## Usage

### In messageProcessor.ts

```ts
import { createTracer, createLLMMetricsTracker, createToolTracker } from "../observability";

const tracer = createTracer({ sessionId, conversationId, organizationId, channel });
const llmMetrics = createLLMMetricsTracker(tracer);
const toolTracker = createToolTracker(tracer);

const span = tracer.startStep("llm_call", "myLlmCall");
// ... do work ...
span.end({ metrics: { tokensIn: 100, tokensOut: 50, costUsd: 0.001 } });

const summary = await tracer.finish(); // persists to storage
```

### Querying Traces

```ts
import { getTrace, listTraces, listTraceSummaries } from "../observability/traceQuery";

const trace = await getTrace(ctx, { traceId: "tr_xxx" });
const summaries = await listTraceSummaries(ctx, { organizationId: "org_xxx", limit: 10 });
```
