# CoreDeskAI — Architecture Review Report

> **Phase 0 — Pre-Implementation Analysis**
> **Date:** July 2026
> **Scope:** Unified observability, billing, tracing, and AI telemetry architecture

---

## Table of Contents

1. [Current Architecture](#1-current-architecture)
2. [Current Request Lifecycle](#2-current-request-lifecycle)
3. [Current Tracing Lifecycle](#3-current-tracing-lifecycle)
4. [Current Billing Lifecycle](#4-current-billing-lifecycle)
5. [Current LLM Invocation Lifecycle](#5-current-llm-invocation-lifecycle)
6. [Current Token Tracking Flow](#6-current-token-tracking-flow)
7. [Current Provider Abstraction](#7-current-provider-abstraction)
8. [Current Observability Stack](#8-current-observability-stack)
9. [Critical Findings](#9-critical-findings)
10. [Architectural Recommendations](#10-architectural-recommendations)

---

## 1. Current Architecture

### 1.1 System Overview

CoreDeskAI is a multi-tenant AI customer support platform with:

- **3 frontend apps**: Web dashboard (Next.js 16), Chat widget (Next.js 15), Embed script (Vite)
- **1 backend**: Convex (serverless, realtime, reactive DB) with ~60 tables
- **3 services**: LiveKit voice agent (Python), Telnyx bridge (Python, DR), WhatsApp Baileys (Node.js, demo)
- **6 channels**: Widget, Phone, WhatsApp, Instagram, Messenger, Telegram
- **5+ LLM providers**: Groq, OpenAI, Anthropic, Google, Mistral (via Vercel AI SDK)

### 1.2 Technology Stack

```
Frontend:   Next.js 15/16 · React 19 · Tailwind CSS v4 · shadcn/ui · Jotai · Zustand
Backend:    Convex (BaaS) · Clerk (auth) · Resend (email) · AWS Secrets Manager
AI/LLM:     Groq (primary) · OpenAI · Anthropic · Google · Mistral · Voyage (embeddings)
Voice:      Telnyx SIP · LiveKit Agents · Deepgram · Azure OpenAI gpt-realtime
Data:       CData SQL-over-REST · Airbyte · Supabase · Google Sheets
Monitoring: Sentry (errors) · Custom tracing (observability/)
```

### 1.3 Key Architectural Patterns

| Pattern | Implementation | Status |
|---------|---------------|--------|
| Multi-tenancy | Every table keyed by `organizationId` | ✅ Consistent |
| Realtime | Convex reactive queries | ✅ Working |
| Event-driven | `ctx.scheduler.runAfter(0, ...)` fire-and-forget | ⚠️ No retry |
| Provider abstraction | `LLMProviderManager` (5 providers) | ⚠️ Groq excluded |
| Policy engine | Two-layer: system checks + org-configurable rules | ✅ Working |
| Workflow engine | v2 with branching, approvals, compensation | ✅ Working |
| Knowledge base | Voyage vectors + TF-IDF fallback + optional reranking | ✅ Working |

---

## 2. Current Request Lifecycle

### 2.1 Inbound Message Flow

```
INBOUND MESSAGE (any channel)
    │
    ▼
incomingMessage.processIncomingMessage
    ├─ 1. Resolve socialChannel → orgId
    ├─ 2. findOrCreateByChannelIdentifier → sessionId
    ├─ 3. Find/create conversation → threadId
    ├─ 4. Save user message to thread
    ├─ 4b. Check for waiting_confirmation workflow
    └─ 5. ctx.runAction → messageProcessor.processMessage
         │
         ▼
    runProcessMessage
         ├─ [0] Suspension guard
         ├─ [0] CREATE TRACER ← trace begins
         ├─ [0] Quota guard (checkQuota query)
         ├─ [0a] Fetch widgetSettings (try/catch, warn)
         ├─ [0a] Resolve active agent (try/catch, warn)
         ├─ [0b] Resolve org chat model (try/catch, warn)
         ├─ [1] Context extraction (runAgent → Groq Llama-8B)
         │   └─ scheduleTokenLog() for extraction usage
         ├─ [1a] Load thread history (try/catch, warn)
         ├─ [1a] Platform language default (FR/AR → Mistral)
         ├─ [1b] Escalation intercept (keyword + LLM rule check)
         │   └─ scheduleTokenLog() for escalation rule
         ├─ [2] Workflow matching (4-tier)
         │
         ├─ IF NO WORKFLOW:
         │   ├─ Retrieval decision (decideRetrievalAsync)
         │   ├─ KB search (tryAnswerFromKnowledge)
         │   │   └─ scheduleTokenLog() for KB interpretation
         │   ├─ Direct tool call (Groq tool-calling)
         │   │   └─ scheduleTokenLog() for tool selection
         │   ├─ MCP/CData fallback
         │   │   └─ scheduleTokenLog() for MCP selection
         │   └─ Fallback formulateReply
         │       └─ scheduleTokenLog()
         │
         ├─ IF WORKFLOW MATCHED:
         │   ├─ Pre-execution guards
         │   ├─ Required entities check
         │   ├─ startWorkflowExecution
         │   └─ executeWorkflowStep
         │       ├─ actionExecutor.executeNextStep
         │       │   ├─ CREATE NESTED TRACER ← orphaned
         │       │   ├─ executeStep (policy/DB/API)
         │       │   └─ executionLogger.log
         │       ├─ Update execution state
         │       └─ sendWorkflowReply
         │           ├─ Language detection (UNBILLED LLM)
         │           ├─ Translation (UNBILLED LLM)
         │           └─ saveThreadMessage
         │
         └─ FINISH TRACER ← trace ends (NOT persisted)
              │
              └─ Return result with trace summary
```

### 2.2 Data Flow Summary

```
Inbound Message
    → messageProcessor.processMessage (action)
        → [Tracing] tracer.startStep / span.end
        → [Billing] scheduleTokenLog → recordTokenUsage (scheduled mutation)
        → [Evaluation] logEventFromMessage (scheduled mutation)
        → [Execution] executionLogger.log (scheduled mutation)
    → Return result
```

---

## 3. Current Tracing Lifecycle

### 3.1 Trace Types

```typescript
Trace = {
  traceId: string;                    // "tr_{timestamp}_{random6}"
  sessionId, conversationId, organizationId;
  startTime, endTime;
  status: "success" | "error" | "partial";
  steps: TraceStep[];
  metadata: { model, channel, language, intent, workflowName };
}

TraceStep = {
  stepId, parentStepId?, type, name;
  startTime, endTime, durationMs;
  status: "success" | "error" | "pending";
  input?, output?, error?;
  metrics?: { tokensIn, tokensOut, costUsd, latencyMs, model };
}

TraceSummary = {
  totalLatencyMs, totalTokens, totalCostUsd;
  stepCount, errorCount;
  breakdown: { llmMs, toolMs, ragMs, orchestrationMs, ... };
}
```

### 3.2 Tracer Creation Points

| Location | File | Line | Status |
|----------|------|------|--------|
| Main pipeline | `messageProcessor.ts` | 204 | Created and finished ✅ |
| Workflow step | `actionExecutor.ts` | 244 | Created but **NEVER FINISHED** ❌ |

### 3.3 Trace Step Coverage in messageProcessor.ts

| Step | Type | Traced? | Token Tracked? |
|------|------|---------|----------------|
| Suspension guard | — | ❌ | N/A |
| Quota check | quota_check | ✅ | N/A |
| Widget settings fetch | — | ❌ | N/A |
| Agent resolution | — | ❌ | N/A |
| Model resolution | — | ❌ | N/A |
| Context extraction | context_extraction | ✅ | ✅ (scheduleTokenLog) |
| Thread history load | — | ❌ | N/A |
| Escalation rule check | — | ❌ | ✅ (scheduleTokenLog) |
| Escalation execution | escalation | ✅ | ✅ (scheduleTokenLog) |
| Workflow matching | workflow_selection | ✅ | N/A |
| KB search (tryAnswerFromKnowledge) | — | ❌ | ✅ (scheduleTokenLog) |
| Tool selection LLM | — | ❌ | ✅ (scheduleTokenLog) |
| Tool execution | — | ❌ | N/A |
| Tool error explanation | — | ❌ | ✅ (scheduleTokenLog) |
| Write confirmation | — | ❌ | ✅ (scheduleTokenLog) |
| MCP tool selection | — | ❌ | ✅ (scheduleTokenLog) |
| MCP tool execution | — | ❌ | N/A |
| MCP reply formulation | — | ❌ | ✅ (scheduleTokenLog) |
| Fallback formulation | response_composition | ✅ | ✅ (scheduleTokenLog) |
| Workflow execution | workflow_execution | ✅ | N/A |
| Missing entities reply | — | ❌ | ✅ (scheduleTokenLog) |

**Coverage: ~30% of LLM calls are traced. 70% have token tracking but no trace step.**

### 3.4 Storage Backend

| Backend | Status | Persistence | Query Support |
|---------|--------|-------------|---------------|
| InMemoryTraceStorage | Exists | ❌ None (lost on restart) | Basic filtering |
| Convex-backed storage | **MISSING** | Would be durable | Would use Convex indexes |
| Production wiring | **MISSING** | Traces are ephemeral | N/A |

### 3.5 Dead Code

| File | Purpose | Status |
|------|---------|--------|
| `formulateReplyTraced.ts` | Traced wrapper for formulateReply | **Never imported** |
| `toolExecutorTraced.ts` | Traced wrapper for tool execution | **Never imported** |
| `llmMetrics.trackLLMCall()` | LLM call tracking helper | Created but **never called** |
| `toolTracker.trackToolCall()` | Tool call tracking helper | Created but **never called** |

---

## 4. Current Billing Lifecycle

### 4.1 Credit Calculation

```typescript
CREDITS_PER_1K_TOKENS = 1
creditsConsumed = Math.ceil(totalTokens / 1000) * CREDITS_PER_1K_TOKENS
```

**Same rate regardless of provider or model.** GPT-4o and Llama-8B cost the same.

### 4.2 Dual Credit Deduction Systems

| System | Trigger | Deduction | Where |
|--------|---------|-----------|-------|
| `recordTokenUsage` | After LLM call (async) | `ceil(tokens/1000)` credits | `tokenTracking.ts:78` |
| `consumeQuota` | Tool call execution | 1 credit per call | `quotas.ts:134` |

Both operate on the **same `orgQuotas.creditsRemaining` field**. This creates:
- **Double deduction risk**: LLM tokens deducted + tool call credit deducted for same request
- **Inconsistent defaults**: `recordTokenUsage` sets `creditLimit: 500`, `consumeQuota` does not
- **Divergent reads**: `public/quotas.ts:getMyQuota` recalculates from logs instead of reading the quota row

### 4.3 Period Reset Bug

```typescript
// tokenTracking.ts:72-75 — reset on expiry
if (isExpired) {
  resetAt: nextMonthReset(),
  toolCallsUsed: 0,
  // BUG: creditAlertsSent NOT reset → alerts never fire again
  // BUG: tokensUsed NOT explicitly reset
}
```

### 4.4 Quota Check vs Deduction Race

```
Time ──────────────────────────────────────────────────────►

Request A: [checkQuota: 2 credits ✅] ──LLM call──► [recordTokenUsage: deduct 500]
Request B: [checkQuota: 2 credits ✅] ──LLM call──► [recordTokenUsage: deduct 500]
                                                           │
                                                           ▼
                                                    credits: 0 (clamped)
                                                    B consumed 500 tokens for FREE
```

The quota check is a **read** at the start of processing. Credits are deducted **asynchronously** after the LLM call completes. Multiple requests can slip through.

### 4.5 Unbilled LLM Calls

| Location | Call | Tokens Consumed | Billed? |
|----------|------|----------------|---------|
| `workflowTriggers.ts:42` | Language detection | ~50-100 | ❌ |
| `workflowTriggers.ts:140` | Translation | ~100-200 | ❌ |
| `workflowTriggers.ts:1512` | classifyAndDetectLanguage | ~100 | ❌ |
| `workflowTriggers.ts:1551` | replyInLanguage | ~200 | ❌ |
| `workflowTriggers.ts:1579` | translateToTarget | ~200 | ❌ |
| `voiceTurn.ts:187` | Campaign answer recording | ~100 | ❌ |
| `voiceTurn.ts:293` | Campaign script turn | ~200 | ❌ |
| `LanguageDetection.ts:22` | Language detection | ~50 | ❌ |
| `selectiveRag.ts:404` | RAG router LLM | ~50 | ❌ |
| `tools/search.ts:77` | KB search interpretation | ~100 | ❌ |
| `extractTextContent.ts:135,163,184` | File processing | ~200-500 | ❌ |
| `private/messages.ts:48` | Operator enhancement | ~100 | ❌ |

**Estimated: 10-20% of LLM tokens are unbilled.**

---

## 5. Current LLM Invocation Lifecycle

### 5.1 Two Parallel Call Paths

```
┌─────────────────────────────────┬──────────────────────────────────────┐
│ PATH A: ORG-SELECTABLE MODEL    │ PATH B: INTERNAL/INFRASTRUCTURE LLM │
│ (Customer-facing replies)       │ (Classification, extraction, tools) │
│                                 │                                      │
│ resolveOrgChatModelConfig()     │ groqClient (groq-sdk)               │
│ → LLMProviderManager.getModel() │ groq() (@ai-sdk/groq)               │
│ → LanguageModelV1               │ openai() (@ai-sdk/openai) [1 call]  │
│ → generateText()                │ Direct SDK instantiation             │
│                                 │ (no abstraction)                     │
│ Used by: formulateReply         │ Used by: everything else (20+ sites) │
└─────────────────────────────────┴──────────────────────────────────────┘
```

### 5.2 Provider Abstraction

| Provider | Manager? | SDK Used | Call Sites |
|----------|----------|----------|------------|
| Groq | ❌ NO | `groq-sdk` (native) + `@ai-sdk/groq` | 20+ |
| OpenAI | ✅ YES | `@ai-sdk/openai` | 2 (1 via manager, 1 direct) |
| Anthropic | ✅ YES | `@ai-sdk/anthropic` | 0 (configured, not used) |
| Google | ✅ YES | `@ai-sdk/google` | 0 (configured, not used) |
| Azure | ✅ YES | `@ai-sdk/azure` | 0 (configured, not used) |
| Mistral | ✅ YES | `@ai-sdk/mistral` | 0 (configured via language default) |

**Groq is the workhorse (20+ call sites) but operates entirely outside the provider abstraction.**

### 5.3 Three Duplicate Groq Clients

| Client | File | Purpose |
|--------|------|---------|
| `new Groq({apiKey})` | `utils/Agent.ts` | Tool selection, MCP selection |
| `new Groq({apiKey})` | `utils/grokClient.ts` | Unused/duplicate |
| `new Groq({apiKey})` | `utils/formulateReply.ts` | Default reply generation |

Plus scattered `@ai-sdk/groq` dynamic imports in `messageProcessor.ts` and `selectiveRag.ts`.

### 5.4 Provider Label Bug

In `messageProcessor.ts`, ALL `scheduleTokenLog` calls hardcode `"groq"` as the provider:

```typescript
// Line 447 — even if org model is OpenAI/Anthropic/etc.
await scheduleTokenLog(ctx, args.organizationId, "groq", escalationResult.usage, ...);
// Line 779 — same issue
await scheduleTokenLog(ctx, args.organizationId, "groq", toolReplyResult.usage, ...);
```

**The actual provider used (which could be OpenAI, Anthropic, etc.) is always logged as "groq".**

### 5.5 Token Usage Extraction

Two formats exist:

| SDK | Usage Format | Extraction Helper |
|-----|-------------|-------------------|
| `groq-sdk` (native) | `usage.prompt_tokens`, `usage.completion_tokens` | `extractGroqUsage()` (exists, unused) |
| `@ai-sdk/*` (AI SDK) | `usage.promptTokens`, `usage.completionTokens` | `extractAISdkUsage()` (exists, unused) |

Each call site does its own inline extraction instead of using the helpers.

---

## 6. Current Token Tracking Flow

### 6.1 Call Chain

```
LLM API Call
    │
    ▼
Usage Object { promptTokens, completionTokens, totalTokens, model }
    │
    ├─[messageProcessor.ts]──► scheduleTokenLog() (local async helper)
    │       │
    │       ▼
    │   ctx.scheduler.runAfter(0, internal.system.tokenTracking.recordTokenUsage, {...})
    │
    ├─[conversationSummary.ts]──► ctx.scheduler.runAfter(0, recordTokenUsage, {...})
    │       (direct call, no scheduleTokenLog wrapper)
    │
    ├─[llmJudge.ts]──► ctx.scheduler.runAfter(0, recordTokenUsage, {...})
    │       (direct call, no scheduleTokenLog wrapper)
    │
    ▼
recordTokenUsage (tokenTracking.ts)
    │
    ├─ (1) INSERT tokenUsageLogs row
    ├─ (2) QUERY orgQuotas (auto-create if missing)
    ├─ (3) PATCH orgQuotas (deduct credits, update token count)
    └─ (4) Check credit thresholds → schedule sendCreditAlert if needed
```

### 6.2 All RecordTokenUsage Call Sites (14 total)

| # | File | CallSite | Provider |
|---|------|----------|----------|
| 1 | `messageProcessor.ts:113` | `messageProcessor:kbSearch` | groq |
| 2 | `messageProcessor.ts:329` | `runAgent:main` (dynamic) | groq |
| 3 | `messageProcessor.ts:418` | `messageProcessor:escalationRule` | groq |
| 4 | `messageProcessor.ts:447` | `messageProcessor:escalation` | groq (wrong if org model) |
| 5 | `messageProcessor.ts:623` | `messageProcessor:toolSelection` | groq |
| 6 | `messageProcessor.ts:721` | `messageProcessor:toolErrorExplanation` | groq |
| 7 | `messageProcessor.ts:763` | `messageProcessor:writeConfirmation` | groq |
| 8 | `messageProcessor.ts:779` | `messageProcessor:toolReply` | groq (wrong if org model) |
| 9 | `messageProcessor.ts:829` | `messageProcessor:mcpSelection` | groq |
| 10 | `messageProcessor.ts:887` | `messageProcessor:mcpReply` | groq (wrong if org model) |
| 11 | `messageProcessor.ts:923` | `messageProcessor:fallback` | groq (wrong if org model) |
| 12 | `messageProcessor.ts:1105` | `messageProcessor:missingEntities` | groq (wrong if org model) |
| 13 | `conversationSummary.ts:72` | `conversationSummary:generate` | groq |
| 14 | `llmJudge.ts:102` | `llmJudge:evaluateConversation` | openai |

### 6.3 Quota Management — All Read/Write Points

**Reads (10 locations):**
- `quotas.ts:18,34,102,155` — CRUD operations
- `tokenTracking.ts:52,165,195,260` — tracking and alerting
- `admin/billing.ts:57` — admin view
- `public/quotas.ts:13` — user view

**Writes (6 locations):**
- `quotas.ts:110,132,162,172` — quota CRUD
- `tokenTracking.ts:60,80` — auto-create and deduct
- `tokenTracking.ts:179` — mark alert sent

---

## 7. Current Provider Abstraction

### 7.1 What Works

- `LLMProviderManager` is a clean singleton for org-selectable models
- `resolveOrgChatModelConfig` has robust fallback chain (agent → org → language default → null)
- `chatCallSettings` correctly handles provider-specific quirks (no frequency penalties for Anthropic/Google)
- Platform language defaults (Mistral for FR/AR) are smart

### 7.2 What's Broken

| Issue | Impact | Location |
|-------|--------|----------|
| No Groq provider in manager | 20+ calls bypass abstraction | `LLMProviderManager.ts` |
| 3 duplicate Groq clients | Inconsistent config, no shared state | `Agent.ts`, `grokClient.ts`, `formulateReply.ts` |
| Two SDKs for Groq | `groq-sdk` (native) + `@ai-sdk/groq` (AI SDK) | Multiple files |
| Provider label hardcoded to "groq" | Wrong billing data | `messageProcessor.ts` (12 sites) |
| Token extraction helpers unused | Each site reimplements extraction | `llmMetrics.ts` |
| No cost-aware routing | Same credit rate for all providers | `tokenTracking.ts` |

---

## 8. Current Observability Stack

### 8.1 Three Independent Logging Systems

| System | Table/Store | Purpose | Persisted? |
|--------|-------------|---------|------------|
| **Tracing** | InMemoryTraceStorage (ephemeral) | Request-level traces with DAG steps | ❌ Lost on restart |
| **Execution Logs** | `executionLogs` table | Step-level audit trail | ✅ Durable |
| **Evaluation Events** | `system_events` table | Per-message analytics | ✅ Durable |
| **Token Usage** | `tokenUsageLogs` table | Billing audit trail | ✅ Durable (90-day retention) |

### 8.2 Triple Logging of Workflow Steps

A single workflow step outcome is written to **three independent systems**:

```
Workflow Step Completes
    │
    ├─► executionLogger.log("step_succeeded")     → executionLogs table
    ├─► logEventFromMessage (scheduled)            → system_events table
    └─► tracer.startStep/end                       → InMemoryTraceStorage (ephemeral)
```

These are **completely independent** and could diverge if any one fails.

### 8.3 Error Swallowing

| Category | Count | Impact |
|----------|-------|--------|
| `catch { /* non-fatal */ }` | 25+ | Debugging impossible |
| `console.warn` + continue | 15+ | Silent degradation |
| Fire-and-forget with `.catch(() => {})` | 5+ | Data loss |
| Silent `try {} catch {}` | 3+ | Invisible failures |

### 8.4 Fire-and-Forget Data Loss Risk

| System | Method | Retry? | Failure Impact |
|--------|--------|--------|----------------|
| Token tracking | `scheduler.runAfter(0, recordTokenUsage)` | ❌ | Tokens unbilled |
| Evaluation events | `scheduler.runAfter(0, logEventFromMessage)` | ❌ | Analytics gap |
| Execution logs | `scheduler.runAfter(0, executionLogger.log)` | ❌ | Audit trail gap |
| Conversation summary | `scheduler.runAfter(0, generateAndStore)` | ❌ | Missing summaries |
| Credit alerts | `scheduler.runAfter(0, sendCreditAlert)` | ❌ | Missed alerts |

---

## 9. Critical Findings

### 9.1 Duplicate Sources of Truth

| Data | Source 1 | Source 2 | Source 3 | Can Diverge? |
|------|----------|----------|----------|--------------|
| Credits remaining | `orgQuotas.creditsRemaining` | `tokenUsageLogs` sum | `public/quotas.ts` recalc | ✅ YES |
| Workflow step outcomes | `executionLogs` | `system_events` | In-memory trace | ✅ YES |
| Token usage | `tokenUsageLogs` | Trace `totalTokens` | — | ✅ YES |
| Quota creation | `recordTokenUsage` auto-create | `consumeQuota` auto-create | — | ✅ YES (different defaults) |

### 9.2 Tight Coupling

| Coupling | Location | Risk |
|----------|----------|------|
| `messageProcessor.ts` directly imports `scheduleTokenLog` | Line 12 | Can't track tokens outside this file |
| `scheduleTokenLog` is a local function, not exported | Line 30-49 | No reuse possible |
| `formulateReply` directly instantiates Groq client | Line 124 | Can't mock or replace |
| Tracer creation hardcoded in `messageProcessor.ts` | Line 204 | Can't inject storage config |
| `actionExecutor` creates its own isolated tracer | Line 244 | No parent-child relationship |

### 9.3 Hidden Assumptions

| Assumption | Reality | Risk |
|------------|---------|------|
| All LLM calls are Groq | 5+ calls use OpenAI/Anthropic/etc. | Wrong billing |
| `checkQuota` prevents overuse | Async deduction allows race | Overconsumption |
| Traces are persisted | No storage backend wired | Data loss |
| `creditAlertsSent` resets monthly | Never reset | Alerts stop firing |
| `scheduleTokenLog` always succeeds | Fire-and-forget, no retry | Tokens unbilled |
| Provider is always "groq" | Org may use OpenAI/Anthropic | Wrong analytics |

### 9.4 Missing Abstractions

| Missing | Impact | Priority |
|---------|--------|----------|
| Unified LLM call interface | 20+ fragmented call sites | HIGH |
| Single token tracking entry point | 14 scattered `recordTokenUsage` calls | HIGH |
| Trace storage backend | All traces ephemeral | HIGH |
| Cost-aware provider routing | Same credit rate for all models | MEDIUM |
| Retry queue for fire-and-forget | Data loss on scheduler failure | MEDIUM |
| Unified event bus | 3 independent logging systems | MEDIUM |

### 9.5 Scalability Bottlenecks

| Bottleneck | Location | Impact |
|------------|----------|--------|
| Quota check-then-act race | `messageProcessor.ts:215` → async deduction | Overconsumption under load |
| `orgQuotas` read-modify-write | `tokenTracking.ts:51-80` | Write contention per org |
| InMemoryTraceStorage linear scan | `traceQuery.ts` | O(n) query performance |
| 25+ silent catch blocks | Throughout codebase | Invisible failures at scale |

### 9.6 Race Conditions

| Race | Severity | Scenario |
|------|----------|----------|
| Quota check vs async deduction | HIGH | 10 concurrent requests all pass quota check, all consume LLM tokens, credits go negative (clamped to 0) |
| Concurrent `recordTokenUsage` | MEDIUM | Multiple scheduled mutations execute sequentially (Convex-safe), but between quota check and deduction, excess tokens consumed |
| `creditAlertsSent` not reset | MEDIUM | Alerts fire once ever per org, even across billing periods |
| Period reset during request | LOW | Request checks before midnight, deduction happens after — allows one extra request |

### 9.7 Failure Scenarios

| Scenario | Current Behavior | Data Lost |
|----------|-----------------|-----------|
| Scheduler failure | All `runAfter(0, ...)` calls silently fail | Tokens, events, logs, summaries |
| `recordTokenUsage` mutation throws | Scheduled function error isolated | Token record + quota not updated |
| `sendWorkflowReply` fails | Entire 170-line function in try/catch | User never sees reply |
| `actionExecutor` trace not finished | Tracer created but never finalized | Trace data leaked |
| `executionLogger.log` fails | Wrapped in `catch { /* non-fatal */ }` | Audit trail gap |
| Convex DB write contention | Quota patches serialized by Convex | Latency spike |

### 9.8 Data Consistency Risks

| Risk | Tables Affected | Scenario |
|------|----------------|----------|
| `creditsRemaining` diverges from `tokenUsageLogs` sum | `orgQuotas`, `tokenUsageLogs` | Admin adjusts credits, or `consumeQuota` double-deducts |
| `tokensUsed` not reset on period expiry | `orgQuotas` | Token counter grows monotonically |
| `toolCallsUsed` reset but `tokensUsed` not | `orgQuotas` | Tool calls reset but token count doesn't |
| `creditAlertsSent` not reset | `orgQuotas` | Alerts never fire in new period |
| `public/quotas.ts` recalculates from logs | `tokenUsageLogs` | 90-day log deletion makes quota display wrong |

---

## 10. Architectural Recommendations

### 10.1 Single Source of Truth — Token Tracking

**Current:** 14 scattered `scheduleTokenLog` calls + 2 direct `recordTokenUsage` calls + 10+ unbilled calls.

**Target:** Single `TokenTracker` class that:
- Wraps all LLM calls (Groq, OpenAI, Anthropic, etc.)
- Automatically extracts usage from both SDK formats
- Logs to `tokenUsageLogs` with correct provider/model
- Deducts from `orgQuotas` atomically
- Emits events for observability

### 10.2 Single Source of Truth — Tracing

**Current:** Ephemeral InMemoryTraceStorage, 30% LLM call coverage, orphaned nested tracers.

**Target:** Convex-backed trace storage + automatic instrumentation:
- Every LLM call creates a trace step (via `TokenTracker` or `LLMCaller`)
- Every tool call creates a trace step (via `ToolExecutor`)
- Parent-child relationships preserved across `actionExecutor`
- Traces persisted to Convex table on `finish()`

### 10.3 Unified LLM Caller

**Current:** 20+ direct Groq calls + 5 provider-managed calls.

**Target:** Single `callLLM()` function that:
- Accepts provider, model, messages, options
- Routes through `LLMProviderManager` (including Groq)
- Returns `{ text, usage: { promptTokens, completionTokens }, provider, model, costUsd }`
- Automatically feeds `TokenTracker` and `Tracer`

### 10.4 Atomic Quota Management

**Current:** Check (read) → LLM call → Deduct (write) with race window.

**Target:** Deduct-before-execute pattern:
- Reserve credits at quota check time (optimistic or pessimistic)
- Adjust actual usage after LLM call completes
- Refund on failure

### 10.5 Event-Driven Architecture

**Current:** 5 independent fire-and-forget schedulers with no retry.

**Target:** Convex-backed event queue:
- All side effects (token logging, evaluation events, execution logs) go through a single event bus
- Retry with exponential backoff
- Dead letter queue for failed events
- At-least-once delivery guarantees

### 10.6 Observability-First Design

**Current:** Tracing, execution logs, and evaluation events are separate systems.

**Target:** Unified event model:
- Single `PipelineEvent` type that covers all observability needs
- Events flow to trace, execution log, evaluation, and billing simultaneously
- No duplicate logging — one event, multiple projections

---

## Appendix A: File Reference

### Core Files

| File | Path | Purpose |
|------|------|---------|
| Message Processor | `packages/backend/convex/system/messageProcessor.ts` | Main pipeline |
| Token Tracking | `packages/backend/convex/system/tokenTracking.ts` | Credit/token management |
| Credit Alerts | `packages/backend/convex/system/creditAlerts.ts` | Threshold notifications |
| Quotas | `packages/backend/convex/system/quotas.ts` | Quota CRUD |
| Billing | `packages/backend/convex/system/billing.ts` | Usage events + rollups |
| Tracer | `packages/backend/convex/observability/tracer.ts` | Core trace engine |
| LLM Metrics | `packages/backend/convex/observability/llmMetrics.ts` | LLM tracking helper |
| Tool Tracker | `packages/backend/convex/observability/toolTracker.ts` | Tool tracking helper |
| Trace Types | `packages/backend/convex/observability/trace.ts` | Types + pricing |
| Trace Storage | `packages/backend/convex/observability/storage/TraceStorage.ts` | Storage interface |
| InMemory Storage | `packages/backend/convex/observability/storage/InMemoryTraceStorage.ts` | Dev storage |
| Action Executor | `packages/backend/convex/lib/actionExecutor/actionExecutor.ts` | Workflow execution |
| Execute Step | `packages/backend/convex/lib/actionExecutor/executeStep.ts` | Step execution |
| Workflow Triggers | `packages/backend/convex/system/workflowTriggers.ts` | Trigger handling |
| Formulate Reply | `packages/backend/convex/utils/formulateReply.ts` | Reply generation |
| Resolve Chat Model | `packages/backend/convex/system/ai/resolveChatModel.ts` | Model routing |
| LLM Manager | `packages/backend/convex/system/ai/llm/LLMProviderManager.ts` | Provider abstraction |
| Selective RAG | `packages/backend/convex/system/ai/selectiveRag.ts` | RAG routing |
| Schema | `packages/backend/convex/schema.ts` | Database schema |
| Cron Handlers | `packages/backend/convex/cronHandlers.ts` | Scheduled tasks |

### Tests

| File | Path | Tests |
|------|------|-------|
| Token Tracking | `packages/backend/convex/system/tokenTracking.test.ts` | 33 |
| Full Suite | `packages/backend/convex/system/fullSuite.test.ts` | 69 |
| Selective RAG | `packages/backend/convex/system/ai/selectiveRag.test.ts` | 14 |
| Evaluation | `packages/backend/convex/lib/evaluationEngine/evaluationTest.test.ts` | 7 |
| Observability | `eval/test-observability.ts` | 5 scenarios |

---

*This report documents the current state of the CoreDeskAI architecture across observability, billing, tracing, and AI telemetry. It identifies 9 critical findings, 6 race conditions, 8 failure scenarios, and 5 data consistency risks that must be addressed in the implementation phases.*
