# CoreDeskAI — Full Project Description

> **Version:** 1.0 · **Last Updated:** July 2026 · **Status:** Active Development

---

## Table of Contents

1. [Executive Summary](#1-executive-summary)
2. [Architecture Overview](#2-architecture-overview)
3. [Monorepo Structure](#3-monorepo-structure)
4. [Tech Stack](#4-tech-stack)
5. [Backend — Convex Functions](#5-backend--convex-functions)
6. [Database Schema](#6-database-schema)
7. [AI & LLM System](#7-ai--llm-system)
8. [Knowledge Base (RAG)](#8-knowledge-base-rag)
9. [Workflow Engine](#9-workflow-engine)
10. [Channel System](#10-channel-system)
11. [Voice & Telephony](#11-voice--telephony)
12. [Custom API Tools & Marketplace](#12-custom-api-tools--marketplace)
13. [Outbound Campaigns](#13-outbound-campaigns)
14. [Evaluation & Observability](#14-evaluation--observability)
15. [Security & Governance](#15-security--governance)
16. [Billing & Quotas](#16-billing--quotas)
17. [Frontend — Web Dashboard](#17-frontend--web-dashboard)
18. [Frontend — Chat Widget](#18-frontend--chat-widget)
19. [Frontend — Embed Script](#19-frontend--embed-script)
20. [Services](#20-services)
21. [Integrations](#21-integrations)
22. [Deployment & CI/CD](#22-deployment--cicd)
23. [Environment Variables](#23-environment-variables)
24. [Running the Project](#24-running-the-project)
25. [Testing](#25-testing)

---

## 1. Executive Summary

**CoreDeskAI** is a **multi-tenant AI-powered customer support and outbound voice automation platform**. It enables organizations to deploy intelligent, omnichannel chat and voice agents that:

- Handle customer conversations across web widget, phone, WhatsApp, Telegram, Instagram, and Messenger
- Execute automated multi-step workflows with human-in-the-loop approvals
- Query knowledge bases via Retrieval-Augmented Generation (RAG)
- Run outbound phone qualification and pitch campaigns
- Provide real-time observability, evaluation, and cost tracking

**Target market:** EU/GDPR-oriented businesses. French is the product default language, with Arabic and English as supported languages. Hosted in EU (Ireland for Convex, Paris region for Vercel).

---

## 2. Architecture Overview

```
┌─────────────────────────────────────────────────────────────────┐
│                        CLIENT LAYER                              │
│  ┌──────────┐  ┌──────────┐  ┌──────────┐  ┌────────────────┐  │
│  │ Web App  │  │ Widget   │  │ Embed    │  │ Voice Agent    │  │
│  │ (Next.js)│  │ (Next.js)│  │ (Vite)   │  │ (LiveKit/Py)   │  │
│  │ :3000    │  │ :3001    │  │ :3002    │  │                │  │
│  └────┬─────┘  └────┬─────┘  └────┬─────┘  └───────┬────────┘  │
│       │              │              │                │            │
│  ┌────┴──────────────┴──────────────┴────────────────┴────┐     │
│  │              CONVEX REALTIME DATABASE                    │     │
│  │  Schema (~60 tables) · Queries · Mutations · Actions    │     │
│  │  HTTP Router · Cron Jobs · Observability                │     │
│  └────┬────────────────────────────────────────────────────┘     │
│       │                                                          │
├───────┼──────────────────────────────────────────────────────────┤
│       │            INTEGRATION LAYER                              │
│  ┌────┴─────────────────────────────────────────────────────┐   │
│  │  AI/LLM: Groq · OpenAI · Anthropic · Google · Mistral    │   │
│  │  RAG: Voyage AI · @convex-dev/rag · TF-IDF fallback      │   │
│  │  Auth: Clerk (orgs, SSO, JWT)                             │   │
│  │  Voice: Telnyx SIP · LiveKit · Deepgram · Azure S2S      │   │
│  │  Social: WhatsApp · Instagram · Messenger · Telegram      │   │
│  │  Data: CData SQL · Airbyte · Supabase · Google Sheets     │   │
│  │  Email: Resend · Gmail SMTP                               │   │
│  │  Secrets: AWS Secrets Manager                             │   │
│  └──────────────────────────────────────────────────────────┘   │
└─────────────────────────────────────────────────────────────────┘
```

### Six-Layer Architecture

| Layer | Responsibility |
|-------|---------------|
| **Channels** | Receive/send messages across widget, phone, social channels |
| **Orchestration** | Message routing, workflow matching, agent resolution |
| **Knowledge** | RAG retrieval, vector search, KB interpretation |
| **Integration** | External API tools, MCP, OAuth connectors |
| **Governance** | Policy engine, approvals, quotas, compliance |
| **Observability** | Tracing, metrics, evaluation, alerts |

---

## 3. Monorepo Structure

```
coredeskai/
├── apps/
│   ├── web/                  Main Next.js dashboard (port 3000)
│   ├── widget/               Embeddable chat widget (port 3001)
│   └── embed/                Lightweight embeddable script (Vite, port 3002)
│
├── packages/
│   ├── backend/              Core Convex backend (schema, functions, HTTP, crons)
│   ├── ui/                   Shared design system (Radix + Tailwind + shadcn/ui)
│   ├── eslint-config/        Shared ESLint configuration
│   └── typescript-config/    Shared TypeScript configuration
│
├── services/
│   ├── voice-livekit/        Python LiveKit agent worker (production voice)
│   ├── voice-bridge/         Python Telnyx media bridge (disaster recovery)
│   └── whatsapp-baileys/     Node.js unofficial WhatsApp channel (demo)
│
├── eval/                     RAG evaluation scripts, comparison benchmarks
├── docs/                     Architecture and operational documentation
└── convex/                   Generated Convex types
```

---

## 4. Tech Stack

### Frontend

| Layer | Technology |
|-------|-----------|
| Framework | **Next.js** 15/16 (App Router) |
| UI Library | **React 19** |
| Styling | **Tailwind CSS v4** + `tw-animate-css` |
| Component Library | **shadcn/ui** (Radix UI primitives) |
| State Management | **Jotai** (widget), **Zustand** (web) |
| Forms | **React Hook Form** + **Zod** validation |
| Charts | **Recharts** |
| Flow Diagrams | **@xyflow/react** (workflow builder) |
| Icons | **Lucide React** |
| Theming | **next-themes** (light/dark/system) |
| Notifications | **Sonner** (toast) |

### Backend

| Layer | Technology |
|-------|-----------|
| Backend-as-a-Service | **Convex** (serverless, realtime, reactive DB) |
| Auth | **Clerk** (organizations, SSO, JWT templates) |
| Webhook Verification | **Svix** (Clerk webhook signatures) |
| Email | **Resend** + **Nodemailer** (Gmail proxy) |
| Secrets | **AWS Secrets Manager** (prod) / Convex DB (dev) |
| Observability | Custom tracing system (traces, steps, DAG) |

### AI / LLM Providers

| Provider | Usage |
|----------|-------|
| **Groq** | Primary LLM — Llama 3.3 70B (chat), Llama 3.1 8B (KB interpreter, intent) |
| **Mistral** | Default for French/Arabic replies (auto-routed by language) |
| **OpenAI** | Configurable per-org; LLM-as-Judge (GPT-4o) |
| **Anthropic** | Configurable per-org |
| **Google (Gemini)** | Configurable per-org; voice agent fallback |
| **Azure OpenAI** | Configurable per-org; `gpt-realtime` for voice (Sweden Central EU) |
| **Voyage AI** | Embeddings (`voyage-3.5-lite`, 1024d) + optional reranker for RAG |
| **Cohere** | Eval comparison only (`embed-v4.0`) |

### Voice / Telephony

| Component | Technology |
|-----------|-----------|
| SIP/Telephony | **Telnyx** (SIP trunking) |
| Voice AI (production) | **LiveKit Agents** + **Deepgram** Nova-3 STT + Groq LLM + Aura-2 TTS |
| Voice AI (target) | **Azure OpenAI** `gpt-realtime` speech-to-speech (Sweden Central) |
| Voice AI (disaster recovery) | Telnyx ↔ Python bridge ↔ Voxtral STT ↔ Mistral/TTS |

### Database

| System | Usage |
|--------|-------|
| **Convex Database** | Primary data store (serverless, realtime, EU Ireland region) |
| **@convex-dev/rag** | Vector embeddings for knowledge base (RAG) |
| **@convex-dev/agent** | Agent framework for Convex |

---

## 5. Backend — Convex Functions

### Directory Layout

```
packages/backend/convex/
├── schema.ts                  # Database schema (~60 tables, 1569 lines)
├── http.ts                    # HTTP router (webhooks, voice, eval, Gmail proxy)
├── crons.ts                   # Scheduled tasks
├── constants.ts               # Shared constants
│
├── actions/
│   └── processMessage.ts      # Context extraction engine (intent, entities, language)
│
├── system/                    # Core business logic (59 files)
│   ├── messageProcessor.ts    # Main message orchestration pipeline
│   ├── incomingMessage.ts     # Inbound message routing
│   ├── channelMessaging.ts    # Outbound message sending
│   ├── channelSwitching.ts    # Cross-channel routing
│   ├── conversations.ts       # Conversation management
│   ├── workflows.ts           # Workflow management
│   ├── workflowTriggers.ts    # Workflow execution engine
│   ├── outboundCalls.ts       # Outbound dialing
│   ├── outboundCampaigns.ts   # Campaign management
│   ├── voiceTurn.ts           # Voice turn orchestration
│   ├── voiceTools.ts          # Voice tool implementations
│   ├── tokenTracking.ts       # Token usage logging
│   ├── creditAlerts.ts        # Credit threshold alerting
│   ├── ai/                    # AI subsystem
│   │   ├── selectiveRag.ts    # RAG decision + retrieval
│   │   ├── rag.ts             # RAG utilities
│   │   ├── resolveChatModel.ts # Model routing
│   │   └── agents/            # Agent implementations
│   └── ...
│
├── lib/                       # Reusable libraries (38 files)
│   ├── auth.ts                # Auth helpers
│   ├── policyEngine.ts        # Policy evaluation
│   ├── workflowMatcher.ts     # Workflow matching
│   ├── channelAdapter.ts      # Base channel adapter
│   ├── customApiToolExecutor.ts # HTTP tool executor
│   ├── mcpClient.ts           # MCP client
│   ├── secrets.ts             # AWS Secrets Manager
│   ├── actionExecutor/        # Workflow execution engine
│   ├── evaluationEngine/      # Evaluation logic
│   └── ...
│
├── evaluation/                # Evaluation framework (7 files)
│   ├── events.ts              # Event logging
│   ├── feedback.ts            # User feedback
│   ├── reports.ts             # Quality reports
│   ├── alerts.ts              # Threshold alerting
│   ├── llmJudge.ts            # LLM-as-Judge evaluation
│   └── retraining.ts          # Model retraining export
│
├── observability/             # Tracing, metrics, logging
│   ├── tracer.ts              # Core tracer engine
│   ├── llmMetrics.ts          # LLM metrics tracker
│   ├── toolTracker.ts         # Tool call tracker
│   ├── trace.ts               # Trace/Step/Summary types
│   └── storage/               # Pluggable trace storage
│
├── public/                    # Client-facing functions (44 files)
├── private/                   # Internal functions (22 files)
├── admin/                     # Platform admin functions
└── utils/                     # Utility functions
```

### HTTP Endpoints

| Endpoint | Method | Purpose |
|----------|--------|---------|
| `/clerk-webhook` | POST | Clerk events (subscription, user, session sync) |
| `/webhooks/meta` | GET/POST | Instagram + Messenger webhooks |
| `/webhooks/whatsapp` | GET/POST | WhatsApp Cloud API webhooks |
| `/webhooks/telegram` | POST | Telegram Bot webhook |
| `/voice/session` | POST | Voice call bootstrap (persona, tools, guards) |
| `/voice/tools/*` | POST | In-call tools (KB, data, workflow, escalate) |
| `/voice/persist` | POST | Idempotent transcript persistence |
| `/voice/turn` | POST | Full voice orchestration pipeline |
| `/voice/llm` | POST | Voice LLM streaming proxy |
| `/evaluation/events/log` | POST | Evaluation event ingestion |
| `/evaluation/feedback/log` | POST | User feedback ingestion |
| `/evaluation/reports/generate` | POST | Trigger report generation |
| `/gmail/send` | POST | Gmail send proxy |

### Cron Jobs

| Schedule | Job | Purpose |
|----------|-----|---------|
| Daily 2am UTC | `clean-old-debug-logs` | Delete debug logs > 7 days |
| Weekly Sun 3am | `clean-old-logs` | Delete all logs > 30 days |
| Daily midnight | `generate-daily-evaluation-report` | Auto-generate evaluation reports per org |
| Every 24h | `generate-daily-billing-rollups` | Recompute billing rollups |
| Daily 6am | `check-credit-alerts` | Safety-net credit threshold checks |
| Every 5 min | `process-due-callbacks` | Requeue voicemail callback attempts |

---

## 6. Database Schema

The schema defines **~60 tables** organized by domain:

### Platform & Auth

| Table | Purpose |
|-------|---------|
| `subscriptions` | Clerk billing subscription state |
| `platformOrgState` | Platform-operator state per tenant (suspended, feature flags) |
| `platformFeatureFlags` | Global feature-flag / canary registry |
| `users` | Basic user table |
| `userPreferences` | Per-user dashboard preferences |

### Agent System

| Table | Purpose |
|-------|---------|
| `agents` | First-class per-org agents (multi-agent with persona, LLM config, tool allowlist, channel routing) |
| `widgetSettings` | Legacy per-org widget/agent config |

### Conversations & Messages

| Table | Purpose |
|-------|---------|
| `conversations` | Chat conversation container (status, channel, metadata) |
| `contactSessions` | Anonymous visitor sessions |
| `messages` | Individual messages (intent, entities, language, metadata) |
| `conversationContexts` | Accumulated context per conversation |
| `pendingConfirmations` | Human-in-the-loop approval gates |

### Tools & Integrations

| Table | Purpose |
|-------|---------|
| `customApiTools` | User-defined HTTP API tools |
| `convexTools` | Built-in Convex DB tools |
| `marketplaceTools` | Shared tool marketplace catalog |
| `installedTools` | Per-org installed marketplace tools |
| `toolReviews` | Marketplace tool reviews |
| `toolExecutionLogs` | Tool execution audit trail |
| `llmConfigurations` | Per-org selectable LLM configs |
| `phoneNumbers` | Phone number assignments |
| `socialChannels` | Social channel connections |
| `orgOAuthTokens` | OAuth tokens for connectors |
| `plugins` | Service plugin configs |

### Workflows

| Table | Purpose |
|-------|---------|
| `agentWorkflows` | Workflow definitions (v1+v2 steps) |
| `workflowExecutions` | Running workflow state |
| `workflowVersions` | Versioned workflow snapshots (draft/published/archived) |
| `workflowTemplates` | Reusable workflow templates |

### Outbound Campaigns & Voice

| Table | Purpose |
|-------|---------|
| `outboundCampaigns` | Outbound calling campaigns |
| `outboundCallAttempts` | Per-contact dial state machine |
| `outboundCallOutcomes` | Post-call disposition |
| `doNotCallList` | DNC registry |
| `callRecordings` | Call recording metadata |
| `voiceMigrationConfigs` | Voice migration policy per org |
| `voiceKpiBaselines` | Baseline metrics per queue/language |
| `voiceAbExperiments` | A/B experiment definitions |
| `voiceTurnEvents` | Per-turn idempotent persistence ledger |
| `voiceCallSignals` | IVR/voicemail detection signals |

### Knowledge Base

| Table | Purpose |
|-------|---------|
| `fileChunks` | Uploaded file chunks for RAG (per-chunk enable toggle) |

### Evaluation & Observability

| Table | Purpose |
|-------|---------|
| `system_events` | Per-message execution events |
| `feedback` | User ratings and corrections |
| `evaluation_reports` | Generated quality reports |
| `evaluation_alerts` | In-app evaluation alerts |
| `executionLogs` | Workflow execution audit trail |
| `dailyOrgInsights` | Aggregated daily metrics |

### Billing & Quotas

| Table | Purpose |
|-------|---------|
| `orgQuotas` | Per-org usage quotas (credits, tokens, tool calls) |
| `tokenUsageLogs` | Token usage audit trail |
| `creditAlerts` | Credit threshold alerts |
| `usageEvents` | Billing metric events |
| `billingRollups` | Monthly billing aggregates |

### Other

| Table | Purpose |
|-------|---------|
| `businessRecords` | Generic business entity store |
| `businessRecordConfig` | Schema for business record types |
| `approvalPolicies` | Org-configurable approval rules |
| `onboardingTemplates` | Pre-built onboarding configs |
| `onboardingProgress` | Per-org onboarding state |
| `sandboxSuites` | Test suites for agent sandbox |
| `sandboxRuns` | Sandbox test run results |
| `orgSecrets` | Dev-mode fallback secret store |
| `airbyteConnections` | Airbyte sync pipelines |
| `cdataConnections` | CData connector sync state |

---

## 7. AI & LLM System

### Multi-Provider LLM Manager

The `LLMProviderManager` singleton supports 5 providers:

| Provider | Models | Usage |
|----------|--------|-------|
| **Groq** | Llama 3.3 70B, Llama 3.1 8B | Primary default; intent classification, KB interpretation, tool selection |
| **Mistral** | mistral-large-latest, mistral-small-latest | French/Arabic default; voice campaign qualification |
| **OpenAI** | GPT-4o, GPT-4o-mini | Configurable per-org; LLM-as-Judge evaluation |
| **Anthropic** | Claude Sonnet, Haiku | Configurable per-org |
| **Google** | Gemini Flash, Pro | Configurable per-org; voice agent fallback |

### Model Resolution Flow

```
1. Check active agent's llmConfigId → specific model row
2. Fall back to org-default llmConfigurations row
3. For French/Arabic without org model → Mistral (platform default)
4. Final fallback → Groq Llama 3.3 70B
```

### Agent System

- **`@convex-dev/agent`**: Core agent component managing thread history, message persistence
- **Multi-agent per organization**: Each agent has distinct persona, LLM config, KB scope, tool allowlist, channel routing
- **Persona presets**: General, Support, Sales, Receptionist, Technical
- **Tone options**: Friendly, Formal, Concise

### Context Extraction Engine

`actions/processMessage.ts` runs the Groq Llama 3.1 8B model for:
- Intent classification (name + category)
- Entity extraction (customer name, email, dates, order numbers, etc.)
- Language detection (fr, en, ar, etc.)
- Missing entity identification

### Message Processing Pipeline

```
Inbound Message (any channel)
    │
    ▼
[0] Suspension Guard → [0a] Resolve Active Agent → [0b] Resolve Chat Model
    │
    ▼
[1] Context Extraction (intent, entities, language)
    │
    ▼
[1a] Load Conversation History
[1b] Escalation Intercept (keyword + org rule classification)
    │
    ▼
[2] Workflow Matching (4-tier: pending → exact → all-keywords → any-keyword)
    │
    ▼
[3] If NO workflow:
    - Selective RAG decision (KB vs tool vs skip)
    - KB search (Voyage vector → TF-IDF fallback)
    - Direct tool call (Groq tool-calling → customApiToolExecutor)
    - MCP/CData fallback
    - Generic LLM reply (formulateReply)
    │
    ▼
[4] If workflow matched:
    - Pre-execution guard (business rules)
    - Start workflow execution
    - Execute steps (engine loop)
    - Formulate reply
    - Save to thread
```

---

## 8. Knowledge Base (RAG)

### Architecture

```
User Query
    │
    ▼
[1] decideRetrieval() — Rules or LLM router
    │  ├─ conversational → skip KB
    │  ├─ live-data → route to tools
    │  └─ informational → KB-first
    │
    ▼
[2] Vector Search (Voyage AI voyage-3.5-lite, 1024d)
    │  ├─ @convex-dev/rag per-org namespace
    │  ├─ Threshold: 0.7 (configurable via KB_VECTOR_SCORE_THRESHOLD)
    │  └─ Optional: Voyage reranker (rerank-2-lite)
    │
    ▼
[3] Source + Chunk Selection
    │  ├─ Rank documents by best chunk score
    │  ├─ Admit runner-up within SOURCE_SCORE_MARGIN (0.05)
    │  ├─ Cap: 2 sources, 4 chunks
    │  └─ Format grouped by source
    │
    ▼
[4] Keyword Fallback (TF-IDF localSearch)
    │  └─ When vector search misses
    │
    ▼
[5] LLM Interpretation (Groq Llama 3.1 8B)
    │  └─ Generates natural language answer from retrieved context
```

### Features

- File upload and chunking with keyword extraction
- Per-chunk enable/disable curation toggle
- Multilingual support (French, English, Arabic, Derja)
- Optional reranking for cross-lingual precision
- Automatic language detection for reply formatting

---

## 9. Workflow Engine

### Workflow Definition (v2)

```typescript
{
  name: string;
  triggerType: "conversation" | "manual" | "scheduled" | "webhook";
  steps: Array<{
    stepId: string;           // Stable identifier
    stepType: "agent" | "tool" | "condition" | "approval" | "transform";
    nextStepId?: string;      // Explicit routing (overrides sequential)
    branches?: Array<{        // For condition steps
      condition: { field, operator, value };
      targetStepId: string;
    }>;
    retryPolicy?: {           // Exponential backoff
      maxRetries: number;
      backoffMs: number;
    };
    approvalConfig?: {        // Human-in-the-loop
      prompt: string;
      timeoutMs: number;
    };
    transformConfig?: {       // Reshape previous step output
      mappings: Array<{ from, to }>;
    };
    compensation?: {          // Rollback action
      toolId: string;
      parameters: Record;
    };
  }>;
}
```

### Execution Engine

- **Condition steps**: Evaluate branches top-down (entity, step-output, step-status conditions)
- **Approval steps**: Pause for human approval via `humanInTheLoop.createApprovalRequest`
- **Transform steps**: Reshape previous step output into row arrays
- **Virtual tools**: `SendMessage` (localized), `CollectEntity` (ask user), `ConvexDB` (direct DB ops)
- **External tools**: Delegates to `ActionExecutor` for custom API tools
- **Placeholder resolution**: `{{entities.X}}`, `{{steps.STEP_ID.output}}`, `{{steps.STEP_ID.result}}`
- **Version locking**: Running executions lock to published version; mid-flight edits never affect active conversations

### Workflow Matching (4-Tier)

1. **Pending workflow name** — exact match (conversation was asking for more info mid-workflow)
2. **Exact name match** — normalized comparison
3. **All keywords match** — transaction/troubleshooting intents only
4. **Any keyword match** — actionable intents only

---

## 10. Channel System

### Supported Channels (6)

| Channel | Adapter | Features |
|---------|---------|----------|
| **Widget** | Built-in | Embeddable script + iframe, infinite-scroll chat |
| **Phone** | `phoneChannelAdapter.ts` | Vapi/Telnyx/SIP, voice call handling |
| **WhatsApp** | `whatsappChannelAdapter.ts` | Cloud API, template messages, 24h session window |
| **Instagram** | `instagramChannelAdapter.ts` | Meta Graph API, DM handling |
| **Messenger** | `messengerChannelAdapter.ts` | Meta Messenger Platform, page-level webhooks |
| **Telegram** | `telegramChannelAdapter.ts` | Bot API, secret_token authentication |

### Channel Adapter Interface

```typescript
interface ChannelAdapter {
  receiveMessage(event: any): Promise<ChannelMessage>;
  sendMessage(recipientId: string, content: MessageContent): Promise<void>;
  downloadMedia(mediaId: string): Promise<Buffer>;
  sendTypingIndicator(recipientId: string): Promise<void>;
  validateWebhook(signature: string, body: string): boolean;
  getChannelType(): ChannelType;
}
```

### Incoming Message Flow

All channel webhooks funnel through `incomingMessage.processIncomingMessage`:
1. Resolve the social channel record
2. Create/update contact session
3. Create/get conversation
4. Call `messageProcessor.processMessage` for the core pipeline

### Agent Routing

- Agents claim specific channels via `agents.channels[]`
- Channel-specific agent takes precedence over org default
- Cross-channel switching supported mid-conversation

---

## 11. Voice & Telephony

### Voice Stack Evolution

```
Vapi (retired) → Telnyx Bridge (disaster recovery) → LiveKit + Realtime S2S (production)
```

### LiveKit Voice Agent (`services/voice-livekit/`)

**Engine selection** (env `VOICE_REALTIME_PROVIDER`):

| Engine | Stack | Latency | Status |
|--------|-------|---------|--------|
| `deepgram` (default) | Deepgram Nova-3 STT + Groq LLM + Deepgram Aura-2 TTS + Silero VAD | Low | Production |
| `azure` | Azure OpenAI `gpt-realtime` speech-to-speech (Sweden Central) | Lowest | Production target |
| `openai` | OpenAI Realtime API | Low | Dev fallback |
| `gemini` | Google Gemini realtime | Low | Dev fallback |

### Voice Tools (In-Call)

| Tool | Purpose |
|------|---------|
| `search_knowledge` | KB search during call |
| `lookup_data` | Live data lookup via CData/custom tools |
| `run_workflow` | Execute workflow during call |
| `escalate_to_human` | Transfer to human agent |
| `record_campaign_answer` | Record qualification answers |

### Voice Features

- Session bootstrap with guards, persona, tool manifest
- Idempotent turn persistence
- Barge-in detection, IVR navigation (DTMF)
- Call recordings with PII redaction and retention
- Voice migration framework with A/B experiments, KPI baselines, rollback thresholds
- Campaign chaining (dial next contact when current call ends)

---

## 12. Custom API Tools & Marketplace

### Custom API Tools

Configurable HTTP tools with:
- Methods: GET, POST, PUT, PATCH, DELETE
- Auth types: API key, bearer, OAuth2, basic, Supabase, Google service account
- Timeout and retry configuration
- Human confirmation gate per tool
- Internal-only tools (hidden from customer chat agent)
- Full execution logging with status codes and timing

### Tool Marketplace

- 150+ community-submitted connectors across categories:
  - Payments, Email, Communication, CRM, Analytics, Storage, AI
  - ERP, Database, Marketing, E-Commerce, Finance, HR, Project Management
- Approval workflow (pending/approved/rejected)
- Install, review, and rating system
- CData-powered: SQL-over-REST for 100+ SaaS connectors

### ConvexDB Virtual Tools

Direct database operations via agent:
- `findOne`, `findMany`, `insert`, `patch`, `delete`
- Configurable filters with parameterized queries

---

## 13. Outbound Campaigns

### Campaign Types

| Type | Description |
|------|-------------|
| **Qualification** | Structured questions with branching (followUpIfYes/followUpIfNo) |
| **Opportunity** | Pitch + booking flow |

### Features

- Contact lists (inline or sourced from CData)
- Calling windows with timezone-aware scheduling
- Retry policies with backoff
- Result destinations: webhook, email, none
- Per-contact attempt tracking with voicemail detection
- Callback scheduling
- Campaign chaining (dial next contact when current call ends)
- Do-not-call list compliance

---

## 14. Evaluation & Observability

### Evaluation System

#### Event Logging

Every processed message logs a `system_events` row with:
- Predicted intent and entities
- Execution result (success/failure)
- Agent and model used
- Channel and language

#### Quality Reports

Automated daily reports with:
- **Quality score** (0-1 composite)
- **Summary metrics**: Intent accuracy, entity extraction accuracy, task success rate, average response time, fallback rate, retry rate, human takeover rate
- **Detected failure types**: Categorized with severity (low/medium/high/critical)
- **Improvement suggestions**: Prioritized with expected impact
- **Flagged conversations**: For human review

#### LLM-as-Judge

OpenAI GPT-4o evaluates conversation quality on 5 dimensions:
- Relevance, Accuracy, Helpfulness, Clarity, Professionalism (each 0-100)
- Detailed reasoning, strengths, weaknesses, improvements

#### Feedback Collection

- User ratings (thumbs up/down)
- Free-text comments
- Corrections (suggested fix for wrong answer)
- Resolution status tracking

#### Retraining Export

Export labeled datasets from:
- Failed events
- User corrections
- Low-rated feedback

### Observability System

#### Architecture

```
Tracer (per request)
├── TraceStep: quota_check
├── TraceStep: context_extraction
│   └── childStep: intent_classification
├── TraceStep: workflow_selection
├── TraceStep: knowledge_base_search
│   ├── childStep: rag_query
│   └── childStep: llm_call
├── TraceStep: workflow_execution
│   ├── childStep: tool_call
│   └── childStep: tool_call
└── TraceStep: response_composition
    └── childStep: llm_call
```

#### Core Types

```typescript
type Trace = {
  traceId: string;
  sessionId, conversationId, organizationId;
  startTime, endTime;
  status: "success" | "error" | "partial";
  steps: TraceStep[];
  metadata: { model, channel, language, intent, workflowName };
};

type TraceSummary = {
  totalLatencyMs, totalTokens, totalCostUsd, stepCount, errorCount;
  breakdown: {
    llmMs, toolMs, ragMs, orchestrationMs,
    intentDetectionMs, workflowSelectionMs, responseCompositionMs
  };
};
```

#### Modules

| Module | Purpose |
|--------|---------|
| `tracer.ts` | Core tracer engine — lifecycle, step handles, auto-close, persist |
| `llmMetrics.ts` | LLM call metrics (tokens, cost, latency) |
| `toolTracker.ts` | Tool execution metrics |
| `formulateReplyTraced.ts` | Traced wrapper around formulateReply |
| `toolExecutorTraced.ts` | Traced tool execution wrapper |
| `traceQuery.ts` | Query helpers for trace data |
| `storage/InMemoryTraceStorage.ts` | Pluggable in-memory storage |

#### Built-in Cost Estimation

Supports pricing for: Groq (Llama 8B/70B), OpenAI (GPT-4o/mini/turbo), Anthropic (Claude Sonnet/Haiku), Google (Gemini Flash/Pro), Mistral (Large/Small), Voyage (3.5-lite), Cohere (embed-v4.0).

#### 5-Scenario Test Suite

| Scenario | Description |
|----------|-------------|
| 1. Simple Greeting | Conversational shortcut — no LLM, no workflow |
| 2. KB Search | TF-IDF + LLM interpretation with child steps |
| 3. Workflow Trigger | Refund request with tool calls (Stripe + email) |
| 4. Escalation | Arabic WhatsApp, human agent handoff |
| 5. Complex Workflow | Multi-step with error + retry + tool calls |

---

## 15. Security & Governance

### Two-Layer Policy Engine

**Layer 1 — System Checks** (platform-enforced):
- Session authentication (existence + expiry)
- Tool permission checks
- Ethics/safety: destructive intent blocking, sensitive tool gating
- High-risk detection: urgency ≥ 0.8 + negative sentiment
- PII exfiltration prevention

**Layer 2 — Org-Configurable Rules** (`approvalPolicies` table):
- Field-based conditions (intent, urgencyScore, sentiment, toolName, entities.*)
- Operators: eq, neq, gt, lt, gte, lte, contains
- Actions: auto_approve, require_approval, block
- Priority-ordered evaluation

### Secrets Management

- **Production**: AWS Secrets Manager with full CRUD + rotation
- **Dev fallback**: `orgSecrets` table (Convex-native)
- Secret types: bearer, api_key, oauth2, basic, supabase, google_service_account

### Platform Admin

- Multi-tenant organization management
- Feature flags with deterministic canary bucketing (percentage-based rollout)
- Per-org feature flag overrides
- Organization suspension capability
- Platform admin authentication (allowlist + JWT claim)

### Compliance

- Do-not-call list for outbound campaigns
- Call recording PII redaction
- GDPR-oriented data retention policies
- Execution logging for audit trails

---

## 16. Billing & Quotas

### Credit System

- Per-org credit allocation with consumption tracking
- Token-based billing (per 1K tokens, ceil-based)
- Tool call counting with limits
- Usage events: voice minutes, transcription minutes, messages

### Quota Management

```typescript
orgQuotas = {
  toolCallsUsed, toolCallsLimit,
  creditsRemaining, creditLimit,
  tokensUsed, tokenLimit,
  creditAlertsSent: { twenty: boolean, ten: boolean, exhausted: boolean }
}
```

### Alert Levels

| Level | Threshold | Action |
|-------|-----------|--------|
| Warning | 20% consumed | Email + in-app alert |
| Critical | 10% remaining | Email + in-app alert |
| Exhausted | 0% remaining | Block all AI processing |

### Billing Rollups

- Daily aggregation of usage events
- Monthly billing rollups with cost estimates
- Overage detection and alerts

---

## 17. Frontend — Web Dashboard

### Pages

| Route | Page | Description |
|-------|------|-------------|
| `/` | Smart redirect | Checks onboarding → conversations or onboarding |
| `/conversations` | Conversations inbox | List + detail with messages |
| `/agent-builder` | Agent Builder | Multi-agent config: persona, tools, responses |
| `/files` | Knowledge Base | File upload/delete |
| `/campaigns` | Campaigns | Outbound campaign management |
| `/calendar` | Calendar | Appointment scheduling |
| `/outbound-outcomes` | Call Outcomes | Analytics dashboard |
| `/call-history` | Call History | Call log |
| `/evaluation` | Evaluation | Quality reports + LLM Judge |
| `/customization` | Widget Designer | Live preview of widget appearance |
| `/integrations` | Integrations Hub | API tools, marketplace, LLM, social, phone |
| `/integrations/marketplace` | Marketplace | 150+ connectors |
| `/integrations/agent-workflows` | Workflow Builder | Visual drag-and-drop workflow designer |
| `/usage` | Usage & Credits | Analytics with charts |
| `/billing` | Billing | Plans & pricing |
| `/tool-testing` | Tool Sandbox | Test tools in isolation |
| `/admin` | Platform Admin | Feature flags, org management |
| `/onboarding` | Onboarding | 3-step wizard |

### Key Modules

| Module | Purpose |
|--------|---------|
| `agent-builder/` | Multi-agent CRUD, persona config, tool allowlist, response rules |
| `integrations/` | Custom API tools, marketplace, LLM configs, social channels |
| `campaigns/` | Campaign list, create, detail |
| `customization/` | Widget Designer with live preview |
| `marketplace/` | 150+ connectors, CData integration, OAuth flows |
| `i18n/` | Internationalization (EN, FR, AR) |

---

## 18. Frontend — Chat Widget

### State Machine

```
loading → (validates org, session, settings) → auth OR selection
selection → chat (start conversation)
selection → inbox (view past conversations)
auth → selection (after creating contact session)
error (terminal)
```

### Screens

| Screen | Description |
|--------|-------------|
| `loading` | 3-step initialization with progress |
| `auth` | Collects name + email, creates contact session |
| `selection` | Hero header with "Start chat" |
| `chat` | Full chat UI with infinite-scroll, suggestions, workflow banners |
| `inbox` | Paginated conversation list with status pills |
| `error` | Error display |

### Design Tokens

CSS custom properties: `--widget-accent`, `--widget-radius`, `--widget-title`

---

## 19. Frontend — Embed Script

Lightweight `<script>` tag that creates:
- Floating action button (60x60px, circular)
- Container (400x600px) with iframe loading the widget
- Global API: `window.EchoWidget.init()`, `.show()`, `.hide()`, `.destroy()`

---

## 20. Services

### Voice LiveKit (`services/voice-livekit/`)

Python LiveKit Agents worker:
- Bootstraps via `POST /voice/session`
- Exposes 5 Convex function tools
- Persists turns via `POST /voice/persist`
- Supports Deepgram, Azure, OpenAI, Gemini engines
- EU-resident via Deepgram EU endpoint

### Voice Bridge (`services/voice-bridge/`)

Python Telnyx media bridge (disaster recovery):
- WebSocket server: Telnyx ↔ Voxtral STT ↔ Mistral/TTS
- Barge-in, half-duplex echo guard, filler phrases
- Dockerized for Fly.io deployment (Paris)

### WhatsApp Baileys (`services/whatsapp-baileys/`)

Node.js unofficial WhatsApp bridge (demo-grade):
- QR code linking
- Forwards messages to Convex webhook
- Against WhatsApp ToS — production uses official Cloud API

---

## 21. Integrations

### OAuth Providers

| Provider | Scopes | Purpose |
|----------|--------|---------|
| **Google** | Drive, Gmail, Sheets, Calendar, Analytics, BigQuery, Contacts, YouTube | Data connectors |
| **GitHub** | repo, read:org, read:user | Code/data access |
| **Microsoft** | OneDrive, SharePoint, Teams, Office365, Exchange, Azure DevOps, Dynamics | Enterprise connectors |
| **Meta** | Instagram, Messenger, Pages | Social channel webhooks |

### CData Connect

SQL-over-REST for 100+ SaaS connectors:
- Auto-discovers catalogs, tables, columns
- Creates custom API tools per catalog
- Shared platform CData account for marketplace

### Airbyte

Batch data sync:
- Gmail connector → Supabase
- Custom source connectors
- Postgres views per org

### Email

- **Resend**: Transactional emails (confirmation, credit alerts, evaluation alerts)
- **Gmail OAuth**: Send proxy via RFC 2822 encoding
- **Nodemailer**: SMTP fallback for escalation emails

---

## 22. Deployment & CI/CD

| Component | Platform | Region |
|-----------|----------|--------|
| Convex Backend | Convex Cloud | EU Ireland |
| Web Dashboard | Vercel | Paris (`cdg1`) |
| Widget | Vercel | Paris (`cdg1`) |
| Voice Bridge | Fly.io | Paris |
| Voice LiveKit | LiveKit Cloud | EU |

### CI/CD Pipeline (`.gitlab-ci.yml`)

- Lint (ESLint)
- Type check (TypeScript)
- Secret scanning
- Deploy (Convex + Vercel)

### Error Monitoring

**Sentry** (`@sentry/nextjs`) for frontend error tracking.

---

## 23. Environment Variables

### Required

| Variable | Purpose |
|----------|---------|
| `GROQ_API_KEY` | Primary LLM (Llama 3.3 70B) |
| `CLERK_SECRET_KEY` | Auth — user/org management |
| `CLERK_WEBHOOK_SECRET` | Clerk webhook verification |
| `NEXT_PUBLIC_CLERK_PUBLISHABLE_KEY` | Frontend auth |
| `NEXT_PUBLIC_CONVEX_URL` | Client-facing Convex URL |

### Strongly Recommended

| Variable | Purpose |
|----------|---------|
| `MISTRAL_API_KEY` | French & Arabic chat replies |
| `RESEND_API_KEY` | Transactional emails |
| `VOYAGE_API_KEY` | Knowledge base vector embeddings |

### Optional (by feature)

| Category | Variables |
|----------|-----------|
| **Voice** | `TELNYX_API_KEY`, `LIVEKIT_URL`, `LIVEKIT_API_KEY`, `LIVEKIT_API_SECRET`, `DEEPGRAM_API_KEY`, `VOICE_TOOL_SECRET` |
| **OAuth** | `GOOGLE_CLIENT_ID/SECRET`, `GITHUB_CLIENT_ID/SECRET`, `MICROSOFT_CLIENT_ID/SECRET`, `OAUTH_CALLBACK_TOKEN` |
| **Data** | `CDATA_EMAIL/PAT`, `SUPABASE_SERVICE_ROLE_KEY`, `AIRBYTE_CLIENT_ID/SECRET/WORKSPACE_ID` |
| **Social** | `META_APP_ID/SECRET`, `WHATSAPP_VERIFY_TOKEN`, `WHATSAPP_APP_SECRET` |
| **Email** | `GMAIL_USER`, `GMAIL_APP_PASSWORD`, `ALERT_EMAIL`, `SLACK_WEBHOOK_URL` |
| **Eval** | `OPENAI_API_KEY` (LLM Judge), `COHERE_API_KEY` (3-way eval) |
| **AWS** | `AWS_REGION`, `AWS_ACCESS_KEY_ID`, `AWS_SECRET_ACCESS_KEY` |

---

## 24. Running the Project

### Prerequisites

- Node.js ≥ 20
- pnpm 10.4.1

### Install Dependencies

```bash
pnpm install
```

### Start Development

```bash
pnpm dev
```

This starts all 3 apps via Turborepo:

| App | URL |
|-----|-----|
| Web Dashboard | http://localhost:3000 |
| Chat Widget | http://localhost:3001 |
| Embed Script | http://localhost:3002 |
| Convex Backend | http://127.0.0.1:3210 |

### Environment Setup

Create `packages/backend/.env.local` with required keys (see Section 23).

---

## 25. Testing

### Unit Tests (Vitest)

```bash
cd packages/backend
pnpm test
```

**14 test suites, 345 tests:**

| Suite | Tests | Coverage |
|-------|-------|----------|
| Evaluation engine | 7 | Report generation, quality scores, failure detection |
| Token tracking | 33 | Credit deduction, thresholds, quota checks, alerts |
| Full suite | 69 | runAgent, formulateReply, messageProcessor, tools, workflows |
| Selective RAG | 14 | Routing rules, LLM router, source/chunk selection |
| Custom API tool | 17 | Auth types, retry, confirmation, validation |
| Schema validation | 33 | All multi-channel tables |
| Conversations | 22 | Channel creation, session linking |
| Channel switching | 23 | Cross-channel routing, context preservation |
| Channel adapter | 44 | Message handling, webhook validation, HMAC |
| Phone adapter | 21 | Transcript parsing, events |
| Channel health | 16 | Health monitoring, alerts |
| Secrets | 29 | AWS Secrets Manager CRUD |
| Workflows | 16 | Step sorting, execution, validation |
| Semantic context | 1 | Output validation |

### Observability Integration Test

```bash
npx tsx eval/test-observability.ts
```

5 scenarios testing the full tracing pipeline.

### RAG Evaluation

```bash
# TF-IDF vs Voyage
npx tsx eval/run-comparison.ts

# TF-IDF vs Voyage vs Cohere
npx tsx eval/run-comparison-3way.ts
```

---

*This document describes the complete CoreDeskAI platform — a multi-tenant, multi-channel AI customer support platform with workflow automation, voice integration, knowledge base RAG, and comprehensive evaluation/observability systems.*
