# TF-IDF vs Voyage AI vs Cohere — Embedding Comparison

## Overview

Comparative evaluation of three retrieval methods on a synthetic Tunisian SaaS customer support knowledge base (25 French chunks, 145 queries across 4 language variants).

---

## Results Summary

| Language | TF-IDF | Voyage | Cohere | Winner |
|----------|--------|--------|--------|--------|
| French | 72% | 100% | 96% | Voyage |
| MSA (Modern Standard Arabic) | 8% | 96% | 100% | Cohere |
| Derja (Arabic script) | 8% | 72% | 84% | Cohere |
| Derja (Arabizi/Latin script) | 24% | 36% | 24% | Voyage |

### Key Findings

- **TF-IDF** fails completely on cross-language retrieval (0% on Arabic queries against French KB)
- **Voyage** excels at French (100%) and Arabizi (36%) but weaker on Arabic script
- **Cohere** wins on MSA (100%) and Derja Arabic (84%) — best for Tunisian Arabic speakers
- **Derja Arabizi** underperforms across all methods (max 36%) — out-of-distribution for embedding models

### Recommendation

For a Tunisian business: **Cohere embed-v4.0** (best on MSA + Derja Arabic, the primary customer languages).

---

## Files Created

### Evaluation Dataset

| File | Description |
|------|-------------|
| `eval/rag-eval-chunks.json` | 25 synthetic French KB chunks covering 14 topics (refund, billing, integrations, security, etc.) |
| `eval/rag-eval-set.json` | 145 queries: 100 valid (25 per language × 4 languages) + 45 distractors (off-topic) |

### Evaluation Scripts

| File | Description |
|------|-------------|
| `eval/run-rag-eval.ts` | TF-IDF baseline evaluation |
| `eval/analyze-results.ts` | Results analysis and summary generation |
| `eval/run-comparison.ts` | Voyage vs TF-IDF head-to-head comparison |
| `eval/run-comparison-3way.ts` | 3-way comparison (TF-IDF vs Voyage vs Cohere) |

### Results

| File | Description |
|------|-------------|
| `eval/results.csv` | TF-IDF raw results (145 rows) |
| `eval/comparison-results.csv` | Voyage vs TF-IDF results |
| `eval/comparison-3way-results.csv` | 3-way raw results |
| `eval/comparison-summary.md` | Voyage vs TF-IDF summary |
| `eval/comparison-3way-summary.md` | 3-way comparison summary |

---

## Files Modified

| File | Change |
|------|--------|
| `packages/backend/convex/system/ai/tools/search.ts` | Switched from TF-IDF (`localSearch`) to Voyage vector search (`voyageEmbedding`) with TF-IDF fallback on error |

---

## Architecture Notes

### What Changed in Production

```
Before:  query → localSearch (TF-IDF) → top chunks → LLM
After:   query → voyageEmbed (vector) → cosine similarity → top chunks → LLM
                                        ↓ (on error)
                               localSearch (TF-IDF fallback)
```

### What Stayed the Same

- `rag.ts` — RAG instance still uses Voyage voyage-3.5-lite (1024d) for document ingestion
- `voyageEmbedding.ts` — Custom Voyage adapter unchanged
- `fileChunks.ts` — Chunk storage unchanged
- `localSearch.ts` — Still used as fallback

### Cohere Status

- **Eval only**: Cohere embed-v4.0 was tested in `run-comparison-3way.ts` but NOT integrated into production
- **Production search**: Uses Voyage (switched from TF-IDF)
- **Production RAG**: Uses Voyage for document embedding

---

## API Keys Used

| Provider | Key (masked) | Purpose |
|----------|-------------|---------|
| Voyage AI | `pa--Vtw...ljw` | Embedding queries + chunks in comparison |
| Cohere | `cohere_hu...H` | Embedding queries + chunks in comparison |
| Groq | `gsk_XFcl...` | LLM calls (context extraction, reply formulation) |

---

## Limitations

1. **Small KB**: Only 25 chunks — real production KB would be much larger
2. **French-only KB**: Tests cross-language retrieval (Arabic queries → French answers)
3. **Synthetic data**: Queries and chunks are AI-generated, not from real customer interactions
4. **Derja Arabizi inconsistency**: Transliteration varies (e.g., "kifach" vs "kif")
5. **No production traffic**: Eval uses controlled queries, not real customer messages
