# TF-IDF vs Voyage Vector Search — Comparison

## Retrieval Accuracy (top-1 matches expected)

| Language | TF-IDF | Voyage | Improvement |
|----------|--------|--------|-------------|
| French | 72.0% (18/25) | 92.0% (23/25) | +5 |
| MSA | 8.0% (2/25) | 80.0% (20/25) | +18 |
| Derja (Arabic) | 8.0% (2/25) | 60.0% (15/25) | +13 |
| Derja (Arabizi) | 24.0% (6/25) | 32.0% (8/25) | +2 |

## Score Distribution (expected chunk)

| Language | Method | Mean | Median | Min | Max |
|----------|--------|------|--------|-----|-----|
| French | TF-IDF | 0.5288 | 0.4370 | 0.1587 | 1.6349 |
| French | Voyage | 0.8414 | 0.8509 | 0.6877 | 0.9038 |
| MSA | TF-IDF | 0.0078 | 0.0000 | 0.0000 | 0.1273 |
| MSA | Voyage | 0.7378 | 0.7434 | 0.6020 | 0.8369 |
| Derja (Arabic) | TF-IDF | 0.0051 | 0.0000 | 0.0000 | 0.1273 |
| Derja (Arabic) | Voyage | 0.6663 | 0.6663 | 0.4682 | 0.7783 |
| Derja (Arabizi) | TF-IDF | 0.0801 | 0.0000 | 0.0000 | 0.5347 |
| Derja (Arabizi) | Voyage | 0.6158 | 0.6097 | 0.4739 | 0.7893 |

## False Negative Rate (expected_chunk_score < threshold)

For TF-IDF, no threshold exists (MIN_SCORE = 0.001). For Voyage, using 0.7.

| Language | TF-IDF FN Rate | Voyage FN Rate |
|----------|----------------|----------------|
| French | 0.0% (0/25) | 4.0% (1/25) |
| MSA | 92.0% (23/25) | 28.0% (7/25) |
| Derja (Arabic) | 96.0% (24/25) | 64.0% (16/25) |
| Derja (Arabizi) | 64.0% (16/25) | 80.0% (20/25) |

## Arabizi vs Arabic Script

| Variant | TF-IDF | Voyage |
|---------|--------|--------|
| Derja (Arabic) | 8.0% | 60.0% |
| Derja (Arabizi) | 24.0% | 32.0% |

## Conclusion

**Recommendation: Switch from TF-IDF to Voyage vector search.**

The Voyage embeddings provide significant improvements for non-French queries while maintaining or improving French retrieval accuracy. This is the minimum viable fix for the multilingual retrieval problem.
