Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
998 changes: 858 additions & 140 deletions docs/eval/baseline-v1.6.json

Large diffs are not rendered by default.

43 changes: 22 additions & 21 deletions docs/eval/baseline-v1.6.md
Original file line number Diff line number Diff line change
Expand Up @@ -11,23 +11,23 @@ Generated by `npm run eval`. The numbers below are harness output — do not edi
| Retrieval | `dense` |
| Ranks | `candidateK=20, threshold=0.5` |
| Context width | `contextK=3` |
| Corpus | `eval/corpus` (13 documents) |
| Split | `all` (44 questions, 38 answerable) |
| Index size | 19 chunks |
| Corpus | `eval/corpus` (21 documents) |
| Split | `all` (63 questions, 53 answerable) |
| Index size | 53 chunks |

## Metrics

| Metric | Value |
| --- | --- |
| Recall@1 | 0.6579 |
| Recall@5 | 0.9211 |
| Recall@10 | 1.0000 |
| MRR | 0.8009 |
| nDCG@10 | 0.8476 |
| Hit rate@5 | 0.9211 |
| MAP@10 | 0.7965 |
| Context precision@3 | 0.3246 |
| Context recall@3 | 0.9211 |
| Recall@1 | 0.5943 |
| Recall@5 | 0.8774 |
| Recall@10 | 0.9198 |
| MRR | 0.7548 |
| nDCG@10 | 0.7804 |
| Hit rate@5 | 0.9057 |
| MAP@10 | 0.7280 |
| Context precision@3 | 0.3082 |
| Context recall@3 | 0.8160 |

### By query type

Expand All @@ -37,10 +37,11 @@ The type comes from `type` in `questions.jsonl`; untagged questions report as

| Type | Questions | Recall@5 | nDCG@10 | Hit rate@5 | MAP@10 |
| --- | --- | --- | --- | --- | --- |
| cross-lingual | 9 | 0.6667 | 0.5032 | 0.6667 | 0.3446 |
| exact | 6 | 1.0000 | 1.0000 | 1.0000 | 1.0000 |
| multi-hop | 2 | 1.0000 | 0.9599 | 1.0000 | 0.9167 |
| semantic | 18 | 1.0000 | 0.9312 | 1.0000 | 0.9074 |
| cross-lingual | 9 | 0.6667 | 0.4173 | 0.6667 | 0.2994 |
| exact | 8 | 1.0000 | 0.8663 | 1.0000 | 0.8229 |
| hard-negative | 7 | 1.0000 | 0.8758 | 1.0000 | 0.8333 |
| multi-hop | 5 | 0.7000 | 0.7218 | 1.0000 | 0.6350 |
| semantic | 21 | 0.9048 | 0.8542 | 0.9048 | 0.8238 |
| zh | 3 | 1.0000 | 1.0000 | 1.0000 | 1.0000 |

### Unanswerable questions
Expand All @@ -62,13 +63,13 @@ with a small second one means the threshold filters nothing and the window is al

| Metric | Value |
| --- | --- |
| Unanswerable questions | 6 |
| Retrieval abstained | 0.0000 (0/6) |
| Mean candidates passing the threshold | 19.00 |
| Unanswerable questions | 10 |
| Retrieval abstained | 0.0000 (0/10) |
| Mean candidates passing the threshold | 20.00 |
| Mean passages in the context window | 3.00 |

Timing is informational only and is **not** frozen: indexing 1583 ms, query
p50 13.38 ms, p95 16.66 ms on the
Timing is informational only and is **not** frozen: indexing 2796 ms, query
p50 13.67 ms, p95 16.34 ms on the
machine that produced this file. Timing and index size depend on hardware and on the
corpus, so they must never be the reason two runs differ.

Expand Down
102 changes: 51 additions & 51 deletions docs/eval/retrieval-v1.6.json
Original file line number Diff line number Diff line change
Expand Up @@ -15,79 +15,79 @@
"id": "dense",
"label": "dense (vector)",
"split": "validation",
"questions": 16,
"answerableCount": 13,
"questions": 24,
"answerableCount": 19,
"chunking": "1000/100",
"chunkCount": 19,
"recallAt1": 0.576923,
"recallAt5": 0.923077,
"recallAt10": 1,
"mrr": 0.778846,
"ndcgAt10": 0.827608,
"hitRateAt5": 0.923077,
"mapAt10": 0.766026,
"contextPrecision": 0.333333,
"contextRecall": 0.923077,
"latencyP95Ms": 16.02
"chunkCount": 53,
"recallAt1": 0.513158,
"recallAt5": 0.868421,
"recallAt10": 0.921053,
"mrr": 0.720175,
"ndcgAt10": 0.755611,
"hitRateAt5": 0.894737,
"mapAt10": 0.69386,
"contextPrecision": 0.315789,
"contextRecall": 0.802632,
"latencyP95Ms": 13.27
},
{
"id": "sparse",
"label": "sparse (BM25)",
"split": "validation",
"questions": 16,
"answerableCount": 13,
"questions": 24,
"answerableCount": 19,
"chunking": "1000/100",
"chunkCount": 19,
"recallAt1": 0.576923,
"recallAt5": 0.615385,
"recallAt10": 0.615385,
"mrr": 0.615385,
"ndcgAt10": 0.609209,
"hitRateAt5": 0.615385,
"mapAt10": 0.602564,
"contextPrecision": 0.230769,
"contextRecall": 0.615385,
"latencyP95Ms": 2.96
"chunkCount": 53,
"recallAt1": 0.460526,
"recallAt5": 0.671053,
"recallAt10": 0.684211,
"mrr": 0.600877,
"ndcgAt10": 0.610495,
"hitRateAt5": 0.684211,
"mapAt10": 0.576817,
"contextPrecision": 0.263158,
"contextRecall": 0.657895,
"latencyP95Ms": 2.26
},
{
"id": "hybrid",
"label": "hybrid (RRF of dense + BM25)",
"split": "validation",
"questions": 16,
"answerableCount": 13,
"questions": 24,
"answerableCount": 19,
"chunking": "1000/100",
"chunkCount": 19,
"recallAt1": 0.576923,
"recallAt5": 0.923077,
"recallAt10": 1,
"mrr": 0.778846,
"ndcgAt10": 0.827608,
"hitRateAt5": 0.923077,
"mapAt10": 0.766026,
"chunkCount": 53,
"recallAt1": 0.513158,
"recallAt5": 0.868421,
"recallAt10": 0.894737,
"mrr": 0.732456,
"ndcgAt10": 0.760265,
"hitRateAt5": 0.894737,
"mapAt10": 0.707018,
"contextPrecision": 0.333333,
"contextRecall": 0.923077,
"latencyP95Ms": 19.47
"contextRecall": 0.855263,
"latencyP95Ms": 13.61
}
],
"test": [
{
"id": "dense",
"label": "dense (vector)",
"split": "test",
"questions": 28,
"answerableCount": 25,
"questions": 39,
"answerableCount": 34,
"chunking": "1000/100",
"chunkCount": 19,
"recallAt1": 0.7,
"recallAt5": 0.92,
"recallAt10": 1,
"mrr": 0.812381,
"ndcgAt10": 0.858056,
"hitRateAt5": 0.92,
"mapAt10": 0.812381,
"contextPrecision": 0.32,
"contextRecall": 0.92,
"latencyP95Ms": 14.49
"chunkCount": 53,
"recallAt1": 0.639706,
"recallAt5": 0.882353,
"recallAt10": 0.919118,
"mrr": 0.774079,
"ndcgAt10": 0.794329,
"hitRateAt5": 0.911765,
"mapAt10": 0.747141,
"contextPrecision": 0.303922,
"contextRecall": 0.823529,
"latencyP95Ms": 12.89
}
]
}
8 changes: 4 additions & 4 deletions docs/eval/retrieval-v1.6.md
Original file line number Diff line number Diff line change
Expand Up @@ -16,15 +16,15 @@ to.

| Strategy | Recall@1 | Recall@5 | MRR | nDCG@10 | MAP@10 | Context P | n | p95 |
| --- | --- | --- | --- | --- | --- | --- | --- | --- |
| dense (vector) | 0.5769 | 0.9231 | 0.7788 | 0.8276 | 0.7660 | 0.3333 | 16 | 16.02 ms |
| sparse (BM25) | 0.5769 | 0.6154 | 0.6154 | 0.6092 | 0.6026 | 0.2308 | 16 | 2.96 ms |
| hybrid (RRF of dense + BM25) | 0.5769 | 0.9231 | 0.7788 | 0.8276 | 0.7660 | 0.3333 | 16 | 19.47 ms |
| dense (vector) | 0.5132 | 0.8684 | 0.7202 | 0.7556 | 0.6939 | 0.3158 | 24 | 13.27 ms |
| sparse (BM25) | 0.4605 | 0.6711 | 0.6009 | 0.6105 | 0.5768 | 0.2632 | 24 | 2.26 ms |
| hybrid (RRF of dense + BM25) | 0.5132 | 0.8684 | 0.7325 | 0.7603 | 0.7070 | 0.3333 | 24 | 13.61 ms |

## Test — reported, not selected

| Strategy | Recall@1 | Recall@5 | MRR | nDCG@10 | MAP@10 | Context P | n | p95 |
| --- | --- | --- | --- | --- | --- | --- | --- | --- |
| dense (vector) | 0.7000 | 0.9200 | 0.8124 | 0.8581 | 0.8124 | 0.3200 | 28 | 14.49 ms |
| dense (vector) | 0.6397 | 0.8824 | 0.7741 | 0.7943 | 0.7471 | 0.3039 | 39 | 12.89 ms |

## Not evaluated

Expand Down
Loading
Loading