Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 4 additions & 0 deletions .prettierignore
Original file line number Diff line number Diff line change
Expand Up @@ -18,3 +18,7 @@ docs/eval/baseline-*.md
# with `npm run eval:chunking`; the formatter must not be a second writer.
docs/eval/chunking-*.json
docs/eval/chunking-*.md

# And the same again for `scripts/eval-retrieval.mjs` (#77).
docs/eval/retrieval-*.json
docs/eval/retrieval-*.md
4 changes: 2 additions & 2 deletions docs/eval/baseline-v1.5.md
Original file line number Diff line number Diff line change
Expand Up @@ -25,8 +25,8 @@ Generated by `npm run eval`. The numbers below are harness output — do not edi
| nDCG@10 | 0.9437 |
| Evidence precision@5 | 0.2133 |

Timing is informational only and is **not** frozen: indexing 2515 ms, query
p50 13.99 ms, p95 23.62 ms on the
Timing is informational only and is **not** frozen: indexing 2235 ms, query
p50 15.12 ms, p95 26.41 ms on the
machine that produced this file. Timing and index size depend on hardware and on the
corpus, so they must never be the reason two runs differ.

Expand Down
48 changes: 48 additions & 0 deletions docs/eval/retrieval-v1.5.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,48 @@
{
"baseline": "dense",
"chunking": "1000/100",
"strategies": [
{
"id": "dense",
"label": "dense (vector)",
"chunking": "1000/100",
"chunkCount": 19,
"recallAt1": 0.833333,
"recallAt5": 1,
"recallAt10": 1,
"mrr": 0.927778,
"ndcgAt10": 0.94375,
"evidencePrecisionAt5": 0.213333,
"indexingMs": 2041,
"latencyP95Ms": 30.33
},
{
"id": "sparse",
"label": "sparse (BM25)",
"chunking": "1000/100",
"chunkCount": 19,
"recallAt1": 0.733333,
"recallAt5": 0.866667,
"recallAt10": 0.866667,
"mrr": 0.805556,
"ndcgAt10": 0.818355,
"evidencePrecisionAt5": 0.186667,
"indexingMs": 2070,
"latencyP95Ms": 2.5
},
{
"id": "hybrid",
"label": "hybrid (RRF of dense + BM25)",
"chunking": "1000/100",
"chunkCount": 19,
"recallAt1": 0.866667,
"recallAt5": 1,
"recallAt10": 1,
"mrr": 0.944444,
"ndcgAt10": 0.956053,
"evidencePrecisionAt5": 0.213333,
"indexingMs": 2134,
"latencyP95Ms": 41.23
}
]
}
47 changes: 47 additions & 0 deletions docs/eval/retrieval-v1.5.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,47 @@
# Retrieval experiments — v1.5 (#77)

Generated by `node scripts/eval-retrieval.mjs`. Numbers are harness output; do not edit them by hand.

## What was measured

Every strategy runs the real RAG eval harness against the same corpus and the same 30
questions as `baseline-v1.5.json`, with chunking held fixed at 1000/100. Only the retrieval strategy changes.

| Strategy | Recall@1 | Recall@5 | MRR | nDCG@10 | Evidence P@5 | Query p95 |
| --- | --- | --- | --- | --- | --- | --- |
| dense (vector) | 0.8333 | 1.0000 | 0.9278 | 0.9437 | 0.2133 | 30.33 ms |
| sparse (BM25) | 0.7333 | 0.8667 | 0.8056 | 0.8184 | 0.1867 | 2.50 ms |
| hybrid (RRF of dense + BM25) | 0.8667 | 1.0000 | 0.9444 | 0.9561 | 0.2133 | 41.23 ms |

## Not evaluated

**Reranking.** The issue lists "hybrid + reranker" as a step, but a cross-encoder
model is not available offline and inventing its numbers would defeat the point of
the harness. It stays open until a model can be pinned the way the embedding model
is.

## Corpus limitation

The rule's Recall@5 condition is **saturated** on this corpus: dense already scores
1.0000, so no strategy can improve it and the rule can therefore never be met here.
The metrics that still discriminate are Recall@1, MRR and nDCG@10. A hybrid result
that is better on all three but equal on Recall@5 is therefore *inconclusive*, not a
negative result, and the default is left unchanged until the comparison can run on a
corpus where Recall@5 is not already perfect.

## Adoption rule

> Adopt a change only if Recall@5 improves and nDCG@10 does not regress. A change
> that trades a large latency increase for a marginal recall gain is a product
> decision, not an automatic win.

## Outcome

Recall@5 is saturated at 1.0000, so the rule's first condition cannot be met by any strategy. **Dense stays the default**, and the non-dense strategies are reported as inconclusive rather than adopted or rejected on a metric that cannot move.

## Reproduce

```bash
npm run eval:prepare # one-time, networked model bootstrap
npm run eval:retrieval # offline; runs every strategy and rewrites this file
```
3 changes: 2 additions & 1 deletion package.json
Original file line number Diff line number Diff line change
Expand Up @@ -38,7 +38,8 @@
"db:generate": "drizzle-kit generate",
"db:migrate": "drizzle-kit migrate",
"db:push": "drizzle-kit push",
"db:studio": "drizzle-kit studio"
"db:studio": "drizzle-kit studio",
"eval:retrieval": "npm run build && node scripts/eval-retrieval.mjs"
},
"//test": [
"`node --test` strips TypeScript types rather than compiling them, and strip-only",
Expand Down
224 changes: 224 additions & 0 deletions scripts/eval-retrieval.mjs
Original file line number Diff line number Diff line change
@@ -0,0 +1,224 @@
#!/usr/bin/env node
/**
* Retrieval experiments for #77.
*
* Runs the real RAG eval harness once per retrieval strategy against the frozen
* chunk baseline (`baseline-v1.5.json`, 1000/100), holding chunking fixed, and
* writes the comparison the issue asks for as a delta against dense.
*
* The harness does the measuring; this script only orchestrates and tabulates.
*
* Usage:
* node scripts/eval-retrieval.mjs
*
* The embedding model must already be prepared (`npm run eval:prepare`).
*/

import { spawn } from 'node:child_process'
import { existsSync, mkdtempSync, mkdirSync, readFileSync, rmSync, writeFileSync } from 'node:fs'
import { tmpdir } from 'node:os'
import { join, resolve } from 'node:path'

const OUT_MD = resolve(readArg('--out=', 'docs/eval/retrieval-v1.5.md'))
const OUT_JSON = OUT_MD.replace(/\.md$/, '.json')

/**
* `dense` is the shipped default and the comparison baseline. `sparse` is BM25 over
* the shared FTS index (#96); `hybrid` is RRF of the two. A reranker is not
* evaluated: it needs a cross-encoder model, which this offline harness does not
* have, and pretending otherwise would be a number invented by the script.
*/
const STRATEGIES = [
{ id: 'dense', label: 'dense (vector)' },
{ id: 'sparse', label: 'sparse (BM25)' },
{ id: 'hybrid', label: 'hybrid (RRF of dense + BM25)' }
]

function readArg(prefix, fallback) {
const arg = process.argv.find((value) => value.startsWith(prefix))
return arg ? arg.slice(prefix.length) : fallback
}

const executable = resolve(
'node_modules/.bin',
process.platform === 'win32' ? 'electron.cmd' : 'electron'
)

if (!existsSync(executable)) {
console.error('[retrieval] could not find the electron binary. Run `npm install` first.')
process.exit(1)
}

function runStrategy(strategy, outDir) {
return new Promise((resolvePromise, reject) => {
const args = [
'.',
'--eval-harness',
'--eval-baseline=v1.5',
`--eval-out=${outDir}`,
`--eval-retrieval=${strategy.id}`
]

const isRoot = typeof process.getuid === 'function' && process.getuid() === 0
if (isRoot || process.env.CI) args.push('--no-sandbox')

const child = spawn(executable, args, {
stdio: ['ignore', 'pipe', 'pipe'],
env: { ...process.env, ELECTRON_DISABLE_SECURITY_WARNINGS: '1' }
})

let stdout = ''
child.stdout.on('data', (data) => {
stdout += data.toString()
})
child.stderr.on('data', () => {})

child.on('error', reject)
child.on('exit', (code) => {
if (code !== 0) {
reject(new Error(`strategy ${strategy.id} exited with code ${code}`))
return
}
const metricsLine = /\[eval\] metrics (\{.*\})/.exec(stdout)
if (!metricsLine) {
reject(new Error(`strategy ${strategy.id} printed no metrics line`))
return
}
try {
resolvePromise(JSON.parse(metricsLine[1]))
} catch (error) {
reject(
new Error(`strategy ${strategy.id} printed an unreadable metrics line: ${error.message}`)
)
}
})
})
}

/** Throughput and p95 are informational and excluded from the deterministic JSON. */
function readTiming(mdPath) {
if (!existsSync(mdPath)) return { indexingMs: null, latencyP95Ms: null }
const text = readFileSync(mdPath, 'utf8')
const indexing = /indexing (\d+) ms/.exec(text)
const p95 = /p95 ([\d.]+) ms/.exec(text)
return {
indexingMs: indexing ? Number(indexing[1]) : null,
latencyP95Ms: p95 ? Number(p95[1]) : null
}
}

const workDir = mkdtempSync(join(tmpdir(), 'knownote-retrieval-'))
const results = []

try {
for (const strategy of STRATEGIES) {
const outDir = join(workDir, strategy.id)
mkdirSync(outDir, { recursive: true })
console.log(`[retrieval] running ${strategy.label}`)
const metrics = await runStrategy(strategy, outDir)

const report = JSON.parse(readFileSync(join(outDir, 'baseline-v1.5.json'), 'utf8'))
results.push({
id: strategy.id,
label: strategy.label,
chunking: `${report.config.chunking.chunkSize}/${report.config.chunking.chunkOverlap}`,
chunkCount: report.config.chunkCount,
...metrics,
...readTiming(join(outDir, 'baseline-v1.5.md'))
})
}
} finally {
rmSync(workDir, { recursive: true, force: true })
}

const baseline = results.find((result) => result.id === 'dense')
if (!baseline) throw new Error('the dense strategy did not run')

const format4 = (value) => value.toFixed(4)

const rows = results.map(
(result) =>
`| ${result.label} | ${format4(result.recallAt1)} | ${format4(result.recallAt5)} | ${format4(result.mrr)} | ${format4(result.ndcgAt10)} | ${format4(result.evidencePrecisionAt5)} | ${result.latencyP95Ms?.toFixed(2)} ms |`
)

/**
* The rule frozen in the baseline report: Recall@5 must improve and nDCG@10 must
* not regress. Latency is reported so a win that costs 5x latency is stated as a
* trade-off, not hidden.
*/
const adopted = results.filter(
(result) =>
result.id !== 'dense' &&
result.recallAt5 > baseline.recallAt5 &&
result.ndcgAt10 >= baseline.ndcgAt10
)
const winner =
adopted.sort((a, b) => b.recallAt5 - a.recallAt5 || b.ndcgAt10 - a.ndcgAt10)[0] ?? null

let outcome
if (winner) {
outcome = `\`${winner.label}\` clears the rule (Recall@5 ${format4(winner.recallAt5)} vs dense ${format4(baseline.recallAt5)}, nDCG@10 ${format4(winner.ndcgAt10)} vs ${format4(baseline.ndcgAt10)}).`
} else if (baseline.recallAt5 === 1) {
outcome = `Recall@5 is saturated at 1.0000, so the rule's first condition cannot be met by any strategy. **Dense stays the default**, and the non-dense strategies are reported as inconclusive rather than adopted or rejected on a metric that cannot move.`
} else {
outcome = `No strategy cleared the rule. **Dense stays the default.** A negative result is the point of the experiment: it is the measurement that says the extra machinery is not worth its cost on this corpus, not a failure to deliver.`
}

const markdown = `# Retrieval experiments — v1.5 (#77)

Generated by \`node scripts/eval-retrieval.mjs\`. Numbers are harness output; do not edit them by hand.

## What was measured

Every strategy runs the real RAG eval harness against the same corpus and the same 30
questions as \`baseline-v1.5.json\`, with chunking held fixed at ${baseline.chunking}. Only the retrieval strategy changes.

| Strategy | Recall@1 | Recall@5 | MRR | nDCG@10 | Evidence P@5 | Query p95 |
| --- | --- | --- | --- | --- | --- | --- |
${rows.join('\n')}

## Not evaluated

**Reranking.** The issue lists "hybrid + reranker" as a step, but a cross-encoder
model is not available offline and inventing its numbers would defeat the point of
the harness. It stays open until a model can be pinned the way the embedding model
is.

## Corpus limitation

The rule's Recall@5 condition is **saturated** on this corpus: dense already scores
1.0000, so no strategy can improve it and the rule can therefore never be met here.
The metrics that still discriminate are Recall@1, MRR and nDCG@10. A hybrid result
that is better on all three but equal on Recall@5 is therefore *inconclusive*, not a
negative result, and the default is left unchanged until the comparison can run on a
corpus where Recall@5 is not already perfect.

## Adoption rule

> Adopt a change only if Recall@5 improves and nDCG@10 does not regress. A change
> that trades a large latency increase for a marginal recall gain is a product
> decision, not an automatic win.

## Outcome

${outcome}

## Reproduce

\`\`\`bash
npm run eval:prepare # one-time, networked model bootstrap
npm run eval:retrieval # offline; runs every strategy and rewrites this file
\`\`\`
`

mkdirSync(resolve(OUT_MD, '..'), { recursive: true })
writeFileSync(
OUT_JSON,
`${JSON.stringify({ baseline: baseline.id, chunking: baseline.chunking, strategies: results }, null, 2)}\n`
)
writeFileSync(OUT_MD, markdown)

console.log(`[retrieval] wrote ${OUT_JSON} and ${OUT_MD}`)
console.log(
`[retrieval] ${winner ? `best clearing strategy: ${winner.label}` : 'no strategy cleared the rule; keep dense'}`
)
8 changes: 6 additions & 2 deletions src/main/eval/harness.ts
Original file line number Diff line number Diff line change
Expand Up @@ -18,6 +18,7 @@ import { documentBlocks, notebooks, chunks } from '../db/schema'
import type { getDatabase } from '../db'
import type { KnowledgeService } from '../services/KnowledgeService'
import { DEFAULT_CHUNK_OPTIONS, type ChunkOptions } from '../services/ChunkingService'
import type { RetrievalStrategy } from '../services/retrieval'
import { LOCAL_EMBEDDING_MODEL } from '../embedding/localModel'
import {
evidencePrecisionAtK,
Expand Down Expand Up @@ -54,6 +55,8 @@ export interface EvalHarnessOptions {
evidenceK: number
/** 分块配置(#78)。实验变体通过它选择策略;缺省时用生产默认值。 */
chunkOptions: ChunkOptions
/** 检索策略(#77):dense / sparse(BM25) / hybrid(RRF)。 */
strategy: RetrievalStrategy
}

const NOTEBOOK_ID = 'eval-notebook'
Expand Down Expand Up @@ -189,7 +192,8 @@ export async function runEvalHarness(
const started = performance.now()
const results = await knowledgeService.search(NOTEBOOK_ID, question.question, {
topK: options.topK,
threshold: options.threshold
threshold: options.threshold,
strategy: options.strategy
})
const latencyMs = performance.now() - started
latencies.push(latencyMs)
Expand Down Expand Up @@ -235,7 +239,7 @@ export async function runEvalHarness(
allowSpanPages: chunking.allowSpanPages,
respectHeadings: chunking.respectHeadings
},
retrieval: 'dense',
retrieval: options.strategy,
topK: options.topK,
threshold: options.threshold,
evidenceK: options.evidenceK,
Expand Down
Loading
Loading