Skip to content

feat(search): a full-text index over chunks and a global search palette - #169

Merged
mrsibe merged 1 commit into
mainfrom
feat/96-fts-search
Sep 28, 2026
Merged

mrsibe merged 1 commit into
mainfrom
feat/96-fts-search

Conversation

@mrsibe

@mrsibe mrsibe commented Sep 28, 2026

Copy link
Copy Markdown
Owner

What does this PR do?

Adds chunks_fts (FTS5/BM25) and a Ctrl/Cmd+K search palette that shows literal and semantic results as two separate signals. This is the sparse index #77 will fuse with dense.

Why?

Both chat retrieval and the search box were vector-only. Embeddings are measurably bad at exact tokens — model names, ids, symbols, numbers — so there was no way to find a literal string, and a note could not be searched at all. #96 also builds the index #77 needs for BM25/hybrid, so the two share an index but not a ranking.

Related issue

Fixes #96
Related to #154, #77

What changed

  • chunks_fts (FTS5). ensureChunksFts() creates it and backfills missing rows on every call, so a library indexed before FTS existed — or a crash between the chunk write and the FTS write — heals itself.
  • Maintained with the chunks: written in the ingestion pipeline, deleted with the document, so re-index and delete stay consistent.
  • KnowledgeService.searchText() returns the same SearchResult shape as the vector path, so provenance and the jump-to-source chain are not duplicated.
  • Injection-proof queries: user text is quoted term by term before MATCH, so -, NEAR(, col:value, a AND b or an unbalanced quote return no rows instead of a syntax error — a syntax error reads as "no results", not "bad query".
  • SearchPalette (Ctrl/Cmd+K): matching text and related passages as labelled groups, never fused. A hit opens the reader at its passage (the locator travels with the result).
  • KnowledgeSearchResult gained an optional locator, and the preload API a searchText method.

How was this tested?

  • npm run typecheck — passes.
  • npm test — 391 pass, including FTS5 integration against the bundled SQLite: BM25 ordering, notebook isolation, document-scoped filter, operator-looking queries not throwing, snippet extraction.
  • npm run check:design — no violations.
  • npm run build — passes.
  • npm run eval — output byte-identical to baseline-v1.5.json (FTS does not touch dense retrieval).
  • electron . --smoke-test — PASS, 27 checks, including a new one: a literal search finds the term after indexing, and finds exactly one row after a re-index.

Not verified

  • The palette was not clicked in a running app. The FTS path is covered by unit tests and the smoke test; the palette markup only by typecheck and the design guard.

Checklist

  • I have reviewed my own changes.
  • npm run typecheck passes.
  • npm run build passes.
  • I have tested the affected user workflow. (FTS via smoke; palette noted above.)
  • I have not included unrelated changes.
  • I have updated documentation when necessary.

Desktop / build changes

  • Not applicable (the FTS table is created at runtime with IF NOT EXISTS; no drizzle migration).

Chat retrieval and the search box were both vector-only. Embeddings are measurably
bad at exact tokens — model names, ids, symbols, numbers — so there was no way to
find a literal string, and a note could not be searched at all.

Build the sparse index #77 also needs, and give it a user-facing surface:

- `chunks_fts` (FTS5) is created and backfilled by `ensureChunksFts()`. The backfill
  runs on every ensure, so a library indexed before FTS existed, or a crash between
  the chunk write and the FTS write, heals itself.
- The index is written with the chunks in the ingestion pipeline and deleted with
  the document, so it tracks re-index and delete.
- `KnowledgeService.searchText()` returns the same `SearchResult` shape as the
  vector path, so provenance and the jump-to-source chain are not duplicated.
- User text is quoted term by term before it reaches `MATCH`: an operator-looking
  query (`-`, `NEAR(`, `col:value`, an unbalanced quote) returns no rows instead of
  a syntax error, which would read as "no results rather than "bad query".
- `SearchPalette` (Ctrl/Cmd+K) shows **matching text** and **related passages** as
  two labelled groups and never fuses them. Chat retrieval (#77) does fuse them,
  because the model needs one ordered list; the two must not share ranking
  semantics, which is why the shared part is the index, not the ranking.

Verified: npm run typecheck; npm test (391 pass, incl. FTS5 integration against the
bundled SQLite); npm run check:design; npm run build; `npm run eval` unchanged and
byte-identical; `electron . --smoke-test` PASS — 27 checks, including a new one that
a literal search finds the term after indexing and still finds exactly one row
after a re-index.

Not verified: the palette was not clicked in a running app (markup is covered by
typecheck and the design guard).
@mrsibe mrsibe added enhancement New feature or request area:retrieval Retrieval quality and evaluation area:ux Interface, workflow, information architecture labels Sep 28, 2026
@mrsibe
mrsibe merged commit b4dee84 into main Sep 28, 2026
4 checks passed
@mrsibe
mrsibe deleted the feat/96-fts-search branch September 28, 2026 10:27
@mrsibe
mrsibe restored the feat/96-fts-search branch September 28, 2026 10:31
@mrsibe
mrsibe deleted the feat/96-fts-search branch September 28, 2026 10:31
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area:retrieval Retrieval quality and evaluation area:ux Interface, workflow, information architecture enhancement New feature or request

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Feat] Global search across sources and notes, with distinct exact and semantic signals

1 participant