Skip to content

Make large warm reconciles and prunes fast, and cancel them on shutdown - #113

Merged
JesseHerrick merged 3 commits into
mainfrom
perf/large-reconcile
Oct 4, 2026
Merged

JesseHerrick merged 3 commits into
mainfrom
perf/large-reconcile

Conversation

@JesseHerrick

@JesseHerrick JesseHerrick commented Oct 3, 2026 •

Copy link
Copy Markdown
Member

Makes a warm reconcile that finds a very large change set fast, makes large prunes faster, and makes dexter stop and SIGTERM end the daemon at once even while such a pass runs.

Summary

A warm start that found a large change set, for example nested checkouts of the whole tree, wrote one transaction per file: each commit wrote every index page that the file touched. A large prune ran three DELETE statements per path. Both took minutes on a large project. During that time, dexter stop and SIGTERM could not end the daemon (the walk and the prune took no cancellation), so it kept the workspace lock without serving its socket, and every editor and CLI call was refused with "owned by another dexter process that is not serving its daemon socket".

Profiles: for the warm add, 59% of CPU was in COMMIT and 18% in row inserts; parsing was about 4%. For the prune, 84% of CPU was in the per-path DELETE statements.

Changes

  • Sweep: reads all stored mtimes with one query (Store.FileStates) instead of one query per file. The walk is the same: one lstat per file.
  • Writes: changed files are parsed on all cores. A small change set is written in batches (2048 files or 500 ms per transaction) with multi-row INSERTs. A change set of at least 4096 files and a quarter of the index uses Store.BeginRebuild: one transaction writes unindexed copies of definitions and refs, copies in the rows of all other files, swaps the copies in, and builds each index with one sort. Readers keep their snapshot until the commit. The rebuild takes the write lock with its first statement, so it waits for busy_timeout like other writers; if it fails, the change set is written in batches. If a batch fails for a reason other than cancellation, its files are written one by one, so only a failing file is lost.
  • Removal: deletes by sets of file ids. When the removed files are a quarter of the index, it rebuilds the tables without them, and falls back to the deletes if the rebuild fails.
  • Cancellation: IndexCoordinator.CancelWork stops the sweep, the prune, every removal, the rebuild, and a cold build (before and during its walk and parse). Runtime.Close calls it first. A file and its rows are always in one transaction, so a canceled pass writes no partial file; the next start finishes the pass from the stored mtimes.
  • Embedded server: single-file writes from saves and watched-file events take the coordinator's mutation lock, as daemon events do, so they wait behind a pass instead of racing its batches for the write lock.

Performance

Synthetic project: 56k Elixir files from the Hex cache, plus four nested copies (223k files).

Case Before After
Warm add of 223k files 3m48s 44–53 s
Prune of 223k files 66 s 23–24 s
Prune of 56k files (half of the index) 16.5 s 16 s
Daemon exit after stop/SIGTERM during these passes 14–225 s (stop failed) 0.07–0.9 s
Warm start, no changes 0.87 s 0.69 s
Warm start, five changed files 1.04 s 0.79 s

A prune that removes about half of the index is not much faster: deleting or copying that many rows costs about the same, and SQLite reads every freed page when it drops a table. It can now be stopped at once.

Known limits

  • SQLite's busy wait cannot be interrupted. A removal that waits for a write lock held by another process returns at the end of busy_timeout (5 s) after a cancel. Only the daemon opens the index in normal use, and inside the daemon all index writers are serialized by the coordinator.
  • A rebuild holds the write lock for its whole transaction, and the WAL grows to about the size of the rebuilt tables while it runs, as the cold build already does.
  • A canceled rebuild loses its progress (one transaction); the next start does it again. Batched passes keep the batches that were committed.

Validation

  • New tests: internal/store/bulk_test.go (deletes and rebuild give the same rows, a rebuild matches the incremental path, a canceled rebuild changes nothing, a file written twice in one batch, a rebuild that waits for a concurrent writer), internal/lsp/reconcile_test.go (both write paths, a cancel in the middle of a pass, a failed write that loses only that file, an embedded file change that waits for a pass), internal/workspace/runtime_test.go (Close during the first pass and during a queued removal, then a reopen finishes the work).
  • An independent review checked the table swap against a fresh schema (columns, foreign keys with ON DELETE CASCADE, partial indexes, integrity_check), readers during a rebuild, cancellation at every stage, and the parameter limits. Its four findings are fixed here.
  • go test ./..., go test -race on store, lsp, workspace and indexer, and golangci-lint pass.

🤖 Generated with Claude Code


Note

High Risk
Changes core SQLite write paths, table-swap rebuilds, and cross-process indexing concurrency; mistakes could corrupt indexes, leave stale data, or worsen shutdown races.

Overview
Speeds up large warm reconciles and prunes by replacing one SQLite transaction per changed file with batched multi-row inserts (2048 files or 500 ms), parallel parsing, and a single upfront FileStates query instead of per-file mtime lookups. When the change set is large enough (≥4096 files and ≥¼ of the index), writes go through Store.BeginRebuild: unindexed scratch tables, copy untouched rows, swap, and rebuild indexes in one transaction, with batch fallback on failure.

Large removals resolve paths to ids and delete in chunks; at the same scale threshold they use the same rebuild-without-removed-files path, with delete fallback.

Shutdown and coordination: IndexCoordinator.CancelWork propagates through cold FullBuild, warm reconcile, prune, removals, and rebuilds (BeginBatchContext / BeginRebuild / RemoveFilesContext). Runtime.Close cancels index work first so the daemon does not hold the workspace lock for minutes. Embedded saves and watched-file events use ReconcileFile (mutation lock) so they queue behind a pass instead of racing batches for the write lock.

Reviewed by Cursor Bugbot for commit 2fcf5ec. Bugbot is set up for automated code reviews on this repo. Configure here.

A warm start that found a large change set (for example, nested
checkouts of the whole tree) wrote one transaction per file. Each
commit wrote every index page that the file touched. A large prune ran
three DELETE statements per path. Both took minutes on a large
monorepo. During that time, stop and SIGTERM could not end the daemon,
and it kept the workspace lock without serving its socket.

Changes:
- The sweep reads all stored mtimes with one query, not one query for
  each file. The walk is the same: one lstat for each file.
- Changed files are parsed on all cores. A small change set is written
  in batches of many files for each transaction, with multi-row
  INSERTs. When a batch fails, its files are written again one at a
  time, so only a file that fails by itself is lost.
- A change set of at least 4096 files and a quarter of the index uses
  Store.BeginRebuild. One transaction writes new, unindexed copies of
  definitions and refs, copies in the rows of all other files, swaps
  the copies in, and builds each index with one sort. Readers keep
  their snapshot until the commit. The rebuild takes the write lock
  with its first statement, so it waits for busy_timeout like other
  writers. If the rebuild fails, the change set is written in batches.
- Removal deletes by sets of file ids. When the removed files are a
  quarter of the index, it rebuilds the tables without them, and falls
  back to the deletes if the rebuild fails.
- IndexCoordinator.CancelWork stops the sweep, the prune, every
  removal, the rebuild, and a cold build (before and during its walk
  and parse). Runtime.Close calls it first. A file and its rows are
  always in one transaction, so a canceled pass writes no partial file.
  The next start finishes the pass from the stored mtimes.
- In an embedded server, single-file writes from saves and watched-file
  events take the coordinator's mutation lock, as daemon events do. They
  wait behind a pass instead of racing its batches for the write lock.

Measured on a synthetic project: 56k Elixir files from the Hex cache,
plus four nested copies (223k files).
- Warm add of 223k files: 3m48s -> 44s to 53s
- Prune of 223k files: 66s -> 23s to 24s
- Prune of 56k files (half of the index): 16.5s -> 16s
- Daemon exit after stop or SIGTERM during these passes: 14s to 225s
  -> 0.07s to 0.9s
- Warm start, no changes: 0.87s -> 0.69s; five changed files:
  1.04s -> 0.79s

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

@cursor cursor Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes and found 1 potential issue.

Fix All in Cursor

❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, have a team admin enable autofix in the Cursor dashboard.

Reviewed by Cursor Bugbot for commit 9827981. Configure here.

Comment thread internal/lsp/reconcile.go
When every changed file of a large change set failed to parse, the
rebuild still copied the old tables, swapped them in and built every
index, although nothing changed: a file that fails to parse is not
written, so its old rows are copied and kept. The next warm start then
did the same for the same files. The rebuild now rolls back when no
file was written.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@JesseHerrick
JesseHerrick merged commit 932d18d into main Oct 4, 2026
5 checks passed
@JesseHerrick
JesseHerrick deleted the perf/large-reconcile branch October 4, 2026 00:48
JesseHerrick added a commit that referenced this pull request Oct 4, 2026
#113 replaced the per-file warm walk in startBackgroundReindex with the
reconcile pipeline in reconcile.go. This branch had its reporting hooks
in that walk, so they move into the pipeline:

- A file that cannot be read or parsed is reported by the parse workers
  (newParsePool takes an error callback).
- A file is reported as written only when its rows are committed: each
  file of a committed batch, a file written alone after a failed batch,
  and each file of a committed rebuild. A file whose write fails is
  reported as failed; a failed rebuild falls back to batches, which
  report their files.
- The progress of a large pass counts the committed files, and
  failures.retain(seen) runs after the prune.
- buildTask.End runs on every early return, as before.

reconcileProgressThreshold is a variable so that a test can lower it.
TestReconcilePathsReportFailuresAndProgress covers the batched and the
rebuild path: a failed write is reported once, progress starts and ends,
and the failure clears once the file is written. It fails when the
pipeline does not report written and failed files.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant