Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
7 changes: 4 additions & 3 deletions docs/architecture.md
Original file line number Diff line number Diff line change
Expand Up @@ -212,7 +212,7 @@ The cold index is bound by the single SQLite writer, not by parsing. On a 70k-fi
Consequences that are easy to undo by accident:

- **Refs are deduplicated in the parser** (`dedupeRefs`), not in the store. Refs are line-granular, so `@spec f(String.t(), String.t())` produces identical rows; ~60k of them on a large monorepo. Identical rows cannot change a result — no query counts refs, and the References handler dedupes by file+line — so the parse workers drop them before they reach the writer.
- **The bulk path batches inserts** into multi-row `INSERT`s (`multiRowInsert`, 900 bound parameters per statement, which is under even the legacy `SQLITE_MAX_VARIABLE_NUMBER` of 999). Incremental reindex keeps the row-at-a-time path, where a file's `DELETE` must stay ordered ahead of its `INSERT`s.
- **Every batch buffers inserts** into multi-row `INSERT`s (`multiRowInsert`, 900 bound parameters per statement, which is under even the legacy `SQLITE_MAX_VARIABLE_NUMBER` of 999). In an incremental batch a file's `DELETE` targets only its own id, so buffered rows of other files are safe; a file written twice in one batch flushes the buffer before its second `DELETE`.
- **Prefix queries use a range, never `LIKE`.** `LIKE` is case-insensitive by default, so SQLite cannot turn `module LIKE 'Prefix.%'` into an index range and scans all refs. `module >= 'Prefix.' AND module < 'Prefix/'` ('/' is '.'+1) uses `idx_refs_module_function` and turns that scan into a range search: on a 3.9M-row index, 11-14x faster with a warm page cache and ~190x faster cold.
- **`idx_refs_function_kind` was retired.** No query leads with `function`; the two that filter on function/kind both lead with `file_path`. It cost 80 MB and a share of every index rebuild. Check `EXPLAIN QUERY PLAN` before adding an index here — index build time is ~40% of a cold index.

Expand All @@ -223,8 +223,9 @@ The largest remaining win is interning `file_path`: every ref row stores a ~122-
- **Tokenizer instead of tree-sitter for indexing** — a hand-rolled tokenizer + walker replaced the original regex-based parser for both file indexing and runtime `__using__` parsing. The tokenizer handles heredocs, sigils, multi-line expressions, and comments as opaque tokens, eliminating fragile line-joining heuristics. Tree-sitter is only used for scope-aware variable operations in files already opened by the editor.
- **SQLite for storage** — single file, fast reads, incremental updates via mtime tracking.
- **Parallel indexing** — the cold build uses all CPU cores for parsing, single writer for SQLite. Both callers share `indexer.FullBuild`; the server used to have a second, serial implementation that parsed one file at a time and committed a transaction per file, which measured ~4x slower on a 10k-file corpus.
- **Two write paths** — `indexer.FullBuild` is insert-only and allocates file ids from a counter, so it is only correct on an empty index and only with no other writer active. Everything else — incremental sweeps included — goes through the per-file path, which deletes a file's rows before reinserting them.
- **`IndexCoordinator.writes` covers every writer, without exception.** A full build holds it for writing; every other writer — a save, a watched-file event, a rename, and the incremental sweep's own walk — holds it for reading. A new writer that skips it can run inside an insert-only build, whose emptiness precondition it then violates. The coordinator belongs to the workspace rather than to a session, so every LSP session attached to one store shares it, and the daemon's ownership lock extends the precondition across processes: `dexter reindex` in a shell now enters the same queue instead of opening a second writer the lock could not see.
- **Three write paths** — `indexer.FullBuild` is insert-only and allocates file ids from a counter, so it is only correct on an empty index and only with no other writer active. A single-file write deletes a file's rows before reinserting them. The warm sweep reads every stored mtime with one query (`Store.FileStates`), walks, and then parses the changed files on every core: a small change set goes through incremental batches (many files per transaction), and a change set of at least 4096 files and a quarter of the index goes through `Store.BeginRebuild`. A rebuild writes new, unindexed copies of `definitions` and `refs`, copies in the rows of every other file, swaps the copies in, and builds each index with one sort, all in one transaction. Random updates of the name-sorted indexes are what made one transaction per file take minutes on a large change set. Removal follows the same rule: `Store.RemoveFileIDs` deletes by id sets, or rebuilds without the removed files when they are a quarter of the index. A failed rebuild falls back to batches (or to id-set deletes), and a failed batch is written again one file at a time, so only a file that fails by itself is lost. A rebuild takes SQLite's write lock with its first statement, because SQLite does not call the busy handler when a deferred transaction upgrades a read to a write.
- **Shutdown cancels index work.** `workspace.Runtime.Close` calls `IndexCoordinator.CancelWork` first. The sweep, the prune, every removal, the rebuild, and a cold build stop; an open transaction rolls back, and a canceled SQL statement is interrupted. A file and its rows are always in one transaction, so nothing is left half written, and the files a canceled pass did not reach keep their old mtime, so the next start does them.
- **`IndexCoordinator.writes` covers every writer, without exception.** A full build holds it for writing; every other writer — a save, a watched-file event, a rename, and the incremental sweep's own walk — holds it for reading. Single-file writes and removals also take `IndexCoordinator.reindexing`, so they wait behind a sweep instead of racing its batches for SQLite's write lock. A new writer that skips it can run inside an insert-only build, whose emptiness precondition it then violates. The coordinator belongs to the workspace rather than to a session, so every LSP session attached to one store shares it, and the daemon's ownership lock extends the precondition across processes: `dexter reindex` in a shell now enters the same queue instead of opening a second writer the lock could not see.
- **Emptiness is decided under the lock, never sampled.** `Server.fullBuild` tests `IsEmpty` while holding the write lock and reports through its `ran` return value, because one save arriving between a sample and the lock invalidates the answer. `Store.IsEmpty` answers `false` when its own query fails, so a database too broken to count is never taken for an empty one.
- **The sweep's prune re-checks the filesystem.** `pruneMissingFiles` deletes only stored paths that are absent from the walk *and* fail to stat. The walk is one traversal, so a file saved after it passed that directory is legitimately missing from `seen`; the stat is also what stops a walker that yields nothing from deleting the whole index.
- **A cold build is one WAL transaction**, so the `-wal` file grows to roughly the size of the index — a few hundred MiB on a large monorepo — until the post-build checkpoint reclaims it. `wal_autocheckpoint` cannot touch frames belonging to an open transaction, and `InProcess` deliberately leaves `journal_mode` alone, so the server cannot avoid this the way `dexter init` does.
Expand Down
49 changes: 45 additions & 4 deletions internal/indexer/indexer.go
Original file line number Diff line number Diff line change
Expand Up @@ -10,6 +10,7 @@
package indexer

import (
"context"
"errors"
"fmt"
"os"
Expand Down Expand Up @@ -40,6 +41,11 @@ type Options struct {
// store.SetBulkPragmas.
InProcess bool

// Context, when set, cancels the build: parsing stops, the bulk
// transaction rolls back, and FullBuild returns the context's error. The
// index is left empty, so the next start builds it again. Optional.
Context context.Context

// Warn reports a recoverable per-file failure. Optional.
//
// It is called from every parse worker, so it must be safe for concurrent
Expand Down Expand Up @@ -146,6 +152,10 @@ func statFilesParallel(paths []string) []fileEntry {
func FullBuild(s *store.Store, projectRoot string, opts Options) (Stats, error) {
var stats Stats
start := time.Now()
ctx := opts.Context
if ctx == nil {
ctx = context.Background()
}

warn := opts.serialWarn()

Expand All @@ -157,6 +167,12 @@ func FullBuild(s *store.Store, projectRoot string, opts Options) (Stats, error)
}
}

// Nothing is written before the bulk transaction, so a cancel up to that
// point returns with the database untouched.
if err := ctx.Err(); err != nil {
return stats, err
}

// Phase 1: collect file paths and mtimes. Both halves run on all cores:
// the traversal fans out per directory, and stat costs one syscall per
// file — ~70k on a large monorepo, which dominated this phase.
Expand All @@ -165,13 +181,19 @@ func FullBuild(s *store.Store, projectRoot string, opts Options) (Stats, error)
if opts.StdlibRoot != "" {
stdlibPaths = dedupeAgainst(parser.CollectElixirFilesParallel(opts.StdlibRoot), filePaths)
}
if err := ctx.Err(); err != nil {
return stats, err
}
files := statFilesParallel(filePaths)
stdlibFiles := statFilesParallel(stdlibPaths)
stats.Walk = time.Since(start)

// Phase 2a: parse stdlib files in parallel. Definitions only — refs are
// not indexed for stdlib.
stdlibResults := parseStdlib(stdlibFiles)
stdlibResults := parseStdlib(ctx, stdlibFiles)
if err := ctx.Err(); err != nil {
return stats, err
}

// Phase 2b: parse project files in parallel, streaming to the writer.
workers := runtime.NumCPU()
Expand Down Expand Up @@ -199,8 +221,13 @@ func FullBuild(s *store.Store, projectRoot string, opts Options) (Stats, error)
}

go func() {
feed:
for _, f := range files {
fileCh <- f
select {
case fileCh <- f:
case <-ctx.Done():
break feed
}
}
close(fileCh)
wg.Wait()
Expand Down Expand Up @@ -242,6 +269,9 @@ func FullBuild(s *store.Store, projectRoot string, opts Options) (Stats, error)

var writeNanos time.Duration
for res := range resultCh {
if err := ctx.Err(); err != nil {
return stats, abortBuild(s, batch, resultCh, err)
}
writeStart := time.Now()
err := batch.IndexFileWithMtimeAndRefs(res.path, res.mtimeNano, res.defs, res.refs)
writeNanos += time.Since(writeStart)
Expand All @@ -255,6 +285,9 @@ func FullBuild(s *store.Store, projectRoot string, opts Options) (Stats, error)
stats.Write = writeNanos
stats.Parse = time.Duration(parseNanos.Load())

if err := ctx.Err(); err != nil {
return stats, abortBuild(s, batch, resultCh, err)
}
commitStart := time.Now()
if err := batch.Commit(); err != nil {
return stats, restoreIndexes(s, fmt.Errorf("commit: %w", err))
Expand Down Expand Up @@ -290,7 +323,7 @@ type stdlibResult struct {
defs []parser.Definition
}

func parseStdlib(stdlibFiles []fileEntry) []stdlibResult {
func parseStdlib(ctx context.Context, stdlibFiles []fileEntry) []stdlibResult {
if len(stdlibFiles) == 0 {
return nil
}
Expand All @@ -305,6 +338,9 @@ func parseStdlib(stdlibFiles []fileEntry) []stdlibResult {
go func() {
defer wg.Done()
for f := range fileCh {
if ctx.Err() != nil {
continue
}
defs, _, err := parser.ParseFile(f.path)
if err != nil {
continue
Expand All @@ -314,8 +350,13 @@ func parseStdlib(stdlibFiles []fileEntry) []stdlibResult {
}()
}
go func() {
feed:
for _, f := range stdlibFiles {
fileCh <- f
select {
case fileCh <- f:
case <-ctx.Done():
break feed
}
}
close(fileCh)
wg.Wait()
Expand Down
Loading
Loading