Skip to content

Show failures and degraded states in the editor - #114

Merged
JesseHerrick merged 9 commits into
mainfrom
feat/loud-failures
Oct 4, 2026
Merged

JesseHerrick merged 9 commits into
mainfrom
feat/loud-failures

Conversation

@JesseHerrick

@JesseHerrick JesseHerrick commented Oct 3, 2026 •

Copy link
Copy Markdown
Member

Shows every failure and degraded state in the editor, at the time it happens and again to every editor that attaches later, and makes dexter lsp explain why it cannot start. Also fixes a case where Dexter deleted an index that another process was writing, and two daemons for one directory on macOS.

Why

Before, almost every problem only went to the log. In daemon mode, the index build runs on the headless service, which has no editor client, so even the existing "building the index" messages went nowhere. A typical case: an editor started a different Dexter build than the one that wrote the index. Dexter logged "Index version mismatch … rebuilding index" and "Failed to open index (attempt 1/3: database is locked), rebuilding from scratch…", spent minutes rebuilding, and the editor only wrote these lines to its LSP log. To the user, the language server looked broken.

How it works

  • One reporter per workspace (internal/notify), on the shared IndexCoordinator. The embedded server, the daemon's headless service and every editor session use it.
    • Set/Clear for conditions (no repeat when the same condition is set again), Notify for one-time events, and Begin/Report/End for work-done progress. Progress is used only when the client declares window.workDoneProgress; otherwise the start and the end are shown as messages.
    • A session attaches in initialized and gets the active conditions and running progress replayed in order. Each editor has its own queue, so a slow editor never delays indexing.
  • dexter lsp startup failures: when the proxy cannot serve (a daemon from another build in either direction, a daemon that does not start, a different spelling of the root, a workspace held by dexter init), it reads the editor's initialize, sends window/showMessage (Error) and answers initialize with an error (-32603, data.retry=false) that says what to do, then exits. Before, it printed to stderr and exited. On success, the proxy is unchanged: it reads nothing from stdin before the byte pipe starts.
  • CLI: lookup, references and reindex results carry the active conditions as notes, printed as note:/error: on stderr. workspace/status returns all conditions. These are optional fields, so the daemon contract does not change.

What the editor shows

All messages start with "Dexter: ". Each condition says what happened, what Dexter does about it, and what the user can do, and says so again when it clears.

  • Index written by a newer or older build, index with no version, damaged index: Warning, with progress for the rebuild; Info when the rebuild is complete.
  • Index that cannot be opened because another process holds it, or because of a permission or disk error: Error (see below).
  • First index build: Info with progress, once per workspace.
  • SQL indexes not restored after the bulk build: Error with the dexter init --force steps.
  • Fast full build failed and Dexter indexes one file at a time: Warning.
  • Files that could not be read, parsed, written or removed: one Warning for all of them, cleared when they are indexed or removed.
  • Native file watching unavailable, the FSEvents fallback, directories that cannot be watched: Warning; Info when coverage is complete again.
  • Standard library not found (from the shared root of the workspace), mix not found (only to the editor whose environment lacks it).
  • Formatter failures, and the Elixir/OTP mismatch, kept for each mix project, so one broken project in an umbrella does not alternate messages for the others.
  • Root is the home directory or not an Elixir project.
  • A rename that could not change some files: Error, only to the editor that asked.

Benign log lines (WAL checkpoints, transient watcher errors already counted by the coverage report, formatter process restarts) stay in the log.

Fixes found on the way

  • A locked index was deleted under the process that held it. The daemon deleted and rebuilt the index on any open error, including SQLITE_BUSY. A test reproduces the damage: a second connection writes in rollback mode, as dexter init and older builds do; the daemon waits for busy_timeout, deletes the files, and the other process's COMMIT returns without an error while its data is gone. Now open errors are classified: a damaged index (SQLITE_CORRUPT, SQLITE_NOTADB) is deleted and rebuilt; a busy or locked index is never deleted, the daemon retries with backoff for up to 30 s, and then fails to start with a message that the proxy shows in the editor; permission, disk and other errors are reported and the index is left in place.
  • Two spellings of one directory got two daemons on macOS. The default macOS file system ignores case, but the workspace identity kept the case that the user typed, so ~/Code/app and ~/code/app got two locks and two daemons for one index. The identity now uses the on-disk spelling (fcntl(F_GETPATH)), so the second editor gets the root-mismatch message.

Note

Medium Risk
Touches daemon handshake, SQLite open/delete policy, and LSP startup paths—user-visible but well-tested; incorrect classification of open errors or reporter races could still cause wrong UX or rare data-loss scenarios.

Overview
Dexter now surfaces failures and degraded states in the editor instead of relying on logs. A shared workspace notify.Reporter drives window/showMessage, work-done progress, and condition replay when another editor attaches; the LSP layer reports index build/rebuild/unavailability, aggregated file-index failures, watcher/stdlib/formatter issues, and per-session rename errors. lookup / references / reindex and workspace/status expose the same conditions as notes / conditions on stderr and the control API.

dexter lsp no longer exits silently when the daemon cannot start: the proxy reads initialize, shows an error message, and returns JSON-RPC -32603 with remediation text (contract mismatch, held workspace, spawn failure, etc.).

Index open behavior is safer: only damaged databases are deleted and rebuilt; locked indexes are waited on (up to ~30s) then reported, and other open errors leave the files in place. On macOS, workspace identity uses the filesystem’s stored path case so two spellings of one directory do not spawn two daemons.

Supporting moves include store project-root helpers, indexer.FileError for per-file failures, formatter OTP mismatch reporting that avoids repeated BEAM restarts, and docs/AGENTS guidance for reporting via the reporter.

Reviewed by Cursor Bugbot for commit b6c0e11. Bugbot is set up for automated code reviews on this repo. Configure here.

Before this change, almost every failure or degraded state of Dexter went
only to the log. An editor writes that log to a file that the user does not
read, so the language server looked broken. For example, an index written by
a newer build was rebuilt for minutes, and the editor showed nothing.

Add internal/notify, the one reporting path. A Reporter lives on the shared
IndexCoordinator, so the embedded server, the daemon's headless service, and
every editor session use the same one. It logs each report as before and
sends it to every attached editor:

- Conditions (Set/Clear) are window/showMessage: Error when Dexter does not
  work, Warning when it works with less. A condition that stops says so.
- Long work (cold build, rebuild, large incremental pass) is work-done
  progress when the client supports it, with messages as the fallback.
- An editor that attaches later receives the active conditions and the work
  in progress. Sessions attach in initialized and detach on close.
- Delivery never blocks: each editor has its own queue.

Report through it: an index from another index version, a damaged index
that is rebuilt, an index that another process holds locked or that cannot
be opened for another cause, an index without its SQL indexes, a fast build
that fell back to the slow path, files that could not be indexed (one
aggregate message, which also ends when their directory is removed), a root
that is the home directory or not a project (now checked by the daemon),
native watching unavailable, the fsnotify fallback, directories that cannot
be watched and their recovery, a workspace with no stdlib (read from the
shared root), a formatter that cannot run or an OTP mismatch (for each Mix
project; syntax errors in user code are excluded), and a rename that could
not change some files. A session with no mix and a failed rename are told
to that editor only. The first build of an empty index is told once for
each workspace.

When `dexter lsp` cannot attach to a daemon, it now reads the initialize
request, sends window/showMessage with the explanation and the fix, answers
initialize with JSON-RPC error -32603, and exits.

lookup, references, and reindex results carry the active index conditions
with their keys, and the CLI prints them on stderr. workspace/status lists
all conditions.

Fix two causes of lost or duplicated index work:

- openStore deleted the index for any open error, also for "database is
  locked". A `dexter init` or an older release that held the index in a
  rollback-journal transaction then lost its work in silence. Now only a
  damaged index (SQLITE_CORRUPT, SQLITE_NOTADB) is deleted. A locked index
  is waited for up to 30 seconds and then fails the daemon start with a
  message; permission, disk-space, and similar errors keep the index.
- On macOS the workspace identity kept the case that the caller typed, so
  two spellings of one directory got two locks and two daemons. The
  identity now uses the case that the file system stores (F_GETPATH), so
  the second spelling gets the root-mismatch message.

Reports are made only on state changes. The walk and the watchers pay one
atomic load for each changed file when no file has failed, and a save pays
one more atomic load to see that the set of failed files did not change.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

@cursor cursor Bot left a comment •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

Comment thread internal/lsp/report.go
When the persistent formatter BEAM failed with an Elixir/OTP mismatch,
formatting fell back to `mix format`. A success of that fallback cleared the
mismatch condition. The next format then started the BEAM again, which
failed again and set the condition again. Each save sent an error and a
"formatting works again" message to every editor, and paid for a failed
BEAM start.

Now:

- Only a format by the persistent BEAM clears the OTP condition of its
  build root. A `mix format` success clears only the report that formatting
  does not work in that Mix project.
- After an OTP mismatch, the build root does not start a BEAM again until
  the elixir or mix binary or the _build directory changes, or for 10
  minutes. It formats through `mix format` in the meantime.
- The condition is a Warning that says that formatting still works through
  the slower `mix format` fallback, and how to fix the mismatch. An OTP
  mismatch from `mix format` itself is an Error: then formatting does not
  work.

Before the reporting change, a sync.Once showed the mismatch once for each
session, but the BEAM still started again on each save. Both are now fixed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

@cursor cursor Bot left a comment •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

Comment thread internal/lsp/report.go Outdated
The BEAM OTP mismatch always set a Warning that formatting still works
through the slower `mix format` fallback. When the fallback failed with the
same mismatch, reportFormatFailure also set an Error that formatting does
not work. The two conditions have different keys, so the editor showed both,
and they contradict each other.

Now a failed BEAM start only records the mismatch for its build root. The
formatter decides what to tell the user after the `mix format` fallback ran:

- The fallback works: the Warning that formatting is only slower.
- The fallback fails with the mismatch: only the Error; the Warning is
  cleared without a message.
- Any other fallback failure, such as a syntax error, changes nothing.

When mix format works again, the Error ends ("formatting works again") and
the Warning applies, one time. The mismatch is recorded before the fallback
runs, so the first save already gets the right condition. A recorded
mismatch still stops a BEAM start on each save, and the Warning still clears
only when the BEAM formats.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

@cursor cursor Bot left a comment •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

Comment thread internal/lsp/report.go
The Warning that the fast formatter cannot start belongs to a build root,
which the Mix projects of an umbrella share. reportBeamOTP cleared it
whenever one project's mix format also failed with the OTP mismatch, and set
it again when another project's mix format worked. Saves that alternated
between two such projects sent the Warning again each time.

Now each remembered BEAM mismatch keeps the last fallback outcome of each Mix
project on its build root. The Warning is active while at least one
project's fallback works. It is cleared, without a message, only when no
project's fallback works; each of those projects then has its own Error. A
fallback failure for another cause changes nothing, and setting an active
Warning sends nothing. The outcomes go away with the remembered mismatch, so
the memory is bounded by the number of Mix projects.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

@cursor cursor Bot left a comment •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

Comment thread internal/lsp/report.go
JesseHerrick and others added 2 commits October 3, 2026 20:12
reportBeamOTP updated the fallback outcomes of a build root and decided
whether the shared OTP Warning applies under beamMu, but did the Set or Clear
after it released the lock. A concurrent format could apply a newer decision
first; the older one could then clear the Warning although another Mix
project's last fallback still worked. The outcomes were also kept for each
session, while the Warning is shared by all sessions.

Now the outcomes are on the IndexCoordinator, next to the reporter, and one
lock covers each update and the Set or Clear that it decides. The reporter's
lock is a leaf (it only queues and logs), so this cannot deadlock. A BEAM
format drops the outcomes of its build root and clears the Warning under the
same lock. reportFormatFailure and the mix format success only set or clear
the condition of their own Mix project from their own result, so they have no
stale decision to apply.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

@cursor cursor Bot left a comment •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

Comment thread internal/lsp/server.go Outdated
JesseHerrick and others added 2 commits October 3, 2026 20:36
A full build recorded the files that it could not read or parse, but it
never removed a file from the set. When the only files of a project failed,
the index stayed empty, and the next pass was a full build again. That build
could index those files, or they could be gone, but the warning that they
could not be indexed stayed.

A full build sees every file, so its failures are now collected during the
build and, when the build succeeds, replace the whole set. A file that was
indexed or is gone drops out, and the condition clears with its "indexed now"
message. When the build fails, the incremental walk that follows sets each
file and drops the ones it did not see, as before. The other paths already
keep the set exact: the incremental walk marks each file and keeps only what
it saw, a single-file write marks that file, and a removal drops the path and
everything below it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
#113 replaced the per-file warm walk in startBackgroundReindex with the
reconcile pipeline in reconcile.go. This branch had its reporting hooks
in that walk, so they move into the pipeline:

- A file that cannot be read or parsed is reported by the parse workers
  (newParsePool takes an error callback).
- A file is reported as written only when its rows are committed: each
  file of a committed batch, a file written alone after a failed batch,
  and each file of a committed rebuild. A file whose write fails is
  reported as failed; a failed rebuild falls back to batches, which
  report their files.
- The progress of a large pass counts the committed files, and
  failures.retain(seen) runs after the prune.
- buildTask.End runs on every early return, as before.

reconcileProgressThreshold is a variable so that a test can lower it.
TestReconcilePathsReportFailuresAndProgress covers the batched and the
rebuild path: a failed write is reported once, progress starts and ends,
and the failure clears once the file is written. It fails when the
pipeline does not report written and failed files.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

@cursor cursor Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes and found 1 potential issue.

Fix All in Cursor

❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, have a team admin enable autofix in the Cursor dashboard.

Reviewed by Cursor Bugbot for commit 93a7805. Configure here.

Comment thread internal/lsp/server.go
When the fast cold build fails, the incremental pass indexes the files.
The cold build already shows its own progress, and the old walk counted
progress only on a warm pass; the move of the reporting hooks into the
reconcile pipeline lost that guard, so the fallback could start a second
"updating the index" progress for the same work. The pass now gets the
progress counter only when it is warm.

A test hook makes the fast build fail; TestColdBuildFallbackShowsOneProgress
fails without the fix.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@JesseHerrick
JesseHerrick merged commit e567b9d into main Oct 4, 2026
5 checks passed
@JesseHerrick
JesseHerrick deleted the feat/loud-failures branch October 4, 2026 01:05
JesseHerrick added a commit that referenced this pull request Oct 4, 2026
Main now has one workspace daemon for every frontend (#105) and one name
navigation for the CLI and the editor (#107). This merge makes `dexter mcp`
a frontend of that daemon, like `dexter lsp` and the CLI:

- The frontend keeps the MCP protocol and the root negotiation. It opens no
  store and starts no watcher. For each root it holds one control
  connection to the daemon (daemon.Ensure), and sends each tool call as the
  new control method "mcp/tool".
- The tool bodies run in the daemon, against its store and its headless
  language service. Definitions and references use LookupName and
  ReferenceNames, the same navigation as the editor and `dexter lookup`.
- A tool call waits for a cold index for a short time, then answers with a
  note when the index is still building or has a warning or an error
  condition (#114). The rename tool refuses until the index is complete.
- dexter_reindex uses the daemon's reindex barrier.
- Negotiated and fallback roots resolve with the CLI's project-root search.
  A launch directory that is not a project is refused as a fallback.

Removed because the daemon now does this work: the MCP file watcher, the
per-root store bindings, the git HEAD watch stop, the blocking reindex,
attached mode (`dexter lsp --mcp-listen`, which needs an in-process LSP),
deliverEdits and its text-edit applier, CollectReferences, and Serve.

The rename tools keep the editor rename machinery. RenameFunction and
RenameModule run it on the headless language service, report the files
changed, moved, and not written, and return when the index shows the
rename.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
JesseHerrick added a commit that referenced this pull request Oct 4, 2026
## What

A built-in MCP server, modeled on [gopls
mcp](https://go.dev/gopls/features/mcp), so AI agents can navigate
Elixir codebases through dexter's index instead of grep:

```sh
claude mcp add dexter -- dexter mcp
```

Ten tools, deliberately coarse and agent-oriented rather than 1:1 LSP
methods, and addressed by module/function name rather than file+position
(Elixir modules are not tied to files, which makes name-based addressing
the natural fit for agents):

| Tool | What it does |
|---|---|
| `dexter_workspace` | project layout, index state and active
conditions, stdlib status |
| `dexter_search` | fuzzy workspace symbol search |
| `dexter_definition` | definition with `@doc`/`@spec` and source
snippet; follows defdelegate chains; at most 20 clauses shown |
| `dexter_references` | references including use-chain injected call
sites |
| `dexter_module_api` | moduledoc, public functions with signatures and
doc summaries, delegates, types, callbacks, submodules |
| `dexter_file_outline` | modules/functions a file defines (fresh
parse); only regular files inside the project, up to 10 MB |
| `dexter_implementations` | behaviour implementors and protocol
defimpls |
| `dexter_call_hierarchy` | incoming/outgoing calls |
| `dexter_reindex` | forces a reindex through the daemon (the index also
updates automatically via the daemon's file watcher) |
| `dexter_rename_symbol` | workspace-wide rename of a module or function
with the editor rename machinery: writes changes, moves
convention-following files, reports every file touched, returns when the
index shows the rename |

Transports: stdio (`dexter mcp`) and streamable HTTP (`dexter mcp
--listen ADDR`). `dexter mcp --instructions` prints an agent-facing
guide covering Elixir-specific behavior (modules vs files, defdelegate
following, use-chain injection, behaviours vs protocols).

Uses the official `github.com/modelcontextprotocol/go-sdk` (v1.6.1,
stable), the same SDK gopls uses. Tool input schemas are inferred from
Go param structs.

## Why a built-in MCP server rather than an LSP bridge?

Agent frontends can already drive `dexter lsp` through a generic LSP
bridge (Claude Code's LSP tool, for example), so the real question is
what built-in tools add over bridging.

A bridge inherits LSP's request shapes. Apart from workspace symbol
search, every operation is position-based: the agent must find the file,
locate the exact line and column, make the call, then open each returned
location. Every step is a round trip, and a wrong position silently
returns nothing. A bridge also cannot expose anything the protocol does
not define.

Built-in tools have neither limit:

- Name-based: `dexter_definition {module: MyApp.Accounts, function:
fetch_user}` answers directly.
- Coarse: `dexter_module_api` summarizes a whole module in one call;
references include source lines, so no second pass.
- Beyond LSP's surface: workspace overview and explicit reindexing have
no LSP method to bridge.
- Elixir-aware: server instructions cover defdelegate, use-chain
injection, and modules vs files.
- Client-agnostic: one-line registration in anything that speaks MCP;
bridges exist only in some clients.

Editors keep the LSP; both share the same index.

## How it works: a frontend of the workspace daemon

`dexter mcp` is a frontend of the shared workspace daemon (#105), like
`dexter lsp` and the CLI.

- **No workspace state in the MCP process.** It opens no index and
starts no watcher. For each workspace root it holds one control
connection to the daemon (`daemon.Ensure`), which also keeps the daemon
alive while the session is open. Freshness comes from the daemon's file
watcher and git HEAD poll. When the daemon goes away, the frontend
reconnects once per call; a rename is never sent twice.
- **Tool bodies run in the daemon.** Each tool call goes to the daemon
as one control method, `mcp/tool`, with the tool name and its arguments.
The daemon runs the tool against its store and its headless language
service. One method (instead of one per tool) keeps each tool's
parameter type in one place, shared by the input schema and the body.
- **Cancellation.** A canceled tool call sends `$/cancel` to the daemon,
which cancels that request's context and frees its slot. A frontend runs
at most 32 calls at a time per workspace, so one agent cannot fill the
daemon's request limit that editors share. The daemon `ContractVersion`
is now 4: a newer frontend replaces an older daemon, and an older
frontend gets a clear upgrade error from a newer daemon.
- **Shared navigation.** `dexter_definition` and `dexter_references` use
`LookupName` and `ReferenceNames`, the same navigation as
go-to-definition, find-references, and `dexter lookup` (use chains,
injected aliases, generated functions and their declaring lines).
- **Index state in the answers (#114).** A tool call waits up to 30 s
for a cold index (bounded by the call's context), then answers from what
is indexed. An answer from an index that is still building, or that has
an active warning or error condition, ends with a note. The rename tool
refuses until the index is complete.
- **Unsaved editor buffers.** Tools that show source text use an
editor's buffer for a file only when that buffer has unsaved edits and
the file on disk is not newer than those edits. Otherwise they read the
disk; when both changed, the answer says that an editor also has unsaved
changes. Index lines are mapped into the buffer with a bounded line
diff, and the answer names the files that came from unsaved buffers.
- **Roots.** Client roots and the fallback root resolve with the CLI's
project-root search, so every frontend reaches the same daemon. A root
that is not a project (for example a home directory) is refused, for
client roots and for the launch directory. When a daemon already serves
the directory through another spelling, the frontend uses the daemon's
root. A session serves the first usable root.
- **HTTP mode.** `--listen` accepts only loopback addresses unless
`--listen-unsafe` is given (which logs a warning). The handler refuses
cross-origin requests and keeps the SDK's localhost Host check. Request
bodies are limited to 4 MB (413), and an idle HTTP session ends after 30
minutes.
- **Rename.** Renames from every frontend run one at a time
(`renameSerial`), so two parallel renames cannot lose each other's
edits. A canceled rename stops if it has not started to write; one
canceled during its writes completes, and the agent is told to check
`git status`. The MCP rename does not yet apply the shared rule for
buffers that an editor has open (never write over unsaved editor work);
the README asks users to save before an agent renames, and the rename
tool adopts that rule when it lands in `internal/lsp`.
- **Recovery hints** say `dexter stop --force`, which works while
frontends are attached.

Co-authored-by: Jesse Herrick <jesse@remote.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant