馃攳 Problem Description
The connect/close/never-reconnect failure class tracked in Gentleman-Programming/gentle-ai#1019 (Engram's MCP connection gets SIGINT'd 2-5 seconds after connecting under Claude Code and never reconnects) has no client-free diagnostic. Today the only way to tell a real defect from ordinary session teardown is manual inspection of Claude Code's mcp-logs-<server>/*.jsonl cache fragments: reporters publish raw logs and every reader re-derives the judgment by hand, thread after thread.
馃挕 Proposed Solution
Add a Go package internal/mcplogs that scans a Claude Code cache fragment, groups log lines by (server, sessionId), and classifies each connection lifecycle under the surviving-witness standard proposed in the gentle-ai#1019 thread: a 4-clause candidate (declared close of 2-5s, cache-clear, zero completed tool calls, no later reconnection) is only a defect-suspect when a peer MCP server in the same session demonstrably outlived the close (a tool call completed after our close, or a declared close beyond a 10s window). Peers that closed within the window mean ordinary session teardown; a candidate with no possible witness either way is indeterminate; everything else is healthy.
Delivery is planned as a chained pair: the parse layer (ScanLifecycles, session lifecycle extraction) first, then the classification layer plus fixtures transcribed from the thread evidence. Client-free by design: no Claude Code invocation anywhere, including tests.
馃摝 Affected Area
MCP Server (tools, transport)
馃攧 Alternatives Considered
Manual log reading (the status quo, which does not scale across reporters). A detector running inside Claude Code itself was considered and rejected: acceptance would be client-bound, while fixture-based validation over reporter-published logs keeps every assertion deterministic.
馃搸 Additional Context
Cross-references: Gentleman-Programming/gentle-ai#1019 (approved; the surviving-witness standard and the negative control come from its thread) and #648, the closed engram-side twin of the connection bug. The fixture set includes jjeg1979's unauthenticated short-session runs as the load-bearing negative control (strict 4-clause matches on multiple servers, zero defect-suspects) and a labeled synthetic defect-positive, since no confirmed defect-positive log exists in the wild yet.
馃攳 Problem Description
The connect/close/never-reconnect failure class tracked in Gentleman-Programming/gentle-ai#1019 (Engram's MCP connection gets SIGINT'd 2-5 seconds after connecting under Claude Code and never reconnects) has no client-free diagnostic. Today the only way to tell a real defect from ordinary session teardown is manual inspection of Claude Code's
mcp-logs-<server>/*.jsonlcache fragments: reporters publish raw logs and every reader re-derives the judgment by hand, thread after thread.馃挕 Proposed Solution
Add a Go package
internal/mcplogsthat scans a Claude Code cache fragment, groups log lines by (server, sessionId), and classifies each connection lifecycle under the surviving-witness standard proposed in the gentle-ai#1019 thread: a 4-clause candidate (declared close of 2-5s, cache-clear, zero completed tool calls, no later reconnection) is only adefect-suspectwhen a peer MCP server in the same session demonstrably outlived the close (a tool call completed after our close, or a declared close beyond a 10s window). Peers that closed within the window mean ordinary sessionteardown; a candidate with no possible witness either way isindeterminate; everything else ishealthy.Delivery is planned as a chained pair: the parse layer (
ScanLifecycles, session lifecycle extraction) first, then the classification layer plus fixtures transcribed from the thread evidence. Client-free by design: no Claude Code invocation anywhere, including tests.馃摝 Affected Area
MCP Server (tools, transport)
馃攧 Alternatives Considered
Manual log reading (the status quo, which does not scale across reporters). A detector running inside Claude Code itself was considered and rejected: acceptance would be client-bound, while fixture-based validation over reporter-published logs keeps every assertion deterministic.
馃搸 Additional Context
Cross-references: Gentleman-Programming/gentle-ai#1019 (approved; the surviving-witness standard and the negative control come from its thread) and #648, the closed engram-side twin of the connection bug. The fixture set includes jjeg1979's unauthenticated short-session runs as the load-bearing negative control (strict 4-clause matches on multiple servers, zero defect-suspects) and a labeled synthetic defect-positive, since no confirmed defect-positive log exists in the wild yet.