From 1dfaf610f561143bf8f862f9f5063933407b9f4a Mon Sep 17 00:00:00 2001 From: Claude Date: Sun, 4 Oct 2026 22:28:12 +0000 Subject: [PATCH 01/15] docs(redaction): Phase 0 findings and choke-point decision Survey of the sync/index/summarize pipeline for the inline secret redaction fork. Picks the archive write as the choke point (option b) and records the three supporting changes it needs: indexer parses the archive, summarizer resume/fork disabled under redaction, staging exports redacted at write. Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_0141gGbp343kbYP6imaJaji2 --- docs/redaction/PHASE0-FINDINGS.md | 184 ++++++++++++++++++++++++++++++ 1 file changed, 184 insertions(+) create mode 100644 docs/redaction/PHASE0-FINDINGS.md diff --git a/docs/redaction/PHASE0-FINDINGS.md b/docs/redaction/PHASE0-FINDINGS.md new file mode 100644 index 00000000..d73e34eb --- /dev/null +++ b/docs/redaction/PHASE0-FINDINGS.md @@ -0,0 +1,184 @@ +# Redaction — Phase 0 findings + +Research note for the inline secret redaction fork. Read-only survey of +`src/`, `cli/`, and `hooks/` at upstream `7e06519` (v1.6.0 + fixes). This +note answers the six Phase 0 questions and records the choke-point decision. + +## Pipeline as it exists today + +``` +SessionStart hook (hooks/hooks.json) + └─ cli/episodic-memory.js sync --background + └─ dist/sync-cli.js + ├─ exportOpencodeSessions() opencode.db → /opencode-transcripts/*.jsonl + └─ syncConversations(src, archive) for each source dir + ├─ copyIfNewer(src → archive) byte-for-byte copy + ├─ parseConversation(archive) → exchanges + │ ├─ generateExchangeEmbedding(user, assistant, toolNames) + │ └─ insertExchange() → SQLite exchanges + tool_calls + vec_exchanges + └─ summarizeConversation(exchanges, sessionId) → -summary.txt + +episodic-memory index (cli/index-conversations.js → dist/index-cli.js → indexer.ts) + └─ copyFileSync(src → archive), parseConversation(**src**), embed, insert, summarize + +episodic-memory import-cursor-history (cursor-legacy.ts) + └─ state.vscdb → /cursor-legacy-export/*.jsonl (a sync source dir) + +episodic-memory index --repair (verify.ts) parses archive files, re-embeds, re-summarizes +``` + +## Answers + +### 1. Where does sync copy into the archive? Byte-for-byte or parsed first? + +Byte-for-byte. `copyIfNewer()` in `src/sync.ts` does `fs.copyFileSync` to +`.tmp.` then `renameSync`, then stamps the source mtime on the +destination (the mtime is the "is the archive current?" check on the next +run). Parsing happens **after** the copy, and in `sync.ts` it parses the +archive file, not the source. + +There are three more copy sites, all in `src/indexer.ts` +(`indexConversations`, `indexSession`, `indexUnprocessed`). They also use +`fs.copyFileSync`, but they **parse the source path**, not the archive. That +matters: redacting the archive copy alone would not reach the index on the +`episodic-memory index` path. + +### 2. One shared parse path or one per source? + +One entry point, `parseConversation()` in `src/parser.ts`. It sniffs the file +and dispatches to one of five per-harness parsers (Claude, Codex, Cursor, +opencode, OMP). More important for redaction: every harness reaches the +archive through the same copy step. opencode and legacy Cursor are first +exported from their SQLite stores into staging JSONL directories under the +plugin's config dir (`opencode-transcripts/`, `cursor-legacy-export/`), and +those staging directories are then ordinary sync sources. + +The staging exports are plugin-owned plaintext copies outside any harness's +own retention, so they count as a fourth sink even though the spec's table +doesn't list them. + +### 3. Where is exchange text written to SQLite, and where is the embedding input built? + +- `insertExchange()` in `src/db.ts` writes `exchanges.user_message`, + `exchanges.assistant_message`, `tool_calls.tool_input` (JSON-stringified + tool input), and `tool_calls.tool_result`. Text search (`search.ts`) runs + `LIKE` over `user_message`/`assistant_message`. +- Embedding input is `generateExchangeEmbedding(userMessage, assistantMessage, + toolNames)` in `src/embeddings.ts`, called from `sync.ts`, `indexer.ts`, + `verify.ts`, and `embedding-migration.ts`. The migration re-embeds from the + **SQLite rows**, so it inherits whatever text is already stored. + +All of these are built from the parsed `ConversationExchange` objects, so a +clean parse input gives clean SQLite text and clean embedding input. + +### 4. Where does the summarizer get its input? + +Usually from the parsed exchanges (`formatConversationText(exchanges)`), and +`sync.ts` parses those from the archive. **But two paths bypass the archive +entirely:** + +- **Claude session resume.** For Claude conversations with at most 15 + exchanges, `summarizeConversation()` calls the Agent SDK with + `resume: sessionId`. The SDK loads the **source** transcript from + `~/.claude/projects/...` and sends it to the model. The prompt carries no + transcript text on this path. +- **Codex fork.** For Codex conversations, `callCodex()` runs + `thread/fork` on the **source** rollout via `codex app-server`. Note that + `getCodexSessionId()` derives the id from the exchanges when no + `sessionId` is passed, so passing `undefined` is not enough to stop it. + +Both paths send unredacted source content to a model endpoint no matter what +the archive holds. Redaction therefore has to turn resume and fork off and +force the transcript-text path, which is built from redacted exchanges. + +### 5. How does the `DO NOT INDEX THIS CHAT` marker work? + +`shouldSkipConversation()` in `src/sync.ts` streams a file in 1 MiB chunks +and looks for any of three markers. If the read fails, it fails closed +(skips the file). Sync calls it on the **archive** path, after the copy, to +gate both indexing and summary queueing. So an excluded conversation is +still copied to the archive; only the index and the summarizer skip it. + +The hint holds: the archive file is the reference that every downstream +stage of `sync` reads. The marker check sits right after the archive write, +and that is where the redaction hook belongs. + +### 6. Must the archive stay valid harness JSONL? + +Yes, on two counts: + +- **Structure.** `show.ts` (CLI `show`, MCP `read`) runs `JSON.parse` on every + line, and one invalid line throws. Harness detection in both `parser.ts` + and `show.ts` keys off line shapes (`type`, `payload`, `role`, ...). +- **Line numbers.** `exchanges.line_start`/`line_end` index into archive lines. + MCP `read` takes `startLine`/`endLine`, and incremental indexing resumes + from `MAX(line_end)` (#152). Redaction must keep **exactly one output line + per input line**. + +So redaction parses each line as JSON, redacts string values, and +re-serializes only the lines that changed. Unchanged lines stay +byte-identical. A line that isn't valid JSON (for example, a partially +written last line) is redacted as raw text, because it was already invalid. + +## Decision: choke point (b), the archive write + +Option (a) doesn't exist. No single function sees every source *before* the +archive write. The only place every harness converges is the copy into the +archive itself. + +So the hook goes at the **archive write**, which becomes a redacting copy +(`copyFileRedacted()` in `src/redaction.ts`), and every downstream stage +reads from the archive. Three supporting changes are needed to make +"downstream reads the archive" actually true: + +1. **`indexer.ts` parses the archive, not the source.** All three copy sites + switch to the redacting copy. They now refresh the archive when the source + is newer, as `sync` already does, so parsing the archive loses no data. +2. **Summarizer resume and fork are disabled while redaction is active.** + `summarizeConversation()` gains an `allowResume` option. `sync.ts`, + `indexer.ts`, and `verify.ts` pass `allowResume: false` when a redactor is + active, which forces the transcript-text path built from redacted + exchanges. Trade-off: Codex-only users then summarize through the Claude + transcript path. If no Claude auth is available, that writes a retryable + error sentinel. `EPISODIC_MEMORY_SKIP_SUMMARIES=1` turns summaries off + entirely. +3. **Staging exports are redacted at write time.** The opencode export + (`opencode-sync.ts`) and the Cursor legacy export (`cursor-legacy.ts`) run + each JSONL line through the same `redactJsonlLine()` before writing. They + are plugin-owned copies, and the archive hook alone would leave a + plaintext copy beside it. + +Everything else (SQLite text, `tool_calls`, embeddings, summaries, `show`, +MCP `read`, embedding migration) reads the archive or the rows built from +it, so it inherits clean text with no further hooks. + +## SOPS + +Decrypted SOPS output (`sops -d`, `sops exec-env`) **drops** the `sops:` +metadata block and comes out as plain YAML, JSON, or dotenv. Its shape is +not reliable. There is no dedicated SOPS rule. Coverage comes from the +value-level rules (storage keys, client secrets, connection strings, private +keys) plus the keyword-assignment rule (`password: ...`, `client_secret=...`). +Encrypted values (`ENC[AES256_GCM,data:...]`) are left alone because they +are safe to keep. + +## Windows + +The redaction code is pure Node `fs` plus regex, and it adds no native +dependencies. The temp-file-and-rename pattern already runs on Windows in +upstream `copyIfNewer`. I could **not** verify sqlite-vec or Transformers.js +on Windows from this Linux container. That is unchanged upstream surface, +and it still needs a manual check on a Windows host before rollout. + +## Other observations + +- `embedding-migration.ts` re-embeds from SQLite text. After a + `redact --rewrite` cleans the rows, the migration can't reintroduce + secrets. +- The summarizer's `SummarizerSdkError` keeps up to 300 characters of the + SDK's `result` text in logs and error sentinels. If a model ever echoed a + secret, it would land there. Feeding the summarizer redacted input removes + that source. +- `verify.ts --repair` re-summarizes without a `sessionId`, but the Codex fork + still fires through `getCodexSessionId()`'s fallback. It needs + `allowResume: false` too. From 9b16810c75a779018a98de1cd1618404e1b64f81 Mon Sep 17 00:00:00 2001 From: Claude Date: Sun, 4 Oct 2026 22:42:53 +0000 Subject: [PATCH 02/15] feat(redaction): redact secrets at the archive write, before index/embed/summarize Adds an inline, rules-driven redaction stage at the one point every harness passes through: the copy into the conversation archive (Phase 0 option b). Secrets become typed tokens ([REDACTED:]); every downstream sink reads the archive, so SQLite text, tool_calls, embeddings, summaries and show/read only see redacted text. - src/redaction.ts: engine, rules loading/validation, JSONL-preserving line redaction, streaming copyFileRedacted. Findings are rule IDs and counts only. Idempotent: tokens never re-match. - src/redaction-rules.ts: bundled defaults ported from gitleaks for this stack (Azure client secret / storage key / SAS / connection strings, private keys, JWT, provider keys, url creds, key=value assignments); git SHA + GUID allowlist; entropy fallback off by default. - sync.ts / indexer.ts: archive copy goes through the redactor; indexer now parses the archive instead of the source. - summarizer: allowResume option; callers disable Claude resume and Codex fork under redaction, since both read the unredacted source transcript. - opencode / legacy Cursor staging exports are redacted at write time. - Fail closed: strict mode (default) aborts sync/index/import when rules fail to load. Env: EPISODIC_MEMORY_REDACTION, _RULES, _STRICT. - `episodic-memory redact --rewrite [--dry-run]` backfills the existing archive, staging dirs and index in place (re-embeds changed rows, deletes summaries built from unredacted text); also --stdin and --print-default-rules. - Tests: per-rule positives, allowlist negatives, idempotency, config, JSONL structure, end-to-end pipeline (archive/SQLite/embedding input/ summarizer input/logs), fail-closed, rewrite. Fake secrets are built at runtime; no literal in the repo matches a rule. Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_0141gGbp343kbYP6imaJaji2 --- CHANGELOG.md | 17 + README.md | 33 ++ cli/episodic-memory.js | 5 + dist/cursor-legacy.d.ts | 3 + dist/cursor-legacy.js | 9 +- dist/indexer.js | 54 +-- dist/opencode-sync.d.ts | 3 + dist/opencode-sync.js | 13 +- dist/redact-cli.d.ts | 1 + dist/redact-cli.js | 123 +++++++ dist/redact-rewrite.d.ts | 40 +++ dist/redact-rewrite.js | 170 +++++++++ dist/redaction-rules.d.ts | 27 ++ dist/redaction-rules.js | 164 +++++++++ dist/redaction.d.ts | 137 +++++++ dist/redaction.js | 484 +++++++++++++++++++++++++ dist/summarizer.d.ts | 11 +- dist/summarizer.js | 7 +- dist/sync-cli.js | 35 +- dist/sync.d.ts | 15 + dist/sync.js | 43 ++- dist/verify.js | 4 +- docs/REDACTION.md | 167 +++++++++ src/cursor-legacy.ts | 11 +- src/indexer.ts | 59 +-- src/opencode-sync.ts | 17 +- src/redact-cli.ts | 141 ++++++++ src/redact-rewrite.ts | 220 ++++++++++++ src/redaction-rules.ts | 168 +++++++++ src/redaction.ts | 618 ++++++++++++++++++++++++++++++++ src/summarizer.ts | 22 +- src/sync-cli.ts | 36 +- src/sync.ts | 52 ++- src/verify.ts | 4 +- test/fake-secrets.ts | 146 ++++++++ test/redact-rewrite.test.ts | 172 +++++++++ test/redaction-pipeline.test.ts | 353 ++++++++++++++++++ test/redaction.test.ts | 526 +++++++++++++++++++++++++++ 38 files changed, 4034 insertions(+), 76 deletions(-) create mode 100644 dist/redact-cli.d.ts create mode 100644 dist/redact-cli.js create mode 100644 dist/redact-rewrite.d.ts create mode 100644 dist/redact-rewrite.js create mode 100644 dist/redaction-rules.d.ts create mode 100644 dist/redaction-rules.js create mode 100644 dist/redaction.d.ts create mode 100644 dist/redaction.js create mode 100644 docs/REDACTION.md create mode 100644 src/redact-cli.ts create mode 100644 src/redact-rewrite.ts create mode 100644 src/redaction-rules.ts create mode 100644 src/redaction.ts create mode 100644 test/fake-secrets.ts create mode 100644 test/redact-rewrite.test.ts create mode 100644 test/redaction-pipeline.test.ts create mode 100644 test/redaction.test.ts diff --git a/CHANGELOG.md b/CHANGELOG.md index 511074ff..5d74d39d 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -5,6 +5,23 @@ All notable changes to this project will be documented in this file. The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.0.0/), and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html). +## [Unreleased] + +### Security + +- Secrets no longer leak into your conversation memory. Azure client secrets, storage keys, SAS signatures, connection-string passwords, private keys, JWTs, GitHub, Anthropic, OpenAI, and AWS keys, and `password: …` style assignments are now replaced with tokens like `[REDACTED:azure-storage-key]` before a conversation is archived, indexed, embedded, or summarized. Before this, every pasted secret was copied into three new places on disk and could be sent to the summarization model. The rest of the conversation stays searchable, and so do git SHAs and Azure tenant, client, and object IDs. Redaction is on by default and fails closed: if the rules can't load, sync won't run rather than store secrets. +- With redaction on, summaries no longer resume Claude Code sessions or fork Codex threads. Both of those let the model read the original, unredacted transcript. Summaries now come from the redacted text. Codex-only users without Claude configured can set `EPISODIC_MEMORY_SKIP_SUMMARIES=1`. + +### Added + +- `episodic-memory redact --rewrite` cleans data you indexed before upgrading. It redacts the archive and the search index in place, re-embeds only the messages that changed, and deletes summaries built from unredacted text so they regenerate. Use `--dry-run` to preview. +- Custom redaction rules via `~/.config/superpowers/redaction-rules.json`, which extends the bundled defaults. Try rules with `episodic-memory redact --stdin`. See `docs/REDACTION.md`. +- New settings: `EPISODIC_MEMORY_REDACTION` (`on`/`off`), `EPISODIC_MEMORY_REDACTION_RULES`, and `EPISODIC_MEMORY_REDACTION_STRICT`. + +### Changed + +- `episodic-memory index` now parses the archived copy instead of the source transcript, and refreshes that copy when the source has grown. This is the same behavior `sync` already had. + ## [1.6.0] - 2026-09-08 Adds a fifth conversation source, an off switch for automatic syncing, and two fixes for real-world resource problems. diff --git a/README.md b/README.md index ee952317..c8105431 100644 --- a/README.md +++ b/README.md @@ -327,6 +327,17 @@ Add to `.claude/hooks/session-end`: episodic-memory sync ``` +### `episodic-memory redact` + +```bash +episodic-memory redact --rewrite --dry-run # what would change +episodic-memory redact --rewrite # redact the existing archive + index in place +episodic-memory redact --stdin < file.jsonl # try the rules on some text +episodic-memory redact --print-default-rules +``` + +See [docs/REDACTION.md](docs/REDACTION.md). + ### `episodic-memory stats` Display index statistics including conversation counts, date ranges, and project breakdown. @@ -415,6 +426,28 @@ open output.html 4. **Index** - Stores in SQLite with sqlite-vec for fast similarity search 5. **Search** - Semantic search using vector similarity or exact text matching +## Secret Redaction + +Secrets in your conversations (Azure client secrets, storage keys, SAS +signatures, connection-string passwords, private keys, JWTs, provider API keys, +`password: …` assignments) are replaced with typed tokens such as +`[REDACTED:azure-storage-key]` **before** anything is archived, indexed, +embedded, or sent to the summarizer. Git SHAs and GUIDs are never redacted, so +they stay searchable. The tokens are searchable too. + +Redaction is on by default and fails closed: if the rules can't be loaded, sync +won't run. After upgrading, run `episodic-memory redact --rewrite` once to clean +data you indexed before. + +| Variable | Default | Meaning | +|---|---|---| +| `EPISODIC_MEMORY_REDACTION` | `on` | `off` disables redaction | +| `EPISODIC_MEMORY_REDACTION_RULES` | `/redaction-rules.json` | Custom rules file (extends the defaults) | +| `EPISODIC_MEMORY_REDACTION_STRICT` | `1` | `0` continues unredacted (with a warning) when rules fail to load | + +See [docs/REDACTION.md](docs/REDACTION.md) for the rule list, the custom-rules +format, and what is and isn't covered. + ## Excluding Conversations Conversations containing this marker anywhere in their content will be archived but not indexed: diff --git a/cli/episodic-memory.js b/cli/episodic-memory.js index 08ebccc8..45c61da1 100755 --- a/cli/episodic-memory.js +++ b/cli/episodic-memory.js @@ -44,6 +44,7 @@ COMMANDS: stats Show index statistics doctor Diagnose Claude Code or Codex integration issues import-cursor-history Export legacy Cursor conversations from state.vscdb for indexing + redact Re-run secret redaction over the archive and index (--rewrite) Run 'episodic-memory --help' for command-specific help. @@ -94,6 +95,10 @@ async function main() { await runScript(join(distDir, 'cursor-import-cli.js'), args); break; + case 'redact': + await runScript(join(distDir, 'redact-cli.js'), args); + break; + case '--help': case '-h': case undefined: diff --git a/dist/cursor-legacy.d.ts b/dist/cursor-legacy.d.ts index a035ba0f..77340197 100644 --- a/dist/cursor-legacy.d.ts +++ b/dist/cursor-legacy.d.ts @@ -1,3 +1,4 @@ +import { type Redactor } from './redaction.js'; export declare function getDefaultCursorVscdbPath(): string | undefined; /** * Collect composer IDs that already have live agent transcripts under @@ -14,6 +15,8 @@ export interface CursorLegacyImportOptions { force?: boolean; /** Report what would be exported without writing files. */ dryRun?: boolean; + /** `undefined` loads from the environment (strict mode throws on bad rules); `null` disables. */ + redactor?: Redactor | null; } export interface CursorLegacyImportResult { exported: number; diff --git a/dist/cursor-legacy.js b/dist/cursor-legacy.js index 811e5637..7b140e0c 100644 --- a/dist/cursor-legacy.js +++ b/dist/cursor-legacy.js @@ -3,6 +3,7 @@ import os from 'os'; import path from 'path'; import Database from 'better-sqlite3'; import { detectCursorCwd } from './parser.js'; +import { loadRedactor, redactJsonlLine } from './redaction.js'; /** * Import legacy Cursor conversations from Cursor's global SQLite store * (state.vscdb) into JSONL files compatible with the Cursor transcript parser. @@ -111,6 +112,7 @@ export function importCursorLegacy(options) { skippedEmpty: 0, errors: [], }; + const redactor = options.redactor === undefined ? loadRedactor() : options.redactor; const liveIds = options.liveTranscriptIds ?? new Set(); const db = new Database(options.dbPath, { readonly: true, fileMustExist: true }); try { @@ -200,9 +202,14 @@ export function importCursorLegacy(options) { if (!options.dryRun) { // Re-serialize with cwd now that it's known (it's derived from the // whole conversation's tool calls). - const finalLines = cwd + const withCwd = cwd ? lines.map(line => JSON.stringify({ ...JSON.parse(line), cwd })) : lines; + // The export dir is a plugin-owned plaintext copy, so redact it at + // write time like the archive (docs/redaction/PHASE0-FINDINGS.md). + const finalLines = redactor + ? withCwd.map(line => redactJsonlLine(line, redactor, { source: 'cursor-legacy', path: outFile })) + : withCwd; fs.mkdirSync(path.dirname(outFile), { recursive: true }); fs.writeFileSync(outFile, finalLines.join('\n') + '\n', 'utf-8'); // Stamp the conversation's end time so mtime-based fallbacks and diff --git a/dist/indexer.js b/dist/indexer.js index 8afa1d35..5b4c803d 100644 --- a/dist/indexer.js +++ b/dist/indexer.js @@ -7,6 +7,8 @@ import { summarizeConversation } from './summarizer.js'; import { getArchiveDir, getExcludedProjects, getConversationSourceDirs, findJsonlFiles, statIfExists } from './paths.js'; import { formatErrorSentinel, shouldQueueForSummary } from './summary-sentinel.js'; import { getMaxMessageBytes, isOversizeExchange } from './message-size.js'; +import { FindingsTally, formatFindings, loadRedactor } from './redaction.js'; +import { copyIfNewer } from './sync.js'; // Set max output tokens for Claude SDK (used by summarizer) process.env.CLAUDE_CODE_MAX_OUTPUT_TOKENS = '20000'; // Increase max listeners for concurrent API calls @@ -25,7 +27,18 @@ async function processBatch(items, processor, concurrency) { function sessionIdForSummary(exchanges) { return exchanges.find(exchange => exchange.sessionId)?.sessionId; } +// Resume/fork would summarize the unredacted source transcript; see sync.ts. +function summarizeOptions(redactor) { + return { allowResume: redactor === null }; +} +function logRedactions(tally) { + if (tally.total > 0) + console.log(` Redaction: ${formatFindings(tally.toArray())}`); +} export async function indexConversations(limitToProject, maxConversations, concurrency = 1, noSummaries = false) { + // Load before touching the archive: strict mode fails closed here. + const redactor = loadRedactor(); + const tally = new FindingsTally(); console.log('Initializing database...'); const db = initDatabase(); console.log('Loading embedding model...'); @@ -73,14 +86,13 @@ export async function indexConversations(limitToProject, maxConversations, concu // Source transcripts can vanish mid-run (Claude Code cleanup). Skip loudly. let exchanges; try { - // Copy to archive (ensure parent dirs exist for subagent files) - if (!fs.existsSync(archivePath)) { - fs.mkdirSync(path.dirname(archivePath), { recursive: true }); - fs.copyFileSync(sourcePath, archivePath); + // Copy (redacted) to the archive, then parse the archive, so the index + // and summaries only ever see redacted text. + if (copyIfNewer(sourcePath, archivePath, redactor, tally)) { console.log(` Archived: ${file}`); } // Parse conversation - exchanges = await parseConversation(sourcePath, project, archivePath); + exchanges = await parseConversation(archivePath, project, archivePath); } catch (error) { console.log(` Skipped ${file} (read failed: ${error instanceof Error ? error.message : error})`); @@ -105,7 +117,7 @@ export async function indexConversations(limitToProject, maxConversations, concu console.log(` Generating ${needsSummary.length} summaries (concurrency: ${concurrency})...`); await processBatch(needsSummary, async (conv) => { try { - const summary = await summarizeConversation(conv.exchanges, sessionIdForSummary(conv.exchanges)); + const summary = await summarizeConversation(conv.exchanges, sessionIdForSummary(conv.exchanges), summarizeOptions(redactor)); fs.writeFileSync(conv.summaryPath, summary, 'utf-8'); const wordCount = summary.split(/\s+/).length; console.log(` ✓ ${conv.file}: ${wordCount} words`); @@ -145,6 +157,7 @@ export async function indexConversations(limitToProject, maxConversations, concu // Check if we hit the limit if (maxConversations && conversationsProcessed >= maxConversations) { console.log(`\nReached limit of ${maxConversations} conversations`); + logRedactions(tally); db.close(); console.log(`✅ Indexing complete! Conversations: ${conversationsProcessed}, Exchanges: ${totalExchanges}`); return; @@ -155,11 +168,14 @@ export async function indexConversations(limitToProject, maxConversations, concu if (oversizeSkipped > 0) { console.log(` Skipped ${oversizeSkipped} oversize exchange(s) (> ${maxMessageBytes} bytes; set EPISODIC_MEMORY_MAX_MESSAGE_BYTES to change) — likely embedded-transcript payloads (#139)`); } + logRedactions(tally); db.close(); console.log(`\n✅ Indexing complete! Conversations: ${conversationsProcessed}, Exchanges: ${totalExchanges}`); } export async function indexSession(sessionId, concurrency = 1, noSummaries = false) { console.log(`Indexing session: ${sessionId}`); + const redactor = loadRedactor(); + const tally = new FindingsTally(); // Find the conversation file for this session const sourceDirs = getConversationSourceDirs(); const ARCHIVE_DIR = getArchiveDir(); @@ -187,11 +203,8 @@ export async function indexSession(sessionId, concurrency = 1, noSummaries = fal // Archive + parse — source may vanish mid-run (Claude Code cleanup). let exchanges; try { - if (!fs.existsSync(archivePath)) { - fs.mkdirSync(path.dirname(archivePath), { recursive: true }); - fs.copyFileSync(sourcePath, archivePath); - } - exchanges = await parseConversation(sourcePath, project, archivePath); + copyIfNewer(sourcePath, archivePath, redactor, tally); + exchanges = await parseConversation(archivePath, project, archivePath); } catch (error) { console.log(`Skipped ${file} (read failed: ${error instanceof Error ? error.message : error})`); @@ -204,7 +217,7 @@ export async function indexSession(sessionId, concurrency = 1, noSummaries = fal if (!noSummaries && shouldQueueForSummary(summaryPath)) { fs.mkdirSync(path.dirname(summaryPath), { recursive: true }); try { - const summary = await summarizeConversation(exchanges, sessionIdForSummary(exchanges)); + const summary = await summarizeConversation(exchanges, sessionIdForSummary(exchanges), summarizeOptions(redactor)); fs.writeFileSync(summaryPath, summary, 'utf-8'); console.log(`Summary: ${summary.split(/\s+/).length} words`); } @@ -235,6 +248,7 @@ export async function indexSession(sessionId, concurrency = 1, noSummaries = fal if (oversizeSkipped > 0) { console.log(` Skipped ${oversizeSkipped} oversize exchange(s) (> ${maxMessageBytes} bytes; set EPISODIC_MEMORY_MAX_MESSAGE_BYTES to change) — likely embedded-transcript payloads (#139)`); } + logRedactions(tally); console.log(`✅ Indexed session ${sessionId}: ${exchanges.length} exchanges`); } db.close(); @@ -254,6 +268,8 @@ export async function indexUnprocessed(concurrency = 1, noSummaries = false) { console.log(`Concurrency: ${concurrency}`); if (noSummaries) console.log('⚠️ Running in no-summaries mode (skipping AI summaries)'); + const redactor = loadRedactor(); + const tally = new FindingsTally(); const db = initDatabase(); await initEmbeddings(); const sourceDirs = getConversationSourceDirs(); @@ -280,15 +296,12 @@ export async function indexUnprocessed(concurrency = 1, noSummaries = false) { // Transcript JSONLs are append-only, so MAX(line_end) tells us where to resume. const hw = db.prepare('SELECT COALESCE(MAX(line_end), 0) as maxLine FROM exchanges WHERE archive_path = ?').get(archivePath); const maxIndexedLine = hw.maxLine; - // Ensure parent dirs exist for subagent files try { - fs.mkdirSync(path.dirname(archivePath), { recursive: true }); - // Refresh the archive when the source may have grown beyond what we've seen. - if (!fs.existsSync(archivePath) || maxIndexedLine > 0) { - fs.copyFileSync(sourcePath, archivePath); - } + // Refresh the (redacted) archive when the source has grown, then parse + // the archive so the index only sees redacted text. + copyIfNewer(sourcePath, archivePath, redactor, tally); // Parse and filter to exchanges past the high-water mark - const exchanges = await parseConversation(sourcePath, project, archivePath); + const exchanges = await parseConversation(archivePath, project, archivePath); const newExchanges = maxIndexedLine > 0 ? exchanges.filter(e => e.lineStart > maxIndexedLine) : exchanges; @@ -303,6 +316,7 @@ export async function indexUnprocessed(concurrency = 1, noSummaries = false) { } } } // end sourceDir loop + logRedactions(tally); if (unprocessed.length === 0) { console.log('✅ All conversations are already processed!'); db.close(); @@ -316,7 +330,7 @@ export async function indexUnprocessed(concurrency = 1, noSummaries = false) { console.log(`Generating ${needsSummary.length} summaries (concurrency: ${concurrency})...\n`); await processBatch(needsSummary, async (conv) => { try { - const summary = await summarizeConversation(conv.exchanges, sessionIdForSummary(conv.exchanges)); + const summary = await summarizeConversation(conv.exchanges, sessionIdForSummary(conv.exchanges), summarizeOptions(redactor)); fs.writeFileSync(conv.summaryPath, summary, 'utf-8'); const wordCount = summary.split(/\s+/).length; console.log(` ✓ ${conv.project}/${conv.file}: ${wordCount} words`); diff --git a/dist/opencode-sync.d.ts b/dist/opencode-sync.d.ts index fff32c4c..64215225 100644 --- a/dist/opencode-sync.d.ts +++ b/dist/opencode-sync.d.ts @@ -1,3 +1,4 @@ +import { type Redactor } from './redaction.js'; export interface OpencodeExportResult { exported: number; skipped: number; @@ -15,4 +16,6 @@ export declare function getOpencodeTranscriptFilePath(transcriptDir: string, inp export declare function exportOpencodeSessions(options?: { dbPath?: string; transcriptDir?: string; + /** `undefined` loads from the environment (strict mode throws on bad rules); `null` disables. */ + redactor?: Redactor | null; }): OpencodeExportResult; diff --git a/dist/opencode-sync.js b/dist/opencode-sync.js index 2f90ab75..c4bd4a8b 100644 --- a/dist/opencode-sync.js +++ b/dist/opencode-sync.js @@ -2,6 +2,7 @@ import fs from 'fs'; import path from 'path'; import Database from 'better-sqlite3'; import { getOpencodeDbPath, getOpencodeTranscriptDir } from './paths.js'; +import { loadRedactor, redactJsonlLine } from './redaction.js'; function safeParseJson(value) { if (!value) return undefined; @@ -40,7 +41,7 @@ function shouldExportSession(filePath, sessionUpdatedMs) { const stat = fs.statSync(filePath); return Math.floor(stat.mtimeMs) < Math.floor(sessionUpdatedMs); } -function writeSessionTranscript(db, session, filePath) { +function writeSessionTranscript(db, session, filePath, redactor) { const messages = db.prepare(` SELECT id, session_id, time_created, time_updated, data FROM message @@ -115,14 +116,20 @@ function writeSessionTranscript(db, session, filePath) { parts, })); } + // The staging transcript is a plugin-owned plaintext copy, so redact it at + // write time like the archive (docs/redaction/PHASE0-FINDINGS.md). + const output = redactor + ? lines.map(line => redactJsonlLine(line, redactor, { source: 'opencode', path: filePath })) + : lines; fs.mkdirSync(path.dirname(filePath), { recursive: true }); const tempPath = `${filePath}.tmp.${process.pid}`; - fs.writeFileSync(tempPath, `${lines.join('\n')}\n`, 'utf-8'); + fs.writeFileSync(tempPath, `${output.join('\n')}\n`, 'utf-8'); fs.renameSync(tempPath, filePath); const mtime = dateFromMillis(session.time_updated); fs.utimesSync(filePath, mtime, mtime); } export function exportOpencodeSessions(options = {}) { + const redactor = options.redactor === undefined ? loadRedactor() : options.redactor; const dbPath = options.dbPath || getOpencodeDbPath(); const transcriptDir = options.transcriptDir || getOpencodeTranscriptDir(); const result = { @@ -168,7 +175,7 @@ export function exportOpencodeSessions(options = {}) { result.skipped++; continue; } - writeSessionTranscript(db, session, filePath); + writeSessionTranscript(db, session, filePath, redactor); result.exported++; } catch (error) { diff --git a/dist/redact-cli.d.ts b/dist/redact-cli.d.ts new file mode 100644 index 00000000..cb0ff5c3 --- /dev/null +++ b/dist/redact-cli.d.ts @@ -0,0 +1 @@ +export {}; diff --git a/dist/redact-cli.js b/dist/redact-cli.js new file mode 100644 index 00000000..97044bb2 --- /dev/null +++ b/dist/redact-cli.js @@ -0,0 +1,123 @@ +import fs from 'fs'; +import { getArchiveDir, getCursorLegacyExportDir, getOpencodeTranscriptDir } from './paths.js'; +import { acquireFileLock, readLockHolder, releaseFileLock } from './file-lock.js'; +import { getSyncLockPath } from './logging.js'; +import { DEFAULT_REDACTION_CONFIG, FindingsTally, formatFindings, getRedactionSettings, loadRedactor, redactJsonlLine, } from './redaction.js'; +const args = process.argv.slice(2); +const HELP = ` +Usage: episodic-memory redact [--rewrite [--dry-run]] [--stdin] [--print-default-rules] + +Secret redaction for the conversation archive and index. + +OPTIONS: + --rewrite Re-run redaction over the existing archive, staging exports, + and search index (in place). Rows that change are re-embedded, + and summaries built from unredacted text are deleted (the next + sync regenerates them). Run once after upgrading, and again + after adding rules. + --dry-run With --rewrite: report what would change, write nothing. + --stdin Redact stdin to stdout, one JSONL/text line at a time, and + print rule counts to stderr. Handy for testing rules. + --print-default-rules Print the bundled rules as JSON (a starting point for + redaction-rules.json). + --help, -h Show this help + +ENVIRONMENT: + EPISODIC_MEMORY_REDACTION on (default) | off + EPISODIC_MEMORY_REDACTION_RULES custom rules file (default: /redaction-rules.json) + EPISODIC_MEMORY_REDACTION_STRICT 1 (default) fails closed on a bad rules file | 0 passes through + +Output names rule IDs and counts only; matched values are never printed. +`; +function fail(message) { + console.error(`episodic-memory: ${message}`); + process.exit(1); +} +function requireRedactor() { + if (!getRedactionSettings().enabled) { + fail('redaction is off (EPISODIC_MEMORY_REDACTION=off); unset it to use this command.'); + } + let redactor; + try { + // Always strict here: a rewrite with no rules would be a silent no-op. + redactor = loadRedactor({ ...process.env, EPISODIC_MEMORY_REDACTION_STRICT: '1' }); + } + catch (error) { + fail(error instanceof Error ? error.message : String(error)); + } + return redactor; +} +async function runRewrite(dryRun) { + const redactor = requireRedactor(); + // Share sync's single-instance lock so a background sync can't write + // between our reads and renames. + const lockPath = getSyncLockPath(); + const lock = acquireFileLock(lockPath); + if (!lock) { + const holder = readLockHolder(lockPath); + fail(`a sync or index run is in progress (${holder !== null ? `pid ${holder}` : 'another process'}); try again when it finishes.`); + } + const release = () => releaseFileLock(lock); + process.on('exit', release); + const { rewriteArchive } = await import('./redact-rewrite.js'); + let embeddingsReady = false; + const embed = async (user, assistant, toolNames) => { + const embeddings = await import('./embeddings.js'); + if (!embeddingsReady) { + await embeddings.initEmbeddings(); + embeddingsReady = true; + } + return embeddings.generateExchangeEmbedding(user, assistant, toolNames); + }; + const archiveDir = getArchiveDir(); + const stagingDirs = [getOpencodeTranscriptDir(), getCursorLegacyExportDir()].filter(d => fs.existsSync(d)); + console.log(`Redacting${dryRun ? ' (dry run)' : ''}: ${archiveDir}`); + for (const dir of stagingDirs) + console.log(` + staging: ${dir}`); + const result = await rewriteArchive({ + archiveDir, + stagingDirs, + redactor, + embed, + dryRun, + log: message => console.log(` ${message}`), + }); + console.log(`\n${dryRun ? 'Would redact' : 'Redacted'} across archive, staging, and index: ${formatFindings(result.findings)}`); + if (dryRun && (result.filesRewritten || result.rowsUpdated || result.summariesRemoved || result.stagingFilesRewritten)) { + console.log('Run without --dry-run to apply.'); + } +} +async function runStdin() { + const redactor = requireRedactor(); + const chunks = []; + for await (const chunk of process.stdin) + chunks.push(chunk); + const input = Buffer.concat(chunks).toString('utf-8'); + const tally = new FindingsTally(); + const output = input.split('\n').map(line => redactJsonlLine(line, redactor, { source: 'stdin', path: '-' }, tally)); + process.stdout.write(output.join('\n')); + console.error(formatFindings(tally.toArray())); +} +async function main() { + if (args.length === 0 || args.includes('--help') || args.includes('-h')) { + console.log(HELP); + return; + } + if (args.includes('--print-default-rules')) { + console.log(JSON.stringify(DEFAULT_REDACTION_CONFIG, null, 2)); + return; + } + if (args.includes('--stdin')) { + await runStdin(); + return; + } + if (args.includes('--rewrite')) { + await runRewrite(args.includes('--dry-run')); + return; + } + fail(`unknown option(s): ${args.join(' ')}. Try: episodic-memory redact --help`); +} +main().catch(error => { + console.error('Error:', error instanceof Error ? error.message : error); + process.exit(1); +}); diff --git a/dist/redact-rewrite.d.ts b/dist/redact-rewrite.d.ts new file mode 100644 index 00000000..0900d8e4 --- /dev/null +++ b/dist/redact-rewrite.d.ts @@ -0,0 +1,40 @@ +import { type RedactionFinding, type Redactor } from './redaction.js'; +/** + * Backfill for `episodic-memory redact --rewrite`: re-run redaction over data + * written before redaction existed (or before a rule was added). + * + * 1. Archive: every .jsonl is redacted in place. Line count and mtime are + * preserved, so index line ranges stay valid and sync still sees the + * archive as current. + * 2. Staging dirs (opencode / legacy Cursor exports): same, in place. + * 3. Index: every exchanges/tool_calls row is redacted in place. Changed rows + * are re-embedded from the redacted text, so no vector is left that was + * derived from a secret. + * 4. Summaries: a `-summary.txt` for a conversation that had findings (in the + * archive or the index), or one that matches a rule itself, is deleted. + * It was generated from unredacted text, and the next sync regenerates it + * from the redacted archive. + * + * Idempotent: a second run finds nothing to change. + */ +export type EmbedFn = (user: string, assistant: string, toolNames?: string[]) => Promise; +export interface RewriteOptions { + archiveDir: string; + redactor: Redactor; + embed: EmbedFn; + /** Count what would change without writing anything. */ + dryRun?: boolean; + /** Plugin-owned staging dirs to redact in place (not indexed). */ + stagingDirs?: string[]; + log?: (message: string) => void; +} +export interface RewriteResult { + filesScanned: number; + filesRewritten: number; + stagingFilesRewritten: number; + rowsUpdated: number; + summariesRemoved: number; + /** Rule IDs and counts only — never matched values. */ + findings: RedactionFinding[]; +} +export declare function rewriteArchive(options: RewriteOptions): Promise; diff --git a/dist/redact-rewrite.js b/dist/redact-rewrite.js new file mode 100644 index 00000000..7c0317f6 --- /dev/null +++ b/dist/redact-rewrite.js @@ -0,0 +1,170 @@ +import fs from 'fs'; +import path from 'path'; +import { initDatabase } from './db.js'; +import { recordReembedded } from './embedding-migration.js'; +import { copyFileRedacted, FindingsTally, redactJsonlLine, } from './redaction.js'; +const SUMMARY_SUFFIX = '-summary.txt'; +const PAGE_SIZE = 500; +function walk(dir) { + const out = []; + let entries; + try { + entries = fs.readdirSync(dir, { withFileTypes: true }); + } + catch { + return out; + } + for (const entry of entries) { + const full = path.join(dir, entry.name); + if (entry.isDirectory()) + out.push(...walk(full)); + else if (entry.isFile()) + out.push(full); + } + return out; +} +function summaryPathFor(jsonlPath) { + return jsonlPath.replace(/\.jsonl$/, SUMMARY_SUFFIX); +} +/** Redact one JSONL file in place. Returns the number of values redacted. */ +function rewriteFileInPlace(file, redactor, dryRun, tally) { + const fileTally = new FindingsTally(); + const temp = `${file}.redact.${process.pid}`; + try { + copyFileRedacted(file, temp, redactor, { source: 'rewrite', path: file }, fileTally); + if (fileTally.total > 0 && !dryRun) { + const stat = fs.statSync(file); + fs.renameSync(temp, file); + // Same rounding as sync's copyIfNewer: never leave the archive older than its source. + fs.utimesSync(file, stat.atimeMs / 1000, Math.ceil(stat.mtimeMs) / 1000); + } + } + finally { + try { + fs.unlinkSync(temp); + } + catch { } + } + tally.add(fileTally.toArray()); + return fileTally.total; +} +export async function rewriteArchive(options) { + const { archiveDir, redactor, embed } = options; + const dryRun = options.dryRun === true; + const log = options.log ?? (() => { }); + const tally = new FindingsTally(); + const result = { + filesScanned: 0, + filesRewritten: 0, + stagingFilesRewritten: 0, + rowsUpdated: 0, + summariesRemoved: 0, + findings: [], + }; + // Conversations whose summary was built from unredacted text. + const staleSummaries = new Set(); + // 1. Archive files. + const archiveFiles = walk(archiveDir); + for (const file of archiveFiles.filter(f => f.endsWith('.jsonl'))) { + result.filesScanned++; + if (rewriteFileInPlace(file, redactor, dryRun, tally) > 0) { + result.filesRewritten++; + staleSummaries.add(summaryPathFor(file)); + } + } + log(`Archive: ${result.filesRewritten} of ${result.filesScanned} file(s) ${dryRun ? 'would be ' : ''}rewritten`); + // 2. Staging dirs. + for (const dir of options.stagingDirs ?? []) { + for (const file of walk(dir).filter(f => f.endsWith('.jsonl'))) { + if (rewriteFileInPlace(file, redactor, dryRun, tally) > 0) + result.stagingFilesRewritten++; + } + } + if (options.stagingDirs?.length) { + log(`Staging exports: ${result.stagingFilesRewritten} file(s) ${dryRun ? 'would be ' : ''}rewritten`); + } + // 3. Index rows. Page by rowid so writes between pages don't disturb the scan. + const db = initDatabase(); + try { + const page = db.prepare('SELECT rowid AS rid, id, user_message, assistant_message, archive_path FROM exchanges WHERE rowid > ? ORDER BY rowid LIMIT ?'); + const toolsFor = db.prepare('SELECT id, tool_name, tool_input, tool_result FROM tool_calls WHERE exchange_id = ? ORDER BY rowid'); + const updateExchange = db.prepare('UPDATE exchanges SET user_message = ?, assistant_message = ? WHERE id = ?'); + const updateTool = db.prepare('UPDATE tool_calls SET tool_input = ?, tool_result = ? WHERE id = ?'); + let lastRowid = 0; + for (;;) { + const rows = page.all(lastRowid, PAGE_SIZE); + if (rows.length === 0) + break; + lastRowid = rows[rows.length - 1].rid; + for (const row of rows) { + const rowTally = new FindingsTally(); + const ctx = { source: 'index', path: row.archive_path }; + const user = redactor.redact(row.user_message, ctx); + const assistant = redactor.redact(row.assistant_message, ctx); + rowTally.add(user.findings); + rowTally.add(assistant.findings); + const tools = toolsFor.all(row.id); + const toolUpdates = []; + for (const tool of tools) { + // tool_input is JSON text; redactJsonlLine keeps it valid JSON. + const input = tool.tool_input === null ? null : redactJsonlLine(tool.tool_input, redactor, ctx, rowTally); + let output = tool.tool_result; + if (output !== null) { + const r = redactor.redact(output, ctx); + rowTally.add(r.findings); + output = r.text; + } + if (input !== tool.tool_input || output !== tool.tool_result) { + toolUpdates.push({ id: tool.id, input, result: output }); + } + } + if (rowTally.total === 0) + continue; + tally.add(rowTally.toArray()); + result.rowsUpdated++; + staleSummaries.add(summaryPathFor(row.archive_path)); + if (dryRun) + continue; + const toolNames = tools.length > 0 ? tools.map(t => t.tool_name) : undefined; + const embedding = await embed(user.text, assistant.text, toolNames); + db.transaction(() => { + updateExchange.run(user.text, assistant.text, row.id); + for (const t of toolUpdates) + updateTool.run(t.input, t.result, t.id); + recordReembedded(db, row.id, embedding); + })(); + } + } + } + finally { + db.close(); + } + log(`Index: ${result.rowsUpdated} exchange(s) ${dryRun ? 'would be ' : ''}redacted and re-embedded`); + // 4. Summaries: stale ones, plus any summary that itself matches a rule. + for (const file of archiveFiles.filter(f => f.endsWith(SUMMARY_SUFFIX))) { + if (staleSummaries.has(file)) + continue; + let text; + try { + text = fs.readFileSync(file, 'utf-8'); + } + catch { + continue; + } + const r = redactor.redact(text, { source: 'summary', path: file }); + if (r.findings.length > 0) { + tally.add(r.findings); + staleSummaries.add(file); + } + } + for (const summary of staleSummaries) { + if (!fs.existsSync(summary)) + continue; + result.summariesRemoved++; + if (!dryRun) + fs.unlinkSync(summary); + } + log(`Summaries: ${result.summariesRemoved} ${dryRun ? 'would be ' : ''}removed (regenerated from redacted text on the next sync)`); + result.findings = tally.toArray(); + return result; +} diff --git a/dist/redaction-rules.d.ts b/dist/redaction-rules.d.ts new file mode 100644 index 00000000..a9fa089b --- /dev/null +++ b/dist/redaction-rules.d.ts @@ -0,0 +1,27 @@ +import type { RedactionConfig } from './redaction.js'; +/** + * Bundled default redaction rules. + * + * Patterns are ported from gitleaks' rule set (https://github.com/gitleaks/gitleaks, + * MIT), trimmed to the credentials that realistically show up in Claude Code / + * Codex transcripts for an Azure-heavy stack, and rewritten for JavaScript + * regex (lookbehind instead of gitleaks' consuming boundary groups, so + * adjacent matches aren't swallowed). + * + * Order matters: rules run top to bottom, and a later rule never re-matches + * inside an earlier rule's `[REDACTED:...]` token. Specific shapes go first so + * the token names the most precise rule; the generic `secret-assignment` + * keyword rule goes last. + * + * `secretGroup` replaces only that capture group, so surrounding context + * (connection-string server names, SAS URL paths, usernames) stays searchable. + * + * `keywords` is a case-insensitive prefilter: a rule only runs on text that + * contains at least one keyword. It keeps the per-string cost low on large + * transcripts and bounds the generic rules' work. + * + * Users extend or override these with `redaction-rules.json`; see + * docs/REDACTION.md. `episodic-memory redact --print-default-rules` dumps + * this object as JSON. + */ +export declare const DEFAULT_REDACTION_CONFIG: RedactionConfig; diff --git a/dist/redaction-rules.js b/dist/redaction-rules.js new file mode 100644 index 00000000..5adc1c46 --- /dev/null +++ b/dist/redaction-rules.js @@ -0,0 +1,164 @@ +/** + * Bundled default redaction rules. + * + * Patterns are ported from gitleaks' rule set (https://github.com/gitleaks/gitleaks, + * MIT), trimmed to the credentials that realistically show up in Claude Code / + * Codex transcripts for an Azure-heavy stack, and rewritten for JavaScript + * regex (lookbehind instead of gitleaks' consuming boundary groups, so + * adjacent matches aren't swallowed). + * + * Order matters: rules run top to bottom, and a later rule never re-matches + * inside an earlier rule's `[REDACTED:...]` token. Specific shapes go first so + * the token names the most precise rule; the generic `secret-assignment` + * keyword rule goes last. + * + * `secretGroup` replaces only that capture group, so surrounding context + * (connection-string server names, SAS URL paths, usernames) stays searchable. + * + * `keywords` is a case-insensitive prefilter: a rule only runs on text that + * contains at least one keyword. It keeps the per-string cost low on large + * transcripts and bounds the generic rules' work. + * + * Users extend or override these with `redaction-rules.json`; see + * docs/REDACTION.md. `episodic-memory redact --print-default-rules` dumps + * this object as JSON. + */ +export const DEFAULT_REDACTION_CONFIG = { + rules: [ + { + id: 'private-key-block', + description: 'PEM/OpenSSH private key block; a truncated block (no END line) is redacted to the end of the text', + pattern: String.raw `-----BEGIN[A-Z0-9 _-]{0,40}PRIVATE KEY(?: BLOCK)?-----[\s\S]*?(?:-----END[A-Z0-9 _-]{0,40}PRIVATE KEY(?: BLOCK)?-----|$)`, + keywords: ['private key'], + }, + { + id: 'connection-string-secret', + description: 'Secret value inside an Azure/ADO.NET connection string; account, server and database names are kept', + pattern: String.raw `\b(?:AccountKey|SharedAccessKey|SharedAccessSignature|SharedSecret|ClientSecret|Password|Pwd)[ \t]*=[ \t]*(?![\[$<{%])([^;"'\s]+)`, + secretGroup: 1, + keywords: ['accountkey', 'sharedaccess', 'sharedsecret', 'clientsecret', 'password', 'pwd'], + }, + { + id: 'azure-sas-token', + description: 'Signature of an Azure SAS token; the resource URL and other SAS parameters are kept', + pattern: String.raw `\bsig=([A-Za-z0-9%+/=_-]{16,})`, + secretGroup: 1, + keywords: ['sig='], + }, + { + id: 'jwt', + description: 'JSON Web Token (Entra ID / Azure access tokens, id tokens)', + pattern: String.raw `\beyJ[A-Za-z0-9_-]{8,}\.eyJ[A-Za-z0-9_-]{8,}\.[A-Za-z0-9_-]{8,}`, + keywords: ['eyj'], + }, + { + id: 'anthropic-api-key', + pattern: String.raw `\bsk-ant-[a-z]{2,10}\d{2}-[A-Za-z0-9_-]{32,}`, + keywords: ['sk-ant-'], + }, + { + id: 'openai-api-key', + pattern: String.raw `\bsk-(?:[a-z]+-)?[A-Za-z0-9_-]{16,}T3BlbkFJ[A-Za-z0-9_-]{16,}`, + keywords: ['t3blbkfj'], + }, + { + id: 'github-token', + description: 'GitHub classic, OAuth, app and fine-grained tokens', + pattern: String.raw `\b(?:gh[pousr]_[A-Za-z0-9]{36,255}|github_pat_[A-Za-z0-9_]{22,255})\b`, + keywords: ['ghp_', 'gho_', 'ghu_', 'ghs_', 'ghr_', 'github_pat_'], + }, + { + id: 'aws-access-key-id', + pattern: String.raw `\b(?:AKIA|ASIA|ABIA|ACCA)[A-Z2-7]{16}\b`, + keywords: ['akia', 'asia', 'abia', 'acca'], + }, + { + id: 'aws-secret-access-key', + pattern: String.raw `aws[_-]?secret[_-]?access[_-]?key["']?[ \t]*[:=][ \t]*["']?([A-Za-z0-9/+]{40})(?![A-Za-z0-9/+])`, + flags: 'i', + secretGroup: 1, + keywords: ['secret'], + }, + { + id: 'slack-token', + pattern: String.raw `\bxox[abposr]-[A-Za-z0-9-]{10,}`, + keywords: ['xox'], + }, + { + id: 'google-api-key', + pattern: String.raw `\bAIza[0-9A-Za-z_-]{35}(?![0-9A-Za-z_-])`, + keywords: ['aiza'], + }, + { + id: 'npm-token', + pattern: String.raw `\bnpm_[A-Za-z0-9]{36}\b`, + keywords: ['npm_'], + }, + { + id: 'azure-client-secret', + description: 'Entra ID (Azure AD) application client secret: 3 chars, a digit, "Q~", 31-34 chars', + pattern: String.raw `(?]+:(?![\[$<{%])([^\s@/"'<>]+)@`, + secretGroup: 1, + keywords: ['://'], + }, + { + id: 'bearer-token', + pattern: String.raw `\bBearer\s+([A-Za-z0-9\-._~+/]{20,}=*)`, + flags: 'i', + secretGroup: 1, + keywords: ['bearer'], + }, + { + id: 'basic-auth', + pattern: String.raw `\bAuthorization[ \t]*:[ \t]*Basic\s+([A-Za-z0-9+/]{8,}={0,2})`, + flags: 'i', + secretGroup: 1, + keywords: ['basic'], + }, + { + id: 'secret-assignment', + description: 'Value assigned to a secret-looking key (password: x, CLIENT_SECRET=x, "apiKey": "x"). ' + + 'Covers decrypted SOPS/YAML/dotenv/JSON. Requires 8+ chars including a digit; ' + + 'skips placeholders (${X}, , %X%, {{x}}) and code (calls, dotted member access)', + pattern: String.raw `\b[A-Za-z0-9_.-]{0,40}?(?:secret|password|passwd|passphrase|api[_-]?key|access[_-]?key|private[_-]?key|token)` + + String.raw `["']?[ \t]*[:=][ \t]*["']?(?![\[$<{%])(?=[^\s"',;&\\]*\d)` + + // Not a dotted identifier path (env.AUTH0_SECRET, this.config.token2): code, not a value. + String.raw `(?![A-Za-z_$][\w$]*(?:\.[A-Za-z_$][\w$]*)+(?:$|[\s"',;&\\)}\]]))` + + String.raw `([^\s"',;&\\()<>{}\[\]]{8,})(?=$|[\s"',;&\\)}\]])`, + flags: 'i', + secretGroup: 1, + keywords: ['secret', 'passw', 'passphrase', 'key', 'token'], + }, + ], + allowlist: [ + { + id: 'git-sha', + description: 'Full or short git commit SHA', + pattern: String.raw `\b[0-9a-f]{7,40}\b`, + }, + { + id: 'guid', + description: 'GUID/UUID (Entra tenant, client and object IDs, subscription IDs)', + pattern: String.raw `\b[0-9a-fA-F]{8}-[0-9a-fA-F]{4}-[0-9a-fA-F]{4}-[0-9a-fA-F]{4}-[0-9a-fA-F]{12}\b`, + }, + ], + entropy: { + enabled: false, + minLength: 32, + threshold: 4.5, + requireKeyword: true, + keywords: ['secret', 'key', 'token', 'password', 'passwd', 'credential', 'signature'], + window: 40, + }, +}; diff --git a/dist/redaction.d.ts b/dist/redaction.d.ts new file mode 100644 index 00000000..d45dcd6b --- /dev/null +++ b/dist/redaction.d.ts @@ -0,0 +1,137 @@ +import { DEFAULT_REDACTION_CONFIG } from './redaction-rules.js'; +/** + * Inline secret redaction. + * + * Secrets are replaced with typed tokens (`[REDACTED:]`) at the archive + * write, the one point every harness's transcripts pass through (see + * docs/redaction/PHASE0-FINDINGS.md). Everything downstream (SQLite text, + * tool_calls, embeddings, summaries, show/read) reads the archive, so it only + * ever sees redacted text. + * + * Invariants: + * - Findings carry rule IDs and counts only. A matched value is never logged, + * returned, or thrown. + * - Idempotent: a `[REDACTED:...]` token never matches a rule, so redacting + * redacted text is a no-op. + * - JSONL structure is preserved: one output line per input line, valid JSON in + * gives valid JSON out, and unchanged lines are byte-identical. + */ +export { DEFAULT_REDACTION_CONFIG }; +/** Fork default. Flip to false for an opt-in upstream build. */ +export declare const REDACTION_ENABLED_BY_DEFAULT = true; +export declare const REDACTION_RULES_FILENAME = "redaction-rules.json"; +export interface RedactionRuleSpec { + id: string; + pattern: string; + /** Extra RegExp flags from [imsu]; `g` and `d` are always added. */ + flags?: string; + /** Case-insensitive prefilter: skip the rule unless the text contains one. */ + keywords?: string[]; + /** Replace only this capture group instead of the whole match. */ + secretGroup?: number; + description?: string; +} +export interface AllowlistSpec { + id: string; + /** Matched against the whole candidate secret; a full match keeps it. */ + pattern: string; + flags?: string; + description?: string; +} +export interface EntropySpec { + enabled: boolean; + minLength: number; + /** Shannon entropy in bits per character. */ + threshold: number; + /** Only fire when a keyword appears shortly before the candidate, on the same line. */ + requireKeyword: boolean; + keywords: string[]; + /** How many characters before the candidate to search for a keyword. */ + window: number; +} +export interface RedactionConfig { + rules: RedactionRuleSpec[]; + allowlist: AllowlistSpec[]; + entropy: EntropySpec; +} +/** Shape of a user `redaction-rules.json`. Every field is optional. */ +export interface RedactionRulesFile { + /** Merge with the bundled defaults (default true). */ + includeDefaults?: boolean; + /** Added after the defaults; a rule with a default's id replaces it in place. */ + rules?: RedactionRuleSpec[]; + /** Default rule ids to turn off. */ + disableRules?: string[]; + /** Added to the default allowlist (same id replaces). */ + allowlist?: AllowlistSpec[]; + entropy?: Partial; +} +export interface RedactionContext { + source: string; + path: string; +} +export interface RedactionFinding { + ruleId: string; + count: number; +} +export interface RedactionResult { + text: string; + findings: RedactionFinding[]; +} +export interface Redactor { + redact(text: string, ctx?: RedactionContext): RedactionResult; + readonly ruleIds: string[]; +} +export interface RedactionSettings { + enabled: boolean; + strict: boolean; + /** Explicit EPISODIC_MEMORY_REDACTION_RULES path, if set. */ + rulesPath?: string; +} +/** Rules could not be loaded or are invalid. Messages describe config only. */ +export declare class RedactionConfigError extends Error { + constructor(message: string); +} +/** + * EPISODIC_MEMORY_REDACTION on (fork default) | off + * EPISODIC_MEMORY_REDACTION_RULES path to a custom rules file + * EPISODIC_MEMORY_REDACTION_STRICT 1 (default) | 0; strict fails closed on a rules-load error + */ +export declare function getRedactionSettings(env?: NodeJS.ProcessEnv): RedactionSettings; +/** + * Merge a rules file with the bundled defaults and validate the result. + * Throws RedactionConfigError on any invalid rule. + */ +export declare function loadRedactionConfig(file: RedactionRulesFile): RedactionConfig; +/** + * Load the active redactor from the environment. + * + * Returns null when redaction is off. When the rules can't be loaded: strict + * mode (the default) throws RedactionConfigError so callers fail closed; + * non-strict mode warns and returns null, so text passes through unredacted. + */ +export declare function loadRedactor(env?: NodeJS.ProcessEnv, warn?: (message: string) => void): Redactor | null; +export declare function createRedactor(config: RedactionConfig): Redactor; +/** Aggregates findings by rule id. Never holds matched values. */ +export declare class FindingsTally { + private counts; + add(findings: RedactionFinding[]): void; + get total(): number; + toArray(): RedactionFinding[]; +} +/** "7 value(s) redacted (azure-storage-key: 2, jwt: 5)" */ +export declare function formatFindings(findings: RedactionFinding[]): string; +/** + * Redact one JSONL line. Valid JSON stays valid (string values are redacted + * and the line is re-serialized only if something changed); a line that isn't + * JSON (e.g. a torn last line mid-write) is redacted as raw text. A trailing + * `\r` is preserved. + */ +export declare function redactJsonlLine(line: string, redactor: Redactor, ctx?: RedactionContext, tally?: FindingsTally): string; +/** + * Stream `src` to `dest`, redacting line by line. Writes `dest` directly; + * callers that need atomicity write to a temp path and rename. Line count and + * trailing-newline shape match the source exactly, so archive line numbers + * (exchanges.line_start/line_end, MCP read ranges) stay valid. + */ +export declare function copyFileRedacted(src: string, dest: string, redactor: Redactor, ctx?: RedactionContext, tally?: FindingsTally): void; diff --git a/dist/redaction.js b/dist/redaction.js new file mode 100644 index 00000000..c4190658 --- /dev/null +++ b/dist/redaction.js @@ -0,0 +1,484 @@ +import fs from 'fs'; +import path from 'path'; +import { StringDecoder } from 'string_decoder'; +import { getSuperpowersDir } from './paths.js'; +import { DEFAULT_REDACTION_CONFIG } from './redaction-rules.js'; +/** + * Inline secret redaction. + * + * Secrets are replaced with typed tokens (`[REDACTED:]`) at the archive + * write, the one point every harness's transcripts pass through (see + * docs/redaction/PHASE0-FINDINGS.md). Everything downstream (SQLite text, + * tool_calls, embeddings, summaries, show/read) reads the archive, so it only + * ever sees redacted text. + * + * Invariants: + * - Findings carry rule IDs and counts only. A matched value is never logged, + * returned, or thrown. + * - Idempotent: a `[REDACTED:...]` token never matches a rule, so redacting + * redacted text is a no-op. + * - JSONL structure is preserved: one output line per input line, valid JSON in + * gives valid JSON out, and unchanged lines are byte-identical. + */ +export { DEFAULT_REDACTION_CONFIG }; +/** Fork default. Flip to false for an opt-in upstream build. */ +export const REDACTION_ENABLED_BY_DEFAULT = true; +export const REDACTION_RULES_FILENAME = 'redaction-rules.json'; +const RULE_ID_PATTERN = /^[a-z0-9][a-z0-9-]*$/; +const TOKEN_PATTERN = /\[REDACTED:[a-z0-9][a-z0-9-]*\]/g; +const TOKEN_PREFIX = '[REDACTED:'; +const ALLOWED_FLAGS = /^[imsu]*$/; +const ENTROPY_RULE_ID = 'high-entropy'; +/** Rules could not be loaded or are invalid. Messages describe config only. */ +export class RedactionConfigError extends Error { + constructor(message) { + super(message); + this.name = 'RedactionConfigError'; + } +} +// --------------------------------------------------------------------------- +// Settings and loading +// --------------------------------------------------------------------------- +const OFF_VALUES = new Set(['off', '0', 'false', 'no', 'disabled']); +const ON_VALUES = new Set(['on', '1', 'true', 'yes', 'enabled']); +/** Unknown values fall back to the default, which is the safe side for both switches. */ +function parseToggle(raw, fallback) { + const value = raw?.trim().toLowerCase(); + if (!value) + return fallback; + if (OFF_VALUES.has(value)) + return false; + if (ON_VALUES.has(value)) + return true; + return fallback; +} +/** + * EPISODIC_MEMORY_REDACTION on (fork default) | off + * EPISODIC_MEMORY_REDACTION_RULES path to a custom rules file + * EPISODIC_MEMORY_REDACTION_STRICT 1 (default) | 0; strict fails closed on a rules-load error + */ +export function getRedactionSettings(env = process.env) { + return { + enabled: parseToggle(env.EPISODIC_MEMORY_REDACTION, REDACTION_ENABLED_BY_DEFAULT), + strict: parseToggle(env.EPISODIC_MEMORY_REDACTION_STRICT, true), + rulesPath: env.EPISODIC_MEMORY_REDACTION_RULES || undefined, + }; +} +function readRulesFile(settings) { + let rulesPath = settings.rulesPath; + if (!rulesPath) { + const candidate = path.join(getSuperpowersDir(), REDACTION_RULES_FILENAME); + if (!fs.existsSync(candidate)) + return {}; + rulesPath = candidate; + } + let raw; + try { + raw = fs.readFileSync(rulesPath, 'utf-8'); + } + catch (error) { + throw new RedactionConfigError(`Redaction rules failed to load from ${rulesPath}: ${error instanceof Error ? error.message : String(error)}`); + } + try { + return JSON.parse(raw); + } + catch (error) { + throw new RedactionConfigError(`Redaction rules file ${rulesPath} is not valid JSON: ${error instanceof Error ? error.message : String(error)}`); + } +} +function replaceById(base, additions) { + const out = [...base]; + for (const item of additions) { + const at = out.findIndex(existing => existing.id === item.id); + if (at >= 0) + out[at] = item; + else + out.push(item); + } + return out; +} +/** + * Merge a rules file with the bundled defaults and validate the result. + * Throws RedactionConfigError on any invalid rule. + */ +export function loadRedactionConfig(file) { + if (!file || typeof file !== 'object' || Array.isArray(file)) { + throw new RedactionConfigError('Redaction rules file must contain a JSON object'); + } + for (const key of ['rules', 'allowlist', 'disableRules']) { + if (file[key] !== undefined && !Array.isArray(file[key])) { + throw new RedactionConfigError(`Redaction rules file: "${key}" must be an array`); + } + } + if (file.entropy !== undefined && (typeof file.entropy !== 'object' || file.entropy === null)) { + throw new RedactionConfigError('Redaction rules file: "entropy" must be an object'); + } + const includeDefaults = file.includeDefaults !== false; + const disabled = new Set(file.disableRules ?? []); + const baseRules = includeDefaults ? DEFAULT_REDACTION_CONFIG.rules : []; + const baseAllowlist = includeDefaults ? DEFAULT_REDACTION_CONFIG.allowlist : []; + const config = { + rules: replaceById(baseRules, file.rules ?? []).filter(rule => !disabled.has(rule.id)), + allowlist: replaceById(baseAllowlist, file.allowlist ?? []), + entropy: { ...DEFAULT_REDACTION_CONFIG.entropy, ...(file.entropy ?? {}) }, + }; + compileConfig(config); // validate + return config; +} +/** + * Load the active redactor from the environment. + * + * Returns null when redaction is off. When the rules can't be loaded: strict + * mode (the default) throws RedactionConfigError so callers fail closed; + * non-strict mode warns and returns null, so text passes through unredacted. + */ +export function loadRedactor(env = process.env, warn = message => console.error(message)) { + const settings = getRedactionSettings(env); + if (!settings.enabled) + return null; + try { + return createRedactor(loadRedactionConfig(readRulesFile(settings))); + } + catch (error) { + const err = error instanceof RedactionConfigError + ? error + : new RedactionConfigError(`Redaction rules failed to load: ${error instanceof Error ? error.message : String(error)}`); + if (settings.strict) + throw err; + warn(`episodic-memory: ${err.message}. EPISODIC_MEMORY_REDACTION_STRICT=0, so conversations ` + + 'will be archived and indexed WITHOUT redaction this run.'); + return null; + } +} +function checkFlags(flags, what) { + const value = flags ?? ''; + if (typeof value !== 'string' || !ALLOWED_FLAGS.test(value)) { + throw new RedactionConfigError(`${what}: flags must be a combination of "imsu"`); + } + return value; +} +function compileRule(spec) { + if (!spec || typeof spec !== 'object') { + throw new RedactionConfigError('Redaction rule must be an object'); + } + if (typeof spec.id !== 'string' || !RULE_ID_PATTERN.test(spec.id)) { + throw new RedactionConfigError(`Redaction rule id ${JSON.stringify(spec.id)} is invalid: use lowercase letters, digits and dashes`); + } + if (spec.id === ENTROPY_RULE_ID) { + throw new RedactionConfigError(`Redaction rule id "${ENTROPY_RULE_ID}" is reserved`); + } + const what = `Redaction rule "${spec.id}"`; + if (typeof spec.pattern !== 'string' || spec.pattern.length === 0) { + throw new RedactionConfigError(`${what}: pattern must be a non-empty string`); + } + const flags = checkFlags(spec.flags, what); + let regex; + let groupCount; + try { + regex = new RegExp(spec.pattern, flags + 'gd'); + groupCount = new RegExp(`(?:${spec.pattern})|`, flags).exec('').length - 1; + } + catch (error) { + throw new RedactionConfigError(`${what}: invalid regular expression (${error instanceof Error ? error.message : String(error)})`); + } + if (new RegExp(spec.pattern, flags).test('')) { + throw new RedactionConfigError(`${what}: pattern matches the empty string`); + } + const secretGroup = spec.secretGroup ?? 0; + if (!Number.isInteger(secretGroup) || secretGroup < 0 || secretGroup > groupCount) { + throw new RedactionConfigError(`${what}: secretGroup ${secretGroup} does not exist in the pattern (${groupCount} group(s))`); + } + let keywords; + if (spec.keywords !== undefined) { + if (!Array.isArray(spec.keywords) || spec.keywords.some(k => typeof k !== 'string' || k.length === 0)) { + throw new RedactionConfigError(`${what}: keywords must be an array of non-empty strings`); + } + keywords = spec.keywords.length > 0 ? spec.keywords.map(k => k.toLowerCase()) : undefined; + } + return { id: spec.id, regex, keywords, secretGroup }; +} +function compileAllowlist(spec) { + const what = `Redaction allowlist entry ${JSON.stringify(spec?.id)}`; + if (!spec || typeof spec.pattern !== 'string' || spec.pattern.length === 0) { + throw new RedactionConfigError(`${what}: pattern must be a non-empty string`); + } + const flags = checkFlags(spec.flags, what); + try { + return new RegExp(`^(?:${spec.pattern})$`, flags); + } + catch (error) { + throw new RedactionConfigError(`${what}: invalid regular expression (${error instanceof Error ? error.message : String(error)})`); + } +} +function compileConfig(config) { + const seen = new Set(); + const rules = config.rules.map(spec => { + const rule = compileRule(spec); + if (seen.has(rule.id)) + throw new RedactionConfigError(`Duplicate redaction rule id "${rule.id}"`); + seen.add(rule.id); + return rule; + }); + const allowlist = config.allowlist.map(compileAllowlist); + const e = config.entropy; + if (typeof e.enabled !== 'boolean' || typeof e.requireKeyword !== 'boolean' || + !Number.isInteger(e.minLength) || e.minLength < 8 || + typeof e.threshold !== 'number' || !(e.threshold > 0) || + !Number.isInteger(e.window) || e.window < 0 || + !Array.isArray(e.keywords) || e.keywords.some(k => typeof k !== 'string' || !k)) { + throw new RedactionConfigError('Redaction entropy settings are invalid (need enabled, requireKeyword: boolean; minLength >= 8; threshold > 0; window >= 0; keywords: string[])'); + } + return { rules, allowlist, entropy: { ...e, keywords: e.keywords.map(k => k.toLowerCase()) } }; +} +// --------------------------------------------------------------------------- +// Redaction engine +// --------------------------------------------------------------------------- +function tokenFor(ruleId) { + return `${TOKEN_PREFIX}${ruleId}]`; +} +function tokenSpans(text) { + const spans = []; + for (const m of text.matchAll(TOKEN_PATTERN)) + spans.push([m.index, m.index + m[0].length]); + return spans; +} +function overlapsAny(spans, start, end) { + for (const [s, e] of spans) + if (start < e && end > s) + return true; + return false; +} +function shannonEntropy(text) { + const counts = new Map(); + for (const ch of text) + counts.set(ch, (counts.get(ch) ?? 0) + 1); + let entropy = 0; + for (const n of counts.values()) { + const p = n / text.length; + entropy -= p * Math.log2(p); + } + return entropy; +} +/** + * Replace each accepted match (or its secretGroup) with a token. A candidate is + * skipped when it overlaps an existing token (idempotency) or fully matches an + * allowlist pattern. + */ +function applyRule(text, rule, isAllowed) { + const spans = text.includes(TOKEN_PREFIX) ? tokenSpans(text) : null; + const token = tokenFor(rule.id); + let out = ''; + let last = 0; + let count = 0; + for (const m of text.matchAll(rule.regex)) { + const range = m.indices?.[rule.secretGroup]; + if (!range) + continue; + const [start, end] = range; + if (end <= start || start < last) + continue; + if (spans && overlapsAny(spans, start, end)) + continue; + if (isAllowed(text.slice(start, end))) + continue; + out += text.slice(last, start) + token; + last = end; + count++; + } + return count === 0 ? { text, count } : { text: out + text.slice(last), count }; +} +function applyEntropy(text, spec, isAllowed) { + const candidates = new RegExp(`[A-Za-z0-9+/=_.~-]{${spec.minLength},}`, 'g'); + const spans = text.includes(TOKEN_PREFIX) ? tokenSpans(text) : null; + const token = tokenFor(ENTROPY_RULE_ID); + let out = ''; + let last = 0; + let count = 0; + for (const m of text.matchAll(candidates)) { + const start = m.index; + const end = start + m[0].length; + if (spans && overlapsAny(spans, start, end)) + continue; + if (isAllowed(m[0])) + continue; + if (shannonEntropy(m[0]) < spec.threshold) + continue; + if (spec.requireKeyword) { + let before = text.slice(Math.max(0, start - spec.window), start); + before = before.slice(before.lastIndexOf('\n') + 1).toLowerCase(); + if (!spec.keywords.some(k => before.includes(k))) + continue; + } + out += text.slice(last, start) + token; + last = end; + count++; + } + return count === 0 ? { text, count } : { text: out + text.slice(last), count }; +} +export function createRedactor(config) { + const compiled = compileConfig(config); + const isAllowed = (secret) => compiled.allowlist.some(re => re.test(secret)); + return { + ruleIds: compiled.rules.map(r => r.id), + redact(text, _ctx) { + if (typeof text !== 'string' || text.length === 0) + return { text, findings: [] }; + let current = text; + let lower = null; + const counts = new Map(); + for (const rule of compiled.rules) { + if (rule.keywords) { + lower ??= current.toLowerCase(); + if (!rule.keywords.some(k => lower.includes(k))) + continue; + } + const r = applyRule(current, rule, isAllowed); + if (r.count > 0) { + current = r.text; + lower = null; + counts.set(rule.id, (counts.get(rule.id) ?? 0) + r.count); + } + } + if (compiled.entropy.enabled) { + const r = applyEntropy(current, compiled.entropy, isAllowed); + if (r.count > 0) { + current = r.text; + counts.set(ENTROPY_RULE_ID, r.count); + } + } + return { text: current, findings: [...counts].map(([ruleId, count]) => ({ ruleId, count })) }; + }, + }; +} +// --------------------------------------------------------------------------- +// Findings +// --------------------------------------------------------------------------- +/** Aggregates findings by rule id. Never holds matched values. */ +export class FindingsTally { + counts = new Map(); + add(findings) { + for (const f of findings) + this.counts.set(f.ruleId, (this.counts.get(f.ruleId) ?? 0) + f.count); + } + get total() { + let n = 0; + for (const c of this.counts.values()) + n += c; + return n; + } + toArray() { + return [...this.counts].map(([ruleId, count]) => ({ ruleId, count })); + } +} +/** "7 value(s) redacted (azure-storage-key: 2, jwt: 5)" */ +export function formatFindings(findings) { + const total = findings.reduce((n, f) => n + f.count, 0); + if (total === 0) + return 'no values redacted'; + const detail = [...findings] + .sort((a, b) => b.count - a.count || a.ruleId.localeCompare(b.ruleId)) + .map(f => `${f.ruleId}: ${f.count}`) + .join(', '); + return `${total} value(s) redacted (${detail})`; +} +// --------------------------------------------------------------------------- +// JSON / JSONL / files +// --------------------------------------------------------------------------- +/** Redact every string value in a parsed JSON tree, in place. Keys are left alone. */ +function redactTree(node, redactor, ctx, tally) { + if (typeof node === 'string') { + const r = redactor.redact(node, ctx); + if (r.findings.length === 0) + return { value: node, changed: false }; + tally?.add(r.findings); + return { value: r.text, changed: r.text !== node }; + } + if (node === null || typeof node !== 'object') + return { value: node, changed: false }; + let changed = false; + if (Array.isArray(node)) { + for (let i = 0; i < node.length; i++) { + const r = redactTree(node[i], redactor, ctx, tally); + if (r.changed) { + node[i] = r.value; + changed = true; + } + } + return { value: node, changed }; + } + const obj = node; + for (const key of Object.keys(obj)) { + const r = redactTree(obj[key], redactor, ctx, tally); + if (r.changed) { + Object.defineProperty(obj, key, { value: r.value, writable: true, enumerable: true, configurable: true }); + changed = true; + } + } + return { value: obj, changed }; +} +/** + * Redact one JSONL line. Valid JSON stays valid (string values are redacted + * and the line is re-serialized only if something changed); a line that isn't + * JSON (e.g. a torn last line mid-write) is redacted as raw text. A trailing + * `\r` is preserved. + */ +export function redactJsonlLine(line, redactor, ctx, tally) { + const cr = line.endsWith('\r') ? '\r' : ''; + const body = cr ? line.slice(0, -1) : line; + if (body.trim().length === 0) + return line; + let parsed; + try { + parsed = JSON.parse(body); + } + catch { + const r = redactor.redact(body, ctx); + if (r.findings.length === 0) + return line; + tally?.add(r.findings); + return r.text + cr; + } + const r = redactTree(parsed, redactor, ctx, tally); + return r.changed ? JSON.stringify(r.value) + cr : line; +} +const COPY_CHUNK_BYTES = 1 << 20; // 1 MiB +/** + * Stream `src` to `dest`, redacting line by line. Writes `dest` directly; + * callers that need atomicity write to a temp path and rename. Line count and + * trailing-newline shape match the source exactly, so archive line numbers + * (exchanges.line_start/line_end, MCP read ranges) stay valid. + */ +export function copyFileRedacted(src, dest, redactor, ctx, tally) { + const fdIn = fs.openSync(src, 'r'); + let fdOut; + try { + fdOut = fs.openSync(dest, 'w'); + const buf = Buffer.allocUnsafe(COPY_CHUNK_BYTES); + const decoder = new StringDecoder('utf8'); + let pending = ''; + let bytesRead; + while ((bytesRead = fs.readSync(fdIn, buf, 0, buf.length, null)) > 0) { + let scanFrom = pending.length; + pending += decoder.write(buf.subarray(0, bytesRead)); + const out = []; + let start = 0; + let nl; + while ((nl = pending.indexOf('\n', scanFrom)) !== -1) { + out.push(redactJsonlLine(pending.slice(start, nl), redactor, ctx, tally), '\n'); + start = nl + 1; + scanFrom = start; + } + if (out.length > 0) + fs.writeSync(fdOut, out.join('')); + pending = pending.slice(start); + } + pending += decoder.end(); + if (pending.length > 0) + fs.writeSync(fdOut, redactJsonlLine(pending, redactor, ctx, tally)); + } + finally { + fs.closeSync(fdIn); + if (fdOut !== undefined) + fs.closeSync(fdOut); + } +} diff --git a/dist/summarizer.d.ts b/dist/summarizer.d.ts index 4c97be52..c4b5595d 100644 --- a/dist/summarizer.d.ts +++ b/dist/summarizer.d.ts @@ -172,4 +172,13 @@ export declare function runCodexCommand(command: CodexSummarizerCommand): Promis * See https://github.com/obra/episodic-memory/issues/98. */ export declare function getCodexModel(_exchanges: ConversationExchange[]): string | undefined; -export declare function summarizeConversation(exchanges: ConversationExchange[], sessionId?: string): Promise; +export interface SummarizeOptions { + /** + * Allow Claude session resume and Codex thread/fork (default true). Both + * paths make the model read the *source* transcript rather than `exchanges`, + * so callers pass false when the exchanges were redacted (see + * docs/redaction/PHASE0-FINDINGS.md) to force the transcript-text path. + */ + allowResume?: boolean; +} +export declare function summarizeConversation(exchanges: ConversationExchange[], sessionId?: string, options?: SummarizeOptions): Promise; diff --git a/dist/summarizer.js b/dist/summarizer.js index 60719809..48e78be2 100644 --- a/dist/summarizer.js +++ b/dist/summarizer.js @@ -617,7 +617,8 @@ function getCodexSessionId(exchanges, sessionId) { export function getCodexModel(_exchanges) { return process.env.EPISODIC_MEMORY_CODEX_MODEL || undefined; } -export async function summarizeConversation(exchanges, sessionId) { +export async function summarizeConversation(exchanges, sessionId, options = {}) { + const allowResume = options.allowResume !== false; // Handle trivial conversations if (exchanges.length === 0) { return 'Trivial conversation with no substantive content.'; @@ -628,7 +629,7 @@ export async function summarizeConversation(exchanges, sessionId) { return 'Trivial conversation with no substantive content.'; } } - const codexSessionId = getCodexSessionId(exchanges, sessionId); + const codexSessionId = allowResume ? getCodexSessionId(exchanges, sessionId) : undefined; if (codexSessionId) { try { const result = await callCodex(buildCodexSummaryPrompt(), codexSessionId, getCodexModel(exchanges)); @@ -652,7 +653,7 @@ export async function summarizeConversation(exchanges, sessionId) { // would fail on every one before the no-resume retry kicks in. Treat // missing harness as Claude for backward compatibility with old archives. const isClaudeSession = exchanges.some(e => e.harness === 'claude' || e.harness === undefined); - const claudeSessionId = !codexSessionId && isClaudeSession ? sessionId : undefined; + const claudeSessionId = allowResume && !codexSessionId && isClaudeSession ? sessionId : undefined; const cwd = claudeSessionId ? exchanges.find(e => e.cwd)?.cwd : undefined; const conversationText = claudeSessionId ? '' // When resuming, no need to include conversation text - it's already in context diff --git a/dist/sync-cli.js b/dist/sync-cli.js index 26172c70..b4450162 100644 --- a/dist/sync-cli.js +++ b/dist/sync-cli.js @@ -9,6 +9,7 @@ import { spawn } from 'child_process'; import fs from 'fs'; import { formatLogLine, getSyncLogPath, getSyncLockPath } from './logging.js'; import { acquireFileLock, readLockHolder, releaseFileLock } from './file-lock.js'; +import { FindingsTally, formatFindings, getRedactionSettings, loadRedactor } from './redaction.js'; const args = process.argv.slice(2); // Reentrancy guard (#87): if this sync was triggered by a SessionStart hook // inside a Claude subprocess that the summarizer just spawned, exit silently. @@ -39,10 +40,13 @@ Sync conversations from Claude Code, Codex, and opencode transcript sources to a This command: 1. Exports opencode sessions from its SQLite database when available -2. Copies new or updated .jsonl files to conversation archive +2. Copies new or updated .jsonl files to conversation archive, redacting secrets 3. Generates embeddings for semantic search 4. Updates the search index +Secrets are replaced with [REDACTED:] tokens before anything is archived, +indexed, embedded, or summarized. See EPISODIC_MEMORY_REDACTION* in the README. + Only processes files that are new or have been modified since last sync. Safe to run multiple times - subsequent runs are fast no-ops. @@ -142,8 +146,21 @@ if (isBackground) { console.log(`Sync started in background. Log: ${logPath}`); process.exit(0); } +// Load redaction rules before anything is exported, copied, or indexed. In +// strict mode (the default) a bad rules file stops the sync here: fail closed +// rather than archive unredacted text. +let redactor; +try { + redactor = loadRedactor(); +} +catch (error) { + console.error(`episodic-memory: ${error instanceof Error ? error.message : String(error)}`); + console.error('episodic-memory: refusing to sync without redaction (strict mode). Fix the rules file, ' + + 'or set EPISODIC_MEMORY_REDACTION_STRICT=0 to sync unredacted, or EPISODIC_MEMORY_REDACTION=off.'); + process.exit(1); +} if (!onlyHarnesses || onlyHarnesses.includes('opencode')) { - const opencodeExport = exportOpencodeSessions(); + const opencodeExport = exportOpencodeSessions({ redactor }); if (opencodeExport.exported > 0 || opencodeExport.skipped > 0) { console.log(`opencode export: ${opencodeExport.exported} exported, ${opencodeExport.skipped} skipped`); } @@ -195,8 +212,10 @@ console.log(`Sources: ${sourceDirs.join(', ')}`); console.log(`Destination: ${destDir}\n`); async function syncAll() { const totals = { copied: 0, skipped: 0, indexed: 0, summarized: 0, errors: [], sourcesWithSummaryWork: 0, totalNeedingSummaries: 0 }; + const redactions = new FindingsTally(); for (const sourceDir of sourceDirs) { - const result = await syncConversations(sourceDir, destDir, { ...syncOptions, summaryLimit }); + const result = await syncConversations(sourceDir, destDir, { ...syncOptions, summaryLimit, redactor }); + redactions.add(result.redactions); totals.copied += result.copied; totals.skipped += result.skipped; totals.indexed += result.indexed; @@ -213,6 +232,16 @@ async function syncAll() { else { console.log(` Summarized: ${totals.summarized}`); } + // Rule IDs and counts only — never matched values. + if (redactor) { + console.log(` Redaction: ${formatFindings(redactions.toArray())}`); + } + else if (getRedactionSettings().enabled) { + console.log(' Redaction: NOT APPLIED (rules failed to load; EPISODIC_MEMORY_REDACTION_STRICT=0)'); + } + else { + console.log(' Redaction: off (EPISODIC_MEMORY_REDACTION=off)'); + } if (totals.errors.length > 0) { console.log(`\n⚠️ Errors: ${totals.errors.length}`); totals.errors.forEach(err => console.log(` ${err.file}: ${err.error}`)); diff --git a/dist/sync.d.ts b/dist/sync.d.ts index 658e98ad..065e5386 100644 --- a/dist/sync.d.ts +++ b/dist/sync.d.ts @@ -1,3 +1,4 @@ +import { FindingsTally, type RedactionFinding, type Redactor } from './redaction.js'; /** * Stream and scan for any exclusion marker, carrying an overlap between * chunks so a marker split across a boundary is still found. A single @@ -17,11 +18,18 @@ export interface SyncResult { file: string; error: string; }>; + /** Rule IDs and counts only — never matched values. */ + redactions: RedactionFinding[]; } export interface SyncOptions { skipIndex?: boolean; skipSummaries?: boolean; summaryLimit?: number; + /** + * Redactor applied at the archive write. `undefined` loads it from the + * environment (loadRedactor); `null` means redaction is off. + */ + redactor?: Redactor | null; } /** * Derive sync options from the process environment. @@ -39,5 +47,12 @@ export interface SyncOptions { * Claude quota and can stall on a permission prompt. */ export declare function buildSyncOptionsFromEnv(env: NodeJS.ProcessEnv): SyncOptions; +/** + * Copy `src` into the archive at `dest` when the archive copy is missing or + * older. This is the redaction choke point: with a redactor, the copy is + * redacted line by line (see redaction.ts), and every downstream stage + * (index, embeddings, summaries, show/read) reads the archive. + */ +export declare function copyIfNewer(src: string, dest: string, redactor?: Redactor | null, tally?: FindingsTally): boolean; export declare function extractSessionIdFromPath(filePath: string): string | null; export declare function syncConversations(sourceDir: string, destDir: string, options?: SyncOptions): Promise; diff --git a/dist/sync.js b/dist/sync.js index 0629761d..380037e1 100644 --- a/dist/sync.js +++ b/dist/sync.js @@ -5,6 +5,7 @@ import { SUMMARIZER_CONTEXT_MARKER } from './constants.js'; import { getExcludedProjects, findJsonlFiles, statIfExists } from './paths.js'; import { formatErrorSentinel, shouldQueueForSummary } from './summary-sentinel.js'; import { getMaxMessageBytes, isOversizeExchange } from './message-size.js'; +import { copyFileRedacted, FindingsTally, loadRedactor } from './redaction.js'; const EXCLUSION_MARKERS = [ 'DO NOT INDEX THIS CHAT', 'Only use NO_INSIGHTS_FOUND', @@ -122,7 +123,13 @@ function hasConversationContent(filePath) { export function buildSyncOptionsFromEnv(env) { return { skipSummaries: env.EPISODIC_MEMORY_SKIP_SUMMARIES === '1' }; } -function copyIfNewer(src, dest) { +/** + * Copy `src` into the archive at `dest` when the archive copy is missing or + * older. This is the redaction choke point: with a redactor, the copy is + * redacted line by line (see redaction.ts), and every downstream stage + * (index, embeddings, summaries, show/read) reads the archive. + */ +export function copyIfNewer(src, dest, redactor = null, tally) { // Ensure destination directory exists const destDir = path.dirname(dest); if (!fs.existsSync(destDir)) { @@ -138,8 +145,22 @@ function copyIfNewer(src, dest) { } // Atomic copy: temp file + rename const tempDest = dest + '.tmp.' + process.pid; - fs.copyFileSync(src, tempDest); - fs.renameSync(tempDest, dest); // Atomic on same filesystem + try { + if (redactor) { + copyFileRedacted(src, tempDest, redactor, { source: 'archive', path: src }, tally); + } + else { + fs.copyFileSync(src, tempDest); + } + fs.renameSync(tempDest, dest); // Atomic on same filesystem + } + catch (error) { + try { + fs.unlinkSync(tempDest); + } + catch { } + throw error; + } // Preserve source mtime: harnesses without per-message timestamps (Cursor // agent transcripts) fall back to file mtime. Round up to the next whole // millisecond — utimes can't always represent the source's sub-millisecond @@ -165,8 +186,14 @@ export async function syncConversations(sourceDir, destDir, options = {}) { skipped: 0, indexed: 0, summarized: 0, - errors: [] + errors: [], + redactions: [] }; + // Load the redactor before touching the archive. In strict mode (the + // default) a rules-load failure throws here, so nothing unredacted is + // archived, indexed, or summarized (fail closed). + const redactor = options.redactor === undefined ? loadRedactor() : options.redactor; + const tally = new FindingsTally(); // Ensure source directory exists if (!fs.existsSync(sourceDir)) { return result; @@ -198,7 +225,7 @@ export async function syncConversations(sourceDir, destDir, options = {}) { result.skipped++; continue; } - const wasCopied = copyIfNewer(srcFile, destFile); + const wasCopied = copyIfNewer(srcFile, destFile, redactor, tally); if (wasCopied) { result.copied++; filesToIndex.push(destFile); @@ -229,6 +256,7 @@ export async function syncConversations(sourceDir, destDir, options = {}) { } } } + result.redactions = tally.toArray(); // Index copied files (unless skipIndex is set) if (!options.skipIndex && filesToIndex.length > 0) { const { parseConversation } = await import('./parser.js'); @@ -343,7 +371,10 @@ export async function syncConversations(sourceDir, destDir, options = {}) { continue; } console.log(` Summarizing ${path.basename(filePath)} (${exchanges.length} exchanges)...`); - const summary = await summarizeConversation(exchanges, sessionId); + // With redaction on, never resume/fork the session: those paths hand + // the model the unredacted source transcript instead of these + // (redacted) exchanges. + const summary = await summarizeConversation(exchanges, sessionId, { allowResume: redactor === null }); const summaryPath = filePath.replace('.jsonl', '-summary.txt'); fs.writeFileSync(summaryPath, summary, 'utf-8'); result.summarized++; diff --git a/dist/verify.js b/dist/verify.js index dd4535d2..d2fba59f 100644 --- a/dist/verify.js +++ b/dist/verify.js @@ -4,6 +4,7 @@ import { parseConversation } from './parser.js'; import { initDatabase, getAllExchanges, getFileLastIndexed } from './db.js'; import { getArchiveDir, getExcludedProjects, findJsonlFiles, statIfExists } from './paths.js'; import { isErroredSentinel } from './summary-sentinel.js'; +import { getRedactionSettings } from './redaction.js'; export async function verifyIndex() { const result = { missing: [], @@ -125,7 +126,8 @@ export async function repairIndex(issues) { } // Generate/update summary const summaryPath = conversationPath.replace('.jsonl', '-summary.txt'); - const summary = await summarizeConversation(exchanges); + // Under redaction, don't let the Codex fork fallback read the source rollout. + const summary = await summarizeConversation(exchanges, undefined, { allowResume: !getRedactionSettings().enabled }); fs.writeFileSync(summaryPath, summary, 'utf-8'); console.log(` Created summary: ${summary.split(/\s+/).length} words`); // Index exchanges diff --git a/docs/REDACTION.md b/docs/REDACTION.md new file mode 100644 index 00000000..85d967c9 --- /dev/null +++ b/docs/REDACTION.md @@ -0,0 +1,167 @@ +# Secret redaction + +episodic-memory replaces secrets in your conversations with typed tokens +**before** it archives, indexes, embeds, or summarizes them: + +``` +AccountKey=Zm9v...== → AccountKey=[REDACTED:connection-string-secret] +"password": "abc1Q~..." → "password": "[REDACTED:azure-client-secret]" +Authorization: Bearer eyJ → Authorization: Bearer [REDACTED:jwt] +``` + +Values are redacted, not dropped. The rest of the conversation stays +searchable, and the token tells you what kind of value was there. You can +search for the tokens too: `episodic-memory search --text "[REDACTED:azure-storage-key]"` +finds every conversation where a storage key was pasted. + +Redaction is **on by default**, and it **fails closed**. If the rules can't be +loaded, sync refuses to run. It won't archive unredacted text. + +## What's covered + +| Where | Redacted? | +|---|---| +| Conversation archive (`~/.config/superpowers/conversation-archive`) | Yes. This is the hook point. | +| SQLite index: message text and tool inputs/outputs | Yes. Built from the archive. | +| Embeddings | Yes. Built from redacted text. | +| Summaries (sent to a model) | Yes. Built from redacted text. Session resume and Codex fork are turned off (see below). | +| opencode and legacy Cursor staging exports | Yes. Redacted when written. | +| `show`, MCP `read` | Yes. They read the archive. | +| Sync logs | Rule IDs and counts only. Matched values are never logged. | +| Claude Code's own `~/.claude/projects`, Codex's `~/.codex/sessions`, etc. | **No.** Those files belong to the harness. Use its retention settings. | + +### Summaries + +Without redaction, the summarizer can *resume* a Claude Code session or *fork* +a Codex thread. Both make the model read the original, unredacted transcript. +With redaction on, summaries always come from the redacted conversation text. +This has one side effect for Codex-only setups: summarization then goes +through the Claude Agent SDK. If you don't have Claude set up, summaries fail +and retry on later syncs. Set `EPISODIC_MEMORY_SKIP_SUMMARIES=1` to turn them +off. Summaries are display-only, so search quality is unaffected. + +## Default rules + +Run `episodic-memory redact --print-default-rules` for the full set. In +summary: + +| Rule ID | Catches | +|---|---| +| `private-key-block` | PEM and OpenSSH private keys, including truncated ones | +| `connection-string-secret` | `AccountKey=`, `SharedAccessKey=`, `SharedAccessSignature=`, `Password=`, `Pwd=` in connection strings. Only the value is redacted; server, account, and database names stay. | +| `azure-sas-token` | The `sig=` of a SAS URL. The URL and other parameters stay. | +| `jwt` | JWTs, including Entra ID / Azure access tokens | +| `anthropic-api-key`, `openai-api-key`, `github-token`, `aws-access-key-id`, `aws-secret-access-key`, `slack-token`, `google-api-key`, `npm-token` | Provider keys with a recognizable prefix | +| `azure-client-secret` | Entra ID app client secrets (the `…Q~…` format) | +| `azure-storage-key` | Standalone 88-character base64 keys (Storage, Cosmos DB, Function keys) | +| `url-credentials` | The password in `scheme://user:password@host` | +| `bearer-token`, `basic-auth` | `Authorization` header values | +| `secret-assignment` | A value assigned to a secret-looking key: `password: …`, `CLIENT_SECRET=…`, `"apiKey": "…"`. Covers decrypted SOPS, YAML, dotenv, and JSON. Needs 8+ characters including a digit, and skips placeholders (`${X}`, ``, `%X%`) and code (`env.X`, `getPassword()`). | + +**Allowlisted** (never redacted, even when a rule matches): git SHAs and +GUIDs. Tenant, client, object, and subscription IDs stay searchable, and so do +commit hashes. + +**Entropy fallback:** off by default. When it's on, a long, high-entropy string +is redacted only if a keyword (`secret`, `key`, `token`, …) appears just before +it in the same text value. + +**SOPS:** decrypted SOPS output drops its `sops:` metadata block, so it has no +reliable shape. It's covered by the value rules plus `secret-assignment`. +Encrypted values (`ENC[AES256_GCM,…]`) are left alone. + +## Configuration + +| Variable | Default | Meaning | +|---|---|---| +| `EPISODIC_MEMORY_REDACTION` | `on` | `off` disables redaction entirely: archive copies become byte-for-byte again, and summaries may resume sessions again. | +| `EPISODIC_MEMORY_REDACTION_RULES` | `/redaction-rules.json` if it exists | Path to a custom rules file. If you set this and the file is missing, that's an error. | +| `EPISODIC_MEMORY_REDACTION_STRICT` | `1` | When the rules fail to load, `1` stops sync, index, and import (exit 1, nothing written). `0` logs a loud warning and continues **unredacted**. | + +`` is `~/.config/superpowers` unless `EPISODIC_MEMORY_CONFIG_DIR`, +`PERSONAL_SUPERPOWERS_DIR`, or `XDG_CONFIG_HOME` say otherwise. + +### Custom rules + +Create `~/.config/superpowers/redaction-rules.json`. By default it extends the +bundled rules: + +```jsonc +{ + // Add rules (a rule with a default's id replaces that default) + "rules": [ + { + "id": "contoso-api-key", // lowercase, digits, dashes + "pattern": "\\bctso_[A-Za-z0-9]{32}\\b", // JavaScript regex + "flags": "i", // optional, any of "imsu" + "keywords": ["ctso_"], // optional prefilter (case-insensitive) + "secretGroup": 0 // optional: redact only this capture group + } + ], + "disableRules": ["basic-auth"], // turn off defaults by id + "allowlist": [ // full-match patterns that are never redacted + { "id": "build-ids", "pattern": "build-[0-9]{8}" } + ], + "entropy": { "enabled": true }, // partial override of the entropy settings + "includeDefaults": true // false = use only this file's rules +} +``` + +Every rule is validated when it loads. An invalid regex, a bad id, a pattern +that matches the empty string, or a `secretGroup` that doesn't exist is a +rules-load failure, which strict mode treats as fatal. To try rules out: + +```bash +echo 'password: hunter2hunter2' | episodic-memory redact --stdin +# password: [REDACTED:secret-assignment] +# 1 value(s) redacted (secret-assignment: 1) (stderr) +``` + +## Cleaning up existing data + +New syncs only redact new or changed files. To redact everything indexed +before you upgraded, or after you add a rule: + +```bash +episodic-memory redact --rewrite --dry-run # report what would change +episodic-memory redact --rewrite # apply +``` + +`--rewrite`: + +- redacts every archive file in place, keeping line numbers and timestamps +- redacts the opencode and Cursor staging exports in place +- redacts every index row in place and re-embeds only the rows that changed +- deletes summaries that were generated from unredacted text, so the next + sync regenerates them from the redacted archive + +It takes the same lock as `sync`, so it won't run while a sync is in progress. +Running it twice is safe: the second run finds nothing to change. + +Before your first sync with this version, consider a one-time scan of your +existing `~/.claude/projects` history. Those source files are never modified. + +## How it works + +The hook point is the copy into the archive. Every harness (Claude Code, +Codex, Cursor, opencode, OMP) passes through that copy, and every later stage +reads the archive rather than the source. The design and the research behind +it are in [redaction/PHASE0-FINDINGS.md](redaction/PHASE0-FINDINGS.md). + +Each archive line is parsed as JSON, every string value is redacted, and only +lines that changed are re-serialized. Lines with no secrets stay byte-for-byte +identical. The archive keeps exactly one line per source line, so index line +ranges and MCP `read` ranges still line up. A line that isn't valid JSON (for +example, a half-written last line) is redacted as plain text. + +## Limitations + +- Pattern rules miss secrets that have no recognizable shape and no + `key: value` context. The entropy fallback helps when it's on, but recall + isn't perfect. +- Each JSON string value is checked on its own. Context split across fields + (`{"name": "DB_PASSWORD", "value": "…"}`) isn't linked, so the `value` is + caught only if its own shape matches a rule. +- JSON object *keys* aren't redacted. Only values are. +- On lines that get redacted, re-serializing can change number formatting for + integers above 2^53. No supported harness writes such numbers. diff --git a/src/cursor-legacy.ts b/src/cursor-legacy.ts index a62f4b56..8c91897f 100644 --- a/src/cursor-legacy.ts +++ b/src/cursor-legacy.ts @@ -3,6 +3,7 @@ import os from 'os'; import path from 'path'; import Database from 'better-sqlite3'; import { detectCursorCwd } from './parser.js'; +import { loadRedactor, redactJsonlLine, type Redactor } from './redaction.js'; /** * Import legacy Cursor conversations from Cursor's global SQLite store @@ -141,6 +142,8 @@ export interface CursorLegacyImportOptions { force?: boolean; /** Report what would be exported without writing files. */ dryRun?: boolean; + /** `undefined` loads from the environment (strict mode throws on bad rules); `null` disables. */ + redactor?: Redactor | null; } export interface CursorLegacyImportResult { @@ -160,6 +163,7 @@ export function importCursorLegacy(options: CursorLegacyImportOptions): CursorLe errors: [], }; + const redactor = options.redactor === undefined ? loadRedactor() : options.redactor; const liveIds = options.liveTranscriptIds ?? new Set(); const db = new Database(options.dbPath, { readonly: true, fileMustExist: true }); @@ -264,9 +268,14 @@ export function importCursorLegacy(options: CursorLegacyImportOptions): CursorLe if (!options.dryRun) { // Re-serialize with cwd now that it's known (it's derived from the // whole conversation's tool calls). - const finalLines = cwd + const withCwd = cwd ? lines.map(line => JSON.stringify({ ...JSON.parse(line), cwd })) : lines; + // The export dir is a plugin-owned plaintext copy, so redact it at + // write time like the archive (docs/redaction/PHASE0-FINDINGS.md). + const finalLines = redactor + ? withCwd.map(line => redactJsonlLine(line, redactor, { source: 'cursor-legacy', path: outFile })) + : withCwd; fs.mkdirSync(path.dirname(outFile), { recursive: true }); fs.writeFileSync(outFile, finalLines.join('\n') + '\n', 'utf-8'); diff --git a/src/indexer.ts b/src/indexer.ts index 173ef679..5ba56ae2 100644 --- a/src/indexer.ts +++ b/src/indexer.ts @@ -9,6 +9,8 @@ import { ConversationExchange } from './types.js'; import { getArchiveDir, getExcludedProjects, getConversationSourceDirs, findJsonlFiles, statIfExists } from './paths.js'; import { formatErrorSentinel, shouldQueueForSummary } from './summary-sentinel.js'; import { getMaxMessageBytes, isOversizeExchange } from './message-size.js'; +import { FindingsTally, formatFindings, loadRedactor, type Redactor } from './redaction.js'; +import { copyIfNewer } from './sync.js'; // Set max output tokens for Claude SDK (used by summarizer) process.env.CLAUDE_CODE_MAX_OUTPUT_TOKENS = '20000'; @@ -38,12 +40,25 @@ function sessionIdForSummary(exchanges: ConversationExchange[]): string | undefi return exchanges.find(exchange => exchange.sessionId)?.sessionId; } +// Resume/fork would summarize the unredacted source transcript; see sync.ts. +function summarizeOptions(redactor: Redactor | null) { + return { allowResume: redactor === null }; +} + +function logRedactions(tally: FindingsTally): void { + if (tally.total > 0) console.log(` Redaction: ${formatFindings(tally.toArray())}`); +} + export async function indexConversations( limitToProject?: string, maxConversations?: number, concurrency: number = 1, noSummaries: boolean = false ): Promise { + // Load before touching the archive: strict mode fails closed here. + const redactor = loadRedactor(); + const tally = new FindingsTally(); + console.log('Initializing database...'); const db = initDatabase(); @@ -113,15 +128,14 @@ export async function indexConversations( // Source transcripts can vanish mid-run (Claude Code cleanup). Skip loudly. let exchanges; try { - // Copy to archive (ensure parent dirs exist for subagent files) - if (!fs.existsSync(archivePath)) { - fs.mkdirSync(path.dirname(archivePath), { recursive: true }); - fs.copyFileSync(sourcePath, archivePath); + // Copy (redacted) to the archive, then parse the archive, so the index + // and summaries only ever see redacted text. + if (copyIfNewer(sourcePath, archivePath, redactor, tally)) { console.log(` Archived: ${file}`); } // Parse conversation - exchanges = await parseConversation(sourcePath, project, archivePath); + exchanges = await parseConversation(archivePath, project, archivePath); } catch (error) { console.log(` Skipped ${file} (read failed: ${error instanceof Error ? error.message : error})`); continue; @@ -150,7 +164,7 @@ export async function indexConversations( await processBatch(needsSummary, async (conv) => { try { - const summary = await summarizeConversation(conv.exchanges, sessionIdForSummary(conv.exchanges)); + const summary = await summarizeConversation(conv.exchanges, sessionIdForSummary(conv.exchanges), summarizeOptions(redactor)); fs.writeFileSync(conv.summaryPath, summary, 'utf-8'); const wordCount = summary.split(/\s+/).length; console.log(` ✓ ${conv.file}: ${wordCount} words`); @@ -193,6 +207,7 @@ export async function indexConversations( // Check if we hit the limit if (maxConversations && conversationsProcessed >= maxConversations) { console.log(`\nReached limit of ${maxConversations} conversations`); + logRedactions(tally); db.close(); console.log(`✅ Indexing complete! Conversations: ${conversationsProcessed}, Exchanges: ${totalExchanges}`); return; @@ -205,12 +220,15 @@ export async function indexConversations( console.log(` Skipped ${oversizeSkipped} oversize exchange(s) (> ${maxMessageBytes} bytes; set EPISODIC_MEMORY_MAX_MESSAGE_BYTES to change) — likely embedded-transcript payloads (#139)`); } + logRedactions(tally); db.close(); console.log(`\n✅ Indexing complete! Conversations: ${conversationsProcessed}, Exchanges: ${totalExchanges}`); } export async function indexSession(sessionId: string, concurrency: number = 1, noSummaries: boolean = false): Promise { console.log(`Indexing session: ${sessionId}`); + const redactor = loadRedactor(); + const tally = new FindingsTally(); // Find the conversation file for this session const sourceDirs = getConversationSourceDirs(); @@ -246,11 +264,8 @@ export async function indexSession(sessionId: string, concurrency: number = 1, n // Archive + parse — source may vanish mid-run (Claude Code cleanup). let exchanges; try { - if (!fs.existsSync(archivePath)) { - fs.mkdirSync(path.dirname(archivePath), { recursive: true }); - fs.copyFileSync(sourcePath, archivePath); - } - exchanges = await parseConversation(sourcePath, project, archivePath); + copyIfNewer(sourcePath, archivePath, redactor, tally); + exchanges = await parseConversation(archivePath, project, archivePath); } catch (error) { console.log(`Skipped ${file} (read failed: ${error instanceof Error ? error.message : error})`); db.close(); @@ -263,7 +278,7 @@ export async function indexSession(sessionId: string, concurrency: number = 1, n if (!noSummaries && shouldQueueForSummary(summaryPath)) { fs.mkdirSync(path.dirname(summaryPath), { recursive: true }); try { - const summary = await summarizeConversation(exchanges, sessionIdForSummary(exchanges)); + const summary = await summarizeConversation(exchanges, sessionIdForSummary(exchanges), summarizeOptions(redactor)); fs.writeFileSync(summaryPath, summary, 'utf-8'); console.log(`Summary: ${summary.split(/\s+/).length} words`); } catch (error) { @@ -297,6 +312,7 @@ export async function indexSession(sessionId: string, concurrency: number = 1, n console.log(` Skipped ${oversizeSkipped} oversize exchange(s) (> ${maxMessageBytes} bytes; set EPISODIC_MEMORY_MAX_MESSAGE_BYTES to change) — likely embedded-transcript payloads (#139)`); } + logRedactions(tally); console.log(`✅ Indexed session ${sessionId}: ${exchanges.length} exchanges`); } @@ -317,6 +333,9 @@ export async function indexUnprocessed(concurrency: number = 1, noSummaries: boo if (concurrency > 1) console.log(`Concurrency: ${concurrency}`); if (noSummaries) console.log('⚠️ Running in no-summaries mode (skipping AI summaries)'); + const redactor = loadRedactor(); + const tally = new FindingsTally(); + const db = initDatabase(); await initEmbeddings(); @@ -361,17 +380,13 @@ export async function indexUnprocessed(concurrency: number = 1, noSummaries: boo ).get(archivePath) as { maxLine: number }; const maxIndexedLine = hw.maxLine; - // Ensure parent dirs exist for subagent files try { - fs.mkdirSync(path.dirname(archivePath), { recursive: true }); - - // Refresh the archive when the source may have grown beyond what we've seen. - if (!fs.existsSync(archivePath) || maxIndexedLine > 0) { - fs.copyFileSync(sourcePath, archivePath); - } + // Refresh the (redacted) archive when the source has grown, then parse + // the archive so the index only sees redacted text. + copyIfNewer(sourcePath, archivePath, redactor, tally); // Parse and filter to exchanges past the high-water mark - const exchanges = await parseConversation(sourcePath, project, archivePath); + const exchanges = await parseConversation(archivePath, project, archivePath); const newExchanges = maxIndexedLine > 0 ? exchanges.filter(e => e.lineStart > maxIndexedLine) : exchanges; @@ -386,6 +401,8 @@ export async function indexUnprocessed(concurrency: number = 1, noSummaries: boo } } // end sourceDir loop + logRedactions(tally); + if (unprocessed.length === 0) { console.log('✅ All conversations are already processed!'); db.close(); @@ -402,7 +419,7 @@ export async function indexUnprocessed(concurrency: number = 1, noSummaries: boo await processBatch(needsSummary, async (conv) => { try { - const summary = await summarizeConversation(conv.exchanges, sessionIdForSummary(conv.exchanges)); + const summary = await summarizeConversation(conv.exchanges, sessionIdForSummary(conv.exchanges), summarizeOptions(redactor)); fs.writeFileSync(conv.summaryPath, summary, 'utf-8'); const wordCount = summary.split(/\s+/).length; console.log(` ✓ ${conv.project}/${conv.file}: ${wordCount} words`); diff --git a/src/opencode-sync.ts b/src/opencode-sync.ts index b61f8ef4..60dae458 100644 --- a/src/opencode-sync.ts +++ b/src/opencode-sync.ts @@ -2,6 +2,7 @@ import fs from 'fs'; import path from 'path'; import Database from 'better-sqlite3'; import { getOpencodeDbPath, getOpencodeTranscriptDir } from './paths.js'; +import { loadRedactor, redactJsonlLine, type Redactor } from './redaction.js'; export interface OpencodeExportResult { exported: number; @@ -92,7 +93,8 @@ function shouldExportSession(filePath: string, sessionUpdatedMs: number): boolea function writeSessionTranscript( db: Database.Database, session: OpencodeSessionRow, - filePath: string + filePath: string, + redactor: Redactor | null ): void { const messages = db.prepare(` SELECT id, session_id, time_created, time_updated, data @@ -173,9 +175,15 @@ function writeSessionTranscript( })); } + // The staging transcript is a plugin-owned plaintext copy, so redact it at + // write time like the archive (docs/redaction/PHASE0-FINDINGS.md). + const output = redactor + ? lines.map(line => redactJsonlLine(line, redactor, { source: 'opencode', path: filePath })) + : lines; + fs.mkdirSync(path.dirname(filePath), { recursive: true }); const tempPath = `${filePath}.tmp.${process.pid}`; - fs.writeFileSync(tempPath, `${lines.join('\n')}\n`, 'utf-8'); + fs.writeFileSync(tempPath, `${output.join('\n')}\n`, 'utf-8'); fs.renameSync(tempPath, filePath); const mtime = dateFromMillis(session.time_updated); fs.utimesSync(filePath, mtime, mtime); @@ -184,7 +192,10 @@ function writeSessionTranscript( export function exportOpencodeSessions(options: { dbPath?: string; transcriptDir?: string; + /** `undefined` loads from the environment (strict mode throws on bad rules); `null` disables. */ + redactor?: Redactor | null; } = {}): OpencodeExportResult { + const redactor = options.redactor === undefined ? loadRedactor() : options.redactor; const dbPath = options.dbPath || getOpencodeDbPath(); const transcriptDir = options.transcriptDir || getOpencodeTranscriptDir(); const result: OpencodeExportResult = { @@ -234,7 +245,7 @@ export function exportOpencodeSessions(options: { result.skipped++; continue; } - writeSessionTranscript(db, session, filePath); + writeSessionTranscript(db, session, filePath, redactor); result.exported++; } catch (error) { result.errors.push({ diff --git a/src/redact-cli.ts b/src/redact-cli.ts new file mode 100644 index 00000000..6e75b0a6 --- /dev/null +++ b/src/redact-cli.ts @@ -0,0 +1,141 @@ +import fs from 'fs'; +import { getArchiveDir, getCursorLegacyExportDir, getOpencodeTranscriptDir } from './paths.js'; +import { acquireFileLock, readLockHolder, releaseFileLock } from './file-lock.js'; +import { getSyncLockPath } from './logging.js'; +import { + DEFAULT_REDACTION_CONFIG, + FindingsTally, + formatFindings, + getRedactionSettings, + loadRedactor, + redactJsonlLine, + type Redactor, +} from './redaction.js'; + +const args = process.argv.slice(2); + +const HELP = ` +Usage: episodic-memory redact [--rewrite [--dry-run]] [--stdin] [--print-default-rules] + +Secret redaction for the conversation archive and index. + +OPTIONS: + --rewrite Re-run redaction over the existing archive, staging exports, + and search index (in place). Rows that change are re-embedded, + and summaries built from unredacted text are deleted (the next + sync regenerates them). Run once after upgrading, and again + after adding rules. + --dry-run With --rewrite: report what would change, write nothing. + --stdin Redact stdin to stdout, one JSONL/text line at a time, and + print rule counts to stderr. Handy for testing rules. + --print-default-rules Print the bundled rules as JSON (a starting point for + redaction-rules.json). + --help, -h Show this help + +ENVIRONMENT: + EPISODIC_MEMORY_REDACTION on (default) | off + EPISODIC_MEMORY_REDACTION_RULES custom rules file (default: /redaction-rules.json) + EPISODIC_MEMORY_REDACTION_STRICT 1 (default) fails closed on a bad rules file | 0 passes through + +Output names rule IDs and counts only; matched values are never printed. +`; + +function fail(message: string): never { + console.error(`episodic-memory: ${message}`); + process.exit(1); +} + +function requireRedactor(): Redactor { + if (!getRedactionSettings().enabled) { + fail('redaction is off (EPISODIC_MEMORY_REDACTION=off); unset it to use this command.'); + } + let redactor: Redactor | null; + try { + // Always strict here: a rewrite with no rules would be a silent no-op. + redactor = loadRedactor({ ...process.env, EPISODIC_MEMORY_REDACTION_STRICT: '1' }); + } catch (error) { + fail(error instanceof Error ? error.message : String(error)); + } + return redactor!; +} + +async function runRewrite(dryRun: boolean): Promise { + const redactor = requireRedactor(); + + // Share sync's single-instance lock so a background sync can't write + // between our reads and renames. + const lockPath = getSyncLockPath(); + const lock = acquireFileLock(lockPath); + if (!lock) { + const holder = readLockHolder(lockPath); + fail(`a sync or index run is in progress (${holder !== null ? `pid ${holder}` : 'another process'}); try again when it finishes.`); + } + const release = () => releaseFileLock(lock); + process.on('exit', release); + + const { rewriteArchive } = await import('./redact-rewrite.js'); + let embeddingsReady = false; + const embed = async (user: string, assistant: string, toolNames?: string[]) => { + const embeddings = await import('./embeddings.js'); + if (!embeddingsReady) { + await embeddings.initEmbeddings(); + embeddingsReady = true; + } + return embeddings.generateExchangeEmbedding(user, assistant, toolNames); + }; + + const archiveDir = getArchiveDir(); + const stagingDirs = [getOpencodeTranscriptDir(), getCursorLegacyExportDir()].filter(d => fs.existsSync(d)); + console.log(`Redacting${dryRun ? ' (dry run)' : ''}: ${archiveDir}`); + for (const dir of stagingDirs) console.log(` + staging: ${dir}`); + + const result = await rewriteArchive({ + archiveDir, + stagingDirs, + redactor, + embed, + dryRun, + log: message => console.log(` ${message}`), + }); + + console.log(`\n${dryRun ? 'Would redact' : 'Redacted'} across archive, staging, and index: ${formatFindings(result.findings)}`); + if (dryRun && (result.filesRewritten || result.rowsUpdated || result.summariesRemoved || result.stagingFilesRewritten)) { + console.log('Run without --dry-run to apply.'); + } +} + +async function runStdin(): Promise { + const redactor = requireRedactor(); + const chunks: Buffer[] = []; + for await (const chunk of process.stdin) chunks.push(chunk as Buffer); + const input = Buffer.concat(chunks).toString('utf-8'); + const tally = new FindingsTally(); + const output = input.split('\n').map(line => redactJsonlLine(line, redactor, { source: 'stdin', path: '-' }, tally)); + process.stdout.write(output.join('\n')); + console.error(formatFindings(tally.toArray())); +} + +async function main(): Promise { + if (args.length === 0 || args.includes('--help') || args.includes('-h')) { + console.log(HELP); + return; + } + if (args.includes('--print-default-rules')) { + console.log(JSON.stringify(DEFAULT_REDACTION_CONFIG, null, 2)); + return; + } + if (args.includes('--stdin')) { + await runStdin(); + return; + } + if (args.includes('--rewrite')) { + await runRewrite(args.includes('--dry-run')); + return; + } + fail(`unknown option(s): ${args.join(' ')}. Try: episodic-memory redact --help`); +} + +main().catch(error => { + console.error('Error:', error instanceof Error ? error.message : error); + process.exit(1); +}); diff --git a/src/redact-rewrite.ts b/src/redact-rewrite.ts new file mode 100644 index 00000000..390f74b0 --- /dev/null +++ b/src/redact-rewrite.ts @@ -0,0 +1,220 @@ +import fs from 'fs'; +import path from 'path'; +import { initDatabase } from './db.js'; +import { recordReembedded } from './embedding-migration.js'; +import { + copyFileRedacted, + FindingsTally, + redactJsonlLine, + type RedactionFinding, + type Redactor, +} from './redaction.js'; + +/** + * Backfill for `episodic-memory redact --rewrite`: re-run redaction over data + * written before redaction existed (or before a rule was added). + * + * 1. Archive: every .jsonl is redacted in place. Line count and mtime are + * preserved, so index line ranges stay valid and sync still sees the + * archive as current. + * 2. Staging dirs (opencode / legacy Cursor exports): same, in place. + * 3. Index: every exchanges/tool_calls row is redacted in place. Changed rows + * are re-embedded from the redacted text, so no vector is left that was + * derived from a secret. + * 4. Summaries: a `-summary.txt` for a conversation that had findings (in the + * archive or the index), or one that matches a rule itself, is deleted. + * It was generated from unredacted text, and the next sync regenerates it + * from the redacted archive. + * + * Idempotent: a second run finds nothing to change. + */ + +export type EmbedFn = (user: string, assistant: string, toolNames?: string[]) => Promise; + +export interface RewriteOptions { + archiveDir: string; + redactor: Redactor; + embed: EmbedFn; + /** Count what would change without writing anything. */ + dryRun?: boolean; + /** Plugin-owned staging dirs to redact in place (not indexed). */ + stagingDirs?: string[]; + log?: (message: string) => void; +} + +export interface RewriteResult { + filesScanned: number; + filesRewritten: number; + stagingFilesRewritten: number; + rowsUpdated: number; + summariesRemoved: number; + /** Rule IDs and counts only — never matched values. */ + findings: RedactionFinding[]; +} + +const SUMMARY_SUFFIX = '-summary.txt'; +const PAGE_SIZE = 500; + +function walk(dir: string): string[] { + const out: string[] = []; + let entries: fs.Dirent[]; + try { + entries = fs.readdirSync(dir, { withFileTypes: true }); + } catch { + return out; + } + for (const entry of entries) { + const full = path.join(dir, entry.name); + if (entry.isDirectory()) out.push(...walk(full)); + else if (entry.isFile()) out.push(full); + } + return out; +} + +function summaryPathFor(jsonlPath: string): string { + return jsonlPath.replace(/\.jsonl$/, SUMMARY_SUFFIX); +} + +/** Redact one JSONL file in place. Returns the number of values redacted. */ +function rewriteFileInPlace(file: string, redactor: Redactor, dryRun: boolean, tally: FindingsTally): number { + const fileTally = new FindingsTally(); + const temp = `${file}.redact.${process.pid}`; + try { + copyFileRedacted(file, temp, redactor, { source: 'rewrite', path: file }, fileTally); + if (fileTally.total > 0 && !dryRun) { + const stat = fs.statSync(file); + fs.renameSync(temp, file); + // Same rounding as sync's copyIfNewer: never leave the archive older than its source. + fs.utimesSync(file, stat.atimeMs / 1000, Math.ceil(stat.mtimeMs) / 1000); + } + } finally { + try { fs.unlinkSync(temp); } catch {} + } + tally.add(fileTally.toArray()); + return fileTally.total; +} + +export async function rewriteArchive(options: RewriteOptions): Promise { + const { archiveDir, redactor, embed } = options; + const dryRun = options.dryRun === true; + const log = options.log ?? (() => {}); + const tally = new FindingsTally(); + const result: RewriteResult = { + filesScanned: 0, + filesRewritten: 0, + stagingFilesRewritten: 0, + rowsUpdated: 0, + summariesRemoved: 0, + findings: [], + }; + // Conversations whose summary was built from unredacted text. + const staleSummaries = new Set(); + + // 1. Archive files. + const archiveFiles = walk(archiveDir); + for (const file of archiveFiles.filter(f => f.endsWith('.jsonl'))) { + result.filesScanned++; + if (rewriteFileInPlace(file, redactor, dryRun, tally) > 0) { + result.filesRewritten++; + staleSummaries.add(summaryPathFor(file)); + } + } + log(`Archive: ${result.filesRewritten} of ${result.filesScanned} file(s) ${dryRun ? 'would be ' : ''}rewritten`); + + // 2. Staging dirs. + for (const dir of options.stagingDirs ?? []) { + for (const file of walk(dir).filter(f => f.endsWith('.jsonl'))) { + if (rewriteFileInPlace(file, redactor, dryRun, tally) > 0) result.stagingFilesRewritten++; + } + } + if (options.stagingDirs?.length) { + log(`Staging exports: ${result.stagingFilesRewritten} file(s) ${dryRun ? 'would be ' : ''}rewritten`); + } + + // 3. Index rows. Page by rowid so writes between pages don't disturb the scan. + const db = initDatabase(); + try { + const page = db.prepare( + 'SELECT rowid AS rid, id, user_message, assistant_message, archive_path FROM exchanges WHERE rowid > ? ORDER BY rowid LIMIT ?' + ); + const toolsFor = db.prepare('SELECT id, tool_name, tool_input, tool_result FROM tool_calls WHERE exchange_id = ? ORDER BY rowid'); + const updateExchange = db.prepare('UPDATE exchanges SET user_message = ?, assistant_message = ? WHERE id = ?'); + const updateTool = db.prepare('UPDATE tool_calls SET tool_input = ?, tool_result = ? WHERE id = ?'); + + let lastRowid = 0; + for (;;) { + const rows = page.all(lastRowid, PAGE_SIZE) as Array<{ + rid: number; id: string; user_message: string; assistant_message: string; archive_path: string; + }>; + if (rows.length === 0) break; + lastRowid = rows[rows.length - 1].rid; + + for (const row of rows) { + const rowTally = new FindingsTally(); + const ctx = { source: 'index', path: row.archive_path }; + const user = redactor.redact(row.user_message, ctx); + const assistant = redactor.redact(row.assistant_message, ctx); + rowTally.add(user.findings); + rowTally.add(assistant.findings); + + const tools = toolsFor.all(row.id) as Array<{ id: string; tool_name: string; tool_input: string | null; tool_result: string | null }>; + const toolUpdates: Array<{ id: string; input: string | null; result: string | null }> = []; + for (const tool of tools) { + // tool_input is JSON text; redactJsonlLine keeps it valid JSON. + const input = tool.tool_input === null ? null : redactJsonlLine(tool.tool_input, redactor, ctx, rowTally); + let output = tool.tool_result; + if (output !== null) { + const r = redactor.redact(output, ctx); + rowTally.add(r.findings); + output = r.text; + } + if (input !== tool.tool_input || output !== tool.tool_result) { + toolUpdates.push({ id: tool.id, input, result: output }); + } + } + + if (rowTally.total === 0) continue; + tally.add(rowTally.toArray()); + result.rowsUpdated++; + staleSummaries.add(summaryPathFor(row.archive_path)); + if (dryRun) continue; + + const toolNames = tools.length > 0 ? tools.map(t => t.tool_name) : undefined; + const embedding = await embed(user.text, assistant.text, toolNames); + db.transaction(() => { + updateExchange.run(user.text, assistant.text, row.id); + for (const t of toolUpdates) updateTool.run(t.input, t.result, t.id); + recordReembedded(db, row.id, embedding); + })(); + } + } + } finally { + db.close(); + } + log(`Index: ${result.rowsUpdated} exchange(s) ${dryRun ? 'would be ' : ''}redacted and re-embedded`); + + // 4. Summaries: stale ones, plus any summary that itself matches a rule. + for (const file of archiveFiles.filter(f => f.endsWith(SUMMARY_SUFFIX))) { + if (staleSummaries.has(file)) continue; + let text: string; + try { + text = fs.readFileSync(file, 'utf-8'); + } catch { + continue; + } + const r = redactor.redact(text, { source: 'summary', path: file }); + if (r.findings.length > 0) { + tally.add(r.findings); + staleSummaries.add(file); + } + } + for (const summary of staleSummaries) { + if (!fs.existsSync(summary)) continue; + result.summariesRemoved++; + if (!dryRun) fs.unlinkSync(summary); + } + log(`Summaries: ${result.summariesRemoved} ${dryRun ? 'would be ' : ''}removed (regenerated from redacted text on the next sync)`); + + result.findings = tally.toArray(); + return result; +} diff --git a/src/redaction-rules.ts b/src/redaction-rules.ts new file mode 100644 index 00000000..d73a043f --- /dev/null +++ b/src/redaction-rules.ts @@ -0,0 +1,168 @@ +import type { RedactionConfig } from './redaction.js'; + +/** + * Bundled default redaction rules. + * + * Patterns are ported from gitleaks' rule set (https://github.com/gitleaks/gitleaks, + * MIT), trimmed to the credentials that realistically show up in Claude Code / + * Codex transcripts for an Azure-heavy stack, and rewritten for JavaScript + * regex (lookbehind instead of gitleaks' consuming boundary groups, so + * adjacent matches aren't swallowed). + * + * Order matters: rules run top to bottom, and a later rule never re-matches + * inside an earlier rule's `[REDACTED:...]` token. Specific shapes go first so + * the token names the most precise rule; the generic `secret-assignment` + * keyword rule goes last. + * + * `secretGroup` replaces only that capture group, so surrounding context + * (connection-string server names, SAS URL paths, usernames) stays searchable. + * + * `keywords` is a case-insensitive prefilter: a rule only runs on text that + * contains at least one keyword. It keeps the per-string cost low on large + * transcripts and bounds the generic rules' work. + * + * Users extend or override these with `redaction-rules.json`; see + * docs/REDACTION.md. `episodic-memory redact --print-default-rules` dumps + * this object as JSON. + */ +export const DEFAULT_REDACTION_CONFIG: RedactionConfig = { + rules: [ + { + id: 'private-key-block', + description: 'PEM/OpenSSH private key block; a truncated block (no END line) is redacted to the end of the text', + pattern: String.raw`-----BEGIN[A-Z0-9 _-]{0,40}PRIVATE KEY(?: BLOCK)?-----[\s\S]*?(?:-----END[A-Z0-9 _-]{0,40}PRIVATE KEY(?: BLOCK)?-----|$)`, + keywords: ['private key'], + }, + { + id: 'connection-string-secret', + description: 'Secret value inside an Azure/ADO.NET connection string; account, server and database names are kept', + pattern: String.raw`\b(?:AccountKey|SharedAccessKey|SharedAccessSignature|SharedSecret|ClientSecret|Password|Pwd)[ \t]*=[ \t]*(?![\[$<{%])([^;"'\s]+)`, + secretGroup: 1, + keywords: ['accountkey', 'sharedaccess', 'sharedsecret', 'clientsecret', 'password', 'pwd'], + }, + { + id: 'azure-sas-token', + description: 'Signature of an Azure SAS token; the resource URL and other SAS parameters are kept', + pattern: String.raw`\bsig=([A-Za-z0-9%+/=_-]{16,})`, + secretGroup: 1, + keywords: ['sig='], + }, + { + id: 'jwt', + description: 'JSON Web Token (Entra ID / Azure access tokens, id tokens)', + pattern: String.raw`\beyJ[A-Za-z0-9_-]{8,}\.eyJ[A-Za-z0-9_-]{8,}\.[A-Za-z0-9_-]{8,}`, + keywords: ['eyj'], + }, + { + id: 'anthropic-api-key', + pattern: String.raw`\bsk-ant-[a-z]{2,10}\d{2}-[A-Za-z0-9_-]{32,}`, + keywords: ['sk-ant-'], + }, + { + id: 'openai-api-key', + pattern: String.raw`\bsk-(?:[a-z]+-)?[A-Za-z0-9_-]{16,}T3BlbkFJ[A-Za-z0-9_-]{16,}`, + keywords: ['t3blbkfj'], + }, + { + id: 'github-token', + description: 'GitHub classic, OAuth, app and fine-grained tokens', + pattern: String.raw`\b(?:gh[pousr]_[A-Za-z0-9]{36,255}|github_pat_[A-Za-z0-9_]{22,255})\b`, + keywords: ['ghp_', 'gho_', 'ghu_', 'ghs_', 'ghr_', 'github_pat_'], + }, + { + id: 'aws-access-key-id', + pattern: String.raw`\b(?:AKIA|ASIA|ABIA|ACCA)[A-Z2-7]{16}\b`, + keywords: ['akia', 'asia', 'abia', 'acca'], + }, + { + id: 'aws-secret-access-key', + pattern: String.raw`aws[_-]?secret[_-]?access[_-]?key["']?[ \t]*[:=][ \t]*["']?([A-Za-z0-9/+]{40})(?![A-Za-z0-9/+])`, + flags: 'i', + secretGroup: 1, + keywords: ['secret'], + }, + { + id: 'slack-token', + pattern: String.raw`\bxox[abposr]-[A-Za-z0-9-]{10,}`, + keywords: ['xox'], + }, + { + id: 'google-api-key', + pattern: String.raw`\bAIza[0-9A-Za-z_-]{35}(?![0-9A-Za-z_-])`, + keywords: ['aiza'], + }, + { + id: 'npm-token', + pattern: String.raw`\bnpm_[A-Za-z0-9]{36}\b`, + keywords: ['npm_'], + }, + { + id: 'azure-client-secret', + description: 'Entra ID (Azure AD) application client secret: 3 chars, a digit, "Q~", 31-34 chars', + pattern: String.raw`(?]+:(?![\[$<{%])([^\s@/"'<>]+)@`, + secretGroup: 1, + keywords: ['://'], + }, + { + id: 'bearer-token', + pattern: String.raw`\bBearer\s+([A-Za-z0-9\-._~+/]{20,}=*)`, + flags: 'i', + secretGroup: 1, + keywords: ['bearer'], + }, + { + id: 'basic-auth', + pattern: String.raw`\bAuthorization[ \t]*:[ \t]*Basic\s+([A-Za-z0-9+/]{8,}={0,2})`, + flags: 'i', + secretGroup: 1, + keywords: ['basic'], + }, + { + id: 'secret-assignment', + description: + 'Value assigned to a secret-looking key (password: x, CLIENT_SECRET=x, "apiKey": "x"). ' + + 'Covers decrypted SOPS/YAML/dotenv/JSON. Requires 8+ chars including a digit; ' + + 'skips placeholders (${X}, , %X%, {{x}}) and code (calls, dotted member access)', + pattern: + String.raw`\b[A-Za-z0-9_.-]{0,40}?(?:secret|password|passwd|passphrase|api[_-]?key|access[_-]?key|private[_-]?key|token)` + + String.raw`["']?[ \t]*[:=][ \t]*["']?(?![\[$<{%])(?=[^\s"',;&\\]*\d)` + + // Not a dotted identifier path (env.AUTH0_SECRET, this.config.token2): code, not a value. + String.raw`(?![A-Za-z_$][\w$]*(?:\.[A-Za-z_$][\w$]*)+(?:$|[\s"',;&\\)}\]]))` + + String.raw`([^\s"',;&\\()<>{}\[\]]{8,})(?=$|[\s"',;&\\)}\]])`, + flags: 'i', + secretGroup: 1, + keywords: ['secret', 'passw', 'passphrase', 'key', 'token'], + }, + ], + allowlist: [ + { + id: 'git-sha', + description: 'Full or short git commit SHA', + pattern: String.raw`\b[0-9a-f]{7,40}\b`, + }, + { + id: 'guid', + description: 'GUID/UUID (Entra tenant, client and object IDs, subscription IDs)', + pattern: String.raw`\b[0-9a-fA-F]{8}-[0-9a-fA-F]{4}-[0-9a-fA-F]{4}-[0-9a-fA-F]{4}-[0-9a-fA-F]{12}\b`, + }, + ], + entropy: { + enabled: false, + minLength: 32, + threshold: 4.5, + requireKeyword: true, + keywords: ['secret', 'key', 'token', 'password', 'passwd', 'credential', 'signature'], + window: 40, + }, +}; diff --git a/src/redaction.ts b/src/redaction.ts new file mode 100644 index 00000000..ccd95a12 --- /dev/null +++ b/src/redaction.ts @@ -0,0 +1,618 @@ +import fs from 'fs'; +import path from 'path'; +import { StringDecoder } from 'string_decoder'; +import { getSuperpowersDir } from './paths.js'; +import { DEFAULT_REDACTION_CONFIG } from './redaction-rules.js'; + +/** + * Inline secret redaction. + * + * Secrets are replaced with typed tokens (`[REDACTED:]`) at the archive + * write, the one point every harness's transcripts pass through (see + * docs/redaction/PHASE0-FINDINGS.md). Everything downstream (SQLite text, + * tool_calls, embeddings, summaries, show/read) reads the archive, so it only + * ever sees redacted text. + * + * Invariants: + * - Findings carry rule IDs and counts only. A matched value is never logged, + * returned, or thrown. + * - Idempotent: a `[REDACTED:...]` token never matches a rule, so redacting + * redacted text is a no-op. + * - JSONL structure is preserved: one output line per input line, valid JSON in + * gives valid JSON out, and unchanged lines are byte-identical. + */ + +export { DEFAULT_REDACTION_CONFIG }; + +/** Fork default. Flip to false for an opt-in upstream build. */ +export const REDACTION_ENABLED_BY_DEFAULT = true; + +export const REDACTION_RULES_FILENAME = 'redaction-rules.json'; + +const RULE_ID_PATTERN = /^[a-z0-9][a-z0-9-]*$/; +const TOKEN_PATTERN = /\[REDACTED:[a-z0-9][a-z0-9-]*\]/g; +const TOKEN_PREFIX = '[REDACTED:'; +const ALLOWED_FLAGS = /^[imsu]*$/; +const ENTROPY_RULE_ID = 'high-entropy'; + +export interface RedactionRuleSpec { + id: string; + pattern: string; + /** Extra RegExp flags from [imsu]; `g` and `d` are always added. */ + flags?: string; + /** Case-insensitive prefilter: skip the rule unless the text contains one. */ + keywords?: string[]; + /** Replace only this capture group instead of the whole match. */ + secretGroup?: number; + description?: string; +} + +export interface AllowlistSpec { + id: string; + /** Matched against the whole candidate secret; a full match keeps it. */ + pattern: string; + flags?: string; + description?: string; +} + +export interface EntropySpec { + enabled: boolean; + minLength: number; + /** Shannon entropy in bits per character. */ + threshold: number; + /** Only fire when a keyword appears shortly before the candidate, on the same line. */ + requireKeyword: boolean; + keywords: string[]; + /** How many characters before the candidate to search for a keyword. */ + window: number; +} + +export interface RedactionConfig { + rules: RedactionRuleSpec[]; + allowlist: AllowlistSpec[]; + entropy: EntropySpec; +} + +/** Shape of a user `redaction-rules.json`. Every field is optional. */ +export interface RedactionRulesFile { + /** Merge with the bundled defaults (default true). */ + includeDefaults?: boolean; + /** Added after the defaults; a rule with a default's id replaces it in place. */ + rules?: RedactionRuleSpec[]; + /** Default rule ids to turn off. */ + disableRules?: string[]; + /** Added to the default allowlist (same id replaces). */ + allowlist?: AllowlistSpec[]; + entropy?: Partial; +} + +export interface RedactionContext { + source: string; + path: string; +} + +export interface RedactionFinding { + ruleId: string; + count: number; +} + +export interface RedactionResult { + text: string; + findings: RedactionFinding[]; +} + +export interface Redactor { + redact(text: string, ctx?: RedactionContext): RedactionResult; + readonly ruleIds: string[]; +} + +export interface RedactionSettings { + enabled: boolean; + strict: boolean; + /** Explicit EPISODIC_MEMORY_REDACTION_RULES path, if set. */ + rulesPath?: string; +} + +/** Rules could not be loaded or are invalid. Messages describe config only. */ +export class RedactionConfigError extends Error { + constructor(message: string) { + super(message); + this.name = 'RedactionConfigError'; + } +} + +// --------------------------------------------------------------------------- +// Settings and loading +// --------------------------------------------------------------------------- + +const OFF_VALUES = new Set(['off', '0', 'false', 'no', 'disabled']); +const ON_VALUES = new Set(['on', '1', 'true', 'yes', 'enabled']); + +/** Unknown values fall back to the default, which is the safe side for both switches. */ +function parseToggle(raw: string | undefined, fallback: boolean): boolean { + const value = raw?.trim().toLowerCase(); + if (!value) return fallback; + if (OFF_VALUES.has(value)) return false; + if (ON_VALUES.has(value)) return true; + return fallback; +} + +/** + * EPISODIC_MEMORY_REDACTION on (fork default) | off + * EPISODIC_MEMORY_REDACTION_RULES path to a custom rules file + * EPISODIC_MEMORY_REDACTION_STRICT 1 (default) | 0; strict fails closed on a rules-load error + */ +export function getRedactionSettings(env: NodeJS.ProcessEnv = process.env): RedactionSettings { + return { + enabled: parseToggle(env.EPISODIC_MEMORY_REDACTION, REDACTION_ENABLED_BY_DEFAULT), + strict: parseToggle(env.EPISODIC_MEMORY_REDACTION_STRICT, true), + rulesPath: env.EPISODIC_MEMORY_REDACTION_RULES || undefined, + }; +} + +function readRulesFile(settings: RedactionSettings): RedactionRulesFile { + let rulesPath = settings.rulesPath; + if (!rulesPath) { + const candidate = path.join(getSuperpowersDir(), REDACTION_RULES_FILENAME); + if (!fs.existsSync(candidate)) return {}; + rulesPath = candidate; + } + let raw: string; + try { + raw = fs.readFileSync(rulesPath, 'utf-8'); + } catch (error) { + throw new RedactionConfigError( + `Redaction rules failed to load from ${rulesPath}: ${error instanceof Error ? error.message : String(error)}` + ); + } + try { + return JSON.parse(raw) as RedactionRulesFile; + } catch (error) { + throw new RedactionConfigError( + `Redaction rules file ${rulesPath} is not valid JSON: ${error instanceof Error ? error.message : String(error)}` + ); + } +} + +function replaceById(base: T[], additions: T[]): T[] { + const out = [...base]; + for (const item of additions) { + const at = out.findIndex(existing => existing.id === item.id); + if (at >= 0) out[at] = item; + else out.push(item); + } + return out; +} + +/** + * Merge a rules file with the bundled defaults and validate the result. + * Throws RedactionConfigError on any invalid rule. + */ +export function loadRedactionConfig(file: RedactionRulesFile): RedactionConfig { + if (!file || typeof file !== 'object' || Array.isArray(file)) { + throw new RedactionConfigError('Redaction rules file must contain a JSON object'); + } + for (const key of ['rules', 'allowlist', 'disableRules'] as const) { + if (file[key] !== undefined && !Array.isArray(file[key])) { + throw new RedactionConfigError(`Redaction rules file: "${key}" must be an array`); + } + } + if (file.entropy !== undefined && (typeof file.entropy !== 'object' || file.entropy === null)) { + throw new RedactionConfigError('Redaction rules file: "entropy" must be an object'); + } + + const includeDefaults = file.includeDefaults !== false; + const disabled = new Set(file.disableRules ?? []); + const baseRules = includeDefaults ? DEFAULT_REDACTION_CONFIG.rules : []; + const baseAllowlist = includeDefaults ? DEFAULT_REDACTION_CONFIG.allowlist : []; + + const config: RedactionConfig = { + rules: replaceById(baseRules, file.rules ?? []).filter(rule => !disabled.has(rule.id)), + allowlist: replaceById(baseAllowlist, file.allowlist ?? []), + entropy: { ...DEFAULT_REDACTION_CONFIG.entropy, ...(file.entropy ?? {}) }, + }; + compileConfig(config); // validate + return config; +} + +/** + * Load the active redactor from the environment. + * + * Returns null when redaction is off. When the rules can't be loaded: strict + * mode (the default) throws RedactionConfigError so callers fail closed; + * non-strict mode warns and returns null, so text passes through unredacted. + */ +export function loadRedactor( + env: NodeJS.ProcessEnv = process.env, + warn: (message: string) => void = message => console.error(message) +): Redactor | null { + const settings = getRedactionSettings(env); + if (!settings.enabled) return null; + try { + return createRedactor(loadRedactionConfig(readRulesFile(settings))); + } catch (error) { + const err = error instanceof RedactionConfigError + ? error + : new RedactionConfigError(`Redaction rules failed to load: ${error instanceof Error ? error.message : String(error)}`); + if (settings.strict) throw err; + warn( + `episodic-memory: ${err.message}. EPISODIC_MEMORY_REDACTION_STRICT=0, so conversations ` + + 'will be archived and indexed WITHOUT redaction this run.' + ); + return null; + } +} + +// --------------------------------------------------------------------------- +// Compilation +// --------------------------------------------------------------------------- + +interface CompiledRule { + id: string; + regex: RegExp; + keywords?: string[]; + secretGroup: number; +} + +interface CompiledConfig { + rules: CompiledRule[]; + allowlist: RegExp[]; + entropy: EntropySpec; +} + +function checkFlags(flags: string | undefined, what: string): string { + const value = flags ?? ''; + if (typeof value !== 'string' || !ALLOWED_FLAGS.test(value)) { + throw new RedactionConfigError(`${what}: flags must be a combination of "imsu"`); + } + return value; +} + +function compileRule(spec: RedactionRuleSpec): CompiledRule { + if (!spec || typeof spec !== 'object') { + throw new RedactionConfigError('Redaction rule must be an object'); + } + if (typeof spec.id !== 'string' || !RULE_ID_PATTERN.test(spec.id)) { + throw new RedactionConfigError( + `Redaction rule id ${JSON.stringify(spec.id)} is invalid: use lowercase letters, digits and dashes` + ); + } + if (spec.id === ENTROPY_RULE_ID) { + throw new RedactionConfigError(`Redaction rule id "${ENTROPY_RULE_ID}" is reserved`); + } + const what = `Redaction rule "${spec.id}"`; + if (typeof spec.pattern !== 'string' || spec.pattern.length === 0) { + throw new RedactionConfigError(`${what}: pattern must be a non-empty string`); + } + const flags = checkFlags(spec.flags, what); + + let regex: RegExp; + let groupCount: number; + try { + regex = new RegExp(spec.pattern, flags + 'gd'); + groupCount = new RegExp(`(?:${spec.pattern})|`, flags).exec('')!.length - 1; + } catch (error) { + throw new RedactionConfigError(`${what}: invalid regular expression (${error instanceof Error ? error.message : String(error)})`); + } + if (new RegExp(spec.pattern, flags).test('')) { + throw new RedactionConfigError(`${what}: pattern matches the empty string`); + } + + const secretGroup = spec.secretGroup ?? 0; + if (!Number.isInteger(secretGroup) || secretGroup < 0 || secretGroup > groupCount) { + throw new RedactionConfigError(`${what}: secretGroup ${secretGroup} does not exist in the pattern (${groupCount} group(s))`); + } + + let keywords: string[] | undefined; + if (spec.keywords !== undefined) { + if (!Array.isArray(spec.keywords) || spec.keywords.some(k => typeof k !== 'string' || k.length === 0)) { + throw new RedactionConfigError(`${what}: keywords must be an array of non-empty strings`); + } + keywords = spec.keywords.length > 0 ? spec.keywords.map(k => k.toLowerCase()) : undefined; + } + + return { id: spec.id, regex, keywords, secretGroup }; +} + +function compileAllowlist(spec: AllowlistSpec): RegExp { + const what = `Redaction allowlist entry ${JSON.stringify(spec?.id)}`; + if (!spec || typeof spec.pattern !== 'string' || spec.pattern.length === 0) { + throw new RedactionConfigError(`${what}: pattern must be a non-empty string`); + } + const flags = checkFlags(spec.flags, what); + try { + return new RegExp(`^(?:${spec.pattern})$`, flags); + } catch (error) { + throw new RedactionConfigError(`${what}: invalid regular expression (${error instanceof Error ? error.message : String(error)})`); + } +} + +function compileConfig(config: RedactionConfig): CompiledConfig { + const seen = new Set(); + const rules = config.rules.map(spec => { + const rule = compileRule(spec); + if (seen.has(rule.id)) throw new RedactionConfigError(`Duplicate redaction rule id "${rule.id}"`); + seen.add(rule.id); + return rule; + }); + const allowlist = config.allowlist.map(compileAllowlist); + + const e = config.entropy; + if ( + typeof e.enabled !== 'boolean' || typeof e.requireKeyword !== 'boolean' || + !Number.isInteger(e.minLength) || e.minLength < 8 || + typeof e.threshold !== 'number' || !(e.threshold > 0) || + !Number.isInteger(e.window) || e.window < 0 || + !Array.isArray(e.keywords) || e.keywords.some(k => typeof k !== 'string' || !k) + ) { + throw new RedactionConfigError( + 'Redaction entropy settings are invalid (need enabled, requireKeyword: boolean; minLength >= 8; threshold > 0; window >= 0; keywords: string[])' + ); + } + return { rules, allowlist, entropy: { ...e, keywords: e.keywords.map(k => k.toLowerCase()) } }; +} + +// --------------------------------------------------------------------------- +// Redaction engine +// --------------------------------------------------------------------------- + +function tokenFor(ruleId: string): string { + return `${TOKEN_PREFIX}${ruleId}]`; +} + +function tokenSpans(text: string): Array<[number, number]> { + const spans: Array<[number, number]> = []; + for (const m of text.matchAll(TOKEN_PATTERN)) spans.push([m.index!, m.index! + m[0].length]); + return spans; +} + +function overlapsAny(spans: Array<[number, number]>, start: number, end: number): boolean { + for (const [s, e] of spans) if (start < e && end > s) return true; + return false; +} + +function shannonEntropy(text: string): number { + const counts = new Map(); + for (const ch of text) counts.set(ch, (counts.get(ch) ?? 0) + 1); + let entropy = 0; + for (const n of counts.values()) { + const p = n / text.length; + entropy -= p * Math.log2(p); + } + return entropy; +} + +/** + * Replace each accepted match (or its secretGroup) with a token. A candidate is + * skipped when it overlaps an existing token (idempotency) or fully matches an + * allowlist pattern. + */ +function applyRule(text: string, rule: CompiledRule, isAllowed: (s: string) => boolean): { text: string; count: number } { + const spans = text.includes(TOKEN_PREFIX) ? tokenSpans(text) : null; + const token = tokenFor(rule.id); + let out = ''; + let last = 0; + let count = 0; + for (const m of text.matchAll(rule.regex)) { + const range = m.indices?.[rule.secretGroup]; + if (!range) continue; + const [start, end] = range; + if (end <= start || start < last) continue; + if (spans && overlapsAny(spans, start, end)) continue; + if (isAllowed(text.slice(start, end))) continue; + out += text.slice(last, start) + token; + last = end; + count++; + } + return count === 0 ? { text, count } : { text: out + text.slice(last), count }; +} + +function applyEntropy(text: string, spec: EntropySpec, isAllowed: (s: string) => boolean): { text: string; count: number } { + const candidates = new RegExp(`[A-Za-z0-9+/=_.~-]{${spec.minLength},}`, 'g'); + const spans = text.includes(TOKEN_PREFIX) ? tokenSpans(text) : null; + const token = tokenFor(ENTROPY_RULE_ID); + let out = ''; + let last = 0; + let count = 0; + for (const m of text.matchAll(candidates)) { + const start = m.index!; + const end = start + m[0].length; + if (spans && overlapsAny(spans, start, end)) continue; + if (isAllowed(m[0])) continue; + if (shannonEntropy(m[0]) < spec.threshold) continue; + if (spec.requireKeyword) { + let before = text.slice(Math.max(0, start - spec.window), start); + before = before.slice(before.lastIndexOf('\n') + 1).toLowerCase(); + if (!spec.keywords.some(k => before.includes(k))) continue; + } + out += text.slice(last, start) + token; + last = end; + count++; + } + return count === 0 ? { text, count } : { text: out + text.slice(last), count }; +} + +export function createRedactor(config: RedactionConfig): Redactor { + const compiled = compileConfig(config); + const isAllowed = (secret: string) => compiled.allowlist.some(re => re.test(secret)); + + return { + ruleIds: compiled.rules.map(r => r.id), + redact(text: string, _ctx?: RedactionContext): RedactionResult { + if (typeof text !== 'string' || text.length === 0) return { text, findings: [] }; + let current = text; + let lower: string | null = null; + const counts = new Map(); + + for (const rule of compiled.rules) { + if (rule.keywords) { + lower ??= current.toLowerCase(); + if (!rule.keywords.some(k => lower!.includes(k))) continue; + } + const r = applyRule(current, rule, isAllowed); + if (r.count > 0) { + current = r.text; + lower = null; + counts.set(rule.id, (counts.get(rule.id) ?? 0) + r.count); + } + } + + if (compiled.entropy.enabled) { + const r = applyEntropy(current, compiled.entropy, isAllowed); + if (r.count > 0) { + current = r.text; + counts.set(ENTROPY_RULE_ID, r.count); + } + } + + return { text: current, findings: [...counts].map(([ruleId, count]) => ({ ruleId, count })) }; + }, + }; +} + +// --------------------------------------------------------------------------- +// Findings +// --------------------------------------------------------------------------- + +/** Aggregates findings by rule id. Never holds matched values. */ +export class FindingsTally { + private counts = new Map(); + + add(findings: RedactionFinding[]): void { + for (const f of findings) this.counts.set(f.ruleId, (this.counts.get(f.ruleId) ?? 0) + f.count); + } + + get total(): number { + let n = 0; + for (const c of this.counts.values()) n += c; + return n; + } + + toArray(): RedactionFinding[] { + return [...this.counts].map(([ruleId, count]) => ({ ruleId, count })); + } +} + +/** "7 value(s) redacted (azure-storage-key: 2, jwt: 5)" */ +export function formatFindings(findings: RedactionFinding[]): string { + const total = findings.reduce((n, f) => n + f.count, 0); + if (total === 0) return 'no values redacted'; + const detail = [...findings] + .sort((a, b) => b.count - a.count || a.ruleId.localeCompare(b.ruleId)) + .map(f => `${f.ruleId}: ${f.count}`) + .join(', '); + return `${total} value(s) redacted (${detail})`; +} + +// --------------------------------------------------------------------------- +// JSON / JSONL / files +// --------------------------------------------------------------------------- + +/** Redact every string value in a parsed JSON tree, in place. Keys are left alone. */ +function redactTree( + node: unknown, + redactor: Redactor, + ctx: RedactionContext | undefined, + tally: FindingsTally | undefined +): { value: unknown; changed: boolean } { + if (typeof node === 'string') { + const r = redactor.redact(node, ctx); + if (r.findings.length === 0) return { value: node, changed: false }; + tally?.add(r.findings); + return { value: r.text, changed: r.text !== node }; + } + if (node === null || typeof node !== 'object') return { value: node, changed: false }; + + let changed = false; + if (Array.isArray(node)) { + for (let i = 0; i < node.length; i++) { + const r = redactTree(node[i], redactor, ctx, tally); + if (r.changed) { node[i] = r.value; changed = true; } + } + return { value: node, changed }; + } + const obj = node as Record; + for (const key of Object.keys(obj)) { + const r = redactTree(obj[key], redactor, ctx, tally); + if (r.changed) { + Object.defineProperty(obj, key, { value: r.value, writable: true, enumerable: true, configurable: true }); + changed = true; + } + } + return { value: obj, changed }; +} + +/** + * Redact one JSONL line. Valid JSON stays valid (string values are redacted + * and the line is re-serialized only if something changed); a line that isn't + * JSON (e.g. a torn last line mid-write) is redacted as raw text. A trailing + * `\r` is preserved. + */ +export function redactJsonlLine( + line: string, + redactor: Redactor, + ctx?: RedactionContext, + tally?: FindingsTally +): string { + const cr = line.endsWith('\r') ? '\r' : ''; + const body = cr ? line.slice(0, -1) : line; + if (body.trim().length === 0) return line; + + let parsed: unknown; + try { + parsed = JSON.parse(body); + } catch { + const r = redactor.redact(body, ctx); + if (r.findings.length === 0) return line; + tally?.add(r.findings); + return r.text + cr; + } + + const r = redactTree(parsed, redactor, ctx, tally); + return r.changed ? JSON.stringify(r.value) + cr : line; +} + +const COPY_CHUNK_BYTES = 1 << 20; // 1 MiB + +/** + * Stream `src` to `dest`, redacting line by line. Writes `dest` directly; + * callers that need atomicity write to a temp path and rename. Line count and + * trailing-newline shape match the source exactly, so archive line numbers + * (exchanges.line_start/line_end, MCP read ranges) stay valid. + */ +export function copyFileRedacted( + src: string, + dest: string, + redactor: Redactor, + ctx?: RedactionContext, + tally?: FindingsTally +): void { + const fdIn = fs.openSync(src, 'r'); + let fdOut: number | undefined; + try { + fdOut = fs.openSync(dest, 'w'); + const buf = Buffer.allocUnsafe(COPY_CHUNK_BYTES); + const decoder = new StringDecoder('utf8'); + let pending = ''; + let bytesRead: number; + while ((bytesRead = fs.readSync(fdIn, buf, 0, buf.length, null)) > 0) { + let scanFrom = pending.length; + pending += decoder.write(buf.subarray(0, bytesRead)); + const out: string[] = []; + let start = 0; + let nl: number; + while ((nl = pending.indexOf('\n', scanFrom)) !== -1) { + out.push(redactJsonlLine(pending.slice(start, nl), redactor, ctx, tally), '\n'); + start = nl + 1; + scanFrom = start; + } + if (out.length > 0) fs.writeSync(fdOut, out.join('')); + pending = pending.slice(start); + } + pending += decoder.end(); + if (pending.length > 0) fs.writeSync(fdOut, redactJsonlLine(pending, redactor, ctx, tally)); + } finally { + fs.closeSync(fdIn); + if (fdOut !== undefined) fs.closeSync(fdOut); + } +} diff --git a/src/summarizer.ts b/src/summarizer.ts index d7e3cc04..23b0f31d 100644 --- a/src/summarizer.ts +++ b/src/summarizer.ts @@ -702,7 +702,23 @@ export function getCodexModel(_exchanges: ConversationExchange[]): string | unde return process.env.EPISODIC_MEMORY_CODEX_MODEL || undefined; } -export async function summarizeConversation(exchanges: ConversationExchange[], sessionId?: string): Promise { +export interface SummarizeOptions { + /** + * Allow Claude session resume and Codex thread/fork (default true). Both + * paths make the model read the *source* transcript rather than `exchanges`, + * so callers pass false when the exchanges were redacted (see + * docs/redaction/PHASE0-FINDINGS.md) to force the transcript-text path. + */ + allowResume?: boolean; +} + +export async function summarizeConversation( + exchanges: ConversationExchange[], + sessionId?: string, + options: SummarizeOptions = {} +): Promise { + const allowResume = options.allowResume !== false; + // Handle trivial conversations if (exchanges.length === 0) { return 'Trivial conversation with no substantive content.'; @@ -715,7 +731,7 @@ export async function summarizeConversation(exchanges: ConversationExchange[], s } } - const codexSessionId = getCodexSessionId(exchanges, sessionId); + const codexSessionId = allowResume ? getCodexSessionId(exchanges, sessionId) : undefined; if (codexSessionId) { try { const result = await callCodex(buildCodexSummaryPrompt(), codexSessionId, getCodexModel(exchanges)); @@ -742,7 +758,7 @@ export async function summarizeConversation(exchanges: ConversationExchange[], s const isClaudeSession = exchanges.some( e => e.harness === 'claude' || e.harness === undefined ); - const claudeSessionId = !codexSessionId && isClaudeSession ? sessionId : undefined; + const claudeSessionId = allowResume && !codexSessionId && isClaudeSession ? sessionId : undefined; const cwd = claudeSessionId ? exchanges.find(e => e.cwd)?.cwd : undefined; const conversationText = claudeSessionId ? '' // When resuming, no need to include conversation text - it's already in context diff --git a/src/sync-cli.ts b/src/sync-cli.ts index 297b20f8..b12ca46b 100644 --- a/src/sync-cli.ts +++ b/src/sync-cli.ts @@ -14,6 +14,7 @@ import { spawn } from 'child_process'; import fs from 'fs'; import { formatLogLine, getSyncLogPath, getSyncLockPath } from './logging.js'; import { acquireFileLock, readLockHolder, releaseFileLock } from './file-lock.js'; +import { FindingsTally, formatFindings, getRedactionSettings, loadRedactor, type Redactor } from './redaction.js'; const args = process.argv.slice(2); @@ -48,10 +49,13 @@ Sync conversations from Claude Code, Codex, and opencode transcript sources to a This command: 1. Exports opencode sessions from its SQLite database when available -2. Copies new or updated .jsonl files to conversation archive +2. Copies new or updated .jsonl files to conversation archive, redacting secrets 3. Generates embeddings for semantic search 4. Updates the search index +Secrets are replaced with [REDACTED:] tokens before anything is archived, +indexed, embedded, or summarized. See EPISODIC_MEMORY_REDACTION* in the README. + Only processes files that are new or have been modified since last sync. Safe to run multiple times - subsequent runs are fast no-ops. @@ -157,8 +161,23 @@ if (isBackground) { process.exit(0); } +// Load redaction rules before anything is exported, copied, or indexed. In +// strict mode (the default) a bad rules file stops the sync here: fail closed +// rather than archive unredacted text. +let redactor: Redactor | null; +try { + redactor = loadRedactor(); +} catch (error) { + console.error(`episodic-memory: ${error instanceof Error ? error.message : String(error)}`); + console.error( + 'episodic-memory: refusing to sync without redaction (strict mode). Fix the rules file, ' + + 'or set EPISODIC_MEMORY_REDACTION_STRICT=0 to sync unredacted, or EPISODIC_MEMORY_REDACTION=off.' + ); + process.exit(1); +} + if (!onlyHarnesses || onlyHarnesses.includes('opencode')) { - const opencodeExport = exportOpencodeSessions(); + const opencodeExport = exportOpencodeSessions({ redactor }); if (opencodeExport.exported > 0 || opencodeExport.skipped > 0) { console.log(`opencode export: ${opencodeExport.exported} exported, ${opencodeExport.skipped} skipped`); } @@ -215,8 +234,11 @@ console.log(`Destination: ${destDir}\n`); async function syncAll() { const totals = { copied: 0, skipped: 0, indexed: 0, summarized: 0, errors: [] as Array<{file: string; error: string}>, sourcesWithSummaryWork: 0, totalNeedingSummaries: 0 }; + const redactions = new FindingsTally(); + for (const sourceDir of sourceDirs) { - const result = await syncConversations(sourceDir, destDir, { ...syncOptions, summaryLimit }); + const result = await syncConversations(sourceDir, destDir, { ...syncOptions, summaryLimit, redactor }); + redactions.add(result.redactions); totals.copied += result.copied; totals.skipped += result.skipped; totals.indexed += result.indexed; @@ -233,6 +255,14 @@ async function syncAll() { } else { console.log(` Summarized: ${totals.summarized}`); } + // Rule IDs and counts only — never matched values. + if (redactor) { + console.log(` Redaction: ${formatFindings(redactions.toArray())}`); + } else if (getRedactionSettings().enabled) { + console.log(' Redaction: NOT APPLIED (rules failed to load; EPISODIC_MEMORY_REDACTION_STRICT=0)'); + } else { + console.log(' Redaction: off (EPISODIC_MEMORY_REDACTION=off)'); + } if (totals.errors.length > 0) { console.log(`\n⚠️ Errors: ${totals.errors.length}`); diff --git a/src/sync.ts b/src/sync.ts index d9ee6810..b2de456c 100644 --- a/src/sync.ts +++ b/src/sync.ts @@ -5,6 +5,7 @@ import { SUMMARIZER_CONTEXT_MARKER } from './constants.js'; import { getExcludedProjects, findJsonlFiles, statIfExists } from './paths.js'; import { formatErrorSentinel, shouldQueueForSummary } from './summary-sentinel.js'; import { getMaxMessageBytes, isOversizeExchange } from './message-size.js'; +import { copyFileRedacted, FindingsTally, loadRedactor, type RedactionFinding, type Redactor } from './redaction.js'; const EXCLUSION_MARKERS = [ 'DO NOT INDEX THIS CHAT', @@ -105,12 +106,19 @@ export interface SyncResult { indexed: number; summarized: number; errors: Array<{ file: string; error: string }>; + /** Rule IDs and counts only — never matched values. */ + redactions: RedactionFinding[]; } export interface SyncOptions { skipIndex?: boolean; skipSummaries?: boolean; summaryLimit?: number; // Max summaries to generate per run (default: 10) + /** + * Redactor applied at the archive write. `undefined` loads it from the + * environment (loadRedactor); `null` means redaction is off. + */ + redactor?: Redactor | null; } /** @@ -132,7 +140,18 @@ export function buildSyncOptionsFromEnv(env: NodeJS.ProcessEnv): SyncOptions { return { skipSummaries: env.EPISODIC_MEMORY_SKIP_SUMMARIES === '1' }; } -function copyIfNewer(src: string, dest: string): boolean { +/** + * Copy `src` into the archive at `dest` when the archive copy is missing or + * older. This is the redaction choke point: with a redactor, the copy is + * redacted line by line (see redaction.ts), and every downstream stage + * (index, embeddings, summaries, show/read) reads the archive. + */ +export function copyIfNewer( + src: string, + dest: string, + redactor: Redactor | null = null, + tally?: FindingsTally +): boolean { // Ensure destination directory exists const destDir = path.dirname(dest); if (!fs.existsSync(destDir)) { @@ -150,8 +169,17 @@ function copyIfNewer(src: string, dest: string): boolean { // Atomic copy: temp file + rename const tempDest = dest + '.tmp.' + process.pid; - fs.copyFileSync(src, tempDest); - fs.renameSync(tempDest, dest); // Atomic on same filesystem + try { + if (redactor) { + copyFileRedacted(src, tempDest, redactor, { source: 'archive', path: src }, tally); + } else { + fs.copyFileSync(src, tempDest); + } + fs.renameSync(tempDest, dest); // Atomic on same filesystem + } catch (error) { + try { fs.unlinkSync(tempDest); } catch {} + throw error; + } // Preserve source mtime: harnesses without per-message timestamps (Cursor // agent transcripts) fall back to file mtime. Round up to the next whole @@ -184,9 +212,16 @@ export async function syncConversations( skipped: 0, indexed: 0, summarized: 0, - errors: [] + errors: [], + redactions: [] }; + // Load the redactor before touching the archive. In strict mode (the + // default) a rules-load failure throws here, so nothing unredacted is + // archived, indexed, or summarized (fail closed). + const redactor = options.redactor === undefined ? loadRedactor() : options.redactor; + const tally = new FindingsTally(); + // Ensure source directory exists if (!fs.existsSync(sourceDir)) { return result; @@ -226,7 +261,7 @@ export async function syncConversations( continue; } - const wasCopied = copyIfNewer(srcFile, destFile); + const wasCopied = copyIfNewer(srcFile, destFile, redactor, tally); if (wasCopied) { result.copied++; filesToIndex.push(destFile); @@ -257,6 +292,8 @@ export async function syncConversations( } } + result.redactions = tally.toArray(); + // Index copied files (unless skipIndex is set) if (!options.skipIndex && filesToIndex.length > 0) { const { parseConversation } = await import('./parser.js'); @@ -395,7 +432,10 @@ export async function syncConversations( } console.log(` Summarizing ${path.basename(filePath)} (${exchanges.length} exchanges)...`); - const summary = await summarizeConversation(exchanges, sessionId); + // With redaction on, never resume/fork the session: those paths hand + // the model the unredacted source transcript instead of these + // (redacted) exchanges. + const summary = await summarizeConversation(exchanges, sessionId, { allowResume: redactor === null }); const summaryPath = filePath.replace('.jsonl', '-summary.txt'); fs.writeFileSync(summaryPath, summary, 'utf-8'); diff --git a/src/verify.ts b/src/verify.ts index 0dd2c5db..f0b86110 100644 --- a/src/verify.ts +++ b/src/verify.ts @@ -4,6 +4,7 @@ import { parseConversation } from './parser.js'; import { initDatabase, getAllExchanges, getFileLastIndexed } from './db.js'; import { getArchiveDir, getExcludedProjects, findJsonlFiles, statIfExists } from './paths.js'; import { isErroredSentinel } from './summary-sentinel.js'; +import { getRedactionSettings } from './redaction.js'; export interface VerificationResult { missing: Array<{ path: string; reason: string }>; @@ -160,7 +161,8 @@ export async function repairIndex(issues: VerificationResult): Promise { // Generate/update summary const summaryPath = conversationPath.replace('.jsonl', '-summary.txt'); - const summary = await summarizeConversation(exchanges); + // Under redaction, don't let the Codex fork fallback read the source rollout. + const summary = await summarizeConversation(exchanges, undefined, { allowResume: !getRedactionSettings().enabled }); fs.writeFileSync(summaryPath, summary, 'utf-8'); console.log(` Created summary: ${summary.split(/\s+/).length} words`); diff --git a/test/fake-secrets.ts b/test/fake-secrets.ts new file mode 100644 index 00000000..15deb68c --- /dev/null +++ b/test/fake-secrets.ts @@ -0,0 +1,146 @@ +/** + * Fake-secret builders for redaction tests. + * + * Nothing in this file is a real credential, and no string literal here matches + * a redaction rule on its own: every secret is assembled at runtime from a + * seeded PRNG plus split prefixes. That keeps GitHub push protection and other + * scanners quiet while still producing values in the exact real-world shape. + */ + +const UPPER = 'ABCDEFGHIJKLMNOPQRSTUVWXYZ'; +const LOWER = 'abcdefghijklmnopqrstuvwxyz'; +const DIGITS = '0123456789'; +const ALNUM = UPPER + LOWER + DIGITS; +const B64 = ALNUM + '+/'; +const B64URL = ALNUM + '-_'; +const BASE32 = UPPER + '234567'; + +/** Deterministic PRNG (mulberry32) so failures reproduce. */ +export function makeRng(seed: number): () => number { + let a = seed >>> 0; + return () => { + a = (a + 0x6d2b79f5) >>> 0; + let t = a; + t = Math.imul(t ^ (t >>> 15), t | 1); + t ^= t + Math.imul(t ^ (t >>> 7), t | 61); + return ((t ^ (t >>> 14)) >>> 0) / 4294967296; + }; +} + +export class FakeSecrets { + private rng: () => number; + + constructor(seed = 1337) { + this.rng = makeRng(seed); + } + + chars(charset: string, n: number): string { + let out = ''; + for (let i = 0; i < n; i++) out += charset[Math.floor(this.rng() * charset.length)]; + return out; + } + + /** Always contains at least one digit and one letter. */ + private mixed(charset: string, n: number): string { + return this.chars(charset, n - 2) + this.chars(DIGITS, 1) + this.chars(LOWER, 1); + } + + azureClientSecret(): string { + // <3 chars>Q~<34 chars> — Entra ID client secret shape. + return this.chars(ALNUM, 3) + this.chars(DIGITS, 1) + 'Q' + '~' + this.chars(ALNUM + '_~.-', 33) + 'x'; + } + + azureStorageKey(): string { + // 64 random bytes, base64: 86 chars + '=='. + return this.chars(ALNUM, 1) + this.chars(B64, 84) + this.chars(ALNUM, 1) + '=' + '='; + } + + sasSignature(): string { + return this.chars(ALNUM, 40) + '%2B' + this.chars(ALNUM, 3) + '%3D'; + } + + privateKeyBlock(kind = 'RSA'): string { + const body: string[] = []; + for (let i = 0; i < 6; i++) body.push(this.chars(B64, 64)); + const dashes = '-'.repeat(5); + return `${dashes}BEGIN ${kind} PRIVATE` + ` KEY${dashes}\n${body.join('\n')}\n${dashes}END ${kind} PRIVATE` + ` KEY${dashes}`; + } + + jwt(): string { + const head = 'ey' + 'J' + this.chars(B64URL, 30); + const payload = 'ey' + 'J' + this.chars(B64URL, 60); + return `${head}.${payload}.${this.chars(B64URL, 43)}`; + } + + anthropicKey(): string { + return 'sk-' + 'ant-' + 'api03-' + this.chars(B64URL, 93) + 'AA'; + } + + openAiKey(): string { + return 'sk-' + 'proj-' + this.chars(ALNUM, 20) + 'T3Blbk' + 'FJ' + this.chars(ALNUM, 20); + } + + githubToken(): string { + return 'gh' + 'p_' + this.chars(ALNUM, 36); + } + + githubFineGrainedPat(): string { + return 'github' + '_pat_' + this.chars(ALNUM, 22) + '_' + this.chars(ALNUM, 59); + } + + awsAccessKeyId(): string { + return 'AK' + 'IA' + this.chars(BASE32, 16); + } + + awsSecretAccessKey(): string { + return this.mixed(ALNUM + '/+', 40); + } + + slackToken(): string { + return 'xo' + 'xb-' + this.chars(DIGITS, 12) + '-' + this.chars(DIGITS, 12) + '-' + this.chars(ALNUM, 24); + } + + googleApiKey(): string { + return 'AI' + 'za' + this.chars(ALNUM + '_-', 35); + } + + npmToken(): string { + return 'np' + 'm_' + this.chars(ALNUM, 36); + } + + /** Starts with a letter so it never looks like a `$VAR`/`%VAR%` placeholder. */ + password(): string { + return this.chars(LOWER, 1) + this.mixed(ALNUM + '!#%^*', 15); + } + + bearerOpaque(): string { + return this.mixed(ALNUM, 40); + } + + basicAuth(): string { + return this.chars(ALNUM, 30) + '=' + '='; + } + + /** High-entropy blob with no recognizable prefix. */ + opaqueHighEntropy(n = 48): string { + return this.mixed(ALNUM + '+/', n); + } + + gitSha(): string { + return this.chars('0123456789abcdef', 40); + } + + guid(): string { + const h = (n: number) => this.chars('0123456789abcdef', n); + return `${h(8)}-${h(4)}-${h(4)}-${h(4)}-${h(12)}`; + } + + base64Blob(n: number): string { + return this.chars(B64, n); + } +} + +/** Every `[REDACTED:...]` token in a string. */ +export function redactionTokens(text: string): string[] { + return text.match(/\[REDACTED:[a-z0-9-]+\]/g) ?? []; +} diff --git a/test/redact-rewrite.test.ts b/test/redact-rewrite.test.ts new file mode 100644 index 00000000..10c63f9a --- /dev/null +++ b/test/redact-rewrite.test.ts @@ -0,0 +1,172 @@ +import { describe, it, expect, beforeEach, afterEach, vi } from 'vitest'; +import { mkdtempSync, mkdirSync, writeFileSync, readFileSync, rmSync, existsSync, statSync, utimesSync } from 'fs'; +import { join } from 'path'; +import { tmpdir } from 'os'; +import Database from 'better-sqlite3'; +import * as sqliteVec from 'sqlite-vec'; + +import { initDatabase, insertExchange } from '../src/db.js'; +import { parseConversation } from '../src/parser.js'; +import { rewriteArchive } from '../src/redact-rewrite.js'; +import { createRedactor, DEFAULT_REDACTION_CONFIG } from '../src/redaction.js'; +import { EMBEDDING_VERSION } from '../src/embedding-migration.js'; +import { FakeSecrets } from './fake-secrets.js'; + +const SESSION = '9a8b7c6d-1111-4222-8333-944455556666'; + +function transcript(secret: string, password: string, clean: string): string { + const base = { sessionId: SESSION, isSidechain: false, cwd: '/work/legacy' }; + return [ + { ...base, type: 'user', uuid: 'u1', parentUuid: null, timestamp: '2025-06-01T00:00:00.000Z', message: { role: 'user', content: `use AccountKey=${secret} please` } }, + { + ...base, type: 'assistant', uuid: 'a1', parentUuid: 'u1', timestamp: '2025-06-01T00:00:01.000Z', + message: { role: 'assistant', content: [ + { type: 'text', text: 'Running it.' }, + { type: 'tool_use', id: 'toolu_1', name: 'Bash', input: { command: `mysql --password=${password}` } }, + ] }, + }, + { ...base, type: 'user', uuid: 'u2', parentUuid: 'a1', timestamp: '2025-06-01T00:00:02.000Z', message: { role: 'user', content: [{ type: 'tool_result', tool_use_id: 'toolu_1', content: `ok password: ${password}` }] } }, + { ...base, type: 'assistant', uuid: 'a2', parentUuid: 'u2', timestamp: '2025-06-01T00:00:03.000Z', message: { role: 'assistant', content: [{ type: 'text', text: clean }] } }, + ].map(l => JSON.stringify(l)).join('\n') + '\n'; +} + +describe('redact --rewrite backfill', () => { + let root: string; + let archiveDir: string; + let dbPath: string; + const savedEnv = { ...process.env }; + const embed = vi.fn(async (..._args: unknown[]) => new Array(384).fill(0.5)); + + beforeEach(() => { + root = mkdtempSync(join(tmpdir(), 'episodic-memory-redact-rewrite-')); + archiveDir = join(root, 'archive'); + dbPath = join(root, 'db.sqlite'); + process.env.TEST_DB_PATH = dbPath; + embed.mockClear(); + }); + + afterEach(() => { + process.env = { ...savedEnv }; + rmSync(root, { recursive: true, force: true }); + }); + + /** Simulate a pre-redaction install: unredacted archive, rows, and summary. */ + async function seedLegacy(fake: FakeSecrets) { + const secret = fake.azureStorageKey(); + const password = fake.password(); + const projectDir = join(archiveDir, '-work-legacy'); + mkdirSync(join(projectDir, SESSION, 'subagents'), { recursive: true }); + const dirty = join(projectDir, `${SESSION}.jsonl`); + const clean = join(projectDir, 'clean-session.jsonl'); + const nested = join(projectDir, SESSION, 'subagents', 'agent-abc.jsonl'); + writeFileSync(dirty, transcript(secret, password, 'All done.')); + writeFileSync(nested, transcript(secret, password, 'Sub-agent done.')); + writeFileSync(clean, transcript('[REDACTED:connection-string-secret]', '[REDACTED:secret-assignment]', 'Nothing here.')); + writeFileSync(dirty.replace('.jsonl', '-summary.txt'), 'Configured storage access.'); + writeFileSync(clean.replace('.jsonl', '-summary.txt'), 'A clean summary.'); + const old = new Date('2025-06-02T00:00:00Z'); + for (const f of [dirty, clean, nested]) utimesSync(f, old, old); + + const db = initDatabase(); + for (const file of [dirty, clean, nested]) { + for (const ex of await parseConversation(file, '-work-legacy', file)) { + insertExchange(db, ex, new Array(384).fill(0), ex.toolCalls?.map(t => t.toolName)); + } + } + // Pretend these rows predate the current encoder bookkeeping. + db.prepare('UPDATE exchanges SET embedding_version = 0').run(); + db.close(); + return { secret, password, dirty, clean, nested }; + } + + function dbDump(): string { + const db = new Database(dbPath, { readonly: true }); + sqliteVec.load(db); + try { + return JSON.stringify([ + db.prepare('SELECT user_message, assistant_message FROM exchanges').all(), + db.prepare('SELECT tool_input, tool_result FROM tool_calls').all(), + ]); + } finally { + db.close(); + } + } + + it('rewrites the archive and index in place, then is a no-op on a second run', async () => { + const fake = new FakeSecrets(808); + const { secret, password, dirty, clean, nested } = await seedLegacy(fake); + const lineCount = readFileSync(dirty, 'utf-8').split('\n').length; + const mtimeBefore = statSync(dirty).mtimeMs; + const cleanBytes = readFileSync(clean, 'utf-8'); + expect(dbDump()).toContain(secret); + + const redactor = createRedactor(DEFAULT_REDACTION_CONFIG); + const result = await rewriteArchive({ archiveDir, redactor, embed }); + + // Archive: secrets gone, structure and mtime preserved, nested files covered. + for (const f of [dirty, nested]) { + const text = readFileSync(f, 'utf-8'); + expect(text).not.toContain(secret); + expect(text).not.toContain(password); + text.split('\n').filter(Boolean).forEach(l => expect(() => JSON.parse(l)).not.toThrow()); + } + expect(readFileSync(dirty, 'utf-8').split('\n').length).toBe(lineCount); + expect(Math.abs(statSync(dirty).mtimeMs - mtimeBefore)).toBeLessThan(2); + expect(readFileSync(clean, 'utf-8')).toBe(cleanBytes); + + // Summaries generated from unredacted text are removed; clean ones stay. + expect(existsSync(dirty.replace('.jsonl', '-summary.txt'))).toBe(false); + expect(existsSync(clean.replace('.jsonl', '-summary.txt'))).toBe(true); + + // Index: text columns clean, affected rows re-embedded with the current version. + const dump = dbDump(); + expect(dump).not.toContain(secret); + expect(dump).not.toContain(password); + expect(dump).toContain('[REDACTED:'); + expect(embed).toHaveBeenCalled(); + expectNoSecretInCalls(embed.mock.calls, [secret, password]); + + const db = new Database(dbPath, { readonly: true }); + // Only rows that had secrets are touched; the clean file's rows keep their old version. + const versions = db.prepare( + `SELECT embedding_version v, COUNT(*) n FROM exchanges WHERE archive_path IN (?, ?) AND (user_message LIKE '%REDACTED%' OR assistant_message LIKE '%REDACTED%') GROUP BY v` + ).all(dirty, nested) as Array<{ v: number }>; + const untouched = db.prepare('SELECT DISTINCT embedding_version v FROM exchanges WHERE archive_path = ?').all(clean) as Array<{ v: number }>; + db.close(); + expect(versions.map(r => r.v)).toEqual([EMBEDDING_VERSION]); + expect(untouched.map(r => r.v)).toEqual([0]); + + expect(result.filesRewritten).toBe(2); + expect(result.summariesRemoved).toBe(1); + expect(result.rowsUpdated).toBeGreaterThan(0); + expect(JSON.stringify(result)).not.toContain(secret); + + // Second run: nothing left to do. + embed.mockClear(); + const second = await rewriteArchive({ archiveDir, redactor, embed }); + expect(second.filesRewritten).toBe(0); + expect(second.rowsUpdated).toBe(0); + expect(second.summariesRemoved).toBe(0); + expect(embed).not.toHaveBeenCalled(); + }); + + it('--dry-run reports counts without touching anything', async () => { + const fake = new FakeSecrets(909); + const { secret, dirty } = await seedLegacy(fake); + const before = readFileSync(dirty, 'utf-8'); + + const result = await rewriteArchive({ archiveDir, redactor: createRedactor(DEFAULT_REDACTION_CONFIG), embed, dryRun: true }); + + expect(result.filesRewritten).toBe(2); + expect(result.rowsUpdated).toBeGreaterThan(0); + expect(readFileSync(dirty, 'utf-8')).toBe(before); + expect(existsSync(dirty.replace('.jsonl', '-summary.txt'))).toBe(true); + expect(dbDump()).toContain(secret); + expect(embed).not.toHaveBeenCalled(); + }); +}); + +function expectNoSecretInCalls(calls: unknown[][], secrets: string[]) { + const text = JSON.stringify(calls); + for (const s of secrets) expect(text.includes(s)).toBe(false); +} diff --git a/test/redaction-pipeline.test.ts b/test/redaction-pipeline.test.ts new file mode 100644 index 00000000..edd4d1cb --- /dev/null +++ b/test/redaction-pipeline.test.ts @@ -0,0 +1,353 @@ +import { describe, it, expect, beforeEach, afterEach, vi } from 'vitest'; +import { + mkdtempSync, mkdirSync, writeFileSync, readFileSync, rmSync, existsSync, readdirSync, statSync, +} from 'fs'; +import { join } from 'path'; +import { tmpdir } from 'os'; +import Database from 'better-sqlite3'; +import * as sqliteVec from 'sqlite-vec'; + +// Capture every embedding input and every summarizer call. sync.ts loads both +// via dynamic import and indexer.ts statically; vi.mock intercepts both. +const { embedSpy, summarizeSpy } = vi.hoisted(() => ({ + embedSpy: vi.fn(async (..._args: unknown[]) => new Array(384).fill(0)), + summarizeSpy: vi.fn(async (..._args: unknown[]) => 'A summary.'), +})); +vi.mock('../src/embeddings.js', () => ({ + initEmbeddings: vi.fn(async () => {}), + generateExchangeEmbedding: embedSpy, + generateQueryEmbedding: vi.fn(async () => new Array(384).fill(0)), + generateEmbedding: vi.fn(), + initEmbeddingsFailed: false, +})); +vi.mock('../src/summarizer.js', async () => { + const actual = await vi.importActual('../src/summarizer.js'); + return { ...actual, summarizeConversation: summarizeSpy }; +}); + +import { syncConversations } from '../src/sync.js'; +import { indexUnprocessed } from '../src/indexer.js'; +import { searchConversations } from '../src/search.js'; +import { formatConversationAsMarkdown } from '../src/show.js'; +import { exportOpencodeSessions } from '../src/opencode-sync.js'; +import { createRedactor, DEFAULT_REDACTION_CONFIG } from '../src/redaction.js'; +import { FakeSecrets } from './fake-secrets.js'; + +const SESSION = '4f1c2b3a-1111-4222-8333-944455556666'; + +interface Seed { + secrets: string[]; + sha: string; + guid: string; +} + +function seedSecrets(seed: number): Seed { + const fake = new FakeSecrets(seed); + return { + secrets: [ + fake.azureClientSecret(), + fake.azureStorageKey(), + fake.sasSignature(), + fake.githubToken(), + fake.anthropicKey(), + fake.password(), + fake.jwt(), + fake.privateKeyBlock().split('\n')[2], + ], + sha: fake.gitSha(), + guid: fake.guid(), + }; +} + +function claudeTranscript(s: Seed): string { + const [clientSecret, storageKey, sasSig, ghToken, antKey, password, jwt] = s.secrets; + const keyBlock = (() => { + // Rebuild the block whose body line is s.secrets[7]. + const dashes = '-'.repeat(5); + return `${dashes}BEGIN RSA PRIVATE` + ` KEY${dashes}\n${s.secrets[7]}\n${dashes}END RSA PRIVATE` + ` KEY${dashes}`; + })(); + const ts = (n: number) => new Date(Date.UTC(2026, 0, 1, 0, n)).toISOString(); + const base = { sessionId: SESSION, isSidechain: false, cwd: '/work/contoso', gitBranch: 'main', version: '2.0.0' }; + const lines = [ + { + ...base, type: 'user', uuid: 'u1', parentUuid: null, timestamp: ts(1), + message: { + role: 'user', + content: `Deploy with client secret ${clientSecret} for tenant ${s.guid}; last good commit ${s.sha}`, + }, + }, + { + ...base, type: 'assistant', uuid: 'a1', parentUuid: 'u1', timestamp: ts(2), + message: { + role: 'assistant', + content: [ + { type: 'text', text: 'Checking the storage account keys.' }, + { + type: 'tool_use', id: 'toolu_1', name: 'Bash', + input: { command: `az storage blob upload --connection-string "AccountName=contosodata;AccountKey=${storageKey}" --sas-token "sv=2022-11-02&sig=${sasSig}"` }, + }, + ], + }, + }, + { + ...base, type: 'user', uuid: 'u2', parentUuid: 'a1', timestamp: ts(3), + message: { + role: 'user', + content: [{ + type: 'tool_result', tool_use_id: 'toolu_1', + content: `GITHUB_TOKEN=${ghToken}\nANTHROPIC_API_KEY=${antKey}\npassword: ${password}\n${keyBlock}`, + }], + }, + }, + { + ...base, type: 'assistant', uuid: 'a2', parentUuid: 'u2', timestamp: ts(4), + message: { role: 'assistant', content: [{ type: 'text', text: `Uploaded. The bearer was Bearer ${jwt}.` }] }, + }, + ]; + return lines.map(l => JSON.stringify(l)).join('\n') + '\n'; +} + +function codexRollout(s: Seed): string { + const [clientSecret, storageKey] = s.secrets; + const lines = [ + { timestamp: '2026-05-12T18:00:00.000Z', type: 'session_meta', payload: { id: '019e4c75-d5bf-7c71-9df7-77f5fb86b711', cwd: '/work/contoso', cli_version: '0.130.0', model_provider: 'openai' } }, + { timestamp: '2026-05-12T18:00:02.000Z', type: 'response_item', payload: { type: 'message', role: 'user', content: [{ type: 'input_text', text: `Rotate ${clientSecret} please` }] } }, + { timestamp: '2026-05-12T18:00:04.000Z', type: 'response_item', payload: { type: 'function_call', name: 'exec_command', arguments: JSON.stringify({ cmd: `echo AccountKey=${storageKey}` }), call_id: 'c1' } }, + { timestamp: '2026-05-12T18:00:05.000Z', type: 'response_item', payload: { type: 'function_call_output', call_id: 'c1', output: `AccountKey=${storageKey}` } }, + { timestamp: '2026-05-12T18:00:06.000Z', type: 'response_item', payload: { type: 'message', role: 'assistant', content: [{ type: 'output_text', text: 'Rotated.' }] } }, + ]; + return lines.map(l => JSON.stringify(l)).join('\n') + '\n'; +} + +function walkFiles(dir: string): string[] { + if (!existsSync(dir)) return []; + const out: string[] = []; + for (const entry of readdirSync(dir, { withFileTypes: true })) { + const p = join(dir, entry.name); + if (entry.isDirectory()) out.push(...walkFiles(p)); + else out.push(p); + } + return out; +} + +function dbText(dbPath: string): string { + if (!existsSync(dbPath)) return ''; + const db = new Database(dbPath, { readonly: true }); + sqliteVec.load(db); + try { + const ex = db.prepare('SELECT user_message, assistant_message FROM exchanges').all(); + const tc = db.prepare('SELECT tool_input, tool_result FROM tool_calls').all(); + return JSON.stringify([ex, tc]); + } finally { + db.close(); + } +} + +function expectNoSecrets(haystack: string, s: Seed, where: string) { + for (const secret of s.secrets) { + expect(haystack.includes(secret), `${where} leaked a seeded secret`).toBe(false); + } +} + +describe('redaction pipeline', () => { + let root: string; + let sourceDir: string; + let archiveDir: string; + let dbPath: string; + let logged: string[]; + const savedEnv = { ...process.env }; + + beforeEach(() => { + root = mkdtempSync(join(tmpdir(), 'episodic-memory-redaction-pipeline-')); + sourceDir = join(root, 'source'); + archiveDir = join(root, 'archive'); + dbPath = join(root, 'db.sqlite'); + mkdirSync(join(sourceDir, '-work-contoso'), { recursive: true }); + process.env.TEST_DB_PATH = dbPath; + process.env.TEST_PROJECTS_DIR = sourceDir; + process.env.TEST_ARCHIVE_DIR = archiveDir; + process.env.EPISODIC_MEMORY_CONFIG_DIR = join(root, 'config'); + delete process.env.EPISODIC_MEMORY_REDACTION; + delete process.env.EPISODIC_MEMORY_REDACTION_RULES; + delete process.env.EPISODIC_MEMORY_REDACTION_STRICT; + delete process.env.EPISODIC_MEMORY_SKIP_SUMMARIES; + + embedSpy.mockClear(); + summarizeSpy.mockClear(); + logged = []; + for (const method of ['log', 'error', 'warn', 'info'] as const) { + vi.spyOn(console, method).mockImplementation((...args: unknown[]) => { + logged.push(args.map(a => (typeof a === 'string' ? a : JSON.stringify(a))).join(' ')); + }); + } + }); + + afterEach(() => { + vi.restoreAllMocks(); + process.env = { ...savedEnv }; + rmSync(root, { recursive: true, force: true }); + }); + + it('a seeded secret appears in zero of: archive, SQLite, embedding input, summarizer input, logs', async () => { + const s = seedSecrets(2024); + const claudeSrc = join(sourceDir, '-work-contoso', `${SESSION}.jsonl`); + writeFileSync(claudeSrc, claudeTranscript(s)); + writeFileSync(join(sourceDir, '-work-contoso', 'rollout-2026-05-12T18-00-00-019e4c75-d5bf-7c71-9df7-77f5fb86b711.jsonl'), codexRollout(s)); + + const result = await syncConversations(sourceDir, archiveDir); + expect(result.errors).toEqual([]); + expect(result.copied).toBe(2); + expect(result.indexed).toBe(2); + expect(result.summarized).toBe(2); + expect(result.redactions?.length).toBeGreaterThan(0); + expect(JSON.stringify(result.redactions)).not.toMatch(/[A-Za-z0-9+/]{30,}/); + + // (1) archive + const archived = walkFiles(archiveDir).filter(f => f.endsWith('.jsonl')); + expect(archived.length).toBe(2); + for (const file of archived) { + const text = readFileSync(file, 'utf-8'); + expectNoSecrets(text, s, `archive ${file}`); + expect(text).toContain('[REDACTED:'); + } + + // (2) SQLite text columns (exchanges + tool_calls) + const stored = dbText(dbPath); + expect(stored.length).toBeGreaterThan(0); + expectNoSecrets(stored, s, 'SQLite'); + expect(stored).toContain('[REDACTED:'); + + // (3) embedding input + expect(embedSpy).toHaveBeenCalled(); + expectNoSecrets(JSON.stringify(embedSpy.mock.calls), s, 'embedding input'); + + // (4) summarizer input — and resume/fork must be off, since those paths + // would hand the model the unredacted source transcript. + expect(summarizeSpy).toHaveBeenCalledTimes(2); + expectNoSecrets(JSON.stringify(summarizeSpy.mock.calls), s, 'summarizer input'); + for (const call of summarizeSpy.mock.calls) { + expect(call[2]).toMatchObject({ allowResume: false }); + } + + // (5) logs + expectNoSecrets(logged.join('\n'), s, 'console output'); + }); + + it('archive stays valid JSONL with the same line count, and show/read still render it', async () => { + const s = seedSecrets(7); + const src = join(sourceDir, '-work-contoso', `${SESSION}.jsonl`); + writeFileSync(src, claudeTranscript(s)); + await syncConversations(sourceDir, archiveDir, { skipSummaries: true }); + + const archived = join(archiveDir, '-work-contoso', `${SESSION}.jsonl`); + const srcLines = readFileSync(src, 'utf-8').split('\n'); + const outLines = readFileSync(archived, 'utf-8').split('\n'); + expect(outLines.length).toBe(srcLines.length); + outLines.filter(Boolean).forEach(l => expect(() => JSON.parse(l)).not.toThrow()); + + // The archive keeps the source mtime so the next sync treats it as current. + expect(Math.abs(statSync(archived).mtimeMs - statSync(src).mtimeMs)).toBeLessThan(2); + const again = await syncConversations(sourceDir, archiveDir, { skipSummaries: true }); + expect(again.copied).toBe(0); + + const md = formatConversationAsMarkdown(readFileSync(archived, 'utf-8')); + expect(md).toContain('[REDACTED:azure-client-secret]'); + expectNoSecrets(md, s, 'show output'); + }); + + it('git SHAs and GUIDs remain findable via text search', async () => { + const s = seedSecrets(31); + writeFileSync(join(sourceDir, '-work-contoso', `${SESSION}.jsonl`), claudeTranscript(s)); + await syncConversations(sourceDir, archiveDir, { skipSummaries: true }); + + const bySha = await searchConversations(s.sha, { mode: 'text' }); + expect(bySha.length).toBe(1); + const byGuid = await searchConversations(s.guid, { mode: 'text' }); + expect(byGuid.length).toBe(1); + const byToken = await searchConversations('[REDACTED:azure-client-secret]', { mode: 'text' }); + expect(byToken.length).toBe(1); + }); + + it('fails closed in strict mode: a corrupt rules file leaves no new archive or index content', async () => { + const s = seedSecrets(5); + writeFileSync(join(sourceDir, '-work-contoso', `${SESSION}.jsonl`), claudeTranscript(s)); + const rules = join(root, 'rules.json'); + writeFileSync(rules, '{ "rules": [ { "id": "x", "pattern": "(" } ] }'); + process.env.EPISODIC_MEMORY_REDACTION_RULES = rules; + + await expect(syncConversations(sourceDir, archiveDir)).rejects.toThrow(/redaction/i); + expect(walkFiles(archiveDir)).toEqual([]); + expect(dbText(dbPath)).toBe(''); + expect(embedSpy).not.toHaveBeenCalled(); + expect(summarizeSpy).not.toHaveBeenCalled(); + }); + + it('EPISODIC_MEMORY_REDACTION=off restores the byte-for-byte copy and resume-capable summaries', async () => { + process.env.EPISODIC_MEMORY_REDACTION = 'off'; + const s = seedSecrets(8); + const src = join(sourceDir, '-work-contoso', `${SESSION}.jsonl`); + writeFileSync(src, claudeTranscript(s)); + await syncConversations(sourceDir, archiveDir); + expect(readFileSync(join(archiveDir, '-work-contoso', `${SESSION}.jsonl`), 'utf-8')).toBe(readFileSync(src, 'utf-8')); + expect(summarizeSpy.mock.calls[0][2]).toMatchObject({ allowResume: true }); + }); + + it('the `index` path (indexUnprocessed) parses the redacted archive, not the source', async () => { + const s = seedSecrets(77); + writeFileSync(join(sourceDir, '-work-contoso', `${SESSION}.jsonl`), claudeTranscript(s)); + + await indexUnprocessed(1, false); + + const archived = walkFiles(archiveDir).filter(f => f.endsWith('.jsonl')); + expect(archived.length).toBe(1); + expectNoSecrets(readFileSync(archived[0], 'utf-8'), s, 'archive (index path)'); + expectNoSecrets(dbText(dbPath), s, 'SQLite (index path)'); + expectNoSecrets(JSON.stringify(embedSpy.mock.calls), s, 'embedding input (index path)'); + expect(summarizeSpy).toHaveBeenCalledTimes(1); + expectNoSecrets(JSON.stringify(summarizeSpy.mock.calls), s, 'summarizer input (index path)'); + expect(summarizeSpy.mock.calls[0][2]).toMatchObject({ allowResume: false }); + }); +}); + +describe('redaction: opencode staging export', () => { + let root: string; + const savedEnv = { ...process.env }; + + beforeEach(() => { + root = mkdtempSync(join(tmpdir(), 'episodic-memory-redaction-opencode-')); + process.env.EPISODIC_MEMORY_OPENCODE_DB_PATH = join(root, 'opencode.db'); + process.env.EPISODIC_MEMORY_OPENCODE_TRANSCRIPT_DIR = join(root, 'transcripts'); + }); + + afterEach(() => { + process.env = { ...savedEnv }; + rmSync(root, { recursive: true, force: true }); + }); + + it('redacts secrets before the staging JSONL is written', () => { + const fake = new FakeSecrets(55); + const secret = fake.azureStorageKey(); + const db = new Database(join(root, 'opencode.db')); + db.exec(` + CREATE TABLE project (id TEXT PRIMARY KEY, worktree TEXT NOT NULL, name TEXT, time_created INTEGER NOT NULL, time_updated INTEGER NOT NULL, sandboxes TEXT NOT NULL); + CREATE TABLE session (id TEXT PRIMARY KEY, project_id TEXT NOT NULL, slug TEXT NOT NULL, directory TEXT NOT NULL, title TEXT NOT NULL, version TEXT NOT NULL, time_created INTEGER NOT NULL, time_updated INTEGER NOT NULL, agent TEXT, model TEXT); + CREATE TABLE message (id TEXT PRIMARY KEY, session_id TEXT NOT NULL, time_created INTEGER NOT NULL, time_updated INTEGER NOT NULL, data TEXT NOT NULL); + CREATE TABLE part (id TEXT PRIMARY KEY, message_id TEXT NOT NULL, session_id TEXT NOT NULL, time_created INTEGER NOT NULL, time_updated INTEGER NOT NULL, data TEXT NOT NULL); + `); + db.prepare(`INSERT INTO project VALUES ('p1', '/work/x', 'X', 1700000000000, 1700000000000, '[]')`).run(); + db.prepare(`INSERT INTO session VALUES ('ses_1', 'p1', 's', '/work/x', 'T', '1.0.0', 1700000000000, 1700000004000, 'build', NULL)`).run(); + db.prepare(`INSERT INTO message VALUES ('m1', 'ses_1', 1700000001000, 1700000001000, ?)`).run(JSON.stringify({ role: 'user' })); + db.prepare(`INSERT INTO part VALUES ('pt1', 'm1', 'ses_1', 1700000001000, 1700000001000, ?)`).run( + JSON.stringify({ type: 'text', text: `AccountKey=${secret}` }) + ); + db.close(); + + const result = exportOpencodeSessions({ redactor: createRedactor(DEFAULT_REDACTION_CONFIG) }); + expect(result.exported).toBe(1); + const files = walkFiles(join(root, 'transcripts')); + expect(files.length).toBe(1); + const text = readFileSync(files[0], 'utf-8'); + expect(text).not.toContain(secret); + expect(text).toContain('[REDACTED:connection-string-secret]'); + }); +}); diff --git a/test/redaction.test.ts b/test/redaction.test.ts new file mode 100644 index 00000000..95466d4e --- /dev/null +++ b/test/redaction.test.ts @@ -0,0 +1,526 @@ +import { describe, it, expect, beforeEach, afterEach, vi } from 'vitest'; +import { mkdtempSync, writeFileSync, readFileSync, rmSync, mkdirSync } from 'fs'; +import { join } from 'path'; +import { tmpdir } from 'os'; + +import { + createRedactor, + loadRedactor, + loadRedactionConfig, + getRedactionSettings, + redactJsonlLine, + copyFileRedacted, + FindingsTally, + RedactionConfigError, + DEFAULT_REDACTION_CONFIG, + type Redactor, +} from '../src/redaction.js'; +import { FakeSecrets, redactionTokens } from './fake-secrets.js'; + +const ctx = { source: 'test', path: '/tmp/x.jsonl' }; + +function defaults(): Redactor { + return createRedactor(DEFAULT_REDACTION_CONFIG); +} + +/** Assert `secret` is gone, the expected token is present, and findings name the rule. */ +function expectRedacted(redactor: Redactor, text: string, secret: string, ruleId: string) { + const result = redactor.redact(text, ctx); + expect(result.text).not.toContain(secret); + expect(result.text).toContain(`[REDACTED:${ruleId}]`); + expect(result.findings.find(f => f.ruleId === ruleId)?.count ?? 0).toBeGreaterThan(0); + // Findings carry rule IDs and counts, never the matched value. + expect(JSON.stringify(result.findings)).not.toContain(secret); + return result; +} + +describe('redaction: default rules (positive)', () => { + let fake: FakeSecrets; + let redactor: Redactor; + + beforeEach(() => { + fake = new FakeSecrets(42); + redactor = defaults(); + }); + + it('azure-client-secret in pasted `az ad sp credential reset` output', () => { + const secret = fake.azureClientSecret(); + const text = `{\n "appId": "${fake.guid()}",\n "password": "${secret}",\n "tenant": "${fake.guid()}"\n}`; + expectRedacted(redactor, text, secret, 'azure-client-secret'); + }); + + it('azure-client-secret bare in prose', () => { + const secret = fake.azureClientSecret(); + expectRedacted(redactor, `use ${secret} as the client secret`, secret, 'azure-client-secret'); + }); + + it('azure-storage-key in `az storage account keys list` output', () => { + const secret = fake.azureStorageKey(); + const text = `[\n {\n "keyName": "key1",\n "permissions": "FULL",\n "value": "${secret}"\n }\n]`; + expectRedacted(redactor, text, secret, 'azure-storage-key'); + }); + + it('connection-string-secret redacts only the value, keeping account and endpoint searchable', () => { + const key = fake.azureStorageKey(); + const text = `DefaultEndpointsProtocol=https;AccountName=contosodata;AccountKey=${key};EndpointSuffix=core.windows.net`; + const result = expectRedacted(redactor, text, key, 'connection-string-secret'); + expect(result.text).toContain('AccountName=contosodata'); + expect(result.text).toContain('EndpointSuffix=core.windows.net'); + }); + + it('connection-string-secret in an appsettings.json blob (SQL Server password field)', () => { + const pw = fake.password(); + const text = JSON.stringify({ + ConnectionStrings: { + Default: `Server=tcp:contoso-sql.database.windows.net,1433;Database=orders;User ID=app;Password=${pw};Encrypt=True;`, + }, + }, null, 2); + const result = expectRedacted(redactor, text, pw, 'connection-string-secret'); + expect(result.text).toContain('contoso-sql.database.windows.net'); + expect(result.text).toContain('Database=orders'); + }); + + it('connection-string-secret for Service Bus SharedAccessKey', () => { + const key = fake.chars('ABCDEFGHIJKLMNOPQRSTUVWXYZabcdefghijklmnopqrstuvwxyz0123456789+/', 43) + '='; + const text = `Endpoint=sb://contoso.servicebus.windows.net/;SharedAccessKeyName=RootManageSharedAccessKey;SharedAccessKey=${key}`; + const result = expectRedacted(redactor, text, key, 'connection-string-secret'); + expect(result.text).toContain('SharedAccessKeyName=RootManageSharedAccessKey'); + }); + + it('azure-sas-token redacts the sig while keeping the blob URL', () => { + const sig = fake.sasSignature(); + const text = `https://contoso.blob.core.windows.net/backups/db.bak?sv=2022-11-02&ss=b&srt=co&sp=rl&se=2026-12-31T00:00:00Z&sig=${sig}`; + const result = expectRedacted(redactor, text, sig, 'azure-sas-token'); + expect(result.text).toContain('https://contoso.blob.core.windows.net/backups/db.bak?sv=2022-11-02'); + }); + + it('private-key-block (multi-line, PEM)', () => { + const block = fake.privateKeyBlock('RSA'); + const body = block.split('\n')[2]; + const result = expectRedacted(redactor, `cat key.pem\n${block}\n$ `, body, 'private-key-block'); + expect(result.text).not.toContain('PRIVATE KEY'); + }); + + it('private-key-block (OPENSSH, escaped newlines as in nested JSON)', () => { + const block = fake.privateKeyBlock('OPENSSH').replace(/\n/g, '\\n'); + const body = block.split('\\n')[3]; + expectRedacted(redactor, `{"content":"${block}"}`, body, 'private-key-block'); + }); + + it('private-key-block truncated (no END marker) is still redacted to end of text', () => { + const block = fake.privateKeyBlock('EC'); + const truncated = block.split('\n').slice(0, 4).join('\n'); + const body = truncated.split('\n')[2]; + expectRedacted(redactor, truncated, body, 'private-key-block'); + }); + + it('jwt (Azure access token from `az account get-access-token`)', () => { + const token = fake.jwt(); + const text = `{\n "accessToken": "${token}",\n "expiresOn": "2026-10-04 12:00:00.000000",\n "tokenType": "Bearer"\n}`; + const result = redactor.redact(text, ctx); + expect(result.text).not.toContain(token); + expect(redactionTokens(result.text).length).toBeGreaterThan(0); + }); + + it('bearer-token in an Authorization header', () => { + const token = fake.bearerOpaque(); + expectRedacted(redactor, `curl -H "Authorization: Bearer ${token}" https://api.example.com`, token, 'bearer-token'); + }); + + it('basic-auth in an Authorization header', () => { + const token = fake.basicAuth(); + expectRedacted(redactor, `Authorization: Basic ${token}`, token, 'basic-auth'); + }); + + it('anthropic-api-key', () => { + const key = fake.anthropicKey(); + expectRedacted(redactor, `export ANTHROPIC_API_KEY=${key}`, key, 'anthropic-api-key'); + }); + + it('openai-api-key', () => { + const key = fake.openAiKey(); + expectRedacted(redactor, `OPENAI_API_KEY="${key}"`, key, 'openai-api-key'); + }); + + it('github-token (classic) and fine-grained PAT', () => { + const classic = fake.githubToken(); + expectRedacted(redactor, `git remote set-url origin https://${classic}@github.com/o/r.git`, classic, 'github-token'); + const fine = fake.githubFineGrainedPat(); + expectRedacted(redactor, `GH_TOKEN=${fine}`, fine, 'github-token'); + }); + + it('aws-access-key-id and aws-secret-access-key in ~/.aws/credentials', () => { + const id = fake.awsAccessKeyId(); + const secret = fake.awsSecretAccessKey(); + const text = `[default]\naws_access_key_id = ${id}\naws_secret_access_key = ${secret}\n`; + expectRedacted(redactor, text, id, 'aws-access-key-id'); + expectRedacted(redactor, text, secret, 'aws-secret-access-key'); + }); + + it('slack-token, google-api-key, npm-token', () => { + const slack = fake.slackToken(); + expectRedacted(redactor, `SLACK_BOT_TOKEN=${slack}`, slack, 'slack-token'); + const google = fake.googleApiKey(); + expectRedacted(redactor, `key=${google}&q=coffee`, google, 'google-api-key'); + const npm = fake.npmToken(); + expectRedacted(redactor, `//registry.npmjs.org/:_authToken=${npm}`, npm, 'npm-token'); + }); + + it('url-credentials redacts only the password', () => { + const pw = fake.password().replace(/[#%^*!]/g, 'x'); + const text = `DATABASE_URL=postgres://app_user:${pw}@db.internal:5432/orders`; + const result = expectRedacted(redactor, text, pw, 'url-credentials'); + expect(result.text).toContain('app_user'); + expect(result.text).toContain('@db.internal:5432/orders'); + }); + + it('secret-assignment covers env, YAML, and JSON shapes (decrypted SOPS output)', () => { + const a = fake.password(); + const b = fake.password(); + const c = fake.password(); + const yaml = `database:\n host: db.internal\n password: ${a}\nazure:\n client_secret: "${b}"\n`; + const env = `AZURE_CLIENT_SECRET=${c}`; + for (const [text, secret] of [[yaml, a], [yaml, b], [env, c]] as const) { + const result = redactor.redact(text, ctx); + expect(result.text).not.toContain(secret); + } + expect(redactor.redact(yaml, ctx).text).toContain('host: db.internal'); + }); + + it('secret-assignment redacts a JSON "clientSecret" field', () => { + const secret = fake.password(); + expectRedacted(redactor, JSON.stringify({ clientId: fake.guid(), clientSecret: secret }), secret, 'secret-assignment'); + }); +}); + +describe('redaction: allowlist and false-positive resistance (negative)', () => { + let fake: FakeSecrets; + let redactor: Redactor; + + beforeEach(() => { + fake = new FakeSecrets(7); + redactor = defaults(); + }); + + function expectUntouched(text: string) { + const result = redactor.redact(text, ctx); + expect(result.text).toBe(text); + expect(result.findings).toEqual([]); + } + + it('git SHAs (full and short) survive', () => { + const sha = fake.gitSha(); + expectUntouched(`commit ${sha}\nAuthor: someone\n\n fix: thing (${sha.slice(0, 7)})`); + }); + + it('a git SHA assigned to a token-like key is still allowlisted', () => { + const sha = fake.gitSha(); + expectUntouched(`token: ${sha}`); + }); + + it('GUIDs (tenant, client, object IDs) survive', () => { + expectUntouched(JSON.stringify({ tenantId: fake.guid(), clientId: fake.guid(), objectId: fake.guid() })); + }); + + it('ARM resource IDs survive', () => { + expectUntouched( + `/subscriptions/${fake.guid()}/resourceGroups/rg-prod-eastus/providers/Microsoft.Storage/storageAccounts/contosodata` + ); + }); + + it('error codes and HTTP statuses survive', () => { + expectUntouched('AADSTS7000215: Invalid client secret provided. Status: 401 Unauthorized. AuthorizationFailed (403)'); + }); + + it('base64 image data survives', () => { + expectUntouched(JSON.stringify({ type: 'image', source: { type: 'base64', media_type: 'image/png', data: 'iVBORw0KGgo' + fake.base64Blob(20000) + '==' } })); + }); + + it('npm integrity hashes (sha512, 88 chars) survive', () => { + expectUntouched(`"integrity": "sha512-${fake.base64Blob(86)}=="`); + }); + + it('minified JS survives', () => { + expectUntouched( + 'function a(e,t){var n=e.password,r=t.token;return n&&r?{password:n,token:r,expires_in:3600}:null}' + + 'const s=new URLSearchParams({grant_type:"client_credentials"});' + ); + }); + + it('code that reads secrets (not literals) survives', () => { + expectUntouched('const password = getPassword();\nconst token = process.env.GITHUB_TOKEN;\npassword = os.environ["DB_PASSWORD"]'); + }); + + it('member-access references to secrets survive, even with digits (env.AUTH0_CLIENT_SECRET)', () => { + expectUntouched('client_id: env.AUTH0_CLIENT_ID,\n client_secret: env.AUTH0_CLIENT_SECRET,\n token: this.config.apiToken2;'); + }); + + it('placeholders and SOPS-encrypted values survive', () => { + expectUntouched('password: ${DB_PASSWORD}\nclient_secret: \napi_key: ENC[AES256_GCM,data:abc,iv:def,tag:ghi,type:str]'); + }); + + it('PWD and other env noise survive', () => { + expectUntouched('PWD=/home/user1/src/project2\nOLDPWD=/home/user1\nmax_tokens: 4096\ntokenizer: bert-base-uncased'); + }); + + it('prose about tokens and passwords survives', () => { + expectUntouched('Rotate the bearer token every 90 days. Basic authentication is disabled. The password policy requires 14 characters.'); + }); +}); + +describe('redaction: idempotency', () => { + it('redacting already-redacted text is a no-op and tokens match no rule', () => { + const fake = new FakeSecrets(99); + const redactor = defaults(); + const text = [ + `AccountKey=${fake.azureStorageKey()}`, + `password: ${fake.password()}`, + `Authorization: Bearer ${fake.jwt()}`, + fake.privateKeyBlock(), + `client secret ${fake.azureClientSecret()}`, + `https://u:${fake.password().replace(/[#%^*!]/g, 'y')}@host/x`, + ].join('\n'); + const once = redactor.redact(text, ctx); + expect(once.findings.length).toBeGreaterThan(0); + const twice = redactor.redact(once.text, ctx); + expect(twice.text).toBe(once.text); + expect(twice.findings).toEqual([]); + }); + + it('every token for every default rule id is inert', () => { + const redactor = defaults(); + for (const id of redactor.ruleIds) { + for (const shape of [`[REDACTED:${id}]`, `password: [REDACTED:${id}]`, `Bearer [REDACTED:${id}]`, `AccountKey=[REDACTED:${id}];`]) { + const result = redactor.redact(shape, ctx); + expect(result.findings, `${id} in ${shape}`).toEqual([]); + } + } + }); +}); + +describe('redaction: entropy fallback', () => { + it('is off by default', () => { + const fake = new FakeSecrets(3); + const blob = fake.opaqueHighEntropy(); + const result = defaults().redact(`the signing value is ${blob}`, ctx); + expect(result.text).toContain(blob); + }); + + it('when enabled, fires near a keyword but not elsewhere or on allowlisted shapes', () => { + const fake = new FakeSecrets(4); + const redactor = createRedactor({ + ...DEFAULT_REDACTION_CONFIG, + entropy: { ...DEFAULT_REDACTION_CONFIG.entropy, enabled: true }, + }); + const blob = fake.opaqueHighEntropy(48); + const near = redactor.redact(`signing key is ${blob}`, ctx); + expect(near.text).not.toContain(blob); + expect(near.text).toContain('[REDACTED:high-entropy]'); + + const far = redactor.redact(`the build artifact digest ${fake.opaqueHighEntropy(48)}`, ctx); + expect(far.findings).toEqual([]); + + const sha = fake.gitSha(); + expect(redactor.redact(`key ${sha}`, ctx).text).toContain(sha); + }); +}); + +describe('redaction: JSONL line handling', () => { + const redactor = defaults(); + const fake = new FakeSecrets(11); + + it('redacts string values inside JSON and keeps the line valid JSON', () => { + const key = fake.azureStorageKey(); + const line = JSON.stringify({ + type: 'user', + message: { role: 'user', content: [{ type: 'tool_result', tool_use_id: 't1', content: `AccountKey=${key}` }] }, + uuid: 'u1', + }); + const out = redactJsonlLine(line, redactor, ctx); + expect(out).not.toContain(key); + const parsed = JSON.parse(out); + expect(parsed.uuid).toBe('u1'); + expect(parsed.message.content[0].content).toBe('AccountKey=[REDACTED:connection-string-secret]'); + }); + + it('redacts tool inputs (nested objects) too', () => { + const pw = fake.password(); + const line = JSON.stringify({ + type: 'assistant', + message: { role: 'assistant', content: [{ type: 'tool_use', id: 't1', name: 'Bash', input: { command: `mysql -u root --password=${pw}` } }] }, + }); + const out = redactJsonlLine(line, redactor, ctx); + expect(out).not.toContain(pw); + expect(() => JSON.parse(out)).not.toThrow(); + }); + + it('returns unchanged lines byte-for-byte (formatting preserved)', () => { + const line = '{"type":"user", "message":{"role":"user","content":"caf\\u00e9 hello"}}'; + expect(redactJsonlLine(line, redactor, ctx)).toBe(line); + }); + + it('redacts non-JSON lines (e.g. a torn last line) as raw text', () => { + const key = fake.azureStorageKey(); + const torn = `{"type":"user","message":{"content":"AccountKey=${key}`; + const out = redactJsonlLine(torn, redactor, ctx); + expect(out).not.toContain(key); + }); + + it('preserves CRLF line endings', () => { + const pw = fake.password(); + const line = JSON.stringify({ content: `password: ${pw}` }) + '\r'; + const out = redactJsonlLine(line, redactor, ctx); + expect(out.endsWith('\r')).toBe(true); + expect(out).not.toContain(pw); + }); + + it('tallies findings by rule id', () => { + const tally = new FindingsTally(); + redactJsonlLine(JSON.stringify({ a: `password: ${fake.password()}`, b: `password: ${fake.password()}` }), redactor, ctx, tally); + expect(tally.toArray()).toEqual([{ ruleId: 'secret-assignment', count: 2 }]); + expect(tally.total).toBe(2); + }); +}); + +describe('redaction: copyFileRedacted', () => { + let dir: string; + beforeEach(() => { dir = mkdtempSync(join(tmpdir(), 'episodic-memory-redact-copy-')); }); + afterEach(() => { rmSync(dir, { recursive: true, force: true }); }); + + it('streams a file, keeps one output line per input line, and keeps the trailing-newline shape', () => { + const fake = new FakeSecrets(5); + const redactor = defaults(); + const secrets: string[] = []; + const lines: string[] = []; + for (let i = 0; i < 3000; i++) { + if (i % 500 === 0) { + const s = fake.azureClientSecret(); + secrets.push(s); + lines.push(JSON.stringify({ type: 'user', i, message: { content: `secret is ${s}` } })); + } else { + lines.push(JSON.stringify({ type: 'assistant', i, message: { content: 'x'.repeat(i % 700) } })); + } + } + const src = join(dir, 'src.jsonl'); + const dest = join(dir, 'dest.jsonl'); + writeFileSync(src, lines.join('\n')); // no trailing newline + const tally = new FindingsTally(); + copyFileRedacted(src, dest, redactor, ctx, tally); + const out = readFileSync(dest, 'utf-8'); + for (const s of secrets) expect(out).not.toContain(s); + expect(out.split('\n').length).toBe(lines.length); + expect(out.endsWith('\n')).toBe(false); + expect(tally.total).toBe(secrets.length); + out.split('\n').forEach(l => expect(() => JSON.parse(l)).not.toThrow()); + }); + + it('handles multi-byte UTF-8 across chunk boundaries', () => { + const redactor = defaults(); + const src = join(dir, 'utf8.jsonl'); + const dest = join(dir, 'utf8-out.jsonl'); + const line = JSON.stringify({ content: '日本語テキスト🙂'.repeat(200_000) }); + writeFileSync(src, line + '\n' + line + '\n'); + copyFileRedacted(src, dest, redactor, ctx); + expect(readFileSync(dest, 'utf-8')).toBe(readFileSync(src, 'utf-8')); + }); +}); + +describe('redaction: settings and rules loading', () => { + let dir: string; + const saved = { ...process.env }; + + beforeEach(() => { + dir = mkdtempSync(join(tmpdir(), 'episodic-memory-redact-cfg-')); + process.env.EPISODIC_MEMORY_CONFIG_DIR = join(dir, 'config'); + mkdirSync(process.env.EPISODIC_MEMORY_CONFIG_DIR, { recursive: true }); + delete process.env.EPISODIC_MEMORY_REDACTION; + delete process.env.EPISODIC_MEMORY_REDACTION_RULES; + delete process.env.EPISODIC_MEMORY_REDACTION_STRICT; + }); + + afterEach(() => { + process.env = { ...saved }; + rmSync(dir, { recursive: true, force: true }); + }); + + it('is on and strict by default', () => { + expect(getRedactionSettings(process.env)).toMatchObject({ enabled: true, strict: true }); + }); + + it('EPISODIC_MEMORY_REDACTION=off disables it', () => { + process.env.EPISODIC_MEMORY_REDACTION = 'off'; + expect(getRedactionSettings(process.env).enabled).toBe(false); + expect(loadRedactor(process.env)).toBeNull(); + }); + + it('loads the bundled defaults when no rules file exists', () => { + const redactor = loadRedactor(process.env)!; + expect(redactor.ruleIds).toEqual(DEFAULT_REDACTION_CONFIG.rules.map(r => r.id)); + }); + + it('picks up redaction-rules.json from the config dir and merges it with defaults', () => { + writeFileSync(join(process.env.EPISODIC_MEMORY_CONFIG_DIR!, 'redaction-rules.json'), JSON.stringify({ + rules: [{ id: 'internal-ticket-secret', pattern: 'TKT-SECRET-[0-9]{6}' }], + disableRules: ['basic-auth'], + })); + const redactor = loadRedactor(process.env)!; + expect(redactor.ruleIds).toContain('internal-ticket-secret'); + expect(redactor.ruleIds).toContain('azure-storage-key'); + expect(redactor.ruleIds).not.toContain('basic-auth'); + expect(redactor.redact('ref TKT-SECRET-123456', ctx).text).toBe('ref [REDACTED:internal-ticket-secret]'); + }); + + it('EPISODIC_MEMORY_REDACTION_RULES points at a custom file; includeDefaults:false replaces the defaults', () => { + const file = join(dir, 'custom.json'); + writeFileSync(file, JSON.stringify({ includeDefaults: false, rules: [{ id: 'only-this', pattern: 'zzz[0-9]+' }] })); + process.env.EPISODIC_MEMORY_REDACTION_RULES = file; + const redactor = loadRedactor(process.env)!; + expect(redactor.ruleIds).toEqual(['only-this']); + }); + + it('strict mode throws RedactionConfigError on a corrupt rules file', () => { + const file = join(dir, 'bad.json'); + writeFileSync(file, '{ "rules": [ this is not json'); + process.env.EPISODIC_MEMORY_REDACTION_RULES = file; + expect(() => loadRedactor(process.env)).toThrow(RedactionConfigError); + }); + + it('strict mode throws when the custom rules path does not exist', () => { + process.env.EPISODIC_MEMORY_REDACTION_RULES = join(dir, 'missing.json'); + expect(() => loadRedactor(process.env)).toThrow(RedactionConfigError); + }); + + it('non-strict mode warns and returns null (pass-through) on a corrupt rules file', () => { + const file = join(dir, 'bad.json'); + writeFileSync(file, 'nope'); + process.env.EPISODIC_MEMORY_REDACTION_RULES = file; + process.env.EPISODIC_MEMORY_REDACTION_STRICT = '0'; + const warn = vi.fn(); + expect(loadRedactor(process.env, warn)).toBeNull(); + expect(warn).toHaveBeenCalled(); + }); + + it('rejects invalid rules: bad regex, bad id, empty-matching pattern, bad secretGroup', () => { + const bad = [ + { id: 'bad-regex', pattern: '(unclosed' }, + { id: 'Bad Id!', pattern: 'abc' }, + { id: 'empty-match', pattern: 'a*' }, + { id: 'bad-group', pattern: 'abc', secretGroup: 2 }, + ]; + for (const rule of bad) { + expect(() => loadRedactionConfig({ includeDefaults: false, rules: [rule] } as any), rule.id).toThrow(RedactionConfigError); + } + }); + + it('rule-load errors never echo text being redacted (only config details)', () => { + const file = join(dir, 'bad.json'); + writeFileSync(file, JSON.stringify({ rules: [{ id: 'x', pattern: '(' }] })); + process.env.EPISODIC_MEMORY_REDACTION_RULES = file; + try { + loadRedactor(process.env); + expect.unreachable(); + } catch (e) { + expect((e as Error).message).toContain('x'); + } + }); +}); From 0c5789063d752ce36fc8c7d0e96bdf1ae5a993d7 Mon Sep 17 00:00:00 2001 From: Claude Date: Sun, 4 Oct 2026 23:02:53 +0000 Subject: [PATCH 03/15] fix(redaction): use key context so .NET/Azure secrets without a shape are caught The first cut only redacted values whose own shape matched a rule, plus an unquoted key=value rule that required a digit. Common .NET/Azure leaks all slipped through: appsettings.json fields, web.config , publish-profile userPWD, `az webapp config appsettings list` and `az keyvault secret show` output, C# `ClientSecret = "..."`, and structured MCP results (parsed JSON lost field names; keys were never checked). Text rules (key context, no digit requirement, placeholder-aware): - azure-keyvault-secret, name-value-secret (either order), xml-appsettings-secret (either order), xml-secret-element, xml-secret-attribute, quoted-secret-assignment. - These opt out of the SHA/GUID allowlist: a GUID in a password slot (legacy create-for-rbac) is a secret. Structured JSON (secretFields, on by default, configurable keyPattern): - a string whose field name ends in a secret keyword is redacted whole; - {name, value} pairs with a secret-looking name and Key Vault bundles have their value redacted; - keys that are themselves secrets are renamed to their token. Harness fields (thinking signature, usage, token_count, apiKeySource) are untouched. Engine: secretGroup may list several groups (first that matched); per-rule useAllowlist. connection-string-secret no longer allows spaces around '=', fixing a C# false positive (options.Password = configuration[...]). Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_0141gGbp343kbYP6imaJaji2 --- CHANGELOG.md | 2 +- dist/redaction-rules.d.ts | 25 ----- dist/redaction-rules.js | 91 ++++++++++++++++- dist/redaction.d.ts | 28 +++++- dist/redaction.js | 132 ++++++++++++++++++++++--- docs/REDACTION.md | 67 ++++++++++--- src/redaction-rules.ts | 105 +++++++++++++++++++- src/redaction.ts | 167 ++++++++++++++++++++++++++++---- test/fake-secrets.ts | 7 +- test/redaction-pipeline.test.ts | 37 +++++++ test/redaction.test.ts | 165 ++++++++++++++++++++++++++++++- 11 files changed, 747 insertions(+), 79 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 5d74d39d..a467864f 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -9,7 +9,7 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0 ### Security -- Secrets no longer leak into your conversation memory. Azure client secrets, storage keys, SAS signatures, connection-string passwords, private keys, JWTs, GitHub, Anthropic, OpenAI, and AWS keys, and `password: …` style assignments are now replaced with tokens like `[REDACTED:azure-storage-key]` before a conversation is archived, indexed, embedded, or summarized. Before this, every pasted secret was copied into three new places on disk and could be sent to the summarization model. The rest of the conversation stays searchable, and so do git SHAs and Azure tenant, client, and object IDs. Redaction is on by default and fails closed: if the rules can't load, sync won't run rather than store secrets. +- Secrets no longer leak into your conversation memory. Azure client secrets, storage keys, SAS signatures, connection-string passwords, private keys, JWTs, GitHub, Anthropic, OpenAI, and AWS keys, and `password: …` style assignments are now replaced with tokens like `[REDACTED:azure-storage-key]` before a conversation is archived, indexed, embedded, or summarized. Before this, every pasted secret was copied into three new places on disk and could be sent to the summarization model. The rest of the conversation stays searchable, and so do git SHAs and Azure tenant, client, and object IDs. Redaction is on by default and fails closed: if the rules can't load, sync won't run rather than store secrets. .NET and Azure shapes where only the setting's name gives the secret away are covered too, even when the value has no recognizable format: `appsettings.json` fields, `web.config` ``, publish-profile passwords, `az webapp config appsettings list` and `az keyvault secret show` output, C# `ClientSecret = "…"` assignments, and structured MCP results. - With redaction on, summaries no longer resume Claude Code sessions or fork Codex threads. Both of those let the model read the original, unredacted transcript. Summaries now come from the redacted text. Codex-only users without Claude configured can set `EPISODIC_MEMORY_SKIP_SUMMARIES=1`. ### Added diff --git a/dist/redaction-rules.d.ts b/dist/redaction-rules.d.ts index a9fa089b..d9c5ad24 100644 --- a/dist/redaction-rules.d.ts +++ b/dist/redaction-rules.d.ts @@ -1,27 +1,2 @@ import type { RedactionConfig } from './redaction.js'; -/** - * Bundled default redaction rules. - * - * Patterns are ported from gitleaks' rule set (https://github.com/gitleaks/gitleaks, - * MIT), trimmed to the credentials that realistically show up in Claude Code / - * Codex transcripts for an Azure-heavy stack, and rewritten for JavaScript - * regex (lookbehind instead of gitleaks' consuming boundary groups, so - * adjacent matches aren't swallowed). - * - * Order matters: rules run top to bottom, and a later rule never re-matches - * inside an earlier rule's `[REDACTED:...]` token. Specific shapes go first so - * the token names the most precise rule; the generic `secret-assignment` - * keyword rule goes last. - * - * `secretGroup` replaces only that capture group, so surrounding context - * (connection-string server names, SAS URL paths, usernames) stays searchable. - * - * `keywords` is a case-insensitive prefilter: a rule only runs on text that - * contains at least one keyword. It keeps the per-string cost low on large - * transcripts and bounds the generic rules' work. - * - * Users extend or override these with `redaction-rules.json`; see - * docs/REDACTION.md. `episodic-memory redact --print-default-rules` dumps - * this object as JSON. - */ export declare const DEFAULT_REDACTION_CONFIG: RedactionConfig; diff --git a/dist/redaction-rules.js b/dist/redaction-rules.js index 5adc1c46..153da735 100644 --- a/dist/redaction-rules.js +++ b/dist/redaction-rules.js @@ -23,6 +23,23 @@ * docs/REDACTION.md. `episodic-memory redact --print-default-rules` dumps * this object as JSON. */ +// Names that mark the value next to them as a secret, for the key-context +// rules below. A name must END in one of these (optionally plus digits), so +// tokenType, secretName, passwordPolicy, TokenEndpoint and maxTokens don't +// count. `pwd` is deliberately absent (the PWD env var); `userPWD` is handled +// by xml-secret-attribute. +const SECRET_NAME = String.raw `(?:secret|password|passwd|passphrase|api[_-]?key|access[_-]?key|account[_-]?key|private[_-]?key|` + + String.raw `shared[_-]?(?:access[_-]?)?key|primary[_-]?key|secondary[_-]?key|master[_-]?key|signing[_-]?key|` + + String.raw `subscription[_-]?key|client[_-]?key|encryption[_-]?key|token|credentials?)\d*`; +// Value is not a template/placeholder: ${X} $(X) $X #{X}# {{x}} %X% __X__, +// an existing token, a type name (Swagger "string"), or a mask (*****). +const NOT_PLACEHOLDER = String.raw `(?!\[REDACTED:|\$\{|\$\(|\$[A-Za-z_]\w*["'<\s]|#\{|\{\{|<[^>]*>|%[A-Za-z_]\w*%|__[A-Za-z0-9_]+__|` + + String.raw `(?:string|null|true|false|none|undefined|\*+)["'<])`; +// A JSON string value (escapes allowed), captured. +const JSON_STRING_VALUE = String.raw `"${NOT_PLACEHOLDER}([^"\\]+(?:\\.[^"\\]*)*)"`; +// Sibling members inside the same object: strings, or anything but braces/quotes. +const SAME_OBJECT = String.raw `(?:[^{}"]|"(?:[^"\\]|\\.)*"){0,400}?`; +const KEY_VAULT_ID = String.raw `"id"[ \t]*:[ \t]*"https://[^"\s]+\.vault\.(?:azure\.net|azure\.cn|usgovcloudapi\.net|microsoftazure\.de)/secrets/[^"\s]*"`; export const DEFAULT_REDACTION_CONFIG = { rules: [ { @@ -34,7 +51,7 @@ export const DEFAULT_REDACTION_CONFIG = { { id: 'connection-string-secret', description: 'Secret value inside an Azure/ADO.NET connection string; account, server and database names are kept', - pattern: String.raw `\b(?:AccountKey|SharedAccessKey|SharedAccessSignature|SharedSecret|ClientSecret|Password|Pwd)[ \t]*=[ \t]*(?![\[$<{%])([^;"'\s]+)`, + pattern: String.raw `\b(?:AccountKey|SharedAccessKey|SharedAccessSignature|SharedSecret|ClientSecret|Password|Pwd)=(?![\[$<{%])([^;"'\s]+)`, secretGroup: 1, keywords: ['accountkey', 'sharedaccess', 'sharedsecret', 'clientsecret', 'password', 'pwd'], }, @@ -126,6 +143,70 @@ export const DEFAULT_REDACTION_CONFIG = { secretGroup: 1, keywords: ['basic'], }, + { + id: 'azure-keyvault-secret', + description: 'The "value" of a Key Vault secret bundle (az keyvault secret show/set, SDK JSON), whatever the secret is named', + pattern: KEY_VAULT_ID + String.raw `(?:[^{}]|\{[^{}]*\}){0,800}?"value"[ \t]*:[ \t]*` + JSON_STRING_VALUE + + String.raw `|"value"[ \t]*:[ \t]*` + JSON_STRING_VALUE + String.raw `(?=(?:[^{}]|\{[^{}]*\}){0,800}?` + KEY_VAULT_ID + ')', + secretGroup: [1, 2], + useAllowlist: false, + keywords: ['.vault.'], + }, + { + id: 'name-value-secret', + description: 'The "value" of a {"name"/"key": , "value": ...} object, in either order: ' + + 'az webapp/functionapp config appsettings list, Kubernetes env, ARM/Bicep parameters', + pattern: String.raw `"(?:name|key)"[ \t]*:[ \t]*"[^"\\]{0,80}?` + SECRET_NAME + '"' + SAME_OBJECT + + String.raw `"value"[ \t]*:[ \t]*` + JSON_STRING_VALUE + + String.raw `|"value"[ \t]*:[ \t]*` + JSON_STRING_VALUE + '(?=' + SAME_OBJECT + + String.raw `"(?:name|key)"[ \t]*:[ \t]*"[^"\\]{0,80}?` + SECRET_NAME + '")', + flags: 'i', + secretGroup: [1, 2], + useAllowlist: false, + keywords: ['"value"'], + }, + { + id: 'xml-appsettings-secret', + description: 'web.config / app.config , either attribute order', + pattern: String.raw `]*?\bkey[ \t]*=[ \t]*"[^"]{0,80}?` + SECRET_NAME + String.raw `"[^>]*?\bvalue[ \t]*=[ \t]*"` + + NOT_PLACEHOLDER + String.raw `([^"]+)"` + + String.raw `|]*?\bvalue[ \t]*=[ \t]*"` + NOT_PLACEHOLDER + String.raw `([^"]+)"(?=[^>]*?\bkey[ \t]*=[ \t]*"[^"]{0,80}?` + + SECRET_NAME + '")', + flags: 'i', + secretGroup: [1, 2], + useAllowlist: false, + keywords: [']*)?>` + NOT_PLACEHOLDER + String.raw `([^<]{1,4096})`, + flags: 'i', + secretGroup: 2, + useAllowlist: false, + keywords: ['; + secretFields?: Partial; } export interface RedactionContext { source: string; @@ -81,6 +103,8 @@ export interface RedactionResult { export interface Redactor { redact(text: string, ctx?: RedactionContext): RedactionResult; readonly ruleIds: string[]; + /** True when a JSON key / setting name marks its value as a secret (secretFields). */ + isSecretField(name: string): boolean; } export interface RedactionSettings { enabled: boolean; diff --git a/dist/redaction.js b/dist/redaction.js index c4190658..83f76944 100644 --- a/dist/redaction.js +++ b/dist/redaction.js @@ -29,6 +29,11 @@ const TOKEN_PATTERN = /\[REDACTED:[a-z0-9][a-z0-9-]*\]/g; const TOKEN_PREFIX = '[REDACTED:'; const ALLOWED_FLAGS = /^[imsu]*$/; const ENTROPY_RULE_ID = 'high-entropy'; +const FIELD_RULE_ID = 'secret-field'; +const RESERVED_IDS = new Set([ENTROPY_RULE_ID, FIELD_RULE_ID]); +/** Values that are templates or type names, not secrets. Never redacted by field context. */ +const PLACEHOLDER_VALUE = /^(?:\$\{.*\}|\$\(.*\)|\$[A-Za-z_]\w*|#\{.*\}#?|\{\{.*\}\}|<[^>]*>|%[A-Za-z_]\w*%|__[A-Za-z0-9_]+__|string|null|true|false|none|undefined|\*+)$/i; +const KEY_VAULT_SECRET_ID = /^https:\/\/[^/\s]+\.vault\.(?:azure\.net|azure\.cn|usgovcloudapi\.net|microsoftazure\.de)\/secrets\//i; /** Rules could not be loaded or are invalid. Messages describe config only. */ export class RedactionConfigError extends Error { constructor(message) { @@ -110,8 +115,10 @@ export function loadRedactionConfig(file) { throw new RedactionConfigError(`Redaction rules file: "${key}" must be an array`); } } - if (file.entropy !== undefined && (typeof file.entropy !== 'object' || file.entropy === null)) { - throw new RedactionConfigError('Redaction rules file: "entropy" must be an object'); + for (const key of ['entropy', 'secretFields']) { + if (file[key] !== undefined && (typeof file[key] !== 'object' || file[key] === null || Array.isArray(file[key]))) { + throw new RedactionConfigError(`Redaction rules file: "${key}" must be an object`); + } } const includeDefaults = file.includeDefaults !== false; const disabled = new Set(file.disableRules ?? []); @@ -121,6 +128,7 @@ export function loadRedactionConfig(file) { rules: replaceById(baseRules, file.rules ?? []).filter(rule => !disabled.has(rule.id)), allowlist: replaceById(baseAllowlist, file.allowlist ?? []), entropy: { ...DEFAULT_REDACTION_CONFIG.entropy, ...(file.entropy ?? {}) }, + secretFields: { ...DEFAULT_REDACTION_CONFIG.secretFields, ...(file.secretFields ?? {}) }, }; compileConfig(config); // validate return config; @@ -164,8 +172,8 @@ function compileRule(spec) { if (typeof spec.id !== 'string' || !RULE_ID_PATTERN.test(spec.id)) { throw new RedactionConfigError(`Redaction rule id ${JSON.stringify(spec.id)} is invalid: use lowercase letters, digits and dashes`); } - if (spec.id === ENTROPY_RULE_ID) { - throw new RedactionConfigError(`Redaction rule id "${ENTROPY_RULE_ID}" is reserved`); + if (RESERVED_IDS.has(spec.id)) { + throw new RedactionConfigError(`Redaction rule id "${spec.id}" is reserved`); } const what = `Redaction rule "${spec.id}"`; if (typeof spec.pattern !== 'string' || spec.pattern.length === 0) { @@ -184,9 +192,17 @@ function compileRule(spec) { if (new RegExp(spec.pattern, flags).test('')) { throw new RedactionConfigError(`${what}: pattern matches the empty string`); } - const secretGroup = spec.secretGroup ?? 0; - if (!Number.isInteger(secretGroup) || secretGroup < 0 || secretGroup > groupCount) { - throw new RedactionConfigError(`${what}: secretGroup ${secretGroup} does not exist in the pattern (${groupCount} group(s))`); + const secretGroups = Array.isArray(spec.secretGroup) ? spec.secretGroup : [spec.secretGroup ?? 0]; + if (secretGroups.length === 0) { + throw new RedactionConfigError(`${what}: secretGroup must not be an empty array`); + } + for (const group of secretGroups) { + if (!Number.isInteger(group) || group < 0 || group > groupCount) { + throw new RedactionConfigError(`${what}: secretGroup ${group} does not exist in the pattern (${groupCount} group(s))`); + } + } + if (spec.useAllowlist !== undefined && typeof spec.useAllowlist !== 'boolean') { + throw new RedactionConfigError(`${what}: useAllowlist must be a boolean`); } let keywords; if (spec.keywords !== undefined) { @@ -195,7 +211,7 @@ function compileRule(spec) { } keywords = spec.keywords.length > 0 ? spec.keywords.map(k => k.toLowerCase()) : undefined; } - return { id: spec.id, regex, keywords, secretGroup }; + return { id: spec.id, regex, keywords, secretGroups, useAllowlist: spec.useAllowlist !== false }; } function compileAllowlist(spec) { const what = `Redaction allowlist entry ${JSON.stringify(spec?.id)}`; @@ -228,7 +244,22 @@ function compileConfig(config) { !Array.isArray(e.keywords) || e.keywords.some(k => typeof k !== 'string' || !k)) { throw new RedactionConfigError('Redaction entropy settings are invalid (need enabled, requireKeyword: boolean; minLength >= 8; threshold > 0; window >= 0; keywords: string[])'); } - return { rules, allowlist, entropy: { ...e, keywords: e.keywords.map(k => k.toLowerCase()) } }; + const f = config.secretFields; + if (!f || typeof f.enabled !== 'boolean' || typeof f.keyPattern !== 'string' || f.keyPattern.length === 0) { + throw new RedactionConfigError('Redaction secretFields settings are invalid (need enabled: boolean; keyPattern: non-empty string)'); + } + let secretField = null; + if (f.enabled) { + try { + secretField = new RegExp(`(?:${f.keyPattern})$`); + } + catch (error) { + throw new RedactionConfigError(`Redaction secretFields.keyPattern: invalid regular expression (${error instanceof Error ? error.message : String(error)})`); + } + if (secretField.test('')) + throw new RedactionConfigError('Redaction secretFields.keyPattern matches the empty string'); + } + return { rules, allowlist, entropy: { ...e, keywords: e.keywords.map(k => k.toLowerCase()) }, secretField }; } // --------------------------------------------------------------------------- // Redaction engine @@ -271,7 +302,12 @@ function applyRule(text, rule, isAllowed) { let last = 0; let count = 0; for (const m of text.matchAll(rule.regex)) { - const range = m.indices?.[rule.secretGroup]; + let range; + for (const group of rule.secretGroups) { + range = m.indices?.[group]; + if (range) + break; + } if (!range) continue; const [start, end] = range; @@ -279,7 +315,7 @@ function applyRule(text, rule, isAllowed) { continue; if (spans && overlapsAny(spans, start, end)) continue; - if (isAllowed(text.slice(start, end))) + if (rule.useAllowlist && isAllowed(text.slice(start, end))) continue; out += text.slice(last, start) + token; last = end; @@ -320,6 +356,12 @@ export function createRedactor(config) { const isAllowed = (secret) => compiled.allowlist.some(re => re.test(secret)); return { ruleIds: compiled.rules.map(r => r.id), + isSecretField(name) { + if (!compiled.secretField || typeof name !== 'string') + return false; + const normalized = name.toLowerCase().replace(/[^a-z0-9]/g, '').replace(/\d+$/, ''); + return normalized.length > 0 && compiled.secretField.test(normalized); + }, redact(text, _ctx) { if (typeof text !== 'string' || text.length === 0) return { text, findings: [] }; @@ -384,7 +426,10 @@ export function formatFindings(findings) { // --------------------------------------------------------------------------- // JSON / JSONL / files // --------------------------------------------------------------------------- -/** Redact every string value in a parsed JSON tree, in place. Keys are left alone. */ +/** + * Redact a parsed JSON tree in place: string values by the text rules, values + * by their field name (secretFields), and keys that are themselves secrets. + */ function redactTree(node, redactor, ctx, tally) { if (typeof node === 'string') { const r = redactor.redact(node, ctx); @@ -407,15 +452,74 @@ function redactTree(node, redactor, ctx, tally) { return { value: node, changed }; } const obj = node; - for (const key of Object.keys(obj)) { + const keys = Object.keys(obj); + for (const key of keys) { const r = redactTree(obj[key], redactor, ctx, tally); if (r.changed) { - Object.defineProperty(obj, key, { value: r.value, writable: true, enumerable: true, configurable: true }); + setOwn(obj, key, r.value); changed = true; } } + // Field-name context: the key (or a {name, value} pair's name, or a Key + // Vault secret id) says the value is a secret even when its shape matches + // no rule — e.g. a letters-only clientSecret field in a structured + // MCP result. Runs after the text rules, so a value they already redacted + // in part (a connection string) keeps its searchable remainder. + const redactWhole = (key) => { + if (!isRedactableWhole(obj[key])) + return; + setOwn(obj, key, tokenFor(FIELD_RULE_ID)); + tally?.add([{ ruleId: FIELD_RULE_ID, count: 1 }]); + changed = true; + }; + for (const key of keys) { + if (redactor.isSecretField(key)) + redactWhole(key); + } + const lower = new Map(keys.map(k => [k.toLowerCase(), k])); + const valueKey = lower.get('value'); + if (valueKey !== undefined) { + const nameKey = lower.get('name') ?? lower.get('key'); + const nameValue = nameKey !== undefined ? obj[nameKey] : undefined; + const id = lower.has('id') ? obj[lower.get('id')] : undefined; + if ((typeof nameValue === 'string' && redactor.isSecretField(nameValue)) || + (typeof id === 'string' && KEY_VAULT_SECRET_ID.test(id))) { + redactWhole(valueKey); + } + } + // Keys themselves: a secret used as a key (a token-keyed cache) is renamed + // to its token. Harness keys never match a rule, so structure is unchanged. + let renamed = null; + for (let i = 0; i < keys.length; i++) { + const r = redactor.redact(keys[i], ctx); + if (r.findings.length === 0) + continue; + tally?.add(r.findings); + renamed ??= keys.map(k => [k, obj[k]]); + renamed[i][0] = r.text; + } + if (renamed) { + const rebuilt = {}; + for (const [key, value] of renamed) { + let unique = key; + for (let n = 2; Object.prototype.hasOwnProperty.call(rebuilt, unique); n++) + unique = `${key}#${n}`; + setOwn(rebuilt, unique, value); + } + return { value: rebuilt, changed: true }; + } return { value: obj, changed }; } +function setOwn(obj, key, value) { + // defineProperty so a "__proto__" key from JSON.parse stays an own property. + Object.defineProperty(obj, key, { value, writable: true, enumerable: true, configurable: true }); +} +function isRedactableWhole(value) { + return typeof value === 'string' && + value.trim().length > 0 && + !value.includes(TOKEN_PREFIX) && + !PLACEHOLDER_VALUE.test(value.trim()); +} /** * Redact one JSONL line. Valid JSON stays valid (string values are redacted * and the line is re-serialized only if something changed); a line that isn't diff --git a/docs/REDACTION.md b/docs/REDACTION.md index 85d967c9..f3bafbe1 100644 --- a/docs/REDACTION.md +++ b/docs/REDACTION.md @@ -4,9 +4,10 @@ episodic-memory replaces secrets in your conversations with typed tokens **before** it archives, indexes, embeds, or summarizes them: ``` -AccountKey=Zm9v...== → AccountKey=[REDACTED:connection-string-secret] -"password": "abc1Q~..." → "password": "[REDACTED:azure-client-secret]" -Authorization: Bearer eyJ → Authorization: Bearer [REDACTED:jwt] +AccountKey=<88-char key> → AccountKey=[REDACTED:connection-string-secret] +"ClientSecret": "" → "ClientSecret": "[REDACTED:quoted-secret-assignment]" + → +Authorization: Bearer → Authorization: Bearer [REDACTED:jwt] ``` Values are redacted, not dropped. The rest of the conversation stays @@ -56,11 +57,32 @@ summary: | `azure-storage-key` | Standalone 88-character base64 keys (Storage, Cosmos DB, Function keys) | | `url-credentials` | The password in `scheme://user:password@host` | | `bearer-token`, `basic-auth` | `Authorization` header values | -| `secret-assignment` | A value assigned to a secret-looking key: `password: …`, `CLIENT_SECRET=…`, `"apiKey": "…"`. Covers decrypted SOPS, YAML, dotenv, and JSON. Needs 8+ characters including a digit, and skips placeholders (`${X}`, ``, `%X%`) and code (`env.X`, `getPassword()`). | - -**Allowlisted** (never redacted, even when a rule matches): git SHAs and -GUIDs. Tenant, client, object, and subscription IDs stay searchable, and so do -commit hashes. +| `azure-keyvault-secret` | The `value` of a Key Vault secret bundle (`az keyvault secret show`, SDK JSON), whatever the secret is named | +| `name-value-secret` | The `value` of a `{"name": , "value": …}` object, in either order: `az webapp`/`functionapp config appsettings list`, Kubernetes `env`, ARM/Bicep parameters | +| `xml-appsettings-secret` | `web.config` / `app.config` ``, either attribute order | +| `xml-secret-element` | `…`, `…` and the like | +| `xml-secret-attribute` | Secret-named XML attributes, e.g. `userPWD="…"` in Azure publish profiles | +| `quoted-secret-assignment` | A **quoted** value assigned to a secret-looking name: `"ClientSecret": "…"` (appsettings.json and any JSON in tool output), `ClientSecret = "…"` (C#), `apiKey: '…'` (JS/YAML/Python). No digit or minimum entropy is required; the value must have no spaces. | +| `secret-assignment` | An **unquoted** value assigned to a secret-looking key: `password: …`, `CLIENT_SECRET=…`. Covers decrypted SOPS, YAML, and dotenv. Needs 8+ characters including a digit, and skips code (`env.X`, `getPassword()`). | +| `secret-field` | Not a text rule: in parsed JSON (transcript lines, structured MCP/tool results, tool inputs), a string whose **field name** is secret-looking, or the `value` of a `{name, value}` pair or Key Vault bundle, is redacted whole. See `secretFields` below. | + +A **secret-looking name** ends in `secret`, `password`, `passwd`, +`passphrase`, `apikey`, `accesskey`, `accountkey`, `privatekey`, `sharedkey`, +`primarykey`, `secondarykey`, `masterkey`, `signingkey`, `subscriptionkey`, +`clientkey`, `encryptionkey`, `token`, or `credential(s)`, in any case and +with any separators (`AzureAd:ClientSecret`, `Stripe__ApiKey`, `DB_PASSWORD2`). +Because it has to *end* in one of these, `TokenEndpoint`, `secretName`, +`passwordPolicy`, `tokenType`, and `maxTokens` don't count. + +Key-context rules leave **placeholders** alone: `${X}`, `$(X)`, `#{X}#` +(Azure DevOps token replacement), `{{x}}`, ``, `%X%`, `__X__`, and type +names or masks like `"string"` and `"*****"`. All other rules skip the same +`${X}`, ``, and `%X%` forms. + +**Allowlisted:** git SHAs and GUIDs are not redacted by shape-based rules, +so tenant, client, object, and subscription IDs and commit hashes stay +searchable. Key-context rules ignore the allowlist on purpose: a GUID in a +`password` slot (older `az ad sp create-for-rbac` output) *is* a secret. **Entropy fallback:** off by default. When it's on, a long, high-entropy string is redacted only if a keyword (`secret`, `key`, `token`, …) appears just before @@ -95,7 +117,8 @@ bundled rules: "pattern": "\\bctso_[A-Za-z0-9]{32}\\b", // JavaScript regex "flags": "i", // optional, any of "imsu" "keywords": ["ctso_"], // optional prefilter (case-insensitive) - "secretGroup": 0 // optional: redact only this capture group + "secretGroup": 0, // optional: redact only this group ([1, 2] = first that matched) + "useAllowlist": true // optional: false = redact even SHA/GUID-shaped values } ], "disableRules": ["basic-auth"], // turn off defaults by id @@ -103,6 +126,12 @@ bundled rules: { "id": "build-ids", "pattern": "build-[0-9]{8}" } ], "entropy": { "enabled": true }, // partial override of the entropy settings + "secretFields": { // field-name context in parsed JSON + "enabled": true, + // matched against the END of the key lowercased with separators removed; + // this example adds "connectionstring" to redact whole connection strings + "keyPattern": "secret|password|passwd|userpwd|passphrase|apikey|accesskey|accountkey|privatekey|sharedkey|primarykey|secondarykey|masterkey|signingkey|subscriptionkey|clientkey|encryptionkey|token|credentials?|connectionstring" + }, "includeDefaults": true // false = use only this file's rules } ``` @@ -148,8 +177,10 @@ Codex, Cursor, opencode, OMP) passes through that copy, and every later stage reads the archive rather than the source. The design and the research behind it are in [redaction/PHASE0-FINDINGS.md](redaction/PHASE0-FINDINGS.md). -Each archive line is parsed as JSON, every string value is redacted, and only -lines that changed are re-serialized. Lines with no secrets stay byte-for-byte +Each archive line is parsed as JSON. Every string value goes through the text +rules, values are redacted whole when their field name marks them as secrets +(`secretFields`), and a key that is itself a secret is renamed to its token. +Only lines that changed are re-serialized. Lines with no secrets stay byte-for-byte identical. The archive keeps exactly one line per source line, so index line ranges and MCP `read` ranges still line up. A line that isn't valid JSON (for example, a half-written last line) is redacted as plain text. @@ -159,9 +190,15 @@ example, a half-written last line) is redacted as plain text. - Pattern rules miss secrets that have no recognizable shape and no `key: value` context. The entropy fallback helps when it's on, but recall isn't perfect. -- Each JSON string value is checked on its own. Context split across fields - (`{"name": "DB_PASSWORD", "value": "…"}`) isn't linked, so the `value` is - caught only if its own shape matches a rule. -- JSON object *keys* aren't redacted. Only values are. +- A secret-looking name is required for the key-context rules. A setting + named `Stripe` or `ConnectionStrings__Default` with a shapeless value is + only caught if a shape rule matches the value. Connection strings are + handled by `connection-string-secret`, which keeps server names searchable. + Add your own names via `secretFields.keyPattern` or a custom rule. +- Quoted values containing spaces (multi-word passphrases) aren't caught by + `quoted-secret-assignment`. The rule excludes them so that UI labels like + `ErrorMessage = "Invalid password"` aren't redacted. +- Positional secrets in code (`new ClientSecretCredential(t, c, "…")`) are + caught only by shape, e.g. `azure-client-secret`. - On lines that get redacted, re-serializing can change number formatting for integers above 2^53. No supported harness writes such numbers. diff --git a/src/redaction-rules.ts b/src/redaction-rules.ts index d73a043f..02c41467 100644 --- a/src/redaction-rules.ts +++ b/src/redaction-rules.ts @@ -25,6 +25,29 @@ import type { RedactionConfig } from './redaction.js'; * docs/REDACTION.md. `episodic-memory redact --print-default-rules` dumps * this object as JSON. */ +// Names that mark the value next to them as a secret, for the key-context +// rules below. A name must END in one of these (optionally plus digits), so +// tokenType, secretName, passwordPolicy, TokenEndpoint and maxTokens don't +// count. `pwd` is deliberately absent (the PWD env var); `userPWD` is handled +// by xml-secret-attribute. +const SECRET_NAME = + String.raw`(?:secret|password|passwd|passphrase|api[_-]?key|access[_-]?key|account[_-]?key|private[_-]?key|` + + String.raw`shared[_-]?(?:access[_-]?)?key|primary[_-]?key|secondary[_-]?key|master[_-]?key|signing[_-]?key|` + + String.raw`subscription[_-]?key|client[_-]?key|encryption[_-]?key|token|credentials?)\d*`; + +// Value is not a template/placeholder: ${X} $(X) $X #{X}# {{x}} %X% __X__, +// an existing token, a type name (Swagger "string"), or a mask (*****). +const NOT_PLACEHOLDER = + String.raw`(?!\[REDACTED:|\$\{|\$\(|\$[A-Za-z_]\w*["'<\s]|#\{|\{\{|<[^>]*>|%[A-Za-z_]\w*%|__[A-Za-z0-9_]+__|` + + String.raw`(?:string|null|true|false|none|undefined|\*+)["'<])`; + +// A JSON string value (escapes allowed), captured. +const JSON_STRING_VALUE = String.raw`"${NOT_PLACEHOLDER}([^"\\]+(?:\\.[^"\\]*)*)"`; +// Sibling members inside the same object: strings, or anything but braces/quotes. +const SAME_OBJECT = String.raw`(?:[^{}"]|"(?:[^"\\]|\\.)*"){0,400}?`; +const KEY_VAULT_ID = + String.raw`"id"[ \t]*:[ \t]*"https://[^"\s]+\.vault\.(?:azure\.net|azure\.cn|usgovcloudapi\.net|microsoftazure\.de)/secrets/[^"\s]*"`; + export const DEFAULT_REDACTION_CONFIG: RedactionConfig = { rules: [ { @@ -36,7 +59,7 @@ export const DEFAULT_REDACTION_CONFIG: RedactionConfig = { { id: 'connection-string-secret', description: 'Secret value inside an Azure/ADO.NET connection string; account, server and database names are kept', - pattern: String.raw`\b(?:AccountKey|SharedAccessKey|SharedAccessSignature|SharedSecret|ClientSecret|Password|Pwd)[ \t]*=[ \t]*(?![\[$<{%])([^;"'\s]+)`, + pattern: String.raw`\b(?:AccountKey|SharedAccessKey|SharedAccessSignature|SharedSecret|ClientSecret|Password|Pwd)=(?![\[$<{%])([^;"'\s]+)`, secretGroup: 1, keywords: ['accountkey', 'sharedaccess', 'sharedsecret', 'clientsecret', 'password', 'pwd'], }, @@ -128,6 +151,77 @@ export const DEFAULT_REDACTION_CONFIG: RedactionConfig = { secretGroup: 1, keywords: ['basic'], }, + { + id: 'azure-keyvault-secret', + description: 'The "value" of a Key Vault secret bundle (az keyvault secret show/set, SDK JSON), whatever the secret is named', + pattern: + KEY_VAULT_ID + String.raw`(?:[^{}]|\{[^{}]*\}){0,800}?"value"[ \t]*:[ \t]*` + JSON_STRING_VALUE + + String.raw`|"value"[ \t]*:[ \t]*` + JSON_STRING_VALUE + String.raw`(?=(?:[^{}]|\{[^{}]*\}){0,800}?` + KEY_VAULT_ID + ')', + secretGroup: [1, 2], + useAllowlist: false, + keywords: ['.vault.'], + }, + { + id: 'name-value-secret', + description: + 'The "value" of a {"name"/"key": , "value": ...} object, in either order: ' + + 'az webapp/functionapp config appsettings list, Kubernetes env, ARM/Bicep parameters', + pattern: + String.raw`"(?:name|key)"[ \t]*:[ \t]*"[^"\\]{0,80}?` + SECRET_NAME + '"' + SAME_OBJECT + + String.raw`"value"[ \t]*:[ \t]*` + JSON_STRING_VALUE + + String.raw`|"value"[ \t]*:[ \t]*` + JSON_STRING_VALUE + '(?=' + SAME_OBJECT + + String.raw`"(?:name|key)"[ \t]*:[ \t]*"[^"\\]{0,80}?` + SECRET_NAME + '")', + flags: 'i', + secretGroup: [1, 2], + useAllowlist: false, + keywords: ['"value"'], + }, + { + id: 'xml-appsettings-secret', + description: 'web.config / app.config , either attribute order', + pattern: + String.raw`]*?\bkey[ \t]*=[ \t]*"[^"]{0,80}?` + SECRET_NAME + String.raw`"[^>]*?\bvalue[ \t]*=[ \t]*"` + + NOT_PLACEHOLDER + String.raw`([^"]+)"` + + String.raw`|]*?\bvalue[ \t]*=[ \t]*"` + NOT_PLACEHOLDER + String.raw`([^"]+)"(?=[^>]*?\bkey[ \t]*=[ \t]*"[^"]{0,80}?` + + SECRET_NAME + '")', + flags: 'i', + secretGroup: [1, 2], + useAllowlist: false, + keywords: [']*)?>` + NOT_PLACEHOLDER + String.raw`([^<]{1,4096})`, + flags: 'i', + secretGroup: 2, + useAllowlist: false, + keywords: [']*>|%[A-Za-z_]\w*%|__[A-Za-z0-9_]+__|string|null|true|false|none|undefined|\*+)$/i; +const KEY_VAULT_SECRET_ID = /^https:\/\/[^/\s]+\.vault\.(?:azure\.net|azure\.cn|usgovcloudapi\.net|microsoftazure\.de)\/secrets\//i; export interface RedactionRuleSpec { id: string; @@ -42,8 +49,13 @@ export interface RedactionRuleSpec { flags?: string; /** Case-insensitive prefilter: skip the rule unless the text contains one. */ keywords?: string[]; - /** Replace only this capture group instead of the whole match. */ - secretGroup?: number; + /** + * Replace only this capture group instead of the whole match. An array means + * "the first of these groups that participated" (for either-order patterns). + */ + secretGroup?: number | number[]; + /** Apply the allowlist to this rule's matches (default true). Key-context rules turn it off: a GUID in a password slot is a secret. */ + useAllowlist?: boolean; description?: string; } @@ -67,10 +79,27 @@ export interface EntropySpec { window: number; } +/** + * Field-name context for parsed JSON (transcript lines, structured tool/MCP + * results, tool inputs). A string value is redacted whole when its own key, or + * the `name`/`key` of a `{name, value}` pair, matches `keyPattern`. + */ +export interface SecretFieldsSpec { + enabled: boolean; + /** + * Matched (anchored at the end) against the key lowercased with everything + * but letters and digits removed and trailing digits dropped, so + * `AzureAd:ClientSecret`, `client_secret` and `DB_PASSWORD2` all normalize to + * something ending in a keyword. + */ + keyPattern: string; +} + export interface RedactionConfig { rules: RedactionRuleSpec[]; allowlist: AllowlistSpec[]; entropy: EntropySpec; + secretFields: SecretFieldsSpec; } /** Shape of a user `redaction-rules.json`. Every field is optional. */ @@ -84,6 +113,7 @@ export interface RedactionRulesFile { /** Added to the default allowlist (same id replaces). */ allowlist?: AllowlistSpec[]; entropy?: Partial; + secretFields?: Partial; } export interface RedactionContext { @@ -104,6 +134,8 @@ export interface RedactionResult { export interface Redactor { redact(text: string, ctx?: RedactionContext): RedactionResult; readonly ruleIds: string[]; + /** True when a JSON key / setting name marks its value as a secret (secretFields). */ + isSecretField(name: string): boolean; } export interface RedactionSettings { @@ -197,8 +229,10 @@ export function loadRedactionConfig(file: RedactionRulesFile): RedactionConfig { throw new RedactionConfigError(`Redaction rules file: "${key}" must be an array`); } } - if (file.entropy !== undefined && (typeof file.entropy !== 'object' || file.entropy === null)) { - throw new RedactionConfigError('Redaction rules file: "entropy" must be an object'); + for (const key of ['entropy', 'secretFields'] as const) { + if (file[key] !== undefined && (typeof file[key] !== 'object' || file[key] === null || Array.isArray(file[key]))) { + throw new RedactionConfigError(`Redaction rules file: "${key}" must be an object`); + } } const includeDefaults = file.includeDefaults !== false; @@ -210,6 +244,7 @@ export function loadRedactionConfig(file: RedactionRulesFile): RedactionConfig { rules: replaceById(baseRules, file.rules ?? []).filter(rule => !disabled.has(rule.id)), allowlist: replaceById(baseAllowlist, file.allowlist ?? []), entropy: { ...DEFAULT_REDACTION_CONFIG.entropy, ...(file.entropy ?? {}) }, + secretFields: { ...DEFAULT_REDACTION_CONFIG.secretFields, ...(file.secretFields ?? {}) }, }; compileConfig(config); // validate return config; @@ -251,13 +286,16 @@ interface CompiledRule { id: string; regex: RegExp; keywords?: string[]; - secretGroup: number; + secretGroups: number[]; + useAllowlist: boolean; } interface CompiledConfig { rules: CompiledRule[]; allowlist: RegExp[]; entropy: EntropySpec; + /** null when secretFields is disabled. */ + secretField: RegExp | null; } function checkFlags(flags: string | undefined, what: string): string { @@ -277,8 +315,8 @@ function compileRule(spec: RedactionRuleSpec): CompiledRule { `Redaction rule id ${JSON.stringify(spec.id)} is invalid: use lowercase letters, digits and dashes` ); } - if (spec.id === ENTROPY_RULE_ID) { - throw new RedactionConfigError(`Redaction rule id "${ENTROPY_RULE_ID}" is reserved`); + if (RESERVED_IDS.has(spec.id)) { + throw new RedactionConfigError(`Redaction rule id "${spec.id}" is reserved`); } const what = `Redaction rule "${spec.id}"`; if (typeof spec.pattern !== 'string' || spec.pattern.length === 0) { @@ -298,9 +336,17 @@ function compileRule(spec: RedactionRuleSpec): CompiledRule { throw new RedactionConfigError(`${what}: pattern matches the empty string`); } - const secretGroup = spec.secretGroup ?? 0; - if (!Number.isInteger(secretGroup) || secretGroup < 0 || secretGroup > groupCount) { - throw new RedactionConfigError(`${what}: secretGroup ${secretGroup} does not exist in the pattern (${groupCount} group(s))`); + const secretGroups = Array.isArray(spec.secretGroup) ? spec.secretGroup : [spec.secretGroup ?? 0]; + if (secretGroups.length === 0) { + throw new RedactionConfigError(`${what}: secretGroup must not be an empty array`); + } + for (const group of secretGroups) { + if (!Number.isInteger(group) || group < 0 || group > groupCount) { + throw new RedactionConfigError(`${what}: secretGroup ${group} does not exist in the pattern (${groupCount} group(s))`); + } + } + if (spec.useAllowlist !== undefined && typeof spec.useAllowlist !== 'boolean') { + throw new RedactionConfigError(`${what}: useAllowlist must be a boolean`); } let keywords: string[] | undefined; @@ -311,7 +357,7 @@ function compileRule(spec: RedactionRuleSpec): CompiledRule { keywords = spec.keywords.length > 0 ? spec.keywords.map(k => k.toLowerCase()) : undefined; } - return { id: spec.id, regex, keywords, secretGroup }; + return { id: spec.id, regex, keywords, secretGroups, useAllowlist: spec.useAllowlist !== false }; } function compileAllowlist(spec: AllowlistSpec): RegExp { @@ -349,7 +395,21 @@ function compileConfig(config: RedactionConfig): CompiledConfig { 'Redaction entropy settings are invalid (need enabled, requireKeyword: boolean; minLength >= 8; threshold > 0; window >= 0; keywords: string[])' ); } - return { rules, allowlist, entropy: { ...e, keywords: e.keywords.map(k => k.toLowerCase()) } }; + const f = config.secretFields; + if (!f || typeof f.enabled !== 'boolean' || typeof f.keyPattern !== 'string' || f.keyPattern.length === 0) { + throw new RedactionConfigError('Redaction secretFields settings are invalid (need enabled: boolean; keyPattern: non-empty string)'); + } + let secretField: RegExp | null = null; + if (f.enabled) { + try { + secretField = new RegExp(`(?:${f.keyPattern})$`); + } catch (error) { + throw new RedactionConfigError(`Redaction secretFields.keyPattern: invalid regular expression (${error instanceof Error ? error.message : String(error)})`); + } + if (secretField.test('')) throw new RedactionConfigError('Redaction secretFields.keyPattern matches the empty string'); + } + + return { rules, allowlist, entropy: { ...e, keywords: e.keywords.map(k => k.toLowerCase()) }, secretField }; } // --------------------------------------------------------------------------- @@ -394,12 +454,16 @@ function applyRule(text: string, rule: CompiledRule, isAllowed: (s: string) => b let last = 0; let count = 0; for (const m of text.matchAll(rule.regex)) { - const range = m.indices?.[rule.secretGroup]; + let range: [number, number] | undefined; + for (const group of rule.secretGroups) { + range = m.indices?.[group]; + if (range) break; + } if (!range) continue; const [start, end] = range; if (end <= start || start < last) continue; if (spans && overlapsAny(spans, start, end)) continue; - if (isAllowed(text.slice(start, end))) continue; + if (rule.useAllowlist && isAllowed(text.slice(start, end))) continue; out += text.slice(last, start) + token; last = end; count++; @@ -438,6 +502,11 @@ export function createRedactor(config: RedactionConfig): Redactor { return { ruleIds: compiled.rules.map(r => r.id), + isSecretField(name: string): boolean { + if (!compiled.secretField || typeof name !== 'string') return false; + const normalized = name.toLowerCase().replace(/[^a-z0-9]/g, '').replace(/\d+$/, ''); + return normalized.length > 0 && compiled.secretField.test(normalized); + }, redact(text: string, _ctx?: RedactionContext): RedactionResult { if (typeof text !== 'string' || text.length === 0) return { text, findings: [] }; let current = text; @@ -508,7 +577,10 @@ export function formatFindings(findings: RedactionFinding[]): string { // JSON / JSONL / files // --------------------------------------------------------------------------- -/** Redact every string value in a parsed JSON tree, in place. Keys are left alone. */ +/** + * Redact a parsed JSON tree in place: string values by the text rules, values + * by their field name (secretFields), and keys that are themselves secrets. + */ function redactTree( node: unknown, redactor: Redactor, @@ -532,16 +604,77 @@ function redactTree( return { value: node, changed }; } const obj = node as Record; - for (const key of Object.keys(obj)) { + const keys = Object.keys(obj); + for (const key of keys) { const r = redactTree(obj[key], redactor, ctx, tally); if (r.changed) { - Object.defineProperty(obj, key, { value: r.value, writable: true, enumerable: true, configurable: true }); + setOwn(obj, key, r.value); changed = true; } } + + // Field-name context: the key (or a {name, value} pair's name, or a Key + // Vault secret id) says the value is a secret even when its shape matches + // no rule — e.g. a letters-only clientSecret field in a structured + // MCP result. Runs after the text rules, so a value they already redacted + // in part (a connection string) keeps its searchable remainder. + const redactWhole = (key: string) => { + if (!isRedactableWhole(obj[key])) return; + setOwn(obj, key, tokenFor(FIELD_RULE_ID)); + tally?.add([{ ruleId: FIELD_RULE_ID, count: 1 }]); + changed = true; + }; + for (const key of keys) { + if (redactor.isSecretField(key)) redactWhole(key); + } + const lower = new Map(keys.map(k => [k.toLowerCase(), k])); + const valueKey = lower.get('value'); + if (valueKey !== undefined) { + const nameKey = lower.get('name') ?? lower.get('key'); + const nameValue = nameKey !== undefined ? obj[nameKey] : undefined; + const id = lower.has('id') ? obj[lower.get('id')!] : undefined; + if ( + (typeof nameValue === 'string' && redactor.isSecretField(nameValue)) || + (typeof id === 'string' && KEY_VAULT_SECRET_ID.test(id)) + ) { + redactWhole(valueKey); + } + } + + // Keys themselves: a secret used as a key (a token-keyed cache) is renamed + // to its token. Harness keys never match a rule, so structure is unchanged. + let renamed: Array<[string, unknown]> | null = null; + for (let i = 0; i < keys.length; i++) { + const r = redactor.redact(keys[i], ctx); + if (r.findings.length === 0) continue; + tally?.add(r.findings); + renamed ??= keys.map(k => [k, obj[k]]); + renamed[i][0] = r.text; + } + if (renamed) { + const rebuilt: Record = {}; + for (const [key, value] of renamed) { + let unique = key; + for (let n = 2; Object.prototype.hasOwnProperty.call(rebuilt, unique); n++) unique = `${key}#${n}`; + setOwn(rebuilt, unique, value); + } + return { value: rebuilt, changed: true }; + } return { value: obj, changed }; } +function setOwn(obj: Record, key: string, value: unknown): void { + // defineProperty so a "__proto__" key from JSON.parse stays an own property. + Object.defineProperty(obj, key, { value, writable: true, enumerable: true, configurable: true }); +} + +function isRedactableWhole(value: unknown): value is string { + return typeof value === 'string' && + value.trim().length > 0 && + !value.includes(TOKEN_PREFIX) && + !PLACEHOLDER_VALUE.test(value.trim()); +} + /** * Redact one JSONL line. Valid JSON stays valid (string values are redacted * and the line is re-serialized only if something changed); a line that isn't diff --git a/test/fake-secrets.ts b/test/fake-secrets.ts index 15deb68c..0eebca9e 100644 --- a/test/fake-secrets.ts +++ b/test/fake-secrets.ts @@ -113,7 +113,12 @@ export class FakeSecrets { return this.chars(LOWER, 1) + this.mixed(ALNUM + '!#%^*', 15); } - bearerOpaque(): string { + /** Letters only: a real password with no digit, which digit-gated rules miss. */ + passphrase(): string { + return this.chars(UPPER, 1) + this.chars(LOWER, 7) + this.chars(UPPER, 1) + this.chars(LOWER, 9); + } + + bearerOpaque(): string { return this.mixed(ALNUM, 40); } diff --git a/test/redaction-pipeline.test.ts b/test/redaction-pipeline.test.ts index edd4d1cb..89f60ea4 100644 --- a/test/redaction-pipeline.test.ts +++ b/test/redaction-pipeline.test.ts @@ -292,6 +292,43 @@ describe('redaction pipeline', () => { expect(summarizeSpy.mock.calls[0][2]).toMatchObject({ allowResume: true }); }); + it('.NET/Azure key-context secrets (no recognizable shape) never reach archive or SQLite', async () => { + const fake = new FakeSecrets(1999); + const appsettingsSecret = fake.passphrase(); + const appSettingValue = fake.passphrase(); + const vaultValue = fake.passphrase(); + const mcpSecret = fake.passphrase(); + const seed: Seed = { secrets: [appsettingsSecret, appSettingValue, vaultValue, mcpSecret], sha: fake.gitSha(), guid: fake.guid() }; + const base = { sessionId: SESSION, isSidechain: false, cwd: '/work/contoso' }; + const lines = [ + { ...base, type: 'user', uuid: 'u1', parentUuid: null, timestamp: '2026-01-01T00:00:00Z', message: { role: 'user', content: 'Why does the API fail to get a token?' } }, + { ...base, type: 'assistant', uuid: 'a1', parentUuid: 'u1', timestamp: '2026-01-01T00:00:01Z', message: { role: 'assistant', content: [ + { type: 'tool_use', id: 't1', name: 'Read', input: { file_path: '/work/contoso/appsettings.Development.json' } }, + ] } }, + { ...base, type: 'user', uuid: 'u2', parentUuid: 'a1', timestamp: '2026-01-01T00:00:02Z', + message: { role: 'user', content: [{ type: 'tool_result', tool_use_id: 't1', content: JSON.stringify({ AzureAd: { TenantId: seed.guid, ClientSecret: appsettingsSecret } }, null, 2) }] } }, + { ...base, type: 'assistant', uuid: 'a2', parentUuid: 'u2', timestamp: '2026-01-01T00:00:03Z', message: { role: 'assistant', content: [ + { type: 'tool_use', id: 't2', name: 'Bash', input: { command: 'az webapp config appsettings list -g rg -n app && az keyvault secret show --vault-name kv -n Db' } }, + ] } }, + { ...base, type: 'user', uuid: 'u3', parentUuid: 'a2', timestamp: '2026-01-01T00:00:04Z', + message: { role: 'user', content: [{ type: 'tool_result', tool_use_id: 't2', content: + JSON.stringify([{ name: 'Stripe__ApiKey', slotSetting: false, value: appSettingValue }], null, 2) + '\n' + + JSON.stringify({ id: 'https://kv.vault.azure.net/secrets/Db/0123', value: vaultValue }, null, 2) }] }, + toolUseResult: { structuredContent: { name: 'GraphClientSecret', value: mcpSecret } } }, + { ...base, type: 'assistant', uuid: 'a3', parentUuid: 'u3', timestamp: '2026-01-01T00:00:05Z', message: { role: 'assistant', content: [{ type: 'text', text: 'The client secret has expired.' }] } }, + ]; + writeFileSync(join(sourceDir, '-work-contoso', `${SESSION}.jsonl`), lines.map(l => JSON.stringify(l)).join('\n') + '\n'); + + const result = await syncConversations(sourceDir, archiveDir, { skipSummaries: true }); + expect(result.errors).toEqual([]); + const archived = readFileSync(join(archiveDir, '-work-contoso', `${SESSION}.jsonl`), 'utf-8'); + expectNoSecrets(archived, seed, 'archive'); + expectNoSecrets(dbText(dbPath), seed, 'SQLite'); + expectNoSecrets(JSON.stringify(embedSpy.mock.calls), seed, 'embedding input'); + expect(archived).toContain(seed.guid); + archived.split('\n').filter(Boolean).forEach(l => expect(() => JSON.parse(l)).not.toThrow()); + }); + it('the `index` path (indexUnprocessed) parses the redacted archive, not the source', async () => { const s = seedSecrets(77); writeFileSync(join(sourceDir, '-work-contoso', `${SESSION}.jsonl`), claudeTranscript(s)); diff --git a/test/redaction.test.ts b/test/redaction.test.ts index 95466d4e..8e824c3d 100644 --- a/test/redaction.test.ts +++ b/test/redaction.test.ts @@ -187,9 +187,9 @@ describe('redaction: default rules (positive)', () => { expect(redactor.redact(yaml, ctx).text).toContain('host: db.internal'); }); - it('secret-assignment redacts a JSON "clientSecret" field', () => { + it('quoted-secret-assignment redacts a JSON "clientSecret" field', () => { const secret = fake.password(); - expectRedacted(redactor, JSON.stringify({ clientId: fake.guid(), clientSecret: secret }), secret, 'secret-assignment'); + expectRedacted(redactor, JSON.stringify({ clientId: fake.guid(), clientSecret: secret }), secret, 'quoted-secret-assignment'); }); }); @@ -524,3 +524,164 @@ describe('redaction: settings and rules loading', () => { } }); }); + +describe('redaction: key context (.NET / Azure shapes)', () => { + let fake: FakeSecrets; + let redactor: Redactor; + beforeEach(() => { + fake = new FakeSecrets(2112); + redactor = defaults(); + }); + + function expectGone(text: string, secret: string, ruleId?: string) { + const result = redactor.redact(text, ctx); + expect(result.text, text).not.toContain(secret); + if (ruleId) expect(result.text).toContain(`[REDACTED:${ruleId}]`); + return result; + } + + it('appsettings.json fields, no digit required when the key is quoted JSON', () => { + const pw = fake.passphrase(); + const text = JSON.stringify({ AzureAd: { TenantId: fake.guid(), ClientId: fake.guid(), ClientSecret: pw } }, null, 2); + const r = expectGone(text, pw, 'quoted-secret-assignment'); + expect(r.text).toContain('"TenantId"'); + }); + + it('C# object initializers and locals', () => { + const a = fake.passphrase(); + const b = fake.passphrase(); + expectGone(`var cred = new ClientSecretCredential(tenant, client) { ClientSecret = "${a}", };`, a, 'quoted-secret-assignment'); + expectGone(`string apiKey = "${b}";`, b, 'quoted-secret-assignment'); + }); + + it('`az webapp config appsettings list` name/value pairs (name first and value first)', () => { + const a = fake.passphrase(); + const b = fake.passphrase(); + const text = JSON.stringify([ + { name: 'WEBSITE_RUN_FROM_PACKAGE', slotSetting: false, value: '1' }, + { name: 'Stripe__ApiKey', slotSetting: false, value: a }, + { value: b, slotSetting: true, name: 'DB_PASSWORD' }, + ], null, 2); + const r = expectGone(text, a, 'name-value-secret'); + expect(r.text).not.toContain(b); + expect(r.text).toContain('WEBSITE_RUN_FROM_PACKAGE'); + expect(r.text).toContain('"value": "1"'); + }); + + it('Kubernetes env entries', () => { + const a = fake.passphrase(); + expectGone(`env:\n - {"name": "ConnectionStrings__Redis", "value": "redis:6380"}\n - {"name": "JWT_SIGNING_KEY", "value": "${a}"}`, a); + }); + + it('`az keyvault secret show` value (any secret name, nested attributes)', () => { + const a = fake.passphrase(); + const text = JSON.stringify({ + attributes: { created: '2026-01-01T00:00:00+00:00', enabled: true, recoveryLevel: 'Recoverable' }, + contentType: null, + id: 'https://contoso-kv.vault.azure.net/secrets/StorageThing/0123456789abcdef0123456789abcdef', + name: 'StorageThing', + tags: {}, + value: a, + }, null, 2); + const r = expectGone(text, a, 'azure-keyvault-secret'); + expect(r.text).toContain('contoso-kv.vault.azure.net/secrets/StorageThing'); + }); + + it('web.config appSettings in either attribute order', () => { + const a = fake.passphrase(); + const b = fake.passphrase(); + expectGone(`\n \n \n \n`, a, 'xml-appsettings-secret'); + const r = redactor.redact(``, ctx); + expect(r.text).not.toContain(b); + expect(redactor.redact('', ctx).findings).toEqual([]); + }); + + it('publish profiles and XML secret attributes/elements', () => { + const a = fake.passphrase(); + const b = fake.passphrase(); + const r = expectGone(``, a, 'xml-secret-attribute'); + expect(r.text).toContain('userName="$contoso-app"'); + expect(r.text).toContain('https://contoso-app.azurewebsites.net'); + expectGone(`${b}`, b, 'xml-secret-element'); + }); + + it('a GUID in a password slot is redacted (legacy create-for-rbac), but GUID IDs elsewhere are kept', () => { + const pw = fake.guid(); + const tenant = fake.guid(); + const r = expectGone(JSON.stringify({ appId: fake.guid(), password: pw, tenant }), pw); + expect(r.text).toContain(tenant); + }); + + it('leaves placeholders, labels and non-secret neighbours alone', () => { + const untouched = [ + '{"ClientSecret": "#{ClientSecret}#", "ApiKey": "__API_KEY__", "Password": "$(DbPassword)", "Token": "${TOKEN}", "Secret": ""}', + 'ErrorMessage = "Invalid password or username";', + '{"tokenType": "Bearer", "secretName": "db-password", "maxTokens": "4096", "passwordPolicy": "strict"}', + 'options.Password = configuration["Db:Password"];', + '', + '{"name": "TokenEndpoint", "value": "https://login.microsoftonline.com/common/oauth2/v2.0/token"}', + '' + 'a1b2c3d4-0000-1111-2222-333344445555' + '', + 'PWD="/home/user1/src"', + ]; + for (const text of untouched) { + const r = redactor.redact(text, ctx); + expect(r.findings, text).toEqual([]); + } + }); +}); + +describe('redaction: structured JSON (field names as context, keys redacted)', () => { + const redactor = createRedactor(DEFAULT_REDACTION_CONFIG); + const fake = new FakeSecrets(4242); + + it('redacts a value whose own field name is secret-looking, with no shape match', () => { + const pw = fake.passphrase(); + const line = JSON.stringify({ type: 'user', toolUseResult: { structuredContent: { clientId: 'app', clientSecret: pw } } }); + const out = redactJsonlLine(line, redactor, ctx); + expect(out).not.toContain(pw); + expect(JSON.parse(out).toolUseResult.structuredContent).toEqual({ clientId: 'app', clientSecret: '[REDACTED:secret-field]' }); + }); + + it('redacts the value of a {name, value} object whose name is secret-looking', () => { + const pw = fake.passphrase(); + const line = JSON.stringify({ result: [{ name: 'Db__Password', value: pw }, { name: 'Region', value: 'eastus' }] }); + const out = JSON.parse(redactJsonlLine(line, redactor, ctx)); + expect(out.result).toEqual([{ name: 'Db__Password', value: '[REDACTED:secret-field]' }, { name: 'Region', value: 'eastus' }]); + }); + + it('redacts a structured Key Vault secret bundle value', () => { + const pw = fake.passphrase(); + const line = JSON.stringify({ result: { id: 'https://kv.vault.azure.net/secrets/Anything/1', value: pw, attributes: { enabled: true } } }); + expect(redactJsonlLine(line, redactor, ctx)).not.toContain(pw); + }); + + it('redacts secrets used as object keys', () => { + const token = fake.githubToken(); + const line = JSON.stringify({ cache: { [token]: { user: 'octocat' } } }); + const out = redactJsonlLine(line, redactor, ctx); + expect(out).not.toContain(token); + expect(JSON.parse(out).cache['[REDACTED:github-token]']).toEqual({ user: 'octocat' }); + }); + + it('leaves harness structure alone (thinking signatures, usage, token_count, apiKeySource)', () => { + const lines = [ + JSON.stringify({ type: 'assistant', message: { content: [{ type: 'thinking', thinking: 'hmm', signature: fake.base64Blob(200) }], usage: { input_tokens: 12, output_tokens: 3 } } }), + JSON.stringify({ type: 'event_msg', payload: { type: 'token_count', info: { total_token_usage: { input_tokens: 1 } } } }), + JSON.stringify({ type: 'system', subtype: 'init', apiKeySource: 'none', tokenizer: 'cl100k' }), + JSON.stringify({ secretName: 'db-password', passwordPolicy: 'strict', password: '${DB_PASSWORD}', token: '' }), + ]; + for (const line of lines) expect(redactJsonlLine(line, redactor, ctx), line).toBe(line); + }); + + it('is idempotent on structured redactions', () => { + const line = JSON.stringify({ a: { clientSecret: fake.passphrase() }, b: [{ name: 'API_KEY', value: fake.passphrase() }] }); + const once = redactJsonlLine(line, redactor, ctx); + expect(redactJsonlLine(once, redactor, ctx)).toBe(once); + }); + + it('can be turned off via secretFields.enabled=false', () => { + const r = createRedactor({ ...DEFAULT_REDACTION_CONFIG, secretFields: { ...DEFAULT_REDACTION_CONFIG.secretFields, enabled: false } }); + const pw = fake.passphrase(); + expect(redactJsonlLine(JSON.stringify({ x: { clientSecret: pw } }), r, ctx)).toContain(pw); + }); +}); From a6a47970af65801a47e70037d5b0c12001ff741c Mon Sep 17 00:00:00 2001 From: Steven Molen <533340+darthmolen@users.noreply.github.com> Date: Mon, 5 Oct 2026 08:00:54 -0500 Subject: [PATCH 04/15] docs(verify): explain why repair keys resume off the redaction setting repairIndex disables summarizer resume when redaction is enabled, while the indexer disables it only when a redactor loaded. They differ only when non-strict rules fail to load, and there verify is the stricter one. Co-Authored-By: Claude Opus 5.5 --- dist/verify.js | 3 +++ src/verify.ts | 3 +++ 2 files changed, 6 insertions(+) diff --git a/dist/verify.js b/dist/verify.js index d2fba59f..dbf2d60f 100644 --- a/dist/verify.js +++ b/dist/verify.js @@ -127,6 +127,9 @@ export async function repairIndex(issues) { // Generate/update summary const summaryPath = conversationPath.replace('.jsonl', '-summary.txt'); // Under redaction, don't let the Codex fork fallback read the source rollout. + // This keys off the setting, not the loaded redactor as indexer.ts does, which + // is stricter: they differ only when non-strict rules fail to load, and then + // this disables a resume the indexer would allow. It never enables one. const summary = await summarizeConversation(exchanges, undefined, { allowResume: !getRedactionSettings().enabled }); fs.writeFileSync(summaryPath, summary, 'utf-8'); console.log(` Created summary: ${summary.split(/\s+/).length} words`); diff --git a/src/verify.ts b/src/verify.ts index f0b86110..836eec0c 100644 --- a/src/verify.ts +++ b/src/verify.ts @@ -162,6 +162,9 @@ export async function repairIndex(issues: VerificationResult): Promise { // Generate/update summary const summaryPath = conversationPath.replace('.jsonl', '-summary.txt'); // Under redaction, don't let the Codex fork fallback read the source rollout. + // This keys off the setting, not the loaded redactor as indexer.ts does, which + // is stricter: they differ only when non-strict rules fail to load, and then + // this disables a resume the indexer would allow. It never enables one. const summary = await summarizeConversation(exchanges, undefined, { allowResume: !getRedactionSettings().enabled }); fs.writeFileSync(summaryPath, summary, 'utf-8'); console.log(` Created summary: ${summary.split(/\s+/).length} words`); From 6a49ea12b2b942e4c5435c30e32a1c6b57609b4f Mon Sep 17 00:00:00 2001 From: Steven Molen <533340+darthmolen@users.noreply.github.com> Date: Mon, 5 Oct 2026 08:17:21 -0500 Subject: [PATCH 05/15] fix(redaction): create redacted copies with the source file's mode copyFileRedacted opened its output with the 0666 & umask default, so a 0600 transcript became world-readable in the archive, and redact --rewrite widened existing archive files the same way. copyFileSync, used before redaction, kept the source mode. Co-Authored-By: Claude Opus 5.5 --- dist/redaction.js | 4 +++- src/redaction.ts | 4 +++- test/redaction.test.ts | 12 +++++++++++- 3 files changed, 17 insertions(+), 3 deletions(-) diff --git a/dist/redaction.js b/dist/redaction.js index 83f76944..d2679fbc 100644 --- a/dist/redaction.js +++ b/dist/redaction.js @@ -556,7 +556,9 @@ export function copyFileRedacted(src, dest, redactor, ctx, tally) { const fdIn = fs.openSync(src, 'r'); let fdOut; try { - fdOut = fs.openSync(dest, 'w'); + // Create with the source's mode, as copyFileSync would, so a 0600 + // transcript doesn't become world-readable in the archive. + fdOut = fs.openSync(dest, 'w', fs.fstatSync(fdIn).mode & 0o777); const buf = Buffer.allocUnsafe(COPY_CHUNK_BYTES); const decoder = new StringDecoder('utf8'); let pending = ''; diff --git a/src/redaction.ts b/src/redaction.ts index afdeaea5..df284b56 100644 --- a/src/redaction.ts +++ b/src/redaction.ts @@ -723,7 +723,9 @@ export function copyFileRedacted( const fdIn = fs.openSync(src, 'r'); let fdOut: number | undefined; try { - fdOut = fs.openSync(dest, 'w'); + // Create with the source's mode, as copyFileSync would, so a 0600 + // transcript doesn't become world-readable in the archive. + fdOut = fs.openSync(dest, 'w', fs.fstatSync(fdIn).mode & 0o777); const buf = Buffer.allocUnsafe(COPY_CHUNK_BYTES); const decoder = new StringDecoder('utf8'); let pending = ''; diff --git a/test/redaction.test.ts b/test/redaction.test.ts index 8e824c3d..65e7b5d9 100644 --- a/test/redaction.test.ts +++ b/test/redaction.test.ts @@ -1,5 +1,5 @@ import { describe, it, expect, beforeEach, afterEach, vi } from 'vitest'; -import { mkdtempSync, writeFileSync, readFileSync, rmSync, mkdirSync } from 'fs'; +import { mkdtempSync, writeFileSync, readFileSync, rmSync, mkdirSync, chmodSync, statSync } from 'fs'; import { join } from 'path'; import { tmpdir } from 'os'; @@ -423,6 +423,16 @@ describe('redaction: copyFileRedacted', () => { copyFileRedacted(src, dest, redactor, ctx); expect(readFileSync(dest, 'utf-8')).toBe(readFileSync(src, 'utf-8')); }); + + // Windows ignores POSIX mode bits. + it.skipIf(process.platform === 'win32')('creates dest with the source mode, not the 0666 default', () => { + const src = join(dir, 'private.jsonl'); + const dest = join(dir, 'private-out.jsonl'); + writeFileSync(src, '{"type":"user"}\n'); + chmodSync(src, 0o600); + copyFileRedacted(src, dest, defaults(), ctx); + expect(statSync(dest).mode & 0o777).toBe(0o600); + }); }); describe('redaction: settings and rules loading', () => { From fff43262b24b55583bdc11a96700f2552e77b98e Mon Sep 17 00:00:00 2001 From: Steven Molen <533340+darthmolen@users.noreply.github.com> Date: Mon, 5 Oct 2026 08:17:21 -0500 Subject: [PATCH 06/15] fix(indexer): force the archive refresh for already-indexed transcripts indexUnprocessed used to re-copy every indexed transcript and parse the source. Parsing the redacted archive instead made it depend on copyIfNewer's mtime check, which skips an append that lands within the mtime granularity, so those exchanges were never indexed. copyIfNewer takes a force flag, and indexUnprocessed sets it once a file is indexed. Co-Authored-By: Claude Opus 5.5 --- dist/indexer.js | 7 ++++--- dist/sync.d.ts | 6 +++++- dist/sync.js | 8 ++++++-- src/indexer.ts | 7 ++++--- src/sync.ts | 9 +++++++-- test/redaction-pipeline.test.ts | 31 +++++++++++++++++++++++++++++++ 6 files changed, 57 insertions(+), 11 deletions(-) diff --git a/dist/indexer.js b/dist/indexer.js index 5b4c803d..0fb4e2fc 100644 --- a/dist/indexer.js +++ b/dist/indexer.js @@ -297,9 +297,10 @@ export async function indexUnprocessed(concurrency = 1, noSummaries = false) { const hw = db.prepare('SELECT COALESCE(MAX(line_end), 0) as maxLine FROM exchanges WHERE archive_path = ?').get(archivePath); const maxIndexedLine = hw.maxLine; try { - // Refresh the (redacted) archive when the source has grown, then parse - // the archive so the index only sees redacted text. - copyIfNewer(sourcePath, archivePath, redactor, tally); + // Refresh the (redacted) archive, then parse the archive so the index + // only sees redacted text. Force the refresh once the file is indexed: + // an append inside the mtime granularity would otherwise be skipped. + copyIfNewer(sourcePath, archivePath, redactor, tally, maxIndexedLine > 0); // Parse and filter to exchanges past the high-water mark const exchanges = await parseConversation(archivePath, project, archivePath); const newExchanges = maxIndexedLine > 0 diff --git a/dist/sync.d.ts b/dist/sync.d.ts index 065e5386..36b5afb5 100644 --- a/dist/sync.d.ts +++ b/dist/sync.d.ts @@ -52,7 +52,11 @@ export declare function buildSyncOptionsFromEnv(env: NodeJS.ProcessEnv): SyncOpt * older. This is the redaction choke point: with a redactor, the copy is * redacted line by line (see redaction.ts), and every downstream stage * (index, embeddings, summaries, show/read) reads the archive. + * + * `force` copies even when the mtimes say the archive is current: an append + * within the mtime granularity (or the millisecond the rounding below adds) + * leaves the source no newer than the archive. */ -export declare function copyIfNewer(src: string, dest: string, redactor?: Redactor | null, tally?: FindingsTally): boolean; +export declare function copyIfNewer(src: string, dest: string, redactor?: Redactor | null, tally?: FindingsTally, force?: boolean): boolean; export declare function extractSessionIdFromPath(filePath: string): string | null; export declare function syncConversations(sourceDir: string, destDir: string, options?: SyncOptions): Promise; diff --git a/dist/sync.js b/dist/sync.js index 380037e1..6bb83b75 100644 --- a/dist/sync.js +++ b/dist/sync.js @@ -128,15 +128,19 @@ export function buildSyncOptionsFromEnv(env) { * older. This is the redaction choke point: with a redactor, the copy is * redacted line by line (see redaction.ts), and every downstream stage * (index, embeddings, summaries, show/read) reads the archive. + * + * `force` copies even when the mtimes say the archive is current: an append + * within the mtime granularity (or the millisecond the rounding below adds) + * leaves the source no newer than the archive. */ -export function copyIfNewer(src, dest, redactor = null, tally) { +export function copyIfNewer(src, dest, redactor = null, tally, force = false) { // Ensure destination directory exists const destDir = path.dirname(dest); if (!fs.existsSync(destDir)) { fs.mkdirSync(destDir, { recursive: true }); } // Check if destination exists and is up-to-date - if (fs.existsSync(dest)) { + if (!force && fs.existsSync(dest)) { const srcStat = fs.statSync(src); const destStat = fs.statSync(dest); if (destStat.mtimeMs >= srcStat.mtimeMs) { diff --git a/src/indexer.ts b/src/indexer.ts index 5ba56ae2..df9a5986 100644 --- a/src/indexer.ts +++ b/src/indexer.ts @@ -381,9 +381,10 @@ export async function indexUnprocessed(concurrency: number = 1, noSummaries: boo const maxIndexedLine = hw.maxLine; try { - // Refresh the (redacted) archive when the source has grown, then parse - // the archive so the index only sees redacted text. - copyIfNewer(sourcePath, archivePath, redactor, tally); + // Refresh the (redacted) archive, then parse the archive so the index + // only sees redacted text. Force the refresh once the file is indexed: + // an append inside the mtime granularity would otherwise be skipped. + copyIfNewer(sourcePath, archivePath, redactor, tally, maxIndexedLine > 0); // Parse and filter to exchanges past the high-water mark const exchanges = await parseConversation(archivePath, project, archivePath); diff --git a/src/sync.ts b/src/sync.ts index b2de456c..62fa00ec 100644 --- a/src/sync.ts +++ b/src/sync.ts @@ -145,12 +145,17 @@ export function buildSyncOptionsFromEnv(env: NodeJS.ProcessEnv): SyncOptions { * older. This is the redaction choke point: with a redactor, the copy is * redacted line by line (see redaction.ts), and every downstream stage * (index, embeddings, summaries, show/read) reads the archive. + * + * `force` copies even when the mtimes say the archive is current: an append + * within the mtime granularity (or the millisecond the rounding below adds) + * leaves the source no newer than the archive. */ export function copyIfNewer( src: string, dest: string, redactor: Redactor | null = null, - tally?: FindingsTally + tally?: FindingsTally, + force = false ): boolean { // Ensure destination directory exists const destDir = path.dirname(dest); @@ -159,7 +164,7 @@ export function copyIfNewer( } // Check if destination exists and is up-to-date - if (fs.existsSync(dest)) { + if (!force && fs.existsSync(dest)) { const srcStat = fs.statSync(src); const destStat = fs.statSync(dest); if (destStat.mtimeMs >= srcStat.mtimeMs) { diff --git a/test/redaction-pipeline.test.ts b/test/redaction-pipeline.test.ts index 89f60ea4..770a9cd0 100644 --- a/test/redaction-pipeline.test.ts +++ b/test/redaction-pipeline.test.ts @@ -1,6 +1,7 @@ import { describe, it, expect, beforeEach, afterEach, vi } from 'vitest'; import { mkdtempSync, mkdirSync, writeFileSync, readFileSync, rmSync, existsSync, readdirSync, statSync, + appendFileSync, utimesSync, } from 'fs'; import { join } from 'path'; import { tmpdir } from 'os'; @@ -344,6 +345,36 @@ describe('redaction pipeline', () => { expectNoSecrets(JSON.stringify(summarizeSpy.mock.calls), s, 'summarizer input (index path)'); expect(summarizeSpy.mock.calls[0][2]).toMatchObject({ allowResume: false }); }); + + it('the `index` path refreshes an indexed archive even when the append left the source mtime no newer', async () => { + const fake = new FakeSecrets(78); + const secret = fake.azureClientSecret(); + const base = { sessionId: SESSION, isSidechain: false, cwd: '/work/contoso', gitBranch: 'main', version: '2.0.0' }; + const exchange = (n: number, text: string) => [ + { ...base, type: 'user', uuid: `u${n}`, parentUuid: n > 1 ? `a${n - 1}` : null, + timestamp: new Date(Date.UTC(2026, 0, 1, 0, n)).toISOString(), + message: { role: 'user', content: text } }, + { ...base, type: 'assistant', uuid: `a${n}`, parentUuid: `u${n}`, + timestamp: new Date(Date.UTC(2026, 0, 1, 0, n, 30)).toISOString(), + message: { role: 'assistant', content: [{ type: 'text', text: `Reply ${n}.` }] } }, + ].map(l => JSON.stringify(l) + '\n').join(''); + const src = join(sourceDir, '-work-contoso', `${SESSION}.jsonl`); + + writeFileSync(src, exchange(1, 'First question.')); + await indexUnprocessed(1, true); + const archived = walkFiles(archiveDir).filter(f => f.endsWith('.jsonl'))[0]; + const archivedMtime = statSync(archived).mtime; + + // Append, then pin the source mtime to the archive's: copyIfNewer's mtime + // check alone would treat the archive as current. + appendFileSync(src, exchange(2, `Second question with ${secret} as the secret`)); + utimesSync(src, archivedMtime, archivedMtime); + await indexUnprocessed(1, true); + + const text = dbText(dbPath); + expect(text).toContain('Second question with [REDACTED:'); + expect(text).not.toContain(secret); + }); }); describe('redaction: opencode staging export', () => { From c931a142ab6279cf72309a98b0b482db9f2c7910 Mon Sep 17 00:00:00 2001 From: Steven Molen <533340+darthmolen@users.noreply.github.com> Date: Mon, 5 Oct 2026 08:17:21 -0500 Subject: [PATCH 07/15] fix(redaction): match an Entra client secret that ends a sentence The rule's trailing lookahead excluded '.', which is also in the secret alphabet, so a 40-char secret followed by a period could match neither with nor without it and was left in clear text. Co-Authored-By: Claude Opus 5.5 --- dist/redaction-rules.js | 4 ++-- src/redaction-rules.ts | 4 ++-- test/redaction.test.ts | 6 ++++++ 3 files changed, 10 insertions(+), 4 deletions(-) diff --git a/dist/redaction-rules.js b/dist/redaction-rules.js index 153da735..27eecdcb 100644 --- a/dist/redaction-rules.js +++ b/dist/redaction-rules.js @@ -113,8 +113,8 @@ export const DEFAULT_REDACTION_CONFIG = { }, { id: 'azure-client-secret', - description: 'Entra ID (Azure AD) application client secret: 3 chars, a digit, "Q~", 31-34 chars', - pattern: String.raw `(? { expectRedacted(redactor, `use ${secret} as the client secret`, secret, 'azure-client-secret'); }); + it('azure-client-secret at the end of a sentence', () => { + const secret = fake.azureClientSecret(); + const out = expectRedacted(redactor, `The client secret is ${secret}.`, secret, 'azure-client-secret'); + expect(out.text).toBe('The client secret is [REDACTED:azure-client-secret].'); + }); + it('azure-storage-key in `az storage account keys list` output', () => { const secret = fake.azureStorageKey(); const text = `[\n {\n "keyName": "key1",\n "permissions": "FULL",\n "value": "${secret}"\n }\n]`; From 3b7edc25ecf222b44ee3457db785fc73ec6ddd8e Mon Sep 17 00:00:00 2001 From: Steven Molen <533340+darthmolen@users.noreply.github.com> Date: Mon, 5 Oct 2026 08:36:40 -0500 Subject: [PATCH 08/15] fix(redaction): match connection-string keywords in any case ADO.NET keywords are case-insensitive, but the rule only matched the listed capitalization, so `password=` and `PASSWORD=` values were kept. A letters-only value isn't caught by secret-assignment either, which needs a digit. A path after `Pwd=` is still left alone: that's the shell PWD variable. Co-Authored-By: Claude Opus 5.5 --- dist/redaction-rules.js | 5 +++-- docs/REDACTION.md | 2 +- src/redaction-rules.ts | 5 +++-- test/redaction.test.ts | 9 +++++++++ 4 files changed, 16 insertions(+), 5 deletions(-) diff --git a/dist/redaction-rules.js b/dist/redaction-rules.js index 27eecdcb..46b62b8c 100644 --- a/dist/redaction-rules.js +++ b/dist/redaction-rules.js @@ -50,8 +50,9 @@ export const DEFAULT_REDACTION_CONFIG = { }, { id: 'connection-string-secret', - description: 'Secret value inside an Azure/ADO.NET connection string; account, server and database names are kept', - pattern: String.raw `\b(?:AccountKey|SharedAccessKey|SharedAccessSignature|SharedSecret|ClientSecret|Password|Pwd)=(?![\[$<{%])([^;"'\s]+)`, + description: 'Secret value inside an Azure/ADO.NET connection string; account, server and database names are kept. Keywords are case-insensitive, as in ADO.NET; a path after Pwd= is the shell PWD variable and is kept', + pattern: String.raw `\b(?:(?:AccountKey|SharedAccessKey|SharedAccessSignature|SharedSecret|ClientSecret|Password)=|Pwd=(?![/~]))(?![\[$<{%])([^;"'\s]+)`, + flags: 'i', secretGroup: 1, keywords: ['accountkey', 'sharedaccess', 'sharedsecret', 'clientsecret', 'password', 'pwd'], }, diff --git a/docs/REDACTION.md b/docs/REDACTION.md index f3bafbe1..3e8a796b 100644 --- a/docs/REDACTION.md +++ b/docs/REDACTION.md @@ -49,7 +49,7 @@ summary: | Rule ID | Catches | |---|---| | `private-key-block` | PEM and OpenSSH private keys, including truncated ones | -| `connection-string-secret` | `AccountKey=`, `SharedAccessKey=`, `SharedAccessSignature=`, `Password=`, `Pwd=` in connection strings. Only the value is redacted; server, account, and database names stay. | +| `connection-string-secret` | `AccountKey=`, `SharedAccessKey=`, `SharedAccessSignature=`, `Password=`, `Pwd=` in connection strings, in any case. Only the value is redacted; server, account, and database names stay. | | `azure-sas-token` | The `sig=` of a SAS URL. The URL and other parameters stay. | | `jwt` | JWTs, including Entra ID / Azure access tokens | | `anthropic-api-key`, `openai-api-key`, `github-token`, `aws-access-key-id`, `aws-secret-access-key`, `slack-token`, `google-api-key`, `npm-token` | Provider keys with a recognizable prefix | diff --git a/src/redaction-rules.ts b/src/redaction-rules.ts index 1814e214..b0d6d99a 100644 --- a/src/redaction-rules.ts +++ b/src/redaction-rules.ts @@ -58,8 +58,9 @@ export const DEFAULT_REDACTION_CONFIG: RedactionConfig = { }, { id: 'connection-string-secret', - description: 'Secret value inside an Azure/ADO.NET connection string; account, server and database names are kept', - pattern: String.raw`\b(?:AccountKey|SharedAccessKey|SharedAccessSignature|SharedSecret|ClientSecret|Password|Pwd)=(?![\[$<{%])([^;"'\s]+)`, + description: 'Secret value inside an Azure/ADO.NET connection string; account, server and database names are kept. Keywords are case-insensitive, as in ADO.NET; a path after Pwd= is the shell PWD variable and is kept', + pattern: String.raw`\b(?:(?:AccountKey|SharedAccessKey|SharedAccessSignature|SharedSecret|ClientSecret|Password)=|Pwd=(?![/~]))(?![\[$<{%])([^;"'\s]+)`, + flags: 'i', secretGroup: 1, keywords: ['accountkey', 'sharedaccess', 'sharedsecret', 'clientsecret', 'password', 'pwd'], }, diff --git a/test/redaction.test.ts b/test/redaction.test.ts index 1bde7140..bf6ed5e2 100644 --- a/test/redaction.test.ts +++ b/test/redaction.test.ts @@ -93,6 +93,15 @@ describe('redaction: default rules (positive)', () => { expect(result.text).toContain('SharedAccessKeyName=RootManageSharedAccessKey'); }); + it('connection-string-secret keywords are case-insensitive, as in ADO.NET', () => { + // Letters only: secret-assignment needs a digit, so this rule is the only catch. + const pw = fake.chars('abcdefghijklmnopqrstuvwxyz', 19); + for (const keyword of ['password', 'PASSWORD', 'pwd', 'accountkey']) { + const result = expectRedacted(redactor, `Server=db;Database=app;${keyword}=${pw};`, pw, 'connection-string-secret'); + expect(result.text).toContain(`${keyword}=[REDACTED:connection-string-secret];`); + } + }); + it('azure-sas-token redacts the sig while keeping the blob URL', () => { const sig = fake.sasSignature(); const text = `https://contoso.blob.core.windows.net/backups/db.bak?sv=2022-11-02&ss=b&srt=co&sp=rl&se=2026-12-31T00:00:00Z&sig=${sig}`; From 70592f547fa9f9aa0d93329ca6c9b6ec1d73e0e7 Mon Sep 17 00:00:00 2001 From: Steven Molen <533340+darthmolen@users.noreply.github.com> Date: Mon, 5 Oct 2026 08:37:25 -0500 Subject: [PATCH 09/15] fix(redaction): redact the part of a match outside an existing token applyRule skipped any match that overlapped a token, so in partly redacted text (an older rule set, before redact --rewrite) `Password=[REDACTED:jwt]` kept the secret. The existing tokens are now kept and each stretch between them with a letter or digit is replaced. A stretch already fully covered yields no change, so a second pass is still a no-op. Co-Authored-By: Claude Opus 5.5 --- dist/redaction.js | 49 ++++++++++++++++++++++++++++++++++++----- src/redaction.ts | 50 ++++++++++++++++++++++++++++++++++++++---- test/redaction.test.ts | 12 ++++++++++ 3 files changed, 102 insertions(+), 9 deletions(-) diff --git a/dist/redaction.js b/dist/redaction.js index d2679fbc..222d5eb7 100644 --- a/dist/redaction.js +++ b/dist/redaction.js @@ -279,6 +279,37 @@ function overlapsAny(spans, start, end) { return true; return false; } +/** + * Rewrite text[start, end) keeping the existing tokens in it and replacing each + * stretch between them that has a letter or digit with `token`. Returns null + * when nothing outside the tokens needs redacting, which keeps this idempotent. + */ +function redactAroundTokens(text, spans, start, end, token) { + let out = ''; + let pos = start; + let replaced = false; + const flush = (to) => { + const piece = text.slice(pos, to); + if (/[A-Za-z0-9]/.test(piece)) { + out += token; + replaced = true; + } + else { + out += piece; + } + }; + for (const [s, e] of spans) { + if (e <= start || s >= end) + continue; + if (s > pos) + flush(s); + out += text.slice(Math.max(s, pos), Math.min(e, end)); + pos = Math.min(e, end); + } + if (pos < end) + flush(end); + return replaced ? out : null; +} function shannonEntropy(text) { const counts = new Map(); for (const ch of text) @@ -291,9 +322,10 @@ function shannonEntropy(text) { return entropy; } /** - * Replace each accepted match (or its secretGroup) with a token. A candidate is - * skipped when it overlaps an existing token (idempotency) or fully matches an - * allowlist pattern. + * Replace each accepted match (or its secretGroup) with a token. A candidate + * that overlaps existing tokens keeps them, and only the text around them is + * redacted (so re-running is a no-op). A candidate is skipped when it fully + * matches an allowlist pattern. */ function applyRule(text, rule, isAllowed) { const spans = text.includes(TOKEN_PREFIX) ? tokenSpans(text) : null; @@ -313,10 +345,17 @@ function applyRule(text, rule, isAllowed) { const [start, end] = range; if (end <= start || start < last) continue; - if (spans && overlapsAny(spans, start, end)) - continue; if (rule.useAllowlist && isAllowed(text.slice(start, end))) continue; + if (spans && overlapsAny(spans, start, end)) { + const rewritten = redactAroundTokens(text, spans, start, end, token); + if (rewritten === null) + continue; + out += text.slice(last, start) + rewritten; + last = end; + count++; + continue; + } out += text.slice(last, start) + token; last = end; count++; diff --git a/src/redaction.ts b/src/redaction.ts index df284b56..b60afc4f 100644 --- a/src/redaction.ts +++ b/src/redaction.ts @@ -431,6 +431,40 @@ function overlapsAny(spans: Array<[number, number]>, start: number, end: number) return false; } +/** + * Rewrite text[start, end) keeping the existing tokens in it and replacing each + * stretch between them that has a letter or digit with `token`. Returns null + * when nothing outside the tokens needs redacting, which keeps this idempotent. + */ +function redactAroundTokens( + text: string, + spans: Array<[number, number]>, + start: number, + end: number, + token: string +): string | null { + let out = ''; + let pos = start; + let replaced = false; + const flush = (to: number) => { + const piece = text.slice(pos, to); + if (/[A-Za-z0-9]/.test(piece)) { + out += token; + replaced = true; + } else { + out += piece; + } + }; + for (const [s, e] of spans) { + if (e <= start || s >= end) continue; + if (s > pos) flush(s); + out += text.slice(Math.max(s, pos), Math.min(e, end)); + pos = Math.min(e, end); + } + if (pos < end) flush(end); + return replaced ? out : null; +} + function shannonEntropy(text: string): number { const counts = new Map(); for (const ch of text) counts.set(ch, (counts.get(ch) ?? 0) + 1); @@ -443,9 +477,10 @@ function shannonEntropy(text: string): number { } /** - * Replace each accepted match (or its secretGroup) with a token. A candidate is - * skipped when it overlaps an existing token (idempotency) or fully matches an - * allowlist pattern. + * Replace each accepted match (or its secretGroup) with a token. A candidate + * that overlaps existing tokens keeps them, and only the text around them is + * redacted (so re-running is a no-op). A candidate is skipped when it fully + * matches an allowlist pattern. */ function applyRule(text: string, rule: CompiledRule, isAllowed: (s: string) => boolean): { text: string; count: number } { const spans = text.includes(TOKEN_PREFIX) ? tokenSpans(text) : null; @@ -462,8 +497,15 @@ function applyRule(text: string, rule: CompiledRule, isAllowed: (s: string) => b if (!range) continue; const [start, end] = range; if (end <= start || start < last) continue; - if (spans && overlapsAny(spans, start, end)) continue; if (rule.useAllowlist && isAllowed(text.slice(start, end))) continue; + if (spans && overlapsAny(spans, start, end)) { + const rewritten = redactAroundTokens(text, spans, start, end, token); + if (rewritten === null) continue; + out += text.slice(last, start) + rewritten; + last = end; + count++; + continue; + } out += text.slice(last, start) + token; last = end; count++; diff --git a/test/redaction.test.ts b/test/redaction.test.ts index bf6ed5e2..b75bbfde 100644 --- a/test/redaction.test.ts +++ b/test/redaction.test.ts @@ -311,6 +311,18 @@ describe('redaction: idempotency', () => { } } }); + + it('a match that runs into an existing token redacts the part outside it, keeping the token', () => { + // Text already partly redacted, e.g. by an older rule set before `redact --rewrite`. + const redactor = defaults(); + const fake = new FakeSecrets(12); + const prefix = fake.chars('abcdefghijklmnopqrstuvwxyz0123456789', 10); + const once = redactor.redact(`Server=db;Password=${prefix}[REDACTED:jwt];`, ctx); + expect(once.text).toBe('Server=db;Password=[REDACTED:connection-string-secret][REDACTED:jwt];'); + const twice = redactor.redact(once.text, ctx); + expect(twice.text).toBe(once.text); + expect(twice.findings).toEqual([]); + }); }); describe('redaction: entropy fallback', () => { From 64e005904137afa8d39a157707b5241a2770e968 Mon Sep 17 00:00:00 2001 From: Steven Molen <533340+darthmolen@users.noreply.github.com> Date: Mon, 5 Oct 2026 08:38:02 -0500 Subject: [PATCH 10/15] fix(redaction): redact a secret field whole even if it holds a token The field-name rule skipped any value containing a token, so in {"password": " "} the JWT rule's token kept the field from being replaced and the word stayed. Only values made up entirely of tokens are skipped now. Co-Authored-By: Claude Opus 5.5 --- dist/redaction.js | 9 ++++++--- src/redaction.ts | 10 +++++++--- test/redaction.test.ts | 8 ++++++++ 3 files changed, 21 insertions(+), 6 deletions(-) diff --git a/dist/redaction.js b/dist/redaction.js index 222d5eb7..4f728424 100644 --- a/dist/redaction.js +++ b/dist/redaction.js @@ -502,8 +502,8 @@ function redactTree(node, redactor, ctx, tally) { // Field-name context: the key (or a {name, value} pair's name, or a Key // Vault secret id) says the value is a secret even when its shape matches // no rule — e.g. a letters-only clientSecret field in a structured - // MCP result. Runs after the text rules, so a value they already redacted - // in part (a connection string) keeps its searchable remainder. + // MCP result. Runs after the text rules; a value they redacted only in part + // is still replaced whole, since the field name says all of it is secret. const redactWhole = (key) => { if (!isRedactableWhole(obj[key])) return; @@ -553,10 +553,13 @@ function setOwn(obj, key, value) { // defineProperty so a "__proto__" key from JSON.parse stays an own property. Object.defineProperty(obj, key, { value, writable: true, enumerable: true, configurable: true }); } +// A value that is only tokens is already redacted; anything else in a secret +// field, even alongside a token, is replaced whole. +const ONLY_TOKENS = /^\s*(?:\[REDACTED:[a-z0-9][a-z0-9-]*\]\s*)+$/; function isRedactableWhole(value) { return typeof value === 'string' && value.trim().length > 0 && - !value.includes(TOKEN_PREFIX) && + !ONLY_TOKENS.test(value) && !PLACEHOLDER_VALUE.test(value.trim()); } /** diff --git a/src/redaction.ts b/src/redaction.ts index b60afc4f..0a6563bb 100644 --- a/src/redaction.ts +++ b/src/redaction.ts @@ -658,8 +658,8 @@ function redactTree( // Field-name context: the key (or a {name, value} pair's name, or a Key // Vault secret id) says the value is a secret even when its shape matches // no rule — e.g. a letters-only clientSecret field in a structured - // MCP result. Runs after the text rules, so a value they already redacted - // in part (a connection string) keeps its searchable remainder. + // MCP result. Runs after the text rules; a value they redacted only in part + // is still replaced whole, since the field name says all of it is secret. const redactWhole = (key: string) => { if (!isRedactableWhole(obj[key])) return; setOwn(obj, key, tokenFor(FIELD_RULE_ID)); @@ -710,10 +710,14 @@ function setOwn(obj: Record, key: string, value: unknown): void Object.defineProperty(obj, key, { value, writable: true, enumerable: true, configurable: true }); } +// A value that is only tokens is already redacted; anything else in a secret +// field, even alongside a token, is replaced whole. +const ONLY_TOKENS = /^\s*(?:\[REDACTED:[a-z0-9][a-z0-9-]*\]\s*)+$/; + function isRedactableWhole(value: unknown): value is string { return typeof value === 'string' && value.trim().length > 0 && - !value.includes(TOKEN_PREFIX) && + !ONLY_TOKENS.test(value) && !PLACEHOLDER_VALUE.test(value.trim()); } diff --git a/test/redaction.test.ts b/test/redaction.test.ts index b75bbfde..745e1c17 100644 --- a/test/redaction.test.ts +++ b/test/redaction.test.ts @@ -710,6 +710,14 @@ describe('redaction: structured JSON (field names as context, keys redacted)', ( for (const line of lines) expect(redactJsonlLine(line, redactor, ctx), line).toBe(line); }); + it('redacts a secret field whole when a text rule already redacted only part of it', () => { + const word = fake.chars('abcdefghijklmnopqrstuvwxyz', 12); + const line = JSON.stringify({ password: `${word} ${fake.jwt()}` }); + const out = redactJsonlLine(line, redactor, ctx); + expect(out).not.toContain(word); + expect(JSON.parse(out).password).toBe('[REDACTED:secret-field]'); + }); + it('is idempotent on structured redactions', () => { const line = JSON.stringify({ a: { clientSecret: fake.passphrase() }, b: [{ name: 'API_KEY', value: fake.passphrase() }] }); const once = redactJsonlLine(line, redactor, ctx); From 1cb16af9c42acba1164bb70db9f3fb3ed2730224 Mon Sep 17 00:00:00 2001 From: Steven Molen <533340+darthmolen@users.noreply.github.com> Date: Mon, 5 Oct 2026 08:38:38 -0500 Subject: [PATCH 11/15] fix(verify): load the redaction rules before index repair repairIndex checked only whether redaction was enabled, so with a corrupt rules file in strict mode it still deleted, re-indexed and summarized, while sync and index abort. It now loads the redactor first, which throws in strict mode, and derives allowResume from the result as the indexer does. Co-Authored-By: Claude Opus 5.5 --- dist/verify.js | 10 +++++----- src/verify.ts | 11 ++++++----- test/verify.test.ts | 31 +++++++++++++++++++++++++++++++ 3 files changed, 42 insertions(+), 10 deletions(-) diff --git a/dist/verify.js b/dist/verify.js index dbf2d60f..54df6a85 100644 --- a/dist/verify.js +++ b/dist/verify.js @@ -4,7 +4,7 @@ import { parseConversation } from './parser.js'; import { initDatabase, getAllExchanges, getFileLastIndexed } from './db.js'; import { getArchiveDir, getExcludedProjects, findJsonlFiles, statIfExists } from './paths.js'; import { isErroredSentinel } from './summary-sentinel.js'; -import { getRedactionSettings } from './redaction.js'; +import { loadRedactor } from './redaction.js'; export async function verifyIndex() { const result = { missing: [], @@ -94,6 +94,9 @@ export async function verifyIndex() { } export async function repairIndex(issues) { console.log('Repairing index...'); + // Load before touching the index: strict mode fails closed here, as in sync + // and index. + const redactor = loadRedactor(); // To avoid circular dependencies, we import the indexer functions dynamically const { initDatabase, insertExchange, deleteExchange } = await import('./db.js'); const { parseConversation } = await import('./parser.js'); @@ -127,10 +130,7 @@ export async function repairIndex(issues) { // Generate/update summary const summaryPath = conversationPath.replace('.jsonl', '-summary.txt'); // Under redaction, don't let the Codex fork fallback read the source rollout. - // This keys off the setting, not the loaded redactor as indexer.ts does, which - // is stricter: they differ only when non-strict rules fail to load, and then - // this disables a resume the indexer would allow. It never enables one. - const summary = await summarizeConversation(exchanges, undefined, { allowResume: !getRedactionSettings().enabled }); + const summary = await summarizeConversation(exchanges, undefined, { allowResume: redactor === null }); fs.writeFileSync(summaryPath, summary, 'utf-8'); console.log(` Created summary: ${summary.split(/\s+/).length} words`); // Index exchanges diff --git a/src/verify.ts b/src/verify.ts index 836eec0c..bc9b6dd1 100644 --- a/src/verify.ts +++ b/src/verify.ts @@ -4,7 +4,7 @@ import { parseConversation } from './parser.js'; import { initDatabase, getAllExchanges, getFileLastIndexed } from './db.js'; import { getArchiveDir, getExcludedProjects, findJsonlFiles, statIfExists } from './paths.js'; import { isErroredSentinel } from './summary-sentinel.js'; -import { getRedactionSettings } from './redaction.js'; +import { loadRedactor } from './redaction.js'; export interface VerificationResult { missing: Array<{ path: string; reason: string }>; @@ -122,6 +122,10 @@ export async function verifyIndex(): Promise { export async function repairIndex(issues: VerificationResult): Promise { console.log('Repairing index...'); + // Load before touching the index: strict mode fails closed here, as in sync + // and index. + const redactor = loadRedactor(); + // To avoid circular dependencies, we import the indexer functions dynamically const { initDatabase, insertExchange, deleteExchange } = await import('./db.js'); const { parseConversation } = await import('./parser.js'); @@ -162,10 +166,7 @@ export async function repairIndex(issues: VerificationResult): Promise { // Generate/update summary const summaryPath = conversationPath.replace('.jsonl', '-summary.txt'); // Under redaction, don't let the Codex fork fallback read the source rollout. - // This keys off the setting, not the loaded redactor as indexer.ts does, which - // is stricter: they differ only when non-strict rules fail to load, and then - // this disables a resume the indexer would allow. It never enables one. - const summary = await summarizeConversation(exchanges, undefined, { allowResume: !getRedactionSettings().enabled }); + const summary = await summarizeConversation(exchanges, undefined, { allowResume: redactor === null }); fs.writeFileSync(summaryPath, summary, 'utf-8'); console.log(` Created summary: ${summary.split(/\s+/).length} words`); diff --git a/test/verify.test.ts b/test/verify.test.ts index ae70da81..dac46ea1 100644 --- a/test/verify.test.ts +++ b/test/verify.test.ts @@ -281,6 +281,37 @@ describe('repairIndex', () => { dbAfter.close(); }); + it('fails closed on a corrupt rules file in strict mode, before changing the index', async () => { + const db = initDatabase(); + insertExchange(db, { + id: 'orphan-strict-1', + project: 'deleted-project', + timestamp: '2024-01-01T00:00:00Z', + userMessage: 'Deleted', + assistantMessage: 'Still indexed', + archivePath: path.join(archiveDir, 'deleted-project', 'deleted.jsonl'), + lineStart: 1, + lineEnd: 2 + }, new Array(384).fill(0.1)); + db.close(); + + const rulesPath = path.join(testDir, 'redaction-rules.json'); + fs.writeFileSync(rulesPath, '{ not json'); + process.env.EPISODIC_MEMORY_REDACTION_RULES = rulesPath; + try { + const issues = await verifyIndex(); + expect(issues.orphaned.length).toBe(1); + await expect(repairIndex(issues)).rejects.toThrow(/redaction/i); + } finally { + delete process.env.EPISODIC_MEMORY_REDACTION_RULES; + } + + const dbAfter = initDatabase(); + const row = dbAfter.prepare(`SELECT COUNT(*) as count FROM exchanges WHERE id = ?`).get('orphan-strict-1') as { count: number }; + expect(row.count).toBe(1); + dbAfter.close(); + }); + it('re-indexes outdated files during repair', { timeout: 30000 }, async () => { // Create conversation file with summary const projectArchive = path.join(archiveDir, 'test-project'); From 505996feba0ceb443a4efd2ffa1dc7f685eeb3ed Mon Sep 17 00:00:00 2001 From: Steven Molen <533340+darthmolen@users.noreply.github.com> Date: Mon, 5 Oct 2026 09:03:19 -0500 Subject: [PATCH 12/15] feat(redact): read-only dry run and --report to review hits `redact --rewrite --dry-run` opened the index with initDatabase(), which runs schema migrations and switches to WAL, so a dry run changed the database it promised not to touch (and created one if missing). A dry run now opens the existing index read-only, or skips the index if there is none. `--report` (with --dry-run) lists each value the rewrite would redact: file and line or index row, the rule, the value's shape (length, character classes, Shannon entropy) and the redacted text around it. The value itself is never printed. It uses an optional onMatch callback on RedactionContext, so it runs the real redaction path. Co-Authored-By: Claude Opus 5.5 --- CHANGELOG.md | 2 +- README.md | 1 + dist/db.d.ts | 6 ++ dist/db.js | 11 +++ dist/redact-cli.js | 19 +++- dist/redact-rewrite.d.ts | 18 +++- dist/redact-rewrite.js | 171 +++++++++++++++++++++++++----------- dist/redaction.d.ts | 17 ++++ dist/redaction.js | 38 ++++++-- docs/REDACTION.md | 24 ++++- src/db.ts | 11 +++ src/redact-cli.ts | 18 +++- src/redact-rewrite.ts | 127 ++++++++++++++++++++++---- src/redaction.ts | 56 ++++++++++-- test/redact-rewrite.test.ts | 47 +++++++++- 15 files changed, 472 insertions(+), 94 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index a467864f..223da791 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -14,7 +14,7 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0 ### Added -- `episodic-memory redact --rewrite` cleans data you indexed before upgrading. It redacts the archive and the search index in place, re-embeds only the messages that changed, and deletes summaries built from unredacted text so they regenerate. Use `--dry-run` to preview. +- `episodic-memory redact --rewrite` cleans data you indexed before upgrading. It redacts the archive and the search index in place, re-embeds only the messages that changed, and deletes summaries built from unredacted text so they regenerate. Use `--dry-run` to preview (it writes nothing, not even a schema migration), and add `--report` to list each value it would redact, by location, rule and shape, without printing the value. - Custom redaction rules via `~/.config/superpowers/redaction-rules.json`, which extends the bundled defaults. Try rules with `episodic-memory redact --stdin`. See `docs/REDACTION.md`. - New settings: `EPISODIC_MEMORY_REDACTION` (`on`/`off`), `EPISODIC_MEMORY_REDACTION_RULES`, and `EPISODIC_MEMORY_REDACTION_STRICT`. diff --git a/README.md b/README.md index c8105431..f59264ee 100644 --- a/README.md +++ b/README.md @@ -331,6 +331,7 @@ episodic-memory sync ```bash episodic-memory redact --rewrite --dry-run # what would change +episodic-memory redact --rewrite --dry-run --report # each hit: location, rule, value shape episodic-memory redact --rewrite # redact the existing archive + index in place episodic-memory redact --stdin < file.jsonl # try the rules on some text episodic-memory redact --print-default-rules diff --git a/dist/db.d.ts b/dist/db.d.ts index 3f590fff..9d8ed98a 100644 --- a/dist/db.d.ts +++ b/dist/db.d.ts @@ -13,6 +13,12 @@ export declare function migrateSchema(db: Database.Database): void; * 3. Recreates the table with ON DELETE CASCADE and copies surviving rows. */ export declare function migrateToolCallsCascade(db: Database.Database): void; +/** + * Open the existing index read-only, without creating it, migrating it, or + * changing its journal mode. For dry runs that must not write. Returns null + * when there is no index yet. + */ +export declare function openDatabaseReadOnly(): Database.Database | null; export declare function initDatabase(): Database.Database; export declare function insertExchange(db: Database.Database, exchange: ConversationExchange, embedding: number[], toolNames?: string[]): void; export declare function getAllExchanges(db: Database.Database): Array<{ diff --git a/dist/db.js b/dist/db.js index 83f546ad..38fbcde0 100644 --- a/dist/db.js +++ b/dist/db.js @@ -90,6 +90,17 @@ export function migrateToolCallsCascade(db) { db.pragma('foreign_keys = ON'); console.log(' tool_calls migration complete.'); } +/** + * Open the existing index read-only, without creating it, migrating it, or + * changing its journal mode. For dry runs that must not write. Returns null + * when there is no index yet. + */ +export function openDatabaseReadOnly() { + const dbPath = getDbPath(); + if (!fs.existsSync(dbPath)) + return null; + return new Database(dbPath, { readonly: true, fileMustExist: true }); +} export function initDatabase() { const dbPath = getDbPath(); // Ensure directory exists diff --git a/dist/redact-cli.js b/dist/redact-cli.js index 97044bb2..ee142efa 100644 --- a/dist/redact-cli.js +++ b/dist/redact-cli.js @@ -5,7 +5,7 @@ import { getSyncLockPath } from './logging.js'; import { DEFAULT_REDACTION_CONFIG, FindingsTally, formatFindings, getRedactionSettings, loadRedactor, redactJsonlLine, } from './redaction.js'; const args = process.argv.slice(2); const HELP = ` -Usage: episodic-memory redact [--rewrite [--dry-run]] [--stdin] [--print-default-rules] +Usage: episodic-memory redact [--rewrite [--dry-run [--report]]] [--stdin] [--print-default-rules] Secret redaction for the conversation archive and index. @@ -16,6 +16,11 @@ OPTIONS: sync regenerates them). Run once after upgrading, and again after adding rules. --dry-run With --rewrite: report what would change, write nothing. + The index is opened read-only and is not migrated. + --report With --rewrite --dry-run: list every value that would be + redacted: where it is, the rule, its shape (length, + character classes, entropy) and the redacted text around + it. Use it to spot false positives before applying. --stdin Redact stdin to stdout, one JSONL/text line at a time, and print rule counts to stderr. Handy for testing rules. --print-default-rules Print the bundled rules as JSON (a starting point for @@ -27,7 +32,7 @@ ENVIRONMENT: EPISODIC_MEMORY_REDACTION_RULES custom rules file (default: /redaction-rules.json) EPISODIC_MEMORY_REDACTION_STRICT 1 (default) fails closed on a bad rules file | 0 passes through -Output names rule IDs and counts only; matched values are never printed. +Output names rule IDs, counts and value shapes only; matched values are never printed. `; function fail(message) { console.error(`episodic-memory: ${message}`); @@ -47,7 +52,9 @@ function requireRedactor() { } return redactor; } -async function runRewrite(dryRun) { +async function runRewrite(dryRun, report) { + if (report && !dryRun) + fail('--report needs --dry-run: review the hits, then apply without it.'); const redactor = requireRedactor(); // Share sync's single-instance lock so a background sync can't write // between our reads and renames. @@ -80,6 +87,10 @@ async function runRewrite(dryRun) { redactor, embed, dryRun, + report: report + ? hit => console.log(` ${hit.location} ${hit.ruleId} ${hit.shape} + ${hit.context}`) + : undefined, log: message => console.log(` ${message}`), }); console.log(`\n${dryRun ? 'Would redact' : 'Redacted'} across archive, staging, and index: ${formatFindings(result.findings)}`); @@ -112,7 +123,7 @@ async function main() { return; } if (args.includes('--rewrite')) { - await runRewrite(args.includes('--dry-run')); + await runRewrite(args.includes('--dry-run'), args.includes('--report')); return; } fail(`unknown option(s): ${args.join(' ')}. Try: episodic-memory redact --help`); diff --git a/dist/redact-rewrite.d.ts b/dist/redact-rewrite.d.ts index 0900d8e4..c067931d 100644 --- a/dist/redact-rewrite.d.ts +++ b/dist/redact-rewrite.d.ts @@ -15,7 +15,8 @@ import { type RedactionFinding, type Redactor } from './redaction.js'; * It was generated from unredacted text, and the next sync regenerates it * from the redacted archive. * - * Idempotent: a second run finds nothing to change. + * Idempotent: a second run finds nothing to change. A dry run opens the index + * read-only, so it doesn't create, migrate or otherwise touch it. */ export type EmbedFn = (user: string, assistant: string, toolNames?: string[]) => Promise; export interface RewriteOptions { @@ -24,6 +25,11 @@ export interface RewriteOptions { embed: EmbedFn; /** Count what would change without writing anything. */ dryRun?: boolean; + /** + * Called once per value that would be redacted, for reviewing hits before + * applying them. Never receives the value itself. + */ + report?: (hit: RedactionHit) => void; /** Plugin-owned staging dirs to redact in place (not indexed). */ stagingDirs?: string[]; log?: (message: string) => void; @@ -37,4 +43,14 @@ export interface RewriteResult { /** Rule IDs and counts only — never matched values. */ findings: RedactionFinding[]; } +/** One redacted value, described without revealing it. */ +export interface RedactionHit { + /** Archive-relative file and 1-based line, or `index:# `. */ + location: string; + ruleId: string; + /** Length, character classes and entropy (see describeShape). */ + shape: string; + /** Redacted text around the token, on one line. */ + context: string; +} export declare function rewriteArchive(options: RewriteOptions): Promise; diff --git a/dist/redact-rewrite.js b/dist/redact-rewrite.js index 7c0317f6..0c6f4678 100644 --- a/dist/redact-rewrite.js +++ b/dist/redact-rewrite.js @@ -1,9 +1,11 @@ import fs from 'fs'; import path from 'path'; -import { initDatabase } from './db.js'; +import { initDatabase, openDatabaseReadOnly } from './db.js'; import { recordReembedded } from './embedding-migration.js'; -import { copyFileRedacted, FindingsTally, redactJsonlLine, } from './redaction.js'; +import { copyFileRedacted, describeShape, findRedactionTokens, FindingsTally, redactJsonlLine, } from './redaction.js'; const SUMMARY_SUFFIX = '-summary.txt'; +const CONTEXT_BEFORE = 60; +const CONTEXT_AFTER = 20; const PAGE_SIZE = 500; function walk(dir) { const out = []; @@ -26,6 +28,62 @@ function walk(dir) { function summaryPathFor(jsonlPath) { return jsonlPath.replace(/\.jsonl$/, SUMMARY_SUFFIX); } +/** + * Pair each value a rule redacted with its token in the redacted text and + * report it. Tokens already in the original are skipped. Pairing is by rule and + * order, so the context is right per rule but approximate when one rule hits + * a line more than once around existing tokens. + */ +function reportHits(original, redacted, matches, location, report) { + const queues = new Map(); + for (const m of matches) { + const q = queues.get(m.ruleId); + if (q) + q.push(m.value); + else + queues.set(m.ruleId, [m.value]); + } + const preexisting = new Map(); + for (const t of findRedactionTokens(original)) + preexisting.set(t.ruleId, (preexisting.get(t.ruleId) ?? 0) + 1); + const oneLine = (text) => text.replace(/\s+/g, ' '); + for (const t of findRedactionTokens(redacted)) { + const skip = preexisting.get(t.ruleId) ?? 0; + if (skip > 0) { + preexisting.set(t.ruleId, skip - 1); + continue; + } + const value = queues.get(t.ruleId)?.shift(); + if (value === undefined) + continue; + const before = redacted.slice(Math.max(0, t.start - CONTEXT_BEFORE), t.start); + const after = redacted.slice(t.end, t.end + CONTEXT_AFTER); + report({ + location, + ruleId: t.ruleId, + shape: describeShape(value), + context: oneLine(`${t.start > CONTEXT_BEFORE ? '…' : ''}${before}${redacted.slice(t.start, t.end)}${after}`), + }); + } +} +/** Redact `text` with `run`, reporting each hit when `report` is set. */ +function redactReporting(text, ctx, location, report, run) { + if (!report) + return run(ctx); + const matches = []; + const out = run({ ...ctx, onMatch: (ruleId, value) => matches.push({ ruleId, value }) }); + if (matches.length > 0) + reportHits(text, out, matches, location, report); + return out; +} +/** Report every hit in a JSONL file, line by line. Writes nothing. */ +function reportFile(file, label, redactor, report) { + const lines = fs.readFileSync(file, 'utf-8').split('\n'); + const ctx = { source: 'rewrite', path: file }; + lines.forEach((line, i) => { + redactReporting(line, ctx, `${label}:${i + 1}`, report, c => redactJsonlLine(line, redactor, c)); + }); +} /** Redact one JSONL file in place. Returns the number of values redacted. */ function rewriteFileInPlace(file, redactor, dryRun, tally) { const fileTally = new FindingsTally(); @@ -70,75 +128,82 @@ export async function rewriteArchive(options) { if (rewriteFileInPlace(file, redactor, dryRun, tally) > 0) { result.filesRewritten++; staleSummaries.add(summaryPathFor(file)); + if (options.report) + reportFile(file, path.relative(archiveDir, file), redactor, options.report); } } log(`Archive: ${result.filesRewritten} of ${result.filesScanned} file(s) ${dryRun ? 'would be ' : ''}rewritten`); // 2. Staging dirs. for (const dir of options.stagingDirs ?? []) { for (const file of walk(dir).filter(f => f.endsWith('.jsonl'))) { - if (rewriteFileInPlace(file, redactor, dryRun, tally) > 0) + if (rewriteFileInPlace(file, redactor, dryRun, tally) > 0) { result.stagingFilesRewritten++; + if (options.report) + reportFile(file, `staging:${path.relative(dir, file)}`, redactor, options.report); + } } } if (options.stagingDirs?.length) { log(`Staging exports: ${result.stagingFilesRewritten} file(s) ${dryRun ? 'would be ' : ''}rewritten`); } // 3. Index rows. Page by rowid so writes between pages don't disturb the scan. - const db = initDatabase(); - try { - const page = db.prepare('SELECT rowid AS rid, id, user_message, assistant_message, archive_path FROM exchanges WHERE rowid > ? ORDER BY rowid LIMIT ?'); - const toolsFor = db.prepare('SELECT id, tool_name, tool_input, tool_result FROM tool_calls WHERE exchange_id = ? ORDER BY rowid'); - const updateExchange = db.prepare('UPDATE exchanges SET user_message = ?, assistant_message = ? WHERE id = ?'); - const updateTool = db.prepare('UPDATE tool_calls SET tool_input = ?, tool_result = ? WHERE id = ?'); - let lastRowid = 0; - for (;;) { - const rows = page.all(lastRowid, PAGE_SIZE); - if (rows.length === 0) - break; - lastRowid = rows[rows.length - 1].rid; - for (const row of rows) { - const rowTally = new FindingsTally(); - const ctx = { source: 'index', path: row.archive_path }; - const user = redactor.redact(row.user_message, ctx); - const assistant = redactor.redact(row.assistant_message, ctx); - rowTally.add(user.findings); - rowTally.add(assistant.findings); - const tools = toolsFor.all(row.id); - const toolUpdates = []; - for (const tool of tools) { - // tool_input is JSON text; redactJsonlLine keeps it valid JSON. - const input = tool.tool_input === null ? null : redactJsonlLine(tool.tool_input, redactor, ctx, rowTally); - let output = tool.tool_result; - if (output !== null) { - const r = redactor.redact(output, ctx); + // A dry run reads the index as it is; opening it normally would migrate it. + const db = dryRun ? openDatabaseReadOnly() : initDatabase(); + if (db) + try { + const page = db.prepare('SELECT rowid AS rid, id, user_message, assistant_message, archive_path FROM exchanges WHERE rowid > ? ORDER BY rowid LIMIT ?'); + const toolsFor = db.prepare('SELECT id, tool_name, tool_input, tool_result FROM tool_calls WHERE exchange_id = ? ORDER BY rowid'); + const updateExchange = db.prepare('UPDATE exchanges SET user_message = ?, assistant_message = ? WHERE id = ?'); + const updateTool = db.prepare('UPDATE tool_calls SET tool_input = ?, tool_result = ? WHERE id = ?'); + let lastRowid = 0; + for (;;) { + const rows = page.all(lastRowid, PAGE_SIZE); + if (rows.length === 0) + break; + lastRowid = rows[rows.length - 1].rid; + for (const row of rows) { + const rowTally = new FindingsTally(); + const ctx = { source: 'index', path: row.archive_path }; + const where = `index:${path.relative(archiveDir, row.archive_path)}#${row.id}`; + const redactText = (text, field) => redactReporting(text, ctx, `${where} ${field}`, options.report, c => { + const r = redactor.redact(text, c); rowTally.add(r.findings); - output = r.text; - } - if (input !== tool.tool_input || output !== tool.tool_result) { - toolUpdates.push({ id: tool.id, input, result: output }); + return r.text; + }); + const user = { text: redactText(row.user_message, 'user') }; + const assistant = { text: redactText(row.assistant_message, 'assistant') }; + const tools = toolsFor.all(row.id); + const toolUpdates = []; + for (const tool of tools) { + // tool_input is JSON text; redactJsonlLine keeps it valid JSON. + const toolInput = tool.tool_input; + const input = toolInput === null ? null : redactReporting(toolInput, ctx, `${where} ${tool.tool_name} input`, options.report, c => redactJsonlLine(toolInput, redactor, c, rowTally)); + const output = tool.tool_result === null ? null : redactText(tool.tool_result, `${tool.tool_name} result`); + if (input !== tool.tool_input || output !== tool.tool_result) { + toolUpdates.push({ id: tool.id, input, result: output }); + } } + if (rowTally.total === 0) + continue; + tally.add(rowTally.toArray()); + result.rowsUpdated++; + staleSummaries.add(summaryPathFor(row.archive_path)); + if (dryRun) + continue; + const toolNames = tools.length > 0 ? tools.map(t => t.tool_name) : undefined; + const embedding = await embed(user.text, assistant.text, toolNames); + db.transaction(() => { + updateExchange.run(user.text, assistant.text, row.id); + for (const t of toolUpdates) + updateTool.run(t.input, t.result, t.id); + recordReembedded(db, row.id, embedding); + })(); } - if (rowTally.total === 0) - continue; - tally.add(rowTally.toArray()); - result.rowsUpdated++; - staleSummaries.add(summaryPathFor(row.archive_path)); - if (dryRun) - continue; - const toolNames = tools.length > 0 ? tools.map(t => t.tool_name) : undefined; - const embedding = await embed(user.text, assistant.text, toolNames); - db.transaction(() => { - updateExchange.run(user.text, assistant.text, row.id); - for (const t of toolUpdates) - updateTool.run(t.input, t.result, t.id); - recordReembedded(db, row.id, embedding); - })(); } } - } - finally { - db.close(); - } + finally { + db.close(); + } log(`Index: ${result.rowsUpdated} exchange(s) ${dryRun ? 'would be ' : ''}redacted and re-embedded`); // 4. Summaries: stale ones, plus any summary that itself matches a rule. for (const file of archiveFiles.filter(f => f.endsWith(SUMMARY_SUFFIX))) { diff --git a/dist/redaction.d.ts b/dist/redaction.d.ts index 6e96ecd9..e6906172 100644 --- a/dist/redaction.d.ts +++ b/dist/redaction.d.ts @@ -91,6 +91,11 @@ export interface RedactionRulesFile { export interface RedactionContext { source: string; path: string; + /** + * Called with each value as it is redacted. For local review tooling + * (`redact --report`) only: the value is the secret itself, so never log it. + */ + onMatch?: (ruleId: string, value: string) => void; } export interface RedactionFinding { ruleId: string; @@ -135,6 +140,18 @@ export declare function loadRedactionConfig(file: RedactionRulesFile): Redaction * non-strict mode warns and returns null, so text passes through unredacted. */ export declare function loadRedactor(env?: NodeJS.ProcessEnv, warn?: (message: string) => void): Redactor | null; +/** + * Describe a value without revealing it: length, character classes and + * Shannon entropy (bits per char), e.g. "len=40 aA9- H=4.9". Lets someone + * reviewing hits tell a random key from a word or a placeholder. + */ +export declare function describeShape(value: string): string; +/** Each `[REDACTED:]` token in `text`, left to right. */ +export declare function findRedactionTokens(text: string): Array<{ + ruleId: string; + start: number; + end: number; +}>; export declare function createRedactor(config: RedactionConfig): Redactor; /** Aggregates findings by rule id. Never holds matched values. */ export declare class FindingsTally { diff --git a/dist/redaction.js b/dist/redaction.js index 4f728424..f3fd20bb 100644 --- a/dist/redaction.js +++ b/dist/redaction.js @@ -310,6 +310,27 @@ function redactAroundTokens(text, spans, start, end, token) { flush(end); return replaced ? out : null; } +/** + * Describe a value without revealing it: length, character classes and + * Shannon entropy (bits per char), e.g. "len=40 aA9- H=4.9". Lets someone + * reviewing hits tell a random key from a word or a placeholder. + */ +export function describeShape(value) { + const classes = (/[a-z]/.test(value) ? 'a' : '') + + (/[A-Z]/.test(value) ? 'A' : '') + + (/[0-9]/.test(value) ? '9' : '') + + (/[^A-Za-z0-9\s]/.test(value) ? '-' : '') + + (/\s/.test(value) ? '_' : ''); + return `len=${value.length} ${classes || '?'} H=${shannonEntropy(value).toFixed(1)}`; +} +/** Each `[REDACTED:]` token in `text`, left to right. */ +export function findRedactionTokens(text) { + const out = []; + for (const m of text.matchAll(TOKEN_PATTERN)) { + out.push({ ruleId: m[0].slice(TOKEN_PREFIX.length, -1), start: m.index, end: m.index + m[0].length }); + } + return out; +} function shannonEntropy(text) { const counts = new Map(); for (const ch of text) @@ -327,7 +348,7 @@ function shannonEntropy(text) { * redacted (so re-running is a no-op). A candidate is skipped when it fully * matches an allowlist pattern. */ -function applyRule(text, rule, isAllowed) { +function applyRule(text, rule, isAllowed, onMatch) { const spans = text.includes(TOKEN_PREFIX) ? tokenSpans(text) : null; const token = tokenFor(rule.id); let out = ''; @@ -351,18 +372,20 @@ function applyRule(text, rule, isAllowed) { const rewritten = redactAroundTokens(text, spans, start, end, token); if (rewritten === null) continue; + onMatch?.(rule.id, text.slice(start, end).replace(TOKEN_PATTERN, '')); out += text.slice(last, start) + rewritten; last = end; count++; continue; } + onMatch?.(rule.id, text.slice(start, end)); out += text.slice(last, start) + token; last = end; count++; } return count === 0 ? { text, count } : { text: out + text.slice(last), count }; } -function applyEntropy(text, spec, isAllowed) { +function applyEntropy(text, spec, isAllowed, onMatch) { const candidates = new RegExp(`[A-Za-z0-9+/=_.~-]{${spec.minLength},}`, 'g'); const spans = text.includes(TOKEN_PREFIX) ? tokenSpans(text) : null; const token = tokenFor(ENTROPY_RULE_ID); @@ -384,6 +407,7 @@ function applyEntropy(text, spec, isAllowed) { if (!spec.keywords.some(k => before.includes(k))) continue; } + onMatch?.(ENTROPY_RULE_ID, m[0]); out += text.slice(last, start) + token; last = end; count++; @@ -401,7 +425,7 @@ export function createRedactor(config) { const normalized = name.toLowerCase().replace(/[^a-z0-9]/g, '').replace(/\d+$/, ''); return normalized.length > 0 && compiled.secretField.test(normalized); }, - redact(text, _ctx) { + redact(text, ctx) { if (typeof text !== 'string' || text.length === 0) return { text, findings: [] }; let current = text; @@ -413,7 +437,7 @@ export function createRedactor(config) { if (!rule.keywords.some(k => lower.includes(k))) continue; } - const r = applyRule(current, rule, isAllowed); + const r = applyRule(current, rule, isAllowed, ctx?.onMatch); if (r.count > 0) { current = r.text; lower = null; @@ -421,7 +445,7 @@ export function createRedactor(config) { } } if (compiled.entropy.enabled) { - const r = applyEntropy(current, compiled.entropy, isAllowed); + const r = applyEntropy(current, compiled.entropy, isAllowed, ctx?.onMatch); if (r.count > 0) { current = r.text; counts.set(ENTROPY_RULE_ID, r.count); @@ -505,8 +529,10 @@ function redactTree(node, redactor, ctx, tally) { // MCP result. Runs after the text rules; a value they redacted only in part // is still replaced whole, since the field name says all of it is secret. const redactWhole = (key) => { - if (!isRedactableWhole(obj[key])) + const value = obj[key]; + if (!isRedactableWhole(value)) return; + ctx?.onMatch?.(FIELD_RULE_ID, value); setOwn(obj, key, tokenFor(FIELD_RULE_ID)); tally?.add([{ ruleId: FIELD_RULE_ID, count: 1 }]); changed = true; diff --git a/docs/REDACTION.md b/docs/REDACTION.md index 3e8a796b..67bb15e7 100644 --- a/docs/REDACTION.md +++ b/docs/REDACTION.md @@ -152,10 +152,30 @@ New syncs only redact new or changed files. To redact everything indexed before you upgraded, or after you add a rule: ```bash -episodic-memory redact --rewrite --dry-run # report what would change -episodic-memory redact --rewrite # apply +episodic-memory redact --rewrite --dry-run # report what would change +episodic-memory redact --rewrite --dry-run --report # list each hit to review +episodic-memory redact --rewrite # apply ``` +A dry run writes nothing, not even a schema migration: it opens the index +read-only. + +`--report` lists every value the rewrite would redact, so you can check for +false positives before applying. Each hit shows where it is, the rule, the +value's shape, and the redacted text around it. The value itself is never +printed: + +``` + -work-contoso/4f1c….jsonl:212 quoted-secret-assignment len=40 aA9- H=4.9 + …"AzureAd": { "ClientId": "…", "ClientSecret": "[REDACTED:quoted-secret-assignment]", "TenantId… +``` + +The shape is the length, the character classes (`a` lowercase, `A` uppercase, +`9` digits, `-` symbols, `_` whitespace) and the Shannon entropy in bits per +character. A random key is long with high entropy (about 4.5 or more); a word or +a placeholder is short or low. If a rule fires on something that isn't a secret, +turn it off or narrow it in `redaction-rules.json`. + `--rewrite`: - redacts every archive file in place, keeping line numbers and timestamps diff --git a/src/db.ts b/src/db.ts index 8c7a73ad..a3811e2c 100644 --- a/src/db.ts +++ b/src/db.ts @@ -104,6 +104,17 @@ export function migrateToolCallsCascade(db: Database.Database): void { console.log(' tool_calls migration complete.'); } +/** + * Open the existing index read-only, without creating it, migrating it, or + * changing its journal mode. For dry runs that must not write. Returns null + * when there is no index yet. + */ +export function openDatabaseReadOnly(): Database.Database | null { + const dbPath = getDbPath(); + if (!fs.existsSync(dbPath)) return null; + return new Database(dbPath, { readonly: true, fileMustExist: true }); +} + export function initDatabase(): Database.Database { const dbPath = getDbPath(); diff --git a/src/redact-cli.ts b/src/redact-cli.ts index 6e75b0a6..01c33ea4 100644 --- a/src/redact-cli.ts +++ b/src/redact-cli.ts @@ -15,7 +15,7 @@ import { const args = process.argv.slice(2); const HELP = ` -Usage: episodic-memory redact [--rewrite [--dry-run]] [--stdin] [--print-default-rules] +Usage: episodic-memory redact [--rewrite [--dry-run [--report]]] [--stdin] [--print-default-rules] Secret redaction for the conversation archive and index. @@ -26,6 +26,11 @@ OPTIONS: sync regenerates them). Run once after upgrading, and again after adding rules. --dry-run With --rewrite: report what would change, write nothing. + The index is opened read-only and is not migrated. + --report With --rewrite --dry-run: list every value that would be + redacted: where it is, the rule, its shape (length, + character classes, entropy) and the redacted text around + it. Use it to spot false positives before applying. --stdin Redact stdin to stdout, one JSONL/text line at a time, and print rule counts to stderr. Handy for testing rules. --print-default-rules Print the bundled rules as JSON (a starting point for @@ -37,7 +42,7 @@ ENVIRONMENT: EPISODIC_MEMORY_REDACTION_RULES custom rules file (default: /redaction-rules.json) EPISODIC_MEMORY_REDACTION_STRICT 1 (default) fails closed on a bad rules file | 0 passes through -Output names rule IDs and counts only; matched values are never printed. +Output names rule IDs, counts and value shapes only; matched values are never printed. `; function fail(message: string): never { @@ -59,7 +64,8 @@ function requireRedactor(): Redactor { return redactor!; } -async function runRewrite(dryRun: boolean): Promise { +async function runRewrite(dryRun: boolean, report: boolean): Promise { + if (report && !dryRun) fail('--report needs --dry-run: review the hits, then apply without it.'); const redactor = requireRedactor(); // Share sync's single-instance lock so a background sync can't write @@ -95,6 +101,10 @@ async function runRewrite(dryRun: boolean): Promise { redactor, embed, dryRun, + report: report + ? hit => console.log(` ${hit.location} ${hit.ruleId} ${hit.shape} + ${hit.context}`) + : undefined, log: message => console.log(` ${message}`), }); @@ -129,7 +139,7 @@ async function main(): Promise { return; } if (args.includes('--rewrite')) { - await runRewrite(args.includes('--dry-run')); + await runRewrite(args.includes('--dry-run'), args.includes('--report')); return; } fail(`unknown option(s): ${args.join(' ')}. Try: episodic-memory redact --help`); diff --git a/src/redact-rewrite.ts b/src/redact-rewrite.ts index 390f74b0..4ef696b7 100644 --- a/src/redact-rewrite.ts +++ b/src/redact-rewrite.ts @@ -1,11 +1,14 @@ import fs from 'fs'; import path from 'path'; -import { initDatabase } from './db.js'; +import { initDatabase, openDatabaseReadOnly } from './db.js'; import { recordReembedded } from './embedding-migration.js'; import { copyFileRedacted, + describeShape, + findRedactionTokens, FindingsTally, redactJsonlLine, + type RedactionContext, type RedactionFinding, type Redactor, } from './redaction.js'; @@ -26,7 +29,8 @@ import { * It was generated from unredacted text, and the next sync regenerates it * from the redacted archive. * - * Idempotent: a second run finds nothing to change. + * Idempotent: a second run finds nothing to change. A dry run opens the index + * read-only, so it doesn't create, migrate or otherwise touch it. */ export type EmbedFn = (user: string, assistant: string, toolNames?: string[]) => Promise; @@ -37,6 +41,11 @@ export interface RewriteOptions { embed: EmbedFn; /** Count what would change without writing anything. */ dryRun?: boolean; + /** + * Called once per value that would be redacted, for reviewing hits before + * applying them. Never receives the value itself. + */ + report?: (hit: RedactionHit) => void; /** Plugin-owned staging dirs to redact in place (not indexed). */ stagingDirs?: string[]; log?: (message: string) => void; @@ -52,7 +61,20 @@ export interface RewriteResult { findings: RedactionFinding[]; } +/** One redacted value, described without revealing it. */ +export interface RedactionHit { + /** Archive-relative file and 1-based line, or `index:# `. */ + location: string; + ruleId: string; + /** Length, character classes and entropy (see describeShape). */ + shape: string; + /** Redacted text around the token, on one line. */ + context: string; +} + const SUMMARY_SUFFIX = '-summary.txt'; +const CONTEXT_BEFORE = 60; +const CONTEXT_AFTER = 20; const PAGE_SIZE = 500; function walk(dir: string): string[] { @@ -75,6 +97,70 @@ function summaryPathFor(jsonlPath: string): string { return jsonlPath.replace(/\.jsonl$/, SUMMARY_SUFFIX); } +type Match = { ruleId: string; value: string }; + +/** + * Pair each value a rule redacted with its token in the redacted text and + * report it. Tokens already in the original are skipped. Pairing is by rule and + * order, so the context is right per rule but approximate when one rule hits + * a line more than once around existing tokens. + */ +function reportHits( + original: string, + redacted: string, + matches: Match[], + location: string, + report: (hit: RedactionHit) => void +): void { + const queues = new Map(); + for (const m of matches) { + const q = queues.get(m.ruleId); + if (q) q.push(m.value); else queues.set(m.ruleId, [m.value]); + } + const preexisting = new Map(); + for (const t of findRedactionTokens(original)) preexisting.set(t.ruleId, (preexisting.get(t.ruleId) ?? 0) + 1); + + const oneLine = (text: string) => text.replace(/\s+/g, ' '); + for (const t of findRedactionTokens(redacted)) { + const skip = preexisting.get(t.ruleId) ?? 0; + if (skip > 0) { preexisting.set(t.ruleId, skip - 1); continue; } + const value = queues.get(t.ruleId)?.shift(); + if (value === undefined) continue; + const before = redacted.slice(Math.max(0, t.start - CONTEXT_BEFORE), t.start); + const after = redacted.slice(t.end, t.end + CONTEXT_AFTER); + report({ + location, + ruleId: t.ruleId, + shape: describeShape(value), + context: oneLine(`${t.start > CONTEXT_BEFORE ? '…' : ''}${before}${redacted.slice(t.start, t.end)}${after}`), + }); + } +} + +/** Redact `text` with `run`, reporting each hit when `report` is set. */ +function redactReporting( + text: string, + ctx: RedactionContext, + location: string, + report: ((hit: RedactionHit) => void) | undefined, + run: (ctx: RedactionContext) => string +): string { + if (!report) return run(ctx); + const matches: Match[] = []; + const out = run({ ...ctx, onMatch: (ruleId, value) => matches.push({ ruleId, value }) }); + if (matches.length > 0) reportHits(text, out, matches, location, report); + return out; +} + +/** Report every hit in a JSONL file, line by line. Writes nothing. */ +function reportFile(file: string, label: string, redactor: Redactor, report: (hit: RedactionHit) => void): void { + const lines = fs.readFileSync(file, 'utf-8').split('\n'); + const ctx = { source: 'rewrite', path: file }; + lines.forEach((line, i) => { + redactReporting(line, ctx, `${label}:${i + 1}`, report, c => redactJsonlLine(line, redactor, c)); + }); +} + /** Redact one JSONL file in place. Returns the number of values redacted. */ function rewriteFileInPlace(file: string, redactor: Redactor, dryRun: boolean, tally: FindingsTally): number { const fileTally = new FindingsTally(); @@ -117,6 +203,7 @@ export async function rewriteArchive(options: RewriteOptions): Promise 0) { result.filesRewritten++; staleSummaries.add(summaryPathFor(file)); + if (options.report) reportFile(file, path.relative(archiveDir, file), redactor, options.report); } } log(`Archive: ${result.filesRewritten} of ${result.filesScanned} file(s) ${dryRun ? 'would be ' : ''}rewritten`); @@ -124,7 +211,10 @@ export async function rewriteArchive(options: RewriteOptions): Promise f.endsWith('.jsonl'))) { - if (rewriteFileInPlace(file, redactor, dryRun, tally) > 0) result.stagingFilesRewritten++; + if (rewriteFileInPlace(file, redactor, dryRun, tally) > 0) { + result.stagingFilesRewritten++; + if (options.report) reportFile(file, `staging:${path.relative(dir, file)}`, redactor, options.report); + } } } if (options.stagingDirs?.length) { @@ -132,8 +222,9 @@ export async function rewriteArchive(options: RewriteOptions): Promise ? ORDER BY rowid LIMIT ?' ); @@ -152,22 +243,26 @@ export async function rewriteArchive(options: RewriteOptions): Promise + redactReporting(text, ctx, `${where} ${field}`, options.report, c => { + const r = redactor.redact(text, c); + rowTally.add(r.findings); + return r.text; + }); + const user = { text: redactText(row.user_message, 'user') }; + const assistant = { text: redactText(row.assistant_message, 'assistant') }; const tools = toolsFor.all(row.id) as Array<{ id: string; tool_name: string; tool_input: string | null; tool_result: string | null }>; const toolUpdates: Array<{ id: string; input: string | null; result: string | null }> = []; for (const tool of tools) { // tool_input is JSON text; redactJsonlLine keeps it valid JSON. - const input = tool.tool_input === null ? null : redactJsonlLine(tool.tool_input, redactor, ctx, rowTally); - let output = tool.tool_result; - if (output !== null) { - const r = redactor.redact(output, ctx); - rowTally.add(r.findings); - output = r.text; - } + const toolInput = tool.tool_input; + const input = toolInput === null ? null : redactReporting( + toolInput, ctx, `${where} ${tool.tool_name} input`, options.report, + c => redactJsonlLine(toolInput, redactor, c, rowTally) + ); + const output = tool.tool_result === null ? null : redactText(tool.tool_result, `${tool.tool_name} result`); if (input !== tool.tool_input || output !== tool.tool_result) { toolUpdates.push({ id: tool.id, input, result: output }); } diff --git a/src/redaction.ts b/src/redaction.ts index 0a6563bb..52b815d8 100644 --- a/src/redaction.ts +++ b/src/redaction.ts @@ -119,6 +119,11 @@ export interface RedactionRulesFile { export interface RedactionContext { source: string; path: string; + /** + * Called with each value as it is redacted. For local review tooling + * (`redact --report`) only: the value is the secret itself, so never log it. + */ + onMatch?: (ruleId: string, value: string) => void; } export interface RedactionFinding { @@ -465,6 +470,30 @@ function redactAroundTokens( return replaced ? out : null; } +/** + * Describe a value without revealing it: length, character classes and + * Shannon entropy (bits per char), e.g. "len=40 aA9- H=4.9". Lets someone + * reviewing hits tell a random key from a word or a placeholder. + */ +export function describeShape(value: string): string { + const classes = + (/[a-z]/.test(value) ? 'a' : '') + + (/[A-Z]/.test(value) ? 'A' : '') + + (/[0-9]/.test(value) ? '9' : '') + + (/[^A-Za-z0-9\s]/.test(value) ? '-' : '') + + (/\s/.test(value) ? '_' : ''); + return `len=${value.length} ${classes || '?'} H=${shannonEntropy(value).toFixed(1)}`; +} + +/** Each `[REDACTED:]` token in `text`, left to right. */ +export function findRedactionTokens(text: string): Array<{ ruleId: string; start: number; end: number }> { + const out: Array<{ ruleId: string; start: number; end: number }> = []; + for (const m of text.matchAll(TOKEN_PATTERN)) { + out.push({ ruleId: m[0].slice(TOKEN_PREFIX.length, -1), start: m.index!, end: m.index! + m[0].length }); + } + return out; +} + function shannonEntropy(text: string): number { const counts = new Map(); for (const ch of text) counts.set(ch, (counts.get(ch) ?? 0) + 1); @@ -482,7 +511,12 @@ function shannonEntropy(text: string): number { * redacted (so re-running is a no-op). A candidate is skipped when it fully * matches an allowlist pattern. */ -function applyRule(text: string, rule: CompiledRule, isAllowed: (s: string) => boolean): { text: string; count: number } { +function applyRule( + text: string, + rule: CompiledRule, + isAllowed: (s: string) => boolean, + onMatch?: RedactionContext['onMatch'] +): { text: string; count: number } { const spans = text.includes(TOKEN_PREFIX) ? tokenSpans(text) : null; const token = tokenFor(rule.id); let out = ''; @@ -501,11 +535,13 @@ function applyRule(text: string, rule: CompiledRule, isAllowed: (s: string) => b if (spans && overlapsAny(spans, start, end)) { const rewritten = redactAroundTokens(text, spans, start, end, token); if (rewritten === null) continue; + onMatch?.(rule.id, text.slice(start, end).replace(TOKEN_PATTERN, '')); out += text.slice(last, start) + rewritten; last = end; count++; continue; } + onMatch?.(rule.id, text.slice(start, end)); out += text.slice(last, start) + token; last = end; count++; @@ -513,7 +549,12 @@ function applyRule(text: string, rule: CompiledRule, isAllowed: (s: string) => b return count === 0 ? { text, count } : { text: out + text.slice(last), count }; } -function applyEntropy(text: string, spec: EntropySpec, isAllowed: (s: string) => boolean): { text: string; count: number } { +function applyEntropy( + text: string, + spec: EntropySpec, + isAllowed: (s: string) => boolean, + onMatch?: RedactionContext['onMatch'] +): { text: string; count: number } { const candidates = new RegExp(`[A-Za-z0-9+/=_.~-]{${spec.minLength},}`, 'g'); const spans = text.includes(TOKEN_PREFIX) ? tokenSpans(text) : null; const token = tokenFor(ENTROPY_RULE_ID); @@ -531,6 +572,7 @@ function applyEntropy(text: string, spec: EntropySpec, isAllowed: (s: string) => before = before.slice(before.lastIndexOf('\n') + 1).toLowerCase(); if (!spec.keywords.some(k => before.includes(k))) continue; } + onMatch?.(ENTROPY_RULE_ID, m[0]); out += text.slice(last, start) + token; last = end; count++; @@ -549,7 +591,7 @@ export function createRedactor(config: RedactionConfig): Redactor { const normalized = name.toLowerCase().replace(/[^a-z0-9]/g, '').replace(/\d+$/, ''); return normalized.length > 0 && compiled.secretField.test(normalized); }, - redact(text: string, _ctx?: RedactionContext): RedactionResult { + redact(text: string, ctx?: RedactionContext): RedactionResult { if (typeof text !== 'string' || text.length === 0) return { text, findings: [] }; let current = text; let lower: string | null = null; @@ -560,7 +602,7 @@ export function createRedactor(config: RedactionConfig): Redactor { lower ??= current.toLowerCase(); if (!rule.keywords.some(k => lower!.includes(k))) continue; } - const r = applyRule(current, rule, isAllowed); + const r = applyRule(current, rule, isAllowed, ctx?.onMatch); if (r.count > 0) { current = r.text; lower = null; @@ -569,7 +611,7 @@ export function createRedactor(config: RedactionConfig): Redactor { } if (compiled.entropy.enabled) { - const r = applyEntropy(current, compiled.entropy, isAllowed); + const r = applyEntropy(current, compiled.entropy, isAllowed, ctx?.onMatch); if (r.count > 0) { current = r.text; counts.set(ENTROPY_RULE_ID, r.count); @@ -661,7 +703,9 @@ function redactTree( // MCP result. Runs after the text rules; a value they redacted only in part // is still replaced whole, since the field name says all of it is secret. const redactWhole = (key: string) => { - if (!isRedactableWhole(obj[key])) return; + const value = obj[key]; + if (!isRedactableWhole(value)) return; + ctx?.onMatch?.(FIELD_RULE_ID, value); setOwn(obj, key, tokenFor(FIELD_RULE_ID)); tally?.add([{ ruleId: FIELD_RULE_ID, count: 1 }]); changed = true; diff --git a/test/redact-rewrite.test.ts b/test/redact-rewrite.test.ts index 10c63f9a..12a05bc8 100644 --- a/test/redact-rewrite.test.ts +++ b/test/redact-rewrite.test.ts @@ -1,6 +1,6 @@ import { describe, it, expect, beforeEach, afterEach, vi } from 'vitest'; import { mkdtempSync, mkdirSync, writeFileSync, readFileSync, rmSync, existsSync, statSync, utimesSync } from 'fs'; -import { join } from 'path'; +import { join, sep } from 'path'; import { tmpdir } from 'os'; import Database from 'better-sqlite3'; import * as sqliteVec from 'sqlite-vec'; @@ -164,6 +164,51 @@ describe('redact --rewrite backfill', () => { expect(dbDump()).toContain(secret); expect(embed).not.toHaveBeenCalled(); }); + + it('--dry-run leaves an older-schema index byte-identical instead of migrating it', async () => { + await seedLegacy(new FakeSecrets(910)); + const legacy = new Database(dbPath); + legacy.exec('ALTER TABLE exchanges DROP COLUMN embedding_version'); + legacy.close(); + const before = readFileSync(dbPath); + + const result = await rewriteArchive({ archiveDir, redactor: createRedactor(DEFAULT_REDACTION_CONFIG), embed, dryRun: true }); + + expect(result.rowsUpdated).toBeGreaterThan(0); + expect(readFileSync(dbPath).equals(before)).toBe(true); + const db = new Database(dbPath, { readonly: true }); + const columns = (db.prepare('PRAGMA table_info(exchanges)').all() as Array<{ name: string }>).map(c => c.name); + db.close(); + expect(columns).not.toContain('embedding_version'); + }); + + it('--dry-run does not create an index that does not exist', async () => { + mkdirSync(join(archiveDir, '-work-legacy'), { recursive: true }); + const result = await rewriteArchive({ archiveDir, redactor: createRedactor(DEFAULT_REDACTION_CONFIG), embed, dryRun: true }); + expect(result.rowsUpdated).toBe(0); + expect(existsSync(dbPath)).toBe(false); + }); + + it('report lists one hit per redacted value, with location, rule and shape but never the value', async () => { + const { secret, password } = await seedLegacy(new FakeSecrets(911)); + const hits: Array<{ location: string; ruleId: string; shape: string; context: string }> = []; + + const result = await rewriteArchive({ + archiveDir, redactor: createRedactor(DEFAULT_REDACTION_CONFIG), embed, dryRun: true, report: hit => hits.push(hit), + }); + + const total = result.findings.reduce((n, f) => n + f.count, 0); + expect(hits.length).toBe(total); + const text = JSON.stringify(hits); + expect(text).not.toContain(secret); + expect(text).not.toContain(password); + + const archiveHit = hits.find(h => h.location === `-work-legacy${sep}${SESSION}.jsonl:1`); + expect(archiveHit?.ruleId).toBe('connection-string-secret'); + expect(archiveHit!.shape).toMatch(new RegExp(`^len=${secret.length} \\S+ H=\\d+\\.\\d$`)); + expect(archiveHit!.context).toContain('AccountKey=[REDACTED:connection-string-secret]'); + expect(hits.some(h => h.location.startsWith('index:') && h.location.includes(' user'))).toBe(true); + }); }); function expectNoSecretInCalls(calls: unknown[][], secrets: string[]) { From c46a8edcb0392b01f8fbe83f58d921c2a6f81310 Mon Sep 17 00:00:00 2001 From: Steven Molen <533340+darthmolen@users.noreply.github.com> Date: Mon, 5 Oct 2026 09:12:37 -0500 Subject: [PATCH 13/15] fix(redact): pair report hits with their tokens exactly The report matched values to tokens by rule and order, skipping as many tokens as the original line already had. A new hit before an existing token of the same rule took that token's context. A secret field redacted whole after a text rule hit part of it lost the text rule's hit, and its shape described the partly redacted value. onMatch may now return the token to write. In report mode each hit gets a numbered one ([REDACTED:--hit]), so it is found exactly; unnumbered tokens were already there. A field value is restored from the numbered tokens inside it, and hits it swallowed are reported at its token. rewriteArchive also refuses report without dryRun: after a real rewrite the files would have nothing left to report. Co-Authored-By: Claude Opus 5.5 --- dist/redact-rewrite.js | 104 +++++++++++++++++++++++------------- dist/redaction.d.ts | 4 +- dist/redaction.js | 23 ++++---- src/redact-rewrite.ts | 96 ++++++++++++++++++++++----------- src/redaction.ts | 32 ++++++----- test/redact-rewrite.test.ts | 43 +++++++++++++++ 6 files changed, 208 insertions(+), 94 deletions(-) diff --git a/dist/redact-rewrite.js b/dist/redact-rewrite.js index 0c6f4678..ae62e1d9 100644 --- a/dist/redact-rewrite.js +++ b/dist/redact-rewrite.js @@ -28,60 +28,85 @@ function walk(dir) { function summaryPathFor(jsonlPath) { return jsonlPath.replace(/\.jsonl$/, SUMMARY_SUFFIX); } +// In report mode each hit is written as its own numbered token, e.g. +// `[REDACTED:jwt--hit3]`, so it can be found in the output exactly; tokens that +// were already in the text have no number. Numbers are stripped for display. +const HIT_TOKEN = /\[REDACTED:([a-z0-9][a-z0-9-]*?)--hit(\d+)\]/g; +/** `value` with any numbered tokens in it replaced by the values they stand for. */ +function restoreHits(value, matches) { + return value.replace(HIT_TOKEN, (_token, _ruleId, n) => matches[Number(n)].value); +} /** - * Pair each value a rule redacted with its token in the redacted text and - * report it. Tokens already in the original are skipped. Pairing is by rule and - * order, so the context is right per rule but approximate when one rule hits - * a line more than once around existing tokens. + * Report each hit with the redacted text around its token. A hit whose token + * was swallowed by a later one (a secret field redacted whole after a text rule + * hit part of it) is reported at the token that swallowed it. */ -function reportHits(original, redacted, matches, location, report) { - const queues = new Map(); - for (const m of matches) { - const q = queues.get(m.ruleId); - if (q) - q.push(m.value); - else - queues.set(m.ruleId, [m.value]); - } - const preexisting = new Map(); - for (const t of findRedactionTokens(original)) - preexisting.set(t.ruleId, (preexisting.get(t.ruleId) ?? 0) + 1); - const oneLine = (text) => text.replace(/\s+/g, ' '); +function reportHits(redacted, matches, location, report) { + let display = ''; + let pos = 0; + const spans = new Map(); for (const t of findRedactionTokens(redacted)) { - const skip = preexisting.get(t.ruleId) ?? 0; - if (skip > 0) { - preexisting.set(t.ruleId, skip - 1); - continue; + const numbered = /^(.*)--hit(\d+)$/.exec(t.ruleId); + display += redacted.slice(pos, t.start); + const shown = numbered ? `[REDACTED:${numbered[1]}]` : redacted.slice(t.start, t.end); + if (numbered && !spans.has(Number(numbered[2]))) { + spans.set(Number(numbered[2]), [display.length, display.length + shown.length]); } - const value = queues.get(t.ruleId)?.shift(); - if (value === undefined) - continue; - const before = redacted.slice(Math.max(0, t.start - CONTEXT_BEFORE), t.start); - const after = redacted.slice(t.end, t.end + CONTEXT_AFTER); + display += shown; + pos = t.end; + } + display += redacted.slice(pos); + const spanOf = (n) => { + for (let i = n; i !== undefined; i = matches[i].absorbedBy) { + const span = spans.get(i); + if (span) + return span; + } + return undefined; + }; + const oneLine = (text) => text.replace(/\s+/g, ' '); + const hits = matches.map((m, n) => ({ m, span: spanOf(n) })); + hits.sort((a, b) => (a.span?.[0] ?? Infinity) - (b.span?.[0] ?? Infinity)); + for (const { m, span } of hits) { + const [start, end] = span ?? [0, 0]; + const before = display.slice(Math.max(0, start - CONTEXT_BEFORE), start); report({ location, - ruleId: t.ruleId, - shape: describeShape(value), - context: oneLine(`${t.start > CONTEXT_BEFORE ? '…' : ''}${before}${redacted.slice(t.start, t.end)}${after}`), + ruleId: m.ruleId, + shape: describeShape(m.value), + context: span + ? oneLine(`${start > CONTEXT_BEFORE ? '…' : ''}${before}${display.slice(start, end + CONTEXT_AFTER)}`) + : '', }); } + return display; } -/** Redact `text` with `run`, reporting each hit when `report` is set. */ -function redactReporting(text, ctx, location, report, run) { +/** + * Redact with `run`, reporting each hit when `report` is set. Returns + * the redacted text with ordinary tokens either way. + */ +function redactReporting(ctx, location, report, run) { if (!report) return run(ctx); const matches = []; - const out = run({ ...ctx, onMatch: (ruleId, value) => matches.push({ ruleId, value }) }); - if (matches.length > 0) - reportHits(text, out, matches, location, report); - return out; + const out = run({ + ...ctx, + onMatch: (ruleId, value) => { + const n = matches.length; + for (const [, , k] of value.matchAll(HIT_TOKEN)) + matches[Number(k)].absorbedBy = n; + matches.push({ ruleId, value: restoreHits(value, matches) }); + return `[REDACTED:${ruleId}--hit${n}]`; + }, + }); + return matches.length > 0 ? reportHits(out, matches, location, report) : out; } /** Report every hit in a JSONL file, line by line. Writes nothing. */ function reportFile(file, label, redactor, report) { const lines = fs.readFileSync(file, 'utf-8').split('\n'); const ctx = { source: 'rewrite', path: file }; lines.forEach((line, i) => { - redactReporting(line, ctx, `${label}:${i + 1}`, report, c => redactJsonlLine(line, redactor, c)); + redactReporting(ctx, `${label}:${i + 1}`, report, c => redactJsonlLine(line, redactor, c)); }); } /** Redact one JSONL file in place. Returns the number of values redacted. */ @@ -109,6 +134,9 @@ function rewriteFileInPlace(file, redactor, dryRun, tally) { export async function rewriteArchive(options) { const { archiveDir, redactor, embed } = options; const dryRun = options.dryRun === true; + // Reporting reads files before redaction; after a real rewrite there'd be nothing to find. + if (options.report && !dryRun) + throw new Error('rewriteArchive: report requires dryRun'); const log = options.log ?? (() => { }); const tally = new FindingsTally(); const result = { @@ -165,7 +193,7 @@ export async function rewriteArchive(options) { const rowTally = new FindingsTally(); const ctx = { source: 'index', path: row.archive_path }; const where = `index:${path.relative(archiveDir, row.archive_path)}#${row.id}`; - const redactText = (text, field) => redactReporting(text, ctx, `${where} ${field}`, options.report, c => { + const redactText = (text, field) => redactReporting(ctx, `${where} ${field}`, options.report, c => { const r = redactor.redact(text, c); rowTally.add(r.findings); return r.text; @@ -177,7 +205,7 @@ export async function rewriteArchive(options) { for (const tool of tools) { // tool_input is JSON text; redactJsonlLine keeps it valid JSON. const toolInput = tool.tool_input; - const input = toolInput === null ? null : redactReporting(toolInput, ctx, `${where} ${tool.tool_name} input`, options.report, c => redactJsonlLine(toolInput, redactor, c, rowTally)); + const input = toolInput === null ? null : redactReporting(ctx, `${where} ${tool.tool_name} input`, options.report, c => redactJsonlLine(toolInput, redactor, c, rowTally)); const output = tool.tool_result === null ? null : redactText(tool.tool_result, `${tool.tool_name} result`); if (input !== tool.tool_input || output !== tool.tool_result) { toolUpdates.push({ id: tool.id, input, result: output }); diff --git a/dist/redaction.d.ts b/dist/redaction.d.ts index e6906172..b19ee59b 100644 --- a/dist/redaction.d.ts +++ b/dist/redaction.d.ts @@ -94,8 +94,10 @@ export interface RedactionContext { /** * Called with each value as it is redacted. For local review tooling * (`redact --report`) only: the value is the secret itself, so never log it. + * May return the token to write instead of `[REDACTED:]`, e.g. one + * numbered per hit; it must still have the token's shape. */ - onMatch?: (ruleId: string, value: string) => void; + onMatch?: (ruleId: string, value: string) => string | void; } export interface RedactionFinding { ruleId: string; diff --git a/dist/redaction.js b/dist/redaction.js index f3fd20bb..f654ae08 100644 --- a/dist/redaction.js +++ b/dist/redaction.js @@ -267,6 +267,12 @@ function compileConfig(config) { function tokenFor(ruleId) { return `${TOKEN_PREFIX}${ruleId}]`; } +const WHOLE_TOKEN = /^\[REDACTED:[a-z0-9][a-z0-9-]*\]$/; +/** The token for one redacted value: onMatch's, if it returns a valid one. */ +function tokenForMatch(ruleId, value, onMatch) { + const custom = onMatch?.(ruleId, value); + return typeof custom === 'string' && WHOLE_TOKEN.test(custom) ? custom : tokenFor(ruleId); +} function tokenSpans(text) { const spans = []; for (const m of text.matchAll(TOKEN_PATTERN)) @@ -291,7 +297,7 @@ function redactAroundTokens(text, spans, start, end, token) { const flush = (to) => { const piece = text.slice(pos, to); if (/[A-Za-z0-9]/.test(piece)) { - out += token; + out += token(); replaced = true; } else { @@ -350,7 +356,6 @@ function shannonEntropy(text) { */ function applyRule(text, rule, isAllowed, onMatch) { const spans = text.includes(TOKEN_PREFIX) ? tokenSpans(text) : null; - const token = tokenFor(rule.id); let out = ''; let last = 0; let count = 0; @@ -369,17 +374,16 @@ function applyRule(text, rule, isAllowed, onMatch) { if (rule.useAllowlist && isAllowed(text.slice(start, end))) continue; if (spans && overlapsAny(spans, start, end)) { - const rewritten = redactAroundTokens(text, spans, start, end, token); + let token; + const rewritten = redactAroundTokens(text, spans, start, end, () => (token ??= tokenForMatch(rule.id, text.slice(start, end).replace(TOKEN_PATTERN, ''), onMatch))); if (rewritten === null) continue; - onMatch?.(rule.id, text.slice(start, end).replace(TOKEN_PATTERN, '')); out += text.slice(last, start) + rewritten; last = end; count++; continue; } - onMatch?.(rule.id, text.slice(start, end)); - out += text.slice(last, start) + token; + out += text.slice(last, start) + tokenForMatch(rule.id, text.slice(start, end), onMatch); last = end; count++; } @@ -388,7 +392,6 @@ function applyRule(text, rule, isAllowed, onMatch) { function applyEntropy(text, spec, isAllowed, onMatch) { const candidates = new RegExp(`[A-Za-z0-9+/=_.~-]{${spec.minLength},}`, 'g'); const spans = text.includes(TOKEN_PREFIX) ? tokenSpans(text) : null; - const token = tokenFor(ENTROPY_RULE_ID); let out = ''; let last = 0; let count = 0; @@ -407,8 +410,7 @@ function applyEntropy(text, spec, isAllowed, onMatch) { if (!spec.keywords.some(k => before.includes(k))) continue; } - onMatch?.(ENTROPY_RULE_ID, m[0]); - out += text.slice(last, start) + token; + out += text.slice(last, start) + tokenForMatch(ENTROPY_RULE_ID, m[0], onMatch); last = end; count++; } @@ -532,8 +534,7 @@ function redactTree(node, redactor, ctx, tally) { const value = obj[key]; if (!isRedactableWhole(value)) return; - ctx?.onMatch?.(FIELD_RULE_ID, value); - setOwn(obj, key, tokenFor(FIELD_RULE_ID)); + setOwn(obj, key, tokenForMatch(FIELD_RULE_ID, value, ctx?.onMatch)); tally?.add([{ ruleId: FIELD_RULE_ID, count: 1 }]); changed = true; }; diff --git a/src/redact-rewrite.ts b/src/redact-rewrite.ts index 4ef696b7..03cede53 100644 --- a/src/redact-rewrite.ts +++ b/src/redact-rewrite.ts @@ -97,49 +97,74 @@ function summaryPathFor(jsonlPath: string): string { return jsonlPath.replace(/\.jsonl$/, SUMMARY_SUFFIX); } -type Match = { ruleId: string; value: string }; +type Match = { ruleId: string; value: string; absorbedBy?: number }; + +// In report mode each hit is written as its own numbered token, e.g. +// `[REDACTED:jwt--hit3]`, so it can be found in the output exactly; tokens that +// were already in the text have no number. Numbers are stripped for display. +const HIT_TOKEN = /\[REDACTED:([a-z0-9][a-z0-9-]*?)--hit(\d+)\]/g; + +/** `value` with any numbered tokens in it replaced by the values they stand for. */ +function restoreHits(value: string, matches: Match[]): string { + return value.replace(HIT_TOKEN, (_token, _ruleId, n: string) => matches[Number(n)].value); +} /** - * Pair each value a rule redacted with its token in the redacted text and - * report it. Tokens already in the original are skipped. Pairing is by rule and - * order, so the context is right per rule but approximate when one rule hits - * a line more than once around existing tokens. + * Report each hit with the redacted text around its token. A hit whose token + * was swallowed by a later one (a secret field redacted whole after a text rule + * hit part of it) is reported at the token that swallowed it. */ function reportHits( - original: string, redacted: string, matches: Match[], location: string, report: (hit: RedactionHit) => void -): void { - const queues = new Map(); - for (const m of matches) { - const q = queues.get(m.ruleId); - if (q) q.push(m.value); else queues.set(m.ruleId, [m.value]); +): string { + let display = ''; + let pos = 0; + const spans = new Map(); + for (const t of findRedactionTokens(redacted)) { + const numbered = /^(.*)--hit(\d+)$/.exec(t.ruleId); + display += redacted.slice(pos, t.start); + const shown = numbered ? `[REDACTED:${numbered[1]}]` : redacted.slice(t.start, t.end); + if (numbered && !spans.has(Number(numbered[2]))) { + spans.set(Number(numbered[2]), [display.length, display.length + shown.length]); + } + display += shown; + pos = t.end; } - const preexisting = new Map(); - for (const t of findRedactionTokens(original)) preexisting.set(t.ruleId, (preexisting.get(t.ruleId) ?? 0) + 1); + display += redacted.slice(pos); + const spanOf = (n: number): [number, number] | undefined => { + for (let i: number | undefined = n; i !== undefined; i = matches[i].absorbedBy) { + const span = spans.get(i); + if (span) return span; + } + return undefined; + }; const oneLine = (text: string) => text.replace(/\s+/g, ' '); - for (const t of findRedactionTokens(redacted)) { - const skip = preexisting.get(t.ruleId) ?? 0; - if (skip > 0) { preexisting.set(t.ruleId, skip - 1); continue; } - const value = queues.get(t.ruleId)?.shift(); - if (value === undefined) continue; - const before = redacted.slice(Math.max(0, t.start - CONTEXT_BEFORE), t.start); - const after = redacted.slice(t.end, t.end + CONTEXT_AFTER); + const hits = matches.map((m, n) => ({ m, span: spanOf(n) })); + hits.sort((a, b) => (a.span?.[0] ?? Infinity) - (b.span?.[0] ?? Infinity)); + for (const { m, span } of hits) { + const [start, end] = span ?? [0, 0]; + const before = display.slice(Math.max(0, start - CONTEXT_BEFORE), start); report({ location, - ruleId: t.ruleId, - shape: describeShape(value), - context: oneLine(`${t.start > CONTEXT_BEFORE ? '…' : ''}${before}${redacted.slice(t.start, t.end)}${after}`), + ruleId: m.ruleId, + shape: describeShape(m.value), + context: span + ? oneLine(`${start > CONTEXT_BEFORE ? '…' : ''}${before}${display.slice(start, end + CONTEXT_AFTER)}`) + : '', }); } + return display; } -/** Redact `text` with `run`, reporting each hit when `report` is set. */ +/** + * Redact with `run`, reporting each hit when `report` is set. Returns + * the redacted text with ordinary tokens either way. + */ function redactReporting( - text: string, ctx: RedactionContext, location: string, report: ((hit: RedactionHit) => void) | undefined, @@ -147,9 +172,16 @@ function redactReporting( ): string { if (!report) return run(ctx); const matches: Match[] = []; - const out = run({ ...ctx, onMatch: (ruleId, value) => matches.push({ ruleId, value }) }); - if (matches.length > 0) reportHits(text, out, matches, location, report); - return out; + const out = run({ + ...ctx, + onMatch: (ruleId, value) => { + const n = matches.length; + for (const [, , k] of value.matchAll(HIT_TOKEN)) matches[Number(k)].absorbedBy = n; + matches.push({ ruleId, value: restoreHits(value, matches) }); + return `[REDACTED:${ruleId}--hit${n}]`; + }, + }); + return matches.length > 0 ? reportHits(out, matches, location, report) : out; } /** Report every hit in a JSONL file, line by line. Writes nothing. */ @@ -157,7 +189,7 @@ function reportFile(file: string, label: string, redactor: Redactor, report: (hi const lines = fs.readFileSync(file, 'utf-8').split('\n'); const ctx = { source: 'rewrite', path: file }; lines.forEach((line, i) => { - redactReporting(line, ctx, `${label}:${i + 1}`, report, c => redactJsonlLine(line, redactor, c)); + redactReporting(ctx, `${label}:${i + 1}`, report, c => redactJsonlLine(line, redactor, c)); }); } @@ -183,6 +215,8 @@ function rewriteFileInPlace(file: string, redactor: Redactor, dryRun: boolean, t export async function rewriteArchive(options: RewriteOptions): Promise { const { archiveDir, redactor, embed } = options; const dryRun = options.dryRun === true; + // Reporting reads files before redaction; after a real rewrite there'd be nothing to find. + if (options.report && !dryRun) throw new Error('rewriteArchive: report requires dryRun'); const log = options.log ?? (() => {}); const tally = new FindingsTally(); const result: RewriteResult = { @@ -245,7 +279,7 @@ export async function rewriteArchive(options: RewriteOptions): Promise - redactReporting(text, ctx, `${where} ${field}`, options.report, c => { + redactReporting(ctx, `${where} ${field}`, options.report, c => { const r = redactor.redact(text, c); rowTally.add(r.findings); return r.text; @@ -259,7 +293,7 @@ export async function rewriteArchive(options: RewriteOptions): Promise redactJsonlLine(toolInput, redactor, c, rowTally) ); const output = tool.tool_result === null ? null : redactText(tool.tool_result, `${tool.tool_name} result`); diff --git a/src/redaction.ts b/src/redaction.ts index 52b815d8..1411650d 100644 --- a/src/redaction.ts +++ b/src/redaction.ts @@ -122,8 +122,10 @@ export interface RedactionContext { /** * Called with each value as it is redacted. For local review tooling * (`redact --report`) only: the value is the secret itself, so never log it. + * May return the token to write instead of `[REDACTED:]`, e.g. one + * numbered per hit; it must still have the token's shape. */ - onMatch?: (ruleId: string, value: string) => void; + onMatch?: (ruleId: string, value: string) => string | void; } export interface RedactionFinding { @@ -425,6 +427,14 @@ function tokenFor(ruleId: string): string { return `${TOKEN_PREFIX}${ruleId}]`; } +const WHOLE_TOKEN = /^\[REDACTED:[a-z0-9][a-z0-9-]*\]$/; + +/** The token for one redacted value: onMatch's, if it returns a valid one. */ +function tokenForMatch(ruleId: string, value: string, onMatch: RedactionContext['onMatch']): string { + const custom = onMatch?.(ruleId, value); + return typeof custom === 'string' && WHOLE_TOKEN.test(custom) ? custom : tokenFor(ruleId); +} + function tokenSpans(text: string): Array<[number, number]> { const spans: Array<[number, number]> = []; for (const m of text.matchAll(TOKEN_PATTERN)) spans.push([m.index!, m.index! + m[0].length]); @@ -446,7 +456,7 @@ function redactAroundTokens( spans: Array<[number, number]>, start: number, end: number, - token: string + token: () => string ): string | null { let out = ''; let pos = start; @@ -454,7 +464,7 @@ function redactAroundTokens( const flush = (to: number) => { const piece = text.slice(pos, to); if (/[A-Za-z0-9]/.test(piece)) { - out += token; + out += token(); replaced = true; } else { out += piece; @@ -518,7 +528,6 @@ function applyRule( onMatch?: RedactionContext['onMatch'] ): { text: string; count: number } { const spans = text.includes(TOKEN_PREFIX) ? tokenSpans(text) : null; - const token = tokenFor(rule.id); let out = ''; let last = 0; let count = 0; @@ -533,16 +542,16 @@ function applyRule( if (end <= start || start < last) continue; if (rule.useAllowlist && isAllowed(text.slice(start, end))) continue; if (spans && overlapsAny(spans, start, end)) { - const rewritten = redactAroundTokens(text, spans, start, end, token); + let token: string | undefined; + const rewritten = redactAroundTokens(text, spans, start, end, () => + (token ??= tokenForMatch(rule.id, text.slice(start, end).replace(TOKEN_PATTERN, ''), onMatch))); if (rewritten === null) continue; - onMatch?.(rule.id, text.slice(start, end).replace(TOKEN_PATTERN, '')); out += text.slice(last, start) + rewritten; last = end; count++; continue; } - onMatch?.(rule.id, text.slice(start, end)); - out += text.slice(last, start) + token; + out += text.slice(last, start) + tokenForMatch(rule.id, text.slice(start, end), onMatch); last = end; count++; } @@ -557,7 +566,6 @@ function applyEntropy( ): { text: string; count: number } { const candidates = new RegExp(`[A-Za-z0-9+/=_.~-]{${spec.minLength},}`, 'g'); const spans = text.includes(TOKEN_PREFIX) ? tokenSpans(text) : null; - const token = tokenFor(ENTROPY_RULE_ID); let out = ''; let last = 0; let count = 0; @@ -572,8 +580,7 @@ function applyEntropy( before = before.slice(before.lastIndexOf('\n') + 1).toLowerCase(); if (!spec.keywords.some(k => before.includes(k))) continue; } - onMatch?.(ENTROPY_RULE_ID, m[0]); - out += text.slice(last, start) + token; + out += text.slice(last, start) + tokenForMatch(ENTROPY_RULE_ID, m[0], onMatch); last = end; count++; } @@ -705,8 +712,7 @@ function redactTree( const redactWhole = (key: string) => { const value = obj[key]; if (!isRedactableWhole(value)) return; - ctx?.onMatch?.(FIELD_RULE_ID, value); - setOwn(obj, key, tokenFor(FIELD_RULE_ID)); + setOwn(obj, key, tokenForMatch(FIELD_RULE_ID, value, ctx?.onMatch)); tally?.add([{ ruleId: FIELD_RULE_ID, count: 1 }]); changed = true; }; diff --git a/test/redact-rewrite.test.ts b/test/redact-rewrite.test.ts index 12a05bc8..b794a2fb 100644 --- a/test/redact-rewrite.test.ts +++ b/test/redact-rewrite.test.ts @@ -209,6 +209,49 @@ describe('redact --rewrite backfill', () => { expect(archiveHit!.context).toContain('AccountKey=[REDACTED:connection-string-secret]'); expect(hits.some(h => h.location.startsWith('index:') && h.location.includes(' user'))).toBe(true); }); + + it('report refuses to run without dryRun, since files would already be rewritten', async () => { + mkdirSync(archiveDir, { recursive: true }); + await expect(rewriteArchive({ + archiveDir, redactor: createRedactor(DEFAULT_REDACTION_CONFIG), embed, report: () => {}, + })).rejects.toThrow(/dryRun/); + }); + + /** Dry-run report over one archive file holding `line`. */ + async function reportLine(line: string) { + mkdirSync(join(archiveDir, '-work-x'), { recursive: true }); + writeFileSync(join(archiveDir, '-work-x', 's.jsonl'), line + '\n'); + const hits: Array<{ location: string; ruleId: string; shape: string; context: string }> = []; + await rewriteArchive({ archiveDir, redactor: createRedactor(DEFAULT_REDACTION_CONFIG), embed, dryRun: true, report: h => hits.push(h) }); + return hits; + } + + it('report pairs a new hit with its own token, not an earlier token of the same rule', async () => { + const pw = new FakeSecrets(912).chars('abcdefghijklmnopqrstuvwxyz', 14); + const hits = await reportLine( + JSON.stringify({ content: `first: Server=a;password=${pw}; later: Server=b;Password=[REDACTED:connection-string-secret];` }) + ); + expect(hits).toHaveLength(1); + expect(hits[0].shape).toMatch(/^len=14 a /); + expect(hits[0].context).toContain('first: Server=a;password=[REDACTED:connection-string-secret]'); + expect(hits[0].context).not.toContain('--hit'); + }); + + it('report keeps a hit swallowed by a whole secret field, and shapes the field from its original value', async () => { + const fake = new FakeSecrets(913); + const word = fake.chars('abcdefghijklmnopqrstuvwxyz', 9); + const jwt = fake.jwt(); + const hits = await reportLine(JSON.stringify({ password: `${word} ${jwt}` })); + + expect(hits.map(h => h.ruleId).sort()).toEqual(['jwt', 'secret-field']); + const field = hits.find(h => h.ruleId === 'secret-field')!; + expect(field.shape).toMatch(new RegExp(`^len=${word.length + 1 + jwt.length} `)); + const swallowed = hits.find(h => h.ruleId === 'jwt')!; + expect(swallowed.shape).toMatch(new RegExp(`^len=${jwt.length} `)); + expect(swallowed.context).toBe(field.context); + expect(field.context).toContain('"password":"[REDACTED:secret-field]"'); + expect(JSON.stringify(hits)).not.toContain(word); + }); }); function expectNoSecretInCalls(calls: unknown[][], secrets: string[]) { From 87e806aa65377762f69c0d02cfb4167c950900a5 Mon Sep 17 00:00:00 2001 From: Steven Molen <533340+darthmolen@users.noreply.github.com> Date: Mon, 5 Oct 2026 09:31:33 -0500 Subject: [PATCH 14/15] fix(redaction): require a letter or digit in a connection-string value Markdown like `Password=`, `Pwd=` matched with "`," as the value, so docs and chat about connection strings were mangled and inflated the counts (found dogfooding on a real archive). A real value always has a letter or digit. Co-Authored-By: Claude Opus 5.5 --- dist/redaction-rules.js | 4 ++-- src/redaction-rules.ts | 4 ++-- test/redaction.test.ts | 4 ++++ 3 files changed, 8 insertions(+), 4 deletions(-) diff --git a/dist/redaction-rules.js b/dist/redaction-rules.js index 46b62b8c..6295381c 100644 --- a/dist/redaction-rules.js +++ b/dist/redaction-rules.js @@ -50,8 +50,8 @@ export const DEFAULT_REDACTION_CONFIG = { }, { id: 'connection-string-secret', - description: 'Secret value inside an Azure/ADO.NET connection string; account, server and database names are kept. Keywords are case-insensitive, as in ADO.NET; a path after Pwd= is the shell PWD variable and is kept', - pattern: String.raw `\b(?:(?:AccountKey|SharedAccessKey|SharedAccessSignature|SharedSecret|ClientSecret|Password)=|Pwd=(?![/~]))(?![\[$<{%])([^;"'\s]+)`, + description: 'Secret value inside an Azure/ADO.NET connection string; account, server and database names are kept. Keywords are case-insensitive, as in ADO.NET; a path after Pwd= is the shell PWD variable and is kept; the value needs a letter or digit, so Markdown like `Password=` is kept', + pattern: String.raw `\b(?:(?:AccountKey|SharedAccessKey|SharedAccessSignature|SharedSecret|ClientSecret|Password)=|Pwd=(?![/~]))(?![\[$<{%])(?=[^;"'\s]*[A-Za-z0-9])([^;"'\s]+)`, flags: 'i', secretGroup: 1, keywords: ['accountkey', 'sharedaccess', 'sharedsecret', 'clientsecret', 'password', 'pwd'], diff --git a/src/redaction-rules.ts b/src/redaction-rules.ts index b0d6d99a..7f5490b0 100644 --- a/src/redaction-rules.ts +++ b/src/redaction-rules.ts @@ -58,8 +58,8 @@ export const DEFAULT_REDACTION_CONFIG: RedactionConfig = { }, { id: 'connection-string-secret', - description: 'Secret value inside an Azure/ADO.NET connection string; account, server and database names are kept. Keywords are case-insensitive, as in ADO.NET; a path after Pwd= is the shell PWD variable and is kept', - pattern: String.raw`\b(?:(?:AccountKey|SharedAccessKey|SharedAccessSignature|SharedSecret|ClientSecret|Password)=|Pwd=(?![/~]))(?![\[$<{%])([^;"'\s]+)`, + description: 'Secret value inside an Azure/ADO.NET connection string; account, server and database names are kept. Keywords are case-insensitive, as in ADO.NET; a path after Pwd= is the shell PWD variable and is kept; the value needs a letter or digit, so Markdown like `Password=` is kept', + pattern: String.raw`\b(?:(?:AccountKey|SharedAccessKey|SharedAccessSignature|SharedSecret|ClientSecret|Password)=|Pwd=(?![/~]))(?![\[$<{%])(?=[^;"'\s]*[A-Za-z0-9])([^;"'\s]+)`, flags: 'i', secretGroup: 1, keywords: ['accountkey', 'sharedaccess', 'sharedsecret', 'clientsecret', 'password', 'pwd'], diff --git a/test/redaction.test.ts b/test/redaction.test.ts index 745e1c17..3ac42229 100644 --- a/test/redaction.test.ts +++ b/test/redaction.test.ts @@ -278,6 +278,10 @@ describe('redaction: allowlist and false-positive resistance (negative)', () => expectUntouched('PWD=/home/user1/src/project2\nOLDPWD=/home/user1\nmax_tokens: 4096\ntokenizer: bert-base-uncased'); }); + it('Markdown about connection-string keywords survives', () => { + expectUntouched('Set `AccountKey=`, `SharedAccessKey=`, `Password=` or `Pwd=` in the connection string.'); + }); + it('prose about tokens and passwords survives', () => { expectUntouched('Rotate the bearer token every 90 days. Basic authentication is disabled. The password policy requires 14 characters.'); }); From 46c6a8a94943ac6c45b3ca09aba19f3d8158f33b Mon Sep 17 00:00:00 2001 From: Steven Molen <533340+darthmolen@users.noreply.github.com> Date: Mon, 5 Oct 2026 09:46:33 -0500 Subject: [PATCH 15/15] docs(redaction): fold the design note into one summary-level guide docs/REDACTION.md now covers where redaction happens (with links to the choke points), the three detection layers, the guarantees, the backfill, configuration and limits, including how this differs from NER. It picks up behavior added since the first draft: index repair failing closed, partial tokens, whole-field replacement, kept file modes, and the connection-string exclusions. PHASE0-FINDINGS.md described the build process rather than the result, so its one lasting point (why the archive write is the choke point) moved into the guide and source comments now point there. Co-Authored-By: Claude Opus 5.5 --- dist/cursor-legacy.js | 2 +- dist/opencode-sync.js | 2 +- dist/redaction.d.ts | 2 +- dist/redaction.js | 2 +- dist/summarizer.d.ts | 2 +- docs/REDACTION.md | 337 +++++++++++++----------------- docs/redaction/PHASE0-FINDINGS.md | 184 ---------------- src/cursor-legacy.ts | 2 +- src/opencode-sync.ts | 2 +- src/redaction.ts | 2 +- src/summarizer.ts | 2 +- 11 files changed, 153 insertions(+), 386 deletions(-) delete mode 100644 docs/redaction/PHASE0-FINDINGS.md diff --git a/dist/cursor-legacy.js b/dist/cursor-legacy.js index 7b140e0c..e6ea46d6 100644 --- a/dist/cursor-legacy.js +++ b/dist/cursor-legacy.js @@ -206,7 +206,7 @@ export function importCursorLegacy(options) { ? lines.map(line => JSON.stringify({ ...JSON.parse(line), cwd })) : lines; // The export dir is a plugin-owned plaintext copy, so redact it at - // write time like the archive (docs/redaction/PHASE0-FINDINGS.md). + // write time like the archive (docs/REDACTION.md). const finalLines = redactor ? withCwd.map(line => redactJsonlLine(line, redactor, { source: 'cursor-legacy', path: outFile })) : withCwd; diff --git a/dist/opencode-sync.js b/dist/opencode-sync.js index c4bd4a8b..33916055 100644 --- a/dist/opencode-sync.js +++ b/dist/opencode-sync.js @@ -117,7 +117,7 @@ function writeSessionTranscript(db, session, filePath, redactor) { })); } // The staging transcript is a plugin-owned plaintext copy, so redact it at - // write time like the archive (docs/redaction/PHASE0-FINDINGS.md). + // write time like the archive (docs/REDACTION.md). const output = redactor ? lines.map(line => redactJsonlLine(line, redactor, { source: 'opencode', path: filePath })) : lines; diff --git a/dist/redaction.d.ts b/dist/redaction.d.ts index b19ee59b..d4af59bd 100644 --- a/dist/redaction.d.ts +++ b/dist/redaction.d.ts @@ -4,7 +4,7 @@ import { DEFAULT_REDACTION_CONFIG } from './redaction-rules.js'; * * Secrets are replaced with typed tokens (`[REDACTED:]`) at the archive * write, the one point every harness's transcripts pass through (see - * docs/redaction/PHASE0-FINDINGS.md). Everything downstream (SQLite text, + * docs/REDACTION.md). Everything downstream (SQLite text, * tool_calls, embeddings, summaries, show/read) reads the archive, so it only * ever sees redacted text. * diff --git a/dist/redaction.js b/dist/redaction.js index f654ae08..e4f206d5 100644 --- a/dist/redaction.js +++ b/dist/redaction.js @@ -8,7 +8,7 @@ import { DEFAULT_REDACTION_CONFIG } from './redaction-rules.js'; * * Secrets are replaced with typed tokens (`[REDACTED:]`) at the archive * write, the one point every harness's transcripts pass through (see - * docs/redaction/PHASE0-FINDINGS.md). Everything downstream (SQLite text, + * docs/REDACTION.md). Everything downstream (SQLite text, * tool_calls, embeddings, summaries, show/read) reads the archive, so it only * ever sees redacted text. * diff --git a/dist/summarizer.d.ts b/dist/summarizer.d.ts index c4b5595d..32272d9e 100644 --- a/dist/summarizer.d.ts +++ b/dist/summarizer.d.ts @@ -177,7 +177,7 @@ export interface SummarizeOptions { * Allow Claude session resume and Codex thread/fork (default true). Both * paths make the model read the *source* transcript rather than `exchanges`, * so callers pass false when the exchanges were redacted (see - * docs/redaction/PHASE0-FINDINGS.md) to force the transcript-text path. + * docs/REDACTION.md) to force the transcript-text path. */ allowResume?: boolean; } diff --git a/docs/REDACTION.md b/docs/REDACTION.md index 67bb15e7..6fe53eac 100644 --- a/docs/REDACTION.md +++ b/docs/REDACTION.md @@ -1,224 +1,175 @@ # Secret redaction episodic-memory replaces secrets in your conversations with typed tokens -**before** it archives, indexes, embeds, or summarizes them: +before it archives, indexes, embeds, or summarizes them: -``` +```text AccountKey=<88-char key> → AccountKey=[REDACTED:connection-string-secret] "ClientSecret": "" → "ClientSecret": "[REDACTED:quoted-secret-assignment]" → Authorization: Bearer → Authorization: Bearer [REDACTED:jwt] ``` -Values are redacted, not dropped. The rest of the conversation stays -searchable, and the token tells you what kind of value was there. You can -search for the tokens too: `episodic-memory search --text "[REDACTED:azure-storage-key]"` -finds every conversation where a storage key was pasted. +Only the value is replaced, so the rest of the conversation stays searchable, +and the token says what kind of value was there. +`episodic-memory search --text "[REDACTED:azure-storage-key]"` finds every +conversation where a storage key was pasted. + +Redaction is on by default and fails closed: if the rules can't load, sync +doesn't run. -Redaction is **on by default**, and it **fails closed**. If the rules can't be -loaded, sync refuses to run. It won't archive unredacted text. +## Where it happens -## What's covered +Every harness's transcripts (Claude Code, Codex, Cursor, opencode, OMP) are +copied into the conversation archive before anything else reads them. That +copy is the one point they all pass through, so it is where redaction happens +([`copyIfNewer`](../src/sync.ts) → [`copyFileRedacted`](../src/redaction.ts)). +Everything downstream reads the archive: | Where | Redacted? | |---|---| -| Conversation archive (`~/.config/superpowers/conversation-archive`) | Yes. This is the hook point. | -| SQLite index: message text and tool inputs/outputs | Yes. Built from the archive. | -| Embeddings | Yes. Built from redacted text. | -| Summaries (sent to a model) | Yes. Built from redacted text. Session resume and Codex fork are turned off (see below). | -| opencode and legacy Cursor staging exports | Yes. Redacted when written. | -| `show`, MCP `read` | Yes. They read the archive. | -| Sync logs | Rule IDs and counts only. Matched values are never logged. | -| Claude Code's own `~/.claude/projects`, Codex's `~/.codex/sessions`, etc. | **No.** Those files belong to the harness. Use its retention settings. | - -### Summaries - -Without redaction, the summarizer can *resume* a Claude Code session or *fork* -a Codex thread. Both make the model read the original, unredacted transcript. -With redaction on, summaries always come from the redacted conversation text. -This has one side effect for Codex-only setups: summarization then goes -through the Claude Agent SDK. If you don't have Claude set up, summaries fail -and retry on later syncs. Set `EPISODIC_MEMORY_SKIP_SUMMARIES=1` to turn them -off. Summaries are display-only, so search quality is unaffected. - -## Default rules - -Run `episodic-memory redact --print-default-rules` for the full set. In -summary: - -| Rule ID | Catches | +| Conversation archive (`~/.config/superpowers/conversation-archive`) | Yes, line by line as it's copied | +| SQLite index: message text and tool inputs/outputs | Yes, parsed from the archive ([indexer.ts](../src/indexer.ts) included) | +| Embeddings | Yes, built from redacted text | +| Summaries (sent to a model) | Yes. Summarizer resume and Codex fork are turned off, because both read the original transcript ([summarizer.ts](../src/summarizer.ts)) | +| opencode and legacy Cursor staging exports | Yes, when written ([opencode-sync.ts](../src/opencode-sync.ts), [cursor-legacy.ts](../src/cursor-legacy.ts)) | +| `show`, MCP `read` | Yes, they read the archive | +| Logs | Rule IDs and counts only | +| The harness's own files (`~/.claude/projects`, `~/.codex/sessions`, …) | **No.** They belong to the harness; use its retention settings | + +With summarizer resume off, Codex-only setups summarize through the Claude +Agent SDK. Without Claude configured, set `EPISODIC_MEMORY_SKIP_SUMMARIES=1`. +Summaries are display-only, so search is unaffected. + +## How a secret is recognized + +The engine is [src/redaction.ts](../src/redaction.ts) and the default rules are +[src/redaction-rules.ts](../src/redaction-rules.ts). Each archive line is parsed +as JSON and passes through three layers: + +1. **Shape.** Values whose format gives them away: private keys, JWTs, + provider-prefixed keys, Entra client secrets (`…Q~…`), 88-character Azure + storage keys, SAS `sig=` values. +2. **Key context.** Values with no recognizable format, caught by the name + they're assigned to: `Password=` in a connection string, `"ClientSecret": "…"`, + ``, `{"name": "DB_PASSWORD", "value": "…"}`, + `password: …`. +3. **Field name.** In parsed JSON, any string whose field name is + secret-looking is replaced whole, including tool inputs and MCP results. + +A **secret-looking name** *ends* in `secret`, `password`, `passwd`, +`passphrase`, `apikey`, `accesskey`, `accountkey`, `privatekey`, `token`, +`credential(s)` or a similar key word, in any case and with any separators +(`AzureAd:ClientSecret`, `DB_PASSWORD2`). `TokenEndpoint` and `passwordPolicy` +don't count. Placeholders (`${X}`, `$(X)`, `#{X}#`, `{{x}}`, ``, `%X%`, +`"string"`, `"*****"`) are left alone. + +Git SHAs and GUIDs are allowlisted for the shape rules, so commit hashes and +tenant, client and object IDs stay searchable. Key-context rules ignore the +allowlist: a GUID in a `password` slot is a secret. An entropy fallback exists +but is off by default. + +| Rules | Catch | |---|---| -| `private-key-block` | PEM and OpenSSH private keys, including truncated ones | -| `connection-string-secret` | `AccountKey=`, `SharedAccessKey=`, `SharedAccessSignature=`, `Password=`, `Pwd=` in connection strings, in any case. Only the value is redacted; server, account, and database names stay. | -| `azure-sas-token` | The `sig=` of a SAS URL. The URL and other parameters stay. | -| `jwt` | JWTs, including Entra ID / Azure access tokens | -| `anthropic-api-key`, `openai-api-key`, `github-token`, `aws-access-key-id`, `aws-secret-access-key`, `slack-token`, `google-api-key`, `npm-token` | Provider keys with a recognizable prefix | -| `azure-client-secret` | Entra ID app client secrets (the `…Q~…` format) | -| `azure-storage-key` | Standalone 88-character base64 keys (Storage, Cosmos DB, Function keys) | -| `url-credentials` | The password in `scheme://user:password@host` | -| `bearer-token`, `basic-auth` | `Authorization` header values | -| `azure-keyvault-secret` | The `value` of a Key Vault secret bundle (`az keyvault secret show`, SDK JSON), whatever the secret is named | -| `name-value-secret` | The `value` of a `{"name": , "value": …}` object, in either order: `az webapp`/`functionapp config appsettings list`, Kubernetes `env`, ARM/Bicep parameters | -| `xml-appsettings-secret` | `web.config` / `app.config` ``, either attribute order | -| `xml-secret-element` | `…`, `…` and the like | -| `xml-secret-attribute` | Secret-named XML attributes, e.g. `userPWD="…"` in Azure publish profiles | -| `quoted-secret-assignment` | A **quoted** value assigned to a secret-looking name: `"ClientSecret": "…"` (appsettings.json and any JSON in tool output), `ClientSecret = "…"` (C#), `apiKey: '…'` (JS/YAML/Python). No digit or minimum entropy is required; the value must have no spaces. | -| `secret-assignment` | An **unquoted** value assigned to a secret-looking key: `password: …`, `CLIENT_SECRET=…`. Covers decrypted SOPS, YAML, and dotenv. Needs 8+ characters including a digit, and skips code (`env.X`, `getPassword()`). | -| `secret-field` | Not a text rule: in parsed JSON (transcript lines, structured MCP/tool results, tool inputs), a string whose **field name** is secret-looking, or the `value` of a `{name, value}` pair or Key Vault bundle, is redacted whole. See `secretFields` below. | - -A **secret-looking name** ends in `secret`, `password`, `passwd`, -`passphrase`, `apikey`, `accesskey`, `accountkey`, `privatekey`, `sharedkey`, -`primarykey`, `secondarykey`, `masterkey`, `signingkey`, `subscriptionkey`, -`clientkey`, `encryptionkey`, `token`, or `credential(s)`, in any case and -with any separators (`AzureAd:ClientSecret`, `Stripe__ApiKey`, `DB_PASSWORD2`). -Because it has to *end* in one of these, `TokenEndpoint`, `secretName`, -`passwordPolicy`, `tokenType`, and `maxTokens` don't count. - -Key-context rules leave **placeholders** alone: `${X}`, `$(X)`, `#{X}#` -(Azure DevOps token replacement), `{{x}}`, ``, `%X%`, `__X__`, and type -names or masks like `"string"` and `"*****"`. All other rules skip the same -`${X}`, ``, and `%X%` forms. - -**Allowlisted:** git SHAs and GUIDs are not redacted by shape-based rules, -so tenant, client, object, and subscription IDs and commit hashes stay -searchable. Key-context rules ignore the allowlist on purpose: a GUID in a -`password` slot (older `az ad sp create-for-rbac` output) *is* a secret. - -**Entropy fallback:** off by default. When it's on, a long, high-entropy string -is redacted only if a keyword (`secret`, `key`, `token`, …) appears just before -it in the same text value. - -**SOPS:** decrypted SOPS output drops its `sops:` metadata block, so it has no -reliable shape. It's covered by the value rules plus `secret-assignment`. -Encrypted values (`ENC[AES256_GCM,…]`) are left alone. - -## Configuration - -| Variable | Default | Meaning | -|---|---|---| -| `EPISODIC_MEMORY_REDACTION` | `on` | `off` disables redaction entirely: archive copies become byte-for-byte again, and summaries may resume sessions again. | -| `EPISODIC_MEMORY_REDACTION_RULES` | `/redaction-rules.json` if it exists | Path to a custom rules file. If you set this and the file is missing, that's an error. | -| `EPISODIC_MEMORY_REDACTION_STRICT` | `1` | When the rules fail to load, `1` stops sync, index, and import (exit 1, nothing written). `0` logs a loud warning and continues **unredacted**. | - -`` is `~/.config/superpowers` unless `EPISODIC_MEMORY_CONFIG_DIR`, -`PERSONAL_SUPERPOWERS_DIR`, or `XDG_CONFIG_HOME` say otherwise. - -### Custom rules - -Create `~/.config/superpowers/redaction-rules.json`. By default it extends the -bundled rules: - -```jsonc -{ - // Add rules (a rule with a default's id replaces that default) - "rules": [ - { - "id": "contoso-api-key", // lowercase, digits, dashes - "pattern": "\\bctso_[A-Za-z0-9]{32}\\b", // JavaScript regex - "flags": "i", // optional, any of "imsu" - "keywords": ["ctso_"], // optional prefilter (case-insensitive) - "secretGroup": 0, // optional: redact only this group ([1, 2] = first that matched) - "useAllowlist": true // optional: false = redact even SHA/GUID-shaped values - } - ], - "disableRules": ["basic-auth"], // turn off defaults by id - "allowlist": [ // full-match patterns that are never redacted - { "id": "build-ids", "pattern": "build-[0-9]{8}" } - ], - "entropy": { "enabled": true }, // partial override of the entropy settings - "secretFields": { // field-name context in parsed JSON - "enabled": true, - // matched against the END of the key lowercased with separators removed; - // this example adds "connectionstring" to redact whole connection strings - "keyPattern": "secret|password|passwd|userpwd|passphrase|apikey|accesskey|accountkey|privatekey|sharedkey|primarykey|secondarykey|masterkey|signingkey|subscriptionkey|clientkey|encryptionkey|token|credentials?|connectionstring" - }, - "includeDefaults": true // false = use only this file's rules -} -``` - -Every rule is validated when it loads. An invalid regex, a bad id, a pattern -that matches the empty string, or a `secretGroup` that doesn't exist is a -rules-load failure, which strict mode treats as fatal. To try rules out: - -```bash -echo 'password: hunter2hunter2' | episodic-memory redact --stdin -# password: [REDACTED:secret-assignment] -# 1 value(s) redacted (secret-assignment: 1) (stderr) -``` +| `private-key-block`, `jwt`, `anthropic-api-key`, `openai-api-key`, `github-token`, `aws-access-key-id`, `aws-secret-access-key`, `slack-token`, `google-api-key`, `npm-token`, `azure-client-secret`, `azure-storage-key`, `azure-sas-token` | Shape | +| `connection-string-secret` | `AccountKey=`, `SharedAccessKey=`, `Password=`, `Pwd=` and similar in connection strings, any case | +| `url-credentials`, `bearer-token`, `basic-auth` | `scheme://user:password@host`, `Authorization` headers | +| `azure-keyvault-secret`, `name-value-secret` | Key Vault bundles; `{name, value}` pairs (`az … appsettings list`, Kubernetes `env`, ARM parameters) | +| `xml-appsettings-secret`, `xml-secret-element`, `xml-secret-attribute` | `web.config`, `…`, publish-profile `userPWD="…"` | +| `quoted-secret-assignment` | `"ClientSecret": "…"`, `ClientSecret = "…"`, `apiKey: '…'`; no spaces in the value | +| `secret-assignment` | `password: …`, `CLIENT_SECRET=…` (YAML, dotenv, decrypted SOPS); 8+ characters with a digit, code like `env.X` skipped | +| `secret-field` | Layer 3: a secret-named JSON field, replaced whole | + +`episodic-memory redact --print-default-rules` prints the full definitions. + +## Guarantees + +- **Fails closed.** In strict mode (the default), sync, index, `index repair` + and import stop if the rules can't load. +- **Idempotent.** Tokens never match again. When a match runs into an existing + token, only the text outside it is redacted, so re-running after adding a + rule is safe. +- **Structure-preserving.** One output line per input line, valid JSON stays + valid, and lines with no secrets stay byte-for-byte identical. Index line + ranges and MCP `read` ranges stay correct. Archive copies keep the source + file's permissions. +- **Values are never printed.** Logs and reports name rules and counts; the + review report adds the value's shape, never the value. ## Cleaning up existing data -New syncs only redact new or changed files. To redact everything indexed -before you upgraded, or after you add a rule: +New syncs redact what they copy. To redact what was archived and indexed +before, or after you add a rule: ```bash -episodic-memory redact --rewrite --dry-run # report what would change -episodic-memory redact --rewrite --dry-run --report # list each hit to review +episodic-memory redact --rewrite --dry-run # what would change; writes nothing +episodic-memory redact --rewrite --dry-run --report # each hit, to check for false positives episodic-memory redact --rewrite # apply ``` -A dry run writes nothing, not even a schema migration: it opens the index +`--rewrite` ([src/redact-rewrite.ts](../src/redact-rewrite.ts)) redacts the +archive, the staging exports and the index in place, re-embeds only the rows +that changed, and deletes summaries built from unredacted text so the next +sync regenerates them. It takes the sync lock. A dry run opens the index read-only. -`--report` lists every value the rewrite would redact, so you can check for -false positives before applying. Each hit shows where it is, the rule, the -value's shape, and the redacted text around it. The value itself is never -printed: +`--report` prints one hit per value: -``` +```text -work-contoso/4f1c….jsonl:212 quoted-secret-assignment len=40 aA9- H=4.9 …"AzureAd": { "ClientId": "…", "ClientSecret": "[REDACTED:quoted-secret-assignment]", "TenantId… ``` -The shape is the length, the character classes (`a` lowercase, `A` uppercase, -`9` digits, `-` symbols, `_` whitespace) and the Shannon entropy in bits per -character. A random key is long with high entropy (about 4.5 or more); a word or -a placeholder is short or low. If a rule fires on something that isn't a secret, -turn it off or narrow it in `redaction-rules.json`. - -`--rewrite`: - -- redacts every archive file in place, keeping line numbers and timestamps -- redacts the opencode and Cursor staging exports in place -- redacts every index row in place and re-embeds only the rows that changed -- deletes summaries that were generated from unredacted text, so the next - sync regenerates them from the redacted archive - -It takes the same lock as `sync`, so it won't run while a sync is in progress. -Running it twice is safe: the second run finds nothing to change. - -Before your first sync with this version, consider a one-time scan of your -existing `~/.claude/projects` history. Those source files are never modified. - -## How it works - -The hook point is the copy into the archive. Every harness (Claude Code, -Codex, Cursor, opencode, OMP) passes through that copy, and every later stage -reads the archive rather than the source. The design and the research behind -it are in [redaction/PHASE0-FINDINGS.md](redaction/PHASE0-FINDINGS.md). - -Each archive line is parsed as JSON. Every string value goes through the text -rules, values are redacted whole when their field name marks them as secrets -(`secretFields`), and a key that is itself a secret is renamed to its token. -Only lines that changed are re-serialized. Lines with no secrets stay byte-for-byte -identical. The archive keeps exactly one line per source line, so index line -ranges and MCP `read` ranges still line up. A line that isn't valid JSON (for -example, a half-written last line) is redacted as plain text. - -## Limitations - -- Pattern rules miss secrets that have no recognizable shape and no - `key: value` context. The entropy fallback helps when it's on, but recall - isn't perfect. -- A secret-looking name is required for the key-context rules. A setting - named `Stripe` or `ConnectionStrings__Default` with a shapeless value is - only caught if a shape rule matches the value. Connection strings are - handled by `connection-string-secret`, which keeps server names searchable. - Add your own names via `secretFields.keyPattern` or a custom rule. -- Quoted values containing spaces (multi-word passphrases) aren't caught by - `quoted-secret-assignment`. The rule excludes them so that UI labels like - `ErrorMessage = "Invalid password"` aren't redacted. -- Positional secrets in code (`new ClientSecretCredential(t, c, "…")`) are - caught only by shape, e.g. `azure-client-secret`. -- On lines that get redacted, re-serializing can change number formatting for - integers above 2^53. No supported harness writes such numbers. +The shape is the length, the character classes (`a` lower, `A` upper, `9` +digits, `-` symbols, `_` spaces) and the entropy in bits per character. A +random key is long with entropy around 4.5 or more; a word or placeholder is +short or low. + +## Configuration + +| Variable | Default | Meaning | +|---|---|---| +| `EPISODIC_MEMORY_REDACTION` | `on` | `off` turns redaction off: byte-for-byte archive copies, summarizer resume allowed | +| `EPISODIC_MEMORY_REDACTION_RULES` | `/redaction-rules.json` if present | Custom rules file; a missing file you named is an error | +| `EPISODIC_MEMORY_REDACTION_STRICT` | `1` | `0` continues **unredacted**, with a warning, when the rules fail to load | + +Custom rules extend the defaults: + +```jsonc +{ + "rules": [ // a rule with a default's id replaces it + { "id": "contoso-api-key", "pattern": "\\bctso_[A-Za-z0-9]{32}\\b", "keywords": ["ctso_"] } + ], + "disableRules": ["basic-auth"], + "allowlist": [{ "id": "build-ids", "pattern": "build-[0-9]{8}" }], + "entropy": { "enabled": true }, + "secretFields": { "keyPattern": "secret|password|token|credentials?|connectionstring" }, + "includeDefaults": true +} +``` + +A rule can also set `flags` (`imsu`), `secretGroup` (redact only that capture +group) and `useAllowlist: false`. Every rule is validated on load; an invalid +one is a load failure, which strict mode treats as fatal. Try rules with +`echo 'password: hunter2hunter2' | episodic-memory redact --stdin`. + +## Limits + +This is pattern and context matching, not named-entity recognition: a value +is caught by its format, by the name it's assigned to, or by the field it's +in. That covers how secrets reach transcripts in practice (pasted config, CLI +output, connection strings, tool results), and it is deterministic, cheap +enough for every line of every sync, and each token names the rule that fired. +What it misses: + +- A secret in prose with no name or shape: "the password is hunter2". +- An unquoted letters-only value (`password: hunter`): `secret-assignment` + needs a digit so ordinary English doesn't trigger it. +- Code-like values (`config.password`, `getPassword()`), skipped on purpose. +- Quoted values with spaces, so UI strings like `"Invalid password"` survive. +- Settings with a shapeless value and a name that isn't secret-looking + (`Stripe`, `ConnectionStrings__Default`). Add the name to + `secretFields.keyPattern` or write a rule. + +The entropy fallback narrows some of these gaps, at the cost of false +positives. diff --git a/docs/redaction/PHASE0-FINDINGS.md b/docs/redaction/PHASE0-FINDINGS.md deleted file mode 100644 index d73e34eb..00000000 --- a/docs/redaction/PHASE0-FINDINGS.md +++ /dev/null @@ -1,184 +0,0 @@ -# Redaction — Phase 0 findings - -Research note for the inline secret redaction fork. Read-only survey of -`src/`, `cli/`, and `hooks/` at upstream `7e06519` (v1.6.0 + fixes). This -note answers the six Phase 0 questions and records the choke-point decision. - -## Pipeline as it exists today - -``` -SessionStart hook (hooks/hooks.json) - └─ cli/episodic-memory.js sync --background - └─ dist/sync-cli.js - ├─ exportOpencodeSessions() opencode.db → /opencode-transcripts/*.jsonl - └─ syncConversations(src, archive) for each source dir - ├─ copyIfNewer(src → archive) byte-for-byte copy - ├─ parseConversation(archive) → exchanges - │ ├─ generateExchangeEmbedding(user, assistant, toolNames) - │ └─ insertExchange() → SQLite exchanges + tool_calls + vec_exchanges - └─ summarizeConversation(exchanges, sessionId) → -summary.txt - -episodic-memory index (cli/index-conversations.js → dist/index-cli.js → indexer.ts) - └─ copyFileSync(src → archive), parseConversation(**src**), embed, insert, summarize - -episodic-memory import-cursor-history (cursor-legacy.ts) - └─ state.vscdb → /cursor-legacy-export/*.jsonl (a sync source dir) - -episodic-memory index --repair (verify.ts) parses archive files, re-embeds, re-summarizes -``` - -## Answers - -### 1. Where does sync copy into the archive? Byte-for-byte or parsed first? - -Byte-for-byte. `copyIfNewer()` in `src/sync.ts` does `fs.copyFileSync` to -`.tmp.` then `renameSync`, then stamps the source mtime on the -destination (the mtime is the "is the archive current?" check on the next -run). Parsing happens **after** the copy, and in `sync.ts` it parses the -archive file, not the source. - -There are three more copy sites, all in `src/indexer.ts` -(`indexConversations`, `indexSession`, `indexUnprocessed`). They also use -`fs.copyFileSync`, but they **parse the source path**, not the archive. That -matters: redacting the archive copy alone would not reach the index on the -`episodic-memory index` path. - -### 2. One shared parse path or one per source? - -One entry point, `parseConversation()` in `src/parser.ts`. It sniffs the file -and dispatches to one of five per-harness parsers (Claude, Codex, Cursor, -opencode, OMP). More important for redaction: every harness reaches the -archive through the same copy step. opencode and legacy Cursor are first -exported from their SQLite stores into staging JSONL directories under the -plugin's config dir (`opencode-transcripts/`, `cursor-legacy-export/`), and -those staging directories are then ordinary sync sources. - -The staging exports are plugin-owned plaintext copies outside any harness's -own retention, so they count as a fourth sink even though the spec's table -doesn't list them. - -### 3. Where is exchange text written to SQLite, and where is the embedding input built? - -- `insertExchange()` in `src/db.ts` writes `exchanges.user_message`, - `exchanges.assistant_message`, `tool_calls.tool_input` (JSON-stringified - tool input), and `tool_calls.tool_result`. Text search (`search.ts`) runs - `LIKE` over `user_message`/`assistant_message`. -- Embedding input is `generateExchangeEmbedding(userMessage, assistantMessage, - toolNames)` in `src/embeddings.ts`, called from `sync.ts`, `indexer.ts`, - `verify.ts`, and `embedding-migration.ts`. The migration re-embeds from the - **SQLite rows**, so it inherits whatever text is already stored. - -All of these are built from the parsed `ConversationExchange` objects, so a -clean parse input gives clean SQLite text and clean embedding input. - -### 4. Where does the summarizer get its input? - -Usually from the parsed exchanges (`formatConversationText(exchanges)`), and -`sync.ts` parses those from the archive. **But two paths bypass the archive -entirely:** - -- **Claude session resume.** For Claude conversations with at most 15 - exchanges, `summarizeConversation()` calls the Agent SDK with - `resume: sessionId`. The SDK loads the **source** transcript from - `~/.claude/projects/...` and sends it to the model. The prompt carries no - transcript text on this path. -- **Codex fork.** For Codex conversations, `callCodex()` runs - `thread/fork` on the **source** rollout via `codex app-server`. Note that - `getCodexSessionId()` derives the id from the exchanges when no - `sessionId` is passed, so passing `undefined` is not enough to stop it. - -Both paths send unredacted source content to a model endpoint no matter what -the archive holds. Redaction therefore has to turn resume and fork off and -force the transcript-text path, which is built from redacted exchanges. - -### 5. How does the `DO NOT INDEX THIS CHAT` marker work? - -`shouldSkipConversation()` in `src/sync.ts` streams a file in 1 MiB chunks -and looks for any of three markers. If the read fails, it fails closed -(skips the file). Sync calls it on the **archive** path, after the copy, to -gate both indexing and summary queueing. So an excluded conversation is -still copied to the archive; only the index and the summarizer skip it. - -The hint holds: the archive file is the reference that every downstream -stage of `sync` reads. The marker check sits right after the archive write, -and that is where the redaction hook belongs. - -### 6. Must the archive stay valid harness JSONL? - -Yes, on two counts: - -- **Structure.** `show.ts` (CLI `show`, MCP `read`) runs `JSON.parse` on every - line, and one invalid line throws. Harness detection in both `parser.ts` - and `show.ts` keys off line shapes (`type`, `payload`, `role`, ...). -- **Line numbers.** `exchanges.line_start`/`line_end` index into archive lines. - MCP `read` takes `startLine`/`endLine`, and incremental indexing resumes - from `MAX(line_end)` (#152). Redaction must keep **exactly one output line - per input line**. - -So redaction parses each line as JSON, redacts string values, and -re-serializes only the lines that changed. Unchanged lines stay -byte-identical. A line that isn't valid JSON (for example, a partially -written last line) is redacted as raw text, because it was already invalid. - -## Decision: choke point (b), the archive write - -Option (a) doesn't exist. No single function sees every source *before* the -archive write. The only place every harness converges is the copy into the -archive itself. - -So the hook goes at the **archive write**, which becomes a redacting copy -(`copyFileRedacted()` in `src/redaction.ts`), and every downstream stage -reads from the archive. Three supporting changes are needed to make -"downstream reads the archive" actually true: - -1. **`indexer.ts` parses the archive, not the source.** All three copy sites - switch to the redacting copy. They now refresh the archive when the source - is newer, as `sync` already does, so parsing the archive loses no data. -2. **Summarizer resume and fork are disabled while redaction is active.** - `summarizeConversation()` gains an `allowResume` option. `sync.ts`, - `indexer.ts`, and `verify.ts` pass `allowResume: false` when a redactor is - active, which forces the transcript-text path built from redacted - exchanges. Trade-off: Codex-only users then summarize through the Claude - transcript path. If no Claude auth is available, that writes a retryable - error sentinel. `EPISODIC_MEMORY_SKIP_SUMMARIES=1` turns summaries off - entirely. -3. **Staging exports are redacted at write time.** The opencode export - (`opencode-sync.ts`) and the Cursor legacy export (`cursor-legacy.ts`) run - each JSONL line through the same `redactJsonlLine()` before writing. They - are plugin-owned copies, and the archive hook alone would leave a - plaintext copy beside it. - -Everything else (SQLite text, `tool_calls`, embeddings, summaries, `show`, -MCP `read`, embedding migration) reads the archive or the rows built from -it, so it inherits clean text with no further hooks. - -## SOPS - -Decrypted SOPS output (`sops -d`, `sops exec-env`) **drops** the `sops:` -metadata block and comes out as plain YAML, JSON, or dotenv. Its shape is -not reliable. There is no dedicated SOPS rule. Coverage comes from the -value-level rules (storage keys, client secrets, connection strings, private -keys) plus the keyword-assignment rule (`password: ...`, `client_secret=...`). -Encrypted values (`ENC[AES256_GCM,data:...]`) are left alone because they -are safe to keep. - -## Windows - -The redaction code is pure Node `fs` plus regex, and it adds no native -dependencies. The temp-file-and-rename pattern already runs on Windows in -upstream `copyIfNewer`. I could **not** verify sqlite-vec or Transformers.js -on Windows from this Linux container. That is unchanged upstream surface, -and it still needs a manual check on a Windows host before rollout. - -## Other observations - -- `embedding-migration.ts` re-embeds from SQLite text. After a - `redact --rewrite` cleans the rows, the migration can't reintroduce - secrets. -- The summarizer's `SummarizerSdkError` keeps up to 300 characters of the - SDK's `result` text in logs and error sentinels. If a model ever echoed a - secret, it would land there. Feeding the summarizer redacted input removes - that source. -- `verify.ts --repair` re-summarizes without a `sessionId`, but the Codex fork - still fires through `getCodexSessionId()`'s fallback. It needs - `allowResume: false` too. diff --git a/src/cursor-legacy.ts b/src/cursor-legacy.ts index 8c91897f..e37366b1 100644 --- a/src/cursor-legacy.ts +++ b/src/cursor-legacy.ts @@ -272,7 +272,7 @@ export function importCursorLegacy(options: CursorLegacyImportOptions): CursorLe ? lines.map(line => JSON.stringify({ ...JSON.parse(line), cwd })) : lines; // The export dir is a plugin-owned plaintext copy, so redact it at - // write time like the archive (docs/redaction/PHASE0-FINDINGS.md). + // write time like the archive (docs/REDACTION.md). const finalLines = redactor ? withCwd.map(line => redactJsonlLine(line, redactor, { source: 'cursor-legacy', path: outFile })) : withCwd; diff --git a/src/opencode-sync.ts b/src/opencode-sync.ts index 60dae458..eca6876c 100644 --- a/src/opencode-sync.ts +++ b/src/opencode-sync.ts @@ -176,7 +176,7 @@ function writeSessionTranscript( } // The staging transcript is a plugin-owned plaintext copy, so redact it at - // write time like the archive (docs/redaction/PHASE0-FINDINGS.md). + // write time like the archive (docs/REDACTION.md). const output = redactor ? lines.map(line => redactJsonlLine(line, redactor, { source: 'opencode', path: filePath })) : lines; diff --git a/src/redaction.ts b/src/redaction.ts index 1411650d..999f0136 100644 --- a/src/redaction.ts +++ b/src/redaction.ts @@ -9,7 +9,7 @@ import { DEFAULT_REDACTION_CONFIG } from './redaction-rules.js'; * * Secrets are replaced with typed tokens (`[REDACTED:]`) at the archive * write, the one point every harness's transcripts pass through (see - * docs/redaction/PHASE0-FINDINGS.md). Everything downstream (SQLite text, + * docs/REDACTION.md). Everything downstream (SQLite text, * tool_calls, embeddings, summaries, show/read) reads the archive, so it only * ever sees redacted text. * diff --git a/src/summarizer.ts b/src/summarizer.ts index 23b0f31d..801f3743 100644 --- a/src/summarizer.ts +++ b/src/summarizer.ts @@ -707,7 +707,7 @@ export interface SummarizeOptions { * Allow Claude session resume and Codex thread/fork (default true). Both * paths make the model read the *source* transcript rather than `exchanges`, * so callers pass false when the exchanges were redacted (see - * docs/redaction/PHASE0-FINDINGS.md) to force the transcript-text path. + * docs/REDACTION.md) to force the transcript-text path. */ allowResume?: boolean; }