diff --git a/CHANGELOG.md b/CHANGELOG.md index 511074ff..223da791 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -5,6 +5,23 @@ All notable changes to this project will be documented in this file. The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.0.0/), and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html). +## [Unreleased] + +### Security + +- Secrets no longer leak into your conversation memory. Azure client secrets, storage keys, SAS signatures, connection-string passwords, private keys, JWTs, GitHub, Anthropic, OpenAI, and AWS keys, and `password: …` style assignments are now replaced with tokens like `[REDACTED:azure-storage-key]` before a conversation is archived, indexed, embedded, or summarized. Before this, every pasted secret was copied into three new places on disk and could be sent to the summarization model. The rest of the conversation stays searchable, and so do git SHAs and Azure tenant, client, and object IDs. Redaction is on by default and fails closed: if the rules can't load, sync won't run rather than store secrets. .NET and Azure shapes where only the setting's name gives the secret away are covered too, even when the value has no recognizable format: `appsettings.json` fields, `web.config` ``, publish-profile passwords, `az webapp config appsettings list` and `az keyvault secret show` output, C# `ClientSecret = "…"` assignments, and structured MCP results. +- With redaction on, summaries no longer resume Claude Code sessions or fork Codex threads. Both of those let the model read the original, unredacted transcript. Summaries now come from the redacted text. Codex-only users without Claude configured can set `EPISODIC_MEMORY_SKIP_SUMMARIES=1`. + +### Added + +- `episodic-memory redact --rewrite` cleans data you indexed before upgrading. It redacts the archive and the search index in place, re-embeds only the messages that changed, and deletes summaries built from unredacted text so they regenerate. Use `--dry-run` to preview (it writes nothing, not even a schema migration), and add `--report` to list each value it would redact, by location, rule and shape, without printing the value. +- Custom redaction rules via `~/.config/superpowers/redaction-rules.json`, which extends the bundled defaults. Try rules with `episodic-memory redact --stdin`. See `docs/REDACTION.md`. +- New settings: `EPISODIC_MEMORY_REDACTION` (`on`/`off`), `EPISODIC_MEMORY_REDACTION_RULES`, and `EPISODIC_MEMORY_REDACTION_STRICT`. + +### Changed + +- `episodic-memory index` now parses the archived copy instead of the source transcript, and refreshes that copy when the source has grown. This is the same behavior `sync` already had. + ## [1.6.0] - 2026-09-08 Adds a fifth conversation source, an off switch for automatic syncing, and two fixes for real-world resource problems. diff --git a/README.md b/README.md index ee952317..f59264ee 100644 --- a/README.md +++ b/README.md @@ -327,6 +327,18 @@ Add to `.claude/hooks/session-end`: episodic-memory sync ``` +### `episodic-memory redact` + +```bash +episodic-memory redact --rewrite --dry-run # what would change +episodic-memory redact --rewrite --dry-run --report # each hit: location, rule, value shape +episodic-memory redact --rewrite # redact the existing archive + index in place +episodic-memory redact --stdin < file.jsonl # try the rules on some text +episodic-memory redact --print-default-rules +``` + +See [docs/REDACTION.md](docs/REDACTION.md). + ### `episodic-memory stats` Display index statistics including conversation counts, date ranges, and project breakdown. @@ -415,6 +427,28 @@ open output.html 4. **Index** - Stores in SQLite with sqlite-vec for fast similarity search 5. **Search** - Semantic search using vector similarity or exact text matching +## Secret Redaction + +Secrets in your conversations (Azure client secrets, storage keys, SAS +signatures, connection-string passwords, private keys, JWTs, provider API keys, +`password: …` assignments) are replaced with typed tokens such as +`[REDACTED:azure-storage-key]` **before** anything is archived, indexed, +embedded, or sent to the summarizer. Git SHAs and GUIDs are never redacted, so +they stay searchable. The tokens are searchable too. + +Redaction is on by default and fails closed: if the rules can't be loaded, sync +won't run. After upgrading, run `episodic-memory redact --rewrite` once to clean +data you indexed before. + +| Variable | Default | Meaning | +|---|---|---| +| `EPISODIC_MEMORY_REDACTION` | `on` | `off` disables redaction | +| `EPISODIC_MEMORY_REDACTION_RULES` | `/redaction-rules.json` | Custom rules file (extends the defaults) | +| `EPISODIC_MEMORY_REDACTION_STRICT` | `1` | `0` continues unredacted (with a warning) when rules fail to load | + +See [docs/REDACTION.md](docs/REDACTION.md) for the rule list, the custom-rules +format, and what is and isn't covered. + ## Excluding Conversations Conversations containing this marker anywhere in their content will be archived but not indexed: diff --git a/cli/episodic-memory.js b/cli/episodic-memory.js index 08ebccc8..45c61da1 100755 --- a/cli/episodic-memory.js +++ b/cli/episodic-memory.js @@ -44,6 +44,7 @@ COMMANDS: stats Show index statistics doctor Diagnose Claude Code or Codex integration issues import-cursor-history Export legacy Cursor conversations from state.vscdb for indexing + redact Re-run secret redaction over the archive and index (--rewrite) Run 'episodic-memory --help' for command-specific help. @@ -94,6 +95,10 @@ async function main() { await runScript(join(distDir, 'cursor-import-cli.js'), args); break; + case 'redact': + await runScript(join(distDir, 'redact-cli.js'), args); + break; + case '--help': case '-h': case undefined: diff --git a/dist/cursor-legacy.d.ts b/dist/cursor-legacy.d.ts index a035ba0f..77340197 100644 --- a/dist/cursor-legacy.d.ts +++ b/dist/cursor-legacy.d.ts @@ -1,3 +1,4 @@ +import { type Redactor } from './redaction.js'; export declare function getDefaultCursorVscdbPath(): string | undefined; /** * Collect composer IDs that already have live agent transcripts under @@ -14,6 +15,8 @@ export interface CursorLegacyImportOptions { force?: boolean; /** Report what would be exported without writing files. */ dryRun?: boolean; + /** `undefined` loads from the environment (strict mode throws on bad rules); `null` disables. */ + redactor?: Redactor | null; } export interface CursorLegacyImportResult { exported: number; diff --git a/dist/cursor-legacy.js b/dist/cursor-legacy.js index 811e5637..e6ea46d6 100644 --- a/dist/cursor-legacy.js +++ b/dist/cursor-legacy.js @@ -3,6 +3,7 @@ import os from 'os'; import path from 'path'; import Database from 'better-sqlite3'; import { detectCursorCwd } from './parser.js'; +import { loadRedactor, redactJsonlLine } from './redaction.js'; /** * Import legacy Cursor conversations from Cursor's global SQLite store * (state.vscdb) into JSONL files compatible with the Cursor transcript parser. @@ -111,6 +112,7 @@ export function importCursorLegacy(options) { skippedEmpty: 0, errors: [], }; + const redactor = options.redactor === undefined ? loadRedactor() : options.redactor; const liveIds = options.liveTranscriptIds ?? new Set(); const db = new Database(options.dbPath, { readonly: true, fileMustExist: true }); try { @@ -200,9 +202,14 @@ export function importCursorLegacy(options) { if (!options.dryRun) { // Re-serialize with cwd now that it's known (it's derived from the // whole conversation's tool calls). - const finalLines = cwd + const withCwd = cwd ? lines.map(line => JSON.stringify({ ...JSON.parse(line), cwd })) : lines; + // The export dir is a plugin-owned plaintext copy, so redact it at + // write time like the archive (docs/REDACTION.md). + const finalLines = redactor + ? withCwd.map(line => redactJsonlLine(line, redactor, { source: 'cursor-legacy', path: outFile })) + : withCwd; fs.mkdirSync(path.dirname(outFile), { recursive: true }); fs.writeFileSync(outFile, finalLines.join('\n') + '\n', 'utf-8'); // Stamp the conversation's end time so mtime-based fallbacks and diff --git a/dist/db.d.ts b/dist/db.d.ts index 3f590fff..9d8ed98a 100644 --- a/dist/db.d.ts +++ b/dist/db.d.ts @@ -13,6 +13,12 @@ export declare function migrateSchema(db: Database.Database): void; * 3. Recreates the table with ON DELETE CASCADE and copies surviving rows. */ export declare function migrateToolCallsCascade(db: Database.Database): void; +/** + * Open the existing index read-only, without creating it, migrating it, or + * changing its journal mode. For dry runs that must not write. Returns null + * when there is no index yet. + */ +export declare function openDatabaseReadOnly(): Database.Database | null; export declare function initDatabase(): Database.Database; export declare function insertExchange(db: Database.Database, exchange: ConversationExchange, embedding: number[], toolNames?: string[]): void; export declare function getAllExchanges(db: Database.Database): Array<{ diff --git a/dist/db.js b/dist/db.js index 83f546ad..38fbcde0 100644 --- a/dist/db.js +++ b/dist/db.js @@ -90,6 +90,17 @@ export function migrateToolCallsCascade(db) { db.pragma('foreign_keys = ON'); console.log(' tool_calls migration complete.'); } +/** + * Open the existing index read-only, without creating it, migrating it, or + * changing its journal mode. For dry runs that must not write. Returns null + * when there is no index yet. + */ +export function openDatabaseReadOnly() { + const dbPath = getDbPath(); + if (!fs.existsSync(dbPath)) + return null; + return new Database(dbPath, { readonly: true, fileMustExist: true }); +} export function initDatabase() { const dbPath = getDbPath(); // Ensure directory exists diff --git a/dist/indexer.js b/dist/indexer.js index 8afa1d35..0fb4e2fc 100644 --- a/dist/indexer.js +++ b/dist/indexer.js @@ -7,6 +7,8 @@ import { summarizeConversation } from './summarizer.js'; import { getArchiveDir, getExcludedProjects, getConversationSourceDirs, findJsonlFiles, statIfExists } from './paths.js'; import { formatErrorSentinel, shouldQueueForSummary } from './summary-sentinel.js'; import { getMaxMessageBytes, isOversizeExchange } from './message-size.js'; +import { FindingsTally, formatFindings, loadRedactor } from './redaction.js'; +import { copyIfNewer } from './sync.js'; // Set max output tokens for Claude SDK (used by summarizer) process.env.CLAUDE_CODE_MAX_OUTPUT_TOKENS = '20000'; // Increase max listeners for concurrent API calls @@ -25,7 +27,18 @@ async function processBatch(items, processor, concurrency) { function sessionIdForSummary(exchanges) { return exchanges.find(exchange => exchange.sessionId)?.sessionId; } +// Resume/fork would summarize the unredacted source transcript; see sync.ts. +function summarizeOptions(redactor) { + return { allowResume: redactor === null }; +} +function logRedactions(tally) { + if (tally.total > 0) + console.log(` Redaction: ${formatFindings(tally.toArray())}`); +} export async function indexConversations(limitToProject, maxConversations, concurrency = 1, noSummaries = false) { + // Load before touching the archive: strict mode fails closed here. + const redactor = loadRedactor(); + const tally = new FindingsTally(); console.log('Initializing database...'); const db = initDatabase(); console.log('Loading embedding model...'); @@ -73,14 +86,13 @@ export async function indexConversations(limitToProject, maxConversations, concu // Source transcripts can vanish mid-run (Claude Code cleanup). Skip loudly. let exchanges; try { - // Copy to archive (ensure parent dirs exist for subagent files) - if (!fs.existsSync(archivePath)) { - fs.mkdirSync(path.dirname(archivePath), { recursive: true }); - fs.copyFileSync(sourcePath, archivePath); + // Copy (redacted) to the archive, then parse the archive, so the index + // and summaries only ever see redacted text. + if (copyIfNewer(sourcePath, archivePath, redactor, tally)) { console.log(` Archived: ${file}`); } // Parse conversation - exchanges = await parseConversation(sourcePath, project, archivePath); + exchanges = await parseConversation(archivePath, project, archivePath); } catch (error) { console.log(` Skipped ${file} (read failed: ${error instanceof Error ? error.message : error})`); @@ -105,7 +117,7 @@ export async function indexConversations(limitToProject, maxConversations, concu console.log(` Generating ${needsSummary.length} summaries (concurrency: ${concurrency})...`); await processBatch(needsSummary, async (conv) => { try { - const summary = await summarizeConversation(conv.exchanges, sessionIdForSummary(conv.exchanges)); + const summary = await summarizeConversation(conv.exchanges, sessionIdForSummary(conv.exchanges), summarizeOptions(redactor)); fs.writeFileSync(conv.summaryPath, summary, 'utf-8'); const wordCount = summary.split(/\s+/).length; console.log(` ✓ ${conv.file}: ${wordCount} words`); @@ -145,6 +157,7 @@ export async function indexConversations(limitToProject, maxConversations, concu // Check if we hit the limit if (maxConversations && conversationsProcessed >= maxConversations) { console.log(`\nReached limit of ${maxConversations} conversations`); + logRedactions(tally); db.close(); console.log(`✅ Indexing complete! Conversations: ${conversationsProcessed}, Exchanges: ${totalExchanges}`); return; @@ -155,11 +168,14 @@ export async function indexConversations(limitToProject, maxConversations, concu if (oversizeSkipped > 0) { console.log(` Skipped ${oversizeSkipped} oversize exchange(s) (> ${maxMessageBytes} bytes; set EPISODIC_MEMORY_MAX_MESSAGE_BYTES to change) — likely embedded-transcript payloads (#139)`); } + logRedactions(tally); db.close(); console.log(`\n✅ Indexing complete! Conversations: ${conversationsProcessed}, Exchanges: ${totalExchanges}`); } export async function indexSession(sessionId, concurrency = 1, noSummaries = false) { console.log(`Indexing session: ${sessionId}`); + const redactor = loadRedactor(); + const tally = new FindingsTally(); // Find the conversation file for this session const sourceDirs = getConversationSourceDirs(); const ARCHIVE_DIR = getArchiveDir(); @@ -187,11 +203,8 @@ export async function indexSession(sessionId, concurrency = 1, noSummaries = fal // Archive + parse — source may vanish mid-run (Claude Code cleanup). let exchanges; try { - if (!fs.existsSync(archivePath)) { - fs.mkdirSync(path.dirname(archivePath), { recursive: true }); - fs.copyFileSync(sourcePath, archivePath); - } - exchanges = await parseConversation(sourcePath, project, archivePath); + copyIfNewer(sourcePath, archivePath, redactor, tally); + exchanges = await parseConversation(archivePath, project, archivePath); } catch (error) { console.log(`Skipped ${file} (read failed: ${error instanceof Error ? error.message : error})`); @@ -204,7 +217,7 @@ export async function indexSession(sessionId, concurrency = 1, noSummaries = fal if (!noSummaries && shouldQueueForSummary(summaryPath)) { fs.mkdirSync(path.dirname(summaryPath), { recursive: true }); try { - const summary = await summarizeConversation(exchanges, sessionIdForSummary(exchanges)); + const summary = await summarizeConversation(exchanges, sessionIdForSummary(exchanges), summarizeOptions(redactor)); fs.writeFileSync(summaryPath, summary, 'utf-8'); console.log(`Summary: ${summary.split(/\s+/).length} words`); } @@ -235,6 +248,7 @@ export async function indexSession(sessionId, concurrency = 1, noSummaries = fal if (oversizeSkipped > 0) { console.log(` Skipped ${oversizeSkipped} oversize exchange(s) (> ${maxMessageBytes} bytes; set EPISODIC_MEMORY_MAX_MESSAGE_BYTES to change) — likely embedded-transcript payloads (#139)`); } + logRedactions(tally); console.log(`✅ Indexed session ${sessionId}: ${exchanges.length} exchanges`); } db.close(); @@ -254,6 +268,8 @@ export async function indexUnprocessed(concurrency = 1, noSummaries = false) { console.log(`Concurrency: ${concurrency}`); if (noSummaries) console.log('⚠️ Running in no-summaries mode (skipping AI summaries)'); + const redactor = loadRedactor(); + const tally = new FindingsTally(); const db = initDatabase(); await initEmbeddings(); const sourceDirs = getConversationSourceDirs(); @@ -280,15 +296,13 @@ export async function indexUnprocessed(concurrency = 1, noSummaries = false) { // Transcript JSONLs are append-only, so MAX(line_end) tells us where to resume. const hw = db.prepare('SELECT COALESCE(MAX(line_end), 0) as maxLine FROM exchanges WHERE archive_path = ?').get(archivePath); const maxIndexedLine = hw.maxLine; - // Ensure parent dirs exist for subagent files try { - fs.mkdirSync(path.dirname(archivePath), { recursive: true }); - // Refresh the archive when the source may have grown beyond what we've seen. - if (!fs.existsSync(archivePath) || maxIndexedLine > 0) { - fs.copyFileSync(sourcePath, archivePath); - } + // Refresh the (redacted) archive, then parse the archive so the index + // only sees redacted text. Force the refresh once the file is indexed: + // an append inside the mtime granularity would otherwise be skipped. + copyIfNewer(sourcePath, archivePath, redactor, tally, maxIndexedLine > 0); // Parse and filter to exchanges past the high-water mark - const exchanges = await parseConversation(sourcePath, project, archivePath); + const exchanges = await parseConversation(archivePath, project, archivePath); const newExchanges = maxIndexedLine > 0 ? exchanges.filter(e => e.lineStart > maxIndexedLine) : exchanges; @@ -303,6 +317,7 @@ export async function indexUnprocessed(concurrency = 1, noSummaries = false) { } } } // end sourceDir loop + logRedactions(tally); if (unprocessed.length === 0) { console.log('✅ All conversations are already processed!'); db.close(); @@ -316,7 +331,7 @@ export async function indexUnprocessed(concurrency = 1, noSummaries = false) { console.log(`Generating ${needsSummary.length} summaries (concurrency: ${concurrency})...\n`); await processBatch(needsSummary, async (conv) => { try { - const summary = await summarizeConversation(conv.exchanges, sessionIdForSummary(conv.exchanges)); + const summary = await summarizeConversation(conv.exchanges, sessionIdForSummary(conv.exchanges), summarizeOptions(redactor)); fs.writeFileSync(conv.summaryPath, summary, 'utf-8'); const wordCount = summary.split(/\s+/).length; console.log(` ✓ ${conv.project}/${conv.file}: ${wordCount} words`); diff --git a/dist/opencode-sync.d.ts b/dist/opencode-sync.d.ts index fff32c4c..64215225 100644 --- a/dist/opencode-sync.d.ts +++ b/dist/opencode-sync.d.ts @@ -1,3 +1,4 @@ +import { type Redactor } from './redaction.js'; export interface OpencodeExportResult { exported: number; skipped: number; @@ -15,4 +16,6 @@ export declare function getOpencodeTranscriptFilePath(transcriptDir: string, inp export declare function exportOpencodeSessions(options?: { dbPath?: string; transcriptDir?: string; + /** `undefined` loads from the environment (strict mode throws on bad rules); `null` disables. */ + redactor?: Redactor | null; }): OpencodeExportResult; diff --git a/dist/opencode-sync.js b/dist/opencode-sync.js index 2f90ab75..33916055 100644 --- a/dist/opencode-sync.js +++ b/dist/opencode-sync.js @@ -2,6 +2,7 @@ import fs from 'fs'; import path from 'path'; import Database from 'better-sqlite3'; import { getOpencodeDbPath, getOpencodeTranscriptDir } from './paths.js'; +import { loadRedactor, redactJsonlLine } from './redaction.js'; function safeParseJson(value) { if (!value) return undefined; @@ -40,7 +41,7 @@ function shouldExportSession(filePath, sessionUpdatedMs) { const stat = fs.statSync(filePath); return Math.floor(stat.mtimeMs) < Math.floor(sessionUpdatedMs); } -function writeSessionTranscript(db, session, filePath) { +function writeSessionTranscript(db, session, filePath, redactor) { const messages = db.prepare(` SELECT id, session_id, time_created, time_updated, data FROM message @@ -115,14 +116,20 @@ function writeSessionTranscript(db, session, filePath) { parts, })); } + // The staging transcript is a plugin-owned plaintext copy, so redact it at + // write time like the archive (docs/REDACTION.md). + const output = redactor + ? lines.map(line => redactJsonlLine(line, redactor, { source: 'opencode', path: filePath })) + : lines; fs.mkdirSync(path.dirname(filePath), { recursive: true }); const tempPath = `${filePath}.tmp.${process.pid}`; - fs.writeFileSync(tempPath, `${lines.join('\n')}\n`, 'utf-8'); + fs.writeFileSync(tempPath, `${output.join('\n')}\n`, 'utf-8'); fs.renameSync(tempPath, filePath); const mtime = dateFromMillis(session.time_updated); fs.utimesSync(filePath, mtime, mtime); } export function exportOpencodeSessions(options = {}) { + const redactor = options.redactor === undefined ? loadRedactor() : options.redactor; const dbPath = options.dbPath || getOpencodeDbPath(); const transcriptDir = options.transcriptDir || getOpencodeTranscriptDir(); const result = { @@ -168,7 +175,7 @@ export function exportOpencodeSessions(options = {}) { result.skipped++; continue; } - writeSessionTranscript(db, session, filePath); + writeSessionTranscript(db, session, filePath, redactor); result.exported++; } catch (error) { diff --git a/dist/redact-cli.d.ts b/dist/redact-cli.d.ts new file mode 100644 index 00000000..cb0ff5c3 --- /dev/null +++ b/dist/redact-cli.d.ts @@ -0,0 +1 @@ +export {}; diff --git a/dist/redact-cli.js b/dist/redact-cli.js new file mode 100644 index 00000000..ee142efa --- /dev/null +++ b/dist/redact-cli.js @@ -0,0 +1,134 @@ +import fs from 'fs'; +import { getArchiveDir, getCursorLegacyExportDir, getOpencodeTranscriptDir } from './paths.js'; +import { acquireFileLock, readLockHolder, releaseFileLock } from './file-lock.js'; +import { getSyncLockPath } from './logging.js'; +import { DEFAULT_REDACTION_CONFIG, FindingsTally, formatFindings, getRedactionSettings, loadRedactor, redactJsonlLine, } from './redaction.js'; +const args = process.argv.slice(2); +const HELP = ` +Usage: episodic-memory redact [--rewrite [--dry-run [--report]]] [--stdin] [--print-default-rules] + +Secret redaction for the conversation archive and index. + +OPTIONS: + --rewrite Re-run redaction over the existing archive, staging exports, + and search index (in place). Rows that change are re-embedded, + and summaries built from unredacted text are deleted (the next + sync regenerates them). Run once after upgrading, and again + after adding rules. + --dry-run With --rewrite: report what would change, write nothing. + The index is opened read-only and is not migrated. + --report With --rewrite --dry-run: list every value that would be + redacted: where it is, the rule, its shape (length, + character classes, entropy) and the redacted text around + it. Use it to spot false positives before applying. + --stdin Redact stdin to stdout, one JSONL/text line at a time, and + print rule counts to stderr. Handy for testing rules. + --print-default-rules Print the bundled rules as JSON (a starting point for + redaction-rules.json). + --help, -h Show this help + +ENVIRONMENT: + EPISODIC_MEMORY_REDACTION on (default) | off + EPISODIC_MEMORY_REDACTION_RULES custom rules file (default: /redaction-rules.json) + EPISODIC_MEMORY_REDACTION_STRICT 1 (default) fails closed on a bad rules file | 0 passes through + +Output names rule IDs, counts and value shapes only; matched values are never printed. +`; +function fail(message) { + console.error(`episodic-memory: ${message}`); + process.exit(1); +} +function requireRedactor() { + if (!getRedactionSettings().enabled) { + fail('redaction is off (EPISODIC_MEMORY_REDACTION=off); unset it to use this command.'); + } + let redactor; + try { + // Always strict here: a rewrite with no rules would be a silent no-op. + redactor = loadRedactor({ ...process.env, EPISODIC_MEMORY_REDACTION_STRICT: '1' }); + } + catch (error) { + fail(error instanceof Error ? error.message : String(error)); + } + return redactor; +} +async function runRewrite(dryRun, report) { + if (report && !dryRun) + fail('--report needs --dry-run: review the hits, then apply without it.'); + const redactor = requireRedactor(); + // Share sync's single-instance lock so a background sync can't write + // between our reads and renames. + const lockPath = getSyncLockPath(); + const lock = acquireFileLock(lockPath); + if (!lock) { + const holder = readLockHolder(lockPath); + fail(`a sync or index run is in progress (${holder !== null ? `pid ${holder}` : 'another process'}); try again when it finishes.`); + } + const release = () => releaseFileLock(lock); + process.on('exit', release); + const { rewriteArchive } = await import('./redact-rewrite.js'); + let embeddingsReady = false; + const embed = async (user, assistant, toolNames) => { + const embeddings = await import('./embeddings.js'); + if (!embeddingsReady) { + await embeddings.initEmbeddings(); + embeddingsReady = true; + } + return embeddings.generateExchangeEmbedding(user, assistant, toolNames); + }; + const archiveDir = getArchiveDir(); + const stagingDirs = [getOpencodeTranscriptDir(), getCursorLegacyExportDir()].filter(d => fs.existsSync(d)); + console.log(`Redacting${dryRun ? ' (dry run)' : ''}: ${archiveDir}`); + for (const dir of stagingDirs) + console.log(` + staging: ${dir}`); + const result = await rewriteArchive({ + archiveDir, + stagingDirs, + redactor, + embed, + dryRun, + report: report + ? hit => console.log(` ${hit.location} ${hit.ruleId} ${hit.shape} + ${hit.context}`) + : undefined, + log: message => console.log(` ${message}`), + }); + console.log(`\n${dryRun ? 'Would redact' : 'Redacted'} across archive, staging, and index: ${formatFindings(result.findings)}`); + if (dryRun && (result.filesRewritten || result.rowsUpdated || result.summariesRemoved || result.stagingFilesRewritten)) { + console.log('Run without --dry-run to apply.'); + } +} +async function runStdin() { + const redactor = requireRedactor(); + const chunks = []; + for await (const chunk of process.stdin) + chunks.push(chunk); + const input = Buffer.concat(chunks).toString('utf-8'); + const tally = new FindingsTally(); + const output = input.split('\n').map(line => redactJsonlLine(line, redactor, { source: 'stdin', path: '-' }, tally)); + process.stdout.write(output.join('\n')); + console.error(formatFindings(tally.toArray())); +} +async function main() { + if (args.length === 0 || args.includes('--help') || args.includes('-h')) { + console.log(HELP); + return; + } + if (args.includes('--print-default-rules')) { + console.log(JSON.stringify(DEFAULT_REDACTION_CONFIG, null, 2)); + return; + } + if (args.includes('--stdin')) { + await runStdin(); + return; + } + if (args.includes('--rewrite')) { + await runRewrite(args.includes('--dry-run'), args.includes('--report')); + return; + } + fail(`unknown option(s): ${args.join(' ')}. Try: episodic-memory redact --help`); +} +main().catch(error => { + console.error('Error:', error instanceof Error ? error.message : error); + process.exit(1); +}); diff --git a/dist/redact-rewrite.d.ts b/dist/redact-rewrite.d.ts new file mode 100644 index 00000000..c067931d --- /dev/null +++ b/dist/redact-rewrite.d.ts @@ -0,0 +1,56 @@ +import { type RedactionFinding, type Redactor } from './redaction.js'; +/** + * Backfill for `episodic-memory redact --rewrite`: re-run redaction over data + * written before redaction existed (or before a rule was added). + * + * 1. Archive: every .jsonl is redacted in place. Line count and mtime are + * preserved, so index line ranges stay valid and sync still sees the + * archive as current. + * 2. Staging dirs (opencode / legacy Cursor exports): same, in place. + * 3. Index: every exchanges/tool_calls row is redacted in place. Changed rows + * are re-embedded from the redacted text, so no vector is left that was + * derived from a secret. + * 4. Summaries: a `-summary.txt` for a conversation that had findings (in the + * archive or the index), or one that matches a rule itself, is deleted. + * It was generated from unredacted text, and the next sync regenerates it + * from the redacted archive. + * + * Idempotent: a second run finds nothing to change. A dry run opens the index + * read-only, so it doesn't create, migrate or otherwise touch it. + */ +export type EmbedFn = (user: string, assistant: string, toolNames?: string[]) => Promise; +export interface RewriteOptions { + archiveDir: string; + redactor: Redactor; + embed: EmbedFn; + /** Count what would change without writing anything. */ + dryRun?: boolean; + /** + * Called once per value that would be redacted, for reviewing hits before + * applying them. Never receives the value itself. + */ + report?: (hit: RedactionHit) => void; + /** Plugin-owned staging dirs to redact in place (not indexed). */ + stagingDirs?: string[]; + log?: (message: string) => void; +} +export interface RewriteResult { + filesScanned: number; + filesRewritten: number; + stagingFilesRewritten: number; + rowsUpdated: number; + summariesRemoved: number; + /** Rule IDs and counts only — never matched values. */ + findings: RedactionFinding[]; +} +/** One redacted value, described without revealing it. */ +export interface RedactionHit { + /** Archive-relative file and 1-based line, or `index:# `. */ + location: string; + ruleId: string; + /** Length, character classes and entropy (see describeShape). */ + shape: string; + /** Redacted text around the token, on one line. */ + context: string; +} +export declare function rewriteArchive(options: RewriteOptions): Promise; diff --git a/dist/redact-rewrite.js b/dist/redact-rewrite.js new file mode 100644 index 00000000..ae62e1d9 --- /dev/null +++ b/dist/redact-rewrite.js @@ -0,0 +1,263 @@ +import fs from 'fs'; +import path from 'path'; +import { initDatabase, openDatabaseReadOnly } from './db.js'; +import { recordReembedded } from './embedding-migration.js'; +import { copyFileRedacted, describeShape, findRedactionTokens, FindingsTally, redactJsonlLine, } from './redaction.js'; +const SUMMARY_SUFFIX = '-summary.txt'; +const CONTEXT_BEFORE = 60; +const CONTEXT_AFTER = 20; +const PAGE_SIZE = 500; +function walk(dir) { + const out = []; + let entries; + try { + entries = fs.readdirSync(dir, { withFileTypes: true }); + } + catch { + return out; + } + for (const entry of entries) { + const full = path.join(dir, entry.name); + if (entry.isDirectory()) + out.push(...walk(full)); + else if (entry.isFile()) + out.push(full); + } + return out; +} +function summaryPathFor(jsonlPath) { + return jsonlPath.replace(/\.jsonl$/, SUMMARY_SUFFIX); +} +// In report mode each hit is written as its own numbered token, e.g. +// `[REDACTED:jwt--hit3]`, so it can be found in the output exactly; tokens that +// were already in the text have no number. Numbers are stripped for display. +const HIT_TOKEN = /\[REDACTED:([a-z0-9][a-z0-9-]*?)--hit(\d+)\]/g; +/** `value` with any numbered tokens in it replaced by the values they stand for. */ +function restoreHits(value, matches) { + return value.replace(HIT_TOKEN, (_token, _ruleId, n) => matches[Number(n)].value); +} +/** + * Report each hit with the redacted text around its token. A hit whose token + * was swallowed by a later one (a secret field redacted whole after a text rule + * hit part of it) is reported at the token that swallowed it. + */ +function reportHits(redacted, matches, location, report) { + let display = ''; + let pos = 0; + const spans = new Map(); + for (const t of findRedactionTokens(redacted)) { + const numbered = /^(.*)--hit(\d+)$/.exec(t.ruleId); + display += redacted.slice(pos, t.start); + const shown = numbered ? `[REDACTED:${numbered[1]}]` : redacted.slice(t.start, t.end); + if (numbered && !spans.has(Number(numbered[2]))) { + spans.set(Number(numbered[2]), [display.length, display.length + shown.length]); + } + display += shown; + pos = t.end; + } + display += redacted.slice(pos); + const spanOf = (n) => { + for (let i = n; i !== undefined; i = matches[i].absorbedBy) { + const span = spans.get(i); + if (span) + return span; + } + return undefined; + }; + const oneLine = (text) => text.replace(/\s+/g, ' '); + const hits = matches.map((m, n) => ({ m, span: spanOf(n) })); + hits.sort((a, b) => (a.span?.[0] ?? Infinity) - (b.span?.[0] ?? Infinity)); + for (const { m, span } of hits) { + const [start, end] = span ?? [0, 0]; + const before = display.slice(Math.max(0, start - CONTEXT_BEFORE), start); + report({ + location, + ruleId: m.ruleId, + shape: describeShape(m.value), + context: span + ? oneLine(`${start > CONTEXT_BEFORE ? '…' : ''}${before}${display.slice(start, end + CONTEXT_AFTER)}`) + : '', + }); + } + return display; +} +/** + * Redact with `run`, reporting each hit when `report` is set. Returns + * the redacted text with ordinary tokens either way. + */ +function redactReporting(ctx, location, report, run) { + if (!report) + return run(ctx); + const matches = []; + const out = run({ + ...ctx, + onMatch: (ruleId, value) => { + const n = matches.length; + for (const [, , k] of value.matchAll(HIT_TOKEN)) + matches[Number(k)].absorbedBy = n; + matches.push({ ruleId, value: restoreHits(value, matches) }); + return `[REDACTED:${ruleId}--hit${n}]`; + }, + }); + return matches.length > 0 ? reportHits(out, matches, location, report) : out; +} +/** Report every hit in a JSONL file, line by line. Writes nothing. */ +function reportFile(file, label, redactor, report) { + const lines = fs.readFileSync(file, 'utf-8').split('\n'); + const ctx = { source: 'rewrite', path: file }; + lines.forEach((line, i) => { + redactReporting(ctx, `${label}:${i + 1}`, report, c => redactJsonlLine(line, redactor, c)); + }); +} +/** Redact one JSONL file in place. Returns the number of values redacted. */ +function rewriteFileInPlace(file, redactor, dryRun, tally) { + const fileTally = new FindingsTally(); + const temp = `${file}.redact.${process.pid}`; + try { + copyFileRedacted(file, temp, redactor, { source: 'rewrite', path: file }, fileTally); + if (fileTally.total > 0 && !dryRun) { + const stat = fs.statSync(file); + fs.renameSync(temp, file); + // Same rounding as sync's copyIfNewer: never leave the archive older than its source. + fs.utimesSync(file, stat.atimeMs / 1000, Math.ceil(stat.mtimeMs) / 1000); + } + } + finally { + try { + fs.unlinkSync(temp); + } + catch { } + } + tally.add(fileTally.toArray()); + return fileTally.total; +} +export async function rewriteArchive(options) { + const { archiveDir, redactor, embed } = options; + const dryRun = options.dryRun === true; + // Reporting reads files before redaction; after a real rewrite there'd be nothing to find. + if (options.report && !dryRun) + throw new Error('rewriteArchive: report requires dryRun'); + const log = options.log ?? (() => { }); + const tally = new FindingsTally(); + const result = { + filesScanned: 0, + filesRewritten: 0, + stagingFilesRewritten: 0, + rowsUpdated: 0, + summariesRemoved: 0, + findings: [], + }; + // Conversations whose summary was built from unredacted text. + const staleSummaries = new Set(); + // 1. Archive files. + const archiveFiles = walk(archiveDir); + for (const file of archiveFiles.filter(f => f.endsWith('.jsonl'))) { + result.filesScanned++; + if (rewriteFileInPlace(file, redactor, dryRun, tally) > 0) { + result.filesRewritten++; + staleSummaries.add(summaryPathFor(file)); + if (options.report) + reportFile(file, path.relative(archiveDir, file), redactor, options.report); + } + } + log(`Archive: ${result.filesRewritten} of ${result.filesScanned} file(s) ${dryRun ? 'would be ' : ''}rewritten`); + // 2. Staging dirs. + for (const dir of options.stagingDirs ?? []) { + for (const file of walk(dir).filter(f => f.endsWith('.jsonl'))) { + if (rewriteFileInPlace(file, redactor, dryRun, tally) > 0) { + result.stagingFilesRewritten++; + if (options.report) + reportFile(file, `staging:${path.relative(dir, file)}`, redactor, options.report); + } + } + } + if (options.stagingDirs?.length) { + log(`Staging exports: ${result.stagingFilesRewritten} file(s) ${dryRun ? 'would be ' : ''}rewritten`); + } + // 3. Index rows. Page by rowid so writes between pages don't disturb the scan. + // A dry run reads the index as it is; opening it normally would migrate it. + const db = dryRun ? openDatabaseReadOnly() : initDatabase(); + if (db) + try { + const page = db.prepare('SELECT rowid AS rid, id, user_message, assistant_message, archive_path FROM exchanges WHERE rowid > ? ORDER BY rowid LIMIT ?'); + const toolsFor = db.prepare('SELECT id, tool_name, tool_input, tool_result FROM tool_calls WHERE exchange_id = ? ORDER BY rowid'); + const updateExchange = db.prepare('UPDATE exchanges SET user_message = ?, assistant_message = ? WHERE id = ?'); + const updateTool = db.prepare('UPDATE tool_calls SET tool_input = ?, tool_result = ? WHERE id = ?'); + let lastRowid = 0; + for (;;) { + const rows = page.all(lastRowid, PAGE_SIZE); + if (rows.length === 0) + break; + lastRowid = rows[rows.length - 1].rid; + for (const row of rows) { + const rowTally = new FindingsTally(); + const ctx = { source: 'index', path: row.archive_path }; + const where = `index:${path.relative(archiveDir, row.archive_path)}#${row.id}`; + const redactText = (text, field) => redactReporting(ctx, `${where} ${field}`, options.report, c => { + const r = redactor.redact(text, c); + rowTally.add(r.findings); + return r.text; + }); + const user = { text: redactText(row.user_message, 'user') }; + const assistant = { text: redactText(row.assistant_message, 'assistant') }; + const tools = toolsFor.all(row.id); + const toolUpdates = []; + for (const tool of tools) { + // tool_input is JSON text; redactJsonlLine keeps it valid JSON. + const toolInput = tool.tool_input; + const input = toolInput === null ? null : redactReporting(ctx, `${where} ${tool.tool_name} input`, options.report, c => redactJsonlLine(toolInput, redactor, c, rowTally)); + const output = tool.tool_result === null ? null : redactText(tool.tool_result, `${tool.tool_name} result`); + if (input !== tool.tool_input || output !== tool.tool_result) { + toolUpdates.push({ id: tool.id, input, result: output }); + } + } + if (rowTally.total === 0) + continue; + tally.add(rowTally.toArray()); + result.rowsUpdated++; + staleSummaries.add(summaryPathFor(row.archive_path)); + if (dryRun) + continue; + const toolNames = tools.length > 0 ? tools.map(t => t.tool_name) : undefined; + const embedding = await embed(user.text, assistant.text, toolNames); + db.transaction(() => { + updateExchange.run(user.text, assistant.text, row.id); + for (const t of toolUpdates) + updateTool.run(t.input, t.result, t.id); + recordReembedded(db, row.id, embedding); + })(); + } + } + } + finally { + db.close(); + } + log(`Index: ${result.rowsUpdated} exchange(s) ${dryRun ? 'would be ' : ''}redacted and re-embedded`); + // 4. Summaries: stale ones, plus any summary that itself matches a rule. + for (const file of archiveFiles.filter(f => f.endsWith(SUMMARY_SUFFIX))) { + if (staleSummaries.has(file)) + continue; + let text; + try { + text = fs.readFileSync(file, 'utf-8'); + } + catch { + continue; + } + const r = redactor.redact(text, { source: 'summary', path: file }); + if (r.findings.length > 0) { + tally.add(r.findings); + staleSummaries.add(file); + } + } + for (const summary of staleSummaries) { + if (!fs.existsSync(summary)) + continue; + result.summariesRemoved++; + if (!dryRun) + fs.unlinkSync(summary); + } + log(`Summaries: ${result.summariesRemoved} ${dryRun ? 'would be ' : ''}removed (regenerated from redacted text on the next sync)`); + result.findings = tally.toArray(); + return result; +} diff --git a/dist/redaction-rules.d.ts b/dist/redaction-rules.d.ts new file mode 100644 index 00000000..d9c5ad24 --- /dev/null +++ b/dist/redaction-rules.d.ts @@ -0,0 +1,2 @@ +import type { RedactionConfig } from './redaction.js'; +export declare const DEFAULT_REDACTION_CONFIG: RedactionConfig; diff --git a/dist/redaction-rules.js b/dist/redaction-rules.js new file mode 100644 index 00000000..6295381c --- /dev/null +++ b/dist/redaction-rules.js @@ -0,0 +1,254 @@ +/** + * Bundled default redaction rules. + * + * Patterns are ported from gitleaks' rule set (https://github.com/gitleaks/gitleaks, + * MIT), trimmed to the credentials that realistically show up in Claude Code / + * Codex transcripts for an Azure-heavy stack, and rewritten for JavaScript + * regex (lookbehind instead of gitleaks' consuming boundary groups, so + * adjacent matches aren't swallowed). + * + * Order matters: rules run top to bottom, and a later rule never re-matches + * inside an earlier rule's `[REDACTED:...]` token. Specific shapes go first so + * the token names the most precise rule; the generic `secret-assignment` + * keyword rule goes last. + * + * `secretGroup` replaces only that capture group, so surrounding context + * (connection-string server names, SAS URL paths, usernames) stays searchable. + * + * `keywords` is a case-insensitive prefilter: a rule only runs on text that + * contains at least one keyword. It keeps the per-string cost low on large + * transcripts and bounds the generic rules' work. + * + * Users extend or override these with `redaction-rules.json`; see + * docs/REDACTION.md. `episodic-memory redact --print-default-rules` dumps + * this object as JSON. + */ +// Names that mark the value next to them as a secret, for the key-context +// rules below. A name must END in one of these (optionally plus digits), so +// tokenType, secretName, passwordPolicy, TokenEndpoint and maxTokens don't +// count. `pwd` is deliberately absent (the PWD env var); `userPWD` is handled +// by xml-secret-attribute. +const SECRET_NAME = String.raw `(?:secret|password|passwd|passphrase|api[_-]?key|access[_-]?key|account[_-]?key|private[_-]?key|` + + String.raw `shared[_-]?(?:access[_-]?)?key|primary[_-]?key|secondary[_-]?key|master[_-]?key|signing[_-]?key|` + + String.raw `subscription[_-]?key|client[_-]?key|encryption[_-]?key|token|credentials?)\d*`; +// Value is not a template/placeholder: ${X} $(X) $X #{X}# {{x}} %X% __X__, +// an existing token, a type name (Swagger "string"), or a mask (*****). +const NOT_PLACEHOLDER = String.raw `(?!\[REDACTED:|\$\{|\$\(|\$[A-Za-z_]\w*["'<\s]|#\{|\{\{|<[^>]*>|%[A-Za-z_]\w*%|__[A-Za-z0-9_]+__|` + + String.raw `(?:string|null|true|false|none|undefined|\*+)["'<])`; +// A JSON string value (escapes allowed), captured. +const JSON_STRING_VALUE = String.raw `"${NOT_PLACEHOLDER}([^"\\]+(?:\\.[^"\\]*)*)"`; +// Sibling members inside the same object: strings, or anything but braces/quotes. +const SAME_OBJECT = String.raw `(?:[^{}"]|"(?:[^"\\]|\\.)*"){0,400}?`; +const KEY_VAULT_ID = String.raw `"id"[ \t]*:[ \t]*"https://[^"\s]+\.vault\.(?:azure\.net|azure\.cn|usgovcloudapi\.net|microsoftazure\.de)/secrets/[^"\s]*"`; +export const DEFAULT_REDACTION_CONFIG = { + rules: [ + { + id: 'private-key-block', + description: 'PEM/OpenSSH private key block; a truncated block (no END line) is redacted to the end of the text', + pattern: String.raw `-----BEGIN[A-Z0-9 _-]{0,40}PRIVATE KEY(?: BLOCK)?-----[\s\S]*?(?:-----END[A-Z0-9 _-]{0,40}PRIVATE KEY(?: BLOCK)?-----|$)`, + keywords: ['private key'], + }, + { + id: 'connection-string-secret', + description: 'Secret value inside an Azure/ADO.NET connection string; account, server and database names are kept. Keywords are case-insensitive, as in ADO.NET; a path after Pwd= is the shell PWD variable and is kept; the value needs a letter or digit, so Markdown like `Password=` is kept', + pattern: String.raw `\b(?:(?:AccountKey|SharedAccessKey|SharedAccessSignature|SharedSecret|ClientSecret|Password)=|Pwd=(?![/~]))(?![\[$<{%])(?=[^;"'\s]*[A-Za-z0-9])([^;"'\s]+)`, + flags: 'i', + secretGroup: 1, + keywords: ['accountkey', 'sharedaccess', 'sharedsecret', 'clientsecret', 'password', 'pwd'], + }, + { + id: 'azure-sas-token', + description: 'Signature of an Azure SAS token; the resource URL and other SAS parameters are kept', + pattern: String.raw `\bsig=([A-Za-z0-9%+/=_-]{16,})`, + secretGroup: 1, + keywords: ['sig='], + }, + { + id: 'jwt', + description: 'JSON Web Token (Entra ID / Azure access tokens, id tokens)', + pattern: String.raw `\beyJ[A-Za-z0-9_-]{8,}\.eyJ[A-Za-z0-9_-]{8,}\.[A-Za-z0-9_-]{8,}`, + keywords: ['eyj'], + }, + { + id: 'anthropic-api-key', + pattern: String.raw `\bsk-ant-[a-z]{2,10}\d{2}-[A-Za-z0-9_-]{32,}`, + keywords: ['sk-ant-'], + }, + { + id: 'openai-api-key', + pattern: String.raw `\bsk-(?:[a-z]+-)?[A-Za-z0-9_-]{16,}T3BlbkFJ[A-Za-z0-9_-]{16,}`, + keywords: ['t3blbkfj'], + }, + { + id: 'github-token', + description: 'GitHub classic, OAuth, app and fine-grained tokens', + pattern: String.raw `\b(?:gh[pousr]_[A-Za-z0-9]{36,255}|github_pat_[A-Za-z0-9_]{22,255})\b`, + keywords: ['ghp_', 'gho_', 'ghu_', 'ghs_', 'ghr_', 'github_pat_'], + }, + { + id: 'aws-access-key-id', + pattern: String.raw `\b(?:AKIA|ASIA|ABIA|ACCA)[A-Z2-7]{16}\b`, + keywords: ['akia', 'asia', 'abia', 'acca'], + }, + { + id: 'aws-secret-access-key', + pattern: String.raw `aws[_-]?secret[_-]?access[_-]?key["']?[ \t]*[:=][ \t]*["']?([A-Za-z0-9/+]{40})(?![A-Za-z0-9/+])`, + flags: 'i', + secretGroup: 1, + keywords: ['secret'], + }, + { + id: 'slack-token', + pattern: String.raw `\bxox[abposr]-[A-Za-z0-9-]{10,}`, + keywords: ['xox'], + }, + { + id: 'google-api-key', + pattern: String.raw `\bAIza[0-9A-Za-z_-]{35}(?![0-9A-Za-z_-])`, + keywords: ['aiza'], + }, + { + id: 'npm-token', + pattern: String.raw `\bnpm_[A-Za-z0-9]{36}\b`, + keywords: ['npm_'], + }, + { + id: 'azure-client-secret', + description: 'Entra ID (Azure AD) application client secret: 3 chars, a digit, "Q~", 31-34 chars. A trailing "." may follow, so a secret ending a sentence still matches', + pattern: String.raw `(?]+:(?![\[$<{%])([^\s@/"'<>]+)@`, + secretGroup: 1, + keywords: ['://'], + }, + { + id: 'bearer-token', + pattern: String.raw `\bBearer\s+([A-Za-z0-9\-._~+/]{20,}=*)`, + flags: 'i', + secretGroup: 1, + keywords: ['bearer'], + }, + { + id: 'basic-auth', + pattern: String.raw `\bAuthorization[ \t]*:[ \t]*Basic\s+([A-Za-z0-9+/]{8,}={0,2})`, + flags: 'i', + secretGroup: 1, + keywords: ['basic'], + }, + { + id: 'azure-keyvault-secret', + description: 'The "value" of a Key Vault secret bundle (az keyvault secret show/set, SDK JSON), whatever the secret is named', + pattern: KEY_VAULT_ID + String.raw `(?:[^{}]|\{[^{}]*\}){0,800}?"value"[ \t]*:[ \t]*` + JSON_STRING_VALUE + + String.raw `|"value"[ \t]*:[ \t]*` + JSON_STRING_VALUE + String.raw `(?=(?:[^{}]|\{[^{}]*\}){0,800}?` + KEY_VAULT_ID + ')', + secretGroup: [1, 2], + useAllowlist: false, + keywords: ['.vault.'], + }, + { + id: 'name-value-secret', + description: 'The "value" of a {"name"/"key": , "value": ...} object, in either order: ' + + 'az webapp/functionapp config appsettings list, Kubernetes env, ARM/Bicep parameters', + pattern: String.raw `"(?:name|key)"[ \t]*:[ \t]*"[^"\\]{0,80}?` + SECRET_NAME + '"' + SAME_OBJECT + + String.raw `"value"[ \t]*:[ \t]*` + JSON_STRING_VALUE + + String.raw `|"value"[ \t]*:[ \t]*` + JSON_STRING_VALUE + '(?=' + SAME_OBJECT + + String.raw `"(?:name|key)"[ \t]*:[ \t]*"[^"\\]{0,80}?` + SECRET_NAME + '")', + flags: 'i', + secretGroup: [1, 2], + useAllowlist: false, + keywords: ['"value"'], + }, + { + id: 'xml-appsettings-secret', + description: 'web.config / app.config , either attribute order', + pattern: String.raw `]*?\bkey[ \t]*=[ \t]*"[^"]{0,80}?` + SECRET_NAME + String.raw `"[^>]*?\bvalue[ \t]*=[ \t]*"` + + NOT_PLACEHOLDER + String.raw `([^"]+)"` + + String.raw `|]*?\bvalue[ \t]*=[ \t]*"` + NOT_PLACEHOLDER + String.raw `([^"]+)"(?=[^>]*?\bkey[ \t]*=[ \t]*"[^"]{0,80}?` + + SECRET_NAME + '")', + flags: 'i', + secretGroup: [1, 2], + useAllowlist: false, + keywords: [']*)?>` + NOT_PLACEHOLDER + String.raw `([^<]{1,4096})`, + flags: 'i', + secretGroup: 2, + useAllowlist: false, + keywords: [', %X%, {{x}}) and code (calls, dotted member access)', + pattern: String.raw `\b[A-Za-z0-9_.-]{0,40}?(?:secret|password|passwd|passphrase|api[_-]?key|access[_-]?key|private[_-]?key|token)` + + String.raw `["']?[ \t]*[:=][ \t]*["']?(?![\[$<{%])(?=[^\s"',;&\\]*\d)` + + // Not a dotted identifier path (env.AUTH0_SECRET, this.config.token2): code, not a value. + String.raw `(?![A-Za-z_$][\w$]*(?:\.[A-Za-z_$][\w$]*)+(?:$|[\s"',;&\\)}\]]))` + + String.raw `([^\s"',;&\\()<>{}\[\]]{8,})(?=$|[\s"',;&\\)}\]])`, + flags: 'i', + secretGroup: 1, + keywords: ['secret', 'passw', 'passphrase', 'key', 'token'], + }, + ], + allowlist: [ + { + id: 'git-sha', + description: 'Full or short git commit SHA', + pattern: String.raw `\b[0-9a-f]{7,40}\b`, + }, + { + id: 'guid', + description: 'GUID/UUID (Entra tenant, client and object IDs, subscription IDs)', + pattern: String.raw `\b[0-9a-fA-F]{8}-[0-9a-fA-F]{4}-[0-9a-fA-F]{4}-[0-9a-fA-F]{4}-[0-9a-fA-F]{12}\b`, + }, + ], + entropy: { + enabled: false, + minLength: 32, + threshold: 4.5, + requireKeyword: true, + keywords: ['secret', 'key', 'token', 'password', 'passwd', 'credential', 'signature'], + window: 40, + }, + secretFields: { + enabled: true, + // Normalized key (lowercase letters/digits only, trailing digits dropped) + // must end with one of these. Harness keys (signature, apiKeySource, + // input_tokens, token_count) don't. + keyPattern: 'secret|password|passwd|userpwd|passphrase|apikey|accesskey|accountkey|privatekey|sharedkey|primarykey|' + + 'secondarykey|masterkey|signingkey|subscriptionkey|clientkey|encryptionkey|token|credentials?', + }, +}; diff --git a/dist/redaction.d.ts b/dist/redaction.d.ts new file mode 100644 index 00000000..d4af59bd --- /dev/null +++ b/dist/redaction.d.ts @@ -0,0 +1,180 @@ +import { DEFAULT_REDACTION_CONFIG } from './redaction-rules.js'; +/** + * Inline secret redaction. + * + * Secrets are replaced with typed tokens (`[REDACTED:]`) at the archive + * write, the one point every harness's transcripts pass through (see + * docs/REDACTION.md). Everything downstream (SQLite text, + * tool_calls, embeddings, summaries, show/read) reads the archive, so it only + * ever sees redacted text. + * + * Invariants: + * - Findings carry rule IDs and counts only. A matched value is never logged, + * returned, or thrown. + * - Idempotent: a `[REDACTED:...]` token never matches a rule, so redacting + * redacted text is a no-op. + * - JSONL structure is preserved: one output line per input line, valid JSON in + * gives valid JSON out, and unchanged lines are byte-identical. + */ +export { DEFAULT_REDACTION_CONFIG }; +/** Fork default. Flip to false for an opt-in upstream build. */ +export declare const REDACTION_ENABLED_BY_DEFAULT = true; +export declare const REDACTION_RULES_FILENAME = "redaction-rules.json"; +export interface RedactionRuleSpec { + id: string; + pattern: string; + /** Extra RegExp flags from [imsu]; `g` and `d` are always added. */ + flags?: string; + /** Case-insensitive prefilter: skip the rule unless the text contains one. */ + keywords?: string[]; + /** + * Replace only this capture group instead of the whole match. An array means + * "the first of these groups that participated" (for either-order patterns). + */ + secretGroup?: number | number[]; + /** Apply the allowlist to this rule's matches (default true). Key-context rules turn it off: a GUID in a password slot is a secret. */ + useAllowlist?: boolean; + description?: string; +} +export interface AllowlistSpec { + id: string; + /** Matched against the whole candidate secret; a full match keeps it. */ + pattern: string; + flags?: string; + description?: string; +} +export interface EntropySpec { + enabled: boolean; + minLength: number; + /** Shannon entropy in bits per character. */ + threshold: number; + /** Only fire when a keyword appears shortly before the candidate, on the same line. */ + requireKeyword: boolean; + keywords: string[]; + /** How many characters before the candidate to search for a keyword. */ + window: number; +} +/** + * Field-name context for parsed JSON (transcript lines, structured tool/MCP + * results, tool inputs). A string value is redacted whole when its own key, or + * the `name`/`key` of a `{name, value}` pair, matches `keyPattern`. + */ +export interface SecretFieldsSpec { + enabled: boolean; + /** + * Matched (anchored at the end) against the key lowercased with everything + * but letters and digits removed and trailing digits dropped, so + * `AzureAd:ClientSecret`, `client_secret` and `DB_PASSWORD2` all normalize to + * something ending in a keyword. + */ + keyPattern: string; +} +export interface RedactionConfig { + rules: RedactionRuleSpec[]; + allowlist: AllowlistSpec[]; + entropy: EntropySpec; + secretFields: SecretFieldsSpec; +} +/** Shape of a user `redaction-rules.json`. Every field is optional. */ +export interface RedactionRulesFile { + /** Merge with the bundled defaults (default true). */ + includeDefaults?: boolean; + /** Added after the defaults; a rule with a default's id replaces it in place. */ + rules?: RedactionRuleSpec[]; + /** Default rule ids to turn off. */ + disableRules?: string[]; + /** Added to the default allowlist (same id replaces). */ + allowlist?: AllowlistSpec[]; + entropy?: Partial; + secretFields?: Partial; +} +export interface RedactionContext { + source: string; + path: string; + /** + * Called with each value as it is redacted. For local review tooling + * (`redact --report`) only: the value is the secret itself, so never log it. + * May return the token to write instead of `[REDACTED:]`, e.g. one + * numbered per hit; it must still have the token's shape. + */ + onMatch?: (ruleId: string, value: string) => string | void; +} +export interface RedactionFinding { + ruleId: string; + count: number; +} +export interface RedactionResult { + text: string; + findings: RedactionFinding[]; +} +export interface Redactor { + redact(text: string, ctx?: RedactionContext): RedactionResult; + readonly ruleIds: string[]; + /** True when a JSON key / setting name marks its value as a secret (secretFields). */ + isSecretField(name: string): boolean; +} +export interface RedactionSettings { + enabled: boolean; + strict: boolean; + /** Explicit EPISODIC_MEMORY_REDACTION_RULES path, if set. */ + rulesPath?: string; +} +/** Rules could not be loaded or are invalid. Messages describe config only. */ +export declare class RedactionConfigError extends Error { + constructor(message: string); +} +/** + * EPISODIC_MEMORY_REDACTION on (fork default) | off + * EPISODIC_MEMORY_REDACTION_RULES path to a custom rules file + * EPISODIC_MEMORY_REDACTION_STRICT 1 (default) | 0; strict fails closed on a rules-load error + */ +export declare function getRedactionSettings(env?: NodeJS.ProcessEnv): RedactionSettings; +/** + * Merge a rules file with the bundled defaults and validate the result. + * Throws RedactionConfigError on any invalid rule. + */ +export declare function loadRedactionConfig(file: RedactionRulesFile): RedactionConfig; +/** + * Load the active redactor from the environment. + * + * Returns null when redaction is off. When the rules can't be loaded: strict + * mode (the default) throws RedactionConfigError so callers fail closed; + * non-strict mode warns and returns null, so text passes through unredacted. + */ +export declare function loadRedactor(env?: NodeJS.ProcessEnv, warn?: (message: string) => void): Redactor | null; +/** + * Describe a value without revealing it: length, character classes and + * Shannon entropy (bits per char), e.g. "len=40 aA9- H=4.9". Lets someone + * reviewing hits tell a random key from a word or a placeholder. + */ +export declare function describeShape(value: string): string; +/** Each `[REDACTED:]` token in `text`, left to right. */ +export declare function findRedactionTokens(text: string): Array<{ + ruleId: string; + start: number; + end: number; +}>; +export declare function createRedactor(config: RedactionConfig): Redactor; +/** Aggregates findings by rule id. Never holds matched values. */ +export declare class FindingsTally { + private counts; + add(findings: RedactionFinding[]): void; + get total(): number; + toArray(): RedactionFinding[]; +} +/** "7 value(s) redacted (azure-storage-key: 2, jwt: 5)" */ +export declare function formatFindings(findings: RedactionFinding[]): string; +/** + * Redact one JSONL line. Valid JSON stays valid (string values are redacted + * and the line is re-serialized only if something changed); a line that isn't + * JSON (e.g. a torn last line mid-write) is redacted as raw text. A trailing + * `\r` is preserved. + */ +export declare function redactJsonlLine(line: string, redactor: Redactor, ctx?: RedactionContext, tally?: FindingsTally): string; +/** + * Stream `src` to `dest`, redacting line by line. Writes `dest` directly; + * callers that need atomicity write to a temp path and rename. Line count and + * trailing-newline shape match the source exactly, so archive line numbers + * (exchanges.line_start/line_end, MCP read ranges) stay valid. + */ +export declare function copyFileRedacted(src: string, dest: string, redactor: Redactor, ctx?: RedactionContext, tally?: FindingsTally): void; diff --git a/dist/redaction.js b/dist/redaction.js new file mode 100644 index 00000000..e4f206d5 --- /dev/null +++ b/dist/redaction.js @@ -0,0 +1,659 @@ +import fs from 'fs'; +import path from 'path'; +import { StringDecoder } from 'string_decoder'; +import { getSuperpowersDir } from './paths.js'; +import { DEFAULT_REDACTION_CONFIG } from './redaction-rules.js'; +/** + * Inline secret redaction. + * + * Secrets are replaced with typed tokens (`[REDACTED:]`) at the archive + * write, the one point every harness's transcripts pass through (see + * docs/REDACTION.md). Everything downstream (SQLite text, + * tool_calls, embeddings, summaries, show/read) reads the archive, so it only + * ever sees redacted text. + * + * Invariants: + * - Findings carry rule IDs and counts only. A matched value is never logged, + * returned, or thrown. + * - Idempotent: a `[REDACTED:...]` token never matches a rule, so redacting + * redacted text is a no-op. + * - JSONL structure is preserved: one output line per input line, valid JSON in + * gives valid JSON out, and unchanged lines are byte-identical. + */ +export { DEFAULT_REDACTION_CONFIG }; +/** Fork default. Flip to false for an opt-in upstream build. */ +export const REDACTION_ENABLED_BY_DEFAULT = true; +export const REDACTION_RULES_FILENAME = 'redaction-rules.json'; +const RULE_ID_PATTERN = /^[a-z0-9][a-z0-9-]*$/; +const TOKEN_PATTERN = /\[REDACTED:[a-z0-9][a-z0-9-]*\]/g; +const TOKEN_PREFIX = '[REDACTED:'; +const ALLOWED_FLAGS = /^[imsu]*$/; +const ENTROPY_RULE_ID = 'high-entropy'; +const FIELD_RULE_ID = 'secret-field'; +const RESERVED_IDS = new Set([ENTROPY_RULE_ID, FIELD_RULE_ID]); +/** Values that are templates or type names, not secrets. Never redacted by field context. */ +const PLACEHOLDER_VALUE = /^(?:\$\{.*\}|\$\(.*\)|\$[A-Za-z_]\w*|#\{.*\}#?|\{\{.*\}\}|<[^>]*>|%[A-Za-z_]\w*%|__[A-Za-z0-9_]+__|string|null|true|false|none|undefined|\*+)$/i; +const KEY_VAULT_SECRET_ID = /^https:\/\/[^/\s]+\.vault\.(?:azure\.net|azure\.cn|usgovcloudapi\.net|microsoftazure\.de)\/secrets\//i; +/** Rules could not be loaded or are invalid. Messages describe config only. */ +export class RedactionConfigError extends Error { + constructor(message) { + super(message); + this.name = 'RedactionConfigError'; + } +} +// --------------------------------------------------------------------------- +// Settings and loading +// --------------------------------------------------------------------------- +const OFF_VALUES = new Set(['off', '0', 'false', 'no', 'disabled']); +const ON_VALUES = new Set(['on', '1', 'true', 'yes', 'enabled']); +/** Unknown values fall back to the default, which is the safe side for both switches. */ +function parseToggle(raw, fallback) { + const value = raw?.trim().toLowerCase(); + if (!value) + return fallback; + if (OFF_VALUES.has(value)) + return false; + if (ON_VALUES.has(value)) + return true; + return fallback; +} +/** + * EPISODIC_MEMORY_REDACTION on (fork default) | off + * EPISODIC_MEMORY_REDACTION_RULES path to a custom rules file + * EPISODIC_MEMORY_REDACTION_STRICT 1 (default) | 0; strict fails closed on a rules-load error + */ +export function getRedactionSettings(env = process.env) { + return { + enabled: parseToggle(env.EPISODIC_MEMORY_REDACTION, REDACTION_ENABLED_BY_DEFAULT), + strict: parseToggle(env.EPISODIC_MEMORY_REDACTION_STRICT, true), + rulesPath: env.EPISODIC_MEMORY_REDACTION_RULES || undefined, + }; +} +function readRulesFile(settings) { + let rulesPath = settings.rulesPath; + if (!rulesPath) { + const candidate = path.join(getSuperpowersDir(), REDACTION_RULES_FILENAME); + if (!fs.existsSync(candidate)) + return {}; + rulesPath = candidate; + } + let raw; + try { + raw = fs.readFileSync(rulesPath, 'utf-8'); + } + catch (error) { + throw new RedactionConfigError(`Redaction rules failed to load from ${rulesPath}: ${error instanceof Error ? error.message : String(error)}`); + } + try { + return JSON.parse(raw); + } + catch (error) { + throw new RedactionConfigError(`Redaction rules file ${rulesPath} is not valid JSON: ${error instanceof Error ? error.message : String(error)}`); + } +} +function replaceById(base, additions) { + const out = [...base]; + for (const item of additions) { + const at = out.findIndex(existing => existing.id === item.id); + if (at >= 0) + out[at] = item; + else + out.push(item); + } + return out; +} +/** + * Merge a rules file with the bundled defaults and validate the result. + * Throws RedactionConfigError on any invalid rule. + */ +export function loadRedactionConfig(file) { + if (!file || typeof file !== 'object' || Array.isArray(file)) { + throw new RedactionConfigError('Redaction rules file must contain a JSON object'); + } + for (const key of ['rules', 'allowlist', 'disableRules']) { + if (file[key] !== undefined && !Array.isArray(file[key])) { + throw new RedactionConfigError(`Redaction rules file: "${key}" must be an array`); + } + } + for (const key of ['entropy', 'secretFields']) { + if (file[key] !== undefined && (typeof file[key] !== 'object' || file[key] === null || Array.isArray(file[key]))) { + throw new RedactionConfigError(`Redaction rules file: "${key}" must be an object`); + } + } + const includeDefaults = file.includeDefaults !== false; + const disabled = new Set(file.disableRules ?? []); + const baseRules = includeDefaults ? DEFAULT_REDACTION_CONFIG.rules : []; + const baseAllowlist = includeDefaults ? DEFAULT_REDACTION_CONFIG.allowlist : []; + const config = { + rules: replaceById(baseRules, file.rules ?? []).filter(rule => !disabled.has(rule.id)), + allowlist: replaceById(baseAllowlist, file.allowlist ?? []), + entropy: { ...DEFAULT_REDACTION_CONFIG.entropy, ...(file.entropy ?? {}) }, + secretFields: { ...DEFAULT_REDACTION_CONFIG.secretFields, ...(file.secretFields ?? {}) }, + }; + compileConfig(config); // validate + return config; +} +/** + * Load the active redactor from the environment. + * + * Returns null when redaction is off. When the rules can't be loaded: strict + * mode (the default) throws RedactionConfigError so callers fail closed; + * non-strict mode warns and returns null, so text passes through unredacted. + */ +export function loadRedactor(env = process.env, warn = message => console.error(message)) { + const settings = getRedactionSettings(env); + if (!settings.enabled) + return null; + try { + return createRedactor(loadRedactionConfig(readRulesFile(settings))); + } + catch (error) { + const err = error instanceof RedactionConfigError + ? error + : new RedactionConfigError(`Redaction rules failed to load: ${error instanceof Error ? error.message : String(error)}`); + if (settings.strict) + throw err; + warn(`episodic-memory: ${err.message}. EPISODIC_MEMORY_REDACTION_STRICT=0, so conversations ` + + 'will be archived and indexed WITHOUT redaction this run.'); + return null; + } +} +function checkFlags(flags, what) { + const value = flags ?? ''; + if (typeof value !== 'string' || !ALLOWED_FLAGS.test(value)) { + throw new RedactionConfigError(`${what}: flags must be a combination of "imsu"`); + } + return value; +} +function compileRule(spec) { + if (!spec || typeof spec !== 'object') { + throw new RedactionConfigError('Redaction rule must be an object'); + } + if (typeof spec.id !== 'string' || !RULE_ID_PATTERN.test(spec.id)) { + throw new RedactionConfigError(`Redaction rule id ${JSON.stringify(spec.id)} is invalid: use lowercase letters, digits and dashes`); + } + if (RESERVED_IDS.has(spec.id)) { + throw new RedactionConfigError(`Redaction rule id "${spec.id}" is reserved`); + } + const what = `Redaction rule "${spec.id}"`; + if (typeof spec.pattern !== 'string' || spec.pattern.length === 0) { + throw new RedactionConfigError(`${what}: pattern must be a non-empty string`); + } + const flags = checkFlags(spec.flags, what); + let regex; + let groupCount; + try { + regex = new RegExp(spec.pattern, flags + 'gd'); + groupCount = new RegExp(`(?:${spec.pattern})|`, flags).exec('').length - 1; + } + catch (error) { + throw new RedactionConfigError(`${what}: invalid regular expression (${error instanceof Error ? error.message : String(error)})`); + } + if (new RegExp(spec.pattern, flags).test('')) { + throw new RedactionConfigError(`${what}: pattern matches the empty string`); + } + const secretGroups = Array.isArray(spec.secretGroup) ? spec.secretGroup : [spec.secretGroup ?? 0]; + if (secretGroups.length === 0) { + throw new RedactionConfigError(`${what}: secretGroup must not be an empty array`); + } + for (const group of secretGroups) { + if (!Number.isInteger(group) || group < 0 || group > groupCount) { + throw new RedactionConfigError(`${what}: secretGroup ${group} does not exist in the pattern (${groupCount} group(s))`); + } + } + if (spec.useAllowlist !== undefined && typeof spec.useAllowlist !== 'boolean') { + throw new RedactionConfigError(`${what}: useAllowlist must be a boolean`); + } + let keywords; + if (spec.keywords !== undefined) { + if (!Array.isArray(spec.keywords) || spec.keywords.some(k => typeof k !== 'string' || k.length === 0)) { + throw new RedactionConfigError(`${what}: keywords must be an array of non-empty strings`); + } + keywords = spec.keywords.length > 0 ? spec.keywords.map(k => k.toLowerCase()) : undefined; + } + return { id: spec.id, regex, keywords, secretGroups, useAllowlist: spec.useAllowlist !== false }; +} +function compileAllowlist(spec) { + const what = `Redaction allowlist entry ${JSON.stringify(spec?.id)}`; + if (!spec || typeof spec.pattern !== 'string' || spec.pattern.length === 0) { + throw new RedactionConfigError(`${what}: pattern must be a non-empty string`); + } + const flags = checkFlags(spec.flags, what); + try { + return new RegExp(`^(?:${spec.pattern})$`, flags); + } + catch (error) { + throw new RedactionConfigError(`${what}: invalid regular expression (${error instanceof Error ? error.message : String(error)})`); + } +} +function compileConfig(config) { + const seen = new Set(); + const rules = config.rules.map(spec => { + const rule = compileRule(spec); + if (seen.has(rule.id)) + throw new RedactionConfigError(`Duplicate redaction rule id "${rule.id}"`); + seen.add(rule.id); + return rule; + }); + const allowlist = config.allowlist.map(compileAllowlist); + const e = config.entropy; + if (typeof e.enabled !== 'boolean' || typeof e.requireKeyword !== 'boolean' || + !Number.isInteger(e.minLength) || e.minLength < 8 || + typeof e.threshold !== 'number' || !(e.threshold > 0) || + !Number.isInteger(e.window) || e.window < 0 || + !Array.isArray(e.keywords) || e.keywords.some(k => typeof k !== 'string' || !k)) { + throw new RedactionConfigError('Redaction entropy settings are invalid (need enabled, requireKeyword: boolean; minLength >= 8; threshold > 0; window >= 0; keywords: string[])'); + } + const f = config.secretFields; + if (!f || typeof f.enabled !== 'boolean' || typeof f.keyPattern !== 'string' || f.keyPattern.length === 0) { + throw new RedactionConfigError('Redaction secretFields settings are invalid (need enabled: boolean; keyPattern: non-empty string)'); + } + let secretField = null; + if (f.enabled) { + try { + secretField = new RegExp(`(?:${f.keyPattern})$`); + } + catch (error) { + throw new RedactionConfigError(`Redaction secretFields.keyPattern: invalid regular expression (${error instanceof Error ? error.message : String(error)})`); + } + if (secretField.test('')) + throw new RedactionConfigError('Redaction secretFields.keyPattern matches the empty string'); + } + return { rules, allowlist, entropy: { ...e, keywords: e.keywords.map(k => k.toLowerCase()) }, secretField }; +} +// --------------------------------------------------------------------------- +// Redaction engine +// --------------------------------------------------------------------------- +function tokenFor(ruleId) { + return `${TOKEN_PREFIX}${ruleId}]`; +} +const WHOLE_TOKEN = /^\[REDACTED:[a-z0-9][a-z0-9-]*\]$/; +/** The token for one redacted value: onMatch's, if it returns a valid one. */ +function tokenForMatch(ruleId, value, onMatch) { + const custom = onMatch?.(ruleId, value); + return typeof custom === 'string' && WHOLE_TOKEN.test(custom) ? custom : tokenFor(ruleId); +} +function tokenSpans(text) { + const spans = []; + for (const m of text.matchAll(TOKEN_PATTERN)) + spans.push([m.index, m.index + m[0].length]); + return spans; +} +function overlapsAny(spans, start, end) { + for (const [s, e] of spans) + if (start < e && end > s) + return true; + return false; +} +/** + * Rewrite text[start, end) keeping the existing tokens in it and replacing each + * stretch between them that has a letter or digit with `token`. Returns null + * when nothing outside the tokens needs redacting, which keeps this idempotent. + */ +function redactAroundTokens(text, spans, start, end, token) { + let out = ''; + let pos = start; + let replaced = false; + const flush = (to) => { + const piece = text.slice(pos, to); + if (/[A-Za-z0-9]/.test(piece)) { + out += token(); + replaced = true; + } + else { + out += piece; + } + }; + for (const [s, e] of spans) { + if (e <= start || s >= end) + continue; + if (s > pos) + flush(s); + out += text.slice(Math.max(s, pos), Math.min(e, end)); + pos = Math.min(e, end); + } + if (pos < end) + flush(end); + return replaced ? out : null; +} +/** + * Describe a value without revealing it: length, character classes and + * Shannon entropy (bits per char), e.g. "len=40 aA9- H=4.9". Lets someone + * reviewing hits tell a random key from a word or a placeholder. + */ +export function describeShape(value) { + const classes = (/[a-z]/.test(value) ? 'a' : '') + + (/[A-Z]/.test(value) ? 'A' : '') + + (/[0-9]/.test(value) ? '9' : '') + + (/[^A-Za-z0-9\s]/.test(value) ? '-' : '') + + (/\s/.test(value) ? '_' : ''); + return `len=${value.length} ${classes || '?'} H=${shannonEntropy(value).toFixed(1)}`; +} +/** Each `[REDACTED:]` token in `text`, left to right. */ +export function findRedactionTokens(text) { + const out = []; + for (const m of text.matchAll(TOKEN_PATTERN)) { + out.push({ ruleId: m[0].slice(TOKEN_PREFIX.length, -1), start: m.index, end: m.index + m[0].length }); + } + return out; +} +function shannonEntropy(text) { + const counts = new Map(); + for (const ch of text) + counts.set(ch, (counts.get(ch) ?? 0) + 1); + let entropy = 0; + for (const n of counts.values()) { + const p = n / text.length; + entropy -= p * Math.log2(p); + } + return entropy; +} +/** + * Replace each accepted match (or its secretGroup) with a token. A candidate + * that overlaps existing tokens keeps them, and only the text around them is + * redacted (so re-running is a no-op). A candidate is skipped when it fully + * matches an allowlist pattern. + */ +function applyRule(text, rule, isAllowed, onMatch) { + const spans = text.includes(TOKEN_PREFIX) ? tokenSpans(text) : null; + let out = ''; + let last = 0; + let count = 0; + for (const m of text.matchAll(rule.regex)) { + let range; + for (const group of rule.secretGroups) { + range = m.indices?.[group]; + if (range) + break; + } + if (!range) + continue; + const [start, end] = range; + if (end <= start || start < last) + continue; + if (rule.useAllowlist && isAllowed(text.slice(start, end))) + continue; + if (spans && overlapsAny(spans, start, end)) { + let token; + const rewritten = redactAroundTokens(text, spans, start, end, () => (token ??= tokenForMatch(rule.id, text.slice(start, end).replace(TOKEN_PATTERN, ''), onMatch))); + if (rewritten === null) + continue; + out += text.slice(last, start) + rewritten; + last = end; + count++; + continue; + } + out += text.slice(last, start) + tokenForMatch(rule.id, text.slice(start, end), onMatch); + last = end; + count++; + } + return count === 0 ? { text, count } : { text: out + text.slice(last), count }; +} +function applyEntropy(text, spec, isAllowed, onMatch) { + const candidates = new RegExp(`[A-Za-z0-9+/=_.~-]{${spec.minLength},}`, 'g'); + const spans = text.includes(TOKEN_PREFIX) ? tokenSpans(text) : null; + let out = ''; + let last = 0; + let count = 0; + for (const m of text.matchAll(candidates)) { + const start = m.index; + const end = start + m[0].length; + if (spans && overlapsAny(spans, start, end)) + continue; + if (isAllowed(m[0])) + continue; + if (shannonEntropy(m[0]) < spec.threshold) + continue; + if (spec.requireKeyword) { + let before = text.slice(Math.max(0, start - spec.window), start); + before = before.slice(before.lastIndexOf('\n') + 1).toLowerCase(); + if (!spec.keywords.some(k => before.includes(k))) + continue; + } + out += text.slice(last, start) + tokenForMatch(ENTROPY_RULE_ID, m[0], onMatch); + last = end; + count++; + } + return count === 0 ? { text, count } : { text: out + text.slice(last), count }; +} +export function createRedactor(config) { + const compiled = compileConfig(config); + const isAllowed = (secret) => compiled.allowlist.some(re => re.test(secret)); + return { + ruleIds: compiled.rules.map(r => r.id), + isSecretField(name) { + if (!compiled.secretField || typeof name !== 'string') + return false; + const normalized = name.toLowerCase().replace(/[^a-z0-9]/g, '').replace(/\d+$/, ''); + return normalized.length > 0 && compiled.secretField.test(normalized); + }, + redact(text, ctx) { + if (typeof text !== 'string' || text.length === 0) + return { text, findings: [] }; + let current = text; + let lower = null; + const counts = new Map(); + for (const rule of compiled.rules) { + if (rule.keywords) { + lower ??= current.toLowerCase(); + if (!rule.keywords.some(k => lower.includes(k))) + continue; + } + const r = applyRule(current, rule, isAllowed, ctx?.onMatch); + if (r.count > 0) { + current = r.text; + lower = null; + counts.set(rule.id, (counts.get(rule.id) ?? 0) + r.count); + } + } + if (compiled.entropy.enabled) { + const r = applyEntropy(current, compiled.entropy, isAllowed, ctx?.onMatch); + if (r.count > 0) { + current = r.text; + counts.set(ENTROPY_RULE_ID, r.count); + } + } + return { text: current, findings: [...counts].map(([ruleId, count]) => ({ ruleId, count })) }; + }, + }; +} +// --------------------------------------------------------------------------- +// Findings +// --------------------------------------------------------------------------- +/** Aggregates findings by rule id. Never holds matched values. */ +export class FindingsTally { + counts = new Map(); + add(findings) { + for (const f of findings) + this.counts.set(f.ruleId, (this.counts.get(f.ruleId) ?? 0) + f.count); + } + get total() { + let n = 0; + for (const c of this.counts.values()) + n += c; + return n; + } + toArray() { + return [...this.counts].map(([ruleId, count]) => ({ ruleId, count })); + } +} +/** "7 value(s) redacted (azure-storage-key: 2, jwt: 5)" */ +export function formatFindings(findings) { + const total = findings.reduce((n, f) => n + f.count, 0); + if (total === 0) + return 'no values redacted'; + const detail = [...findings] + .sort((a, b) => b.count - a.count || a.ruleId.localeCompare(b.ruleId)) + .map(f => `${f.ruleId}: ${f.count}`) + .join(', '); + return `${total} value(s) redacted (${detail})`; +} +// --------------------------------------------------------------------------- +// JSON / JSONL / files +// --------------------------------------------------------------------------- +/** + * Redact a parsed JSON tree in place: string values by the text rules, values + * by their field name (secretFields), and keys that are themselves secrets. + */ +function redactTree(node, redactor, ctx, tally) { + if (typeof node === 'string') { + const r = redactor.redact(node, ctx); + if (r.findings.length === 0) + return { value: node, changed: false }; + tally?.add(r.findings); + return { value: r.text, changed: r.text !== node }; + } + if (node === null || typeof node !== 'object') + return { value: node, changed: false }; + let changed = false; + if (Array.isArray(node)) { + for (let i = 0; i < node.length; i++) { + const r = redactTree(node[i], redactor, ctx, tally); + if (r.changed) { + node[i] = r.value; + changed = true; + } + } + return { value: node, changed }; + } + const obj = node; + const keys = Object.keys(obj); + for (const key of keys) { + const r = redactTree(obj[key], redactor, ctx, tally); + if (r.changed) { + setOwn(obj, key, r.value); + changed = true; + } + } + // Field-name context: the key (or a {name, value} pair's name, or a Key + // Vault secret id) says the value is a secret even when its shape matches + // no rule — e.g. a letters-only clientSecret field in a structured + // MCP result. Runs after the text rules; a value they redacted only in part + // is still replaced whole, since the field name says all of it is secret. + const redactWhole = (key) => { + const value = obj[key]; + if (!isRedactableWhole(value)) + return; + setOwn(obj, key, tokenForMatch(FIELD_RULE_ID, value, ctx?.onMatch)); + tally?.add([{ ruleId: FIELD_RULE_ID, count: 1 }]); + changed = true; + }; + for (const key of keys) { + if (redactor.isSecretField(key)) + redactWhole(key); + } + const lower = new Map(keys.map(k => [k.toLowerCase(), k])); + const valueKey = lower.get('value'); + if (valueKey !== undefined) { + const nameKey = lower.get('name') ?? lower.get('key'); + const nameValue = nameKey !== undefined ? obj[nameKey] : undefined; + const id = lower.has('id') ? obj[lower.get('id')] : undefined; + if ((typeof nameValue === 'string' && redactor.isSecretField(nameValue)) || + (typeof id === 'string' && KEY_VAULT_SECRET_ID.test(id))) { + redactWhole(valueKey); + } + } + // Keys themselves: a secret used as a key (a token-keyed cache) is renamed + // to its token. Harness keys never match a rule, so structure is unchanged. + let renamed = null; + for (let i = 0; i < keys.length; i++) { + const r = redactor.redact(keys[i], ctx); + if (r.findings.length === 0) + continue; + tally?.add(r.findings); + renamed ??= keys.map(k => [k, obj[k]]); + renamed[i][0] = r.text; + } + if (renamed) { + const rebuilt = {}; + for (const [key, value] of renamed) { + let unique = key; + for (let n = 2; Object.prototype.hasOwnProperty.call(rebuilt, unique); n++) + unique = `${key}#${n}`; + setOwn(rebuilt, unique, value); + } + return { value: rebuilt, changed: true }; + } + return { value: obj, changed }; +} +function setOwn(obj, key, value) { + // defineProperty so a "__proto__" key from JSON.parse stays an own property. + Object.defineProperty(obj, key, { value, writable: true, enumerable: true, configurable: true }); +} +// A value that is only tokens is already redacted; anything else in a secret +// field, even alongside a token, is replaced whole. +const ONLY_TOKENS = /^\s*(?:\[REDACTED:[a-z0-9][a-z0-9-]*\]\s*)+$/; +function isRedactableWhole(value) { + return typeof value === 'string' && + value.trim().length > 0 && + !ONLY_TOKENS.test(value) && + !PLACEHOLDER_VALUE.test(value.trim()); +} +/** + * Redact one JSONL line. Valid JSON stays valid (string values are redacted + * and the line is re-serialized only if something changed); a line that isn't + * JSON (e.g. a torn last line mid-write) is redacted as raw text. A trailing + * `\r` is preserved. + */ +export function redactJsonlLine(line, redactor, ctx, tally) { + const cr = line.endsWith('\r') ? '\r' : ''; + const body = cr ? line.slice(0, -1) : line; + if (body.trim().length === 0) + return line; + let parsed; + try { + parsed = JSON.parse(body); + } + catch { + const r = redactor.redact(body, ctx); + if (r.findings.length === 0) + return line; + tally?.add(r.findings); + return r.text + cr; + } + const r = redactTree(parsed, redactor, ctx, tally); + return r.changed ? JSON.stringify(r.value) + cr : line; +} +const COPY_CHUNK_BYTES = 1 << 20; // 1 MiB +/** + * Stream `src` to `dest`, redacting line by line. Writes `dest` directly; + * callers that need atomicity write to a temp path and rename. Line count and + * trailing-newline shape match the source exactly, so archive line numbers + * (exchanges.line_start/line_end, MCP read ranges) stay valid. + */ +export function copyFileRedacted(src, dest, redactor, ctx, tally) { + const fdIn = fs.openSync(src, 'r'); + let fdOut; + try { + // Create with the source's mode, as copyFileSync would, so a 0600 + // transcript doesn't become world-readable in the archive. + fdOut = fs.openSync(dest, 'w', fs.fstatSync(fdIn).mode & 0o777); + const buf = Buffer.allocUnsafe(COPY_CHUNK_BYTES); + const decoder = new StringDecoder('utf8'); + let pending = ''; + let bytesRead; + while ((bytesRead = fs.readSync(fdIn, buf, 0, buf.length, null)) > 0) { + let scanFrom = pending.length; + pending += decoder.write(buf.subarray(0, bytesRead)); + const out = []; + let start = 0; + let nl; + while ((nl = pending.indexOf('\n', scanFrom)) !== -1) { + out.push(redactJsonlLine(pending.slice(start, nl), redactor, ctx, tally), '\n'); + start = nl + 1; + scanFrom = start; + } + if (out.length > 0) + fs.writeSync(fdOut, out.join('')); + pending = pending.slice(start); + } + pending += decoder.end(); + if (pending.length > 0) + fs.writeSync(fdOut, redactJsonlLine(pending, redactor, ctx, tally)); + } + finally { + fs.closeSync(fdIn); + if (fdOut !== undefined) + fs.closeSync(fdOut); + } +} diff --git a/dist/summarizer.d.ts b/dist/summarizer.d.ts index 4c97be52..32272d9e 100644 --- a/dist/summarizer.d.ts +++ b/dist/summarizer.d.ts @@ -172,4 +172,13 @@ export declare function runCodexCommand(command: CodexSummarizerCommand): Promis * See https://github.com/obra/episodic-memory/issues/98. */ export declare function getCodexModel(_exchanges: ConversationExchange[]): string | undefined; -export declare function summarizeConversation(exchanges: ConversationExchange[], sessionId?: string): Promise; +export interface SummarizeOptions { + /** + * Allow Claude session resume and Codex thread/fork (default true). Both + * paths make the model read the *source* transcript rather than `exchanges`, + * so callers pass false when the exchanges were redacted (see + * docs/REDACTION.md) to force the transcript-text path. + */ + allowResume?: boolean; +} +export declare function summarizeConversation(exchanges: ConversationExchange[], sessionId?: string, options?: SummarizeOptions): Promise; diff --git a/dist/summarizer.js b/dist/summarizer.js index 60719809..48e78be2 100644 --- a/dist/summarizer.js +++ b/dist/summarizer.js @@ -617,7 +617,8 @@ function getCodexSessionId(exchanges, sessionId) { export function getCodexModel(_exchanges) { return process.env.EPISODIC_MEMORY_CODEX_MODEL || undefined; } -export async function summarizeConversation(exchanges, sessionId) { +export async function summarizeConversation(exchanges, sessionId, options = {}) { + const allowResume = options.allowResume !== false; // Handle trivial conversations if (exchanges.length === 0) { return 'Trivial conversation with no substantive content.'; @@ -628,7 +629,7 @@ export async function summarizeConversation(exchanges, sessionId) { return 'Trivial conversation with no substantive content.'; } } - const codexSessionId = getCodexSessionId(exchanges, sessionId); + const codexSessionId = allowResume ? getCodexSessionId(exchanges, sessionId) : undefined; if (codexSessionId) { try { const result = await callCodex(buildCodexSummaryPrompt(), codexSessionId, getCodexModel(exchanges)); @@ -652,7 +653,7 @@ export async function summarizeConversation(exchanges, sessionId) { // would fail on every one before the no-resume retry kicks in. Treat // missing harness as Claude for backward compatibility with old archives. const isClaudeSession = exchanges.some(e => e.harness === 'claude' || e.harness === undefined); - const claudeSessionId = !codexSessionId && isClaudeSession ? sessionId : undefined; + const claudeSessionId = allowResume && !codexSessionId && isClaudeSession ? sessionId : undefined; const cwd = claudeSessionId ? exchanges.find(e => e.cwd)?.cwd : undefined; const conversationText = claudeSessionId ? '' // When resuming, no need to include conversation text - it's already in context diff --git a/dist/sync-cli.js b/dist/sync-cli.js index 26172c70..b4450162 100644 --- a/dist/sync-cli.js +++ b/dist/sync-cli.js @@ -9,6 +9,7 @@ import { spawn } from 'child_process'; import fs from 'fs'; import { formatLogLine, getSyncLogPath, getSyncLockPath } from './logging.js'; import { acquireFileLock, readLockHolder, releaseFileLock } from './file-lock.js'; +import { FindingsTally, formatFindings, getRedactionSettings, loadRedactor } from './redaction.js'; const args = process.argv.slice(2); // Reentrancy guard (#87): if this sync was triggered by a SessionStart hook // inside a Claude subprocess that the summarizer just spawned, exit silently. @@ -39,10 +40,13 @@ Sync conversations from Claude Code, Codex, and opencode transcript sources to a This command: 1. Exports opencode sessions from its SQLite database when available -2. Copies new or updated .jsonl files to conversation archive +2. Copies new or updated .jsonl files to conversation archive, redacting secrets 3. Generates embeddings for semantic search 4. Updates the search index +Secrets are replaced with [REDACTED:] tokens before anything is archived, +indexed, embedded, or summarized. See EPISODIC_MEMORY_REDACTION* in the README. + Only processes files that are new or have been modified since last sync. Safe to run multiple times - subsequent runs are fast no-ops. @@ -142,8 +146,21 @@ if (isBackground) { console.log(`Sync started in background. Log: ${logPath}`); process.exit(0); } +// Load redaction rules before anything is exported, copied, or indexed. In +// strict mode (the default) a bad rules file stops the sync here: fail closed +// rather than archive unredacted text. +let redactor; +try { + redactor = loadRedactor(); +} +catch (error) { + console.error(`episodic-memory: ${error instanceof Error ? error.message : String(error)}`); + console.error('episodic-memory: refusing to sync without redaction (strict mode). Fix the rules file, ' + + 'or set EPISODIC_MEMORY_REDACTION_STRICT=0 to sync unredacted, or EPISODIC_MEMORY_REDACTION=off.'); + process.exit(1); +} if (!onlyHarnesses || onlyHarnesses.includes('opencode')) { - const opencodeExport = exportOpencodeSessions(); + const opencodeExport = exportOpencodeSessions({ redactor }); if (opencodeExport.exported > 0 || opencodeExport.skipped > 0) { console.log(`opencode export: ${opencodeExport.exported} exported, ${opencodeExport.skipped} skipped`); } @@ -195,8 +212,10 @@ console.log(`Sources: ${sourceDirs.join(', ')}`); console.log(`Destination: ${destDir}\n`); async function syncAll() { const totals = { copied: 0, skipped: 0, indexed: 0, summarized: 0, errors: [], sourcesWithSummaryWork: 0, totalNeedingSummaries: 0 }; + const redactions = new FindingsTally(); for (const sourceDir of sourceDirs) { - const result = await syncConversations(sourceDir, destDir, { ...syncOptions, summaryLimit }); + const result = await syncConversations(sourceDir, destDir, { ...syncOptions, summaryLimit, redactor }); + redactions.add(result.redactions); totals.copied += result.copied; totals.skipped += result.skipped; totals.indexed += result.indexed; @@ -213,6 +232,16 @@ async function syncAll() { else { console.log(` Summarized: ${totals.summarized}`); } + // Rule IDs and counts only — never matched values. + if (redactor) { + console.log(` Redaction: ${formatFindings(redactions.toArray())}`); + } + else if (getRedactionSettings().enabled) { + console.log(' Redaction: NOT APPLIED (rules failed to load; EPISODIC_MEMORY_REDACTION_STRICT=0)'); + } + else { + console.log(' Redaction: off (EPISODIC_MEMORY_REDACTION=off)'); + } if (totals.errors.length > 0) { console.log(`\n⚠️ Errors: ${totals.errors.length}`); totals.errors.forEach(err => console.log(` ${err.file}: ${err.error}`)); diff --git a/dist/sync.d.ts b/dist/sync.d.ts index 658e98ad..36b5afb5 100644 --- a/dist/sync.d.ts +++ b/dist/sync.d.ts @@ -1,3 +1,4 @@ +import { FindingsTally, type RedactionFinding, type Redactor } from './redaction.js'; /** * Stream and scan for any exclusion marker, carrying an overlap between * chunks so a marker split across a boundary is still found. A single @@ -17,11 +18,18 @@ export interface SyncResult { file: string; error: string; }>; + /** Rule IDs and counts only — never matched values. */ + redactions: RedactionFinding[]; } export interface SyncOptions { skipIndex?: boolean; skipSummaries?: boolean; summaryLimit?: number; + /** + * Redactor applied at the archive write. `undefined` loads it from the + * environment (loadRedactor); `null` means redaction is off. + */ + redactor?: Redactor | null; } /** * Derive sync options from the process environment. @@ -39,5 +47,16 @@ export interface SyncOptions { * Claude quota and can stall on a permission prompt. */ export declare function buildSyncOptionsFromEnv(env: NodeJS.ProcessEnv): SyncOptions; +/** + * Copy `src` into the archive at `dest` when the archive copy is missing or + * older. This is the redaction choke point: with a redactor, the copy is + * redacted line by line (see redaction.ts), and every downstream stage + * (index, embeddings, summaries, show/read) reads the archive. + * + * `force` copies even when the mtimes say the archive is current: an append + * within the mtime granularity (or the millisecond the rounding below adds) + * leaves the source no newer than the archive. + */ +export declare function copyIfNewer(src: string, dest: string, redactor?: Redactor | null, tally?: FindingsTally, force?: boolean): boolean; export declare function extractSessionIdFromPath(filePath: string): string | null; export declare function syncConversations(sourceDir: string, destDir: string, options?: SyncOptions): Promise; diff --git a/dist/sync.js b/dist/sync.js index 0629761d..6bb83b75 100644 --- a/dist/sync.js +++ b/dist/sync.js @@ -5,6 +5,7 @@ import { SUMMARIZER_CONTEXT_MARKER } from './constants.js'; import { getExcludedProjects, findJsonlFiles, statIfExists } from './paths.js'; import { formatErrorSentinel, shouldQueueForSummary } from './summary-sentinel.js'; import { getMaxMessageBytes, isOversizeExchange } from './message-size.js'; +import { copyFileRedacted, FindingsTally, loadRedactor } from './redaction.js'; const EXCLUSION_MARKERS = [ 'DO NOT INDEX THIS CHAT', 'Only use NO_INSIGHTS_FOUND', @@ -122,14 +123,24 @@ function hasConversationContent(filePath) { export function buildSyncOptionsFromEnv(env) { return { skipSummaries: env.EPISODIC_MEMORY_SKIP_SUMMARIES === '1' }; } -function copyIfNewer(src, dest) { +/** + * Copy `src` into the archive at `dest` when the archive copy is missing or + * older. This is the redaction choke point: with a redactor, the copy is + * redacted line by line (see redaction.ts), and every downstream stage + * (index, embeddings, summaries, show/read) reads the archive. + * + * `force` copies even when the mtimes say the archive is current: an append + * within the mtime granularity (or the millisecond the rounding below adds) + * leaves the source no newer than the archive. + */ +export function copyIfNewer(src, dest, redactor = null, tally, force = false) { // Ensure destination directory exists const destDir = path.dirname(dest); if (!fs.existsSync(destDir)) { fs.mkdirSync(destDir, { recursive: true }); } // Check if destination exists and is up-to-date - if (fs.existsSync(dest)) { + if (!force && fs.existsSync(dest)) { const srcStat = fs.statSync(src); const destStat = fs.statSync(dest); if (destStat.mtimeMs >= srcStat.mtimeMs) { @@ -138,8 +149,22 @@ function copyIfNewer(src, dest) { } // Atomic copy: temp file + rename const tempDest = dest + '.tmp.' + process.pid; - fs.copyFileSync(src, tempDest); - fs.renameSync(tempDest, dest); // Atomic on same filesystem + try { + if (redactor) { + copyFileRedacted(src, tempDest, redactor, { source: 'archive', path: src }, tally); + } + else { + fs.copyFileSync(src, tempDest); + } + fs.renameSync(tempDest, dest); // Atomic on same filesystem + } + catch (error) { + try { + fs.unlinkSync(tempDest); + } + catch { } + throw error; + } // Preserve source mtime: harnesses without per-message timestamps (Cursor // agent transcripts) fall back to file mtime. Round up to the next whole // millisecond — utimes can't always represent the source's sub-millisecond @@ -165,8 +190,14 @@ export async function syncConversations(sourceDir, destDir, options = {}) { skipped: 0, indexed: 0, summarized: 0, - errors: [] + errors: [], + redactions: [] }; + // Load the redactor before touching the archive. In strict mode (the + // default) a rules-load failure throws here, so nothing unredacted is + // archived, indexed, or summarized (fail closed). + const redactor = options.redactor === undefined ? loadRedactor() : options.redactor; + const tally = new FindingsTally(); // Ensure source directory exists if (!fs.existsSync(sourceDir)) { return result; @@ -198,7 +229,7 @@ export async function syncConversations(sourceDir, destDir, options = {}) { result.skipped++; continue; } - const wasCopied = copyIfNewer(srcFile, destFile); + const wasCopied = copyIfNewer(srcFile, destFile, redactor, tally); if (wasCopied) { result.copied++; filesToIndex.push(destFile); @@ -229,6 +260,7 @@ export async function syncConversations(sourceDir, destDir, options = {}) { } } } + result.redactions = tally.toArray(); // Index copied files (unless skipIndex is set) if (!options.skipIndex && filesToIndex.length > 0) { const { parseConversation } = await import('./parser.js'); @@ -343,7 +375,10 @@ export async function syncConversations(sourceDir, destDir, options = {}) { continue; } console.log(` Summarizing ${path.basename(filePath)} (${exchanges.length} exchanges)...`); - const summary = await summarizeConversation(exchanges, sessionId); + // With redaction on, never resume/fork the session: those paths hand + // the model the unredacted source transcript instead of these + // (redacted) exchanges. + const summary = await summarizeConversation(exchanges, sessionId, { allowResume: redactor === null }); const summaryPath = filePath.replace('.jsonl', '-summary.txt'); fs.writeFileSync(summaryPath, summary, 'utf-8'); result.summarized++; diff --git a/dist/verify.js b/dist/verify.js index dd4535d2..54df6a85 100644 --- a/dist/verify.js +++ b/dist/verify.js @@ -4,6 +4,7 @@ import { parseConversation } from './parser.js'; import { initDatabase, getAllExchanges, getFileLastIndexed } from './db.js'; import { getArchiveDir, getExcludedProjects, findJsonlFiles, statIfExists } from './paths.js'; import { isErroredSentinel } from './summary-sentinel.js'; +import { loadRedactor } from './redaction.js'; export async function verifyIndex() { const result = { missing: [], @@ -93,6 +94,9 @@ export async function verifyIndex() { } export async function repairIndex(issues) { console.log('Repairing index...'); + // Load before touching the index: strict mode fails closed here, as in sync + // and index. + const redactor = loadRedactor(); // To avoid circular dependencies, we import the indexer functions dynamically const { initDatabase, insertExchange, deleteExchange } = await import('./db.js'); const { parseConversation } = await import('./parser.js'); @@ -125,7 +129,8 @@ export async function repairIndex(issues) { } // Generate/update summary const summaryPath = conversationPath.replace('.jsonl', '-summary.txt'); - const summary = await summarizeConversation(exchanges); + // Under redaction, don't let the Codex fork fallback read the source rollout. + const summary = await summarizeConversation(exchanges, undefined, { allowResume: redactor === null }); fs.writeFileSync(summaryPath, summary, 'utf-8'); console.log(` Created summary: ${summary.split(/\s+/).length} words`); // Index exchanges diff --git a/docs/REDACTION.md b/docs/REDACTION.md new file mode 100644 index 00000000..6fe53eac --- /dev/null +++ b/docs/REDACTION.md @@ -0,0 +1,175 @@ +# Secret redaction + +episodic-memory replaces secrets in your conversations with typed tokens +before it archives, indexes, embeds, or summarizes them: + +```text +AccountKey=<88-char key> → AccountKey=[REDACTED:connection-string-secret] +"ClientSecret": "" → "ClientSecret": "[REDACTED:quoted-secret-assignment]" + → +Authorization: Bearer → Authorization: Bearer [REDACTED:jwt] +``` + +Only the value is replaced, so the rest of the conversation stays searchable, +and the token says what kind of value was there. +`episodic-memory search --text "[REDACTED:azure-storage-key]"` finds every +conversation where a storage key was pasted. + +Redaction is on by default and fails closed: if the rules can't load, sync +doesn't run. + +## Where it happens + +Every harness's transcripts (Claude Code, Codex, Cursor, opencode, OMP) are +copied into the conversation archive before anything else reads them. That +copy is the one point they all pass through, so it is where redaction happens +([`copyIfNewer`](../src/sync.ts) → [`copyFileRedacted`](../src/redaction.ts)). +Everything downstream reads the archive: + +| Where | Redacted? | +|---|---| +| Conversation archive (`~/.config/superpowers/conversation-archive`) | Yes, line by line as it's copied | +| SQLite index: message text and tool inputs/outputs | Yes, parsed from the archive ([indexer.ts](../src/indexer.ts) included) | +| Embeddings | Yes, built from redacted text | +| Summaries (sent to a model) | Yes. Summarizer resume and Codex fork are turned off, because both read the original transcript ([summarizer.ts](../src/summarizer.ts)) | +| opencode and legacy Cursor staging exports | Yes, when written ([opencode-sync.ts](../src/opencode-sync.ts), [cursor-legacy.ts](../src/cursor-legacy.ts)) | +| `show`, MCP `read` | Yes, they read the archive | +| Logs | Rule IDs and counts only | +| The harness's own files (`~/.claude/projects`, `~/.codex/sessions`, …) | **No.** They belong to the harness; use its retention settings | + +With summarizer resume off, Codex-only setups summarize through the Claude +Agent SDK. Without Claude configured, set `EPISODIC_MEMORY_SKIP_SUMMARIES=1`. +Summaries are display-only, so search is unaffected. + +## How a secret is recognized + +The engine is [src/redaction.ts](../src/redaction.ts) and the default rules are +[src/redaction-rules.ts](../src/redaction-rules.ts). Each archive line is parsed +as JSON and passes through three layers: + +1. **Shape.** Values whose format gives them away: private keys, JWTs, + provider-prefixed keys, Entra client secrets (`…Q~…`), 88-character Azure + storage keys, SAS `sig=` values. +2. **Key context.** Values with no recognizable format, caught by the name + they're assigned to: `Password=` in a connection string, `"ClientSecret": "…"`, + ``, `{"name": "DB_PASSWORD", "value": "…"}`, + `password: …`. +3. **Field name.** In parsed JSON, any string whose field name is + secret-looking is replaced whole, including tool inputs and MCP results. + +A **secret-looking name** *ends* in `secret`, `password`, `passwd`, +`passphrase`, `apikey`, `accesskey`, `accountkey`, `privatekey`, `token`, +`credential(s)` or a similar key word, in any case and with any separators +(`AzureAd:ClientSecret`, `DB_PASSWORD2`). `TokenEndpoint` and `passwordPolicy` +don't count. Placeholders (`${X}`, `$(X)`, `#{X}#`, `{{x}}`, ``, `%X%`, +`"string"`, `"*****"`) are left alone. + +Git SHAs and GUIDs are allowlisted for the shape rules, so commit hashes and +tenant, client and object IDs stay searchable. Key-context rules ignore the +allowlist: a GUID in a `password` slot is a secret. An entropy fallback exists +but is off by default. + +| Rules | Catch | +|---|---| +| `private-key-block`, `jwt`, `anthropic-api-key`, `openai-api-key`, `github-token`, `aws-access-key-id`, `aws-secret-access-key`, `slack-token`, `google-api-key`, `npm-token`, `azure-client-secret`, `azure-storage-key`, `azure-sas-token` | Shape | +| `connection-string-secret` | `AccountKey=`, `SharedAccessKey=`, `Password=`, `Pwd=` and similar in connection strings, any case | +| `url-credentials`, `bearer-token`, `basic-auth` | `scheme://user:password@host`, `Authorization` headers | +| `azure-keyvault-secret`, `name-value-secret` | Key Vault bundles; `{name, value}` pairs (`az … appsettings list`, Kubernetes `env`, ARM parameters) | +| `xml-appsettings-secret`, `xml-secret-element`, `xml-secret-attribute` | `web.config`, `…`, publish-profile `userPWD="…"` | +| `quoted-secret-assignment` | `"ClientSecret": "…"`, `ClientSecret = "…"`, `apiKey: '…'`; no spaces in the value | +| `secret-assignment` | `password: …`, `CLIENT_SECRET=…` (YAML, dotenv, decrypted SOPS); 8+ characters with a digit, code like `env.X` skipped | +| `secret-field` | Layer 3: a secret-named JSON field, replaced whole | + +`episodic-memory redact --print-default-rules` prints the full definitions. + +## Guarantees + +- **Fails closed.** In strict mode (the default), sync, index, `index repair` + and import stop if the rules can't load. +- **Idempotent.** Tokens never match again. When a match runs into an existing + token, only the text outside it is redacted, so re-running after adding a + rule is safe. +- **Structure-preserving.** One output line per input line, valid JSON stays + valid, and lines with no secrets stay byte-for-byte identical. Index line + ranges and MCP `read` ranges stay correct. Archive copies keep the source + file's permissions. +- **Values are never printed.** Logs and reports name rules and counts; the + review report adds the value's shape, never the value. + +## Cleaning up existing data + +New syncs redact what they copy. To redact what was archived and indexed +before, or after you add a rule: + +```bash +episodic-memory redact --rewrite --dry-run # what would change; writes nothing +episodic-memory redact --rewrite --dry-run --report # each hit, to check for false positives +episodic-memory redact --rewrite # apply +``` + +`--rewrite` ([src/redact-rewrite.ts](../src/redact-rewrite.ts)) redacts the +archive, the staging exports and the index in place, re-embeds only the rows +that changed, and deletes summaries built from unredacted text so the next +sync regenerates them. It takes the sync lock. A dry run opens the index +read-only. + +`--report` prints one hit per value: + +```text + -work-contoso/4f1c….jsonl:212 quoted-secret-assignment len=40 aA9- H=4.9 + …"AzureAd": { "ClientId": "…", "ClientSecret": "[REDACTED:quoted-secret-assignment]", "TenantId… +``` + +The shape is the length, the character classes (`a` lower, `A` upper, `9` +digits, `-` symbols, `_` spaces) and the entropy in bits per character. A +random key is long with entropy around 4.5 or more; a word or placeholder is +short or low. + +## Configuration + +| Variable | Default | Meaning | +|---|---|---| +| `EPISODIC_MEMORY_REDACTION` | `on` | `off` turns redaction off: byte-for-byte archive copies, summarizer resume allowed | +| `EPISODIC_MEMORY_REDACTION_RULES` | `/redaction-rules.json` if present | Custom rules file; a missing file you named is an error | +| `EPISODIC_MEMORY_REDACTION_STRICT` | `1` | `0` continues **unredacted**, with a warning, when the rules fail to load | + +Custom rules extend the defaults: + +```jsonc +{ + "rules": [ // a rule with a default's id replaces it + { "id": "contoso-api-key", "pattern": "\\bctso_[A-Za-z0-9]{32}\\b", "keywords": ["ctso_"] } + ], + "disableRules": ["basic-auth"], + "allowlist": [{ "id": "build-ids", "pattern": "build-[0-9]{8}" }], + "entropy": { "enabled": true }, + "secretFields": { "keyPattern": "secret|password|token|credentials?|connectionstring" }, + "includeDefaults": true +} +``` + +A rule can also set `flags` (`imsu`), `secretGroup` (redact only that capture +group) and `useAllowlist: false`. Every rule is validated on load; an invalid +one is a load failure, which strict mode treats as fatal. Try rules with +`echo 'password: hunter2hunter2' | episodic-memory redact --stdin`. + +## Limits + +This is pattern and context matching, not named-entity recognition: a value +is caught by its format, by the name it's assigned to, or by the field it's +in. That covers how secrets reach transcripts in practice (pasted config, CLI +output, connection strings, tool results), and it is deterministic, cheap +enough for every line of every sync, and each token names the rule that fired. +What it misses: + +- A secret in prose with no name or shape: "the password is hunter2". +- An unquoted letters-only value (`password: hunter`): `secret-assignment` + needs a digit so ordinary English doesn't trigger it. +- Code-like values (`config.password`, `getPassword()`), skipped on purpose. +- Quoted values with spaces, so UI strings like `"Invalid password"` survive. +- Settings with a shapeless value and a name that isn't secret-looking + (`Stripe`, `ConnectionStrings__Default`). Add the name to + `secretFields.keyPattern` or write a rule. + +The entropy fallback narrows some of these gaps, at the cost of false +positives. diff --git a/src/cursor-legacy.ts b/src/cursor-legacy.ts index a62f4b56..e37366b1 100644 --- a/src/cursor-legacy.ts +++ b/src/cursor-legacy.ts @@ -3,6 +3,7 @@ import os from 'os'; import path from 'path'; import Database from 'better-sqlite3'; import { detectCursorCwd } from './parser.js'; +import { loadRedactor, redactJsonlLine, type Redactor } from './redaction.js'; /** * Import legacy Cursor conversations from Cursor's global SQLite store @@ -141,6 +142,8 @@ export interface CursorLegacyImportOptions { force?: boolean; /** Report what would be exported without writing files. */ dryRun?: boolean; + /** `undefined` loads from the environment (strict mode throws on bad rules); `null` disables. */ + redactor?: Redactor | null; } export interface CursorLegacyImportResult { @@ -160,6 +163,7 @@ export function importCursorLegacy(options: CursorLegacyImportOptions): CursorLe errors: [], }; + const redactor = options.redactor === undefined ? loadRedactor() : options.redactor; const liveIds = options.liveTranscriptIds ?? new Set(); const db = new Database(options.dbPath, { readonly: true, fileMustExist: true }); @@ -264,9 +268,14 @@ export function importCursorLegacy(options: CursorLegacyImportOptions): CursorLe if (!options.dryRun) { // Re-serialize with cwd now that it's known (it's derived from the // whole conversation's tool calls). - const finalLines = cwd + const withCwd = cwd ? lines.map(line => JSON.stringify({ ...JSON.parse(line), cwd })) : lines; + // The export dir is a plugin-owned plaintext copy, so redact it at + // write time like the archive (docs/REDACTION.md). + const finalLines = redactor + ? withCwd.map(line => redactJsonlLine(line, redactor, { source: 'cursor-legacy', path: outFile })) + : withCwd; fs.mkdirSync(path.dirname(outFile), { recursive: true }); fs.writeFileSync(outFile, finalLines.join('\n') + '\n', 'utf-8'); diff --git a/src/db.ts b/src/db.ts index 8c7a73ad..a3811e2c 100644 --- a/src/db.ts +++ b/src/db.ts @@ -104,6 +104,17 @@ export function migrateToolCallsCascade(db: Database.Database): void { console.log(' tool_calls migration complete.'); } +/** + * Open the existing index read-only, without creating it, migrating it, or + * changing its journal mode. For dry runs that must not write. Returns null + * when there is no index yet. + */ +export function openDatabaseReadOnly(): Database.Database | null { + const dbPath = getDbPath(); + if (!fs.existsSync(dbPath)) return null; + return new Database(dbPath, { readonly: true, fileMustExist: true }); +} + export function initDatabase(): Database.Database { const dbPath = getDbPath(); diff --git a/src/indexer.ts b/src/indexer.ts index 173ef679..df9a5986 100644 --- a/src/indexer.ts +++ b/src/indexer.ts @@ -9,6 +9,8 @@ import { ConversationExchange } from './types.js'; import { getArchiveDir, getExcludedProjects, getConversationSourceDirs, findJsonlFiles, statIfExists } from './paths.js'; import { formatErrorSentinel, shouldQueueForSummary } from './summary-sentinel.js'; import { getMaxMessageBytes, isOversizeExchange } from './message-size.js'; +import { FindingsTally, formatFindings, loadRedactor, type Redactor } from './redaction.js'; +import { copyIfNewer } from './sync.js'; // Set max output tokens for Claude SDK (used by summarizer) process.env.CLAUDE_CODE_MAX_OUTPUT_TOKENS = '20000'; @@ -38,12 +40,25 @@ function sessionIdForSummary(exchanges: ConversationExchange[]): string | undefi return exchanges.find(exchange => exchange.sessionId)?.sessionId; } +// Resume/fork would summarize the unredacted source transcript; see sync.ts. +function summarizeOptions(redactor: Redactor | null) { + return { allowResume: redactor === null }; +} + +function logRedactions(tally: FindingsTally): void { + if (tally.total > 0) console.log(` Redaction: ${formatFindings(tally.toArray())}`); +} + export async function indexConversations( limitToProject?: string, maxConversations?: number, concurrency: number = 1, noSummaries: boolean = false ): Promise { + // Load before touching the archive: strict mode fails closed here. + const redactor = loadRedactor(); + const tally = new FindingsTally(); + console.log('Initializing database...'); const db = initDatabase(); @@ -113,15 +128,14 @@ export async function indexConversations( // Source transcripts can vanish mid-run (Claude Code cleanup). Skip loudly. let exchanges; try { - // Copy to archive (ensure parent dirs exist for subagent files) - if (!fs.existsSync(archivePath)) { - fs.mkdirSync(path.dirname(archivePath), { recursive: true }); - fs.copyFileSync(sourcePath, archivePath); + // Copy (redacted) to the archive, then parse the archive, so the index + // and summaries only ever see redacted text. + if (copyIfNewer(sourcePath, archivePath, redactor, tally)) { console.log(` Archived: ${file}`); } // Parse conversation - exchanges = await parseConversation(sourcePath, project, archivePath); + exchanges = await parseConversation(archivePath, project, archivePath); } catch (error) { console.log(` Skipped ${file} (read failed: ${error instanceof Error ? error.message : error})`); continue; @@ -150,7 +164,7 @@ export async function indexConversations( await processBatch(needsSummary, async (conv) => { try { - const summary = await summarizeConversation(conv.exchanges, sessionIdForSummary(conv.exchanges)); + const summary = await summarizeConversation(conv.exchanges, sessionIdForSummary(conv.exchanges), summarizeOptions(redactor)); fs.writeFileSync(conv.summaryPath, summary, 'utf-8'); const wordCount = summary.split(/\s+/).length; console.log(` ✓ ${conv.file}: ${wordCount} words`); @@ -193,6 +207,7 @@ export async function indexConversations( // Check if we hit the limit if (maxConversations && conversationsProcessed >= maxConversations) { console.log(`\nReached limit of ${maxConversations} conversations`); + logRedactions(tally); db.close(); console.log(`✅ Indexing complete! Conversations: ${conversationsProcessed}, Exchanges: ${totalExchanges}`); return; @@ -205,12 +220,15 @@ export async function indexConversations( console.log(` Skipped ${oversizeSkipped} oversize exchange(s) (> ${maxMessageBytes} bytes; set EPISODIC_MEMORY_MAX_MESSAGE_BYTES to change) — likely embedded-transcript payloads (#139)`); } + logRedactions(tally); db.close(); console.log(`\n✅ Indexing complete! Conversations: ${conversationsProcessed}, Exchanges: ${totalExchanges}`); } export async function indexSession(sessionId: string, concurrency: number = 1, noSummaries: boolean = false): Promise { console.log(`Indexing session: ${sessionId}`); + const redactor = loadRedactor(); + const tally = new FindingsTally(); // Find the conversation file for this session const sourceDirs = getConversationSourceDirs(); @@ -246,11 +264,8 @@ export async function indexSession(sessionId: string, concurrency: number = 1, n // Archive + parse — source may vanish mid-run (Claude Code cleanup). let exchanges; try { - if (!fs.existsSync(archivePath)) { - fs.mkdirSync(path.dirname(archivePath), { recursive: true }); - fs.copyFileSync(sourcePath, archivePath); - } - exchanges = await parseConversation(sourcePath, project, archivePath); + copyIfNewer(sourcePath, archivePath, redactor, tally); + exchanges = await parseConversation(archivePath, project, archivePath); } catch (error) { console.log(`Skipped ${file} (read failed: ${error instanceof Error ? error.message : error})`); db.close(); @@ -263,7 +278,7 @@ export async function indexSession(sessionId: string, concurrency: number = 1, n if (!noSummaries && shouldQueueForSummary(summaryPath)) { fs.mkdirSync(path.dirname(summaryPath), { recursive: true }); try { - const summary = await summarizeConversation(exchanges, sessionIdForSummary(exchanges)); + const summary = await summarizeConversation(exchanges, sessionIdForSummary(exchanges), summarizeOptions(redactor)); fs.writeFileSync(summaryPath, summary, 'utf-8'); console.log(`Summary: ${summary.split(/\s+/).length} words`); } catch (error) { @@ -297,6 +312,7 @@ export async function indexSession(sessionId: string, concurrency: number = 1, n console.log(` Skipped ${oversizeSkipped} oversize exchange(s) (> ${maxMessageBytes} bytes; set EPISODIC_MEMORY_MAX_MESSAGE_BYTES to change) — likely embedded-transcript payloads (#139)`); } + logRedactions(tally); console.log(`✅ Indexed session ${sessionId}: ${exchanges.length} exchanges`); } @@ -317,6 +333,9 @@ export async function indexUnprocessed(concurrency: number = 1, noSummaries: boo if (concurrency > 1) console.log(`Concurrency: ${concurrency}`); if (noSummaries) console.log('⚠️ Running in no-summaries mode (skipping AI summaries)'); + const redactor = loadRedactor(); + const tally = new FindingsTally(); + const db = initDatabase(); await initEmbeddings(); @@ -361,17 +380,14 @@ export async function indexUnprocessed(concurrency: number = 1, noSummaries: boo ).get(archivePath) as { maxLine: number }; const maxIndexedLine = hw.maxLine; - // Ensure parent dirs exist for subagent files try { - fs.mkdirSync(path.dirname(archivePath), { recursive: true }); - - // Refresh the archive when the source may have grown beyond what we've seen. - if (!fs.existsSync(archivePath) || maxIndexedLine > 0) { - fs.copyFileSync(sourcePath, archivePath); - } + // Refresh the (redacted) archive, then parse the archive so the index + // only sees redacted text. Force the refresh once the file is indexed: + // an append inside the mtime granularity would otherwise be skipped. + copyIfNewer(sourcePath, archivePath, redactor, tally, maxIndexedLine > 0); // Parse and filter to exchanges past the high-water mark - const exchanges = await parseConversation(sourcePath, project, archivePath); + const exchanges = await parseConversation(archivePath, project, archivePath); const newExchanges = maxIndexedLine > 0 ? exchanges.filter(e => e.lineStart > maxIndexedLine) : exchanges; @@ -386,6 +402,8 @@ export async function indexUnprocessed(concurrency: number = 1, noSummaries: boo } } // end sourceDir loop + logRedactions(tally); + if (unprocessed.length === 0) { console.log('✅ All conversations are already processed!'); db.close(); @@ -402,7 +420,7 @@ export async function indexUnprocessed(concurrency: number = 1, noSummaries: boo await processBatch(needsSummary, async (conv) => { try { - const summary = await summarizeConversation(conv.exchanges, sessionIdForSummary(conv.exchanges)); + const summary = await summarizeConversation(conv.exchanges, sessionIdForSummary(conv.exchanges), summarizeOptions(redactor)); fs.writeFileSync(conv.summaryPath, summary, 'utf-8'); const wordCount = summary.split(/\s+/).length; console.log(` ✓ ${conv.project}/${conv.file}: ${wordCount} words`); diff --git a/src/opencode-sync.ts b/src/opencode-sync.ts index b61f8ef4..eca6876c 100644 --- a/src/opencode-sync.ts +++ b/src/opencode-sync.ts @@ -2,6 +2,7 @@ import fs from 'fs'; import path from 'path'; import Database from 'better-sqlite3'; import { getOpencodeDbPath, getOpencodeTranscriptDir } from './paths.js'; +import { loadRedactor, redactJsonlLine, type Redactor } from './redaction.js'; export interface OpencodeExportResult { exported: number; @@ -92,7 +93,8 @@ function shouldExportSession(filePath: string, sessionUpdatedMs: number): boolea function writeSessionTranscript( db: Database.Database, session: OpencodeSessionRow, - filePath: string + filePath: string, + redactor: Redactor | null ): void { const messages = db.prepare(` SELECT id, session_id, time_created, time_updated, data @@ -173,9 +175,15 @@ function writeSessionTranscript( })); } + // The staging transcript is a plugin-owned plaintext copy, so redact it at + // write time like the archive (docs/REDACTION.md). + const output = redactor + ? lines.map(line => redactJsonlLine(line, redactor, { source: 'opencode', path: filePath })) + : lines; + fs.mkdirSync(path.dirname(filePath), { recursive: true }); const tempPath = `${filePath}.tmp.${process.pid}`; - fs.writeFileSync(tempPath, `${lines.join('\n')}\n`, 'utf-8'); + fs.writeFileSync(tempPath, `${output.join('\n')}\n`, 'utf-8'); fs.renameSync(tempPath, filePath); const mtime = dateFromMillis(session.time_updated); fs.utimesSync(filePath, mtime, mtime); @@ -184,7 +192,10 @@ function writeSessionTranscript( export function exportOpencodeSessions(options: { dbPath?: string; transcriptDir?: string; + /** `undefined` loads from the environment (strict mode throws on bad rules); `null` disables. */ + redactor?: Redactor | null; } = {}): OpencodeExportResult { + const redactor = options.redactor === undefined ? loadRedactor() : options.redactor; const dbPath = options.dbPath || getOpencodeDbPath(); const transcriptDir = options.transcriptDir || getOpencodeTranscriptDir(); const result: OpencodeExportResult = { @@ -234,7 +245,7 @@ export function exportOpencodeSessions(options: { result.skipped++; continue; } - writeSessionTranscript(db, session, filePath); + writeSessionTranscript(db, session, filePath, redactor); result.exported++; } catch (error) { result.errors.push({ diff --git a/src/redact-cli.ts b/src/redact-cli.ts new file mode 100644 index 00000000..01c33ea4 --- /dev/null +++ b/src/redact-cli.ts @@ -0,0 +1,151 @@ +import fs from 'fs'; +import { getArchiveDir, getCursorLegacyExportDir, getOpencodeTranscriptDir } from './paths.js'; +import { acquireFileLock, readLockHolder, releaseFileLock } from './file-lock.js'; +import { getSyncLockPath } from './logging.js'; +import { + DEFAULT_REDACTION_CONFIG, + FindingsTally, + formatFindings, + getRedactionSettings, + loadRedactor, + redactJsonlLine, + type Redactor, +} from './redaction.js'; + +const args = process.argv.slice(2); + +const HELP = ` +Usage: episodic-memory redact [--rewrite [--dry-run [--report]]] [--stdin] [--print-default-rules] + +Secret redaction for the conversation archive and index. + +OPTIONS: + --rewrite Re-run redaction over the existing archive, staging exports, + and search index (in place). Rows that change are re-embedded, + and summaries built from unredacted text are deleted (the next + sync regenerates them). Run once after upgrading, and again + after adding rules. + --dry-run With --rewrite: report what would change, write nothing. + The index is opened read-only and is not migrated. + --report With --rewrite --dry-run: list every value that would be + redacted: where it is, the rule, its shape (length, + character classes, entropy) and the redacted text around + it. Use it to spot false positives before applying. + --stdin Redact stdin to stdout, one JSONL/text line at a time, and + print rule counts to stderr. Handy for testing rules. + --print-default-rules Print the bundled rules as JSON (a starting point for + redaction-rules.json). + --help, -h Show this help + +ENVIRONMENT: + EPISODIC_MEMORY_REDACTION on (default) | off + EPISODIC_MEMORY_REDACTION_RULES custom rules file (default: /redaction-rules.json) + EPISODIC_MEMORY_REDACTION_STRICT 1 (default) fails closed on a bad rules file | 0 passes through + +Output names rule IDs, counts and value shapes only; matched values are never printed. +`; + +function fail(message: string): never { + console.error(`episodic-memory: ${message}`); + process.exit(1); +} + +function requireRedactor(): Redactor { + if (!getRedactionSettings().enabled) { + fail('redaction is off (EPISODIC_MEMORY_REDACTION=off); unset it to use this command.'); + } + let redactor: Redactor | null; + try { + // Always strict here: a rewrite with no rules would be a silent no-op. + redactor = loadRedactor({ ...process.env, EPISODIC_MEMORY_REDACTION_STRICT: '1' }); + } catch (error) { + fail(error instanceof Error ? error.message : String(error)); + } + return redactor!; +} + +async function runRewrite(dryRun: boolean, report: boolean): Promise { + if (report && !dryRun) fail('--report needs --dry-run: review the hits, then apply without it.'); + const redactor = requireRedactor(); + + // Share sync's single-instance lock so a background sync can't write + // between our reads and renames. + const lockPath = getSyncLockPath(); + const lock = acquireFileLock(lockPath); + if (!lock) { + const holder = readLockHolder(lockPath); + fail(`a sync or index run is in progress (${holder !== null ? `pid ${holder}` : 'another process'}); try again when it finishes.`); + } + const release = () => releaseFileLock(lock); + process.on('exit', release); + + const { rewriteArchive } = await import('./redact-rewrite.js'); + let embeddingsReady = false; + const embed = async (user: string, assistant: string, toolNames?: string[]) => { + const embeddings = await import('./embeddings.js'); + if (!embeddingsReady) { + await embeddings.initEmbeddings(); + embeddingsReady = true; + } + return embeddings.generateExchangeEmbedding(user, assistant, toolNames); + }; + + const archiveDir = getArchiveDir(); + const stagingDirs = [getOpencodeTranscriptDir(), getCursorLegacyExportDir()].filter(d => fs.existsSync(d)); + console.log(`Redacting${dryRun ? ' (dry run)' : ''}: ${archiveDir}`); + for (const dir of stagingDirs) console.log(` + staging: ${dir}`); + + const result = await rewriteArchive({ + archiveDir, + stagingDirs, + redactor, + embed, + dryRun, + report: report + ? hit => console.log(` ${hit.location} ${hit.ruleId} ${hit.shape} + ${hit.context}`) + : undefined, + log: message => console.log(` ${message}`), + }); + + console.log(`\n${dryRun ? 'Would redact' : 'Redacted'} across archive, staging, and index: ${formatFindings(result.findings)}`); + if (dryRun && (result.filesRewritten || result.rowsUpdated || result.summariesRemoved || result.stagingFilesRewritten)) { + console.log('Run without --dry-run to apply.'); + } +} + +async function runStdin(): Promise { + const redactor = requireRedactor(); + const chunks: Buffer[] = []; + for await (const chunk of process.stdin) chunks.push(chunk as Buffer); + const input = Buffer.concat(chunks).toString('utf-8'); + const tally = new FindingsTally(); + const output = input.split('\n').map(line => redactJsonlLine(line, redactor, { source: 'stdin', path: '-' }, tally)); + process.stdout.write(output.join('\n')); + console.error(formatFindings(tally.toArray())); +} + +async function main(): Promise { + if (args.length === 0 || args.includes('--help') || args.includes('-h')) { + console.log(HELP); + return; + } + if (args.includes('--print-default-rules')) { + console.log(JSON.stringify(DEFAULT_REDACTION_CONFIG, null, 2)); + return; + } + if (args.includes('--stdin')) { + await runStdin(); + return; + } + if (args.includes('--rewrite')) { + await runRewrite(args.includes('--dry-run'), args.includes('--report')); + return; + } + fail(`unknown option(s): ${args.join(' ')}. Try: episodic-memory redact --help`); +} + +main().catch(error => { + console.error('Error:', error instanceof Error ? error.message : error); + process.exit(1); +}); diff --git a/src/redact-rewrite.ts b/src/redact-rewrite.ts new file mode 100644 index 00000000..03cede53 --- /dev/null +++ b/src/redact-rewrite.ts @@ -0,0 +1,349 @@ +import fs from 'fs'; +import path from 'path'; +import { initDatabase, openDatabaseReadOnly } from './db.js'; +import { recordReembedded } from './embedding-migration.js'; +import { + copyFileRedacted, + describeShape, + findRedactionTokens, + FindingsTally, + redactJsonlLine, + type RedactionContext, + type RedactionFinding, + type Redactor, +} from './redaction.js'; + +/** + * Backfill for `episodic-memory redact --rewrite`: re-run redaction over data + * written before redaction existed (or before a rule was added). + * + * 1. Archive: every .jsonl is redacted in place. Line count and mtime are + * preserved, so index line ranges stay valid and sync still sees the + * archive as current. + * 2. Staging dirs (opencode / legacy Cursor exports): same, in place. + * 3. Index: every exchanges/tool_calls row is redacted in place. Changed rows + * are re-embedded from the redacted text, so no vector is left that was + * derived from a secret. + * 4. Summaries: a `-summary.txt` for a conversation that had findings (in the + * archive or the index), or one that matches a rule itself, is deleted. + * It was generated from unredacted text, and the next sync regenerates it + * from the redacted archive. + * + * Idempotent: a second run finds nothing to change. A dry run opens the index + * read-only, so it doesn't create, migrate or otherwise touch it. + */ + +export type EmbedFn = (user: string, assistant: string, toolNames?: string[]) => Promise; + +export interface RewriteOptions { + archiveDir: string; + redactor: Redactor; + embed: EmbedFn; + /** Count what would change without writing anything. */ + dryRun?: boolean; + /** + * Called once per value that would be redacted, for reviewing hits before + * applying them. Never receives the value itself. + */ + report?: (hit: RedactionHit) => void; + /** Plugin-owned staging dirs to redact in place (not indexed). */ + stagingDirs?: string[]; + log?: (message: string) => void; +} + +export interface RewriteResult { + filesScanned: number; + filesRewritten: number; + stagingFilesRewritten: number; + rowsUpdated: number; + summariesRemoved: number; + /** Rule IDs and counts only — never matched values. */ + findings: RedactionFinding[]; +} + +/** One redacted value, described without revealing it. */ +export interface RedactionHit { + /** Archive-relative file and 1-based line, or `index:# `. */ + location: string; + ruleId: string; + /** Length, character classes and entropy (see describeShape). */ + shape: string; + /** Redacted text around the token, on one line. */ + context: string; +} + +const SUMMARY_SUFFIX = '-summary.txt'; +const CONTEXT_BEFORE = 60; +const CONTEXT_AFTER = 20; +const PAGE_SIZE = 500; + +function walk(dir: string): string[] { + const out: string[] = []; + let entries: fs.Dirent[]; + try { + entries = fs.readdirSync(dir, { withFileTypes: true }); + } catch { + return out; + } + for (const entry of entries) { + const full = path.join(dir, entry.name); + if (entry.isDirectory()) out.push(...walk(full)); + else if (entry.isFile()) out.push(full); + } + return out; +} + +function summaryPathFor(jsonlPath: string): string { + return jsonlPath.replace(/\.jsonl$/, SUMMARY_SUFFIX); +} + +type Match = { ruleId: string; value: string; absorbedBy?: number }; + +// In report mode each hit is written as its own numbered token, e.g. +// `[REDACTED:jwt--hit3]`, so it can be found in the output exactly; tokens that +// were already in the text have no number. Numbers are stripped for display. +const HIT_TOKEN = /\[REDACTED:([a-z0-9][a-z0-9-]*?)--hit(\d+)\]/g; + +/** `value` with any numbered tokens in it replaced by the values they stand for. */ +function restoreHits(value: string, matches: Match[]): string { + return value.replace(HIT_TOKEN, (_token, _ruleId, n: string) => matches[Number(n)].value); +} + +/** + * Report each hit with the redacted text around its token. A hit whose token + * was swallowed by a later one (a secret field redacted whole after a text rule + * hit part of it) is reported at the token that swallowed it. + */ +function reportHits( + redacted: string, + matches: Match[], + location: string, + report: (hit: RedactionHit) => void +): string { + let display = ''; + let pos = 0; + const spans = new Map(); + for (const t of findRedactionTokens(redacted)) { + const numbered = /^(.*)--hit(\d+)$/.exec(t.ruleId); + display += redacted.slice(pos, t.start); + const shown = numbered ? `[REDACTED:${numbered[1]}]` : redacted.slice(t.start, t.end); + if (numbered && !spans.has(Number(numbered[2]))) { + spans.set(Number(numbered[2]), [display.length, display.length + shown.length]); + } + display += shown; + pos = t.end; + } + display += redacted.slice(pos); + + const spanOf = (n: number): [number, number] | undefined => { + for (let i: number | undefined = n; i !== undefined; i = matches[i].absorbedBy) { + const span = spans.get(i); + if (span) return span; + } + return undefined; + }; + const oneLine = (text: string) => text.replace(/\s+/g, ' '); + const hits = matches.map((m, n) => ({ m, span: spanOf(n) })); + hits.sort((a, b) => (a.span?.[0] ?? Infinity) - (b.span?.[0] ?? Infinity)); + for (const { m, span } of hits) { + const [start, end] = span ?? [0, 0]; + const before = display.slice(Math.max(0, start - CONTEXT_BEFORE), start); + report({ + location, + ruleId: m.ruleId, + shape: describeShape(m.value), + context: span + ? oneLine(`${start > CONTEXT_BEFORE ? '…' : ''}${before}${display.slice(start, end + CONTEXT_AFTER)}`) + : '', + }); + } + return display; +} + +/** + * Redact with `run`, reporting each hit when `report` is set. Returns + * the redacted text with ordinary tokens either way. + */ +function redactReporting( + ctx: RedactionContext, + location: string, + report: ((hit: RedactionHit) => void) | undefined, + run: (ctx: RedactionContext) => string +): string { + if (!report) return run(ctx); + const matches: Match[] = []; + const out = run({ + ...ctx, + onMatch: (ruleId, value) => { + const n = matches.length; + for (const [, , k] of value.matchAll(HIT_TOKEN)) matches[Number(k)].absorbedBy = n; + matches.push({ ruleId, value: restoreHits(value, matches) }); + return `[REDACTED:${ruleId}--hit${n}]`; + }, + }); + return matches.length > 0 ? reportHits(out, matches, location, report) : out; +} + +/** Report every hit in a JSONL file, line by line. Writes nothing. */ +function reportFile(file: string, label: string, redactor: Redactor, report: (hit: RedactionHit) => void): void { + const lines = fs.readFileSync(file, 'utf-8').split('\n'); + const ctx = { source: 'rewrite', path: file }; + lines.forEach((line, i) => { + redactReporting(ctx, `${label}:${i + 1}`, report, c => redactJsonlLine(line, redactor, c)); + }); +} + +/** Redact one JSONL file in place. Returns the number of values redacted. */ +function rewriteFileInPlace(file: string, redactor: Redactor, dryRun: boolean, tally: FindingsTally): number { + const fileTally = new FindingsTally(); + const temp = `${file}.redact.${process.pid}`; + try { + copyFileRedacted(file, temp, redactor, { source: 'rewrite', path: file }, fileTally); + if (fileTally.total > 0 && !dryRun) { + const stat = fs.statSync(file); + fs.renameSync(temp, file); + // Same rounding as sync's copyIfNewer: never leave the archive older than its source. + fs.utimesSync(file, stat.atimeMs / 1000, Math.ceil(stat.mtimeMs) / 1000); + } + } finally { + try { fs.unlinkSync(temp); } catch {} + } + tally.add(fileTally.toArray()); + return fileTally.total; +} + +export async function rewriteArchive(options: RewriteOptions): Promise { + const { archiveDir, redactor, embed } = options; + const dryRun = options.dryRun === true; + // Reporting reads files before redaction; after a real rewrite there'd be nothing to find. + if (options.report && !dryRun) throw new Error('rewriteArchive: report requires dryRun'); + const log = options.log ?? (() => {}); + const tally = new FindingsTally(); + const result: RewriteResult = { + filesScanned: 0, + filesRewritten: 0, + stagingFilesRewritten: 0, + rowsUpdated: 0, + summariesRemoved: 0, + findings: [], + }; + // Conversations whose summary was built from unredacted text. + const staleSummaries = new Set(); + + // 1. Archive files. + const archiveFiles = walk(archiveDir); + for (const file of archiveFiles.filter(f => f.endsWith('.jsonl'))) { + result.filesScanned++; + if (rewriteFileInPlace(file, redactor, dryRun, tally) > 0) { + result.filesRewritten++; + staleSummaries.add(summaryPathFor(file)); + if (options.report) reportFile(file, path.relative(archiveDir, file), redactor, options.report); + } + } + log(`Archive: ${result.filesRewritten} of ${result.filesScanned} file(s) ${dryRun ? 'would be ' : ''}rewritten`); + + // 2. Staging dirs. + for (const dir of options.stagingDirs ?? []) { + for (const file of walk(dir).filter(f => f.endsWith('.jsonl'))) { + if (rewriteFileInPlace(file, redactor, dryRun, tally) > 0) { + result.stagingFilesRewritten++; + if (options.report) reportFile(file, `staging:${path.relative(dir, file)}`, redactor, options.report); + } + } + } + if (options.stagingDirs?.length) { + log(`Staging exports: ${result.stagingFilesRewritten} file(s) ${dryRun ? 'would be ' : ''}rewritten`); + } + + // 3. Index rows. Page by rowid so writes between pages don't disturb the scan. + // A dry run reads the index as it is; opening it normally would migrate it. + const db = dryRun ? openDatabaseReadOnly() : initDatabase(); + if (db) try { + const page = db.prepare( + 'SELECT rowid AS rid, id, user_message, assistant_message, archive_path FROM exchanges WHERE rowid > ? ORDER BY rowid LIMIT ?' + ); + const toolsFor = db.prepare('SELECT id, tool_name, tool_input, tool_result FROM tool_calls WHERE exchange_id = ? ORDER BY rowid'); + const updateExchange = db.prepare('UPDATE exchanges SET user_message = ?, assistant_message = ? WHERE id = ?'); + const updateTool = db.prepare('UPDATE tool_calls SET tool_input = ?, tool_result = ? WHERE id = ?'); + + let lastRowid = 0; + for (;;) { + const rows = page.all(lastRowid, PAGE_SIZE) as Array<{ + rid: number; id: string; user_message: string; assistant_message: string; archive_path: string; + }>; + if (rows.length === 0) break; + lastRowid = rows[rows.length - 1].rid; + + for (const row of rows) { + const rowTally = new FindingsTally(); + const ctx = { source: 'index', path: row.archive_path }; + const where = `index:${path.relative(archiveDir, row.archive_path)}#${row.id}`; + const redactText = (text: string, field: string) => + redactReporting(ctx, `${where} ${field}`, options.report, c => { + const r = redactor.redact(text, c); + rowTally.add(r.findings); + return r.text; + }); + const user = { text: redactText(row.user_message, 'user') }; + const assistant = { text: redactText(row.assistant_message, 'assistant') }; + + const tools = toolsFor.all(row.id) as Array<{ id: string; tool_name: string; tool_input: string | null; tool_result: string | null }>; + const toolUpdates: Array<{ id: string; input: string | null; result: string | null }> = []; + for (const tool of tools) { + // tool_input is JSON text; redactJsonlLine keeps it valid JSON. + const toolInput = tool.tool_input; + const input = toolInput === null ? null : redactReporting( + ctx, `${where} ${tool.tool_name} input`, options.report, + c => redactJsonlLine(toolInput, redactor, c, rowTally) + ); + const output = tool.tool_result === null ? null : redactText(tool.tool_result, `${tool.tool_name} result`); + if (input !== tool.tool_input || output !== tool.tool_result) { + toolUpdates.push({ id: tool.id, input, result: output }); + } + } + + if (rowTally.total === 0) continue; + tally.add(rowTally.toArray()); + result.rowsUpdated++; + staleSummaries.add(summaryPathFor(row.archive_path)); + if (dryRun) continue; + + const toolNames = tools.length > 0 ? tools.map(t => t.tool_name) : undefined; + const embedding = await embed(user.text, assistant.text, toolNames); + db.transaction(() => { + updateExchange.run(user.text, assistant.text, row.id); + for (const t of toolUpdates) updateTool.run(t.input, t.result, t.id); + recordReembedded(db, row.id, embedding); + })(); + } + } + } finally { + db.close(); + } + log(`Index: ${result.rowsUpdated} exchange(s) ${dryRun ? 'would be ' : ''}redacted and re-embedded`); + + // 4. Summaries: stale ones, plus any summary that itself matches a rule. + for (const file of archiveFiles.filter(f => f.endsWith(SUMMARY_SUFFIX))) { + if (staleSummaries.has(file)) continue; + let text: string; + try { + text = fs.readFileSync(file, 'utf-8'); + } catch { + continue; + } + const r = redactor.redact(text, { source: 'summary', path: file }); + if (r.findings.length > 0) { + tally.add(r.findings); + staleSummaries.add(file); + } + } + for (const summary of staleSummaries) { + if (!fs.existsSync(summary)) continue; + result.summariesRemoved++; + if (!dryRun) fs.unlinkSync(summary); + } + log(`Summaries: ${result.summariesRemoved} ${dryRun ? 'would be ' : ''}removed (regenerated from redacted text on the next sync)`); + + result.findings = tally.toArray(); + return result; +} diff --git a/src/redaction-rules.ts b/src/redaction-rules.ts new file mode 100644 index 00000000..7f5490b0 --- /dev/null +++ b/src/redaction-rules.ts @@ -0,0 +1,272 @@ +import type { RedactionConfig } from './redaction.js'; + +/** + * Bundled default redaction rules. + * + * Patterns are ported from gitleaks' rule set (https://github.com/gitleaks/gitleaks, + * MIT), trimmed to the credentials that realistically show up in Claude Code / + * Codex transcripts for an Azure-heavy stack, and rewritten for JavaScript + * regex (lookbehind instead of gitleaks' consuming boundary groups, so + * adjacent matches aren't swallowed). + * + * Order matters: rules run top to bottom, and a later rule never re-matches + * inside an earlier rule's `[REDACTED:...]` token. Specific shapes go first so + * the token names the most precise rule; the generic `secret-assignment` + * keyword rule goes last. + * + * `secretGroup` replaces only that capture group, so surrounding context + * (connection-string server names, SAS URL paths, usernames) stays searchable. + * + * `keywords` is a case-insensitive prefilter: a rule only runs on text that + * contains at least one keyword. It keeps the per-string cost low on large + * transcripts and bounds the generic rules' work. + * + * Users extend or override these with `redaction-rules.json`; see + * docs/REDACTION.md. `episodic-memory redact --print-default-rules` dumps + * this object as JSON. + */ +// Names that mark the value next to them as a secret, for the key-context +// rules below. A name must END in one of these (optionally plus digits), so +// tokenType, secretName, passwordPolicy, TokenEndpoint and maxTokens don't +// count. `pwd` is deliberately absent (the PWD env var); `userPWD` is handled +// by xml-secret-attribute. +const SECRET_NAME = + String.raw`(?:secret|password|passwd|passphrase|api[_-]?key|access[_-]?key|account[_-]?key|private[_-]?key|` + + String.raw`shared[_-]?(?:access[_-]?)?key|primary[_-]?key|secondary[_-]?key|master[_-]?key|signing[_-]?key|` + + String.raw`subscription[_-]?key|client[_-]?key|encryption[_-]?key|token|credentials?)\d*`; + +// Value is not a template/placeholder: ${X} $(X) $X #{X}# {{x}} %X% __X__, +// an existing token, a type name (Swagger "string"), or a mask (*****). +const NOT_PLACEHOLDER = + String.raw`(?!\[REDACTED:|\$\{|\$\(|\$[A-Za-z_]\w*["'<\s]|#\{|\{\{|<[^>]*>|%[A-Za-z_]\w*%|__[A-Za-z0-9_]+__|` + + String.raw`(?:string|null|true|false|none|undefined|\*+)["'<])`; + +// A JSON string value (escapes allowed), captured. +const JSON_STRING_VALUE = String.raw`"${NOT_PLACEHOLDER}([^"\\]+(?:\\.[^"\\]*)*)"`; +// Sibling members inside the same object: strings, or anything but braces/quotes. +const SAME_OBJECT = String.raw`(?:[^{}"]|"(?:[^"\\]|\\.)*"){0,400}?`; +const KEY_VAULT_ID = + String.raw`"id"[ \t]*:[ \t]*"https://[^"\s]+\.vault\.(?:azure\.net|azure\.cn|usgovcloudapi\.net|microsoftazure\.de)/secrets/[^"\s]*"`; + +export const DEFAULT_REDACTION_CONFIG: RedactionConfig = { + rules: [ + { + id: 'private-key-block', + description: 'PEM/OpenSSH private key block; a truncated block (no END line) is redacted to the end of the text', + pattern: String.raw`-----BEGIN[A-Z0-9 _-]{0,40}PRIVATE KEY(?: BLOCK)?-----[\s\S]*?(?:-----END[A-Z0-9 _-]{0,40}PRIVATE KEY(?: BLOCK)?-----|$)`, + keywords: ['private key'], + }, + { + id: 'connection-string-secret', + description: 'Secret value inside an Azure/ADO.NET connection string; account, server and database names are kept. Keywords are case-insensitive, as in ADO.NET; a path after Pwd= is the shell PWD variable and is kept; the value needs a letter or digit, so Markdown like `Password=` is kept', + pattern: String.raw`\b(?:(?:AccountKey|SharedAccessKey|SharedAccessSignature|SharedSecret|ClientSecret|Password)=|Pwd=(?![/~]))(?![\[$<{%])(?=[^;"'\s]*[A-Za-z0-9])([^;"'\s]+)`, + flags: 'i', + secretGroup: 1, + keywords: ['accountkey', 'sharedaccess', 'sharedsecret', 'clientsecret', 'password', 'pwd'], + }, + { + id: 'azure-sas-token', + description: 'Signature of an Azure SAS token; the resource URL and other SAS parameters are kept', + pattern: String.raw`\bsig=([A-Za-z0-9%+/=_-]{16,})`, + secretGroup: 1, + keywords: ['sig='], + }, + { + id: 'jwt', + description: 'JSON Web Token (Entra ID / Azure access tokens, id tokens)', + pattern: String.raw`\beyJ[A-Za-z0-9_-]{8,}\.eyJ[A-Za-z0-9_-]{8,}\.[A-Za-z0-9_-]{8,}`, + keywords: ['eyj'], + }, + { + id: 'anthropic-api-key', + pattern: String.raw`\bsk-ant-[a-z]{2,10}\d{2}-[A-Za-z0-9_-]{32,}`, + keywords: ['sk-ant-'], + }, + { + id: 'openai-api-key', + pattern: String.raw`\bsk-(?:[a-z]+-)?[A-Za-z0-9_-]{16,}T3BlbkFJ[A-Za-z0-9_-]{16,}`, + keywords: ['t3blbkfj'], + }, + { + id: 'github-token', + description: 'GitHub classic, OAuth, app and fine-grained tokens', + pattern: String.raw`\b(?:gh[pousr]_[A-Za-z0-9]{36,255}|github_pat_[A-Za-z0-9_]{22,255})\b`, + keywords: ['ghp_', 'gho_', 'ghu_', 'ghs_', 'ghr_', 'github_pat_'], + }, + { + id: 'aws-access-key-id', + pattern: String.raw`\b(?:AKIA|ASIA|ABIA|ACCA)[A-Z2-7]{16}\b`, + keywords: ['akia', 'asia', 'abia', 'acca'], + }, + { + id: 'aws-secret-access-key', + pattern: String.raw`aws[_-]?secret[_-]?access[_-]?key["']?[ \t]*[:=][ \t]*["']?([A-Za-z0-9/+]{40})(?![A-Za-z0-9/+])`, + flags: 'i', + secretGroup: 1, + keywords: ['secret'], + }, + { + id: 'slack-token', + pattern: String.raw`\bxox[abposr]-[A-Za-z0-9-]{10,}`, + keywords: ['xox'], + }, + { + id: 'google-api-key', + pattern: String.raw`\bAIza[0-9A-Za-z_-]{35}(?![0-9A-Za-z_-])`, + keywords: ['aiza'], + }, + { + id: 'npm-token', + pattern: String.raw`\bnpm_[A-Za-z0-9]{36}\b`, + keywords: ['npm_'], + }, + { + id: 'azure-client-secret', + description: 'Entra ID (Azure AD) application client secret: 3 chars, a digit, "Q~", 31-34 chars. A trailing "." may follow, so a secret ending a sentence still matches', + pattern: String.raw`(?]+:(?![\[$<{%])([^\s@/"'<>]+)@`, + secretGroup: 1, + keywords: ['://'], + }, + { + id: 'bearer-token', + pattern: String.raw`\bBearer\s+([A-Za-z0-9\-._~+/]{20,}=*)`, + flags: 'i', + secretGroup: 1, + keywords: ['bearer'], + }, + { + id: 'basic-auth', + pattern: String.raw`\bAuthorization[ \t]*:[ \t]*Basic\s+([A-Za-z0-9+/]{8,}={0,2})`, + flags: 'i', + secretGroup: 1, + keywords: ['basic'], + }, + { + id: 'azure-keyvault-secret', + description: 'The "value" of a Key Vault secret bundle (az keyvault secret show/set, SDK JSON), whatever the secret is named', + pattern: + KEY_VAULT_ID + String.raw`(?:[^{}]|\{[^{}]*\}){0,800}?"value"[ \t]*:[ \t]*` + JSON_STRING_VALUE + + String.raw`|"value"[ \t]*:[ \t]*` + JSON_STRING_VALUE + String.raw`(?=(?:[^{}]|\{[^{}]*\}){0,800}?` + KEY_VAULT_ID + ')', + secretGroup: [1, 2], + useAllowlist: false, + keywords: ['.vault.'], + }, + { + id: 'name-value-secret', + description: + 'The "value" of a {"name"/"key": , "value": ...} object, in either order: ' + + 'az webapp/functionapp config appsettings list, Kubernetes env, ARM/Bicep parameters', + pattern: + String.raw`"(?:name|key)"[ \t]*:[ \t]*"[^"\\]{0,80}?` + SECRET_NAME + '"' + SAME_OBJECT + + String.raw`"value"[ \t]*:[ \t]*` + JSON_STRING_VALUE + + String.raw`|"value"[ \t]*:[ \t]*` + JSON_STRING_VALUE + '(?=' + SAME_OBJECT + + String.raw`"(?:name|key)"[ \t]*:[ \t]*"[^"\\]{0,80}?` + SECRET_NAME + '")', + flags: 'i', + secretGroup: [1, 2], + useAllowlist: false, + keywords: ['"value"'], + }, + { + id: 'xml-appsettings-secret', + description: 'web.config / app.config , either attribute order', + pattern: + String.raw`]*?\bkey[ \t]*=[ \t]*"[^"]{0,80}?` + SECRET_NAME + String.raw`"[^>]*?\bvalue[ \t]*=[ \t]*"` + + NOT_PLACEHOLDER + String.raw`([^"]+)"` + + String.raw`|]*?\bvalue[ \t]*=[ \t]*"` + NOT_PLACEHOLDER + String.raw`([^"]+)"(?=[^>]*?\bkey[ \t]*=[ \t]*"[^"]{0,80}?` + + SECRET_NAME + '")', + flags: 'i', + secretGroup: [1, 2], + useAllowlist: false, + keywords: [']*)?>` + NOT_PLACEHOLDER + String.raw`([^<]{1,4096})`, + flags: 'i', + secretGroup: 2, + useAllowlist: false, + keywords: [', %X%, {{x}}) and code (calls, dotted member access)', + pattern: + String.raw`\b[A-Za-z0-9_.-]{0,40}?(?:secret|password|passwd|passphrase|api[_-]?key|access[_-]?key|private[_-]?key|token)` + + String.raw`["']?[ \t]*[:=][ \t]*["']?(?![\[$<{%])(?=[^\s"',;&\\]*\d)` + + // Not a dotted identifier path (env.AUTH0_SECRET, this.config.token2): code, not a value. + String.raw`(?![A-Za-z_$][\w$]*(?:\.[A-Za-z_$][\w$]*)+(?:$|[\s"',;&\\)}\]]))` + + String.raw`([^\s"',;&\\()<>{}\[\]]{8,})(?=$|[\s"',;&\\)}\]])`, + flags: 'i', + secretGroup: 1, + keywords: ['secret', 'passw', 'passphrase', 'key', 'token'], + }, + ], + allowlist: [ + { + id: 'git-sha', + description: 'Full or short git commit SHA', + pattern: String.raw`\b[0-9a-f]{7,40}\b`, + }, + { + id: 'guid', + description: 'GUID/UUID (Entra tenant, client and object IDs, subscription IDs)', + pattern: String.raw`\b[0-9a-fA-F]{8}-[0-9a-fA-F]{4}-[0-9a-fA-F]{4}-[0-9a-fA-F]{4}-[0-9a-fA-F]{12}\b`, + }, + ], + entropy: { + enabled: false, + minLength: 32, + threshold: 4.5, + requireKeyword: true, + keywords: ['secret', 'key', 'token', 'password', 'passwd', 'credential', 'signature'], + window: 40, + }, + secretFields: { + enabled: true, + // Normalized key (lowercase letters/digits only, trailing digits dropped) + // must end with one of these. Harness keys (signature, apiKeySource, + // input_tokens, token_count) don't. + keyPattern: + 'secret|password|passwd|userpwd|passphrase|apikey|accesskey|accountkey|privatekey|sharedkey|primarykey|' + + 'secondarykey|masterkey|signingkey|subscriptionkey|clientkey|encryptionkey|token|credentials?', + }, +}; diff --git a/src/redaction.ts b/src/redaction.ts new file mode 100644 index 00000000..999f0136 --- /dev/null +++ b/src/redaction.ts @@ -0,0 +1,849 @@ +import fs from 'fs'; +import path from 'path'; +import { StringDecoder } from 'string_decoder'; +import { getSuperpowersDir } from './paths.js'; +import { DEFAULT_REDACTION_CONFIG } from './redaction-rules.js'; + +/** + * Inline secret redaction. + * + * Secrets are replaced with typed tokens (`[REDACTED:]`) at the archive + * write, the one point every harness's transcripts pass through (see + * docs/REDACTION.md). Everything downstream (SQLite text, + * tool_calls, embeddings, summaries, show/read) reads the archive, so it only + * ever sees redacted text. + * + * Invariants: + * - Findings carry rule IDs and counts only. A matched value is never logged, + * returned, or thrown. + * - Idempotent: a `[REDACTED:...]` token never matches a rule, so redacting + * redacted text is a no-op. + * - JSONL structure is preserved: one output line per input line, valid JSON in + * gives valid JSON out, and unchanged lines are byte-identical. + */ + +export { DEFAULT_REDACTION_CONFIG }; + +/** Fork default. Flip to false for an opt-in upstream build. */ +export const REDACTION_ENABLED_BY_DEFAULT = true; + +export const REDACTION_RULES_FILENAME = 'redaction-rules.json'; + +const RULE_ID_PATTERN = /^[a-z0-9][a-z0-9-]*$/; +const TOKEN_PATTERN = /\[REDACTED:[a-z0-9][a-z0-9-]*\]/g; +const TOKEN_PREFIX = '[REDACTED:'; +const ALLOWED_FLAGS = /^[imsu]*$/; +const ENTROPY_RULE_ID = 'high-entropy'; +const FIELD_RULE_ID = 'secret-field'; +const RESERVED_IDS = new Set([ENTROPY_RULE_ID, FIELD_RULE_ID]); + +/** Values that are templates or type names, not secrets. Never redacted by field context. */ +const PLACEHOLDER_VALUE = + /^(?:\$\{.*\}|\$\(.*\)|\$[A-Za-z_]\w*|#\{.*\}#?|\{\{.*\}\}|<[^>]*>|%[A-Za-z_]\w*%|__[A-Za-z0-9_]+__|string|null|true|false|none|undefined|\*+)$/i; +const KEY_VAULT_SECRET_ID = /^https:\/\/[^/\s]+\.vault\.(?:azure\.net|azure\.cn|usgovcloudapi\.net|microsoftazure\.de)\/secrets\//i; + +export interface RedactionRuleSpec { + id: string; + pattern: string; + /** Extra RegExp flags from [imsu]; `g` and `d` are always added. */ + flags?: string; + /** Case-insensitive prefilter: skip the rule unless the text contains one. */ + keywords?: string[]; + /** + * Replace only this capture group instead of the whole match. An array means + * "the first of these groups that participated" (for either-order patterns). + */ + secretGroup?: number | number[]; + /** Apply the allowlist to this rule's matches (default true). Key-context rules turn it off: a GUID in a password slot is a secret. */ + useAllowlist?: boolean; + description?: string; +} + +export interface AllowlistSpec { + id: string; + /** Matched against the whole candidate secret; a full match keeps it. */ + pattern: string; + flags?: string; + description?: string; +} + +export interface EntropySpec { + enabled: boolean; + minLength: number; + /** Shannon entropy in bits per character. */ + threshold: number; + /** Only fire when a keyword appears shortly before the candidate, on the same line. */ + requireKeyword: boolean; + keywords: string[]; + /** How many characters before the candidate to search for a keyword. */ + window: number; +} + +/** + * Field-name context for parsed JSON (transcript lines, structured tool/MCP + * results, tool inputs). A string value is redacted whole when its own key, or + * the `name`/`key` of a `{name, value}` pair, matches `keyPattern`. + */ +export interface SecretFieldsSpec { + enabled: boolean; + /** + * Matched (anchored at the end) against the key lowercased with everything + * but letters and digits removed and trailing digits dropped, so + * `AzureAd:ClientSecret`, `client_secret` and `DB_PASSWORD2` all normalize to + * something ending in a keyword. + */ + keyPattern: string; +} + +export interface RedactionConfig { + rules: RedactionRuleSpec[]; + allowlist: AllowlistSpec[]; + entropy: EntropySpec; + secretFields: SecretFieldsSpec; +} + +/** Shape of a user `redaction-rules.json`. Every field is optional. */ +export interface RedactionRulesFile { + /** Merge with the bundled defaults (default true). */ + includeDefaults?: boolean; + /** Added after the defaults; a rule with a default's id replaces it in place. */ + rules?: RedactionRuleSpec[]; + /** Default rule ids to turn off. */ + disableRules?: string[]; + /** Added to the default allowlist (same id replaces). */ + allowlist?: AllowlistSpec[]; + entropy?: Partial; + secretFields?: Partial; +} + +export interface RedactionContext { + source: string; + path: string; + /** + * Called with each value as it is redacted. For local review tooling + * (`redact --report`) only: the value is the secret itself, so never log it. + * May return the token to write instead of `[REDACTED:]`, e.g. one + * numbered per hit; it must still have the token's shape. + */ + onMatch?: (ruleId: string, value: string) => string | void; +} + +export interface RedactionFinding { + ruleId: string; + count: number; +} + +export interface RedactionResult { + text: string; + findings: RedactionFinding[]; +} + +export interface Redactor { + redact(text: string, ctx?: RedactionContext): RedactionResult; + readonly ruleIds: string[]; + /** True when a JSON key / setting name marks its value as a secret (secretFields). */ + isSecretField(name: string): boolean; +} + +export interface RedactionSettings { + enabled: boolean; + strict: boolean; + /** Explicit EPISODIC_MEMORY_REDACTION_RULES path, if set. */ + rulesPath?: string; +} + +/** Rules could not be loaded or are invalid. Messages describe config only. */ +export class RedactionConfigError extends Error { + constructor(message: string) { + super(message); + this.name = 'RedactionConfigError'; + } +} + +// --------------------------------------------------------------------------- +// Settings and loading +// --------------------------------------------------------------------------- + +const OFF_VALUES = new Set(['off', '0', 'false', 'no', 'disabled']); +const ON_VALUES = new Set(['on', '1', 'true', 'yes', 'enabled']); + +/** Unknown values fall back to the default, which is the safe side for both switches. */ +function parseToggle(raw: string | undefined, fallback: boolean): boolean { + const value = raw?.trim().toLowerCase(); + if (!value) return fallback; + if (OFF_VALUES.has(value)) return false; + if (ON_VALUES.has(value)) return true; + return fallback; +} + +/** + * EPISODIC_MEMORY_REDACTION on (fork default) | off + * EPISODIC_MEMORY_REDACTION_RULES path to a custom rules file + * EPISODIC_MEMORY_REDACTION_STRICT 1 (default) | 0; strict fails closed on a rules-load error + */ +export function getRedactionSettings(env: NodeJS.ProcessEnv = process.env): RedactionSettings { + return { + enabled: parseToggle(env.EPISODIC_MEMORY_REDACTION, REDACTION_ENABLED_BY_DEFAULT), + strict: parseToggle(env.EPISODIC_MEMORY_REDACTION_STRICT, true), + rulesPath: env.EPISODIC_MEMORY_REDACTION_RULES || undefined, + }; +} + +function readRulesFile(settings: RedactionSettings): RedactionRulesFile { + let rulesPath = settings.rulesPath; + if (!rulesPath) { + const candidate = path.join(getSuperpowersDir(), REDACTION_RULES_FILENAME); + if (!fs.existsSync(candidate)) return {}; + rulesPath = candidate; + } + let raw: string; + try { + raw = fs.readFileSync(rulesPath, 'utf-8'); + } catch (error) { + throw new RedactionConfigError( + `Redaction rules failed to load from ${rulesPath}: ${error instanceof Error ? error.message : String(error)}` + ); + } + try { + return JSON.parse(raw) as RedactionRulesFile; + } catch (error) { + throw new RedactionConfigError( + `Redaction rules file ${rulesPath} is not valid JSON: ${error instanceof Error ? error.message : String(error)}` + ); + } +} + +function replaceById(base: T[], additions: T[]): T[] { + const out = [...base]; + for (const item of additions) { + const at = out.findIndex(existing => existing.id === item.id); + if (at >= 0) out[at] = item; + else out.push(item); + } + return out; +} + +/** + * Merge a rules file with the bundled defaults and validate the result. + * Throws RedactionConfigError on any invalid rule. + */ +export function loadRedactionConfig(file: RedactionRulesFile): RedactionConfig { + if (!file || typeof file !== 'object' || Array.isArray(file)) { + throw new RedactionConfigError('Redaction rules file must contain a JSON object'); + } + for (const key of ['rules', 'allowlist', 'disableRules'] as const) { + if (file[key] !== undefined && !Array.isArray(file[key])) { + throw new RedactionConfigError(`Redaction rules file: "${key}" must be an array`); + } + } + for (const key of ['entropy', 'secretFields'] as const) { + if (file[key] !== undefined && (typeof file[key] !== 'object' || file[key] === null || Array.isArray(file[key]))) { + throw new RedactionConfigError(`Redaction rules file: "${key}" must be an object`); + } + } + + const includeDefaults = file.includeDefaults !== false; + const disabled = new Set(file.disableRules ?? []); + const baseRules = includeDefaults ? DEFAULT_REDACTION_CONFIG.rules : []; + const baseAllowlist = includeDefaults ? DEFAULT_REDACTION_CONFIG.allowlist : []; + + const config: RedactionConfig = { + rules: replaceById(baseRules, file.rules ?? []).filter(rule => !disabled.has(rule.id)), + allowlist: replaceById(baseAllowlist, file.allowlist ?? []), + entropy: { ...DEFAULT_REDACTION_CONFIG.entropy, ...(file.entropy ?? {}) }, + secretFields: { ...DEFAULT_REDACTION_CONFIG.secretFields, ...(file.secretFields ?? {}) }, + }; + compileConfig(config); // validate + return config; +} + +/** + * Load the active redactor from the environment. + * + * Returns null when redaction is off. When the rules can't be loaded: strict + * mode (the default) throws RedactionConfigError so callers fail closed; + * non-strict mode warns and returns null, so text passes through unredacted. + */ +export function loadRedactor( + env: NodeJS.ProcessEnv = process.env, + warn: (message: string) => void = message => console.error(message) +): Redactor | null { + const settings = getRedactionSettings(env); + if (!settings.enabled) return null; + try { + return createRedactor(loadRedactionConfig(readRulesFile(settings))); + } catch (error) { + const err = error instanceof RedactionConfigError + ? error + : new RedactionConfigError(`Redaction rules failed to load: ${error instanceof Error ? error.message : String(error)}`); + if (settings.strict) throw err; + warn( + `episodic-memory: ${err.message}. EPISODIC_MEMORY_REDACTION_STRICT=0, so conversations ` + + 'will be archived and indexed WITHOUT redaction this run.' + ); + return null; + } +} + +// --------------------------------------------------------------------------- +// Compilation +// --------------------------------------------------------------------------- + +interface CompiledRule { + id: string; + regex: RegExp; + keywords?: string[]; + secretGroups: number[]; + useAllowlist: boolean; +} + +interface CompiledConfig { + rules: CompiledRule[]; + allowlist: RegExp[]; + entropy: EntropySpec; + /** null when secretFields is disabled. */ + secretField: RegExp | null; +} + +function checkFlags(flags: string | undefined, what: string): string { + const value = flags ?? ''; + if (typeof value !== 'string' || !ALLOWED_FLAGS.test(value)) { + throw new RedactionConfigError(`${what}: flags must be a combination of "imsu"`); + } + return value; +} + +function compileRule(spec: RedactionRuleSpec): CompiledRule { + if (!spec || typeof spec !== 'object') { + throw new RedactionConfigError('Redaction rule must be an object'); + } + if (typeof spec.id !== 'string' || !RULE_ID_PATTERN.test(spec.id)) { + throw new RedactionConfigError( + `Redaction rule id ${JSON.stringify(spec.id)} is invalid: use lowercase letters, digits and dashes` + ); + } + if (RESERVED_IDS.has(spec.id)) { + throw new RedactionConfigError(`Redaction rule id "${spec.id}" is reserved`); + } + const what = `Redaction rule "${spec.id}"`; + if (typeof spec.pattern !== 'string' || spec.pattern.length === 0) { + throw new RedactionConfigError(`${what}: pattern must be a non-empty string`); + } + const flags = checkFlags(spec.flags, what); + + let regex: RegExp; + let groupCount: number; + try { + regex = new RegExp(spec.pattern, flags + 'gd'); + groupCount = new RegExp(`(?:${spec.pattern})|`, flags).exec('')!.length - 1; + } catch (error) { + throw new RedactionConfigError(`${what}: invalid regular expression (${error instanceof Error ? error.message : String(error)})`); + } + if (new RegExp(spec.pattern, flags).test('')) { + throw new RedactionConfigError(`${what}: pattern matches the empty string`); + } + + const secretGroups = Array.isArray(spec.secretGroup) ? spec.secretGroup : [spec.secretGroup ?? 0]; + if (secretGroups.length === 0) { + throw new RedactionConfigError(`${what}: secretGroup must not be an empty array`); + } + for (const group of secretGroups) { + if (!Number.isInteger(group) || group < 0 || group > groupCount) { + throw new RedactionConfigError(`${what}: secretGroup ${group} does not exist in the pattern (${groupCount} group(s))`); + } + } + if (spec.useAllowlist !== undefined && typeof spec.useAllowlist !== 'boolean') { + throw new RedactionConfigError(`${what}: useAllowlist must be a boolean`); + } + + let keywords: string[] | undefined; + if (spec.keywords !== undefined) { + if (!Array.isArray(spec.keywords) || spec.keywords.some(k => typeof k !== 'string' || k.length === 0)) { + throw new RedactionConfigError(`${what}: keywords must be an array of non-empty strings`); + } + keywords = spec.keywords.length > 0 ? spec.keywords.map(k => k.toLowerCase()) : undefined; + } + + return { id: spec.id, regex, keywords, secretGroups, useAllowlist: spec.useAllowlist !== false }; +} + +function compileAllowlist(spec: AllowlistSpec): RegExp { + const what = `Redaction allowlist entry ${JSON.stringify(spec?.id)}`; + if (!spec || typeof spec.pattern !== 'string' || spec.pattern.length === 0) { + throw new RedactionConfigError(`${what}: pattern must be a non-empty string`); + } + const flags = checkFlags(spec.flags, what); + try { + return new RegExp(`^(?:${spec.pattern})$`, flags); + } catch (error) { + throw new RedactionConfigError(`${what}: invalid regular expression (${error instanceof Error ? error.message : String(error)})`); + } +} + +function compileConfig(config: RedactionConfig): CompiledConfig { + const seen = new Set(); + const rules = config.rules.map(spec => { + const rule = compileRule(spec); + if (seen.has(rule.id)) throw new RedactionConfigError(`Duplicate redaction rule id "${rule.id}"`); + seen.add(rule.id); + return rule; + }); + const allowlist = config.allowlist.map(compileAllowlist); + + const e = config.entropy; + if ( + typeof e.enabled !== 'boolean' || typeof e.requireKeyword !== 'boolean' || + !Number.isInteger(e.minLength) || e.minLength < 8 || + typeof e.threshold !== 'number' || !(e.threshold > 0) || + !Number.isInteger(e.window) || e.window < 0 || + !Array.isArray(e.keywords) || e.keywords.some(k => typeof k !== 'string' || !k) + ) { + throw new RedactionConfigError( + 'Redaction entropy settings are invalid (need enabled, requireKeyword: boolean; minLength >= 8; threshold > 0; window >= 0; keywords: string[])' + ); + } + const f = config.secretFields; + if (!f || typeof f.enabled !== 'boolean' || typeof f.keyPattern !== 'string' || f.keyPattern.length === 0) { + throw new RedactionConfigError('Redaction secretFields settings are invalid (need enabled: boolean; keyPattern: non-empty string)'); + } + let secretField: RegExp | null = null; + if (f.enabled) { + try { + secretField = new RegExp(`(?:${f.keyPattern})$`); + } catch (error) { + throw new RedactionConfigError(`Redaction secretFields.keyPattern: invalid regular expression (${error instanceof Error ? error.message : String(error)})`); + } + if (secretField.test('')) throw new RedactionConfigError('Redaction secretFields.keyPattern matches the empty string'); + } + + return { rules, allowlist, entropy: { ...e, keywords: e.keywords.map(k => k.toLowerCase()) }, secretField }; +} + +// --------------------------------------------------------------------------- +// Redaction engine +// --------------------------------------------------------------------------- + +function tokenFor(ruleId: string): string { + return `${TOKEN_PREFIX}${ruleId}]`; +} + +const WHOLE_TOKEN = /^\[REDACTED:[a-z0-9][a-z0-9-]*\]$/; + +/** The token for one redacted value: onMatch's, if it returns a valid one. */ +function tokenForMatch(ruleId: string, value: string, onMatch: RedactionContext['onMatch']): string { + const custom = onMatch?.(ruleId, value); + return typeof custom === 'string' && WHOLE_TOKEN.test(custom) ? custom : tokenFor(ruleId); +} + +function tokenSpans(text: string): Array<[number, number]> { + const spans: Array<[number, number]> = []; + for (const m of text.matchAll(TOKEN_PATTERN)) spans.push([m.index!, m.index! + m[0].length]); + return spans; +} + +function overlapsAny(spans: Array<[number, number]>, start: number, end: number): boolean { + for (const [s, e] of spans) if (start < e && end > s) return true; + return false; +} + +/** + * Rewrite text[start, end) keeping the existing tokens in it and replacing each + * stretch between them that has a letter or digit with `token`. Returns null + * when nothing outside the tokens needs redacting, which keeps this idempotent. + */ +function redactAroundTokens( + text: string, + spans: Array<[number, number]>, + start: number, + end: number, + token: () => string +): string | null { + let out = ''; + let pos = start; + let replaced = false; + const flush = (to: number) => { + const piece = text.slice(pos, to); + if (/[A-Za-z0-9]/.test(piece)) { + out += token(); + replaced = true; + } else { + out += piece; + } + }; + for (const [s, e] of spans) { + if (e <= start || s >= end) continue; + if (s > pos) flush(s); + out += text.slice(Math.max(s, pos), Math.min(e, end)); + pos = Math.min(e, end); + } + if (pos < end) flush(end); + return replaced ? out : null; +} + +/** + * Describe a value without revealing it: length, character classes and + * Shannon entropy (bits per char), e.g. "len=40 aA9- H=4.9". Lets someone + * reviewing hits tell a random key from a word or a placeholder. + */ +export function describeShape(value: string): string { + const classes = + (/[a-z]/.test(value) ? 'a' : '') + + (/[A-Z]/.test(value) ? 'A' : '') + + (/[0-9]/.test(value) ? '9' : '') + + (/[^A-Za-z0-9\s]/.test(value) ? '-' : '') + + (/\s/.test(value) ? '_' : ''); + return `len=${value.length} ${classes || '?'} H=${shannonEntropy(value).toFixed(1)}`; +} + +/** Each `[REDACTED:]` token in `text`, left to right. */ +export function findRedactionTokens(text: string): Array<{ ruleId: string; start: number; end: number }> { + const out: Array<{ ruleId: string; start: number; end: number }> = []; + for (const m of text.matchAll(TOKEN_PATTERN)) { + out.push({ ruleId: m[0].slice(TOKEN_PREFIX.length, -1), start: m.index!, end: m.index! + m[0].length }); + } + return out; +} + +function shannonEntropy(text: string): number { + const counts = new Map(); + for (const ch of text) counts.set(ch, (counts.get(ch) ?? 0) + 1); + let entropy = 0; + for (const n of counts.values()) { + const p = n / text.length; + entropy -= p * Math.log2(p); + } + return entropy; +} + +/** + * Replace each accepted match (or its secretGroup) with a token. A candidate + * that overlaps existing tokens keeps them, and only the text around them is + * redacted (so re-running is a no-op). A candidate is skipped when it fully + * matches an allowlist pattern. + */ +function applyRule( + text: string, + rule: CompiledRule, + isAllowed: (s: string) => boolean, + onMatch?: RedactionContext['onMatch'] +): { text: string; count: number } { + const spans = text.includes(TOKEN_PREFIX) ? tokenSpans(text) : null; + let out = ''; + let last = 0; + let count = 0; + for (const m of text.matchAll(rule.regex)) { + let range: [number, number] | undefined; + for (const group of rule.secretGroups) { + range = m.indices?.[group]; + if (range) break; + } + if (!range) continue; + const [start, end] = range; + if (end <= start || start < last) continue; + if (rule.useAllowlist && isAllowed(text.slice(start, end))) continue; + if (spans && overlapsAny(spans, start, end)) { + let token: string | undefined; + const rewritten = redactAroundTokens(text, spans, start, end, () => + (token ??= tokenForMatch(rule.id, text.slice(start, end).replace(TOKEN_PATTERN, ''), onMatch))); + if (rewritten === null) continue; + out += text.slice(last, start) + rewritten; + last = end; + count++; + continue; + } + out += text.slice(last, start) + tokenForMatch(rule.id, text.slice(start, end), onMatch); + last = end; + count++; + } + return count === 0 ? { text, count } : { text: out + text.slice(last), count }; +} + +function applyEntropy( + text: string, + spec: EntropySpec, + isAllowed: (s: string) => boolean, + onMatch?: RedactionContext['onMatch'] +): { text: string; count: number } { + const candidates = new RegExp(`[A-Za-z0-9+/=_.~-]{${spec.minLength},}`, 'g'); + const spans = text.includes(TOKEN_PREFIX) ? tokenSpans(text) : null; + let out = ''; + let last = 0; + let count = 0; + for (const m of text.matchAll(candidates)) { + const start = m.index!; + const end = start + m[0].length; + if (spans && overlapsAny(spans, start, end)) continue; + if (isAllowed(m[0])) continue; + if (shannonEntropy(m[0]) < spec.threshold) continue; + if (spec.requireKeyword) { + let before = text.slice(Math.max(0, start - spec.window), start); + before = before.slice(before.lastIndexOf('\n') + 1).toLowerCase(); + if (!spec.keywords.some(k => before.includes(k))) continue; + } + out += text.slice(last, start) + tokenForMatch(ENTROPY_RULE_ID, m[0], onMatch); + last = end; + count++; + } + return count === 0 ? { text, count } : { text: out + text.slice(last), count }; +} + +export function createRedactor(config: RedactionConfig): Redactor { + const compiled = compileConfig(config); + const isAllowed = (secret: string) => compiled.allowlist.some(re => re.test(secret)); + + return { + ruleIds: compiled.rules.map(r => r.id), + isSecretField(name: string): boolean { + if (!compiled.secretField || typeof name !== 'string') return false; + const normalized = name.toLowerCase().replace(/[^a-z0-9]/g, '').replace(/\d+$/, ''); + return normalized.length > 0 && compiled.secretField.test(normalized); + }, + redact(text: string, ctx?: RedactionContext): RedactionResult { + if (typeof text !== 'string' || text.length === 0) return { text, findings: [] }; + let current = text; + let lower: string | null = null; + const counts = new Map(); + + for (const rule of compiled.rules) { + if (rule.keywords) { + lower ??= current.toLowerCase(); + if (!rule.keywords.some(k => lower!.includes(k))) continue; + } + const r = applyRule(current, rule, isAllowed, ctx?.onMatch); + if (r.count > 0) { + current = r.text; + lower = null; + counts.set(rule.id, (counts.get(rule.id) ?? 0) + r.count); + } + } + + if (compiled.entropy.enabled) { + const r = applyEntropy(current, compiled.entropy, isAllowed, ctx?.onMatch); + if (r.count > 0) { + current = r.text; + counts.set(ENTROPY_RULE_ID, r.count); + } + } + + return { text: current, findings: [...counts].map(([ruleId, count]) => ({ ruleId, count })) }; + }, + }; +} + +// --------------------------------------------------------------------------- +// Findings +// --------------------------------------------------------------------------- + +/** Aggregates findings by rule id. Never holds matched values. */ +export class FindingsTally { + private counts = new Map(); + + add(findings: RedactionFinding[]): void { + for (const f of findings) this.counts.set(f.ruleId, (this.counts.get(f.ruleId) ?? 0) + f.count); + } + + get total(): number { + let n = 0; + for (const c of this.counts.values()) n += c; + return n; + } + + toArray(): RedactionFinding[] { + return [...this.counts].map(([ruleId, count]) => ({ ruleId, count })); + } +} + +/** "7 value(s) redacted (azure-storage-key: 2, jwt: 5)" */ +export function formatFindings(findings: RedactionFinding[]): string { + const total = findings.reduce((n, f) => n + f.count, 0); + if (total === 0) return 'no values redacted'; + const detail = [...findings] + .sort((a, b) => b.count - a.count || a.ruleId.localeCompare(b.ruleId)) + .map(f => `${f.ruleId}: ${f.count}`) + .join(', '); + return `${total} value(s) redacted (${detail})`; +} + +// --------------------------------------------------------------------------- +// JSON / JSONL / files +// --------------------------------------------------------------------------- + +/** + * Redact a parsed JSON tree in place: string values by the text rules, values + * by their field name (secretFields), and keys that are themselves secrets. + */ +function redactTree( + node: unknown, + redactor: Redactor, + ctx: RedactionContext | undefined, + tally: FindingsTally | undefined +): { value: unknown; changed: boolean } { + if (typeof node === 'string') { + const r = redactor.redact(node, ctx); + if (r.findings.length === 0) return { value: node, changed: false }; + tally?.add(r.findings); + return { value: r.text, changed: r.text !== node }; + } + if (node === null || typeof node !== 'object') return { value: node, changed: false }; + + let changed = false; + if (Array.isArray(node)) { + for (let i = 0; i < node.length; i++) { + const r = redactTree(node[i], redactor, ctx, tally); + if (r.changed) { node[i] = r.value; changed = true; } + } + return { value: node, changed }; + } + const obj = node as Record; + const keys = Object.keys(obj); + for (const key of keys) { + const r = redactTree(obj[key], redactor, ctx, tally); + if (r.changed) { + setOwn(obj, key, r.value); + changed = true; + } + } + + // Field-name context: the key (or a {name, value} pair's name, or a Key + // Vault secret id) says the value is a secret even when its shape matches + // no rule — e.g. a letters-only clientSecret field in a structured + // MCP result. Runs after the text rules; a value they redacted only in part + // is still replaced whole, since the field name says all of it is secret. + const redactWhole = (key: string) => { + const value = obj[key]; + if (!isRedactableWhole(value)) return; + setOwn(obj, key, tokenForMatch(FIELD_RULE_ID, value, ctx?.onMatch)); + tally?.add([{ ruleId: FIELD_RULE_ID, count: 1 }]); + changed = true; + }; + for (const key of keys) { + if (redactor.isSecretField(key)) redactWhole(key); + } + const lower = new Map(keys.map(k => [k.toLowerCase(), k])); + const valueKey = lower.get('value'); + if (valueKey !== undefined) { + const nameKey = lower.get('name') ?? lower.get('key'); + const nameValue = nameKey !== undefined ? obj[nameKey] : undefined; + const id = lower.has('id') ? obj[lower.get('id')!] : undefined; + if ( + (typeof nameValue === 'string' && redactor.isSecretField(nameValue)) || + (typeof id === 'string' && KEY_VAULT_SECRET_ID.test(id)) + ) { + redactWhole(valueKey); + } + } + + // Keys themselves: a secret used as a key (a token-keyed cache) is renamed + // to its token. Harness keys never match a rule, so structure is unchanged. + let renamed: Array<[string, unknown]> | null = null; + for (let i = 0; i < keys.length; i++) { + const r = redactor.redact(keys[i], ctx); + if (r.findings.length === 0) continue; + tally?.add(r.findings); + renamed ??= keys.map(k => [k, obj[k]]); + renamed[i][0] = r.text; + } + if (renamed) { + const rebuilt: Record = {}; + for (const [key, value] of renamed) { + let unique = key; + for (let n = 2; Object.prototype.hasOwnProperty.call(rebuilt, unique); n++) unique = `${key}#${n}`; + setOwn(rebuilt, unique, value); + } + return { value: rebuilt, changed: true }; + } + return { value: obj, changed }; +} + +function setOwn(obj: Record, key: string, value: unknown): void { + // defineProperty so a "__proto__" key from JSON.parse stays an own property. + Object.defineProperty(obj, key, { value, writable: true, enumerable: true, configurable: true }); +} + +// A value that is only tokens is already redacted; anything else in a secret +// field, even alongside a token, is replaced whole. +const ONLY_TOKENS = /^\s*(?:\[REDACTED:[a-z0-9][a-z0-9-]*\]\s*)+$/; + +function isRedactableWhole(value: unknown): value is string { + return typeof value === 'string' && + value.trim().length > 0 && + !ONLY_TOKENS.test(value) && + !PLACEHOLDER_VALUE.test(value.trim()); +} + +/** + * Redact one JSONL line. Valid JSON stays valid (string values are redacted + * and the line is re-serialized only if something changed); a line that isn't + * JSON (e.g. a torn last line mid-write) is redacted as raw text. A trailing + * `\r` is preserved. + */ +export function redactJsonlLine( + line: string, + redactor: Redactor, + ctx?: RedactionContext, + tally?: FindingsTally +): string { + const cr = line.endsWith('\r') ? '\r' : ''; + const body = cr ? line.slice(0, -1) : line; + if (body.trim().length === 0) return line; + + let parsed: unknown; + try { + parsed = JSON.parse(body); + } catch { + const r = redactor.redact(body, ctx); + if (r.findings.length === 0) return line; + tally?.add(r.findings); + return r.text + cr; + } + + const r = redactTree(parsed, redactor, ctx, tally); + return r.changed ? JSON.stringify(r.value) + cr : line; +} + +const COPY_CHUNK_BYTES = 1 << 20; // 1 MiB + +/** + * Stream `src` to `dest`, redacting line by line. Writes `dest` directly; + * callers that need atomicity write to a temp path and rename. Line count and + * trailing-newline shape match the source exactly, so archive line numbers + * (exchanges.line_start/line_end, MCP read ranges) stay valid. + */ +export function copyFileRedacted( + src: string, + dest: string, + redactor: Redactor, + ctx?: RedactionContext, + tally?: FindingsTally +): void { + const fdIn = fs.openSync(src, 'r'); + let fdOut: number | undefined; + try { + // Create with the source's mode, as copyFileSync would, so a 0600 + // transcript doesn't become world-readable in the archive. + fdOut = fs.openSync(dest, 'w', fs.fstatSync(fdIn).mode & 0o777); + const buf = Buffer.allocUnsafe(COPY_CHUNK_BYTES); + const decoder = new StringDecoder('utf8'); + let pending = ''; + let bytesRead: number; + while ((bytesRead = fs.readSync(fdIn, buf, 0, buf.length, null)) > 0) { + let scanFrom = pending.length; + pending += decoder.write(buf.subarray(0, bytesRead)); + const out: string[] = []; + let start = 0; + let nl: number; + while ((nl = pending.indexOf('\n', scanFrom)) !== -1) { + out.push(redactJsonlLine(pending.slice(start, nl), redactor, ctx, tally), '\n'); + start = nl + 1; + scanFrom = start; + } + if (out.length > 0) fs.writeSync(fdOut, out.join('')); + pending = pending.slice(start); + } + pending += decoder.end(); + if (pending.length > 0) fs.writeSync(fdOut, redactJsonlLine(pending, redactor, ctx, tally)); + } finally { + fs.closeSync(fdIn); + if (fdOut !== undefined) fs.closeSync(fdOut); + } +} diff --git a/src/summarizer.ts b/src/summarizer.ts index d7e3cc04..801f3743 100644 --- a/src/summarizer.ts +++ b/src/summarizer.ts @@ -702,7 +702,23 @@ export function getCodexModel(_exchanges: ConversationExchange[]): string | unde return process.env.EPISODIC_MEMORY_CODEX_MODEL || undefined; } -export async function summarizeConversation(exchanges: ConversationExchange[], sessionId?: string): Promise { +export interface SummarizeOptions { + /** + * Allow Claude session resume and Codex thread/fork (default true). Both + * paths make the model read the *source* transcript rather than `exchanges`, + * so callers pass false when the exchanges were redacted (see + * docs/REDACTION.md) to force the transcript-text path. + */ + allowResume?: boolean; +} + +export async function summarizeConversation( + exchanges: ConversationExchange[], + sessionId?: string, + options: SummarizeOptions = {} +): Promise { + const allowResume = options.allowResume !== false; + // Handle trivial conversations if (exchanges.length === 0) { return 'Trivial conversation with no substantive content.'; @@ -715,7 +731,7 @@ export async function summarizeConversation(exchanges: ConversationExchange[], s } } - const codexSessionId = getCodexSessionId(exchanges, sessionId); + const codexSessionId = allowResume ? getCodexSessionId(exchanges, sessionId) : undefined; if (codexSessionId) { try { const result = await callCodex(buildCodexSummaryPrompt(), codexSessionId, getCodexModel(exchanges)); @@ -742,7 +758,7 @@ export async function summarizeConversation(exchanges: ConversationExchange[], s const isClaudeSession = exchanges.some( e => e.harness === 'claude' || e.harness === undefined ); - const claudeSessionId = !codexSessionId && isClaudeSession ? sessionId : undefined; + const claudeSessionId = allowResume && !codexSessionId && isClaudeSession ? sessionId : undefined; const cwd = claudeSessionId ? exchanges.find(e => e.cwd)?.cwd : undefined; const conversationText = claudeSessionId ? '' // When resuming, no need to include conversation text - it's already in context diff --git a/src/sync-cli.ts b/src/sync-cli.ts index 297b20f8..b12ca46b 100644 --- a/src/sync-cli.ts +++ b/src/sync-cli.ts @@ -14,6 +14,7 @@ import { spawn } from 'child_process'; import fs from 'fs'; import { formatLogLine, getSyncLogPath, getSyncLockPath } from './logging.js'; import { acquireFileLock, readLockHolder, releaseFileLock } from './file-lock.js'; +import { FindingsTally, formatFindings, getRedactionSettings, loadRedactor, type Redactor } from './redaction.js'; const args = process.argv.slice(2); @@ -48,10 +49,13 @@ Sync conversations from Claude Code, Codex, and opencode transcript sources to a This command: 1. Exports opencode sessions from its SQLite database when available -2. Copies new or updated .jsonl files to conversation archive +2. Copies new or updated .jsonl files to conversation archive, redacting secrets 3. Generates embeddings for semantic search 4. Updates the search index +Secrets are replaced with [REDACTED:] tokens before anything is archived, +indexed, embedded, or summarized. See EPISODIC_MEMORY_REDACTION* in the README. + Only processes files that are new or have been modified since last sync. Safe to run multiple times - subsequent runs are fast no-ops. @@ -157,8 +161,23 @@ if (isBackground) { process.exit(0); } +// Load redaction rules before anything is exported, copied, or indexed. In +// strict mode (the default) a bad rules file stops the sync here: fail closed +// rather than archive unredacted text. +let redactor: Redactor | null; +try { + redactor = loadRedactor(); +} catch (error) { + console.error(`episodic-memory: ${error instanceof Error ? error.message : String(error)}`); + console.error( + 'episodic-memory: refusing to sync without redaction (strict mode). Fix the rules file, ' + + 'or set EPISODIC_MEMORY_REDACTION_STRICT=0 to sync unredacted, or EPISODIC_MEMORY_REDACTION=off.' + ); + process.exit(1); +} + if (!onlyHarnesses || onlyHarnesses.includes('opencode')) { - const opencodeExport = exportOpencodeSessions(); + const opencodeExport = exportOpencodeSessions({ redactor }); if (opencodeExport.exported > 0 || opencodeExport.skipped > 0) { console.log(`opencode export: ${opencodeExport.exported} exported, ${opencodeExport.skipped} skipped`); } @@ -215,8 +234,11 @@ console.log(`Destination: ${destDir}\n`); async function syncAll() { const totals = { copied: 0, skipped: 0, indexed: 0, summarized: 0, errors: [] as Array<{file: string; error: string}>, sourcesWithSummaryWork: 0, totalNeedingSummaries: 0 }; + const redactions = new FindingsTally(); + for (const sourceDir of sourceDirs) { - const result = await syncConversations(sourceDir, destDir, { ...syncOptions, summaryLimit }); + const result = await syncConversations(sourceDir, destDir, { ...syncOptions, summaryLimit, redactor }); + redactions.add(result.redactions); totals.copied += result.copied; totals.skipped += result.skipped; totals.indexed += result.indexed; @@ -233,6 +255,14 @@ async function syncAll() { } else { console.log(` Summarized: ${totals.summarized}`); } + // Rule IDs and counts only — never matched values. + if (redactor) { + console.log(` Redaction: ${formatFindings(redactions.toArray())}`); + } else if (getRedactionSettings().enabled) { + console.log(' Redaction: NOT APPLIED (rules failed to load; EPISODIC_MEMORY_REDACTION_STRICT=0)'); + } else { + console.log(' Redaction: off (EPISODIC_MEMORY_REDACTION=off)'); + } if (totals.errors.length > 0) { console.log(`\n⚠️ Errors: ${totals.errors.length}`); diff --git a/src/sync.ts b/src/sync.ts index d9ee6810..62fa00ec 100644 --- a/src/sync.ts +++ b/src/sync.ts @@ -5,6 +5,7 @@ import { SUMMARIZER_CONTEXT_MARKER } from './constants.js'; import { getExcludedProjects, findJsonlFiles, statIfExists } from './paths.js'; import { formatErrorSentinel, shouldQueueForSummary } from './summary-sentinel.js'; import { getMaxMessageBytes, isOversizeExchange } from './message-size.js'; +import { copyFileRedacted, FindingsTally, loadRedactor, type RedactionFinding, type Redactor } from './redaction.js'; const EXCLUSION_MARKERS = [ 'DO NOT INDEX THIS CHAT', @@ -105,12 +106,19 @@ export interface SyncResult { indexed: number; summarized: number; errors: Array<{ file: string; error: string }>; + /** Rule IDs and counts only — never matched values. */ + redactions: RedactionFinding[]; } export interface SyncOptions { skipIndex?: boolean; skipSummaries?: boolean; summaryLimit?: number; // Max summaries to generate per run (default: 10) + /** + * Redactor applied at the archive write. `undefined` loads it from the + * environment (loadRedactor); `null` means redaction is off. + */ + redactor?: Redactor | null; } /** @@ -132,7 +140,23 @@ export function buildSyncOptionsFromEnv(env: NodeJS.ProcessEnv): SyncOptions { return { skipSummaries: env.EPISODIC_MEMORY_SKIP_SUMMARIES === '1' }; } -function copyIfNewer(src: string, dest: string): boolean { +/** + * Copy `src` into the archive at `dest` when the archive copy is missing or + * older. This is the redaction choke point: with a redactor, the copy is + * redacted line by line (see redaction.ts), and every downstream stage + * (index, embeddings, summaries, show/read) reads the archive. + * + * `force` copies even when the mtimes say the archive is current: an append + * within the mtime granularity (or the millisecond the rounding below adds) + * leaves the source no newer than the archive. + */ +export function copyIfNewer( + src: string, + dest: string, + redactor: Redactor | null = null, + tally?: FindingsTally, + force = false +): boolean { // Ensure destination directory exists const destDir = path.dirname(dest); if (!fs.existsSync(destDir)) { @@ -140,7 +164,7 @@ function copyIfNewer(src: string, dest: string): boolean { } // Check if destination exists and is up-to-date - if (fs.existsSync(dest)) { + if (!force && fs.existsSync(dest)) { const srcStat = fs.statSync(src); const destStat = fs.statSync(dest); if (destStat.mtimeMs >= srcStat.mtimeMs) { @@ -150,8 +174,17 @@ function copyIfNewer(src: string, dest: string): boolean { // Atomic copy: temp file + rename const tempDest = dest + '.tmp.' + process.pid; - fs.copyFileSync(src, tempDest); - fs.renameSync(tempDest, dest); // Atomic on same filesystem + try { + if (redactor) { + copyFileRedacted(src, tempDest, redactor, { source: 'archive', path: src }, tally); + } else { + fs.copyFileSync(src, tempDest); + } + fs.renameSync(tempDest, dest); // Atomic on same filesystem + } catch (error) { + try { fs.unlinkSync(tempDest); } catch {} + throw error; + } // Preserve source mtime: harnesses without per-message timestamps (Cursor // agent transcripts) fall back to file mtime. Round up to the next whole @@ -184,9 +217,16 @@ export async function syncConversations( skipped: 0, indexed: 0, summarized: 0, - errors: [] + errors: [], + redactions: [] }; + // Load the redactor before touching the archive. In strict mode (the + // default) a rules-load failure throws here, so nothing unredacted is + // archived, indexed, or summarized (fail closed). + const redactor = options.redactor === undefined ? loadRedactor() : options.redactor; + const tally = new FindingsTally(); + // Ensure source directory exists if (!fs.existsSync(sourceDir)) { return result; @@ -226,7 +266,7 @@ export async function syncConversations( continue; } - const wasCopied = copyIfNewer(srcFile, destFile); + const wasCopied = copyIfNewer(srcFile, destFile, redactor, tally); if (wasCopied) { result.copied++; filesToIndex.push(destFile); @@ -257,6 +297,8 @@ export async function syncConversations( } } + result.redactions = tally.toArray(); + // Index copied files (unless skipIndex is set) if (!options.skipIndex && filesToIndex.length > 0) { const { parseConversation } = await import('./parser.js'); @@ -395,7 +437,10 @@ export async function syncConversations( } console.log(` Summarizing ${path.basename(filePath)} (${exchanges.length} exchanges)...`); - const summary = await summarizeConversation(exchanges, sessionId); + // With redaction on, never resume/fork the session: those paths hand + // the model the unredacted source transcript instead of these + // (redacted) exchanges. + const summary = await summarizeConversation(exchanges, sessionId, { allowResume: redactor === null }); const summaryPath = filePath.replace('.jsonl', '-summary.txt'); fs.writeFileSync(summaryPath, summary, 'utf-8'); diff --git a/src/verify.ts b/src/verify.ts index 0dd2c5db..bc9b6dd1 100644 --- a/src/verify.ts +++ b/src/verify.ts @@ -4,6 +4,7 @@ import { parseConversation } from './parser.js'; import { initDatabase, getAllExchanges, getFileLastIndexed } from './db.js'; import { getArchiveDir, getExcludedProjects, findJsonlFiles, statIfExists } from './paths.js'; import { isErroredSentinel } from './summary-sentinel.js'; +import { loadRedactor } from './redaction.js'; export interface VerificationResult { missing: Array<{ path: string; reason: string }>; @@ -121,6 +122,10 @@ export async function verifyIndex(): Promise { export async function repairIndex(issues: VerificationResult): Promise { console.log('Repairing index...'); + // Load before touching the index: strict mode fails closed here, as in sync + // and index. + const redactor = loadRedactor(); + // To avoid circular dependencies, we import the indexer functions dynamically const { initDatabase, insertExchange, deleteExchange } = await import('./db.js'); const { parseConversation } = await import('./parser.js'); @@ -160,7 +165,8 @@ export async function repairIndex(issues: VerificationResult): Promise { // Generate/update summary const summaryPath = conversationPath.replace('.jsonl', '-summary.txt'); - const summary = await summarizeConversation(exchanges); + // Under redaction, don't let the Codex fork fallback read the source rollout. + const summary = await summarizeConversation(exchanges, undefined, { allowResume: redactor === null }); fs.writeFileSync(summaryPath, summary, 'utf-8'); console.log(` Created summary: ${summary.split(/\s+/).length} words`); diff --git a/test/fake-secrets.ts b/test/fake-secrets.ts new file mode 100644 index 00000000..0eebca9e --- /dev/null +++ b/test/fake-secrets.ts @@ -0,0 +1,151 @@ +/** + * Fake-secret builders for redaction tests. + * + * Nothing in this file is a real credential, and no string literal here matches + * a redaction rule on its own: every secret is assembled at runtime from a + * seeded PRNG plus split prefixes. That keeps GitHub push protection and other + * scanners quiet while still producing values in the exact real-world shape. + */ + +const UPPER = 'ABCDEFGHIJKLMNOPQRSTUVWXYZ'; +const LOWER = 'abcdefghijklmnopqrstuvwxyz'; +const DIGITS = '0123456789'; +const ALNUM = UPPER + LOWER + DIGITS; +const B64 = ALNUM + '+/'; +const B64URL = ALNUM + '-_'; +const BASE32 = UPPER + '234567'; + +/** Deterministic PRNG (mulberry32) so failures reproduce. */ +export function makeRng(seed: number): () => number { + let a = seed >>> 0; + return () => { + a = (a + 0x6d2b79f5) >>> 0; + let t = a; + t = Math.imul(t ^ (t >>> 15), t | 1); + t ^= t + Math.imul(t ^ (t >>> 7), t | 61); + return ((t ^ (t >>> 14)) >>> 0) / 4294967296; + }; +} + +export class FakeSecrets { + private rng: () => number; + + constructor(seed = 1337) { + this.rng = makeRng(seed); + } + + chars(charset: string, n: number): string { + let out = ''; + for (let i = 0; i < n; i++) out += charset[Math.floor(this.rng() * charset.length)]; + return out; + } + + /** Always contains at least one digit and one letter. */ + private mixed(charset: string, n: number): string { + return this.chars(charset, n - 2) + this.chars(DIGITS, 1) + this.chars(LOWER, 1); + } + + azureClientSecret(): string { + // <3 chars>Q~<34 chars> — Entra ID client secret shape. + return this.chars(ALNUM, 3) + this.chars(DIGITS, 1) + 'Q' + '~' + this.chars(ALNUM + '_~.-', 33) + 'x'; + } + + azureStorageKey(): string { + // 64 random bytes, base64: 86 chars + '=='. + return this.chars(ALNUM, 1) + this.chars(B64, 84) + this.chars(ALNUM, 1) + '=' + '='; + } + + sasSignature(): string { + return this.chars(ALNUM, 40) + '%2B' + this.chars(ALNUM, 3) + '%3D'; + } + + privateKeyBlock(kind = 'RSA'): string { + const body: string[] = []; + for (let i = 0; i < 6; i++) body.push(this.chars(B64, 64)); + const dashes = '-'.repeat(5); + return `${dashes}BEGIN ${kind} PRIVATE` + ` KEY${dashes}\n${body.join('\n')}\n${dashes}END ${kind} PRIVATE` + ` KEY${dashes}`; + } + + jwt(): string { + const head = 'ey' + 'J' + this.chars(B64URL, 30); + const payload = 'ey' + 'J' + this.chars(B64URL, 60); + return `${head}.${payload}.${this.chars(B64URL, 43)}`; + } + + anthropicKey(): string { + return 'sk-' + 'ant-' + 'api03-' + this.chars(B64URL, 93) + 'AA'; + } + + openAiKey(): string { + return 'sk-' + 'proj-' + this.chars(ALNUM, 20) + 'T3Blbk' + 'FJ' + this.chars(ALNUM, 20); + } + + githubToken(): string { + return 'gh' + 'p_' + this.chars(ALNUM, 36); + } + + githubFineGrainedPat(): string { + return 'github' + '_pat_' + this.chars(ALNUM, 22) + '_' + this.chars(ALNUM, 59); + } + + awsAccessKeyId(): string { + return 'AK' + 'IA' + this.chars(BASE32, 16); + } + + awsSecretAccessKey(): string { + return this.mixed(ALNUM + '/+', 40); + } + + slackToken(): string { + return 'xo' + 'xb-' + this.chars(DIGITS, 12) + '-' + this.chars(DIGITS, 12) + '-' + this.chars(ALNUM, 24); + } + + googleApiKey(): string { + return 'AI' + 'za' + this.chars(ALNUM + '_-', 35); + } + + npmToken(): string { + return 'np' + 'm_' + this.chars(ALNUM, 36); + } + + /** Starts with a letter so it never looks like a `$VAR`/`%VAR%` placeholder. */ + password(): string { + return this.chars(LOWER, 1) + this.mixed(ALNUM + '!#%^*', 15); + } + + /** Letters only: a real password with no digit, which digit-gated rules miss. */ + passphrase(): string { + return this.chars(UPPER, 1) + this.chars(LOWER, 7) + this.chars(UPPER, 1) + this.chars(LOWER, 9); + } + + bearerOpaque(): string { + return this.mixed(ALNUM, 40); + } + + basicAuth(): string { + return this.chars(ALNUM, 30) + '=' + '='; + } + + /** High-entropy blob with no recognizable prefix. */ + opaqueHighEntropy(n = 48): string { + return this.mixed(ALNUM + '+/', n); + } + + gitSha(): string { + return this.chars('0123456789abcdef', 40); + } + + guid(): string { + const h = (n: number) => this.chars('0123456789abcdef', n); + return `${h(8)}-${h(4)}-${h(4)}-${h(4)}-${h(12)}`; + } + + base64Blob(n: number): string { + return this.chars(B64, n); + } +} + +/** Every `[REDACTED:...]` token in a string. */ +export function redactionTokens(text: string): string[] { + return text.match(/\[REDACTED:[a-z0-9-]+\]/g) ?? []; +} diff --git a/test/redact-rewrite.test.ts b/test/redact-rewrite.test.ts new file mode 100644 index 00000000..b794a2fb --- /dev/null +++ b/test/redact-rewrite.test.ts @@ -0,0 +1,260 @@ +import { describe, it, expect, beforeEach, afterEach, vi } from 'vitest'; +import { mkdtempSync, mkdirSync, writeFileSync, readFileSync, rmSync, existsSync, statSync, utimesSync } from 'fs'; +import { join, sep } from 'path'; +import { tmpdir } from 'os'; +import Database from 'better-sqlite3'; +import * as sqliteVec from 'sqlite-vec'; + +import { initDatabase, insertExchange } from '../src/db.js'; +import { parseConversation } from '../src/parser.js'; +import { rewriteArchive } from '../src/redact-rewrite.js'; +import { createRedactor, DEFAULT_REDACTION_CONFIG } from '../src/redaction.js'; +import { EMBEDDING_VERSION } from '../src/embedding-migration.js'; +import { FakeSecrets } from './fake-secrets.js'; + +const SESSION = '9a8b7c6d-1111-4222-8333-944455556666'; + +function transcript(secret: string, password: string, clean: string): string { + const base = { sessionId: SESSION, isSidechain: false, cwd: '/work/legacy' }; + return [ + { ...base, type: 'user', uuid: 'u1', parentUuid: null, timestamp: '2025-06-01T00:00:00.000Z', message: { role: 'user', content: `use AccountKey=${secret} please` } }, + { + ...base, type: 'assistant', uuid: 'a1', parentUuid: 'u1', timestamp: '2025-06-01T00:00:01.000Z', + message: { role: 'assistant', content: [ + { type: 'text', text: 'Running it.' }, + { type: 'tool_use', id: 'toolu_1', name: 'Bash', input: { command: `mysql --password=${password}` } }, + ] }, + }, + { ...base, type: 'user', uuid: 'u2', parentUuid: 'a1', timestamp: '2025-06-01T00:00:02.000Z', message: { role: 'user', content: [{ type: 'tool_result', tool_use_id: 'toolu_1', content: `ok password: ${password}` }] } }, + { ...base, type: 'assistant', uuid: 'a2', parentUuid: 'u2', timestamp: '2025-06-01T00:00:03.000Z', message: { role: 'assistant', content: [{ type: 'text', text: clean }] } }, + ].map(l => JSON.stringify(l)).join('\n') + '\n'; +} + +describe('redact --rewrite backfill', () => { + let root: string; + let archiveDir: string; + let dbPath: string; + const savedEnv = { ...process.env }; + const embed = vi.fn(async (..._args: unknown[]) => new Array(384).fill(0.5)); + + beforeEach(() => { + root = mkdtempSync(join(tmpdir(), 'episodic-memory-redact-rewrite-')); + archiveDir = join(root, 'archive'); + dbPath = join(root, 'db.sqlite'); + process.env.TEST_DB_PATH = dbPath; + embed.mockClear(); + }); + + afterEach(() => { + process.env = { ...savedEnv }; + rmSync(root, { recursive: true, force: true }); + }); + + /** Simulate a pre-redaction install: unredacted archive, rows, and summary. */ + async function seedLegacy(fake: FakeSecrets) { + const secret = fake.azureStorageKey(); + const password = fake.password(); + const projectDir = join(archiveDir, '-work-legacy'); + mkdirSync(join(projectDir, SESSION, 'subagents'), { recursive: true }); + const dirty = join(projectDir, `${SESSION}.jsonl`); + const clean = join(projectDir, 'clean-session.jsonl'); + const nested = join(projectDir, SESSION, 'subagents', 'agent-abc.jsonl'); + writeFileSync(dirty, transcript(secret, password, 'All done.')); + writeFileSync(nested, transcript(secret, password, 'Sub-agent done.')); + writeFileSync(clean, transcript('[REDACTED:connection-string-secret]', '[REDACTED:secret-assignment]', 'Nothing here.')); + writeFileSync(dirty.replace('.jsonl', '-summary.txt'), 'Configured storage access.'); + writeFileSync(clean.replace('.jsonl', '-summary.txt'), 'A clean summary.'); + const old = new Date('2025-06-02T00:00:00Z'); + for (const f of [dirty, clean, nested]) utimesSync(f, old, old); + + const db = initDatabase(); + for (const file of [dirty, clean, nested]) { + for (const ex of await parseConversation(file, '-work-legacy', file)) { + insertExchange(db, ex, new Array(384).fill(0), ex.toolCalls?.map(t => t.toolName)); + } + } + // Pretend these rows predate the current encoder bookkeeping. + db.prepare('UPDATE exchanges SET embedding_version = 0').run(); + db.close(); + return { secret, password, dirty, clean, nested }; + } + + function dbDump(): string { + const db = new Database(dbPath, { readonly: true }); + sqliteVec.load(db); + try { + return JSON.stringify([ + db.prepare('SELECT user_message, assistant_message FROM exchanges').all(), + db.prepare('SELECT tool_input, tool_result FROM tool_calls').all(), + ]); + } finally { + db.close(); + } + } + + it('rewrites the archive and index in place, then is a no-op on a second run', async () => { + const fake = new FakeSecrets(808); + const { secret, password, dirty, clean, nested } = await seedLegacy(fake); + const lineCount = readFileSync(dirty, 'utf-8').split('\n').length; + const mtimeBefore = statSync(dirty).mtimeMs; + const cleanBytes = readFileSync(clean, 'utf-8'); + expect(dbDump()).toContain(secret); + + const redactor = createRedactor(DEFAULT_REDACTION_CONFIG); + const result = await rewriteArchive({ archiveDir, redactor, embed }); + + // Archive: secrets gone, structure and mtime preserved, nested files covered. + for (const f of [dirty, nested]) { + const text = readFileSync(f, 'utf-8'); + expect(text).not.toContain(secret); + expect(text).not.toContain(password); + text.split('\n').filter(Boolean).forEach(l => expect(() => JSON.parse(l)).not.toThrow()); + } + expect(readFileSync(dirty, 'utf-8').split('\n').length).toBe(lineCount); + expect(Math.abs(statSync(dirty).mtimeMs - mtimeBefore)).toBeLessThan(2); + expect(readFileSync(clean, 'utf-8')).toBe(cleanBytes); + + // Summaries generated from unredacted text are removed; clean ones stay. + expect(existsSync(dirty.replace('.jsonl', '-summary.txt'))).toBe(false); + expect(existsSync(clean.replace('.jsonl', '-summary.txt'))).toBe(true); + + // Index: text columns clean, affected rows re-embedded with the current version. + const dump = dbDump(); + expect(dump).not.toContain(secret); + expect(dump).not.toContain(password); + expect(dump).toContain('[REDACTED:'); + expect(embed).toHaveBeenCalled(); + expectNoSecretInCalls(embed.mock.calls, [secret, password]); + + const db = new Database(dbPath, { readonly: true }); + // Only rows that had secrets are touched; the clean file's rows keep their old version. + const versions = db.prepare( + `SELECT embedding_version v, COUNT(*) n FROM exchanges WHERE archive_path IN (?, ?) AND (user_message LIKE '%REDACTED%' OR assistant_message LIKE '%REDACTED%') GROUP BY v` + ).all(dirty, nested) as Array<{ v: number }>; + const untouched = db.prepare('SELECT DISTINCT embedding_version v FROM exchanges WHERE archive_path = ?').all(clean) as Array<{ v: number }>; + db.close(); + expect(versions.map(r => r.v)).toEqual([EMBEDDING_VERSION]); + expect(untouched.map(r => r.v)).toEqual([0]); + + expect(result.filesRewritten).toBe(2); + expect(result.summariesRemoved).toBe(1); + expect(result.rowsUpdated).toBeGreaterThan(0); + expect(JSON.stringify(result)).not.toContain(secret); + + // Second run: nothing left to do. + embed.mockClear(); + const second = await rewriteArchive({ archiveDir, redactor, embed }); + expect(second.filesRewritten).toBe(0); + expect(second.rowsUpdated).toBe(0); + expect(second.summariesRemoved).toBe(0); + expect(embed).not.toHaveBeenCalled(); + }); + + it('--dry-run reports counts without touching anything', async () => { + const fake = new FakeSecrets(909); + const { secret, dirty } = await seedLegacy(fake); + const before = readFileSync(dirty, 'utf-8'); + + const result = await rewriteArchive({ archiveDir, redactor: createRedactor(DEFAULT_REDACTION_CONFIG), embed, dryRun: true }); + + expect(result.filesRewritten).toBe(2); + expect(result.rowsUpdated).toBeGreaterThan(0); + expect(readFileSync(dirty, 'utf-8')).toBe(before); + expect(existsSync(dirty.replace('.jsonl', '-summary.txt'))).toBe(true); + expect(dbDump()).toContain(secret); + expect(embed).not.toHaveBeenCalled(); + }); + + it('--dry-run leaves an older-schema index byte-identical instead of migrating it', async () => { + await seedLegacy(new FakeSecrets(910)); + const legacy = new Database(dbPath); + legacy.exec('ALTER TABLE exchanges DROP COLUMN embedding_version'); + legacy.close(); + const before = readFileSync(dbPath); + + const result = await rewriteArchive({ archiveDir, redactor: createRedactor(DEFAULT_REDACTION_CONFIG), embed, dryRun: true }); + + expect(result.rowsUpdated).toBeGreaterThan(0); + expect(readFileSync(dbPath).equals(before)).toBe(true); + const db = new Database(dbPath, { readonly: true }); + const columns = (db.prepare('PRAGMA table_info(exchanges)').all() as Array<{ name: string }>).map(c => c.name); + db.close(); + expect(columns).not.toContain('embedding_version'); + }); + + it('--dry-run does not create an index that does not exist', async () => { + mkdirSync(join(archiveDir, '-work-legacy'), { recursive: true }); + const result = await rewriteArchive({ archiveDir, redactor: createRedactor(DEFAULT_REDACTION_CONFIG), embed, dryRun: true }); + expect(result.rowsUpdated).toBe(0); + expect(existsSync(dbPath)).toBe(false); + }); + + it('report lists one hit per redacted value, with location, rule and shape but never the value', async () => { + const { secret, password } = await seedLegacy(new FakeSecrets(911)); + const hits: Array<{ location: string; ruleId: string; shape: string; context: string }> = []; + + const result = await rewriteArchive({ + archiveDir, redactor: createRedactor(DEFAULT_REDACTION_CONFIG), embed, dryRun: true, report: hit => hits.push(hit), + }); + + const total = result.findings.reduce((n, f) => n + f.count, 0); + expect(hits.length).toBe(total); + const text = JSON.stringify(hits); + expect(text).not.toContain(secret); + expect(text).not.toContain(password); + + const archiveHit = hits.find(h => h.location === `-work-legacy${sep}${SESSION}.jsonl:1`); + expect(archiveHit?.ruleId).toBe('connection-string-secret'); + expect(archiveHit!.shape).toMatch(new RegExp(`^len=${secret.length} \\S+ H=\\d+\\.\\d$`)); + expect(archiveHit!.context).toContain('AccountKey=[REDACTED:connection-string-secret]'); + expect(hits.some(h => h.location.startsWith('index:') && h.location.includes(' user'))).toBe(true); + }); + + it('report refuses to run without dryRun, since files would already be rewritten', async () => { + mkdirSync(archiveDir, { recursive: true }); + await expect(rewriteArchive({ + archiveDir, redactor: createRedactor(DEFAULT_REDACTION_CONFIG), embed, report: () => {}, + })).rejects.toThrow(/dryRun/); + }); + + /** Dry-run report over one archive file holding `line`. */ + async function reportLine(line: string) { + mkdirSync(join(archiveDir, '-work-x'), { recursive: true }); + writeFileSync(join(archiveDir, '-work-x', 's.jsonl'), line + '\n'); + const hits: Array<{ location: string; ruleId: string; shape: string; context: string }> = []; + await rewriteArchive({ archiveDir, redactor: createRedactor(DEFAULT_REDACTION_CONFIG), embed, dryRun: true, report: h => hits.push(h) }); + return hits; + } + + it('report pairs a new hit with its own token, not an earlier token of the same rule', async () => { + const pw = new FakeSecrets(912).chars('abcdefghijklmnopqrstuvwxyz', 14); + const hits = await reportLine( + JSON.stringify({ content: `first: Server=a;password=${pw}; later: Server=b;Password=[REDACTED:connection-string-secret];` }) + ); + expect(hits).toHaveLength(1); + expect(hits[0].shape).toMatch(/^len=14 a /); + expect(hits[0].context).toContain('first: Server=a;password=[REDACTED:connection-string-secret]'); + expect(hits[0].context).not.toContain('--hit'); + }); + + it('report keeps a hit swallowed by a whole secret field, and shapes the field from its original value', async () => { + const fake = new FakeSecrets(913); + const word = fake.chars('abcdefghijklmnopqrstuvwxyz', 9); + const jwt = fake.jwt(); + const hits = await reportLine(JSON.stringify({ password: `${word} ${jwt}` })); + + expect(hits.map(h => h.ruleId).sort()).toEqual(['jwt', 'secret-field']); + const field = hits.find(h => h.ruleId === 'secret-field')!; + expect(field.shape).toMatch(new RegExp(`^len=${word.length + 1 + jwt.length} `)); + const swallowed = hits.find(h => h.ruleId === 'jwt')!; + expect(swallowed.shape).toMatch(new RegExp(`^len=${jwt.length} `)); + expect(swallowed.context).toBe(field.context); + expect(field.context).toContain('"password":"[REDACTED:secret-field]"'); + expect(JSON.stringify(hits)).not.toContain(word); + }); +}); + +function expectNoSecretInCalls(calls: unknown[][], secrets: string[]) { + const text = JSON.stringify(calls); + for (const s of secrets) expect(text.includes(s)).toBe(false); +} diff --git a/test/redaction-pipeline.test.ts b/test/redaction-pipeline.test.ts new file mode 100644 index 00000000..770a9cd0 --- /dev/null +++ b/test/redaction-pipeline.test.ts @@ -0,0 +1,421 @@ +import { describe, it, expect, beforeEach, afterEach, vi } from 'vitest'; +import { + mkdtempSync, mkdirSync, writeFileSync, readFileSync, rmSync, existsSync, readdirSync, statSync, + appendFileSync, utimesSync, +} from 'fs'; +import { join } from 'path'; +import { tmpdir } from 'os'; +import Database from 'better-sqlite3'; +import * as sqliteVec from 'sqlite-vec'; + +// Capture every embedding input and every summarizer call. sync.ts loads both +// via dynamic import and indexer.ts statically; vi.mock intercepts both. +const { embedSpy, summarizeSpy } = vi.hoisted(() => ({ + embedSpy: vi.fn(async (..._args: unknown[]) => new Array(384).fill(0)), + summarizeSpy: vi.fn(async (..._args: unknown[]) => 'A summary.'), +})); +vi.mock('../src/embeddings.js', () => ({ + initEmbeddings: vi.fn(async () => {}), + generateExchangeEmbedding: embedSpy, + generateQueryEmbedding: vi.fn(async () => new Array(384).fill(0)), + generateEmbedding: vi.fn(), + initEmbeddingsFailed: false, +})); +vi.mock('../src/summarizer.js', async () => { + const actual = await vi.importActual('../src/summarizer.js'); + return { ...actual, summarizeConversation: summarizeSpy }; +}); + +import { syncConversations } from '../src/sync.js'; +import { indexUnprocessed } from '../src/indexer.js'; +import { searchConversations } from '../src/search.js'; +import { formatConversationAsMarkdown } from '../src/show.js'; +import { exportOpencodeSessions } from '../src/opencode-sync.js'; +import { createRedactor, DEFAULT_REDACTION_CONFIG } from '../src/redaction.js'; +import { FakeSecrets } from './fake-secrets.js'; + +const SESSION = '4f1c2b3a-1111-4222-8333-944455556666'; + +interface Seed { + secrets: string[]; + sha: string; + guid: string; +} + +function seedSecrets(seed: number): Seed { + const fake = new FakeSecrets(seed); + return { + secrets: [ + fake.azureClientSecret(), + fake.azureStorageKey(), + fake.sasSignature(), + fake.githubToken(), + fake.anthropicKey(), + fake.password(), + fake.jwt(), + fake.privateKeyBlock().split('\n')[2], + ], + sha: fake.gitSha(), + guid: fake.guid(), + }; +} + +function claudeTranscript(s: Seed): string { + const [clientSecret, storageKey, sasSig, ghToken, antKey, password, jwt] = s.secrets; + const keyBlock = (() => { + // Rebuild the block whose body line is s.secrets[7]. + const dashes = '-'.repeat(5); + return `${dashes}BEGIN RSA PRIVATE` + ` KEY${dashes}\n${s.secrets[7]}\n${dashes}END RSA PRIVATE` + ` KEY${dashes}`; + })(); + const ts = (n: number) => new Date(Date.UTC(2026, 0, 1, 0, n)).toISOString(); + const base = { sessionId: SESSION, isSidechain: false, cwd: '/work/contoso', gitBranch: 'main', version: '2.0.0' }; + const lines = [ + { + ...base, type: 'user', uuid: 'u1', parentUuid: null, timestamp: ts(1), + message: { + role: 'user', + content: `Deploy with client secret ${clientSecret} for tenant ${s.guid}; last good commit ${s.sha}`, + }, + }, + { + ...base, type: 'assistant', uuid: 'a1', parentUuid: 'u1', timestamp: ts(2), + message: { + role: 'assistant', + content: [ + { type: 'text', text: 'Checking the storage account keys.' }, + { + type: 'tool_use', id: 'toolu_1', name: 'Bash', + input: { command: `az storage blob upload --connection-string "AccountName=contosodata;AccountKey=${storageKey}" --sas-token "sv=2022-11-02&sig=${sasSig}"` }, + }, + ], + }, + }, + { + ...base, type: 'user', uuid: 'u2', parentUuid: 'a1', timestamp: ts(3), + message: { + role: 'user', + content: [{ + type: 'tool_result', tool_use_id: 'toolu_1', + content: `GITHUB_TOKEN=${ghToken}\nANTHROPIC_API_KEY=${antKey}\npassword: ${password}\n${keyBlock}`, + }], + }, + }, + { + ...base, type: 'assistant', uuid: 'a2', parentUuid: 'u2', timestamp: ts(4), + message: { role: 'assistant', content: [{ type: 'text', text: `Uploaded. The bearer was Bearer ${jwt}.` }] }, + }, + ]; + return lines.map(l => JSON.stringify(l)).join('\n') + '\n'; +} + +function codexRollout(s: Seed): string { + const [clientSecret, storageKey] = s.secrets; + const lines = [ + { timestamp: '2026-05-12T18:00:00.000Z', type: 'session_meta', payload: { id: '019e4c75-d5bf-7c71-9df7-77f5fb86b711', cwd: '/work/contoso', cli_version: '0.130.0', model_provider: 'openai' } }, + { timestamp: '2026-05-12T18:00:02.000Z', type: 'response_item', payload: { type: 'message', role: 'user', content: [{ type: 'input_text', text: `Rotate ${clientSecret} please` }] } }, + { timestamp: '2026-05-12T18:00:04.000Z', type: 'response_item', payload: { type: 'function_call', name: 'exec_command', arguments: JSON.stringify({ cmd: `echo AccountKey=${storageKey}` }), call_id: 'c1' } }, + { timestamp: '2026-05-12T18:00:05.000Z', type: 'response_item', payload: { type: 'function_call_output', call_id: 'c1', output: `AccountKey=${storageKey}` } }, + { timestamp: '2026-05-12T18:00:06.000Z', type: 'response_item', payload: { type: 'message', role: 'assistant', content: [{ type: 'output_text', text: 'Rotated.' }] } }, + ]; + return lines.map(l => JSON.stringify(l)).join('\n') + '\n'; +} + +function walkFiles(dir: string): string[] { + if (!existsSync(dir)) return []; + const out: string[] = []; + for (const entry of readdirSync(dir, { withFileTypes: true })) { + const p = join(dir, entry.name); + if (entry.isDirectory()) out.push(...walkFiles(p)); + else out.push(p); + } + return out; +} + +function dbText(dbPath: string): string { + if (!existsSync(dbPath)) return ''; + const db = new Database(dbPath, { readonly: true }); + sqliteVec.load(db); + try { + const ex = db.prepare('SELECT user_message, assistant_message FROM exchanges').all(); + const tc = db.prepare('SELECT tool_input, tool_result FROM tool_calls').all(); + return JSON.stringify([ex, tc]); + } finally { + db.close(); + } +} + +function expectNoSecrets(haystack: string, s: Seed, where: string) { + for (const secret of s.secrets) { + expect(haystack.includes(secret), `${where} leaked a seeded secret`).toBe(false); + } +} + +describe('redaction pipeline', () => { + let root: string; + let sourceDir: string; + let archiveDir: string; + let dbPath: string; + let logged: string[]; + const savedEnv = { ...process.env }; + + beforeEach(() => { + root = mkdtempSync(join(tmpdir(), 'episodic-memory-redaction-pipeline-')); + sourceDir = join(root, 'source'); + archiveDir = join(root, 'archive'); + dbPath = join(root, 'db.sqlite'); + mkdirSync(join(sourceDir, '-work-contoso'), { recursive: true }); + process.env.TEST_DB_PATH = dbPath; + process.env.TEST_PROJECTS_DIR = sourceDir; + process.env.TEST_ARCHIVE_DIR = archiveDir; + process.env.EPISODIC_MEMORY_CONFIG_DIR = join(root, 'config'); + delete process.env.EPISODIC_MEMORY_REDACTION; + delete process.env.EPISODIC_MEMORY_REDACTION_RULES; + delete process.env.EPISODIC_MEMORY_REDACTION_STRICT; + delete process.env.EPISODIC_MEMORY_SKIP_SUMMARIES; + + embedSpy.mockClear(); + summarizeSpy.mockClear(); + logged = []; + for (const method of ['log', 'error', 'warn', 'info'] as const) { + vi.spyOn(console, method).mockImplementation((...args: unknown[]) => { + logged.push(args.map(a => (typeof a === 'string' ? a : JSON.stringify(a))).join(' ')); + }); + } + }); + + afterEach(() => { + vi.restoreAllMocks(); + process.env = { ...savedEnv }; + rmSync(root, { recursive: true, force: true }); + }); + + it('a seeded secret appears in zero of: archive, SQLite, embedding input, summarizer input, logs', async () => { + const s = seedSecrets(2024); + const claudeSrc = join(sourceDir, '-work-contoso', `${SESSION}.jsonl`); + writeFileSync(claudeSrc, claudeTranscript(s)); + writeFileSync(join(sourceDir, '-work-contoso', 'rollout-2026-05-12T18-00-00-019e4c75-d5bf-7c71-9df7-77f5fb86b711.jsonl'), codexRollout(s)); + + const result = await syncConversations(sourceDir, archiveDir); + expect(result.errors).toEqual([]); + expect(result.copied).toBe(2); + expect(result.indexed).toBe(2); + expect(result.summarized).toBe(2); + expect(result.redactions?.length).toBeGreaterThan(0); + expect(JSON.stringify(result.redactions)).not.toMatch(/[A-Za-z0-9+/]{30,}/); + + // (1) archive + const archived = walkFiles(archiveDir).filter(f => f.endsWith('.jsonl')); + expect(archived.length).toBe(2); + for (const file of archived) { + const text = readFileSync(file, 'utf-8'); + expectNoSecrets(text, s, `archive ${file}`); + expect(text).toContain('[REDACTED:'); + } + + // (2) SQLite text columns (exchanges + tool_calls) + const stored = dbText(dbPath); + expect(stored.length).toBeGreaterThan(0); + expectNoSecrets(stored, s, 'SQLite'); + expect(stored).toContain('[REDACTED:'); + + // (3) embedding input + expect(embedSpy).toHaveBeenCalled(); + expectNoSecrets(JSON.stringify(embedSpy.mock.calls), s, 'embedding input'); + + // (4) summarizer input — and resume/fork must be off, since those paths + // would hand the model the unredacted source transcript. + expect(summarizeSpy).toHaveBeenCalledTimes(2); + expectNoSecrets(JSON.stringify(summarizeSpy.mock.calls), s, 'summarizer input'); + for (const call of summarizeSpy.mock.calls) { + expect(call[2]).toMatchObject({ allowResume: false }); + } + + // (5) logs + expectNoSecrets(logged.join('\n'), s, 'console output'); + }); + + it('archive stays valid JSONL with the same line count, and show/read still render it', async () => { + const s = seedSecrets(7); + const src = join(sourceDir, '-work-contoso', `${SESSION}.jsonl`); + writeFileSync(src, claudeTranscript(s)); + await syncConversations(sourceDir, archiveDir, { skipSummaries: true }); + + const archived = join(archiveDir, '-work-contoso', `${SESSION}.jsonl`); + const srcLines = readFileSync(src, 'utf-8').split('\n'); + const outLines = readFileSync(archived, 'utf-8').split('\n'); + expect(outLines.length).toBe(srcLines.length); + outLines.filter(Boolean).forEach(l => expect(() => JSON.parse(l)).not.toThrow()); + + // The archive keeps the source mtime so the next sync treats it as current. + expect(Math.abs(statSync(archived).mtimeMs - statSync(src).mtimeMs)).toBeLessThan(2); + const again = await syncConversations(sourceDir, archiveDir, { skipSummaries: true }); + expect(again.copied).toBe(0); + + const md = formatConversationAsMarkdown(readFileSync(archived, 'utf-8')); + expect(md).toContain('[REDACTED:azure-client-secret]'); + expectNoSecrets(md, s, 'show output'); + }); + + it('git SHAs and GUIDs remain findable via text search', async () => { + const s = seedSecrets(31); + writeFileSync(join(sourceDir, '-work-contoso', `${SESSION}.jsonl`), claudeTranscript(s)); + await syncConversations(sourceDir, archiveDir, { skipSummaries: true }); + + const bySha = await searchConversations(s.sha, { mode: 'text' }); + expect(bySha.length).toBe(1); + const byGuid = await searchConversations(s.guid, { mode: 'text' }); + expect(byGuid.length).toBe(1); + const byToken = await searchConversations('[REDACTED:azure-client-secret]', { mode: 'text' }); + expect(byToken.length).toBe(1); + }); + + it('fails closed in strict mode: a corrupt rules file leaves no new archive or index content', async () => { + const s = seedSecrets(5); + writeFileSync(join(sourceDir, '-work-contoso', `${SESSION}.jsonl`), claudeTranscript(s)); + const rules = join(root, 'rules.json'); + writeFileSync(rules, '{ "rules": [ { "id": "x", "pattern": "(" } ] }'); + process.env.EPISODIC_MEMORY_REDACTION_RULES = rules; + + await expect(syncConversations(sourceDir, archiveDir)).rejects.toThrow(/redaction/i); + expect(walkFiles(archiveDir)).toEqual([]); + expect(dbText(dbPath)).toBe(''); + expect(embedSpy).not.toHaveBeenCalled(); + expect(summarizeSpy).not.toHaveBeenCalled(); + }); + + it('EPISODIC_MEMORY_REDACTION=off restores the byte-for-byte copy and resume-capable summaries', async () => { + process.env.EPISODIC_MEMORY_REDACTION = 'off'; + const s = seedSecrets(8); + const src = join(sourceDir, '-work-contoso', `${SESSION}.jsonl`); + writeFileSync(src, claudeTranscript(s)); + await syncConversations(sourceDir, archiveDir); + expect(readFileSync(join(archiveDir, '-work-contoso', `${SESSION}.jsonl`), 'utf-8')).toBe(readFileSync(src, 'utf-8')); + expect(summarizeSpy.mock.calls[0][2]).toMatchObject({ allowResume: true }); + }); + + it('.NET/Azure key-context secrets (no recognizable shape) never reach archive or SQLite', async () => { + const fake = new FakeSecrets(1999); + const appsettingsSecret = fake.passphrase(); + const appSettingValue = fake.passphrase(); + const vaultValue = fake.passphrase(); + const mcpSecret = fake.passphrase(); + const seed: Seed = { secrets: [appsettingsSecret, appSettingValue, vaultValue, mcpSecret], sha: fake.gitSha(), guid: fake.guid() }; + const base = { sessionId: SESSION, isSidechain: false, cwd: '/work/contoso' }; + const lines = [ + { ...base, type: 'user', uuid: 'u1', parentUuid: null, timestamp: '2026-01-01T00:00:00Z', message: { role: 'user', content: 'Why does the API fail to get a token?' } }, + { ...base, type: 'assistant', uuid: 'a1', parentUuid: 'u1', timestamp: '2026-01-01T00:00:01Z', message: { role: 'assistant', content: [ + { type: 'tool_use', id: 't1', name: 'Read', input: { file_path: '/work/contoso/appsettings.Development.json' } }, + ] } }, + { ...base, type: 'user', uuid: 'u2', parentUuid: 'a1', timestamp: '2026-01-01T00:00:02Z', + message: { role: 'user', content: [{ type: 'tool_result', tool_use_id: 't1', content: JSON.stringify({ AzureAd: { TenantId: seed.guid, ClientSecret: appsettingsSecret } }, null, 2) }] } }, + { ...base, type: 'assistant', uuid: 'a2', parentUuid: 'u2', timestamp: '2026-01-01T00:00:03Z', message: { role: 'assistant', content: [ + { type: 'tool_use', id: 't2', name: 'Bash', input: { command: 'az webapp config appsettings list -g rg -n app && az keyvault secret show --vault-name kv -n Db' } }, + ] } }, + { ...base, type: 'user', uuid: 'u3', parentUuid: 'a2', timestamp: '2026-01-01T00:00:04Z', + message: { role: 'user', content: [{ type: 'tool_result', tool_use_id: 't2', content: + JSON.stringify([{ name: 'Stripe__ApiKey', slotSetting: false, value: appSettingValue }], null, 2) + '\n' + + JSON.stringify({ id: 'https://kv.vault.azure.net/secrets/Db/0123', value: vaultValue }, null, 2) }] }, + toolUseResult: { structuredContent: { name: 'GraphClientSecret', value: mcpSecret } } }, + { ...base, type: 'assistant', uuid: 'a3', parentUuid: 'u3', timestamp: '2026-01-01T00:00:05Z', message: { role: 'assistant', content: [{ type: 'text', text: 'The client secret has expired.' }] } }, + ]; + writeFileSync(join(sourceDir, '-work-contoso', `${SESSION}.jsonl`), lines.map(l => JSON.stringify(l)).join('\n') + '\n'); + + const result = await syncConversations(sourceDir, archiveDir, { skipSummaries: true }); + expect(result.errors).toEqual([]); + const archived = readFileSync(join(archiveDir, '-work-contoso', `${SESSION}.jsonl`), 'utf-8'); + expectNoSecrets(archived, seed, 'archive'); + expectNoSecrets(dbText(dbPath), seed, 'SQLite'); + expectNoSecrets(JSON.stringify(embedSpy.mock.calls), seed, 'embedding input'); + expect(archived).toContain(seed.guid); + archived.split('\n').filter(Boolean).forEach(l => expect(() => JSON.parse(l)).not.toThrow()); + }); + + it('the `index` path (indexUnprocessed) parses the redacted archive, not the source', async () => { + const s = seedSecrets(77); + writeFileSync(join(sourceDir, '-work-contoso', `${SESSION}.jsonl`), claudeTranscript(s)); + + await indexUnprocessed(1, false); + + const archived = walkFiles(archiveDir).filter(f => f.endsWith('.jsonl')); + expect(archived.length).toBe(1); + expectNoSecrets(readFileSync(archived[0], 'utf-8'), s, 'archive (index path)'); + expectNoSecrets(dbText(dbPath), s, 'SQLite (index path)'); + expectNoSecrets(JSON.stringify(embedSpy.mock.calls), s, 'embedding input (index path)'); + expect(summarizeSpy).toHaveBeenCalledTimes(1); + expectNoSecrets(JSON.stringify(summarizeSpy.mock.calls), s, 'summarizer input (index path)'); + expect(summarizeSpy.mock.calls[0][2]).toMatchObject({ allowResume: false }); + }); + + it('the `index` path refreshes an indexed archive even when the append left the source mtime no newer', async () => { + const fake = new FakeSecrets(78); + const secret = fake.azureClientSecret(); + const base = { sessionId: SESSION, isSidechain: false, cwd: '/work/contoso', gitBranch: 'main', version: '2.0.0' }; + const exchange = (n: number, text: string) => [ + { ...base, type: 'user', uuid: `u${n}`, parentUuid: n > 1 ? `a${n - 1}` : null, + timestamp: new Date(Date.UTC(2026, 0, 1, 0, n)).toISOString(), + message: { role: 'user', content: text } }, + { ...base, type: 'assistant', uuid: `a${n}`, parentUuid: `u${n}`, + timestamp: new Date(Date.UTC(2026, 0, 1, 0, n, 30)).toISOString(), + message: { role: 'assistant', content: [{ type: 'text', text: `Reply ${n}.` }] } }, + ].map(l => JSON.stringify(l) + '\n').join(''); + const src = join(sourceDir, '-work-contoso', `${SESSION}.jsonl`); + + writeFileSync(src, exchange(1, 'First question.')); + await indexUnprocessed(1, true); + const archived = walkFiles(archiveDir).filter(f => f.endsWith('.jsonl'))[0]; + const archivedMtime = statSync(archived).mtime; + + // Append, then pin the source mtime to the archive's: copyIfNewer's mtime + // check alone would treat the archive as current. + appendFileSync(src, exchange(2, `Second question with ${secret} as the secret`)); + utimesSync(src, archivedMtime, archivedMtime); + await indexUnprocessed(1, true); + + const text = dbText(dbPath); + expect(text).toContain('Second question with [REDACTED:'); + expect(text).not.toContain(secret); + }); +}); + +describe('redaction: opencode staging export', () => { + let root: string; + const savedEnv = { ...process.env }; + + beforeEach(() => { + root = mkdtempSync(join(tmpdir(), 'episodic-memory-redaction-opencode-')); + process.env.EPISODIC_MEMORY_OPENCODE_DB_PATH = join(root, 'opencode.db'); + process.env.EPISODIC_MEMORY_OPENCODE_TRANSCRIPT_DIR = join(root, 'transcripts'); + }); + + afterEach(() => { + process.env = { ...savedEnv }; + rmSync(root, { recursive: true, force: true }); + }); + + it('redacts secrets before the staging JSONL is written', () => { + const fake = new FakeSecrets(55); + const secret = fake.azureStorageKey(); + const db = new Database(join(root, 'opencode.db')); + db.exec(` + CREATE TABLE project (id TEXT PRIMARY KEY, worktree TEXT NOT NULL, name TEXT, time_created INTEGER NOT NULL, time_updated INTEGER NOT NULL, sandboxes TEXT NOT NULL); + CREATE TABLE session (id TEXT PRIMARY KEY, project_id TEXT NOT NULL, slug TEXT NOT NULL, directory TEXT NOT NULL, title TEXT NOT NULL, version TEXT NOT NULL, time_created INTEGER NOT NULL, time_updated INTEGER NOT NULL, agent TEXT, model TEXT); + CREATE TABLE message (id TEXT PRIMARY KEY, session_id TEXT NOT NULL, time_created INTEGER NOT NULL, time_updated INTEGER NOT NULL, data TEXT NOT NULL); + CREATE TABLE part (id TEXT PRIMARY KEY, message_id TEXT NOT NULL, session_id TEXT NOT NULL, time_created INTEGER NOT NULL, time_updated INTEGER NOT NULL, data TEXT NOT NULL); + `); + db.prepare(`INSERT INTO project VALUES ('p1', '/work/x', 'X', 1700000000000, 1700000000000, '[]')`).run(); + db.prepare(`INSERT INTO session VALUES ('ses_1', 'p1', 's', '/work/x', 'T', '1.0.0', 1700000000000, 1700000004000, 'build', NULL)`).run(); + db.prepare(`INSERT INTO message VALUES ('m1', 'ses_1', 1700000001000, 1700000001000, ?)`).run(JSON.stringify({ role: 'user' })); + db.prepare(`INSERT INTO part VALUES ('pt1', 'm1', 'ses_1', 1700000001000, 1700000001000, ?)`).run( + JSON.stringify({ type: 'text', text: `AccountKey=${secret}` }) + ); + db.close(); + + const result = exportOpencodeSessions({ redactor: createRedactor(DEFAULT_REDACTION_CONFIG) }); + expect(result.exported).toBe(1); + const files = walkFiles(join(root, 'transcripts')); + expect(files.length).toBe(1); + const text = readFileSync(files[0], 'utf-8'); + expect(text).not.toContain(secret); + expect(text).toContain('[REDACTED:connection-string-secret]'); + }); +}); diff --git a/test/redaction.test.ts b/test/redaction.test.ts new file mode 100644 index 00000000..3ac42229 --- /dev/null +++ b/test/redaction.test.ts @@ -0,0 +1,736 @@ +import { describe, it, expect, beforeEach, afterEach, vi } from 'vitest'; +import { mkdtempSync, writeFileSync, readFileSync, rmSync, mkdirSync, chmodSync, statSync } from 'fs'; +import { join } from 'path'; +import { tmpdir } from 'os'; + +import { + createRedactor, + loadRedactor, + loadRedactionConfig, + getRedactionSettings, + redactJsonlLine, + copyFileRedacted, + FindingsTally, + RedactionConfigError, + DEFAULT_REDACTION_CONFIG, + type Redactor, +} from '../src/redaction.js'; +import { FakeSecrets, redactionTokens } from './fake-secrets.js'; + +const ctx = { source: 'test', path: '/tmp/x.jsonl' }; + +function defaults(): Redactor { + return createRedactor(DEFAULT_REDACTION_CONFIG); +} + +/** Assert `secret` is gone, the expected token is present, and findings name the rule. */ +function expectRedacted(redactor: Redactor, text: string, secret: string, ruleId: string) { + const result = redactor.redact(text, ctx); + expect(result.text).not.toContain(secret); + expect(result.text).toContain(`[REDACTED:${ruleId}]`); + expect(result.findings.find(f => f.ruleId === ruleId)?.count ?? 0).toBeGreaterThan(0); + // Findings carry rule IDs and counts, never the matched value. + expect(JSON.stringify(result.findings)).not.toContain(secret); + return result; +} + +describe('redaction: default rules (positive)', () => { + let fake: FakeSecrets; + let redactor: Redactor; + + beforeEach(() => { + fake = new FakeSecrets(42); + redactor = defaults(); + }); + + it('azure-client-secret in pasted `az ad sp credential reset` output', () => { + const secret = fake.azureClientSecret(); + const text = `{\n "appId": "${fake.guid()}",\n "password": "${secret}",\n "tenant": "${fake.guid()}"\n}`; + expectRedacted(redactor, text, secret, 'azure-client-secret'); + }); + + it('azure-client-secret bare in prose', () => { + const secret = fake.azureClientSecret(); + expectRedacted(redactor, `use ${secret} as the client secret`, secret, 'azure-client-secret'); + }); + + it('azure-client-secret at the end of a sentence', () => { + const secret = fake.azureClientSecret(); + const out = expectRedacted(redactor, `The client secret is ${secret}.`, secret, 'azure-client-secret'); + expect(out.text).toBe('The client secret is [REDACTED:azure-client-secret].'); + }); + + it('azure-storage-key in `az storage account keys list` output', () => { + const secret = fake.azureStorageKey(); + const text = `[\n {\n "keyName": "key1",\n "permissions": "FULL",\n "value": "${secret}"\n }\n]`; + expectRedacted(redactor, text, secret, 'azure-storage-key'); + }); + + it('connection-string-secret redacts only the value, keeping account and endpoint searchable', () => { + const key = fake.azureStorageKey(); + const text = `DefaultEndpointsProtocol=https;AccountName=contosodata;AccountKey=${key};EndpointSuffix=core.windows.net`; + const result = expectRedacted(redactor, text, key, 'connection-string-secret'); + expect(result.text).toContain('AccountName=contosodata'); + expect(result.text).toContain('EndpointSuffix=core.windows.net'); + }); + + it('connection-string-secret in an appsettings.json blob (SQL Server password field)', () => { + const pw = fake.password(); + const text = JSON.stringify({ + ConnectionStrings: { + Default: `Server=tcp:contoso-sql.database.windows.net,1433;Database=orders;User ID=app;Password=${pw};Encrypt=True;`, + }, + }, null, 2); + const result = expectRedacted(redactor, text, pw, 'connection-string-secret'); + expect(result.text).toContain('contoso-sql.database.windows.net'); + expect(result.text).toContain('Database=orders'); + }); + + it('connection-string-secret for Service Bus SharedAccessKey', () => { + const key = fake.chars('ABCDEFGHIJKLMNOPQRSTUVWXYZabcdefghijklmnopqrstuvwxyz0123456789+/', 43) + '='; + const text = `Endpoint=sb://contoso.servicebus.windows.net/;SharedAccessKeyName=RootManageSharedAccessKey;SharedAccessKey=${key}`; + const result = expectRedacted(redactor, text, key, 'connection-string-secret'); + expect(result.text).toContain('SharedAccessKeyName=RootManageSharedAccessKey'); + }); + + it('connection-string-secret keywords are case-insensitive, as in ADO.NET', () => { + // Letters only: secret-assignment needs a digit, so this rule is the only catch. + const pw = fake.chars('abcdefghijklmnopqrstuvwxyz', 19); + for (const keyword of ['password', 'PASSWORD', 'pwd', 'accountkey']) { + const result = expectRedacted(redactor, `Server=db;Database=app;${keyword}=${pw};`, pw, 'connection-string-secret'); + expect(result.text).toContain(`${keyword}=[REDACTED:connection-string-secret];`); + } + }); + + it('azure-sas-token redacts the sig while keeping the blob URL', () => { + const sig = fake.sasSignature(); + const text = `https://contoso.blob.core.windows.net/backups/db.bak?sv=2022-11-02&ss=b&srt=co&sp=rl&se=2026-12-31T00:00:00Z&sig=${sig}`; + const result = expectRedacted(redactor, text, sig, 'azure-sas-token'); + expect(result.text).toContain('https://contoso.blob.core.windows.net/backups/db.bak?sv=2022-11-02'); + }); + + it('private-key-block (multi-line, PEM)', () => { + const block = fake.privateKeyBlock('RSA'); + const body = block.split('\n')[2]; + const result = expectRedacted(redactor, `cat key.pem\n${block}\n$ `, body, 'private-key-block'); + expect(result.text).not.toContain('PRIVATE KEY'); + }); + + it('private-key-block (OPENSSH, escaped newlines as in nested JSON)', () => { + const block = fake.privateKeyBlock('OPENSSH').replace(/\n/g, '\\n'); + const body = block.split('\\n')[3]; + expectRedacted(redactor, `{"content":"${block}"}`, body, 'private-key-block'); + }); + + it('private-key-block truncated (no END marker) is still redacted to end of text', () => { + const block = fake.privateKeyBlock('EC'); + const truncated = block.split('\n').slice(0, 4).join('\n'); + const body = truncated.split('\n')[2]; + expectRedacted(redactor, truncated, body, 'private-key-block'); + }); + + it('jwt (Azure access token from `az account get-access-token`)', () => { + const token = fake.jwt(); + const text = `{\n "accessToken": "${token}",\n "expiresOn": "2026-10-04 12:00:00.000000",\n "tokenType": "Bearer"\n}`; + const result = redactor.redact(text, ctx); + expect(result.text).not.toContain(token); + expect(redactionTokens(result.text).length).toBeGreaterThan(0); + }); + + it('bearer-token in an Authorization header', () => { + const token = fake.bearerOpaque(); + expectRedacted(redactor, `curl -H "Authorization: Bearer ${token}" https://api.example.com`, token, 'bearer-token'); + }); + + it('basic-auth in an Authorization header', () => { + const token = fake.basicAuth(); + expectRedacted(redactor, `Authorization: Basic ${token}`, token, 'basic-auth'); + }); + + it('anthropic-api-key', () => { + const key = fake.anthropicKey(); + expectRedacted(redactor, `export ANTHROPIC_API_KEY=${key}`, key, 'anthropic-api-key'); + }); + + it('openai-api-key', () => { + const key = fake.openAiKey(); + expectRedacted(redactor, `OPENAI_API_KEY="${key}"`, key, 'openai-api-key'); + }); + + it('github-token (classic) and fine-grained PAT', () => { + const classic = fake.githubToken(); + expectRedacted(redactor, `git remote set-url origin https://${classic}@github.com/o/r.git`, classic, 'github-token'); + const fine = fake.githubFineGrainedPat(); + expectRedacted(redactor, `GH_TOKEN=${fine}`, fine, 'github-token'); + }); + + it('aws-access-key-id and aws-secret-access-key in ~/.aws/credentials', () => { + const id = fake.awsAccessKeyId(); + const secret = fake.awsSecretAccessKey(); + const text = `[default]\naws_access_key_id = ${id}\naws_secret_access_key = ${secret}\n`; + expectRedacted(redactor, text, id, 'aws-access-key-id'); + expectRedacted(redactor, text, secret, 'aws-secret-access-key'); + }); + + it('slack-token, google-api-key, npm-token', () => { + const slack = fake.slackToken(); + expectRedacted(redactor, `SLACK_BOT_TOKEN=${slack}`, slack, 'slack-token'); + const google = fake.googleApiKey(); + expectRedacted(redactor, `key=${google}&q=coffee`, google, 'google-api-key'); + const npm = fake.npmToken(); + expectRedacted(redactor, `//registry.npmjs.org/:_authToken=${npm}`, npm, 'npm-token'); + }); + + it('url-credentials redacts only the password', () => { + const pw = fake.password().replace(/[#%^*!]/g, 'x'); + const text = `DATABASE_URL=postgres://app_user:${pw}@db.internal:5432/orders`; + const result = expectRedacted(redactor, text, pw, 'url-credentials'); + expect(result.text).toContain('app_user'); + expect(result.text).toContain('@db.internal:5432/orders'); + }); + + it('secret-assignment covers env, YAML, and JSON shapes (decrypted SOPS output)', () => { + const a = fake.password(); + const b = fake.password(); + const c = fake.password(); + const yaml = `database:\n host: db.internal\n password: ${a}\nazure:\n client_secret: "${b}"\n`; + const env = `AZURE_CLIENT_SECRET=${c}`; + for (const [text, secret] of [[yaml, a], [yaml, b], [env, c]] as const) { + const result = redactor.redact(text, ctx); + expect(result.text).not.toContain(secret); + } + expect(redactor.redact(yaml, ctx).text).toContain('host: db.internal'); + }); + + it('quoted-secret-assignment redacts a JSON "clientSecret" field', () => { + const secret = fake.password(); + expectRedacted(redactor, JSON.stringify({ clientId: fake.guid(), clientSecret: secret }), secret, 'quoted-secret-assignment'); + }); +}); + +describe('redaction: allowlist and false-positive resistance (negative)', () => { + let fake: FakeSecrets; + let redactor: Redactor; + + beforeEach(() => { + fake = new FakeSecrets(7); + redactor = defaults(); + }); + + function expectUntouched(text: string) { + const result = redactor.redact(text, ctx); + expect(result.text).toBe(text); + expect(result.findings).toEqual([]); + } + + it('git SHAs (full and short) survive', () => { + const sha = fake.gitSha(); + expectUntouched(`commit ${sha}\nAuthor: someone\n\n fix: thing (${sha.slice(0, 7)})`); + }); + + it('a git SHA assigned to a token-like key is still allowlisted', () => { + const sha = fake.gitSha(); + expectUntouched(`token: ${sha}`); + }); + + it('GUIDs (tenant, client, object IDs) survive', () => { + expectUntouched(JSON.stringify({ tenantId: fake.guid(), clientId: fake.guid(), objectId: fake.guid() })); + }); + + it('ARM resource IDs survive', () => { + expectUntouched( + `/subscriptions/${fake.guid()}/resourceGroups/rg-prod-eastus/providers/Microsoft.Storage/storageAccounts/contosodata` + ); + }); + + it('error codes and HTTP statuses survive', () => { + expectUntouched('AADSTS7000215: Invalid client secret provided. Status: 401 Unauthorized. AuthorizationFailed (403)'); + }); + + it('base64 image data survives', () => { + expectUntouched(JSON.stringify({ type: 'image', source: { type: 'base64', media_type: 'image/png', data: 'iVBORw0KGgo' + fake.base64Blob(20000) + '==' } })); + }); + + it('npm integrity hashes (sha512, 88 chars) survive', () => { + expectUntouched(`"integrity": "sha512-${fake.base64Blob(86)}=="`); + }); + + it('minified JS survives', () => { + expectUntouched( + 'function a(e,t){var n=e.password,r=t.token;return n&&r?{password:n,token:r,expires_in:3600}:null}' + + 'const s=new URLSearchParams({grant_type:"client_credentials"});' + ); + }); + + it('code that reads secrets (not literals) survives', () => { + expectUntouched('const password = getPassword();\nconst token = process.env.GITHUB_TOKEN;\npassword = os.environ["DB_PASSWORD"]'); + }); + + it('member-access references to secrets survive, even with digits (env.AUTH0_CLIENT_SECRET)', () => { + expectUntouched('client_id: env.AUTH0_CLIENT_ID,\n client_secret: env.AUTH0_CLIENT_SECRET,\n token: this.config.apiToken2;'); + }); + + it('placeholders and SOPS-encrypted values survive', () => { + expectUntouched('password: ${DB_PASSWORD}\nclient_secret: \napi_key: ENC[AES256_GCM,data:abc,iv:def,tag:ghi,type:str]'); + }); + + it('PWD and other env noise survive', () => { + expectUntouched('PWD=/home/user1/src/project2\nOLDPWD=/home/user1\nmax_tokens: 4096\ntokenizer: bert-base-uncased'); + }); + + it('Markdown about connection-string keywords survives', () => { + expectUntouched('Set `AccountKey=`, `SharedAccessKey=`, `Password=` or `Pwd=` in the connection string.'); + }); + + it('prose about tokens and passwords survives', () => { + expectUntouched('Rotate the bearer token every 90 days. Basic authentication is disabled. The password policy requires 14 characters.'); + }); +}); + +describe('redaction: idempotency', () => { + it('redacting already-redacted text is a no-op and tokens match no rule', () => { + const fake = new FakeSecrets(99); + const redactor = defaults(); + const text = [ + `AccountKey=${fake.azureStorageKey()}`, + `password: ${fake.password()}`, + `Authorization: Bearer ${fake.jwt()}`, + fake.privateKeyBlock(), + `client secret ${fake.azureClientSecret()}`, + `https://u:${fake.password().replace(/[#%^*!]/g, 'y')}@host/x`, + ].join('\n'); + const once = redactor.redact(text, ctx); + expect(once.findings.length).toBeGreaterThan(0); + const twice = redactor.redact(once.text, ctx); + expect(twice.text).toBe(once.text); + expect(twice.findings).toEqual([]); + }); + + it('every token for every default rule id is inert', () => { + const redactor = defaults(); + for (const id of redactor.ruleIds) { + for (const shape of [`[REDACTED:${id}]`, `password: [REDACTED:${id}]`, `Bearer [REDACTED:${id}]`, `AccountKey=[REDACTED:${id}];`]) { + const result = redactor.redact(shape, ctx); + expect(result.findings, `${id} in ${shape}`).toEqual([]); + } + } + }); + + it('a match that runs into an existing token redacts the part outside it, keeping the token', () => { + // Text already partly redacted, e.g. by an older rule set before `redact --rewrite`. + const redactor = defaults(); + const fake = new FakeSecrets(12); + const prefix = fake.chars('abcdefghijklmnopqrstuvwxyz0123456789', 10); + const once = redactor.redact(`Server=db;Password=${prefix}[REDACTED:jwt];`, ctx); + expect(once.text).toBe('Server=db;Password=[REDACTED:connection-string-secret][REDACTED:jwt];'); + const twice = redactor.redact(once.text, ctx); + expect(twice.text).toBe(once.text); + expect(twice.findings).toEqual([]); + }); +}); + +describe('redaction: entropy fallback', () => { + it('is off by default', () => { + const fake = new FakeSecrets(3); + const blob = fake.opaqueHighEntropy(); + const result = defaults().redact(`the signing value is ${blob}`, ctx); + expect(result.text).toContain(blob); + }); + + it('when enabled, fires near a keyword but not elsewhere or on allowlisted shapes', () => { + const fake = new FakeSecrets(4); + const redactor = createRedactor({ + ...DEFAULT_REDACTION_CONFIG, + entropy: { ...DEFAULT_REDACTION_CONFIG.entropy, enabled: true }, + }); + const blob = fake.opaqueHighEntropy(48); + const near = redactor.redact(`signing key is ${blob}`, ctx); + expect(near.text).not.toContain(blob); + expect(near.text).toContain('[REDACTED:high-entropy]'); + + const far = redactor.redact(`the build artifact digest ${fake.opaqueHighEntropy(48)}`, ctx); + expect(far.findings).toEqual([]); + + const sha = fake.gitSha(); + expect(redactor.redact(`key ${sha}`, ctx).text).toContain(sha); + }); +}); + +describe('redaction: JSONL line handling', () => { + const redactor = defaults(); + const fake = new FakeSecrets(11); + + it('redacts string values inside JSON and keeps the line valid JSON', () => { + const key = fake.azureStorageKey(); + const line = JSON.stringify({ + type: 'user', + message: { role: 'user', content: [{ type: 'tool_result', tool_use_id: 't1', content: `AccountKey=${key}` }] }, + uuid: 'u1', + }); + const out = redactJsonlLine(line, redactor, ctx); + expect(out).not.toContain(key); + const parsed = JSON.parse(out); + expect(parsed.uuid).toBe('u1'); + expect(parsed.message.content[0].content).toBe('AccountKey=[REDACTED:connection-string-secret]'); + }); + + it('redacts tool inputs (nested objects) too', () => { + const pw = fake.password(); + const line = JSON.stringify({ + type: 'assistant', + message: { role: 'assistant', content: [{ type: 'tool_use', id: 't1', name: 'Bash', input: { command: `mysql -u root --password=${pw}` } }] }, + }); + const out = redactJsonlLine(line, redactor, ctx); + expect(out).not.toContain(pw); + expect(() => JSON.parse(out)).not.toThrow(); + }); + + it('returns unchanged lines byte-for-byte (formatting preserved)', () => { + const line = '{"type":"user", "message":{"role":"user","content":"caf\\u00e9 hello"}}'; + expect(redactJsonlLine(line, redactor, ctx)).toBe(line); + }); + + it('redacts non-JSON lines (e.g. a torn last line) as raw text', () => { + const key = fake.azureStorageKey(); + const torn = `{"type":"user","message":{"content":"AccountKey=${key}`; + const out = redactJsonlLine(torn, redactor, ctx); + expect(out).not.toContain(key); + }); + + it('preserves CRLF line endings', () => { + const pw = fake.password(); + const line = JSON.stringify({ content: `password: ${pw}` }) + '\r'; + const out = redactJsonlLine(line, redactor, ctx); + expect(out.endsWith('\r')).toBe(true); + expect(out).not.toContain(pw); + }); + + it('tallies findings by rule id', () => { + const tally = new FindingsTally(); + redactJsonlLine(JSON.stringify({ a: `password: ${fake.password()}`, b: `password: ${fake.password()}` }), redactor, ctx, tally); + expect(tally.toArray()).toEqual([{ ruleId: 'secret-assignment', count: 2 }]); + expect(tally.total).toBe(2); + }); +}); + +describe('redaction: copyFileRedacted', () => { + let dir: string; + beforeEach(() => { dir = mkdtempSync(join(tmpdir(), 'episodic-memory-redact-copy-')); }); + afterEach(() => { rmSync(dir, { recursive: true, force: true }); }); + + it('streams a file, keeps one output line per input line, and keeps the trailing-newline shape', () => { + const fake = new FakeSecrets(5); + const redactor = defaults(); + const secrets: string[] = []; + const lines: string[] = []; + for (let i = 0; i < 3000; i++) { + if (i % 500 === 0) { + const s = fake.azureClientSecret(); + secrets.push(s); + lines.push(JSON.stringify({ type: 'user', i, message: { content: `secret is ${s}` } })); + } else { + lines.push(JSON.stringify({ type: 'assistant', i, message: { content: 'x'.repeat(i % 700) } })); + } + } + const src = join(dir, 'src.jsonl'); + const dest = join(dir, 'dest.jsonl'); + writeFileSync(src, lines.join('\n')); // no trailing newline + const tally = new FindingsTally(); + copyFileRedacted(src, dest, redactor, ctx, tally); + const out = readFileSync(dest, 'utf-8'); + for (const s of secrets) expect(out).not.toContain(s); + expect(out.split('\n').length).toBe(lines.length); + expect(out.endsWith('\n')).toBe(false); + expect(tally.total).toBe(secrets.length); + out.split('\n').forEach(l => expect(() => JSON.parse(l)).not.toThrow()); + }); + + it('handles multi-byte UTF-8 across chunk boundaries', () => { + const redactor = defaults(); + const src = join(dir, 'utf8.jsonl'); + const dest = join(dir, 'utf8-out.jsonl'); + const line = JSON.stringify({ content: '日本語テキスト🙂'.repeat(200_000) }); + writeFileSync(src, line + '\n' + line + '\n'); + copyFileRedacted(src, dest, redactor, ctx); + expect(readFileSync(dest, 'utf-8')).toBe(readFileSync(src, 'utf-8')); + }); + + // Windows ignores POSIX mode bits. + it.skipIf(process.platform === 'win32')('creates dest with the source mode, not the 0666 default', () => { + const src = join(dir, 'private.jsonl'); + const dest = join(dir, 'private-out.jsonl'); + writeFileSync(src, '{"type":"user"}\n'); + chmodSync(src, 0o600); + copyFileRedacted(src, dest, defaults(), ctx); + expect(statSync(dest).mode & 0o777).toBe(0o600); + }); +}); + +describe('redaction: settings and rules loading', () => { + let dir: string; + const saved = { ...process.env }; + + beforeEach(() => { + dir = mkdtempSync(join(tmpdir(), 'episodic-memory-redact-cfg-')); + process.env.EPISODIC_MEMORY_CONFIG_DIR = join(dir, 'config'); + mkdirSync(process.env.EPISODIC_MEMORY_CONFIG_DIR, { recursive: true }); + delete process.env.EPISODIC_MEMORY_REDACTION; + delete process.env.EPISODIC_MEMORY_REDACTION_RULES; + delete process.env.EPISODIC_MEMORY_REDACTION_STRICT; + }); + + afterEach(() => { + process.env = { ...saved }; + rmSync(dir, { recursive: true, force: true }); + }); + + it('is on and strict by default', () => { + expect(getRedactionSettings(process.env)).toMatchObject({ enabled: true, strict: true }); + }); + + it('EPISODIC_MEMORY_REDACTION=off disables it', () => { + process.env.EPISODIC_MEMORY_REDACTION = 'off'; + expect(getRedactionSettings(process.env).enabled).toBe(false); + expect(loadRedactor(process.env)).toBeNull(); + }); + + it('loads the bundled defaults when no rules file exists', () => { + const redactor = loadRedactor(process.env)!; + expect(redactor.ruleIds).toEqual(DEFAULT_REDACTION_CONFIG.rules.map(r => r.id)); + }); + + it('picks up redaction-rules.json from the config dir and merges it with defaults', () => { + writeFileSync(join(process.env.EPISODIC_MEMORY_CONFIG_DIR!, 'redaction-rules.json'), JSON.stringify({ + rules: [{ id: 'internal-ticket-secret', pattern: 'TKT-SECRET-[0-9]{6}' }], + disableRules: ['basic-auth'], + })); + const redactor = loadRedactor(process.env)!; + expect(redactor.ruleIds).toContain('internal-ticket-secret'); + expect(redactor.ruleIds).toContain('azure-storage-key'); + expect(redactor.ruleIds).not.toContain('basic-auth'); + expect(redactor.redact('ref TKT-SECRET-123456', ctx).text).toBe('ref [REDACTED:internal-ticket-secret]'); + }); + + it('EPISODIC_MEMORY_REDACTION_RULES points at a custom file; includeDefaults:false replaces the defaults', () => { + const file = join(dir, 'custom.json'); + writeFileSync(file, JSON.stringify({ includeDefaults: false, rules: [{ id: 'only-this', pattern: 'zzz[0-9]+' }] })); + process.env.EPISODIC_MEMORY_REDACTION_RULES = file; + const redactor = loadRedactor(process.env)!; + expect(redactor.ruleIds).toEqual(['only-this']); + }); + + it('strict mode throws RedactionConfigError on a corrupt rules file', () => { + const file = join(dir, 'bad.json'); + writeFileSync(file, '{ "rules": [ this is not json'); + process.env.EPISODIC_MEMORY_REDACTION_RULES = file; + expect(() => loadRedactor(process.env)).toThrow(RedactionConfigError); + }); + + it('strict mode throws when the custom rules path does not exist', () => { + process.env.EPISODIC_MEMORY_REDACTION_RULES = join(dir, 'missing.json'); + expect(() => loadRedactor(process.env)).toThrow(RedactionConfigError); + }); + + it('non-strict mode warns and returns null (pass-through) on a corrupt rules file', () => { + const file = join(dir, 'bad.json'); + writeFileSync(file, 'nope'); + process.env.EPISODIC_MEMORY_REDACTION_RULES = file; + process.env.EPISODIC_MEMORY_REDACTION_STRICT = '0'; + const warn = vi.fn(); + expect(loadRedactor(process.env, warn)).toBeNull(); + expect(warn).toHaveBeenCalled(); + }); + + it('rejects invalid rules: bad regex, bad id, empty-matching pattern, bad secretGroup', () => { + const bad = [ + { id: 'bad-regex', pattern: '(unclosed' }, + { id: 'Bad Id!', pattern: 'abc' }, + { id: 'empty-match', pattern: 'a*' }, + { id: 'bad-group', pattern: 'abc', secretGroup: 2 }, + ]; + for (const rule of bad) { + expect(() => loadRedactionConfig({ includeDefaults: false, rules: [rule] } as any), rule.id).toThrow(RedactionConfigError); + } + }); + + it('rule-load errors never echo text being redacted (only config details)', () => { + const file = join(dir, 'bad.json'); + writeFileSync(file, JSON.stringify({ rules: [{ id: 'x', pattern: '(' }] })); + process.env.EPISODIC_MEMORY_REDACTION_RULES = file; + try { + loadRedactor(process.env); + expect.unreachable(); + } catch (e) { + expect((e as Error).message).toContain('x'); + } + }); +}); + +describe('redaction: key context (.NET / Azure shapes)', () => { + let fake: FakeSecrets; + let redactor: Redactor; + beforeEach(() => { + fake = new FakeSecrets(2112); + redactor = defaults(); + }); + + function expectGone(text: string, secret: string, ruleId?: string) { + const result = redactor.redact(text, ctx); + expect(result.text, text).not.toContain(secret); + if (ruleId) expect(result.text).toContain(`[REDACTED:${ruleId}]`); + return result; + } + + it('appsettings.json fields, no digit required when the key is quoted JSON', () => { + const pw = fake.passphrase(); + const text = JSON.stringify({ AzureAd: { TenantId: fake.guid(), ClientId: fake.guid(), ClientSecret: pw } }, null, 2); + const r = expectGone(text, pw, 'quoted-secret-assignment'); + expect(r.text).toContain('"TenantId"'); + }); + + it('C# object initializers and locals', () => { + const a = fake.passphrase(); + const b = fake.passphrase(); + expectGone(`var cred = new ClientSecretCredential(tenant, client) { ClientSecret = "${a}", };`, a, 'quoted-secret-assignment'); + expectGone(`string apiKey = "${b}";`, b, 'quoted-secret-assignment'); + }); + + it('`az webapp config appsettings list` name/value pairs (name first and value first)', () => { + const a = fake.passphrase(); + const b = fake.passphrase(); + const text = JSON.stringify([ + { name: 'WEBSITE_RUN_FROM_PACKAGE', slotSetting: false, value: '1' }, + { name: 'Stripe__ApiKey', slotSetting: false, value: a }, + { value: b, slotSetting: true, name: 'DB_PASSWORD' }, + ], null, 2); + const r = expectGone(text, a, 'name-value-secret'); + expect(r.text).not.toContain(b); + expect(r.text).toContain('WEBSITE_RUN_FROM_PACKAGE'); + expect(r.text).toContain('"value": "1"'); + }); + + it('Kubernetes env entries', () => { + const a = fake.passphrase(); + expectGone(`env:\n - {"name": "ConnectionStrings__Redis", "value": "redis:6380"}\n - {"name": "JWT_SIGNING_KEY", "value": "${a}"}`, a); + }); + + it('`az keyvault secret show` value (any secret name, nested attributes)', () => { + const a = fake.passphrase(); + const text = JSON.stringify({ + attributes: { created: '2026-01-01T00:00:00+00:00', enabled: true, recoveryLevel: 'Recoverable' }, + contentType: null, + id: 'https://contoso-kv.vault.azure.net/secrets/StorageThing/0123456789abcdef0123456789abcdef', + name: 'StorageThing', + tags: {}, + value: a, + }, null, 2); + const r = expectGone(text, a, 'azure-keyvault-secret'); + expect(r.text).toContain('contoso-kv.vault.azure.net/secrets/StorageThing'); + }); + + it('web.config appSettings in either attribute order', () => { + const a = fake.passphrase(); + const b = fake.passphrase(); + expectGone(`\n \n \n \n`, a, 'xml-appsettings-secret'); + const r = redactor.redact(``, ctx); + expect(r.text).not.toContain(b); + expect(redactor.redact('', ctx).findings).toEqual([]); + }); + + it('publish profiles and XML secret attributes/elements', () => { + const a = fake.passphrase(); + const b = fake.passphrase(); + const r = expectGone(``, a, 'xml-secret-attribute'); + expect(r.text).toContain('userName="$contoso-app"'); + expect(r.text).toContain('https://contoso-app.azurewebsites.net'); + expectGone(`${b}`, b, 'xml-secret-element'); + }); + + it('a GUID in a password slot is redacted (legacy create-for-rbac), but GUID IDs elsewhere are kept', () => { + const pw = fake.guid(); + const tenant = fake.guid(); + const r = expectGone(JSON.stringify({ appId: fake.guid(), password: pw, tenant }), pw); + expect(r.text).toContain(tenant); + }); + + it('leaves placeholders, labels and non-secret neighbours alone', () => { + const untouched = [ + '{"ClientSecret": "#{ClientSecret}#", "ApiKey": "__API_KEY__", "Password": "$(DbPassword)", "Token": "${TOKEN}", "Secret": ""}', + 'ErrorMessage = "Invalid password or username";', + '{"tokenType": "Bearer", "secretName": "db-password", "maxTokens": "4096", "passwordPolicy": "strict"}', + 'options.Password = configuration["Db:Password"];', + '', + '{"name": "TokenEndpoint", "value": "https://login.microsoftonline.com/common/oauth2/v2.0/token"}', + '' + 'a1b2c3d4-0000-1111-2222-333344445555' + '', + 'PWD="/home/user1/src"', + ]; + for (const text of untouched) { + const r = redactor.redact(text, ctx); + expect(r.findings, text).toEqual([]); + } + }); +}); + +describe('redaction: structured JSON (field names as context, keys redacted)', () => { + const redactor = createRedactor(DEFAULT_REDACTION_CONFIG); + const fake = new FakeSecrets(4242); + + it('redacts a value whose own field name is secret-looking, with no shape match', () => { + const pw = fake.passphrase(); + const line = JSON.stringify({ type: 'user', toolUseResult: { structuredContent: { clientId: 'app', clientSecret: pw } } }); + const out = redactJsonlLine(line, redactor, ctx); + expect(out).not.toContain(pw); + expect(JSON.parse(out).toolUseResult.structuredContent).toEqual({ clientId: 'app', clientSecret: '[REDACTED:secret-field]' }); + }); + + it('redacts the value of a {name, value} object whose name is secret-looking', () => { + const pw = fake.passphrase(); + const line = JSON.stringify({ result: [{ name: 'Db__Password', value: pw }, { name: 'Region', value: 'eastus' }] }); + const out = JSON.parse(redactJsonlLine(line, redactor, ctx)); + expect(out.result).toEqual([{ name: 'Db__Password', value: '[REDACTED:secret-field]' }, { name: 'Region', value: 'eastus' }]); + }); + + it('redacts a structured Key Vault secret bundle value', () => { + const pw = fake.passphrase(); + const line = JSON.stringify({ result: { id: 'https://kv.vault.azure.net/secrets/Anything/1', value: pw, attributes: { enabled: true } } }); + expect(redactJsonlLine(line, redactor, ctx)).not.toContain(pw); + }); + + it('redacts secrets used as object keys', () => { + const token = fake.githubToken(); + const line = JSON.stringify({ cache: { [token]: { user: 'octocat' } } }); + const out = redactJsonlLine(line, redactor, ctx); + expect(out).not.toContain(token); + expect(JSON.parse(out).cache['[REDACTED:github-token]']).toEqual({ user: 'octocat' }); + }); + + it('leaves harness structure alone (thinking signatures, usage, token_count, apiKeySource)', () => { + const lines = [ + JSON.stringify({ type: 'assistant', message: { content: [{ type: 'thinking', thinking: 'hmm', signature: fake.base64Blob(200) }], usage: { input_tokens: 12, output_tokens: 3 } } }), + JSON.stringify({ type: 'event_msg', payload: { type: 'token_count', info: { total_token_usage: { input_tokens: 1 } } } }), + JSON.stringify({ type: 'system', subtype: 'init', apiKeySource: 'none', tokenizer: 'cl100k' }), + JSON.stringify({ secretName: 'db-password', passwordPolicy: 'strict', password: '${DB_PASSWORD}', token: '' }), + ]; + for (const line of lines) expect(redactJsonlLine(line, redactor, ctx), line).toBe(line); + }); + + it('redacts a secret field whole when a text rule already redacted only part of it', () => { + const word = fake.chars('abcdefghijklmnopqrstuvwxyz', 12); + const line = JSON.stringify({ password: `${word} ${fake.jwt()}` }); + const out = redactJsonlLine(line, redactor, ctx); + expect(out).not.toContain(word); + expect(JSON.parse(out).password).toBe('[REDACTED:secret-field]'); + }); + + it('is idempotent on structured redactions', () => { + const line = JSON.stringify({ a: { clientSecret: fake.passphrase() }, b: [{ name: 'API_KEY', value: fake.passphrase() }] }); + const once = redactJsonlLine(line, redactor, ctx); + expect(redactJsonlLine(once, redactor, ctx)).toBe(once); + }); + + it('can be turned off via secretFields.enabled=false', () => { + const r = createRedactor({ ...DEFAULT_REDACTION_CONFIG, secretFields: { ...DEFAULT_REDACTION_CONFIG.secretFields, enabled: false } }); + const pw = fake.passphrase(); + expect(redactJsonlLine(JSON.stringify({ x: { clientSecret: pw } }), r, ctx)).toContain(pw); + }); +}); diff --git a/test/verify.test.ts b/test/verify.test.ts index ae70da81..dac46ea1 100644 --- a/test/verify.test.ts +++ b/test/verify.test.ts @@ -281,6 +281,37 @@ describe('repairIndex', () => { dbAfter.close(); }); + it('fails closed on a corrupt rules file in strict mode, before changing the index', async () => { + const db = initDatabase(); + insertExchange(db, { + id: 'orphan-strict-1', + project: 'deleted-project', + timestamp: '2024-01-01T00:00:00Z', + userMessage: 'Deleted', + assistantMessage: 'Still indexed', + archivePath: path.join(archiveDir, 'deleted-project', 'deleted.jsonl'), + lineStart: 1, + lineEnd: 2 + }, new Array(384).fill(0.1)); + db.close(); + + const rulesPath = path.join(testDir, 'redaction-rules.json'); + fs.writeFileSync(rulesPath, '{ not json'); + process.env.EPISODIC_MEMORY_REDACTION_RULES = rulesPath; + try { + const issues = await verifyIndex(); + expect(issues.orphaned.length).toBe(1); + await expect(repairIndex(issues)).rejects.toThrow(/redaction/i); + } finally { + delete process.env.EPISODIC_MEMORY_REDACTION_RULES; + } + + const dbAfter = initDatabase(); + const row = dbAfter.prepare(`SELECT COUNT(*) as count FROM exchanges WHERE id = ?`).get('orphan-strict-1') as { count: number }; + expect(row.count).toBe(1); + dbAfter.close(); + }); + it('re-indexes outdated files during repair', { timeout: 30000 }, async () => { // Create conversation file with summary const projectArchive = path.join(archiveDir, 'test-project');