Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
15 commits
Select commit Hold shift + click to select a range
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
17 changes: 17 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -5,6 +5,23 @@ All notable changes to this project will be documented in this file.
The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.0.0/),
and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).

## [Unreleased]

### Security

- Secrets no longer leak into your conversation memory. Azure client secrets, storage keys, SAS signatures, connection-string passwords, private keys, JWTs, GitHub, Anthropic, OpenAI, and AWS keys, and `password: …` style assignments are now replaced with tokens like `[REDACTED:azure-storage-key]` before a conversation is archived, indexed, embedded, or summarized. Before this, every pasted secret was copied into three new places on disk and could be sent to the summarization model. The rest of the conversation stays searchable, and so do git SHAs and Azure tenant, client, and object IDs. Redaction is on by default and fails closed: if the rules can't load, sync won't run rather than store secrets. .NET and Azure shapes where only the setting's name gives the secret away are covered too, even when the value has no recognizable format: `appsettings.json` fields, `web.config` `<add key=… value=…/>`, publish-profile passwords, `az webapp config appsettings list` and `az keyvault secret show` output, C# `ClientSecret = "…"` assignments, and structured MCP results.
- With redaction on, summaries no longer resume Claude Code sessions or fork Codex threads. Both of those let the model read the original, unredacted transcript. Summaries now come from the redacted text. Codex-only users without Claude configured can set `EPISODIC_MEMORY_SKIP_SUMMARIES=1`.

### Added

- `episodic-memory redact --rewrite` cleans data you indexed before upgrading. It redacts the archive and the search index in place, re-embeds only the messages that changed, and deletes summaries built from unredacted text so they regenerate. Use `--dry-run` to preview (it writes nothing, not even a schema migration), and add `--report` to list each value it would redact, by location, rule and shape, without printing the value.
- Custom redaction rules via `~/.config/superpowers/redaction-rules.json`, which extends the bundled defaults. Try rules with `episodic-memory redact --stdin`. See `docs/REDACTION.md`.
- New settings: `EPISODIC_MEMORY_REDACTION` (`on`/`off`), `EPISODIC_MEMORY_REDACTION_RULES`, and `EPISODIC_MEMORY_REDACTION_STRICT`.

### Changed

- `episodic-memory index` now parses the archived copy instead of the source transcript, and refreshes that copy when the source has grown. This is the same behavior `sync` already had.

## [1.6.0] - 2026-09-08

Adds a fifth conversation source, an off switch for automatic syncing, and two fixes for real-world resource problems.
Expand Down
34 changes: 34 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -327,6 +327,18 @@ Add to `.claude/hooks/session-end`:
episodic-memory sync
```

### `episodic-memory redact`

```bash
episodic-memory redact --rewrite --dry-run # what would change
episodic-memory redact --rewrite --dry-run --report # each hit: location, rule, value shape
episodic-memory redact --rewrite # redact the existing archive + index in place
episodic-memory redact --stdin < file.jsonl # try the rules on some text
episodic-memory redact --print-default-rules
```

See [docs/REDACTION.md](docs/REDACTION.md).

### `episodic-memory stats`

Display index statistics including conversation counts, date ranges, and project breakdown.
Expand Down Expand Up @@ -415,6 +427,28 @@ open output.html
4. **Index** - Stores in SQLite with sqlite-vec for fast similarity search
5. **Search** - Semantic search using vector similarity or exact text matching

## Secret Redaction

Secrets in your conversations (Azure client secrets, storage keys, SAS
signatures, connection-string passwords, private keys, JWTs, provider API keys,
`password: …` assignments) are replaced with typed tokens such as
`[REDACTED:azure-storage-key]` **before** anything is archived, indexed,
embedded, or sent to the summarizer. Git SHAs and GUIDs are never redacted, so
they stay searchable. The tokens are searchable too.

Redaction is on by default and fails closed: if the rules can't be loaded, sync
won't run. After upgrading, run `episodic-memory redact --rewrite` once to clean
data you indexed before.

| Variable | Default | Meaning |
|---|---|---|
| `EPISODIC_MEMORY_REDACTION` | `on` | `off` disables redaction |
| `EPISODIC_MEMORY_REDACTION_RULES` | `<config dir>/redaction-rules.json` | Custom rules file (extends the defaults) |
| `EPISODIC_MEMORY_REDACTION_STRICT` | `1` | `0` continues unredacted (with a warning) when rules fail to load |

See [docs/REDACTION.md](docs/REDACTION.md) for the rule list, the custom-rules
format, and what is and isn't covered.

## Excluding Conversations

Conversations containing this marker anywhere in their content will be archived but not indexed:
Expand Down
5 changes: 5 additions & 0 deletions cli/episodic-memory.js
Original file line number Diff line number Diff line change
Expand Up @@ -44,6 +44,7 @@ COMMANDS:
stats Show index statistics
doctor Diagnose Claude Code or Codex integration issues
import-cursor-history Export legacy Cursor conversations from state.vscdb for indexing
redact Re-run secret redaction over the archive and index (--rewrite)

Run 'episodic-memory <command> --help' for command-specific help.

Expand Down Expand Up @@ -94,6 +95,10 @@ async function main() {
await runScript(join(distDir, 'cursor-import-cli.js'), args);
break;

case 'redact':
await runScript(join(distDir, 'redact-cli.js'), args);
break;

case '--help':
case '-h':
case undefined:
Expand Down
3 changes: 3 additions & 0 deletions dist/cursor-legacy.d.ts
Original file line number Diff line number Diff line change
@@ -1,3 +1,4 @@
import { type Redactor } from './redaction.js';
export declare function getDefaultCursorVscdbPath(): string | undefined;
/**
* Collect composer IDs that already have live agent transcripts under
Expand All @@ -14,6 +15,8 @@ export interface CursorLegacyImportOptions {
force?: boolean;
/** Report what would be exported without writing files. */
dryRun?: boolean;
/** `undefined` loads from the environment (strict mode throws on bad rules); `null` disables. */
redactor?: Redactor | null;
}
export interface CursorLegacyImportResult {
exported: number;
Expand Down
9 changes: 8 additions & 1 deletion dist/cursor-legacy.js
Original file line number Diff line number Diff line change
Expand Up @@ -3,6 +3,7 @@ import os from 'os';
import path from 'path';
import Database from 'better-sqlite3';
import { detectCursorCwd } from './parser.js';
import { loadRedactor, redactJsonlLine } from './redaction.js';
/**
* Import legacy Cursor conversations from Cursor's global SQLite store
* (state.vscdb) into JSONL files compatible with the Cursor transcript parser.
Expand Down Expand Up @@ -111,6 +112,7 @@ export function importCursorLegacy(options) {
skippedEmpty: 0,
errors: [],
};
const redactor = options.redactor === undefined ? loadRedactor() : options.redactor;
const liveIds = options.liveTranscriptIds ?? new Set();
const db = new Database(options.dbPath, { readonly: true, fileMustExist: true });
try {
Expand Down Expand Up @@ -200,9 +202,14 @@ export function importCursorLegacy(options) {
if (!options.dryRun) {
// Re-serialize with cwd now that it's known (it's derived from the
// whole conversation's tool calls).
const finalLines = cwd
const withCwd = cwd
? lines.map(line => JSON.stringify({ ...JSON.parse(line), cwd }))
: lines;
// The export dir is a plugin-owned plaintext copy, so redact it at
// write time like the archive (docs/REDACTION.md).
const finalLines = redactor
? withCwd.map(line => redactJsonlLine(line, redactor, { source: 'cursor-legacy', path: outFile }))
: withCwd;
fs.mkdirSync(path.dirname(outFile), { recursive: true });
fs.writeFileSync(outFile, finalLines.join('\n') + '\n', 'utf-8');
// Stamp the conversation's end time so mtime-based fallbacks and
Expand Down
6 changes: 6 additions & 0 deletions dist/db.d.ts
Original file line number Diff line number Diff line change
Expand Up @@ -13,6 +13,12 @@ export declare function migrateSchema(db: Database.Database): void;
* 3. Recreates the table with ON DELETE CASCADE and copies surviving rows.
*/
export declare function migrateToolCallsCascade(db: Database.Database): void;
/**
* Open the existing index read-only, without creating it, migrating it, or
* changing its journal mode. For dry runs that must not write. Returns null
* when there is no index yet.
*/
export declare function openDatabaseReadOnly(): Database.Database | null;
export declare function initDatabase(): Database.Database;
export declare function insertExchange(db: Database.Database, exchange: ConversationExchange, embedding: number[], toolNames?: string[]): void;
export declare function getAllExchanges(db: Database.Database): Array<{
Expand Down
11 changes: 11 additions & 0 deletions dist/db.js
Original file line number Diff line number Diff line change
Expand Up @@ -90,6 +90,17 @@ export function migrateToolCallsCascade(db) {
db.pragma('foreign_keys = ON');
console.log(' tool_calls migration complete.');
}
/**
* Open the existing index read-only, without creating it, migrating it, or
* changing its journal mode. For dry runs that must not write. Returns null
* when there is no index yet.
*/
export function openDatabaseReadOnly() {
const dbPath = getDbPath();
if (!fs.existsSync(dbPath))
return null;
return new Database(dbPath, { readonly: true, fileMustExist: true });
}
export function initDatabase() {
const dbPath = getDbPath();
// Ensure directory exists
Expand Down
55 changes: 35 additions & 20 deletions dist/indexer.js
Original file line number Diff line number Diff line change
Expand Up @@ -7,6 +7,8 @@ import { summarizeConversation } from './summarizer.js';
import { getArchiveDir, getExcludedProjects, getConversationSourceDirs, findJsonlFiles, statIfExists } from './paths.js';
import { formatErrorSentinel, shouldQueueForSummary } from './summary-sentinel.js';
import { getMaxMessageBytes, isOversizeExchange } from './message-size.js';
import { FindingsTally, formatFindings, loadRedactor } from './redaction.js';
import { copyIfNewer } from './sync.js';
// Set max output tokens for Claude SDK (used by summarizer)
process.env.CLAUDE_CODE_MAX_OUTPUT_TOKENS = '20000';
// Increase max listeners for concurrent API calls
Expand All @@ -25,7 +27,18 @@ async function processBatch(items, processor, concurrency) {
function sessionIdForSummary(exchanges) {
return exchanges.find(exchange => exchange.sessionId)?.sessionId;
}
// Resume/fork would summarize the unredacted source transcript; see sync.ts.
function summarizeOptions(redactor) {
return { allowResume: redactor === null };
}
function logRedactions(tally) {
if (tally.total > 0)
console.log(` Redaction: ${formatFindings(tally.toArray())}`);
}
export async function indexConversations(limitToProject, maxConversations, concurrency = 1, noSummaries = false) {
// Load before touching the archive: strict mode fails closed here.
const redactor = loadRedactor();
const tally = new FindingsTally();
console.log('Initializing database...');
const db = initDatabase();
console.log('Loading embedding model...');
Expand Down Expand Up @@ -73,14 +86,13 @@ export async function indexConversations(limitToProject, maxConversations, concu
// Source transcripts can vanish mid-run (Claude Code cleanup). Skip loudly.
let exchanges;
try {
// Copy to archive (ensure parent dirs exist for subagent files)
if (!fs.existsSync(archivePath)) {
fs.mkdirSync(path.dirname(archivePath), { recursive: true });
fs.copyFileSync(sourcePath, archivePath);
// Copy (redacted) to the archive, then parse the archive, so the index
// and summaries only ever see redacted text.
if (copyIfNewer(sourcePath, archivePath, redactor, tally)) {
console.log(` Archived: ${file}`);
}
// Parse conversation
exchanges = await parseConversation(sourcePath, project, archivePath);
exchanges = await parseConversation(archivePath, project, archivePath);
}
catch (error) {
console.log(` Skipped ${file} (read failed: ${error instanceof Error ? error.message : error})`);
Expand All @@ -105,7 +117,7 @@ export async function indexConversations(limitToProject, maxConversations, concu
console.log(` Generating ${needsSummary.length} summaries (concurrency: ${concurrency})...`);
await processBatch(needsSummary, async (conv) => {
try {
const summary = await summarizeConversation(conv.exchanges, sessionIdForSummary(conv.exchanges));
const summary = await summarizeConversation(conv.exchanges, sessionIdForSummary(conv.exchanges), summarizeOptions(redactor));
fs.writeFileSync(conv.summaryPath, summary, 'utf-8');
const wordCount = summary.split(/\s+/).length;
console.log(` ✓ ${conv.file}: ${wordCount} words`);
Expand Down Expand Up @@ -145,6 +157,7 @@ export async function indexConversations(limitToProject, maxConversations, concu
// Check if we hit the limit
if (maxConversations && conversationsProcessed >= maxConversations) {
console.log(`\nReached limit of ${maxConversations} conversations`);
logRedactions(tally);
db.close();
console.log(`✅ Indexing complete! Conversations: ${conversationsProcessed}, Exchanges: ${totalExchanges}`);
return;
Expand All @@ -155,11 +168,14 @@ export async function indexConversations(limitToProject, maxConversations, concu
if (oversizeSkipped > 0) {
console.log(` Skipped ${oversizeSkipped} oversize exchange(s) (> ${maxMessageBytes} bytes; set EPISODIC_MEMORY_MAX_MESSAGE_BYTES to change) — likely embedded-transcript payloads (#139)`);
}
logRedactions(tally);
db.close();
console.log(`\n✅ Indexing complete! Conversations: ${conversationsProcessed}, Exchanges: ${totalExchanges}`);
}
export async function indexSession(sessionId, concurrency = 1, noSummaries = false) {
console.log(`Indexing session: ${sessionId}`);
const redactor = loadRedactor();
const tally = new FindingsTally();
// Find the conversation file for this session
const sourceDirs = getConversationSourceDirs();
const ARCHIVE_DIR = getArchiveDir();
Expand Down Expand Up @@ -187,11 +203,8 @@ export async function indexSession(sessionId, concurrency = 1, noSummaries = fal
// Archive + parse — source may vanish mid-run (Claude Code cleanup).
let exchanges;
try {
if (!fs.existsSync(archivePath)) {
fs.mkdirSync(path.dirname(archivePath), { recursive: true });
fs.copyFileSync(sourcePath, archivePath);
}
exchanges = await parseConversation(sourcePath, project, archivePath);
copyIfNewer(sourcePath, archivePath, redactor, tally);
exchanges = await parseConversation(archivePath, project, archivePath);
}
catch (error) {
console.log(`Skipped ${file} (read failed: ${error instanceof Error ? error.message : error})`);
Expand All @@ -204,7 +217,7 @@ export async function indexSession(sessionId, concurrency = 1, noSummaries = fal
if (!noSummaries && shouldQueueForSummary(summaryPath)) {
fs.mkdirSync(path.dirname(summaryPath), { recursive: true });
try {
const summary = await summarizeConversation(exchanges, sessionIdForSummary(exchanges));
const summary = await summarizeConversation(exchanges, sessionIdForSummary(exchanges), summarizeOptions(redactor));
fs.writeFileSync(summaryPath, summary, 'utf-8');
console.log(`Summary: ${summary.split(/\s+/).length} words`);
}
Expand Down Expand Up @@ -235,6 +248,7 @@ export async function indexSession(sessionId, concurrency = 1, noSummaries = fal
if (oversizeSkipped > 0) {
console.log(` Skipped ${oversizeSkipped} oversize exchange(s) (> ${maxMessageBytes} bytes; set EPISODIC_MEMORY_MAX_MESSAGE_BYTES to change) — likely embedded-transcript payloads (#139)`);
}
logRedactions(tally);
console.log(`✅ Indexed session ${sessionId}: ${exchanges.length} exchanges`);
}
db.close();
Expand All @@ -254,6 +268,8 @@ export async function indexUnprocessed(concurrency = 1, noSummaries = false) {
console.log(`Concurrency: ${concurrency}`);
if (noSummaries)
console.log('⚠️ Running in no-summaries mode (skipping AI summaries)');
const redactor = loadRedactor();
const tally = new FindingsTally();
const db = initDatabase();
await initEmbeddings();
const sourceDirs = getConversationSourceDirs();
Expand All @@ -280,15 +296,13 @@ export async function indexUnprocessed(concurrency = 1, noSummaries = false) {
// Transcript JSONLs are append-only, so MAX(line_end) tells us where to resume.
const hw = db.prepare('SELECT COALESCE(MAX(line_end), 0) as maxLine FROM exchanges WHERE archive_path = ?').get(archivePath);
const maxIndexedLine = hw.maxLine;
// Ensure parent dirs exist for subagent files
try {
fs.mkdirSync(path.dirname(archivePath), { recursive: true });
// Refresh the archive when the source may have grown beyond what we've seen.
if (!fs.existsSync(archivePath) || maxIndexedLine > 0) {
fs.copyFileSync(sourcePath, archivePath);
}
// Refresh the (redacted) archive, then parse the archive so the index
// only sees redacted text. Force the refresh once the file is indexed:
// an append inside the mtime granularity would otherwise be skipped.
copyIfNewer(sourcePath, archivePath, redactor, tally, maxIndexedLine > 0);
// Parse and filter to exchanges past the high-water mark
const exchanges = await parseConversation(sourcePath, project, archivePath);
const exchanges = await parseConversation(archivePath, project, archivePath);
const newExchanges = maxIndexedLine > 0
? exchanges.filter(e => e.lineStart > maxIndexedLine)
: exchanges;
Expand All @@ -303,6 +317,7 @@ export async function indexUnprocessed(concurrency = 1, noSummaries = false) {
}
}
} // end sourceDir loop
logRedactions(tally);
if (unprocessed.length === 0) {
console.log('✅ All conversations are already processed!');
db.close();
Expand All @@ -316,7 +331,7 @@ export async function indexUnprocessed(concurrency = 1, noSummaries = false) {
console.log(`Generating ${needsSummary.length} summaries (concurrency: ${concurrency})...\n`);
await processBatch(needsSummary, async (conv) => {
try {
const summary = await summarizeConversation(conv.exchanges, sessionIdForSummary(conv.exchanges));
const summary = await summarizeConversation(conv.exchanges, sessionIdForSummary(conv.exchanges), summarizeOptions(redactor));
fs.writeFileSync(conv.summaryPath, summary, 'utf-8');
const wordCount = summary.split(/\s+/).length;
console.log(` ✓ ${conv.project}/${conv.file}: ${wordCount} words`);
Expand Down
Loading