Skip to content

[stack 1/8] feat: make prompt cache behavior measurable - #304

Merged
OnlineChef (ChefGroep) merged 5 commits into
mainfrom
stack/20261005-01-prompt-cache
Oct 5, 2026
Merged

OnlineChef (ChefGroep) merged 5 commits into
mainfrom
stack/20261005-01-prompt-cache

Conversation

@ChefGroep

@ChefGroep OnlineChef (ChefGroep) commented Oct 4, 2026 •

Copy link
Copy Markdown
Contributor

Stacked re-landing of #296 (fix/prompt-cache-observability-20261003). Position 1/8; base: main. Merge bottom-up; each layer is a real merge of the original PR head, so the diff shows only that PR plus any conflict resolution (noted in the merge commit message).

Stack

# Layer Original
1/8 #304 feat: make prompt cache behavior measurable #296
2/8 #305 feat: expand providers and make native model discovery account-aware #297
3/8 #306 feat(gui): start auth-inspired OpenCodex redesign #299
4/8 #307 feat(trace): opt-in trace store linked to usage.jsonl #298
5/8 #308 feat(trace): add safe local trace reader #300
6/8 #309 feat(trace): cover compact and response-cache hits #301
7/8 #310 feat(trace): cover Responses WebSocket turns #302
8/8 #311 feat(trace): cover live call-create HTTP #303

Original description:

Summary

Add an observability-first prompt-cache layer for OpenAI Responses routing so cache behavior can be measured before changing provider/model cache policy.

This PR does not inject cache options, enable paid cache writes, deploy anything, add Redis, or change request semantics. It records privacy-safe structural diagnostics from the exact final outbound Responses body and persists the final adapter alongside usage.

Why

The existing usage log already records provider-reported cache reads/writes, but historical rows do not record the adapter or enough outbound request shape to explain a miss.

The live audit found real historical write-heavy behavior on older provider paths, but applying cache policy from that history to the current Responses route would be speculative.

A deterministic read-only run against the retained live usage log, anchored at 2026-10-03T21:42:00Z, measured:

  • 7d window: 215,536,617 successful provider-reported input tokens
  • 184,376,962 cache-read tokens (85.5%)
  • 690,312 cache-write tokens (0.3%)
  • openai/gpt-5.6-terra: 93.2% cache-read ratio
  • openai/gpt-5.6-luna: 5.9% cache-read ratio in the retained 7d cohort

Historical rows still report adapter=unknown because they predate this instrumentation. That is intentional evidence separation, not backfilled inference.

Changes

  • capture prompt-cache diagnostics from the final sanitized outbound Responses body
  • persist final adapter identity on usage rows
  • persist only structural cache metadata:
    • cache mode / TTL / retention flags
    • key presence, never the key value
    • breakpoint count
    • tool count + stable tool-schema fingerprint
    • stable developer/system prefix fingerprint
    • response-format fingerprint / verbosity
    • continuation presence
  • use a stable JSON serializer for fingerprints that preserves array/tool order and treats keys such as __proto__ as ordinary data
  • keep persistence validation at the JSONL boundary and trust typed adapter metadata internally
  • route adapter-derived request diagnostics through one lifecycle seam
  • clear cache metadata when retries/failovers switch to an adapter with no cache observation
  • add scripts/analyze-prompt-cache-usage.ts with strict CLI validation, --now reproducibility, typed cohort/cache-shape dimensions, and the hard proof boundary status=200 && usageStatus=reported
  • reuse the canonical persisted prompt-cache parser in the analyzer rather than duplicating shape parsing
  • add an end-to-end local proof from final outbound adapter body -> request log -> persisted usage.jsonl
  • assert persisted bytes contain no raw cache key, fixed prefix text, or tool name
  • split the new observation/analyzer stages into small pure helpers after CodeFactor identified avoidable method complexity

No prompt/user/tool text, cache-key value, or response content is persisted in the new diagnostics.

Storage boundary / Redis

Redis is deliberately not added here.

The repository architecture already defines the proxy response cache as local Map + LRU + TTL with optional file persistence and states that Redis is an opt-in for multi-process deployments. The current prompt-cache lane is observability over provider-side prompt caching, not a proxy-level shared cache.

The current durable source of truth remains append-only usage.jsonl. /v1/responses also remains excluded from the proxy body cache because Redis would not solve the missing previous_response_id / provider-continuation reconstruction problem.

Redis/Upstash should only become a follow-up when there is measured need for cross-process or cross-host shared hot state, distributed atomic counters/locks, or a shared response cache. If local analytics outgrow JSONL scanning first, SQLite is the lower-complexity next step; central Postgres/Neon only makes sense once cross-host aggregation is required.

Pstack design boundary

The implementation intentionally stays at the adapter/persistence boundaries:

final outbound body -> typed PromptCacheRequestObservation -> request log -> usage.jsonl -> analyzer

A generic cache normalizer in responses/core.ts was rejected because provider wire semantics belong to the adapter. Public OpenAI API semantics and the private ChatGPT Codex backend remain separate concerns.

Verification

Exact head 54e565df91a0721106da3f36476daa14465a3a4b has focused local verification. Its parent e468b0eae7a25b672137f81592d46d0244e7e183 also passed the full repository suite:

  • bun run typecheck ✅
  • focused cache/analyzer/request-log tests on exact head: 134 pass, 0 fail ✅
  • full repository suite on parent e468b0ea: 7020 pass, 11 explicit skips, 0 fail across 524 files ✅
  • 34,881 assertions ✅
  • bun run privacy:scan ✅
  • git diff --check ✅
  • deterministic analyzer reproduced the retained live-log measurement with a fixed --now anchor ✅
  • no paid model/API request was made for this work ✅

The 11 skips are existing environment/capability skips reported by the repository suite, not failures introduced by this PR.

Follow-up after this lands

Use normal traffic on the current Responses routes to compare cache reads/writes by:

adapter + provider + model + surface + promptCache shape

Only after that evidence exists should a second PR consider model/provider-specific prompt_cache_options, retention, or explicit breakpoints.

Summary by CodeRabbit

  • New Features
    • Added prompt-cache insights to usage records, including cache settings, request shape, adapter, and privacy-preserving fingerprints.
    • Added an analyzer that summarizes cache usage by time range and groups results by adapter, provider, model, and request shape. Supports JSON and readable text output.

@coderabbitai

coderabbitai Bot commented Oct 4, 2026 •

Copy link
Copy Markdown
Contributor

Review in Change Stack →

Navigate logical layers of code changes, visualize relationships, and explore their blast radius.

Note

Currently processing new changes in this PR. This may take a few minutes, please wait...

⚙️ Run configuration
  • Configuration used: Repository: GroepOnline/opencodex/.coderabbit.yaml
  • Review profile: ASSERTIVE
  • Plan: Advanced
  • Run ID: d57dc17f-2943-4717-ab53-538cdada5ccf
📥 Commits

Reviewing files that changed from the base of the PR and between 30221c5 and b90389a.

📒 Files selected for processing (12)
  • scripts/analyze-prompt-cache-usage.ts
  • src/adapters/base.ts
  • src/adapters/openai-responses.ts
  • src/prompt-cache/observability.ts
  • src/server/request-log.ts
  • src/server/responses/core.ts
  • src/usage/log.ts
  • tests/openai-responses-passthrough.test.ts
  • tests/prompt-cache-analyzer.test.ts
  • tests/prompt-cache-observability.test.ts
  • tests/request-log.test.ts
  • tests/usage-log.test.ts
 _______________________________________________________________________________
< GAST: GPU-Assisted Security Testing. Because SAST is too static for my taste. >
 -------------------------------------------------------------------------------
  \
   \   \
        \ /\
        ( )
      .( o ).
✨ Finishing Touches
📝 Generate docstrings
  • Commit to this branch
  • Create a new PR
🧪 Generate unit tests (beta)
  • Commit to this branch
  • Create a new PR
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@capy-ai

capy-ai Bot commented Oct 5, 2026

Copy link
Copy Markdown

Capy couldn't review this pull request because OnlineChef's workspace is out of credits, add credits or enable auto-reload to resume automatic reviews.

Open in Capy

@ChefGroep
OnlineChef (ChefGroep) deleted the stack/20261005-01-prompt-cache branch October 5, 2026 04:18
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant