[stack 1/8] feat: make prompt cache behavior measurable - #304
Merged
Merged
Conversation
…yer 01-prompt-cache
Contributor
|
Navigate logical layers of code changes, visualize relationships, and explore their blast radius. Note Currently processing new changes in this PR. This may take a few minutes, please wait... ⚙️ Run configuration
📒 Files selected for processing (12)
✨ Finishing Touches📝 Generate docstrings
🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
This was referenced Oct 4, 2026
OnlineChef (ChefGroep)
marked this pull request as ready for review
October 5, 2026 04:18
|
Capy couldn't review this pull request because OnlineChef's workspace is out of credits, add credits or enable auto-reload to resume automatic reviews. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Stacked re-landing of #296 (
fix/prompt-cache-observability-20261003). Position 1/8; base:main. Merge bottom-up; each layer is a real merge of the original PR head, so the diff shows only that PR plus any conflict resolution (noted in the merge commit message).Stack
Original description:
Summary
Add an observability-first prompt-cache layer for OpenAI Responses routing so cache behavior can be measured before changing provider/model cache policy.
This PR does not inject cache options, enable paid cache writes, deploy anything, add Redis, or change request semantics. It records privacy-safe structural diagnostics from the exact final outbound Responses body and persists the final adapter alongside usage.
Why
The existing usage log already records provider-reported cache reads/writes, but historical rows do not record the adapter or enough outbound request shape to explain a miss.
The live audit found real historical write-heavy behavior on older provider paths, but applying cache policy from that history to the current Responses route would be speculative.
A deterministic read-only run against the retained live usage log, anchored at
2026-10-03T21:42:00Z, measured:openai/gpt-5.6-terra: 93.2% cache-read ratioopenai/gpt-5.6-luna: 5.9% cache-read ratio in the retained 7d cohortHistorical rows still report
adapter=unknownbecause they predate this instrumentation. That is intentional evidence separation, not backfilled inference.Changes
__proto__as ordinary datascripts/analyze-prompt-cache-usage.tswith strict CLI validation,--nowreproducibility, typed cohort/cache-shape dimensions, and the hard proof boundarystatus=200 && usageStatus=reportedusage.jsonlNo prompt/user/tool text, cache-key value, or response content is persisted in the new diagnostics.
Storage boundary / Redis
Redis is deliberately not added here.
The repository architecture already defines the proxy response cache as local
Map + LRU + TTLwith optional file persistence and states that Redis is an opt-in for multi-process deployments. The current prompt-cache lane is observability over provider-side prompt caching, not a proxy-level shared cache.The current durable source of truth remains append-only
usage.jsonl./v1/responsesalso remains excluded from the proxy body cache because Redis would not solve the missingprevious_response_id/ provider-continuation reconstruction problem.Redis/Upstash should only become a follow-up when there is measured need for cross-process or cross-host shared hot state, distributed atomic counters/locks, or a shared response cache. If local analytics outgrow JSONL scanning first, SQLite is the lower-complexity next step; central Postgres/Neon only makes sense once cross-host aggregation is required.
Pstack design boundary
The implementation intentionally stays at the adapter/persistence boundaries:
final outbound body -> typed PromptCacheRequestObservation -> request log -> usage.jsonl -> analyzerA generic cache normalizer in
responses/core.tswas rejected because provider wire semantics belong to the adapter. Public OpenAI API semantics and the private ChatGPT Codex backend remain separate concerns.Verification
Exact head
54e565df91a0721106da3f36476daa14465a3a4bhas focused local verification. Its parente468b0eae7a25b672137f81592d46d0244e7e183also passed the full repository suite:bun run typecheck✅e468b0ea: 7020 pass, 11 explicit skips, 0 fail across 524 files ✅bun run privacy:scan✅git diff --check✅--nowanchor ✅The 11 skips are existing environment/capability skips reported by the repository suite, not failures introduced by this PR.
Follow-up after this lands
Use normal traffic on the current Responses routes to compare cache reads/writes by:
adapter + provider + model + surface + promptCache shapeOnly after that evidence exists should a second PR consider model/provider-specific
prompt_cache_options, retention, or explicit breakpoints.Summary by CodeRabbit