Summary
Explore adding a prompt-cache simulator to OmniToken.
The goal is to let users provide a batch of rendered prompts or message traces and estimate how much provider prompt caching would help or fail.
Why
Prompt caching can significantly reduce LLM input cost and latency, but cache hits depend on exact stable prefixes. Many apps accidentally put dynamic data such as timestamps, request IDs, user metadata, or changing JSON before long stable instructions.
Proposed UX
omni cache-sim -provider openai -model gpt-4o -file prompts.jsonl
Questions
- What input formats should we support first?
- Should MVP target OpenAI-style automatic prefix caching first?
- How should we report recommendations?
- Should this live in core package, CLI only, or both?
Related Ideas
- repeated prefix detection
- dynamic-field detection
- cache hit simulation
- cost savings estimation
Summary
Explore adding a prompt-cache simulator to OmniToken.
The goal is to let users provide a batch of rendered prompts or message traces and estimate how much provider prompt caching would help or fail.
Why
Prompt caching can significantly reduce LLM input cost and latency, but cache hits depend on exact stable prefixes. Many apps accidentally put dynamic data such as timestamps, request IDs, user metadata, or changing JSON before long stable instructions.
Proposed UX
Questions
Related Ideas