Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
16 changes: 9 additions & 7 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -19,7 +19,7 @@

<p align="center"><sub>If OmniToken is useful in your Go project, a star helps others discover it.</sub></p>

OmniToken is built for Go services that need fast local token accounting for prompt sizing, context-window planning, tokenizer experiments, and cache-boundary analysis without CGO, Rust, or Python runtime dependencies in the root module.
OmniToken is built for Go services that need fast local token accounting for prompt sizing, context-window planning, tokenizer experiments, and cacheflow analysis without CGO, Rust, or Python runtime dependencies in the root module.

The root module supports Go 1.23+. Some optional comparison tooling in this repository uses dependencies that require newer Go versions.

Expand All @@ -30,7 +30,7 @@ The root module supports Go 1.23+. Some optional comparison tooling in this repo
- Local `Encode`, `EncodeOrdinary`, `CountTokens`, and `Decode` APIs.
- Zero-allocation `CountTokens` hot path for supported OpenAI BPE workloads.
- Custom WordPiece and SentencePiece-style vocabularies.
- Prompt-cache alignment planner for token block-boundary analysis.
- `cacheflow` package for prompt-cache boundary and trace analysis.
- Optional adapter modules for Gemini, Llama 3, Mistral, Hugging Face `tokenizer.json`, OSS SentencePiece models, and Anthropic message token counting.

## Benchmarks
Expand Down Expand Up @@ -115,21 +115,23 @@ Use `SpecialTokenID` or `SpecialTokens` on `*omnitoken.Engine` when constructing
| Hugging Face WordPiece adapter | Optional module |
| Anthropic message counter | Optional module |

## Cache Alignment
## Cacheflow

```go
import "github.com/ron2111/omnitoken/cacheflow"

engine, err := omnitoken.ForModel("gpt-4o")
if err != nil {
panic(err)
}

report := omnitoken.NewCacheAligner(engine).AlignPromptToProfile(
report := cacheflow.NewAligner(engine).AlignPromptToProfile(
systemPrompt,
omnitoken.CacheProfileOpenAI,
cacheflow.ProfileOpenAI,
)
```

Cache alignment is informational: OmniToken does not edit prompts automatically. See [cache alignment](./docs/cache.md).
Cacheflow is informational: OmniToken does not edit prompts automatically or claim provider billing parity. See [cacheflow](./cacheflow/README.md).

## Custom Models

Expand All @@ -148,7 +150,7 @@ err = omnitoken.RegisterModelPrefix("my-model-", "my_wordpiece")

- [Architecture](./docs/architecture.md)
- [Benchmarks and correctness](./docs/benchmarks.md)
- [Cache alignment](./docs/cache.md)
- [Cacheflow](./cacheflow/README.md)
- [CLI](./docs/cli.md)
- [Adapters](./adapters/README.md)

Expand Down
118 changes: 0 additions & 118 deletions cache.go

This file was deleted.

81 changes: 81 additions & 0 deletions cacheflow/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,81 @@
# Cacheflow

`cacheflow` is OmniToken's dependency-free prompt-cache analysis package.

It is built for Go teams that want to understand whether rendered prompts have stable token prefixes before sending them to OpenAI, Anthropic, Gemini, or another provider. It does not claim billing parity and it does not call provider APIs.

## What It Does

- Counts prompt tokens with any `omnitoken.ModelEngine`.
- Calculates cache-boundary alignment for one prompt.
- Reads JSONL prompt traces.
- Finds common token prefixes across repeated prompts.
- Estimates reusable prefix tokens under a local cache profile.
- Emits best-effort cache-breaker hints for timestamps, UUIDs, request IDs, and dynamic metadata.

## What It Does Not Do

- It does not edit prompts automatically.
- It does not guarantee provider billing behavior.
- It does not require network calls or credentials.
- It does not add dependencies to the root module.

## Align One Prompt

```go
engine, err := omnitoken.ForModel("gpt-4o")
if err != nil {
panic(err)
}

report := cacheflow.NewAligner(engine).AlignPromptToProfile(systemPrompt, cacheflow.ProfileOpenAI)
fmt.Println(report.CurrentTokens, report.PaddingNeeded)
```

## Simulate A Trace

```go
items := []cacheflow.TraceItem{
{ID: "1", Prompt: stablePrefix + "user question one"},
{ID: "2", Prompt: stablePrefix + "user question two"},
}

report := cacheflow.Simulate(engine, items, cacheflow.SimulationOptions{
Profile: cacheflow.ProfileOpenAI,
DetectBreakers: true,
})
fmt.Println(report.ReusablePrefixTokens)
```

## JSONL Format

Raw rendered prompts:

```json
{"id":"1","model":"gpt-4o","prompt":"..."}
{"id":"2","model":"gpt-4o","prompt":"..."}
```

Structured prompt parts:

```json
{"id":"1","model":"gpt-4o","parts":[{"name":"system","stable":true,"text":"..."},{"name":"user","stable":false,"text":"..."}]}
```

For structured parts, `cacheflow` concatenates `parts[].text` in order and uses `stable` only for diagnostics.

## CLI

```powershell
omni cache -model gpt-4o -profile openai "hello world"
omni cache-sim -model gpt-4o -profile openai -input prompts.jsonl -breakers
```

## Profiles

```go
cacheflow.ProfileGeneric
cacheflow.ProfileOpenAI
```

Profiles are local planning helpers. Provider behavior can change, and final usage metadata remains authoritative.
126 changes: 126 additions & 0 deletions cacheflow/align.go
Original file line number Diff line number Diff line change
@@ -0,0 +1,126 @@
// Package cacheflow provides dependency-free prompt-cache planning and trace analysis.
package cacheflow

import (
"fmt"

"github.com/ron2111/omnitoken"
)

// Profile describes local token-boundary rules for cache planning.
type Profile struct {
Name string `json:"name"`
BlockSize int `json:"block_size"`
MinimumTokens int `json:"minimum_tokens"`
}

// Common cache-planning profiles. Provider behavior can change; use these as
// local planning helpers, not billing guarantees.
var (
ProfileGeneric = Profile{Name: "generic", BlockSize: 1024}
ProfileOpenAI = Profile{Name: "openai", BlockSize: 128, MinimumTokens: 1024}
)

// Alignment describes how close a prompt or prefix is to a cache block boundary.
type Alignment struct {
CurrentTokens int `json:"current_tokens"`
BlockSize int `json:"block_size"`
MinimumTokens int `json:"minimum_tokens"`
PreviousBlockSize int `json:"previous_block_size"`
NextBlockSize int `json:"next_block_size"`
Remainder int `json:"remainder"`
PaddingNeeded int `json:"padding_needed"`
TokensUntilMinimum int `json:"tokens_until_minimum"`
IsAligned bool `json:"is_aligned"`
IsEligible bool `json:"is_eligible"`
StrategyHint string `json:"strategy_hint"`
}

// Aligner evaluates prompt lengths against provider cache block sizes.
type Aligner struct {
engine omnitoken.ModelEngine
}

// NewAligner creates a prompt-cache alignment helper for an engine.
func NewAligner(engine omnitoken.ModelEngine) *Aligner {
return &Aligner{engine: engine}
}

// AlignPrompt evaluates prompt length against a custom cache block size.
func (a *Aligner) AlignPrompt(text string, providerBlockSize int) Alignment {
return a.AlignPromptToProfile(text, Profile{Name: "custom", BlockSize: providerBlockSize})
}

// AlignPromptToProfile evaluates prompt length against a cache-planning profile.
func (a *Aligner) AlignPromptToProfile(text string, profile Profile) Alignment {
if a == nil || a.engine == nil {
return Alignment{StrategyHint: "Invalid cache alignment configuration"}
}
return AlignTokenCount(a.engine.CountTokens(text), profile)
}

// AlignTokenCount evaluates an already-computed token count against a profile.
func AlignTokenCount(tokens int, profile Profile) Alignment {
blockSize := profile.BlockSize
if blockSize <= 0 {
return Alignment{CurrentTokens: tokens, StrategyHint: "Invalid cache alignment configuration"}
}
minimum := profile.MinimumTokens
if minimum < 0 {
minimum = 0
}

remainder := tokens % blockSize
previous := tokens - remainder
padding := 0
if remainder == 0 {
previous = tokens
} else {
padding = blockSize - remainder
}
next := tokens + padding
eligible := tokens >= minimum
untilMinimum := 0
if !eligible {
untilMinimum = minimum - tokens
if next < minimum {
next = roundUp(minimum, blockSize)
padding = next - tokens
}
}

return Alignment{
CurrentTokens: tokens,
BlockSize: blockSize,
MinimumTokens: minimum,
PreviousBlockSize: previous,
NextBlockSize: next,
Remainder: remainder,
PaddingNeeded: padding,
TokensUntilMinimum: untilMinimum,
IsAligned: remainder == 0,
IsEligible: eligible,
StrategyHint: strategyHint(tokens, minimum, padding, remainder),
}
}

func roundUp(value int, blockSize int) int {
if value <= 0 {
return 0
}
remainder := value % blockSize
if remainder == 0 {
return value
}
return value + blockSize - remainder
}

func strategyHint(tokens int, minimum int, padding int, remainder int) string {
if minimum > 0 && tokens < minimum {
return fmt.Sprintf("Prompt is %d tokens below the configured cache minimum", minimum-tokens)
}
if padding == 0 && remainder == 0 {
return "Prompt is aligned to the configured cache block boundary"
}
return fmt.Sprintf("Prompt is %d tokens from the next configured cache block boundary", padding)
}
Loading