Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
3 changes: 3 additions & 0 deletions .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -9,3 +9,6 @@ build/
.DS_Store
.pytest_cache/
PUBLISHING.md

# Benchmark receipts from local runs (the committed examples under samples/ are kept)
/bench-receipts/
50 changes: 50 additions & 0 deletions CLAUDE.md
Original file line number Diff line number Diff line change
Expand Up @@ -65,8 +65,21 @@ src/av/
│ ├── ffmpeg.py # ffmpeg/ffprobe wrappers
│ ├── chunker.py # Text chunking for embeddings
│ └── dense_caption.py # Structured event export
├── bench/ # Cost/accuracy frontier measurement
│ ├── cost.py # Per-hour vs per-token cost models (never conflated)
│ ├── datasets.py # Adapters for public benchmarks (no data vendored)
│ ├── fixtures.py # Deterministic ffmpeg fixtures for the ordering gate
│ ├── frames.py # Pinned frame extraction (command recorded in the receipt)
│ ├── receipts.py # Labelled evidence + endpoint redaction
│ ├── runner.py # Sweep axes, noise floor, collapse point
│ ├── vlm.py # Provider-agnostic multi-image calls with usage accounting
│ └── tasks/
│ ├── ordering.py # Temporal-ordering capability gate
│ ├── videoqa.py # Dense vs agentic arms
│ └── events.py # Event recall vs sampling interval
├── providers/
│ ├── base.py # Abstract interfaces
│ ├── deepseek.py # DeepSeek-V4.1-Flash (SGLang) config + image-token model
│ └── openai.py # OpenAI-compatible client (works for all providers)
├── search/
│ ├── semantic.py # FTS5 + cosine reranking
Expand Down Expand Up @@ -117,6 +130,38 @@ On partial failure: `{"status": "complete_with_warnings", ..., "warnings": ["Tra
### `av list` / `av info <id>` / `av transcript <id>` / `av export` / `av open <id>`
See `av <command> --help` for details.

### `av bench`
Measures the cost/accuracy frontier. Headline axes are **tokens per query** and
**accuracy**, plus **dollars per query on hardware you own** — the axis a per-token
API vendor cannot report.

```bash
av bench probe # is tokens-per-frame tunable on this endpoint?
av bench gate --sizes 2,4,8 # can the model order frames at all?
av bench plan --budgets 200,400,800 # predicted token cost per resolution (offline)
av bench prepare minerva ann.json -o task.jsonl
av bench run task.jsonl --arms dense,agentic --cost hourly:25.0:20000
av bench sweep captions.jsonl videos/ --intervals 1,2,5,10,30
av bench noise --repeats 5
av bench cost --tokens-per-frame 1024 --prefill-tok-s 20000 --hourly-usd 25
```

Every subcommand writes a JSON receipt to `./bench-receipts/`.

**Rules that are not optional here:**
- **Run the gate before quoting any score.** A model that cannot order eight frames
is not being measured on temporal understanding, and its throughput is irrelevant.
- **Label every claim.** `measured` / `derived` / `documented` / `community-reported`
/ `untested`. Non-measured claims must cite a source; `Claim` raises otherwise.
- **Publish the noise floor**, and never on a saturated cell.
- **Never conflate `$/hr` and `$/token`.** They are different economics.
- **Never put someone else's published number and ours in one cell as a ratio.**
Different model, different hardware, different methodology — it is a comparison of
approaches, not a head-to-head.
- **No endpoints in source.** Receipts record a hostname; private hosts are redacted.
- **No benchmark data is vendored.** Adapters read files the user fetched, under the
upstream licence.

## Provider Support

| Provider | Transcription | Vision/Chat | Embeddings | Setup |
Expand All @@ -125,6 +170,7 @@ See `av <command> --help` for details.
| OpenAI (API key) | whisper-1 | gpt-4-1 | text-embedding-3-small | Paste `sk-...` |
| Anthropic | -- | claude-sonnet-4-5 | -- | Paste API key |
| Gemini | -- | gemini-2.5-flash | text-embedding-004 | Paste API key |
| DeepSeek-V4.1-Flash | -- | deepseek-v4.1-flash (self-hosted SGLang) | -- | `AV_API_BASE_URL` + `DEEPSEEK_API_KEY` |

When a capability is unavailable (e.g. Anthropic has no Whisper), the pipeline skips that stage and warns.

Expand All @@ -140,6 +186,8 @@ When a capability is unavailable (e.g. Anthropic has no Whisper), the pipeline s
| `AV_EMBED_MODEL` | `text-embedding-3-small` | Embedding model |
| `AV_CHAT_MODEL` | `gpt-4-1` | Chat/RAG model |
| `AV_DB_PATH` | `~/.config/av/av.db` | Database location |
| `DEEPSEEK_API_KEY` | (none) | Key for a self-hosted DeepSeek-V4.1-Flash server |
| `SGLANG_API_KEY` | (none) | Alias for the same, matching SGLang's own naming |

## Database

Expand Down Expand Up @@ -182,5 +230,7 @@ Near-term priorities for contributors:
- [ ] Cross-video search improvements (search across all indexed videos at once)
- [ ] Streaming ingest progress (SSE-style output for long videos)
- [ ] Profile presets for dense captioning (security, retail, meeting, etc.)
- [ ] `av bench` precision measurement (currently only recall against reference windows)
- [ ] Video fetch helper for benchmark task files (yt-dlp, with link-rot reporting)
- [x] CI/CD with GitHub Actions (lint + test on PR) — `.github/workflows/ci.yml`
- [x] PyPI publish workflow — `.github/workflows/publish.yml` + `PUBLISHING.md`
131 changes: 128 additions & 3 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -65,6 +65,12 @@ av sentinel video.mp4 # Detect events (all 4 alert types)
av sentinel video.mp4 --alerts FALL # Fall detection only
av sentinel video.mp4 -p ollama # Self-hosted (free)
av sentinel videos/ -c cam_lobby # Batch with camera tracking

# Benchmarking
av bench probe # What can this deployment actually do?
av bench gate # Can it order frames at all? Run this first.
av bench run task.jsonl # Dense vs agentic, with tokens and dollars
av bench sweep captions.jsonl vids/ # Where does recall collapse as frames thin out?
```

## Sentinel — Surveillance Event Detection
Expand Down Expand Up @@ -102,6 +108,105 @@ Video → 30s chunks (5s overlap)

Built on 107 experiments across 21 vision models. Key insight: structural extraction + temporal rules beats generic "detect anomalies" prompts.

## Bench — Cost/Accuracy Frontier

Selling video understanding on hardware you own means one number decides everything:
**video-hours analysed per dollar**. `av bench` measures it, and measures what it
costs you in accuracy to get there.

Two headline axes, chosen so results read against published agentic-video
comparisons: **tokens per query** and **accuracy**. Alongside them sits the axis an
API vendor cannot report — **dollars per query on your own box** — because per-token
billing and per-hour hardware are different economics and the tool never conflates
them.

### Run the gate first

```bash
av bench gate --sizes 2,4,8
```

Deterministic ffmpeg fixtures carrying a known order, one question, exact-match
scoring. A model that cannot report the order of eight flat colours cannot be
meaningfully scored on long-video reasoning, and any throughput number measured
against it describes a machine doing the wrong thing quickly. The gate costs cents
and it can save the whole exercise.

### Establish what is tunable before sweeping it

```bash
av bench probe
av bench plan --widths 512,768,1024,1536 --budgets 200,400,800
```

`probe` tests two candidate knobs against your live endpoint — the OpenAI `detail`
hint and the resolution actually uploaded — because a server may honour one and
silently ignore the other. If neither moves the per-frame token count, the
tokens-per-frame axis is reported as fixed rather than faked. `plan` predicts the
same thing offline from a published preprocessor algorithm, and shows the two walls
worth knowing: an upscale floor below which shrinking frames buys nothing, and a
token ceiling above which extra resolution is discarded.

### Dense versus agentic

```bash
av bench prepare minerva minerva.json --out task.jsonl --max-questions 40 --max-videos 6
av bench run task.jsonl --arms dense,agentic --cost hourly:25.0:20000
```

The **dense** arm samples the whole window at a fixed rate and asks once. The
**agentic** arm takes a cheap coarse look, decides which moments it needs, then
fetches only those — and is charged for both requests. Nothing else differs between
them.

`av bench prepare` adapts a public benchmark's annotations into the task format.
**No benchmark data ships with av and no videos are downloaded.** Fetch annotations
yourself and mind their licences: MINERVA's are CC BY 4.0, LVBench's are
CC BY-NC-SA with an explicit commercial-use prohibition, and neither grants any
rights to the videos themselves.

### Where does it collapse?

```bash
av bench sweep captions.jsonl videos/ --intervals 1,2,5,10,30 --cost token:0.30:2.50
```

Event detection against sampling interval on real footage. The interval at which
detection collapses is the cheapest safe sampling rate — and it is a per-task
answer, not a global one. Smoke tolerates sparse frames; a door opening does not.

### Noise floor

```bash
av bench noise --repeats 5
```

Runs one unchanged cell repeatedly and publishes the spread. This is the number that
makes every other number readable: a delta smaller than the spread is noise. Point it
at a cell the model does not already solve perfectly — a saturated cell has no
headroom to vary, and the tool says so rather than reporting a meaningless zero.

### Receipts

Every subcommand writes a JSON receipt to `./bench-receipts/` carrying the provider,
the determinism controls, the exact ffmpeg invocations, fixture hashes, the cost
model, and every cell. Claims are labelled `measured`, `derived`, `documented`,
`community-reported`, or `untested`, and a non-measured claim must cite a source.
Endpoints are reduced to a hostname, and private or tunnelled hosts never appear at
all — receipts are meant to be published.

### Cost model

```bash
av bench cost --tokens-per-frame 1024 --context-tokens 1048576 \
--prefill-tok-s 20000 --hourly-usd 25 --kv-bytes-per-token 890 \
--source "your measurements"
```

Pure arithmetic, no API calls, every input recorded. Supply `--cost hourly:RATE` for
hardware you own or `--cost token:IN:OUT` for a vendor API — they are different
shapes and reporting one in the other's units produces a number that means nothing.

## Configuration

### Interactive Setup (Recommended)
Expand All @@ -110,14 +215,21 @@ Built on 107 experiments across 21 vision models. Key insight: structural extrac
av config setup
```

Choose from four providers:
Choose from six providers:

| # | Provider | Auth | Transcription | Embeddings |
|---|----------|------|---------------|------------|
| 1 | **OpenAI (Codex OAuth)** | Auto-detected | Whisper | text-embedding-3-small |
| 2 | **OpenAI (API key)** | `sk-...` key | Whisper | text-embedding-3-small |
| 3 | **Anthropic (Claude)** | API key | Not supported | Not supported |
| 4 | **Google (Gemini)** | API key | Not supported | text-embedding-004 |
| 3 | **PixelML (OpenRouter)** | API key | Not supported | Not supported |
| 4 | **Anthropic (Claude)** | API key | Not supported | Not supported |
| 5 | **Google (Gemini)** | API key | Not supported | text-embedding-004 |
| 6 | **DeepSeek-V4.1-Flash** | Your own endpoint | Not supported | Not supported |

**DeepSeek-V4.1-Flash** talks to an OpenAI-compatible SGLang server that you run.
No endpoint ships with `av` — the preset defaults to SGLang's own local bind
address, and you point `AV_API_BASE_URL` at your deployment. Set `DEEPSEEK_API_KEY`
if your server requires one; leave it unset if it does not.

Config is saved to `~/.config/av/config.json` and persists across sessions.

Expand All @@ -134,6 +246,11 @@ export AV_TRANSCRIBE_MODEL="whisper"
export AV_VISION_MODEL="gpt-4-1"
export AV_EMBED_MODEL="text-embedding-3-small"
export AV_CHAT_MODEL="gpt-4-1"

# Self-hosted DeepSeek-V4.1-Flash via SGLang
export AV_PROVIDER="deepseek"
export AV_API_BASE_URL="http://your-sglang-host:30000/v1"
export DEEPSEEK_API_KEY="..." # only if your server requires one
```

## Requirements
Expand All @@ -156,6 +273,14 @@ export AV_CHAT_MODEL="gpt-4-1"
| `av transcript <id>` | Output transcript (VTT/SRT/text) |
| `av export` | Export as JSONL/VTT/SRT |
| `av open <id> --at <sec>` | Open video at timestamp |
| `av bench gate` | Temporal-ordering capability gate |
| `av bench probe` | Measure a deployment's image-token and multi-image behaviour |
| `av bench plan` | Predict per-frame token cost against resolution (offline) |
| `av bench prepare` | Adapt a public benchmark's annotations into a task file |
| `av bench run` | Dense vs agentic arms, with tokens and dollars |
| `av bench sweep` | Event recall against sampling interval |
| `av bench noise` | Spread across identical runs |
| `av bench cost` | Cost arithmetic with labelled inputs (offline) |
| `av version` | Print version JSON |

## License
Expand Down
48 changes: 48 additions & 0 deletions samples/bench-receipts/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,48 @@
# Sample bench receipts

Real output from `av bench`, committed so the claims in the README have something
behind them. Every file here was produced by the commands named below on
2026-09-10, against **Gemini 2.5 Flash via its OpenAI-compatible endpoint** — chosen
because it was the provider already configured on the machine, not because it is the
subject of the benchmark.

Read these as a demonstration that the harness measures what it says it measures.
They are **not** a model comparison, and no result here is a claim about any other
model or deployment.

| File | Command | What it shows |
|------|---------|---------------|
| `gate-color-2-4-6-8.json` | `av bench gate --kind color --sizes 2,4,6,8` | Exact-order accuracy at 2, 4, 6 and 8 frames |
| `probe-tokens-per-frame.json` | `av bench probe` | Whether per-frame token cost is tunable on this endpoint |
| `noise-floor-saturated.json` | `av bench noise --n 6 --repeats 5` | A saturated cell, and the harness refusing to call its zero spread a noise floor |
| `sweep-two-probes-five-intervals.json` | `av bench sweep .../captions.jsonl .../videos --probes door_activity,person_enters --intervals 1,2,5,10,30 --max-per-probe 8 --cost token:0.30:2.50` | The full frontier: two probes collapsing at different rates |
| `sweep-door-activity.json` | `av bench sweep .../captions.jsonl .../videos --probes door_activity --intervals 1,5,30 --max-per-probe 4 --cost token:0.30:2.50` | Where detection collapses as frames thin out |
| `plan-image-tokens.json` | `av bench plan --budgets 200,400,800` | Predicted per-frame token cost against resolution (offline) |
| `cost-model-arithmetic.json` | `av bench cost --tokens-per-frame 1024 --context-tokens 1048576 --prefill-tok-s 20000 --hourly-usd 25 --kv-bytes-per-token 890` | The frontier arithmetic, with inputs recorded (offline) |

## Reading a receipt

- `claims` carries the labelled statements. `measured` means this run produced it;
`derived` means arithmetic over inputs; `documented` and `community-reported` must
cite a source; `untested` means we did not check.
- `determinism` carries temperature, seed, the exact ffmpeg invocations, and fixture
hashes, so a cell can be regenerated byte-for-byte.
- `provider.endpoint_host` is a hostname only. Private and tunnelled hosts are
redacted to `<private>` — receipts are meant to be published.
- `notes` carries caveats that apply to the whole run, including the reference caveat
on event sweeps.

## Two caveats that apply to the sweep receipt

1. **The reference is model-generated.** Event windows come from the dense captioning
run shipped in `samples/epstein-cctv`, not from human annotation. The metric is
named `reference_recall` for that reason and must not be quoted as recall.
2. **The sample is small.** Four to eight windows per probe. It demonstrates the
shape of the frontier; it does not establish a rate to two significant figures.

## What is deliberately absent

No receipt here was produced against a self-hosted DeepSeek-V4.1-Flash deployment.
The provider and its image-token model are implemented and unit-tested, but the
harness has not been pointed at a running server, so no measured claim about that
model appears anywhere in this repository.
Loading
Loading