Skip to content
Draft
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
10 changes: 10 additions & 0 deletions CLAUDE.md
Original file line number Diff line number Diff line change
Expand Up @@ -133,6 +133,10 @@ automatically runs Jev source relevance, configurable bounded scene grouping, an
synthesis, and a separate answer-support Noul. `--no-refine` preserves legacy RAG
for one request. Refined responses retain answer/citations/confidence and add route,
evidence, refinement, warning, inspected-window, and stage-usage metadata.
With `AV_DJEV_ENDPOINT` set, the same refinement runs against a self-hosted
[djev-spark](https://github.com/mmastrac/djev-spark) server (same
`/v1/systemone` contract); the served model and engine are recorded in usage
receipts, and a djev answer is never presented as a Jev measurement.

### `av list` / `av info <id>` / `av transcript <id>` / `av export` / `av open <id>`
See `av <command> --help` for details.
Expand Down Expand Up @@ -201,6 +205,12 @@ When a capability is unavailable (e.g. Anthropic has no Whisper), the pipeline s
| `AV_TYPESAFE_API_KEY` | (none) | Jev/System One key; enables ask refinement by default |
| `AV_TYPESAFE_ENDPOINT` | `https://api.typesafe.ai/v1/systemone` | Explicit System One endpoint |
| `AV_TYPESAFE_MODEL` | `jev-latest` | System One model |
| `AV_DJEV_ENDPOINT` | (none) | Self-hosted djev-spark `/v1/systemone` endpoint; selects the self-hosted decision lane over hosted Jev |
| `AV_DJEV_API_KEY` | (none) | Bearer key for the djev-spark server (needed only when the server sets `API_KEY`) |
| `AV_DJEV_MODEL` | (none) | Advisory request model; djev-spark ignores it and reports what it served |
| `AV_DJEV_TIMEOUT_SEC` | `180` | Per-attempt djev-spark timeout (cold reads on long states are slow) |
| `AV_DJEV_MAX_RETRIES` | `1` | Explicit retry count for djev-spark requests |
| `AV_DJEV_SEED` | `42` | Sampler seed sent with every djev-spark decision |
| `AV_REFINE_RELEVANCE_MIN` | `0.5` | Minimum source-relevance Noul probability |
| `AV_REFINE_SUPPORT_MIN` | `0.5` | Minimum answer-support Noul probability |
| `AV_REFINE_MAX_SCENES` | `8` | Maximum merged scenes sent to synthesis |
Expand Down
1 change: 1 addition & 0 deletions cookbook/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -6,6 +6,7 @@ Runnable recipes for the open-source **av** CLI live here alongside the code.
|---|---|---|
| [Cost model](cost-model/README.md) | Separate tokens, estimates, unknown costs, ingestion, queries, failures, reservations, and cap accounting | Completed one-question comparison; no aggregate parity claim |
| [Jev-refined ask](jev-refined-ask/README.md) | Build a source-bound transcript sidecar, retrieve indexed moments, refine evidence, answer, and check support | Runnable recipe; no speed, cost, or quality parity claim |
| [djev-spark refined ask](djev-clipping/README.md) | Run the same refined-ask workflow against a self-hosted djev-spark server, with served-model receipts and visible response validation | Adapter + offline compatibility tests complete; no live quality, latency, or parity claim |
| [Sanitized receipts](receipts/README.md) | Completed ASR/baseline, caption smoke/abort, incompatible cap probes, Grok ingestion, and Jev-refined query | No media, transcript/caption corpus, credentials, upload URIs, or private routes; baseline question and returned answer/rationale retained |

The [original public cost notebook](https://github.com/PixelML/cookbook/tree/main/agentic-video/cost-model)
Expand Down
121 changes: 121 additions & 0 deletions cookbook/djev-clipping/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,121 @@
# djev-spark refined ask — self-hosted decisions over the same contract

This recipe runs the [Jev-refined ask](../jev-refined-ask/README.md) workflow
against a **self-hosted [djev-spark](https://github.com/mmastrac/djev-spark)**
server instead of the hosted TypeSafe System One endpoint. Both speak the same
documented `POST /v1/systemone` contract, so `av ask` refinement, evidence
grouping, and answer-support checking work unchanged — only the decision
provider changes.

**Evidence status:** adapter and offline compatibility tests are complete; no
live djev quality or latency measurement has been made in this repository.
Everything below about runtime performance is upstream's own documentation and
is labelled as such.

## What the adapter guarantees

Audited upstream revision: `mmastrac/djev-spark`
`1444f3e927f83ba508e5b28a4fd4fdd9ecd0976b`.

- **Identity honesty.** djev-spark ignores a request's `model` field and
reports the model it actually served. `DjevClient` records
`served_model`, `server_engine`, and the redacted `served_endpoint_host`
into every stage-usage receipt and into the `refinement` metadata of
`av ask` output. A djev answer is never presented as a Jev measurement.
- **Visible rejection.** Responses that are malformed or incomplete for the
questions actually asked — missing answers, answers to unasked questions,
type mismatches, choices outside the offered options, score legends that do
not match the levels, probabilities out of range or not summing to one —
fail with a `djev-spark response failed validation` error instead of being
silently coerced. A question the server skipped (`ask_if`) is preserved as
`null`.
- **Same failure semantics as the hosted lane.** Timeouts, explicit retries
(429/5xx and connection errors, exponential backoff), sanitized error
messages, and provider-accurate fallback warnings ("djev-spark refinement
was unavailable…") match `SystemOneClient` conventions.
- **Reproducibility knob.** Every request carries `AV_DJEV_SEED`
(default `42`), the server-side default seed.

## Configuration

No endpoint ships with av. djev-spark is software you host; an explicit
endpoint selects this lane over hosted Jev (the endpoint wins even if a
TypeSafe key is also set).

| Variable | Example | Purpose |
|---|---|---|
| `AV_DJEV_ENDPOINT` | `http://10.1.2.3:8011/v1/systemone` | Structured server endpoint (compose default port `8011`) |
| `AV_DJEV_API_KEY` | any string | Sent as `Authorization: Bearer …`; needed only when the server sets `API_KEY` |
| `AV_DJEV_MODEL` | `dgemma` | Advisory only; the server ignores it. Omit unless you want it recorded in request logs |
| `AV_DJEV_TIMEOUT_SEC` | `180` | Cold structured reads on long states are slow (see below) |
| `AV_DJEV_MAX_RETRIES` | `1` | Retry count, mirroring the hosted lane's default |
| `AV_DJEV_SEED` | `42` | Sampler seed sent with every decision |

Verify wiring offline:

```bash
av config show | jq '.djev_endpoint, .djev_api_key'
```

Then ask exactly as in the Jev recipe:

```bash
av ask "When does the door open?" --video-id <id>
```

The response's `ask_settings.decision_provider` reads `"djev-spark"`, and
`refinement.served_model` / `refinement.server_engine` carry the identity the
server reported. If the server is unreachable, `route` becomes
`refinement_fallback`, the warning names the provider, and the answer is
produced from raw retrieval with `evidence_status: raw_unjudged`.

## Running the server safely (read before you start it)

These are requirements this project imposes on itself; the upstream defaults
are more permissive.

- **Networking.** Upstream's compose file uses host networking and both
servers listen on all interfaces. Put the box behind a firewall or tailnet
and reach it over a private address.
- **Authentication.** Upstream serves POST routes with no API key unless its
`API_KEY` env var is set. Set one, and set `AV_DJEV_API_KEY` to match.
`/health` (GET) stays open by upstream design.
- **Playground.** Upstream's test page (`TEST_PAGE=1`) is off by default;
leave it off on any shared host.
- **Resources.** The model is DiffusionGemma 26B-A4B NVFP4 and wants a
compatible GPU, verified non-boot storage for weights and build caches, and
the upstream entrypoint's own headroom check. Never place weights, Docker
caches, or media on the av control plane's boot disk. Do not start this
next to workloads you do not own; do not evict anything to make room.
- **Model terms.** Gemma weight licence terms apply to the checkpoint; the
server source files carry Apache-2.0 headers. av implements the wire
protocol independently and distributes no upstream code or weights.

## Performance claims (upstream-documented, untested here)

Upstream's README reports roughly 0.1-second warm reads and, on its 128k
profile, a 104.94-second cold versus 0.44-second warm selection at a
110,707-token state. These are the author's numbers on their hardware — not
measurements from this repository, and not a guarantee that any full video
clips in two seconds. Measure your own cold/warm/export timings before
relying on any latency figure; record p50/p95 with sample counts when you do.

## Comparison protocol (pending)

A fair comparison against hosted Jev and the plain retrieval baseline must use
the same frozen corpus, queries, candidate windows, and transcript/vision text
as the sibling Jev clipping task, with independent labels — not provider
scores as ground truth. Measure relevance, recall, absent-topic false
positives, timing boundaries, duplication, and export validity separately per
lane. That evaluation is **not started here**: it is gated on the shared
fixture contract and on an authorized runtime. This recipe will gain a
results section only from actual recorded runs.

## Known limitations

- Image input (`data:` URLs / multipart parts) is supported by the upstream
server but not sent by this adapter; the primary comparison is text-to-text.
- Clip candidate construction and export belong to the shared clipping
contract, not to this provider adapter.
- Offered-label probabilities are diagnostics; nothing here treats them as
calibrated quality scores.
4 changes: 4 additions & 0 deletions src/av/cli/config_cmd.py
Original file line number Diff line number Diff line change
Expand Up @@ -73,6 +73,10 @@ def config_show() -> None:
"typesafe_api_key": "***" if config.typesafe_api_key else "(not set)",
"typesafe_endpoint": config.typesafe_endpoint,
"typesafe_model": config.typesafe_model,
"djev_endpoint": config.djev_endpoint or "(not set)",
"djev_api_key": "***" if config.djev_api_key else "(not set)",
"djev_model": config.djev_model or "(not set)",
"djev_seed": config.djev_seed,
"refine_enabled": config.refine_enabled,
"refine_relevance_min": config.refine_relevance_min,
"refine_support_min": config.refine_support_min,
Expand Down
19 changes: 19 additions & 0 deletions src/av/core/config.py
Original file line number Diff line number Diff line change
Expand Up @@ -77,6 +77,19 @@ class AVConfig(BaseSettings):
typesafe_model: str = Field(default="jev-latest")
typesafe_timeout_sec: float = Field(default=30.0, gt=0)
typesafe_max_retries: int = Field(default=1, ge=0, le=3)
# Optional self-hosted djev-spark decision endpoint speaking the same
# documented /v1/systemone contract. No default endpoint ships with av:
# an explicit endpoint selects this lane over hosted System One.
djev_endpoint: str = Field(default="")
djev_api_key: str = Field(default="")
# Advisory request model. djev-spark ignores it and reports the model it
# actually served; responses carry that identity into usage records.
djev_model: str = Field(default="")
# Cold structured reads are slow on long states (upstream documents ~105 s
# at 110k tokens), so the default is more generous than the hosted lane's.
djev_timeout_sec: float = Field(default=180.0, gt=0)
djev_max_retries: int = Field(default=1, ge=0, le=3)
djev_seed: int = Field(default=42, ge=0)
refine_enabled: bool = Field(default=True)
refine_relevance_min: float = Field(default=0.5, ge=0, le=1)
refine_support_min: float = Field(default=0.5, ge=0, le=1)
Expand Down Expand Up @@ -130,6 +143,12 @@ def get_config(db_path: Path | None = None) -> AVConfig:
"typesafe_endpoint",
"typesafe_model",
"typesafe_timeout_sec",
"djev_endpoint",
"djev_api_key",
"djev_model",
"djev_timeout_sec",
"djev_max_retries",
"djev_seed",
"typesafe_max_retries",
"refine_enabled",
"refine_relevance_min",
Expand Down
Loading
Loading