From 46550767d66dfa5b2b30650bdc5aaf13b4edf00d Mon Sep 17 00:00:00 2001 From: ModelMirror <273825391+modelmirror@users.noreply.github.com> Date: Mon, 21 Sep 2026 19:44:36 +0000 Subject: [PATCH 1/2] feat(ops): restructure the weekly digest around three production windows MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The weekly performance digest drops its Health questions and Analytics state blocks — the boards are read on the website, and the health bullets restated the production section — and reports production over three windows instead: the week, the trailing month, and the October Term to date. Each block is one shape (cells, events, measured spend, role x stage table) so the three are read the same way. The month block's window is the ex-post spend backstop's own, taken from its config, and the backstop's verdict closes that block: the verdict and the census printed beside it can then never describe different periods. The Term block closes with the forward cells scored under the process in force — a forward cell is minted once at its event and never again, so the Term is the period that count is worth reading over — stated as a figure, with the frozen-shakedown state named rather than shown as a bare zero. `cell_census` and `spend_over` take a `since` cutoff so the Term window can be pinned to the instant its Term opened: an October Term opens at midnight on 1 October and the digest renders mid-morning, so a day count would cut the Term's own first morning out of its own census. Artifact: the `weekly-digest` issue run-ops opens on the Monday tick. Co-Authored-By: Claude Fable 5.1 --- .github/workflows/run-ops.yml | 4 +- docs/cli.md | 2 +- docs/pipeline.md | 46 +-- src/fedcourtsai/cli.py | 76 ++-- src/fedcourtsai/ops.py | 491 +++++++------------------ src/fedcourtsai/spend.py | 26 +- src/fedcourtsai/store.py | 18 +- tests/test_ops.py | 664 +++++++++++++++++----------------- tests/test_spend.py | 23 +- 9 files changed, 588 insertions(+), 762 deletions(-) diff --git a/.github/workflows/run-ops.yml b/.github/workflows/run-ops.yml index 765282771..a658cb0ca 100644 --- a/.github/workflows/run-ops.yml +++ b/.github/workflows/run-ops.yml @@ -27,8 +27,8 @@ name: run-ops # The reading surfaces are digests a maintainer closes, not a standing body this # job edits in place: an unread digest is an open issue, so the backlog is # visible without a reading-state store. The Monday schedule opens the **weekly -# performance digest** as its own `weekly-digest` issue — the health questions, -# the committed boards' state, the week's cells and measured spend, and the +# performance digest** as its own `weekly-digest` issue — the cells and measured +# spend produced over the week, the trailing month and the Term to date, and the # back-test results, every metrics figure carrying its artifact's vintage. A # second job on the daily tick opens the **daily prediction-reading digest**: # one predicted event with every predictor side by side, on its own diff --git a/docs/cli.md b/docs/cli.md index 324f52853..da8294730 100644 --- a/docs/cli.md +++ b/docs/cli.md @@ -145,7 +145,7 @@ committed), plus the spend ledger. | `qp-corpus` | Extract the stored `questions-presented` texts a `qp-topic-v0` labeler reads — a JSON list of `{case_id, docket_number, text}` in `case_id` order, the whole input the text-only vocabulary entitles a labeler to (no docket context, no case name, no outcome; see [qp-topic.md](qp-topic.md)). **Scoped to the labeling population**: the live/historical slice's modern discretionary-cert petitions — the same scope the docket pack's topic section is computed over, narrowed to its QP-bearing part, so every labeled row has a published home and no QP-bearing row in that scope is unlabelable. It deliberately stops there rather than narrowing to the predict-scope segment, which would drop the IFP stream while the hand reference set spans both: the publication gate's coverage floor would go out of reach, and carrying the reference set back in would make an IFP docket number a certain reference-membership tell ([qp-topic.md](qp-topic.md)). `--all` is the unscoped **measurement** form (every stored questions-presented row in the blob, whatever its case), for reading what a given file holds; it is not a labeling selection. The scoped form reads each case's documents through the registered payload path, so it serves from the per-case content store under the corpus split and from the blob otherwise, while `--all` reads the blob's `documents` table directly — which a split blob leaves empty by construction. Opens the corpus **strictly read-only** and never migrates it, so either form is safe against a pulled blob; a row whose case carries no docket number, or whose stored text is empty, is skipped and counted (the docket number is half the key the reference join is checked on, and an empty extraction is nothing to label). **Cuts a batch sized to the labeling ceiling** (`pipeline.qp_topics.LABEL_ROW_CEILING`, derived from the labeler step's own cap): the labeler runs as one headless turn under that cap and only a complete label file yields the labels artifact, so an over-budget extract buys a killed step and no published labels rather than partial coverage — the rows it did write are uploaded as a run artifact and counted in the job summary, for reading rather than for accrual. The batch is derived from committed state alone — every reference-set case in the frame, force-included so the run stays measurable, plus a Term × fee-class-stratified draw of the rows the committed labels artifact does not yet publish, ordered inside each stratum by a seeded hash of `case_id` (`BATCH_ORDER_SEED`) rather than by docket order, which would select on Term and fee class. There is nothing to choose and no flag: the same committed state always cuts the same batch, so repeat dispatches clear the frame batch by batch and each row is labeled once. The scoped form therefore **requires** the committed reference set. Its arithmetic — frame, labeled so far, batch size, reference share, per-stratum pool and draw — goes to stderr and to a `.batch.json` sidecar beside the extract, deliberately never into the extract, which is the labeler's text-only evidentiary input. One number out of that sidecar travels on: the run mode reads the frame size in the extract job and hands it to `qp-topics --frame-rows` as a value, so the labels artifact records the frame each batch was cut from while the draw's shape stays behind. Two states end the run in the extract job rather than after a labeling dispatch has been spent: a frame with nothing left to label outside the reference set is **converged**, and a frame holding under the 90% coverage floor of the reference set could not clear the publication gate whatever the labeler did (the batch carries every reference case in the frame and the labeler labels every row, so that coverage is exact at extract time). Both say which number they are and write nothing. Neither moves the gate; the second enforces it earlier and more cheaply. `--all` keeps the flat over-ceiling refusal, since a measurement form has nothing to accrue against. A working file for one labeler run, **never committed** — it republishes stored petition text and enumerates the ingested corpus, neither of which any committed surface does, so an `--out` is refused outright under either of two boundaries — the work tree (an untracked file in the checkout is one `git add -A` from being committed) and the configured `data_root`, which a non-default setting can place outside it. Write it to a runner-local path such as `$RUNNER_TEMP`. The `run-analytics` mode that dispatches it is in [pipeline.md](pipeline.md). | `--out`, `--corpus`, `--all` | | `qp-topics` | Measure a topic labeler's JSONL against the hand reference set and accrue `data/qp-topics/qp-topics.json`: the artifact accumulates, so what is written is the union of the committed artifact and this batch's new rows, with prior entries carried forward unchanged. A row inside the reference set publishes the **hand set's adjudicated label**, never the labeler's — the labeler's call there is measurement input, scored into the agreement rate and discarded, so a flip on a reference case moves this run's rate and can change no published row; every other row is published once, since later batches exclude what is already published, and a second labeling of one stops the run. Each row records the batch that first published it and its `source`, and a per-batch ledger keeps every contributing run's labeler and agreement figures. Every label is validated against the `qp-topic-v0` vocabulary, and the labels joined to the extract and to the reference on `case_id` **and** `docket_number` (a half-matching pair is a mis-join and stops the run; so does a labels file that is not exactly the extract's case set, since a partial run measures a prefix rather than a sample and turns the reported `n` into a membership probe on an outcome-encoding reference set). Reports overall agreement beside the rate a **constant labeler** would score, floor-gated per-label agreement, the constitutional-rights / criminal-law / civil-procedure confusion matrix, and the deterministic shadow rules' disagreement rate — whose *level* is uninterpretable off the reference set, only its movement between runs. What it reports is **agreement with the v0 reference raters, not accuracy** — the pooled rate spans the disclosed two-block frame, and the per-stream split is derived at measurement review ([qp-topic.md](qp-topic.md)). The publication gate takes two conditions, the agreement rate and how much of the reference set the run covered; failing either, the artifact is not written, the measurement prints, and the command exits non-zero. There is no override flag. `--frame-rows` is the QP-bearing frame this batch was cut from, which neither input holds — the extract is one batch of the frame and the artifact is the union of every batch — so the extract job passes its own count in as a plain integer, never the `.batch.json` sidecar (whose Term × fee-class shape stays off the labeling job). It lands on this batch's ledger entry, where the docket pack reads it to measure the labeled share of the frame instead of bounding it; omitted, the entry records none and the cut keeps its bound. A count under the rows the batch labeled is refused, since the batch is drawn from the frame. The `run-analytics` mode that dispatches it is in [pipeline.md](pipeline.md). | `--labels`, `--texts`, `--labeler`, `--out`, `--frame-rows` | | `tool-usage` | Roll every committed `retrieval_log.json` into an **offered-vs-called** report: which configured MCP tools were never called, which are used by some engines and not others, and call counts per tool / engine / actor. Reads `data/` only — no corpus, no network — so it runs offline and in the gate; the `run-analytics` mode that dispatches it is in [pipeline.md](pipeline.md). Counting contract: call names normalize to `.` (engines spell one MCP tool `mcp__x__y` or `mcp_x_y`); engine built-ins (shell, file IO, web search) are counted apart from what the manifest offers; a **code-mode** engine invokes everything from inside a freeform builtin call, and both idioms its program reaches through — the MCP manifest and the engine's own builtins, which is where such a program does most of its work — are lifted out of that source into rows of their own, so the freeform call contributes its own builtin row plus one row per lifted call *site the scan reached* (a site inside a loop is still one row); a lifted manifest row carries the same `mcp____` spelling a direct item would and normalizes into the offered denominator identically, while a lifted builtin row names the builtin and so is counted apart from that denominator, like any other builtin — gate on the MCP normalization wherever the question is manifest use, and read such an engine's raw call volume as counting the wrapper beside everything it wrapped; the offered denominator is each log's `mcp_tools` snapshot, falling back to the current manifest's advertised set for logs predating that field. A zero means **never called**, not useless, and the report says which. The same walk publishes four further cuts: per-engine **result observability** — a captured `result_digest` is the only evidence the answer side was recorded at all, and a null covers both an empty result and an uncaptured one, so the rate is honestly two-state and the per-tool dead-end rate is withheld for any engine that never captured a result rather than printed as 100%; cells and calls by **mode, role, and actor**; **calls beside cost** per cell, joined from each log's sibling `usage.json`, where a missing record degrades to a null cost and never to free; and **call volume against Brier**, joined to the gradings of each predicted cell and segmented by engine, mode, and forecast moment with the `n` beside every mean. That last block prints a grade, so it is scoped like the boards — blessed processes only unless `--all-versions` — and it publishes a correlation only per (mode, moment) population and only above the floor pre-declared in code as `tool_usage.TOOL_USAGE_CORRELATION_MIN_CELLS`; there is deliberately no pooled coefficient. What it may be read for: [metrics/README.md](../metrics/README.md). | `--out`, `--markdown-out`, `--all-versions` | -| `ops-report` | Roll pipeline health, **substance**, spend & cost, **agent signals**, data health, and open issues wearing a `run:*` fan-out label into the ops report Markdown (and optional JSON). Each section renders only once its feed exists: `--previous` backs the substance deltas, `--live-frontier` the watchlist readiness, `--corpus-validation` data health, `--trigger-issues` the stale fan-out labels (markers left behind, since nothing keys on a label and each stage derives its own backlog); `--digest-out` renders the weekly interrogative digest. What each section reports and how `run-ops` publishes it: [pipeline.md](pipeline.md). | `--runs`, `--json`, `--generated-at`, `--corpus-validation`, `--live-frontier`, `--previous`, `--digest-out`, `--data-health-out`, `--trigger-issues`, `--all-versions` | +| `ops-report` | Roll pipeline health, **substance**, spend & cost, **agent signals**, data health, and open issues wearing a `run:*` fan-out label into the ops report Markdown (and optional JSON). Each section renders only once its feed exists: `--previous` backs the substance deltas, `--live-frontier` the watchlist readiness, `--corpus-validation` data health, `--trigger-issues` the stale fan-out labels (markers left behind, since nothing keys on a label and each stage derives its own backlog); `--digest-out` renders the weekly performance digest. What each section reports and how `run-ops` publishes it: [pipeline.md](pipeline.md). | `--runs`, `--json`, `--generated-at`, `--corpus-validation`, `--live-frontier`, `--previous`, `--digest-out`, `--data-health-out`, `--trigger-issues`, `--all-versions` | | `daily-digest` | Render the **daily prediction-reading digest**: one predicted event with every predictor side by side — the case/event header (court, docket, kind, stage, moment, decision target, resolved-or-not, and the cells' shared mode and snapshot vintage where they agree, else a per-section note), each cell's probability (captioned by the event's stage, since `probability` is P(granted) at cert and interim and P(disturbed) at merits; an unrecorded stage on a petition/appeal event reads as cert, the same normalization the stratified boards apply) and claims, its `predicted_reasoning.md` and `reasoning.md` inline, its `flags.json` notes, and links to the committed cell paths. Reads the committed ledger under `data/` only — no corpus, no model call. Selection is stateless: the newest event no prior digest issue has featured, rotating to the least-recently-featured one when nothing new has landed, derived from the `` marker in prior bodies (`--prior-issues` reads them from a file, `--fetch-prior` from GitHub). The body is bounded — each document capped, then a clamp under GitHub's issue-body limit, with truncated documents linking the full file; with nothing to feature — or with the day already digested under `--post` — it writes no body and exits 0, so the written file is always what was published. `--post` opens it as a fresh `daily-digest` issue (a non-triggering label, created idempotently, once per **day** — the separate `` marker, so a rotated re-read still posts). What `run-ops` does with it: [pipeline.md](pipeline.md). | `--repo`, `--prior-issues`, `--fetch-prior`, `--post`, `--out`, `--title-out`, `--ref`, `--generated-at` | | `stamp-cell` | Stamp a cell's `prediction.json` / `evaluation.json` with its **process version** — and, on `--role predictor`, copy the provisioned `record/context.json` onto the prediction as its frozen `context`, comparing the cell's own `input_snapshot` against the provisioned snapshot file (both reduced to that file's day) and recording the answer in `context.snapshot_uptake`. A disagreement is recorded, never masked and never refused: it also appends a `warning` to the cell's own `predictions///flags.json`, the one agent-owned file this command writes to, so the run's collect job surfaces it (label + content digest of the prompt template and resolved registry config + `pipeline_sha`) — the harness's word, not the agent's. That note says which of the two misses it was, since `unread` covers both: a cell that named another day's file, and a cell that named none. It names the provisioned snapshot either way; on the second it adds the event-level `record/` path a cell resolving the record one directory too deep would have probed, and rules a provisioning outage out without a runner — an unprovisioned cell runs no engine, which leaves a path fault. Which arm a cell takes is inferred from the shape of the string it wrote, erring toward the weaker reading. The same role emits one `::warning::` for a cell carrying no `big_case_score`, saying which shape it was: an explicit null with the one-line `big_case_rationale` the prompt's null branch requires, or silence on both. The rationale is the whole of the difference — the stamp rewrites the record through the model, so an omitted score and a declared null land as the same bytes — and this only warns, since the schema keeps the field optional and nothing downstream imputes a null; it drops the `(predictor, case)` point from that predictor's `big_case` tau-b instead. `collect-plan` re-derives the same split over the run's records as its per-predictor stakes-read census. A post-agent step in both fan-out workflows, before `validate`; a missing artifact is a no-op, a registry/prompt inconsistency fails the cell. The evaluator role stamps every predictor's `evaluation.json`: its `prediction_run_id` (the run id of the scored prediction the stamp resolves — the graded-artifact anchor every scoring join reads first, with the predictor's latest prediction only as the fallback for records stamped before the field existed), and computes each one's `claim_scores` block (`pipeline.claims.score_claims`, over the committed prediction, outcome, and statpack) plus its `base_rate_salience_version` (from the recorded `base_rate_basis`: the scored prediction's frozen `context.salience_version` on the risk-set path, the live scorer's version on the terminal path) — both assigned unconditionally, so an evaluator-authored value never survives. It also writes `correct` on **every** stage, cert included (`pipeline.evaluate.is_correct` over the scored prediction's committed label and the outcome's, routed on the outcome: judgment at merits, disposition elsewhere), cleared to null where either artifact is missing: a label comparison needs no pooled baseline and so no band judgment, which is the whole of the cert exemption below, and it is the leaderboard's first rank key. On a **merits** or **interim** cell it also writes the whole skill record the same way and for the same reason: the `brier_score` recomputed as `(probability - actual_granted)**2` from the scored prediction's committed probability and the committed outcome (`pipeline.evaluate.brier_score`), the `segment_base_rate` (`pipeline.base_rates.merits_base_rate` at the grant Term; `interim_base_rate` at the scored prediction's frozen application Term), and the `brier_skill_score` derived from those two. Neither the Brier nor those pools is the evaluator's to record — the Brier is a formula over committed artifacts, the pools a Term-keyed ratio of published counts with no band to choose — and each is cleared where the harness cannot compute it. Stamping all three off one set of inputs is what makes the skill ratio verifiable rather than merely self-consistent. `base_rate_basis` is cleared with them, since neither pooled rate is a band product for a basis to name. The **cert** trio is left as recorded — which band population the rate is taken over is a judgment about the frozen band, and the leaderboard's coherence check is what holds that arithmetic to its record. A mispaired basis also fails the cell, after the remaining stamps are written — either half: a `risk_set` basis whose version does not resolve (the null is still written, as the record of what resolution produced, but a basis without its version half names a population nothing pins down), and a `terminal` basis where the scored prediction froze a `band` at all (the fallback taken where a risk-set pairing existed — with the band's version beside it, a well-formed rate read against the wrong population; without it, a moved band priced at the terminal rate). The correction is a re-derived evaluation whose rate, basis, and version come off one population together, or nulling `segment_base_rate`, `base_rate_basis`, and `brier_skill_score` together and re-stamping (which clears the version half); never a relabel of the basis under the number as written, which would pair one population's version with a rate read over the other's table. `validate`'s `base_rate_basis_carries_version` holds both halves over the ledger, so a failed cell reaches a maintainer through the run's draft PR rather than a merged one. `--regrade` writes everything above **except** `process_version` and the stamped `prediction_run_id` — both preserved, the version so an older process's judgment is not re-attributed, the identity so the recompute stays a recompute of the run the record graded even after the predictor re-ran the event — for a cell whose committed outcome was corrected after it was graded: the recomputed fields are functions of the committed artifacts, while the prose and judgment beside them were produced under the process the record already names, so re-resolving that version would attribute an older process's work to a newer pre-registration. It is an in-place *recompute* of harness-owned fields, not the re-grade the leaderboard's collapse counts — that sense mints a second `evaluation.json` from a new evaluator run, which is the route for a changed judgment and the wrong one for a corrected outcome. It requires a record that already carries a stamp (a never-stamped cell takes the ordinary stamp instead), and refuses `--role predictor` — a prediction carries no harness-graded field — and `--stamped-at` / `--pipeline-sha`, which set only the version it declines to write. The recomputed claim block and skill fields are pooled from the statpack committed *now*, so re-grade a whole cohort against one pack. Re-grade **every** evaluator on the event: `validate`'s `evaluation_correct_agrees` collapses to the latest runs and requires them to agree, so a partial re-grade fails the ledger, which is the check working (it reaches the `correct` bit only, and only where two or more evaluators graded the cell — see [metrics/README.md](../metrics/README.md)). One divergence `--regrade` cannot converge: gradings that straddle a re-prediction each preserve their own `prediction_run_id`, so the repair there is the ordinary re-stamp, taken promptly and with its stated cost — it re-resolves the identity to the latest prediction and rewrites `process_version`. Every `--regrade` refusal is judged before the first write, so none can leave an event half corrected, and each target's process scope (`frozen` / `alpha`) is echoed as it goes — the operation leaves no `superseded_gradings` trace, so that line is the record of a claimable cell moving. Exit codes split the two failure kinds: **2** is a usage error (a bad `--role`, an unparseable `--stamped-at`, `--regrade` with a flag or role it does not support), **1** is a state the command refuses to write over or leave standing — a `--regrade` target carrying no `process_version`; a `--regrade` that matched no artifact at all (unlike the ordinary stamp's no-op, since a hand-typed run id must not read as a correction that landed); a `--regrade` of a run a newer grading supersedes, which no surface reads; a `--regrade` of a cell whose evaluator-owned `brier_score` no longer reproduces against the corrected outcome, which would leave `correct` moved beside a trio scored against the superseded binary (null the trio plus `base_rate_basis` together, or commit a re-derivation, then re-grade); and the mispaired basis, which fails *after* the remaining cells are written. See [process-version.md](process-version.md), [outcome-decomposition.md](outcome-decomposition.md). | `--court`, `--docket`, `--event`, `--run-id`, `--role`, `--actor`, `--pipeline-sha`, `--stamped-at`, `--regrade` | | `process-digest` | Print an actor's process digest — the value a maintainer blesses into `FROZEN_PROCESS_DIGESTS` (setting `FROZEN_SINCE` beside it; the two move together) to freeze. Each blessed digest is recorded with its **bless moment** — the carrying promotion's merge time, the retroactivity boundary — which is a different moment from the counting instant; see [process-version.md](process-version.md). `--all` prints every enabled predictor and evaluator. | `--role`, `--actor`, `--all` | diff --git a/docs/pipeline.md b/docs/pipeline.md index b3438758c..cab189dca 100644 --- a/docs/pipeline.md +++ b/docs/pipeline.md @@ -112,25 +112,27 @@ unread backlog. Four blocks, in the order a reader needs them: -- **Health questions** — the fixed interrogative bullets (replay calibration, - forward cells scored, watchlist vs next conference, oldest stalled trigger, - spend vs budget): numbers as questions demanding a reaction. These carry the - ops report's un-vintaged framing, which is why the vintage rule below is - scoped to the two blocks that publish figures to quote. -- **Analytics state** — what the committed boards hold. An empty one names the - condition that empties it — which cells the frozen headline ranks, and how - many have reached it — rather than showing a bare zero, and an artifact that - has never landed reads differently from one that landed empty. Plus the - statpack's two headline rates — stated with each one's own denominator, and - with the plain statement that **neither anchors a scored cell**: a forward - cert cell is scored against its own band's strictly-prior-Term risk-set rate, - and the pooled band rate is a fit diagnostic for the ranking constant rather - than a scoring baseline ([salience.md](salience.md)). - **Produced this week** — cells landed by role and stage, how many events they covered, and the week's measured spend, all over one window and one set of - `usage.json` records; then the spend backstop's own (longer) window and how - much of its ceiling the trailing period has consumed. An unenforced ceiling - says so instead of reporting a fraction of a budget that does not exist. + `usage.json` records. The window's bounds are in the heading: the Monday tick + titles its issue for the ISO week that *starts* that morning while the census + covers the seven days before it, so without them the block would describe the + previous week under this week's heading. +- **Produced this month** — the same shape over the trailing window the ex-post + spend backstop is configured with, closing with that backstop's own verdict + and how much of its ceiling the period has consumed. The window is the + backstop's own precisely so the census and the verdict beside it cannot + describe different periods. An unenforced ceiling says so instead of reporting + a fraction of a budget that does not exist. +- **Produced this term** — the same shape again over the October Term to date, + from the 1 October that Term opened to the day the digest is generated. The + cutoff is pinned to that instant rather than counted back in days, or a + mid-morning render would cut the Term's own first morning out of its census. + It closes with the forward cells scored under the process in force — the Term + is the period that count is worth reading over, since a forward cell is minted + once at its event and never again. A frozen scope with nothing scored in + either stratum is named as the shakedown state rather than shown as a bare + zero. - **Backtest results** — the historical replay **per court**, with each court's own always-deny floor beside its accuracy and the pooled row labelled as the mixture it is (`granted` means cert on a SCOTUS row and a motion granted on a @@ -157,10 +159,12 @@ Four blocks, in the order a reader needs them: the number is in, because a caveat one bullet away does not travel when the line is quoted. A board with no entries says so and prints no floor. -**In the analytics and back-test blocks, every figure carries the vintage of the -artifact it came from.** None of those artifacts is refreshed on this schedule — -a board is byte-stable and a statpack moves only when the corpus does — so a -figure without its vintage would silently claim to be this week's. The vintage +**In the back-test block, every figure carries the vintage of the artifact it +came from.** None of those artifacts is refreshed on this schedule — a board is +byte-stable, and the cert back-test moves only when a maintainer dispatches one — +so a figure without its vintage would silently claim to be this week's. The +production blocks need no vintage: they are computed from the committed ledger at +render time. The vintage is the commit that last wrote the file, and a **shallow** checkout yields none: in a depth-1 clone the one grafted commit matches every path, so a pathspec'd `git log` would stamp every board with today's date — the exact misreading the diff --git a/src/fedcourtsai/cli.py b/src/fedcourtsai/cli.py index 8c081b671..2d692c4e5 100644 --- a/src/fedcourtsai/cli.py +++ b/src/fedcourtsai/cli.py @@ -185,6 +185,7 @@ DAILY_DIGEST_MARKER_LINES, WEEKLY_DIGEST_LABEL, WEEKLY_DIGEST_MARKER_LINES, + ProducedWindow, Vintaged, WeeklyAnalytics, WeeklyProduction, @@ -301,7 +302,6 @@ AgentFlags, AgentToolingFeedback, Backtest, - BigCaseBoard, CellFailure, CellMode, CertBacktest, @@ -309,7 +309,6 @@ CertBacktestDispatch, CertBacktestProvenance, ClaimScoreBlock, - ClaimScoreBoard, ConferenceBucket, CorpusValidation, DataHealth, @@ -7739,12 +7738,8 @@ def _vintaged[T: BaseModel](path: Path, model: type[T]) -> Vintaged[T]: def _weekly_analytics(metrics_root: Path) -> WeeklyAnalytics: - """The committed boards the weekly digest reports, each with its own vintage.""" + """The committed replay artifacts the weekly digest's back-test block reports.""" return WeeklyAnalytics( - leaderboard=_vintaged(metrics_root / "leaderboard.json", Leaderboard), - claim_scores=_vintaged(metrics_root / "claim-scores.json", ClaimScoreBoard), - big_cases=_vintaged(metrics_root / "big-cases.json", BigCaseBoard), - statpack=_vintaged(metrics_root / "statpack.json", StatPack), backtest=_vintaged(metrics_root / "backtest.json", Backtest), salience_replay=_vintaged(metrics_root / "salience-replay.json", SalienceReplay), # Never produced by the scheduled refresh — a real-engine replay spends @@ -7780,29 +7775,66 @@ def _parse_when(stamp: str) -> datetime: return parsed if parsed.tzinfo is not None else parsed.replace(tzinfo=UTC) -#: The window the weekly digest's production census and its spend figure share. -#: A week, because that is the period the digest covers; the spend *backstop* -#: keeps its own, longer window, which is why the two are reported separately -#: rather than one being derived from the other. +#: The window the digest's first production block covers. A week, because that is +#: the period the digest is opened for; the month block takes its own window from +#: the spend backstop's config instead, so the backstop's verdict and the census +#: printed beside it can never describe different periods. _WEEKLY_WINDOW_DAYS = 7 +def _produced_over( + usage: Sequence[ModelUsage], + title: str, + *, + window_days: int, + when: datetime, + since: datetime | None = None, + term: int | None = None, +) -> ProducedWindow: + """One production block's census and spend, over one window of the ledger.""" + return ProducedWindow( + title=title, + census=cell_census(usage, window_days=window_days, now=when, since=since), + spend_usd=spend_over(usage, window_days=window_days, now=when, since=since)[0], + window_start=(since or when - timedelta(days=window_days)).date(), + window_end=when.date(), + term=term, + ) + + def _weekly_production(data_root: Path, config_root: Path, when: datetime) -> WeeklyProduction: - """The week's cells and cost, plus the spend backstop's own verdict. + """The digest's three production windows, plus the spend backstop's verdict. - One walk of the ledger for all three figures: the census, the week's spend, - and the backstop's own longer window read the same records, so they cannot - disagree about what they cover and the growing tree is scanned once. + One walk of the ledger for every figure: the three censuses, their three + spend totals, and the backstop's own window all read the same records, so + they cannot disagree about what they cover and the growing tree is scanned + once. + + The Term window is pinned to the instant its Term opened rather than counted + back in days — an October Term opens at midnight on 1 October and the digest + renders mid-morning, so a day count would cut the Term's own first morning + out of its census. """ usage = iter_usage(data_root) - census = cell_census(usage, window_days=_WEEKLY_WINDOW_DAYS, now=when) - spent, _cells = spend_over(usage, window_days=_WEEKLY_WINDOW_DAYS, now=when) + spend_config = load_spend_config(config_root) + term = october_term_year(when.date()) + term_start = datetime(term, 10, 1, tzinfo=UTC) return WeeklyProduction( - census=census, - spend_usd=spent, - backstop=verdict_over(usage, load_spend_config(config_root), now=when), - window_start=(when - timedelta(days=_WEEKLY_WINDOW_DAYS)).date(), - window_end=when.date(), + week=_produced_over( + usage, "Produced this week", window_days=_WEEKLY_WINDOW_DAYS, when=when + ), + month=_produced_over( + usage, "Produced this month", window_days=spend_config.window_days, when=when + ), + term=_produced_over( + usage, + "Produced this term", + window_days=(when.date() - term_start.date()).days, + when=when, + since=term_start, + term=term, + ), + backstop=verdict_over(usage, spend_config, now=when), ) diff --git a/src/fedcourtsai/ops.py b/src/fedcourtsai/ops.py index 5c1e03609..23a1a18cd 100644 --- a/src/fedcourtsai/ops.py +++ b/src/fedcourtsai/ops.py @@ -38,17 +38,13 @@ AgentToolingFeedback, Backtest, BacktestEntry, - BigCaseBoard, CertBacktest, - ClaimScoreBoard, CostEstimate, DataHealth, Evaluation, FlagsDigest, FlagSeverity, ForwardClaimRecord, - FrozenProcessRecord, - Leaderboard, LeakageDigest, LeakageExclusionRecord, LiveFrontier, @@ -520,109 +516,6 @@ def render_substance(digest: SubstanceDigest) -> str: return "\n".join(lines) + "\n" -def _health_questions(report: OpsReport) -> list[str]: - """The digest's fixed interrogative bullets, with this week's answers. - - Deliberately short and interrogative — the numbers demand a reaction rather - than sit available for inspection, which is what the daily ops report they - are drawn from already does. Renders from whatever the report holds, with - explicit absences. - """ - substance = report.substance - lines: list[str] = [] - - # A frozen scope with no scored cells is the shakedown state (nothing blessed - # yet), not a stalled machine — so the "what is blocking?" framing below would - # misread. Detect it once and reframe those questions honestly. - frozen_shakedown = ( - substance is not None - and substance.process_scope == "frozen" - and substance.cells.evaluations_forward == 0 - and substance.cells.evaluations_retrospective == 0 - ) - - if substance is not None and substance.calibration.sample: - cal = substance.calibration - lift = ( - "lift unavailable (no base rate)" - if cal.lift_over_always_deny is None - else f"lift **{cal.lift_over_always_deny:+.1%}** over always-deny" - ) - skill = ( - "" - if cal.mean_brier_skill is None - else f", Brier skill **{cal.mean_brier_skill:+.3f}** vs the segment base rate" - ) - lines.append( - f"- **Replay calibration on {cal.sample} scored cell(s): {lift}{skill} — " - "do you believe it?**" - ) - elif frozen_shakedown: - lines.append( - "- **No frozen-process cells yet — the headline is scoped to the frozen " - "process; run `--all-versions` for the shakedown pool.**" - ) - else: - lines.append("- **No scored replay cells yet — what is blocking the first batch?**") - - if substance is not None: - c = substance.cells - weekly = ( - f"{c.evaluations_forward_delta:+d} this week, {c.evaluations_forward} total" - if c.evaluations_forward_delta is not None - else f"{c.evaluations_forward} total, no prior snapshot to diff" - ) - question = ( - "still shakedown, none frozen yet" - if frozen_shakedown - else "is the live frontier producing?" - ) - lines.append( - f"- **Forward cells scored ({substance.process_scope}): {weekly} — {question}**" - ) - frontier = substance.live_frontier - if frontier is not None and not frontier.skipped: - upcoming = ( - f"{frontier.next_conference_petitions} petition(s) distributed for " - f"**{frontier.next_conference}**" - if frontier.next_conference is not None - else "no upcoming conference on the calendar" - ) - lines.append( - f"- **Watchlist vs next conference: {upcoming}; documents on " - f"{frontier.documents_provisioned}/{frontier.watchlist} — ready?**" - ) - else: - lines.append("- **Watchlist vs next conference: no published snapshot — why not?**") - - if report.open_triggers: - oldest = report.open_triggers[0] - lines.append( - f"- **Oldest stale fan-out label: `{oldest.label}` " - f"({_age(oldest.created_at, report.generated_at)} old) — clear it?**" - ) - else: - lines.append("- **Stale fan-out labels: none.**") - - monthly = ( - "—" - if report.cost.estimated_monthly_usd is None - else f"${report.cost.estimated_monthly_usd:,.0f}/mo" - ) - # Name the model rate here, not just the all-in total: the cumulative figure - # next to a total that used to exclude it was the misreading this line invited. - model_rate = ( - "unrated" - if report.cost.model_monthly_usd is None - else f"~${report.cost.model_monthly_usd:,.0f}/mo" - ) - lines.append( - f"- **Spend vs budget: ${report.spend.estimated_cost_usd:,.2f} model spend cumulative " - f"({model_rate} while running), ~{monthly} projected all-in — within plan?**" - ) - return lines - - @dataclass(frozen=True) class Vintaged[T]: """A committed metrics artifact and the vintage the figures in it carry. @@ -642,7 +535,11 @@ class Vintaged[T]: @dataclass(frozen=True) class WeeklyAnalytics: - """The committed metrics artifacts the weekly digest reports, each vintaged. + """The committed replay artifacts the weekly digest's back-test block reports. + + Each is vintaged: none of them is refreshed on the digest's own schedule, so a + figure without the vintage of the artifact it came from silently claims to be + this week's. ``salience_version_in_force`` is the scorer the pipeline runs **today**, not the one the replay was produced under: a per-band figure means something only @@ -651,10 +548,6 @@ class WeeklyAnalytics: is looking at. """ - leaderboard: Vintaged[Leaderboard] - claim_scores: Vintaged[ClaimScoreBoard] - big_cases: Vintaged[BigCaseBoard] - statpack: Vintaged[StatPack] backtest: Vintaged[Backtest] salience_replay: Vintaged[SalienceReplay] cert_backtest: Vintaged[CertBacktest] @@ -662,21 +555,21 @@ class WeeklyAnalytics: @dataclass(frozen=True) -class WeeklyProduction: - """What the ledger recorded this week: cells produced, and what they cost. +class ProducedWindow: + """One window of production: the cells the ledger recorded, and what they cost. ``census`` and ``spend_usd`` are taken over the *same* window and the same - usage records, so the two numbers describe one set of cells. ``backstop`` is - the ex-post spend gate's own verdict over its own (longer) window, carried - whole rather than reduced to a fraction — an unenforced ceiling has no - fraction, and a digest that printed one anyway would invent a budget. + usage records, so the two numbers describe one set of cells. ``title`` is the + heading the window renders under. ``term`` is set on the October Term window + only, which names the Term it covers beside its bounds. """ + title: str census: RecentCells spend_usd: float - backstop: SpendVerdict | None window_start: date | None = None window_end: date | None = None + term: int | None = None @property def window(self) -> str: @@ -685,16 +578,44 @@ def window(self) -> str: "last 7d" alone is not recoverable by a reader: the Monday tick titles its issue for the ISO week that *starts* that morning while the census covers the seven days before it, so without the bounds the section silently - describes the previous week under this week's heading. + describes the previous week under this week's heading. The Term window + leads with the Term it names and the day that Term opened — the boundary a + reader checks first — and still carries the far bound and the day count, so + all three windows read the same way. """ if self.window_start is None or self.window_end is None: return f"last {self.census.window_days}d" + if self.term is not None: + return ( + f"OT{self.term}, since {self.window_start.isoformat()} — " + f"{self.census.window_days}d to {self.window_end.isoformat()}" + ) return ( f"{self.window_start.isoformat()} to {self.window_end.isoformat()}, " f"the {self.census.window_days}d before this digest" ) +@dataclass(frozen=True) +class WeeklyProduction: + """What the ledger recorded over the digest's three windows, and the spend gate. + + Three windows of one shape — the week, the trailing month, and the October + Term to date — because one week's count alone cannot distinguish a quiet week + inside a producing Term from a pipeline that has stopped. ``backstop`` is the + ex-post spend gate's own verdict, carried whole rather than reduced to a + fraction: an unenforced ceiling has no fraction, and a digest that printed one + anyway would invent a budget. It renders under ``month``, whose window is the + backstop's own, so the verdict and the census printed beside it cover exactly + the same period. + """ + + week: ProducedWindow + month: ProducedWindow + term: ProducedWindow + backstop: SpendVerdict | None = None + + def _cell(value: str) -> str: """One table cell's text, with any pipe escaped. @@ -717,212 +638,15 @@ def _sourced(filename: str, vintaged: Vintaged[object]) -> str: return f"`metrics/{filename}`, {vintage}" -def _leaderboard_lines(vintaged: Vintaged[Leaderboard], generated_at: str) -> list[str]: - """The standings, or the honest reason there are none.""" - board = vintaged.value - where = _sourced("leaderboard.json", vintaged) - if board is None: - return ["- **Leaderboard**: `metrics/leaderboard.json` has never landed."] - if not board.entries: - return [ - f"- **Leaderboard** ({where}): **no predictor ranked** — " - f"{_empty_headline_reason(board.frozen_process, generated_at)} " - f"{board.evaluations_total} evaluation(s) in the `{board.process_scope}` " - f"scope, {board.events_scored} event(s) scored; `--all-versions` is where " - "the shakedown pool shows." - ] - versions = ( - f" · banded by {', '.join(f'`{v}`' for v in board.salience_versions)}" - if board.salience_versions - else "" - ) - regrades = ( - f" · {board.superseded_gradings} superseded grading(s) collapsed away" - if board.superseded_gradings - else "" - ) - rows = [ - f"- **Leaderboard** ({where}): {board.predictors_ranked} predictor(s) over " - f"{board.events_scored} event(s), scope `{board.process_scope}`{versions}" - f"{regrades}. Rank is forward accuracy, then forward Brier.", - "", - # `evaluators` is the panel depth — how many judges scored the predictor — - # not a cell count. Labelling it "cells" would publish "2 cells over 37 - # events", which is not a thing the board says. - "| Rank | Predictor | Judges | Events |", - "| ---: | --- | ---: | ---: |", - ] - rows += [ - f"| {entry.rank} | {_cell(entry.predictor_id)} | {entry.evaluators} " - f"| {entry.events_scored} |" - for entry in board.entries - ] - return rows - - -def _empty_headline_reason(frozen: FrozenProcessRecord | None, generated_at: str) -> str: - """Why an empty frozen board is empty — which is two different states. - - A freeze instant in the *future* means the counting window has not opened: - the headline is empty by construction and no grading could have reached it - however well the pipeline ran. Once the instant has passed, an empty board - means the window is open and nothing has been graded into it — a fact about - production. Collapsing the two into "nothing has been graded yet" reports the - first as the second, which is the same bare-zero misreading the branch exists - to avoid. - """ - if frozen is None or frozen.since is None: - return "the board records no freeze instant, so nothing is in scope to rank." - since = frozen.since.date().isoformat() - now = parse_iso(generated_at) - if now is not None and _as_utc(now) < _as_utc(frozen.since): - return f"the frozen counting window opens **{since}**, so it is empty by construction." - return ( - f"the frozen counting window opened {since} and no stamped grading has " - "reached the ranked population." - ) - - -def _claim_score_lines(vintaged: Vintaged[ClaimScoreBoard]) -> list[str]: - """The claim-score board's state, empty or otherwise.""" - board = vintaged.value - where = _sourced("claim-scores.json", vintaged) - if board is None: - return ["- **Claim scores**: `metrics/claim-scores.json` has never landed."] - if not board.entries: - return [ - f"- **Claim scores** ({where}): **suppressed** — " - f"{board.cells_with_claims} cell(s) carry a claims block inside the " - f"`{board.process_scope}` scope and this surface's population, so no " - "coefficient is computed." - ] - return [ - f"- **Claim scores** ({where}): {len(board.entries)} predictor(s) over " - f"{board.cells_with_claims} cell(s) carrying claims." - ] - - -def _big_case_lines(vintaged: Vintaged[BigCaseBoard]) -> list[str]: - """The big-case board's denominators — and the carve-out that keeps them read right. - - No case is named. The digest is the surface most often quoted out of, and a - case beside a number reads as a finding about that case, which a stakes read - cannot support. The coverage split is the informative part anyway: it says how - many of the board's means rest on the full panel. - """ - board = vintaged.value - where = _sourced("big-cases.json", vintaged) - if board is None: - return ["- **Big-case board**: `metrics/big-cases.json` has never landed."] - if not board.rows: - # What "empty" means depends on the scope the board was built at, and the - # two are not the same report: version-blind it says the ledger carries no - # stakes read at all, while at `frozen` it is the honest "no - # frozen-process stakes reads yet" state and says nothing about the ledger. - reason = ( - "no committed prediction carries a stakes read" - if board.process_scope == "all" - else "no frozen-process stakes reads yet — the ledger may hold plenty, and " - "`--process-scope all` is the census" - ) - return [ - f"- **Big-case board** ({where}, `process_scope: {board.process_scope}`): " - f"empty — {reason}." - ] - # The coverage distribution rather than a "full panel" count: `predictors` is - # read off the cells, so one stray cell from a fourth engine — or a retired - # third — would move a derived count and report a coverage collapse that did - # not happen. The distribution says the same thing and cannot lie that way. - split = ", ".join(f"{entry.n}→{entry.cases}" for entry in board.coverage) - leakage = ( - f" {board.rows_with_leakage_flag} row(s) rest partly on a leakage-flagged read." - if board.rows_with_leakage_flag - else "" - ) - # The scope travels with the count because the two scopes rank different - # populations: a reader comparing this line across builds without it would - # read a selected hold-out as cases the panel stopped caring about. - held = ( - f" {board.cases_out_of_scope} case(s) held off by the scope." - if board.cases_out_of_scope - else "" - ) - return [ - f"- **Big-case board** ({where}, `process_scope: {board.process_scope}`): " - f"{board.cases} case(s) ranked over " - f"{board.scored_reads} scored stakes read(s) of {board.current_reads} from " - f"{len(board.predictors)} predictor(s); cases by scoring predictors {split}." - f"{held}{leakage} A panel opinion about stakes — neither scored nor ranked, and " - "not a statement about cert likelihood." - ] - - -def _base_rate_lines(vintaged: Vintaged[StatPack]) -> list[str]: - """The statpack's two headline rates, each with its own denominator and its limit. - - **Neither anchors a scored cell**, and the line says so. A forward cert cell - is scored against its own salience band's strictly-prior-Term risk-set rate; - the pack-wide band rate here pools every band over the whole walked range with - no own-Term exclusion, which ``docs/salience.md`` registers as a fit - diagnostic for the ranking constant and *not* a scoring baseline — quoting it - as a forecast anchor would breach the leakage guard registered there. The - always-deny figure is the whole modern-cert slice's, not the predicted - segment's, so it is an orientation for reading an accuracy, not a floor any - scored cell is measured against. - """ - pack = vintaged.value - where = _sourced("statpack.json", vintaged) - if pack is None: - return ["- **Base rates**: `metrics/statpack.json` has never landed."] - deny, deny_cases = _deny_base_rate(pack) - grant, grant_cases = _segment_base_rate(pack) - deny_text = ( - f"always-deny **{deny:.0%}** (est. over {deny_cases:,} resolved modern-cert " - "petitions, denial-reweighted)" - if deny is not None and deny_cases is not None - else "always-deny **—** (no cert-stage disposition section)" - ) - grant_text = ( - f"pooled salience-band grant **{grant:.0%}** (over {grant_cases:,} resolved " - "petitions of the scored segment, every band pooled)" - if grant is not None and grant_cases is not None - else "pooled band grant **—** (no salience-band section)" - ) - return [ - f"- **Base rates** ({where}): {deny_text}; {grant_text}. Pack coverage: " - f"{pack.coverage.live_slice_resolved:,} resolved of " - f"{pack.coverage.live_slice_rows:,} live-slice rows.", - " _Neither figure anchors a scored cell. A forward cert cell is scored " - + "against its own band's strictly-prior-Term risk-set rate; the pooled band " - + "rate is a fit diagnostic for the ranking constant, not a scoring baseline " - + "(`docs/salience.md`), and the always-deny rate is taken over a different, " - + "much wider population than the one the gate predicts — an orientation " - + "rather than an effect size._", - ] - - -def _render_analytics_state(analytics: WeeklyAnalytics, generated_at: str) -> list[str]: - """Section 1: what the committed boards say, and what they honestly cannot.""" - return [ - "", - "## Analytics state", - "", - *_leaderboard_lines(analytics.leaderboard, generated_at), - *_claim_score_lines(analytics.claim_scores), - *_big_case_lines(analytics.big_cases), - *_base_rate_lines(analytics.statpack), - ] - - -def _render_production(production: WeeklyProduction) -> list[str]: - """Section 2: cells produced this week, and the measured cost of producing them.""" - census = production.census +def _render_produced(window: ProducedWindow) -> list[str]: + """One production block: the cells that ran inside a window, and their cost.""" + census = window.census lines = [ "", - f"## Produced this week ({production.window})", + f"## {window.title} ({window.window})", "", f"**{census.cells}** cell(s) over **{census.events}** event(s), and " - f"**${production.spend_usd:,.2f}** of measured model spend over the same " + f"**${window.spend_usd:,.2f}** of measured model spend over the same " "records — the recorded `usage.json` ledger, which lags: a cell's usage " "reaches `data/` only when its run's collect PR merges, so this is a floor " "on what was spent, not a real-time figure.", @@ -936,30 +660,57 @@ def _render_production(production: WeeklyProduction) -> list[str]: ] else: lines += ["", "_No cell landed in the window._"] + return lines - backstop = production.backstop + +def _render_backstop(backstop: SpendVerdict | None) -> list[str]: + """The ex-post spend gate's verdict, closing the block whose window it shares.""" if backstop is None: - lines += ["", "_Spend backstop: not evaluated this run._"] - elif not backstop.enforced: - lines += [ + return ["", "_Spend backstop: not evaluated this run._"] + if not backstop.enforced: + return [ "", "_Spend backstop: **no ceiling configured**, so nothing is measured " + "against one and nothing would be deferred._", ] - else: - share = backstop.spent_usd / backstop.ceiling_usd if backstop.ceiling_usd else 0.0 - verdict = ( - "**BREACHED** — a plan seam would defer its matrix" if backstop.breached else "clear" - ) - lines += [ - "", - f"Spend backstop: **${backstop.spent_usd:,.2f}** of " - f"**${backstop.ceiling_usd:,.2f}** over the trailing " - f"{backstop.window_days}d window (**{share:.0%}** consumed, " - f"${backstop.remaining_usd:,.2f} left, {backstop.cells} cell(s)) — {verdict} " - "on a lagging ledger, so the verdict is a floor too.", - ] - return lines + share = backstop.spent_usd / backstop.ceiling_usd if backstop.ceiling_usd else 0.0 + verdict = "**BREACHED** — a plan seam would defer its matrix" if backstop.breached else "clear" + return [ + "", + f"Spend backstop: **${backstop.spent_usd:,.2f}** of " + f"**${backstop.ceiling_usd:,.2f}** over the trailing " + f"{backstop.window_days}d window (**{share:.0%}** consumed, " + f"${backstop.remaining_usd:,.2f} left, {backstop.cells} cell(s)) — {verdict} " + "on a lagging ledger, so the verdict is a floor too.", + ] + + +def _frozen_cells_line(substance: SubstanceDigest | None) -> list[str]: + """Forward cells scored under the process in force, closing the Term block. + + The Term is the period this count is worth reading over: a forward cell is + minted once at its event and never again, so a week's delta alone cannot say + whether the scored population is growing or standing still. Stated as a + figure — the boards are where a reader interrogates it. + """ + if substance is None: + return [] + cells = substance.cells + counted = ( + f"{cells.evaluations_forward_delta:+d} this week, {cells.evaluations_forward} total" + if cells.evaluations_forward_delta is not None + else f"{cells.evaluations_forward} total, no prior snapshot to diff" + ) + # A frozen scope with nothing scored in *either* stratum is the shakedown + # state — nothing blessed yet — not a stalled machine, and a bare zero reads + # as the latter. Name it. + shakedown = ( + substance.process_scope == "frozen" + and cells.evaluations_forward == 0 + and cells.evaluations_retrospective == 0 + ) + tail = " No frozen-process cells yet — still shakedown." if shakedown else "" + return ["", f"Forward cells scored ({substance.process_scope}): {counted}.{tail}"] #: The court whose rows the digest publishes beside the pooled figure. Pooling @@ -1272,7 +1023,7 @@ def _cert_backtest_lines(vintaged: Vintaged[CertBacktest]) -> list[str]: def _render_backtests(analytics: WeeklyAnalytics) -> list[str]: - """Section 3: what the replays say, and which replay has never been run.""" + """The last block: what the replays say, and which replay has never been run.""" return [ "", "## Backtest results", @@ -1344,25 +1095,26 @@ def render_weekly_digest( ) -> str: """The weekly performance digest: the week's substance, on its own issue. - Four blocks, in the order a reader needs them. The **health questions** are - short and interrogative — numbers that demand a reaction rather than sit - available for inspection. **Analytics state** says what the committed boards - hold, with each empty one explaining *why* it is empty rather than showing a - bare zero. **Produced this week** counts the cells that ran and what they - cost, over one window and one set of records. **Backtest results** reports - the replays, and says plainly that the cert back-test has never been run - rather than leaving its absence to be inferred from a missing line. - - In the analytics and back-test blocks every figure carries the vintage of - the artifact it came from, because none of those artifacts is refreshed on - this schedule: a board is byte-stable and a statpack moves only when the - corpus does, so a figure without its vintage silently claims to be this - week's. The health questions keep the ops report's un-vintaged framing — - they are the same bullets that surface there, read as questions rather than - as figures to quote. - - ``analytics`` and ``production`` are optional so the digest degrades to its - questions rather than failing when a feed is absent. + Four blocks, in the order a reader needs them. **Produced this week** counts + the cells that ran and what they cost, over one window and one set of + ``usage.json`` records. **Produced this month** is the same shape over the + trailing window the ex-post spend backstop uses, and closes with that + backstop's own verdict — the census and the verdict beside it then cover + exactly one period. **Produced this term** is the same shape again over the + October Term to date, closing with the forward cells scored under the process + in force, which is the period that count is worth reading over. **Backtest + results** reports the replays, and says plainly that the cert back-test has + never been run rather than leaving its absence to be inferred from a missing + line. + + Every figure in the back-test block carries the vintage of the artifact it + came from, because none of those artifacts is refreshed on this schedule: a + board is byte-stable, so a figure without its vintage silently claims to be + this week's. The production blocks need none — they are computed from the + committed ledger at render time. + + ``analytics`` and ``production`` are optional so the digest degrades to the + blocks it can fill rather than failing when a feed is absent. """ lines = [ weekly_digest_marker(report.generated_at), @@ -1370,22 +1122,21 @@ def render_weekly_digest( "", f"_Generated {report.generated_at}. Close this issue once you have read it — the " "open `weekly-digest` issues are the unread backlog._", - "", - "## Health questions", - "", - *_health_questions(report), ] - if analytics is not None: - lines += _render_analytics_state(analytics, report.generated_at) if production is not None: - lines += _render_production(production) + lines += _render_produced(production.week) + lines += _render_produced(production.month) + lines += _render_backstop(production.backstop) + lines += _render_produced(production.term) + lines += _frozen_cells_line(report.substance) if analytics is not None: lines += _render_backtests(analytics) # The marker line stays verbatim; everything below it is defused in one pass, # exactly as the daily digest's body is. Almost all of this document is - # harness-computed figures, but the predictor ids threaded through the board - # tables are free-form strings from the ledger, and a field added later would - # otherwise arrive untreated. + # harness-computed figures, but the role and stage labels threaded through the + # census tables — and the predictor ids in the back-test ones — are free-form + # strings off the ledger, and a field added later would otherwise arrive + # untreated. prose = _defuse_comments("\n".join(lines[WEEKLY_DIGEST_MARKER_LINES:])) document = "\n".join([*lines[:WEEKLY_DIGEST_MARKER_LINES], prose]) + "\n" if len(document) > _DIGEST_MAX_CHARS: diff --git a/src/fedcourtsai/spend.py b/src/fedcourtsai/spend.py index 37eeddba3..cb2aacb71 100644 --- a/src/fedcourtsai/spend.py +++ b/src/fedcourtsai/spend.py @@ -86,17 +86,29 @@ def remaining_usd(self) -> float: def spend_over( - usage: Iterable[ModelUsage], *, window_days: int, now: datetime | None = None + usage: Iterable[ModelUsage], + *, + window_days: int, + now: datetime | None = None, + since: datetime | None = None, ) -> tuple[float, int]: - """Estimated cost and cell count among ``usage`` inside the trailing window. + """Estimated cost and cell count among ``usage`` inside the window. The pure half of :func:`trailing_spend`, for a caller that has already read - the ledger and would otherwise walk it again — the weekly digest prices two - different windows and counts a census over the same records, and three walks - of a growing tree for one report is three times the work and one more chance - for the three figures to disagree about what they cover. + the ledger and would otherwise walk it again — the weekly digest prices + several windows and counts a census over the same records, and a walk per + figure over a growing tree is that many times the work and that many more + chances for the figures to disagree about what they cover. + + ``since`` pins the cutoff to an instant instead of deriving it from + ``window_days``, for a window whose start is a calendar boundary rather than + a count of days back: an October Term opens at midnight on 1 October, and a + cutoff taken as *N* days before a mid-morning digest would drop the Term's + own first morning. ``window_days`` still names the span the caller reports. """ - cutoff = (now or datetime.now(UTC)) - timedelta(days=window_days) + cutoff = ( + since if since is not None else (now or datetime.now(UTC)) - timedelta(days=window_days) + ) total = 0.0 cells = 0 for record in usage: diff --git a/src/fedcourtsai/store.py b/src/fedcourtsai/store.py index 2a8430a88..cd268c768 100644 --- a/src/fedcourtsai/store.py +++ b/src/fedcourtsai/store.py @@ -1375,18 +1375,24 @@ def recent_cell_census( def cell_census( - usage: Iterable[ModelUsage], *, window_days: int, now: datetime | None = None + usage: Iterable[ModelUsage], + *, + window_days: int, + now: datetime | None = None, + since: datetime | None = None, ) -> RecentCells: - """The cells among ``usage`` inside the trailing window, by role and stage. + """The cells among ``usage`` inside the window, by role and stage. Applies the same cutoff rule :func:`fedcourtsai.spend.spend_over` applies to - the very same records (a naive ``created_at`` reads as UTC), so the count and - the cost a digest reports describe exactly the same set of cells. The stage - comes off the moment register + the very same records (a naive ``created_at`` reads as UTC), ``since`` + included, so the count and the cost a digest reports describe exactly the + same set of cells. The stage comes off the moment register (:func:`fedcourtsai.pipeline.moments.spec_for`) rather than the ledger, since a usage record names its event but not the standard governing it. """ - cutoff = (now or datetime.now(UTC)) - timedelta(days=window_days) + cutoff = ( + since if since is not None else (now or datetime.now(UTC)) - timedelta(days=window_days) + ) counts: Counter[tuple[str, str]] = Counter() events: set[tuple[str, str]] = set() for record in usage: diff --git a/tests/test_ops.py b/tests/test_ops.py index 317b57113..0a3e56422 100644 --- a/tests/test_ops.py +++ b/tests/test_ops.py @@ -1,7 +1,7 @@ import json import re from collections.abc import Sequence -from datetime import UTC, date, datetime +from datetime import UTC, date, datetime, timedelta from pathlib import Path import pytest @@ -20,9 +20,6 @@ BacktestCourtScore, BacktestEntry, BaseRateBucket, - BigCaseBoard, - BigCaseCoverage, - BigCaseRow, CalibrationBin, CertBacktest, CertBacktestCellLoss, @@ -30,7 +27,6 @@ CertBacktestEntry, CertBacktestProvenance, ClaimProbability, - ClaimScoreBoard, ConferenceBucket, CorpusCheck, CorpusValidation, @@ -42,10 +38,7 @@ EventKind, FlagCategory, FlagSeverity, - FrozenProcessRecord, GroupBy, - Leaderboard, - LeaderboardEntry, LeakageAssessment, LedgerValidation, LiveFrontier, @@ -357,12 +350,8 @@ def test_render_surfaces_the_model_rate_not_just_the_cumulative_total() -> None: assert "$60.00 cumulative over 6.0d of ledger" in body assert f"Run-rate **~${all_in}/mo** projected" in body - digest = ops.render_weekly_digest(report) - assert "(~$300/mo while running)" in digest - assert f"~${all_in}/mo projected all-in" in digest - -def test_render_digest_says_unrated_rather_than_implying_zero_spend() -> None: +def test_render_markdown_says_unrated_rather_than_implying_zero_spend() -> None: """The None branch must read as 'not computed', never as a small number.""" report = ops.build_ops_report( generated_at="2026-06-26T12:00:00+00:00", @@ -371,9 +360,8 @@ def test_render_digest_says_unrated_rather_than_implying_zero_spend() -> None: ) assert report.cost.estimated_monthly_usd is None - digest = ops.render_weekly_digest(report) - assert "$500.00 model spend cumulative (unrated while running)" in digest - assert "~— projected all-in" in digest + body = ops.render_markdown(report) + assert "Run-rate **~—** projected · model —" in body def test_render_markdown_smoke() -> None: @@ -1364,93 +1352,6 @@ def test_render_markdown_includes_substance_when_present() -> None: # --- the weekly digest ---------------------------------------------------------- -def test_render_weekly_digest_asks_the_fixed_questions() -> None: - report = ops.build_ops_report( - generated_at="2026-07-11T00:00:00+00:00", - runs=[], - usage=[_usage("a", 1.5)], - substance=ops.summarize_substance( - cell_counts=(6, 4, 3), - stratified_evaluations=[ - (_evaluation("p", correct=1), "retrospective"), - (_evaluation("p", correct=1, run_id="20260702T000000Z"), "forward"), - ], - statpack=_statpack_with_cert_section(denied=95, granted=5), - live_frontier=LiveFrontier( - generated_on=date(2026, 7, 11), - watchlist=40, - next_conference=date(2026, 9, 29), - next_conference_petitions=35, - documents_provisioned=28, - ), - ), - open_triggers=ops.summarize_trigger_issues( - [ - { - "number": 9, - "title": "evaluate: 1 case(s)", - "labels": [{"name": "run:evaluate"}], - "createdAt": "2026-07-08T09:00:00Z", - } - ] - ), - ) - md = ops.render_weekly_digest(report) - assert md.startswith("") - assert "# Weekly performance digest" in md - assert "## Health questions" in md - assert "Replay calibration on 1 scored cell(s)" in md and "do you believe it?" in md - assert "Forward cells scored (frozen): 1 total, no prior snapshot to diff" in md - assert "35 petition(s) distributed for **2026-09-29**" in md and "28/40" in md - assert "Oldest stale fan-out label: `run:evaluate` (2d old) — clear it?" in md - assert "Spend vs budget: $1.50" in md - - -def test_weekly_digest_reports_the_segment_brier_skill_when_present() -> None: - report = ops.build_ops_report( - generated_at="2026-07-11T00:00:00+00:00", - runs=[], - usage=[], - substance=ops.summarize_substance( - cell_counts=(2, 1, 2), - stratified_evaluations=[ - (_evaluation("p", correct=1, brier_skill=0.3), "retrospective") - ], - statpack=_statpack_with_salience_section({"high": (30, 70)}), - ), - ) - md = ops.render_weekly_digest(report) - assert "Brier skill **+0.300** vs the segment base rate" in md - - -def test_render_weekly_digest_all_absent_still_asks() -> None: - report = ops.build_ops_report(generated_at="2026-07-11T00:00:00+00:00", runs=[], usage=[]) - md = ops.render_weekly_digest(report) - assert "No scored replay cells yet" in md - assert "Stale fan-out labels: none" in md - assert "within plan?" in md - - -def test_weekly_digest_reframes_the_shakedown_state_honestly() -> None: - """The frozen-empty shakedown must not read as a stalled machine: the digest's - 'what is blocking?' / 'is the frontier producing?' questions would mislead when - the answer is just 'nothing frozen yet'.""" - report = ops.build_ops_report( - generated_at="2026-07-11T00:00:00+00:00", - runs=[], - usage=[], - # Predictions committed (version-blind census), zero frozen evaluations. - substance=ops.summarize_substance( - cell_counts=(410, 137, 5), stratified_evaluations=[], process_scope="frozen" - ), - ) - md = ops.render_weekly_digest(report) - assert "No frozen-process cells yet" in md - assert "what is blocking the first batch" not in md - assert "still shakedown, none frozen yet" in md - assert "is the live frontier producing?" not in md - - # --- lenient prior snapshots ------------------------------------------------------ @@ -1517,7 +1418,7 @@ def test_ops_report_writes_the_digest_and_reads_the_frontier(tmp_path: Path) -> body = digest_out.read_text() assert body.startswith("") + assert "# Weekly performance digest" in bare + for heading in ("## Produced this", "## Backtest results"): + assert heading not in bare + + backtests_only = ops.render_weekly_digest(_empty_report(), analytics=_analytics()) - assert "## Health questions" in md - assert "## Analytics state" not in md - assert "## Produced this week" not in md - assert "## Backtest results" not in md + assert "## Backtest results" in backtests_only + assert "## Produced this" not in backtests_only + + +def _commit_usage(data_root: Path, *, docket: int, cost: float, created_at: datetime) -> None: + """Commit one `usage.json` at the ledger path `iter_usage` globs.""" + run_id = created_at.strftime("%Y%m%dT%H%M%SZ") + write_json( + data_root + / "cases" + / "scotus" + / str(docket) + / "events" + / "evt-petition-disposition" + / "predictions" + / "claude-baseline" + / run_id + / "usage.json", + ModelUsage( + case_id=f"scotus/{docket}", + event_id="evt-petition-disposition", + run_id=run_id, + role=UsageRole.predictor, + actor_id="claude-baseline", + engine=Engine.claude_code, + model="claude-fable-5", + created_at=created_at, + input_tokens=1_000, + output_tokens=100, + estimated_cost_usd=cost, + ), + ) + + +@pytest.mark.parametrize( + ("when", "term", "start", "days"), + [ + # The pivot is the calendar month, so a late-September digest still + # reports the outgoing Term — the long-conference cohort's own Term. + ("2026-09-21T08:30:00+00:00", 2025, "2025-10-01", 355), + ("2026-10-05T08:30:00+00:00", 2026, "2026-10-01", 4), + ], +) +def test_the_term_block_rolls_on_the_first_of_october( + tmp_path: Path, when: str, term: int, start: str, days: int +) -> None: + production = cli._weekly_production( + tmp_path / "data", tmp_path / "config", datetime.fromisoformat(when) + ) + + assert production.term.term == term + assert production.term.window_start == date.fromisoformat(start) + assert production.term.census.window_days == days + assert production.term.window == f"OT{term}, since {start} — {days}d to {when[:10]}" + + +def test_the_term_census_keeps_the_terms_own_first_morning(tmp_path: Path) -> None: + # The cutoff is the instant the Term opened, not a count of days back from a + # mid-morning render — which would drop 1 October before 08:30 out of the + # Term whose census it belongs in, and pull in 30 September instead. + _commit_usage( + tmp_path / "data", docket=1, cost=3.0, created_at=datetime(2025, 10, 1, 2, 0, tzinfo=UTC) + ) + _commit_usage( + tmp_path / "data", docket=2, cost=5.0, created_at=datetime(2025, 9, 30, 23, 0, tzinfo=UTC) + ) + + production = cli._weekly_production( + tmp_path / "data", tmp_path / "config", datetime(2026, 9, 21, 8, 30, tzinfo=UTC) + ) + + assert production.term.census.cells == 1 + assert production.term.spend_usd == 3.0 + + +def test_the_month_window_is_the_spend_backstops_own(tmp_path: Path) -> None: + # The backstop's verdict renders at the end of the month block, so a month + # window derived independently would print a verdict over one period beside a + # census over another. + config_root = tmp_path / "config" + config_root.mkdir() + (config_root / "tracking.yaml").write_text("spend:\n ceiling_usd: 100.0\n window_days: 21\n") + + production = cli._weekly_production( + tmp_path / "data", config_root, datetime(2026, 9, 21, 8, 30, tzinfo=UTC) + ) + + assert production.month.census.window_days == 21 + assert production.backstop is not None + assert production.backstop.window_days == 21 + assert production.week.census.window_days == 7 + + +def test_the_three_windows_are_counted_over_one_walk_of_the_ledger(tmp_path: Path) -> None: + # One cell inside each nested window: the week's count must not leak into a + # narrower one, and every window must see what falls inside it. + now = datetime(2026, 9, 21, 8, 30, tzinfo=UTC) + _commit_usage(tmp_path / "data", docket=1, cost=1.0, created_at=now - timedelta(days=2)) + _commit_usage(tmp_path / "data", docket=2, cost=2.0, created_at=now - timedelta(days=20)) + _commit_usage(tmp_path / "data", docket=3, cost=4.0, created_at=now - timedelta(days=200)) + _commit_usage(tmp_path / "data", docket=4, cost=8.0, created_at=now - timedelta(days=400)) + + production = cli._weekly_production(tmp_path / "data", tmp_path / "config", now) + + assert (production.week.census.cells, production.week.spend_usd) == (1, 1.0) + assert (production.month.census.cells, production.month.spend_usd) == (2, 3.0) + assert (production.term.census.cells, production.term.spend_usd) == (3, 7.0) + + +def test_the_term_block_closes_with_the_forward_cells_scored() -> None: + # A forward cell is minted once at its event and never again, so the Term is + # the period the count is worth reading over — and it is a figure, not a + # question: the boards are where a reader interrogates it. + report = ops.build_ops_report( + generated_at="2026-09-21T08:30:00+00:00", + runs=[], + usage=[], + substance=ops.summarize_substance( + cell_counts=(6, 4, 3), + stratified_evaluations=[ + (_evaluation("p", correct=1), "retrospective"), + (_evaluation("p", correct=1, run_id="20260702T000000Z"), "forward"), + ], + ), + ) + + lines = ops.render_weekly_digest(report, production=_production()).splitlines() + + figure = lines.index("Forward cells scored (frozen): 1 total, no prior snapshot to diff.") + assert lines.index("## Produced this term (last 355d)") < figure + assert not any(line.startswith("## ") for line in lines[figure:]) + + +def test_the_forward_cell_figure_carries_the_week_over_week_delta() -> None: + prior = ops.build_ops_report( + generated_at="2026-09-14T08:30:00+00:00", + runs=[], + usage=[], + substance=ops.summarize_substance( + cell_counts=(6, 4, 3), + stratified_evaluations=[(_evaluation("p", correct=1, run_id="a"), "forward")], + ), + ) + report = ops.build_ops_report( + generated_at="2026-09-21T08:30:00+00:00", + runs=[], + usage=[], + substance=ops.summarize_substance( + cell_counts=(9, 6, 4), + stratified_evaluations=[ + (_evaluation("p", correct=1, run_id="a"), "forward"), + (_evaluation("p", correct=1, run_id="b"), "forward"), + (_evaluation("p", correct=1, run_id="c"), "forward"), + ], + previous=prior, + ), + ) + + md = ops.render_weekly_digest(report, production=_production()) + + assert "Forward cells scored (frozen): +2 this week, 3 total." in md + + +def test_the_term_block_names_the_shakedown_state_rather_than_a_bare_zero() -> None: + # A frozen scope with nothing scored in either stratum is the registered + # shakedown state — nothing blessed yet — and a bare zero reads as a stalled + # machine instead. + report = ops.build_ops_report( + generated_at="2026-09-21T08:30:00+00:00", + runs=[], + usage=[], + # Predictions committed (the census is version-blind), zero frozen evaluations. + substance=ops.summarize_substance( + cell_counts=(410, 137, 5), stratified_evaluations=[], process_scope="frozen" + ), + ) + + md = ops.render_weekly_digest(report, production=_production()) + + assert ( + "Forward cells scored (frozen): 0 total, no prior snapshot to diff. " + "No frozen-process cells yet — still shakedown." in md + ) + + +def test_an_all_versions_scope_is_not_read_as_the_shakedown_state() -> None: + report = ops.build_ops_report( + generated_at="2026-09-21T08:30:00+00:00", + runs=[], + usage=[], + substance=ops.summarize_substance( + cell_counts=(410, 137, 5), stratified_evaluations=[], process_scope="all" + ), + ) + + md = ops.render_weekly_digest(report, production=_production()) + + assert "Forward cells scored (all): 0 total, no prior snapshot to diff." in md + assert "still shakedown" not in md + + +def test_the_term_block_omits_the_figure_without_a_substance_section() -> None: + md = ops.render_weekly_digest(_empty_report(), production=_production()) + + assert "Forward cells scored" not in md def test_the_weekly_digest_marker_is_one_per_iso_week() -> None: @@ -2987,10 +3020,10 @@ def test_the_weekly_digest_marker_is_one_per_iso_week() -> None: def test_ops_report_renders_every_weekly_section_over_the_committed_tree(tmp_path: Path) -> None: - # The acceptance dry-run, against the repo's own `metrics/` and `data/`: all - # three sections, the honest-empty leaderboard state, the missing - # cert-backtest line, and a vintage beside every metrics-derived figure. - if not Path("metrics/leaderboard.json").exists(): # pragma: no cover - it is committed + # The acceptance dry-run, against the repo's own `metrics/` and `data/`: every + # block, the missing cert-backtest line, and a vintage beside every + # metrics-derived figure. + if not Path("metrics/backtest.json").exists(): # pragma: no cover - it is committed pytest.skip("no committed metrics in this checkout") out = tmp_path / "digest.md" @@ -3003,17 +3036,17 @@ def test_ops_report_renders_every_weekly_section_over_the_committed_tree(tmp_pat md = out.read_text() assert md.startswith("") for heading in ( - "## Health questions", - "## Analytics state", - "## Produced this week", + "## Produced this week (2026-08-26 to 2026-09-02, the 7d before this digest)", + "## Produced this month (", + "## Produced this term (OT2025, since 2025-10-01 — ", "## Backtest results", ): assert heading in md - # Structural only. Asserting the *current* contents of `metrics/` — the empty - # leaderboard, the absent cert back-test — would turn each of those milestones - # into a red required check on `main` the day it is reached; the honest-empty - # and missing-report branches are pinned synthetically above instead. - for artifact in ("leaderboard.json", "claim-scores.json", "statpack.json", "backtest.json"): + # Structural only. Asserting the *current* contents of `metrics/` — the absent + # cert back-test — would turn that milestone into a red required check on + # `main` the day it is reached; the missing-report branch is pinned + # synthetically above instead. + for artifact in ("backtest.json", "salience-replay.json"): assert f"`metrics/{artifact}`, " in md assert len(md) <= 60_000 @@ -3119,47 +3152,6 @@ def test_artifact_vintage_of_an_absent_file_is_unknown(tmp_path: Path) -> None: assert cli._artifact_vintage(tmp_path / "never-landed.json") is None -def test_the_weekly_digest_says_the_frozen_window_has_not_opened_yet() -> None: - # An empty board has two causes and only one of them is about production. A - # freeze instant in the future means the counting window has not opened, so - # no grading could have reached it however well the pipeline ran; reporting - # that as "nothing has been graded" is the bare-zero misreading again. - board = Leaderboard( - process_scope="frozen", - predictors_ranked=0, - evaluations_total=0, - events_scored=0, - frozen_process=FrozenProcessRecord( - since=datetime(2026, 9, 5, tzinfo=UTC), digests=["sha256:abc"] - ), - ) - before = ops.render_weekly_digest( - _empty_report("2026-09-02T08:30:00+00:00"), analytics=_analytics(leaderboard=board) - ) - after = ops.render_weekly_digest( - _empty_report("2026-09-12T08:30:00+00:00"), analytics=_analytics(leaderboard=board) - ) - - assert "the frozen counting window opens **2026-09-05**" in before - assert "empty by construction" in before - assert "the frozen counting window opened 2026-09-05 and no stamped grading" in after - - -def test_the_weekly_digest_refuses_to_anchor_a_score_on_a_pooled_base_rate() -> None: - # docs/salience.md registers the pooled band rate as a fit diagnostic for the - # ranking constant, not a scoring baseline: quoting one as a forecast anchor - # would breach the leakage guard registered there. - md = ops.render_weekly_digest( - _empty_report(), - analytics=_analytics(statpack=_statpack_with_cert_section(denied=95, granted=5)), - ) - - assert "Neither figure anchors a scored cell" in md - assert "strictly-prior-Term risk-set rate" in md - assert "orientation" in md - assert "anchors every score" not in md - - def test_the_weekly_digest_publishes_the_predicted_courts_row_not_only_the_pooled_one() -> None: # A pooled lift can be bought entirely on a docket this pipeline never # predicts: the SCOTUS floor is ~82% and the appellate floors near zero, so @@ -3241,39 +3233,60 @@ def test_the_weekly_digest_dates_the_window_it_counted() -> None: # it, so without bounds the section describes the previous week under this # week's heading. production = ops.WeeklyProduction( - census=RecentCells(rows=[], cells=0, events=0, window_days=7), - spend_usd=0.0, - backstop=None, - window_start=date(2026, 8, 26), - window_end=date(2026, 9, 2), + week=_produced( + "Produced this week", + rows=[], + cells=11, + events=4, + spend=12.5, + start=date(2026, 8, 26), + end=date(2026, 9, 2), + ), + month=_produced( + "Produced this month", + rows=[], + cells=48, + events=19, + spend=61.25, + window_days=30, + start=date(2026, 8, 3), + end=date(2026, 9, 2), + ), + term=_produced( + "Produced this term", + rows=[], + cells=402, + events=173, + spend=910.4, + window_days=336, + start=date(2025, 10, 1), + end=date(2026, 9, 2), + term=2025, + ), ) md = ops.render_weekly_digest(_empty_report(), production=production) - assert "## Produced this week (2026-08-26 to 2026-09-02, the 7d before this digest)" in md + assert [line for line in md.splitlines() if line.startswith("## ")] == [ + "## Produced this week (2026-08-26 to 2026-09-02, the 7d before this digest)", + "## Produced this month (2026-08-03 to 2026-09-02, the 30d before this digest)", + "## Produced this term (OT2025, since 2025-10-01 — 336d to 2026-09-02)", + ] + # Each block prices its own window over its own records. + assert "**11** cell(s) over **4** event(s), and **$12.50**" in md + assert "**48** cell(s) over **19** event(s), and **$61.25**" in md + assert "**402** cell(s) over **173** event(s), and **$910.40**" in md def test_the_weekly_digest_body_is_clamped_under_the_issue_size_limit() -> None: - # Bounded by construction, but the boards it renders grow with the roster, so - # the clamp is what makes the bound a guarantee — and the marker has to - # survive it or the week's idempotency key is gone. - board = Leaderboard( - process_scope="all", - predictors_ranked=4_000, - evaluations_total=4_000, - events_scored=1, - entries=[ - LeaderboardEntry( - predictor_id=f"predictor-{index:05d}", - rank=index + 1, - evaluators=1, - events_scored=1, - ) - for index in range(4_000) - ], - ) + # Bounded by construction, but the census tables grow with the roster, so the + # clamp is what makes the bound a guarantee — and the marker has to survive it + # or the week's idempotency key is gone. + rows = [CellCensusRow(f"predictor-{index:05d}", "cert", 1) for index in range(4_000)] - md = ops.render_weekly_digest(_empty_report(), analytics=_analytics(leaderboard=board)) + md = ops.render_weekly_digest( + _empty_report(), production=_production(rows, cells=4_000, events=4_000) + ) assert len(md) <= 60_000 assert md.startswith("") @@ -3281,24 +3294,11 @@ def test_the_weekly_digest_body_is_clamped_under_the_issue_size_limit() -> None: def test_the_weekly_digest_defuses_a_marker_quoted_by_the_ledger() -> None: - # Predictor ids are free-form strings off the ledger; the same one-pass - # defusing the daily digest applies keeps one from forging a marker. - board = Leaderboard( - process_scope="all", - predictors_ranked=1, - evaluations_total=1, - events_scored=1, - entries=[ - LeaderboardEntry( - predictor_id="", - rank=1, - evaluators=1, - events_scored=1, - ), - ], - ) + # Role and stage labels come off the ledger as free-form strings; the same + # one-pass defusing the daily digest applies keeps one from forging a marker. + rows = [CellCensusRow("", "cert", 1)] - md = ops.render_weekly_digest(_empty_report(), analytics=_analytics(leaderboard=board)) + md = ops.render_weekly_digest(_empty_report(), production=_production(rows, cells=1, events=1)) assert "" not in md assert "<!-- weekly-digest: 2026-W40 -->" in md diff --git a/tests/test_spend.py b/tests/test_spend.py index 9e19c4e9b..c42f6aa67 100644 --- a/tests/test_spend.py +++ b/tests/test_spend.py @@ -15,7 +15,8 @@ from fedcourtsai.config import SpendConfig from fedcourtsai.schemas import Engine, ModelUsage, UsageRole from fedcourtsai.serialize import write_json -from fedcourtsai.spend import check_spend, trailing_spend +from fedcourtsai.spend import check_spend, spend_over, trailing_spend +from fedcourtsai.store import iter_usage NOW = datetime(2026, 7, 27, 12, 0, tzinfo=UTC) @@ -134,3 +135,23 @@ def test_a_naive_created_at_is_read_as_utc(tmp_path: Path) -> None: ) # naive on purpose spent, cells = trailing_spend(tmp_path, window_days=30, now=NOW) assert (spent, cells) == (5.0, 1) + + +def test_a_pinned_since_overrides_the_day_count_cutoff(tmp_path: Path) -> None: + """A window whose start is a calendar boundary cannot be counted back in days. + + An October Term opens at midnight on 1 October while the weekly digest renders + mid-morning, so a cutoff taken as N days before the render lands at 08:30 on + 1 October and drops the Term's own first morning out of its own census. + """ + _usage(tmp_path, docket=1, cost=3.00, created_at=datetime(2025, 10, 1, 2, 0, tzinfo=UTC)) + _usage(tmp_path, docket=2, cost=5.00, created_at=datetime(2025, 9, 30, 23, 0, tzinfo=UTC)) + now = datetime(2026, 9, 21, 8, 30, tzinfo=UTC) + + pinned = spend_over( + iter_usage(tmp_path), window_days=355, now=now, since=datetime(2025, 10, 1, tzinfo=UTC) + ) + counted = spend_over(iter_usage(tmp_path), window_days=355, now=now) + + assert pinned == (3.00, 1) + assert counted == (0.0, 0) From 646645464076cf78c8e2a699b17ad43ab71a9479 Mon Sep 17 00:00:00 2001 From: ModelMirror <273825391+modelmirror@users.noreply.github.com> Date: Mon, 21 Sep 2026 20:15:54 +0000 Subject: [PATCH 2/2] feat(ops): restructure the weekly digest around three production windows The weekly performance digest now reads: produced this week, produced this month (the spend backstop's own trailing window, closing with its verdict), produced this term (from the Term's 1 October, pinned to that instant rather than counted back in days, closing with the forward cells scored under the process in force, stated as a ledger-to-date figure), then the back-test results unchanged. The health questions and the analytics-state block are gone: the boards are read on the site, and the health bullets restated the production block. The ops report, the daily digest and every process-digest input are untouched. Co-Authored-By: Claude Fable 5.1 --- .github/workflows/run-ops.yml | 11 +++---- README.md | 3 +- docs/pipeline.md | 13 +++++--- src/fedcourtsai/cli.py | 13 ++++---- src/fedcourtsai/ops.py | 32 +++++++++++-------- tests/test_ops.py | 59 +++++++++++++++++++++++++++++------ 6 files changed, 90 insertions(+), 41 deletions(-) diff --git a/.github/workflows/run-ops.yml b/.github/workflows/run-ops.yml index a658cb0ca..1c33f0e00 100644 --- a/.github/workflows/run-ops.yml +++ b/.github/workflows/run-ops.yml @@ -49,10 +49,10 @@ on: schedule: - cron: "0 8 * * *" # after run-pull (07:17) so the day's runs are visible # The weekly tick: the same daily job, plus the weekly performance digest - # issue. Two steps are gated on this exact schedule string — the build - # step's WEEKLY_TICK env and the posting step's `if` — and a shape test - # pins both to a cron this block actually declares. Offset from the daily - # 08:00 so the two runs never race on the `ops-metrics` branch. + # issue. The posting step's `if` is gated on this exact schedule string, + # and a shape test pins it to a cron this block actually declares. Offset + # from the daily 08:00 so the two runs never race on the `ops-metrics` + # branch. - cron: "30 8 * * 1" workflow_dispatch: @@ -120,8 +120,7 @@ jobs: # An open issue wearing a run:* fan-out label is a marker somebody left # behind: nothing keys on those labels, so it is queued work to no one # and the report says so rather than letting it read as a round in - # flight — the weekly digest carries the same signal as a health - # question. One gh call per label, merged; the tested + # flight. One gh call per label, merged; the tested # `summarize_trigger_issues` does the filtering/ordering. for label in run:predict run:evaluate; do gh_retry gh issue list --repo "$REPO" --label "$label" --state open --limit 100 \ diff --git a/README.md b/README.md index 1afdb20c1..1b4dc1701 100644 --- a/README.md +++ b/README.md @@ -47,7 +47,8 @@ requests. Plus `run-ops` (a read-only daily operations report, plus two issues the maintainer closes once read: a daily prediction-reading digest — one predicted event with every predictor side by side — and a Monday performance digest carrying the -week's cells, spend, board state, and back-test results) and +cells and spend produced over the week, the trailing month and the Term to +date, plus the back-test results) and `run-analytics` — seven dispatch modes: corpus statistics, the distribution-parse census, the document text-coverage enumeration, the tool-usage roll-up, the metrics refresh, the daily big-case board, and the diff --git a/docs/pipeline.md b/docs/pipeline.md index cab189dca..52b92b6d4 100644 --- a/docs/pipeline.md +++ b/docs/pipeline.md @@ -128,11 +128,14 @@ Four blocks, in the order a reader needs them: from the 1 October that Term opened to the day the digest is generated. The cutoff is pinned to that instant rather than counted back in days, or a mid-morning render would cut the Term's own first morning out of its census. - It closes with the forward cells scored under the process in force — the Term - is the period that count is worth reading over, since a forward cell is minted - once at its event and never again. A frozen scope with nothing scored in - either stratum is named as the shakedown state rather than shown as a bare - zero. + It closes with the forward cells scored under the process in force. That + count is cumulative over the whole ledger, not the Term's, with a delta + against the prior ops-metrics snapshot (a week when the dated snapshot + exists, shorter when the job fell back to the latest one); it sits in the + Term block because a forward cell is minted once at its event and never + again, so the Term is the period it is worth reading beside. A frozen scope + with nothing scored in either stratum is named as the shakedown state rather + than shown as a bare zero. - **Backtest results** — the historical replay **per court**, with each court's own always-deny floor beside its accuracy and the pooled row labelled as the mixture it is (`granted` means cert on a SCOTUS row and a motion granted on a diff --git a/src/fedcourtsai/cli.py b/src/fedcourtsai/cli.py index 2d692c4e5..1ed31db69 100644 --- a/src/fedcourtsai/cli.py +++ b/src/fedcourtsai/cli.py @@ -7920,12 +7920,13 @@ def ops_report( # noqa: PLR0913 - one option per independent read-only feed to stdout (the run-ops job's Actions step summary); ``--json`` writes the structured ``OpsReport``. - ``--digest-out`` renders the **weekly performance digest** — the health - questions, the committed boards' state with each empty one saying why it is - empty, the week's cells and measured spend, and the back-test results, - every metrics-derived figure carrying the vintage of the artifact it came - from. ``--digest-post-repo`` additionally opens it as a `weekly-digest` - issue, once per ISO week. + ``--digest-out`` renders the **weekly performance digest** — the cells and + measured spend produced over the week, the trailing month (closing with the + spend backstop's verdict) and the Term to date (closing with the forward + cells scored under the process in force), then the back-test results, each + back-test figure carrying the vintage of the artifact it came from. + ``post-weekly-digest`` is the separate command that opens it as a + `weekly-digest` issue, once per ISO week. Unlike the leaderboard/back-test roll-ups it is a point-in-time snapshot, so it is surfaced, not committed. diff --git a/src/fedcourtsai/ops.py b/src/fedcourtsai/ops.py index 23a1a18cd..d77a3892f 100644 --- a/src/fedcourtsai/ops.py +++ b/src/fedcourtsai/ops.py @@ -628,11 +628,12 @@ def _cell(value: str) -> str: def _sourced(filename: str, vintaged: Vintaged[object]) -> str: """``` `metrics/x.json`, vintage YYYY-MM-DD ``` — the artifact and when it moved. - Every metrics-derived figure carries this, because none of these artifacts is - refreshed on the digest's own schedule: a board is byte-stable and a statpack - moves only when the corpus does, so a figure without its vintage silently - claims to be this week's. An unknown vintage says so rather than being - omitted — the reader still needs to know the number's age is unestablished. + Every back-test figure carries this, because none of those artifacts is + refreshed on the digest's own schedule: a replay board is byte-stable and the + cert back-test moves only when a maintainer dispatches one, so a figure + without its vintage silently claims to be this week's. An unknown vintage + says so rather than being omitted — the reader still needs to know the + number's age is unestablished. """ vintage = f"vintage {vintaged.vintage}" if vintaged.vintage else "vintage unknown" return f"`metrics/{filename}`, {vintage}" @@ -688,9 +689,11 @@ def _render_backstop(backstop: SpendVerdict | None) -> list[str]: def _frozen_cells_line(substance: SubstanceDigest | None) -> list[str]: """Forward cells scored under the process in force, closing the Term block. - The Term is the period this count is worth reading over: a forward cell is - minted once at its event and never again, so a week's delta alone cannot say - whether the scored population is growing or standing still. Stated as a + The count is cumulative over the whole ledger, not the Term's, and the + rendered line says so, because a line under a Term heading is quoted + without it. It sits in the Term block because a forward cell is minted once + at its event and never again, so the Term is the period it is worth reading + beside; the delta is against the prior ops-metrics snapshot. Stated as a figure — the boards are where a reader interrogates it. """ if substance is None: @@ -710,7 +713,10 @@ def _frozen_cells_line(substance: SubstanceDigest | None) -> list[str]: and cells.evaluations_retrospective == 0 ) tail = " No frozen-process cells yet — still shakedown." if shakedown else "" - return ["", f"Forward cells scored ({substance.process_scope}): {counted}.{tail}"] + return [ + "", + f"Forward cells scored ({substance.process_scope}, ledger to date): {counted}.{tail}", + ] #: The court whose rows the digest publishes beside the pooled figure. Pooling @@ -1133,10 +1139,10 @@ def render_weekly_digest( lines += _render_backtests(analytics) # The marker line stays verbatim; everything below it is defused in one pass, # exactly as the daily digest's body is. Almost all of this document is - # harness-computed figures, but the role and stage labels threaded through the - # census tables — and the predictor ids in the back-test ones — are free-form - # strings off the ledger, and a field added later would otherwise arrive - # untreated. + # harness-computed figures (the census tables' role and stage labels are + # enum-valued), but the predictor ids threaded through the back-test tables + # are free-form strings off the ledger, and a field added later would + # otherwise arrive untreated. prose = _defuse_comments("\n".join(lines[WEEKLY_DIGEST_MARKER_LINES:])) document = "\n".join([*lines[:WEEKLY_DIGEST_MARKER_LINES], prose]) + "\n" if len(document) > _DIGEST_MAX_CHARS: diff --git a/tests/test_ops.py b/tests/test_ops.py index 0a3e56422..75125fde0 100644 --- a/tests/test_ops.py +++ b/tests/test_ops.py @@ -2842,6 +2842,8 @@ def _commit_usage(data_root: Path, *, docket: int, cost: float, created_at: date # reports the outgoing Term — the long-conference cohort's own Term. ("2026-09-21T08:30:00+00:00", 2025, "2025-10-01", 355), ("2026-10-05T08:30:00+00:00", 2026, "2026-10-01", 4), + # The Term's first morning: the window is the day itself, 0d wide. + ("2026-10-01T08:30:00+00:00", 2026, "2026-10-01", 0), ], ) def test_the_term_block_rolls_on_the_first_of_october( @@ -2895,8 +2897,8 @@ def test_the_month_window_is_the_spend_backstops_own(tmp_path: Path) -> None: def test_the_three_windows_are_counted_over_one_walk_of_the_ledger(tmp_path: Path) -> None: - # One cell inside each nested window: the week's count must not leak into a - # narrower one, and every window must see what falls inside it. + # One cell inside each nested window: a wider window must see everything the + # narrower ones do and only what falls inside its own bounds. now = datetime(2026, 9, 21, 8, 30, tzinfo=UTC) _commit_usage(tmp_path / "data", docket=1, cost=1.0, created_at=now - timedelta(days=2)) _commit_usage(tmp_path / "data", docket=2, cost=2.0, created_at=now - timedelta(days=20)) @@ -2910,6 +2912,25 @@ def test_the_three_windows_are_counted_over_one_walk_of_the_ledger(tmp_path: Pat assert (production.term.census.cells, production.term.spend_usd) == (3, 7.0) +def test_the_month_block_and_the_backstop_count_the_same_records(tmp_path: Path) -> None: + # The verdict renders at the end of the month block; both are one window over + # one walk, so a ledger with rows inside and outside it must give the census + # and the verdict identical cells and dollars. + now = datetime(2026, 9, 21, 8, 30, tzinfo=UTC) + config_root = tmp_path / "config" + config_root.mkdir() + (config_root / "tracking.yaml").write_text("spend:\n ceiling_usd: 100.0\n window_days: 30\n") + _commit_usage(tmp_path / "data", docket=1, cost=1.0, created_at=now - timedelta(days=2)) + _commit_usage(tmp_path / "data", docket=2, cost=2.0, created_at=now - timedelta(days=20)) + _commit_usage(tmp_path / "data", docket=3, cost=4.0, created_at=now - timedelta(days=200)) + + production = cli._weekly_production(tmp_path / "data", config_root, now) + + assert production.backstop is not None and production.backstop.enforced + assert production.month.spend_usd == production.backstop.spent_usd == 3.0 + assert production.month.census.cells == production.backstop.cells == 2 + + def test_the_term_block_closes_with_the_forward_cells_scored() -> None: # A forward cell is minted once at its event and never again, so the Term is # the period the count is worth reading over — and it is a figure, not a @@ -2929,7 +2950,9 @@ def test_the_term_block_closes_with_the_forward_cells_scored() -> None: lines = ops.render_weekly_digest(report, production=_production()).splitlines() - figure = lines.index("Forward cells scored (frozen): 1 total, no prior snapshot to diff.") + figure = lines.index( + "Forward cells scored (frozen, ledger to date): 1 total, no prior snapshot to diff." + ) assert lines.index("## Produced this term (last 355d)") < figure assert not any(line.startswith("## ") for line in lines[figure:]) @@ -2961,7 +2984,7 @@ def test_the_forward_cell_figure_carries_the_week_over_week_delta() -> None: md = ops.render_weekly_digest(report, production=_production()) - assert "Forward cells scored (frozen): +2 this week, 3 total." in md + assert "Forward cells scored (frozen, ledger to date): +2 this week, 3 total." in md def test_the_term_block_names_the_shakedown_state_rather_than_a_bare_zero() -> None: @@ -2981,7 +3004,7 @@ def test_the_term_block_names_the_shakedown_state_rather_than_a_bare_zero() -> N md = ops.render_weekly_digest(report, production=_production()) assert ( - "Forward cells scored (frozen): 0 total, no prior snapshot to diff. " + "Forward cells scored (frozen, ledger to date): 0 total, no prior snapshot to diff. " "No frozen-process cells yet — still shakedown." in md ) @@ -2998,7 +3021,7 @@ def test_an_all_versions_scope_is_not_read_as_the_shakedown_state() -> None: md = ops.render_weekly_digest(report, production=_production()) - assert "Forward cells scored (all): 0 total, no prior snapshot to diff." in md + assert "Forward cells scored (all, ledger to date): 0 total, no prior snapshot to diff." in md assert "still shakedown" not in md @@ -3294,11 +3317,27 @@ def test_the_weekly_digest_body_is_clamped_under_the_issue_size_limit() -> None: def test_the_weekly_digest_defuses_a_marker_quoted_by_the_ledger() -> None: - # Role and stage labels come off the ledger as free-form strings; the same - # one-pass defusing the daily digest applies keeps one from forging a marker. - rows = [CellCensusRow("", "cert", 1)] + # Predictor ids come off the ledger as free-form strings and reach the + # back-test tables; the same one-pass defusing the daily digest applies keeps + # one from forging a marker. + board = Backtest( + predictors_evaluated=1, + events_scored=1, + entries=[ + BacktestEntry( + rank=1, + predictor_id="", + events_scored=1, + accuracy=1.0, + granted_accuracy=1.0, + mean_brier_score=0.0, + always_denied_accuracy=1.0, + lift_over_always_denied=0.0, + ) + ], + ) - md = ops.render_weekly_digest(_empty_report(), production=_production(rows, cells=1, events=1)) + md = ops.render_weekly_digest(_empty_report(), analytics=_analytics(backtest=board)) assert "" not in md assert "<!-- weekly-digest: 2026-W40 -->" in md