Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
15 changes: 7 additions & 8 deletions .github/workflows/run-ops.yml
Original file line number Diff line number Diff line change
Expand Up @@ -27,8 +27,8 @@ name: run-ops
# The reading surfaces are digests a maintainer closes, not a standing body this
# job edits in place: an unread digest is an open issue, so the backlog is
# visible without a reading-state store. The Monday schedule opens the **weekly
# performance digest** as its own `weekly-digest` issue — the health questions,
# the committed boards' state, the week's cells and measured spend, and the
# performance digest** as its own `weekly-digest` issue — the cells and measured
# spend produced over the week, the trailing month and the Term to date, and the
# back-test results, every metrics figure carrying its artifact's vintage. A
# second job on the daily tick opens the **daily prediction-reading digest**:
# one predicted event with every predictor side by side, on its own
Expand All @@ -49,10 +49,10 @@ on:
schedule:
- cron: "0 8 * * *" # after run-pull (07:17) so the day's runs are visible
# The weekly tick: the same daily job, plus the weekly performance digest
# issue. Two steps are gated on this exact schedule string — the build
# step's WEEKLY_TICK env and the posting step's `if` — and a shape test
# pins both to a cron this block actually declares. Offset from the daily
# 08:00 so the two runs never race on the `ops-metrics` branch.
# issue. The posting step's `if` is gated on this exact schedule string,
# and a shape test pins it to a cron this block actually declares. Offset
# from the daily 08:00 so the two runs never race on the `ops-metrics`
# branch.
- cron: "30 8 * * 1"
workflow_dispatch:

Expand Down Expand Up @@ -120,8 +120,7 @@ jobs:
# An open issue wearing a run:* fan-out label is a marker somebody left
# behind: nothing keys on those labels, so it is queued work to no one
# and the report says so rather than letting it read as a round in
# flight — the weekly digest carries the same signal as a health
# question. One gh call per label, merged; the tested
# flight. One gh call per label, merged; the tested
# `summarize_trigger_issues` does the filtering/ordering.
for label in run:predict run:evaluate; do
gh_retry gh issue list --repo "$REPO" --label "$label" --state open --limit 100 \
Expand Down
3 changes: 2 additions & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -47,7 +47,8 @@ requests.
Plus `run-ops` (a read-only daily operations report, plus two issues the maintainer
closes once read: a daily prediction-reading digest — one predicted event with
every predictor side by side — and a Monday performance digest carrying the
week's cells, spend, board state, and back-test results) and
cells and spend produced over the week, the trailing month and the Term to
date, plus the back-test results) and
`run-analytics` — seven dispatch modes: corpus statistics, the
distribution-parse census, the document text-coverage enumeration, the
tool-usage roll-up, the metrics refresh, the daily big-case board, and the
Expand Down
2 changes: 1 addition & 1 deletion docs/cli.md

Large diffs are not rendered by default.

49 changes: 28 additions & 21 deletions docs/pipeline.md
Original file line number Diff line number Diff line change
Expand Up @@ -112,25 +112,30 @@ unread backlog.

Four blocks, in the order a reader needs them:

- **Health questions** — the fixed interrogative bullets (replay calibration,
forward cells scored, watchlist vs next conference, oldest stalled trigger,
spend vs budget): numbers as questions demanding a reaction. These carry the
ops report's un-vintaged framing, which is why the vintage rule below is
scoped to the two blocks that publish figures to quote.
- **Analytics state** — what the committed boards hold. An empty one names the
condition that empties it — which cells the frozen headline ranks, and how
many have reached it — rather than showing a bare zero, and an artifact that
has never landed reads differently from one that landed empty. Plus the
statpack's two headline rates — stated with each one's own denominator, and
with the plain statement that **neither anchors a scored cell**: a forward
cert cell is scored against its own band's strictly-prior-Term risk-set rate,
and the pooled band rate is a fit diagnostic for the ranking constant rather
than a scoring baseline ([salience.md](salience.md)).
- **Produced this week** — cells landed by role and stage, how many events they
covered, and the week's measured spend, all over one window and one set of
`usage.json` records; then the spend backstop's own (longer) window and how
much of its ceiling the trailing period has consumed. An unenforced ceiling
says so instead of reporting a fraction of a budget that does not exist.
`usage.json` records. The window's bounds are in the heading: the Monday tick
titles its issue for the ISO week that *starts* that morning while the census
covers the seven days before it, so without them the block would describe the
previous week under this week's heading.
- **Produced this month** — the same shape over the trailing window the ex-post
spend backstop is configured with, closing with that backstop's own verdict
and how much of its ceiling the period has consumed. The window is the
backstop's own precisely so the census and the verdict beside it cannot
describe different periods. An unenforced ceiling says so instead of reporting
a fraction of a budget that does not exist.
- **Produced this term** — the same shape again over the October Term to date,
from the 1 October that Term opened to the day the digest is generated. The
cutoff is pinned to that instant rather than counted back in days, or a
mid-morning render would cut the Term's own first morning out of its census.
It closes with the forward cells scored under the process in force. That
count is cumulative over the whole ledger, not the Term's, with a delta
against the prior ops-metrics snapshot (a week when the dated snapshot
exists, shorter when the job fell back to the latest one); it sits in the
Term block because a forward cell is minted once at its event and never
again, so the Term is the period it is worth reading beside. A frozen scope
with nothing scored in either stratum is named as the shakedown state rather
than shown as a bare zero.
- **Backtest results** — the historical replay **per court**, with each court's
own always-deny floor beside its accuracy and the pooled row labelled as the
mixture it is (`granted` means cert on a SCOTUS row and a motion granted on a
Expand All @@ -157,10 +162,12 @@ Four blocks, in the order a reader needs them:
the number is in, because a caveat one bullet away does not travel when the
line is quoted. A board with no entries says so and prints no floor.

**In the analytics and back-test blocks, every figure carries the vintage of the
artifact it came from.** None of those artifacts is refreshed on this schedule —
a board is byte-stable and a statpack moves only when the corpus does — so a
figure without its vintage would silently claim to be this week's. The vintage
**In the back-test block, every figure carries the vintage of the artifact it
came from.** None of those artifacts is refreshed on this schedule — a board is
byte-stable, and the cert back-test moves only when a maintainer dispatches one —
so a figure without its vintage would silently claim to be this week's. The
production blocks need no vintage: they are computed from the committed ledger at
render time. The vintage
is the commit that last wrote the file, and a **shallow** checkout yields none:
in a depth-1 clone the one grafted commit matches every path, so a pathspec'd
`git log` would stamp every board with today's date — the exact misreading the
Expand Down
89 changes: 61 additions & 28 deletions src/fedcourtsai/cli.py
Original file line number Diff line number Diff line change
Expand Up @@ -185,6 +185,7 @@
DAILY_DIGEST_MARKER_LINES,
WEEKLY_DIGEST_LABEL,
WEEKLY_DIGEST_MARKER_LINES,
ProducedWindow,
Vintaged,
WeeklyAnalytics,
WeeklyProduction,
Expand Down Expand Up @@ -301,15 +302,13 @@
AgentFlags,
AgentToolingFeedback,
Backtest,
BigCaseBoard,
CellFailure,
CellMode,
CertBacktest,
CertBacktestCellLoss,
CertBacktestDispatch,
CertBacktestProvenance,
ClaimScoreBlock,
ClaimScoreBoard,
ConferenceBucket,
CorpusValidation,
DataHealth,
Expand Down Expand Up @@ -7739,12 +7738,8 @@ def _vintaged[T: BaseModel](path: Path, model: type[T]) -> Vintaged[T]:


def _weekly_analytics(metrics_root: Path) -> WeeklyAnalytics:
"""The committed boards the weekly digest reports, each with its own vintage."""
"""The committed replay artifacts the weekly digest's back-test block reports."""
return WeeklyAnalytics(
leaderboard=_vintaged(metrics_root / "leaderboard.json", Leaderboard),
claim_scores=_vintaged(metrics_root / "claim-scores.json", ClaimScoreBoard),
big_cases=_vintaged(metrics_root / "big-cases.json", BigCaseBoard),
statpack=_vintaged(metrics_root / "statpack.json", StatPack),
backtest=_vintaged(metrics_root / "backtest.json", Backtest),
salience_replay=_vintaged(metrics_root / "salience-replay.json", SalienceReplay),
# Never produced by the scheduled refresh — a real-engine replay spends
Expand Down Expand Up @@ -7780,29 +7775,66 @@ def _parse_when(stamp: str) -> datetime:
return parsed if parsed.tzinfo is not None else parsed.replace(tzinfo=UTC)


#: The window the weekly digest's production census and its spend figure share.
#: A week, because that is the period the digest covers; the spend *backstop*
#: keeps its own, longer window, which is why the two are reported separately
#: rather than one being derived from the other.
#: The window the digest's first production block covers. A week, because that is
#: the period the digest is opened for; the month block takes its own window from
#: the spend backstop's config instead, so the backstop's verdict and the census
#: printed beside it can never describe different periods.
_WEEKLY_WINDOW_DAYS = 7


def _produced_over(
usage: Sequence[ModelUsage],
title: str,
*,
window_days: int,
when: datetime,
since: datetime | None = None,
term: int | None = None,
) -> ProducedWindow:
"""One production block's census and spend, over one window of the ledger."""
return ProducedWindow(
title=title,
census=cell_census(usage, window_days=window_days, now=when, since=since),
spend_usd=spend_over(usage, window_days=window_days, now=when, since=since)[0],
window_start=(since or when - timedelta(days=window_days)).date(),
window_end=when.date(),
term=term,
)


def _weekly_production(data_root: Path, config_root: Path, when: datetime) -> WeeklyProduction:
"""The week's cells and cost, plus the spend backstop's own verdict.
"""The digest's three production windows, plus the spend backstop's verdict.

One walk of the ledger for all three figures: the census, the week's spend,
and the backstop's own longer window read the same records, so they cannot
disagree about what they cover and the growing tree is scanned once.
One walk of the ledger for every figure: the three censuses, their three
spend totals, and the backstop's own window all read the same records, so
they cannot disagree about what they cover and the growing tree is scanned
once.

The Term window is pinned to the instant its Term opened rather than counted
back in days — an October Term opens at midnight on 1 October and the digest
renders mid-morning, so a day count would cut the Term's own first morning
out of its census.
"""
usage = iter_usage(data_root)
census = cell_census(usage, window_days=_WEEKLY_WINDOW_DAYS, now=when)
spent, _cells = spend_over(usage, window_days=_WEEKLY_WINDOW_DAYS, now=when)
spend_config = load_spend_config(config_root)
term = october_term_year(when.date())
term_start = datetime(term, 10, 1, tzinfo=UTC)
return WeeklyProduction(
census=census,
spend_usd=spent,
backstop=verdict_over(usage, load_spend_config(config_root), now=when),
window_start=(when - timedelta(days=_WEEKLY_WINDOW_DAYS)).date(),
window_end=when.date(),
week=_produced_over(
usage, "Produced this week", window_days=_WEEKLY_WINDOW_DAYS, when=when
),
month=_produced_over(
usage, "Produced this month", window_days=spend_config.window_days, when=when
),
term=_produced_over(
usage,
"Produced this term",
window_days=(when.date() - term_start.date()).days,
when=when,
since=term_start,
term=term,
),
backstop=verdict_over(usage, spend_config, now=when),
)


Expand Down Expand Up @@ -7888,12 +7920,13 @@ def ops_report( # noqa: PLR0913 - one option per independent read-only feed
to stdout (the run-ops job's Actions step summary); ``--json`` writes the
structured ``OpsReport``.

``--digest-out`` renders the **weekly performance digest** — the health
questions, the committed boards' state with each empty one saying why it is
empty, the week's cells and measured spend, and the back-test results,
every metrics-derived figure carrying the vintage of the artifact it came
from. ``--digest-post-repo`` additionally opens it as a `weekly-digest`
issue, once per ISO week.
``--digest-out`` renders the **weekly performance digest** — the cells and
measured spend produced over the week, the trailing month (closing with the
spend backstop's verdict) and the Term to date (closing with the forward
cells scored under the process in force), then the back-test results, each
back-test figure carrying the vintage of the artifact it came from.
``post-weekly-digest`` is the separate command that opens it as a
`weekly-digest` issue, once per ISO week.

Unlike the leaderboard/back-test roll-ups it is a point-in-time snapshot, so
it is surfaced, not committed.
Expand Down
Loading
Loading