Skip to content

Calibrate the frozen attention candidate before the prospective evaluation phase #396

Description

@atimics

Next slice

Complete the calibration milestone already specified by attention-shadow-release.md, using its frozen candidate and existing durable trial records.

The September 24 study observed 172 versus 143 event selections in 340 review slots on four later dates. This supports the prospective trial; repeated ticker selections and four dates do not establish a general ranking gain or trading profitability.

Acceptance

  • Check actual deployed trial health and saved receipts before changing collection. Coordinate with [P1] Restore worker health and sports feed freshness after repeated production 503s #328 if worker health prevents reliable observations.
  • Keep the frozen model, baseline policy, twenty input features, prediction timing, symmetric event target and unknown-outcome rules bound to every included record.
  • Use ten complete calibration sessions under the existing contract. Keep interrupted captures, fallback cases and missing outcomes visible; do not drop unknown review slots to improve the reported score.
  • Fit the simple calibration map described by the release plan. Save source receipt hashes, cutoff, parameters, reliability results and resulting policy version.
  • Set calibration limits before collecting evaluation predictions. Preserve the existing later-session ranking/coverage/uncertainty rules.
  • Freeze the calibrated artifact before the at-least-twenty-session evaluation phase. Keep calibration dates out of that phase.
  • Report full-board and market-only effects separately, with date-level uncertainty, coverage, unknown-outcome bounds and operational failures.
  • Preserve the existing public policy until the separate final release decision has the required evidence.

Completion

Link the calibration artifact, reproducible report and frozen later-evaluation specification. If ten usable sessions are unavailable, report the exact coverage gap and collection state; no additional model search is required.

This is the next milestone of the existing hobby experiment, with its existing spend and deployment scope. No new general ML platform, new data provider, or broadened claim is part of this issue.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions