Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
3 changes: 3 additions & 0 deletions .github/workflows/reproduce.yml
Original file line number Diff line number Diff line change
Expand Up @@ -45,3 +45,6 @@ jobs:

- name: LSUV data-aware init (4th candidate, underperformed)
run: python scripts/explore_lsuv_init.py

- name: PROBE-mode tie-break for the Xavier/He blind spot (raises he-recall, nets worse regret)
run: python scripts/explore_probe_tiebreak.py
Binary file modified data/meta_dataset.db
Binary file not shown.
157 changes: 29 additions & 128 deletions docs/index.html
Original file line number Diff line number Diff line change
Expand Up @@ -873,134 +873,35 @@ <h2 id="14-end-to-end-experimental-pipeline"><span class="num">14</span>End-to-E
MDU --&gt; FA[Failure Analysis + Retrain]
FA --&gt; NEXT["PRECOG v(n+1)"]

classDef reject fill:#fdecec,stroke:#c8483a,color:#7a2b21;
classDef confirm fill:#e9f7ef,stroke:#1f9d55,color:#155c33;
classDef next fill:#eef2ff,stroke:#1a56db,stroke-width:2px,color:#1a56db;
class STOP reject;
class FT,GT confirm;
class NEXT next;
</pre>
<p>This loop never stops after a single iteration: every PRECOG generation must be compared to the previous one under a strictly identical protocol.</p>
<hr/>
<h2 id="15-test-protocols"><span class="num">15</span>Test Protocols</h2>
<div class="table-scroll"><table>
<thead>
<tr>
<th>Protocol</th>
<th>Question</th>
<th>Main metric</th>
</tr>
</thead>
<tbody>
<tr>
<td><strong>P1 — Ranking</strong></td>
<td>Does PRECOG rank configurations correctly?</td>
<td>Spearman ρ, Kendall τ</td>
</tr>
<tr>
<td><strong>P2 — Top-K</strong></td>
<td>Does it retrieve the best configurations?</td>
<td>Recall@K</td>
</tr>
<tr>
<td><strong>P3 — Convergence</strong></td>
<td>Does the chosen configuration converge faster?</td>
<td>Steps/Time-to-Target</td>
</tr>
<tr>
<td><strong>P4 — Compute</strong></td>
<td>How much compute is saved?</td>
<td>GPU-hours / FLOPs</td>
</tr>
<tr>
<td><strong>P5 — Data efficiency</strong></td>
<td>Same quality with less data?</td>
<td>Samples-to-Target</td>
</tr>
<tr>
<td><strong>P6 — Generalization</strong></td>
<td>Does it work on a never-seen model/dataset?</td>
<td>Out-of-distribution performance</td>
</tr>
</tbody>
</table></div>
<h3 id="151-trainvalidationtest-separation"><span class="num sub">15.1</span>TRAIN/VALIDATION/TEST separation</h3>
<pre class="diagram"><code class="language-text">PRECOG TRAIN → known datasets and architectures, experiment history
PRECOG VALIDATION → different datasets, partially new architectures
PRECOG TEST (locked) → never seen, never used to improve PRECOG
</code></pre>
<h3 id="152-reference-benchmarks-for-the-initial-phase"><span class="num sub">15.2</span>Reference benchmarks for the initial phase</h3>
<ul>
<li><strong>NATS-Bench</strong> (successor to the now-deprecated NAS-Bench-201): a reference architecture space with pre-computed performance (CIFAR-10, CIFAR-100, ImageNet16-120) — useful for testing ranking without having to train every architecture oneself.</li>
<li><strong>NAS-Bench-Suite-Zero / JAHS-Bench / HPO-B</strong>: actively maintained benchmarks, the first specifically designed to evaluate zero-cost proxies (see stack.md §4 for the rationale behind these choices over the now poorly-maintained HPOBench).</li>
<li><strong>Synthetic laboratory</strong> (generated in-house): fully controlled datasets and models (noise, entropy, dimensionality, depth, width), enabling candidate causal variables to be isolated before moving to real benchmarks.</li>
</ul>
<h3 id="153-multi-seed-and-statistical-tests"><span class="num sub">15.3</span>Multi-seed and statistical tests</h3>
<p>Every important experiment is repeated over several seeds, with mean, standard deviation, and confidence interval (95% CI) computed. Comparisons between methods (PRECOG vs. Random, vs. BO, vs. Hyperband, vs. Vizier) use appropriate statistical tests (e.g. a Wilcoxon signed-rank test rather than a t-test when parametric assumptions aren't guaranteed), to avoid declaring superiority based on a lucky seed.</p>
<hr/>
<h2 id="16-metrics-and-objectives-to-be-demonstrated-not-guaranteed"><span class="num">16</span>Metrics and Objectives (to be demonstrated, not guaranteed)</h2>
<div class="table-scroll"><table>
<thead>
<tr>
<th>Metric</th>
<th>Definition</th>
<th>Experimental target</th>
</tr>
</thead>
<tbody>
<tr>
<td>Ranking correlation</td>
<td>Spearman ρ / Kendall τ between PRECOG's ranking and the real ranking</td>
<td>ρ ≥ 0.80 then ≥ 0.90</td>
</tr>
<tr>
<td>Top-K recall</td>
<td>$Recall@K = \|\text{PredictedTopK} \cap \text{TrueTopK}\| / K$</td>
<td>Recall@10 ≥ 80% then ≥ 90%</td>
</tr>
<tr>
<td>Compute reduction</td>
<td>$1 - C_{PRECOG}/C_{baseline}$</td>
<td>≥ 50% then ≥ 70%</td>
</tr>
<tr>
<td>Performance retention</td>
<td>$Performance_{PRECOG}/Performance_{oracle}$</td>
<td>≥ 99% (or a tolerance defined a priori)</td>
</tr>
<tr>
<td>Data efficiency</td>
<td>$Samples_{baseline}/Samples_{PRECOG}$ for equal target performance</td>
<td>≥ 30–50% reduction, to be refined</td>
</tr>
<tr>
<td>Time/Steps-to-Target</td>
<td>Reduction in time/number of steps to reach a target</td>
<td>≥ 50% reduction</td>
</tr>
<tr>
<td>Prediction error (learning curve)</td>
<td>$\lvert \text{Prediction} - \text{Actual} \rvert$</td>
<td>≈ 5–10% depending on the metric</td>
</tr>
<tr>
<td>Generalization</td>
<td>Recall@K on never-seen tasks/architectures/datasets</td>
<td>same order of magnitude as on known data</td>
</tr>
</tbody>
</table></div>
<p>These targets are <strong>progression hypotheses</strong>, formalized as successive gates (§17), never presented as already achieved.</p>
<hr/>
<h2 id="17-progression-gates"><span class="num">17</span>Progression Gates</h2>
<pre class="mermaid">flowchart TD
P[PRECOG] --&gt; G1["Gate 1: ρ ≥ 0.70?"]
G1 --&gt; G2["Gate 2: Recall@10 ≥ 80%?"]
G2 --&gt; G3["Gate 3: Compute reduction ≥ 50%?"]
G3 --&gt; G4["Gate 4: Generalization maintained&lt;br/&gt;(never-seen data)?"]
G4 --&gt; G5["Gate 5: Recall@10 ≥ 90%?"]
G5 --&gt; G6["Gate 6: Compute reduction ≥ 70%?"]
G6 --&gt; ADV["PRECOG 'advanced level'"]
<section>
<h2><span class="num">04</span>A bug we found but can't fix</h2>
<p>
<code>jacob_cov</code> never once recommends "He init" across 60 test
tasks — including the 10 where it's genuinely fastest. The cause is
structural: Xavier and He (zero-biased networks) draw from the same
underlying Gaussian values, differing only by a positive scale — and
<code>jacob_cov</code> only reads activation <em>sign</em>, which scaling
cannot flip.
</p>
<pre><code>jacob_cov(He-init network) == jacob_cov(Xavier-init network)</code></pre>
<p>
exactly, for the same seed — checked across all 312 tasks, max difference
is <code>0.0</code>. Two other proxies (<code>effective_rank</code>,
<code>jacobian_condition_mean</code>) share the property for the same
reason. Three attempted fixes, all failed: raw and population-normalized
tie-breaking with <code>gradient_norm</code> (same structural blindness,
different proxy), and — the one thing that should have worked, since
it stops reading a PURE-mode proxy entirely — a bounded 50-step
<a href="https://github.com/rustnew/precog-trainability/blob/main/docs.md">PROBE-mode</a>
run on just the tied candidates. It does recover some "he" picks
(recall 0%→30% on the 10 tasks where "he" is truly best), but net
<em>regret gets worse</em> (+14.6→+45.4 steps): 50 real training steps
is enough to make "he" look locally better, not enough to see that it
sometimes never converges at all within budget.
</p>
<a class="finding-link" href="https://github.com/rustnew/precog-trainability/blob/main/results/reports/2026-09-02T08-04-49Z_explore_scale_invariance_blindspot.md">Full audit of all 11 proxies →</a>
<a class="finding-link" href="https://github.com/rustnew/precog-trainability/blob/main/results/reports/2026-09-03T12-59-38Z_explore_probe_tiebreak.md">Why the PROBE-mode fix doesn't net out →</a>
</section>

classDef next fill:#eef2ff,stroke:#1a56db,stroke-width:2px,color:#1a56db;
class ADV next;
Expand Down
87 changes: 87 additions & 0 deletions precog/meta_predictor.py
Original file line number Diff line number Diff line change
Expand Up @@ -52,6 +52,7 @@

from precog.meta_knowledge_base import MetaKnowledgeBase
from precog.model import InitMethod
from precog.modes import Mode, TrainingConfig, TrainProtocol, train
from precog.regime import _bucket_noise, _bucket_volume
from precog.trainability import zero_cost_features

Expand Down Expand Up @@ -151,6 +152,7 @@ class Recommendation:
steps_range: tuple[float, float] # +/- 1 std across the ensemble
confidence: float # 1 - (relative spread), clamped to [0, 1]
per_candidate: dict[str, dict] # every candidate's own prediction, for transparency
probe_cost_steps: int = 0 # PROBE-mode budget actually spent on this decision (docs.md §5 cost-accounting)


class MetaPredictor:
Expand Down Expand Up @@ -433,6 +435,91 @@ def secondary_key(k: str) -> float:
)


class ProbeTieBreakPredictor:
"""Third fix attempt for zc_jacobcov's proven blind spot (see
TieBreakHeuristicPredictor above): jacob_cov's binary activation-sign
statistic is *exactly* invariant to the positive rescaling that
separates Xavier from He (max |xavier-he| jacob_cov = 0.0 across all
312 meta-dataset tasks), so no PURE-mode secondary proxy -- raw or
population-normalized gradient_norm, both tried in
scripts/compare_meta_predictors.py -- can ever break that exact tie;
gradient_norm turned out to carry the same he>xavier scale confound
jacob_cov's sign-only statistic doesn't even look at.

This tries the other option named in
results/reports/2026-09-02T08-04-49Z_explore_scale_invariance_blindspot.md:
a minimal PROBE-mode check (docs.md §5: DeltaW != 0, but bounded and
logged, 50-1000 steps by contract) spent *only* on the exact tie
jacob_cov cannot see -- train each tied candidate for `probe_steps`
real steps at its own (learning_rate, batch_size, optimizer) and keep
whichever ends with the lower loss. `last_probe_cost_steps` records
the budget actually spent on the most recent call, so callers can
report it per the Zero-Training Contract's own requirement ("must
always be possible to answer how much PROBE adds over PURE alone, for
what additional cost") -- see scripts/explore_probe_tiebreak.py."""

def __init__(
self,
primary_proxy: str = "jacob_cov",
primary_higher_is_better: bool = False,
tie_tolerance: float = 1e-6,
probe_steps: int = 50,
):
self.primary_proxy = primary_proxy
self.primary_higher_is_better = primary_higher_is_better
self.tie_tolerance = tie_tolerance
self.probe_steps = probe_steps
self.last_probe_cost_steps = 0

def recommend(
self,
features_row: pd.DataFrame,
zero_cost_by_candidate: dict[InitMethod, dict],
architecture=None,
x: torch.Tensor | None = None,
y: torch.Tensor | None = None,
training_by_candidate: dict[InitMethod, TrainingConfig] | None = None,
) -> Recommendation:
per_candidate = {
c.value: {"expected_steps": float("nan"), "std_steps": 0.0, "primary_score": zc[self.primary_proxy]}
for c, zc in zero_cost_by_candidate.items()
}
primary_sign = -1 if self.primary_higher_is_better else 1
primary_values = {k: v["primary_score"] * primary_sign for k, v in per_candidate.items()}
best_primary = min(primary_values.values())
tied = [k for k, v in primary_values.items() if abs(v - best_primary) <= self.tie_tolerance]

self.last_probe_cost_steps = 0
if len(tied) == 1:
best_init_name = tied[0]
else:
if architecture is None or x is None or y is None or training_by_candidate is None:
raise ValueError(
"ProbeTieBreakPredictor needs a live architecture/x/y/training_by_candidate "
"context to actually run the PROBE that breaks the tie -- pass them through, "
"see scripts/explore_probe_tiebreak.py for how the harness wires this up."
)
probe_losses = {}
for k in tied:
training = training_by_candidate[InitMethod(k)]
protocol = TrainProtocol(
mode=Mode.PROBE, max_steps=self.probe_steps, loss_threshold=-1.0, seed=0
)
result = train(architecture, x, y, training, protocol)
probe_losses[k] = result.final_loss
self.last_probe_cost_steps += self.probe_steps
best_init_name = min(probe_losses, key=probe_losses.get)

return Recommendation(
recommended_init=InitMethod(best_init_name),
expected_steps=float("nan"),
steps_range=(float("nan"), float("nan")),
confidence=float("nan"), # this method makes no probabilistic claim -- see docs.md §20
per_candidate=per_candidate,
probe_cost_steps=self.last_probe_cost_steps,
)


class KNNMetaPredictor:
"""Alternative to the RandomForest MetaPredictor (docs.md §19 ablation
spirit): predicts purely from the Meta-Knowledge Base's (§9.6) nearest
Expand Down
1 change: 1 addition & 0 deletions results/gate_evaluations.csv
Original file line number Diff line number Diff line change
Expand Up @@ -232,3 +232,4 @@ evaluation_id,timestamp,generation,gate_number,metric_name,metric_value,threshol
231,2026-09-02 07:46:49,v1-meta-predictor-zc_jacobcov_normtiebreak,2,mean_regret_steps_to_threshold,47.666666666666664,0.0,1,60,regret = steps(predicted_init) - steps(true_best_init); relative_regret=+33.72%; universal_baseline_regret=+22.2; random_baseline_regret=+43.2
232,2026-09-02 08:04:49,v1-scale-invariance-blindspot,0,fraction_proxies_blind_to_xavier_vs_he,0.2727272727272727,0.0,1,312,"blind_proxies=['jacob_cov', 'effective_rank', 'jacobian_condition_mean'], tie_threshold=5% mean relative difference"
233,2026-09-02 19:28:18,v1-trainability-engine-at-scale,1,spearman_rho_gradient_norm_vs_steps_full_scale,0.5398950415558994,0.7,0,936,"full meta-dataset re-check (312 tasks) of gate1_ranking.py's original n=36 (12-task) result; no new training runs, same controlled design (§21)"
234,2026-09-03 12:59:38,v1-probe-tiebreak,1,he_recall_probe_tiebreak,0.3,0.0,1,10,"raw_he_hits=0/10, tiebreak_he_hits=0/10, probe_he_hits=3/10, probe_steps=50, mean_probe_cost_steps=68.3, overhead_pct_of_mean_full_training=37.5"
Loading
Loading