From 98a173fa21f23705521673fdc3e0ed175b4d6fa8 Mon Sep 17 00:00:00 2001 From: rustnew Date: Thu, 3 Sep 2026 21:57:14 +0100 Subject: [PATCH 1/2] Replace project page with the full docs.md specification MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Requested: the GitHub Pages site (docs/index.html) now renders docs.md in full, instead of the shorter curated landing page (the "3 findings" summary from PR #2), so a reader following the project-page link gets the complete vision document / scientific specification / roadmap, not a digest. Converted programmatically (docs.md is the single source of truth -- this page is generated from it, not hand-copied) with Python-Markdown (extra + toc + sane_lists extensions): all 30 sections, ASCII architecture diagrams, tables, and blockquotes preserved verbatim. LaTeX math ($...$/$$...$$) is protected from Markdown's underscore-based emphasis parser before conversion (subscripts like X_{model} would otherwise be mangled) and rendered client-side via MathJax. Layout: sticky table-of-contents sidebar on desktop (the spec has 30 top-level sections plus subsections, too long to scroll blind), a collapsible
TOC on narrow screens, reusing the existing site's CSS design tokens/typography. docs.md's own §0 methodological disclaimer (all figures are experimental targets, not results) stays intact and is the first section on the page, so nothing here reframes the doc as a results report. The prior page's own "3 findings" content (47% top-1, the rho correction, the Xavier/He bug) isn't lost -- it lives in the README and this project's `results/reports/*.md`, linked from the page's header ("Verified results (README)"). Co-Authored-By: Claude Sonnet 5 Claude-Session: https://claude.ai/code/session_01WSc9sb1otU6ssfxDBzeNHG --- docs/index.html | 1391 ++++++++++++++++++++++++++++++++++++++--------- 1 file changed, 1133 insertions(+), 258 deletions(-) diff --git a/docs/index.html b/docs/index.html index bb261f7..ead1517 100644 --- a/docs/index.html +++ b/docs/index.html @@ -3,8 +3,12 @@ -PRECOG — Can Pre-Training Signals Predict Trainability? - +PRECOG — Predictive Configuration & Trainability Engine + + + -
-
-

Research snapshot · v0.1.0

-

PRECOG

-

Can signals computed on an untrained network — before a single optimizer step — predict which hyperparameters will actually train well?

-
- Reproduce CI status - MIT License - Python 3.10+ +
+
+
+

Full specification · docs.md

+

PRECOG

+

Predictive Configuration & Trainability Engine — vision document, scientific specification, and research roadmap.

-
+
+
-
-

TL;DR

-
-
- 47% - top-1 accuracy picking the best init, with zero training and zero learned model — beats every fancier method tried -
-
- 0.670 → 0.395 - our own headline correlation, corrected in public after rechecking at 26× the sample size -
-
- 0.0 - the exact, provable difference our best proxy sees between two of its three candidate choices — a math bug, not a data problem -
-
-
- -
-

01Why this matters

-

- The practical bet is: check a network before you train it — in - milliseconds, on CPU, no GPU-hours spent — and skip the initialization - choice that would have wasted your training run. Two real examples, - pulled directly from our own locked test set, not cherry-picked for - effect beyond "these are honestly the clearest to read": -

- - - - - - - - - - - - - - - - - - -
TaskHeOrthogonalXavierOur pick
nonlinear_product, dim=71600 (never converges)1600 (never converges)627 stepsXavier — correct, and the only one that finishes at all
nonlinear_product, dim=4208 steps140 steps61 stepsXavier — correct, 3.4× fewer steps than the worst choice
-

- Both calls above came from a single forward/backward pass on the - untrained network — no training run spent finding out. -

-

- What this is not: a claim about needing less - data. We tested that directly — actively selecting which - training samples to use instead of random batches — and it made - things worse, not better (0/12 tasks improved). The demonstrated - gain here is narrower and more honest: skip a training run wasted on - a bad initialization, decided for free before training starts. It - also isn't universal — the pick above is only right 47% of the time - across the full test set (see below); these two are real, verified - wins, not the typical case. - Why the "less data" attempt failed → -

-
+
+
+ Contents + +
+ +
+

0. Methodological disclaimer

+

All numerical values cited in this document (Recall@10 ≥ 90%, Spearman ρ ≥ 0.90, compute reduction ≥ 70%, data reduction, prediction error ≤ 5–10%, etc.) are experimental targets to be demonstrated, not results already obtained. The ablation tables shown as examples are expected templates, not real measurements. This document is a research specification, not a results report.

+
+

1. Executive Summary

+

PRECOG is a research system aiming to transform the classic hyperparameter optimization (HPO) problem into a trainability prediction problem: from an untrained model, dataset statistics, and a hardware environment, PRECOG seeks to predict — before any training on real data — a distribution of learning configurations likely to lead to fast, stable, and data-efficient convergence.

+

PRECOG does not replace final training. It precedes and guides the configuration search, drastically reducing the number of full training runs needed to find a good configuration.

+

The project's central statement is:

+
+

PRECOG does not search for the best hyperparameters after training many configurations; it seeks to learn the relationship between a model's initial state, the properties of the problem, and the learning conditions, in order to predict — before any real training — which configurations have the highest probability of leading to fast, efficient convergence.

+
+

PRECOG is designed as a hybrid architecture combining six complementary families of methods (zero-cost proxies, NEAR-style expressivity analysis, initialization theory, meta-learning, Bayesian optimization, adaptive short validation), organized in a closed continuous-improvement loop built on a meta-dataset of experiments.

+
+

2. Motivation

+

Classic hyperparameter optimization (grid search, random search, Bayesian Optimization, Hyperband, PBT, etc.) essentially proceeds by expensive trial and error: every candidate configuration must be partially or fully trained to be evaluated. This cost becomes prohibitive as models grow.

+

Part of the recent literature (zero-cost proxies, training-free NAS, NEAR) shows that it is possible to extract informative signals about the potential quality of an architecture or configuration without full training, sometimes from a single mini-batch. These results remain fragmentary, however: no proxy is universally dominant, cross-domain generalization (vision → NLP → LLM) remains uncertain, and this work almost always focuses on ranking architectures rather than fully predicting a learning configuration (learning rate, batch size, initialization, scheduler, etc.).

+

PRECOG starts from the hypothesis that these signals, combined with each other and enriched by experience accumulated across many past training runs (meta-learning), can be exploited to build a configuration predictor, not merely an architecture ranking.

+
+

3. Scientific Problem

+

3.1 Informal formulation

+
+

Given an untrained model M, a dataset D (characterized only by its statistics, without training on it), and a hardware environment H, can we predict a learning configuration θ (fine-grained architecture, initialization, optimizer, learning rate, batch size, regularization, scheduler) that maximizes the probability of reaching a target performance, while minimizing the compute, time, and amount of data needed?

+
+

3.2 Mathematical formulation

+

$$ +\theta^* = \arg\max_{\theta} \; P\big(\text{Convergence} \geq \text{Target} \mid M, D, H, \theta\big) +$$

+

PRECOG seeks to approximate:

+

$$ +P(\theta^* \mid M, D, H) +$$

+

without updating the real model's weights on real data (see §5 for the strict definition of "without training").

+

3.3 What PRECOG is not

+
    +
  • It is not a NAS (Neural Architecture Search) in the strict sense: PRECOG can suggest architecture adjustments, but its core is the learning configuration.
  • +
  • It is not a simple wrapper around a Bayesian optimizer (e.g. Google Vizier): Vizier/BO is an internal component (the search engine), not the whole system.
  • +
  • It is not a performance guarantee: it is a probabilistic system that must express its uncertainty.
  • +
+
+

4. Research Hypotheses

+
    +
  • H1 (Pre-training signal): the state of an untrained network (gradient, Jacobian, activation, spectrum, initialization statistics) contains exploitable information about its future trainability.
  • +
  • H2 (Non-universality of proxies): no single signal is sufficient on its own; combining several families of signals is more robust than any one alone.
  • +
  • H3 (Transferability via meta-learning): experience accumulated over past (model, dataset, configuration, result) tuples improves prediction on new tuples, via a shared task representation.
  • +
  • H4 (Usefulness of short training): a very short validation run (a few dozen to a few hundred steps) sharply reduces uncertainty on the best predictions, at marginal cost.
  • +
  • H5 (Existence of regimes): the optimal relationships between hyperparameters (e.g. LR* = f(BatchSize)) depend on the learning regime (model size, data noise, architecture), not on a universal constant.
  • +
  • H6 (Correlation ≠ causation): some observed relationships between pre-training signals and final performance are confounded by third variables (the architecture, in particular); some of these must be tested experimentally before being exploited with confidence.
  • +
+

Each of these hypotheses must be tested and potentially refuted by the protocols described in §14.

+
+

5. Operational Definition of "Without Training" — the Three Modes

+

This is the project's most important methodological constraint: it must be unambiguous.

+ + + + + + + + + + + + + + + + + + + + + + + + + + + + + +
ModeDescriptionReal model weight updateUsage
PURE-PRECOGAnalysis of the untrained model and the dataset (statistics, forward passes without learning, zero-cost computations, Jacobian, etc.)ΔW = 0Reference mode for the project's central promise
PROBEVery short, controlled training (e.g. 50–1000 steps, 0.1–1% of the total budget)ΔW ≠ 0, but bounded and loggedValidation/refinement of a PURE prediction
FULL TRAININGComplete trainingΔW ≠ 0, unrestrictedGround-truth generation, never used to "cheat" on the prediction
+

Contract rule (Zero-Training Contract): any benchmark claiming PRECOG's central promise ("predict without training") must be carried out exclusively in PURE mode. PROBE mode is an explicitly, separately measured extension: it must always be possible to answer the question "how much does PROBE add over PURE alone, for what additional cost?".

+

In PURE mode, the operations allowed on the dataset are limited to descriptive statistics (size, dimensionality, approximate entropy, class imbalance, redundancy, estimated noise) and, if needed, to forward passes without backpropagation or weight updates (to measure activations/Jacobian). No optimizer.step() loop is permitted.

+
+

6. Positioning Relative to the State of the Art

+ + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + +
Line of workContribution to PRECOGAcknowledged limitation
Zero-Cost Proxies (training-free NAS)Fast signals (SynFlow, SNIP, GraSP, Jacob-Cov…) from a mini-batchNo proxy dominates everywhere; correlations vary widely by domain
NEAR (effective rank of activations)Training-free expressivity signal, useful for choosing activation/initializationA single signal, insufficient to predict a full configuration
Initialization theory / dynamical isometryFramework for understanding signal and gradient propagationResults mostly established on simplified cases (deep linear networks)
Meta-learning for HPOReuse of past experiments as a priorStrongly depends on the quality and diversity of the meta-dataset
Bayesian Optimization, Hyperband, BOHB, PBT, ASHA, Vizier, OptunaEfficient search engines under a budgetGenerally start from a weak or null prior; evaluation cost still high without a pre-training signal
Freeze-thaw BO / learning-curve predictionProgressive resource allocation, early stoppingAlready requires partial training observations
+

PRECOG positions itself as an upstream prediction layer for these search engines: they remain used as exploration arms, fed by a far more informed prior.

+
+

7. Fundamental Principles

+
    +
  1. Observe before testing. Any information exploitable without training must be exploited before spending compute.
  2. +
  3. Never depend on a single signal. Each family of signals compensates for another's weaknesses (see §9).
  4. +
  5. Predict distributions, not values. PRECOG returns a probable region with a confidence level, never a point value presented as certain.
  6. +
  7. Learn conditional functions, not constants. E.g. LR* = f(Model, Dataset, Initialization, BatchSize, Optimizer), not "LR = 0.001".
  8. +
  9. Measurable economy. PRECOG only has value if its total cost (analysis + any probes) remains far below the cost of classic HPO.
  10. +
  11. Learn from its mistakes. Every gap between prediction and ground truth is valuable data, kept and exploited, not a result to ignore.
  12. +
  13. Correlation ≠ causation. Relationships exploited in production must, as much as possible, be validated by controlled tests.
  14. +
  15. Generalization above all. A high score on an already-seen benchmark has no scientific value until it is reproduced on tasks, architectures, and datasets never encountered before.
  16. +
+
+

8. Full Architecture

+
                         PRECOG
+                            │
+            ┌───────────────┼────────────────┐
+            ▼               ▼                ▼
+       MODEL ENCODER   DATA ENCODER    HARDWARE ENCODER
+            │               │                │
+            └───────────────┼────────────────┘
+                            ▼
+                     TASK REPRESENTATION
+                            │
+                            ▼
+                    TRAINABILITY ENGINE
+                            │
+            ┌───────────────┼───────────────┐
+            ▼               ▼               ▼
+       Zero-Cost          NEAR          Initialization /
+       Proxies                          Gradient / Jacobian
+            │               │               │
+            └───────────────┼───────────────┘
+                            ▼
+                      REGIME DETECTOR
+                            │
+                            ▼
+                    META-KNOWLEDGE BASE
+                     (meta-dataset + task
+                        embeddings)
+                            │
+                            ▼
+                       META-PREDICTOR
+                     (multi-head ensemble)
+                      /              \
+              Prediction         Uncertainty
+                (distribution)    (calibrated)
+                      \              /
+                            ▼
+                  HYPERPARAMETER DISTRIBUTION
+                            │
+              ┌─────────────┴─────────────┐
+              ▼                           ▼
+        Pareto Search                Search Engine
+       (multi-objective)          (BO / Active Learning /
+                                    Diversity)
+              └─────────────┬─────────────┘
+                            ▼
+                    ADAPTIVE SHORT-PROBE
+                     (PROBE mode, optional)
+                            │
+                    ┌───────┴────────┐
+                    ▼                ▼
+                REJECT            CONFIRM
+                    │                │
+                    ▼                ▼
+              (loop back)     FULL TRAINING
+                                     │
+                                     ▼
+                               GROUND TRUTH
+                                     │
+                     ┌───────────────┴───────────────┐
+                     ▼                                ▼
+              META-DATASET UPDATE               FAILURE ANALYSIS
+                     │                                │
+                     └───────────────┬────────────────┘
+                                     ▼
+                       SCIENTIFIC DISCOVERY ENGINE
+                                     │
+                                     ▼
+                              PRECOG v(n+1)
+
+
+

9. Detailed Components

+

9.1 Model Encoder

+

Extracts a descriptor vector $X_{model}$ from the architecture alone (no data): depth, width, number of parameters, FLOPs, activation type, normalization, residual-connection ratio, attention structure, required memory.

+

9.2 Data Encoder

+

Extracts $X_{data}$ from descriptive statistics allowed in PURE mode: size, dimensionality, entropy, estimated noise, class imbalance, feature correlation, redundancy, distribution. Long-term goal: an embedding $Z_D = \text{Encoder}_{data}(D)$ enabling datasets to be compared by similarity.

+

9.3 Hardware Encoder

+

Captures GPU/CPU, memory, bandwidth, numerical precision, batch capacity, interconnect — because the optimal configuration also depends on the execution environment: $\theta^* = f(M, D, H)$.

+

9.4 Trainability Engine

+

The system's analytical core. Computes, without any weight update:

+
    +
  • Zero-Cost Proxies: SynFlow, SNIP, GraSP, Jacob-Cov, gradient and activation statistics on one or a few mini-batches.
  • +
  • NEAR: effective rank of activations before/after the nonlinearity, as an expressivity indicator.
  • +
  • Initialization analysis: variance of activations and gradients, singular values of the Jacobian $J = \partial f(x)/\partial x$, conditioning $\kappa(J) = \sigma_{max}/\sigma_{min}$, link to dynamical isometry.
  • +
  • Curvature (when measurable at low cost): local Hessian approximations.
  • +
+

Combination rule: $Score_{ZC} = f(S_1, S_2, ..., S_n)$, never a single isolated score.

+

9.5 Regime Detector

+

Classifies the (model, dataset, hardware) tuple into a learning regime (e.g. small model/clean data, large model/noisy data, low data volume, long sequences). Produces a regime prior used to constrain the predicted hyperparameter distribution.

+
(Model, Dataset, Hardware) → Regime → Hyperparameter Prior
+
+

9.6 Meta-Knowledge Base

+

A structured base of all past experiments (see §12), with a task embedding mechanism enabling retrieval of the historical experiments closest to a new task, and using that neighborhood as a search prior (experience transfer).

+

9.7 Meta-Predictor

+

A model (or ensemble of models) taking as input:

+

$$ +X = [X_{model}, X_{data}, X_{ZC}, X_{NEAR}, X_{init}, X_{regime}] +$$

+

and producing, for each candidate configuration, a multi-head prediction:

+
    +
  • $\hat{A}$: expected performance
  • +
  • $\hat{T}$: convergence steps/time
  • +
  • $\hat{C}$: expected compute
  • +
  • $\hat{N}$: data needed
  • +
  • an uncertainty attached to each head (e.g. via ensembles, quantile regression, or Bayesian approaches)
  • +
+

The result is never a single value but a distribution, for example:

+
Learning rate
+  recommended = 3.5e-4
+  range       = [2e-4, 6e-4]
+  confidence  = 91%
+
+

9.8 Search Engine (BO + Active Learning + Diversity)

+

The meta-predictor provides an informed prior; the search engine then explores the remaining space. Hybrid acquisition function:

+

$$ +Acquisition = \alpha \cdot \text{Expected Improvement} + \beta \cdot \text{Uncertainty} + \gamma \cdot \text{Diversity} +$$

+

Google Vizier / Optuna / BOHB play the role here of exploration arms, not the system's brain.

+

9.9 Pareto Search (multi-objective optimization)

+

Rather than seeking a single optimum, PRECOG searches for a Pareto front over (performance, compute, data, time, memory, energy):

+
                  Performance
+                       ▲
+                  A ●
+                    \
+                 B ● \
+                       ● C
+                          \
+                           ● D
+                       └──────────────► Cost
+
+

PRECOG can then return several Pareto-optimal configurations, leaving it to the user (human or system) to choose according to their constraints.

+

9.10 Adaptive Short-Probe (PROBE mode)

+

A short training budget allocated dynamically based on uncertainty and intermediate performance:

+
Candidate A → 50 steps → very poor    → STOP
+Candidate B → 50 steps → promising    → 200 steps
+Candidate C → 50 steps → excellent    → 1000 steps
+
+

Formalization: $Budget_i = f(Uncertainty_i, Performance_i)$. This mechanism relies on learning-curve prediction (freeze-thaw) to estimate a time-to-target and decide CONTINUE/STOP.

+

9.11 Decision Policy

+

An explicit policy turning PRECOG from a simple predictor into an experimental optimization agent:

+

$$ +Policy(s_t) \rightarrow \{\text{TRAIN}, \text{STOP}, \text{EXPLORE}, \text{EXPLOIT}, \text{REQUEST MORE DATA}\} +$$

+

9.12 Causal Discovery Module

+

Separates correlation from causation through controlled experiments: with architecture, dataset, and optimizer fixed, a single candidate variable is varied (e.g. the gradient variance induced by initialization) to observe its isolated effect on convergence, rather than concluding from a simple observational correlation.

+

9.13 OOD / Distribution-Shift Detector

+

Estimates $P(\text{known task})$. If a new task is judged far from the meta-dataset, PRECOG must automatically increase the validation budget (PROBE mode) rather than make an overconfident PURE prediction.

+

9.14 Failure Analysis Engine

+

Categorizes every significant prediction error:

+
DATA_SHIFT
+ARCHITECTURE_SHIFT
+INITIALIZATION_FAILURE
+OPTIMIZER_FAILURE
+PROXY_FAILURE
+PREDICTOR_FAILURE
+
+

and feeds the improvement cycle (meta-dataset → meta-predictor retraining).

+

9.15 Scientific Discovery Engine

+

Longer-term goal: turn observed correlations into hypotheses, test those hypotheses through controlled experiments (see 9.12), and derive general principles of trainability from them (e.g. a candidate relationship $LR^* \approx f(\text{BatchSize}, \text{GradientNoise}, \text{ModelScale})$ to be experimentally verified).

+
Experiments → Patterns → Correlations → Hypotheses
+   → Controlled experiments → Causal evidence → New principle
+
+
+

10. Variables and Hyperparameters

+

10.1 Hierarchy of target hyperparameters (of the trained model)

+ + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + +
LevelFamilyVariables
1Architecturedepth, width, hidden dimension, number of heads, activation, normalization, residual connections
2InitializationXavier, He, Orthogonal, variance/scale, bias init, LSUV
3Optimizationoptimizer (SGD, Momentum, Adam, AdamW, RMSProp, Lion), learning rate, batch size, gradient accumulation, momentum
4Schedulingwarmup, scheduler (cosine, linear, exponential, OneCycle), decay, minimum LR
5Regularizationweight decay, dropout, label smoothing
6Datasampling ratio, augmentation, curriculum, amount of data
+

10.2 PRECOG's internal hyperparameters (strictly distinct from the above)

+ + + + + + + + + + + + + + + + + + + + + + + + + +
ComponentInternal hyperparameters
Bayesian Optimizationacquisition function, exploration/exploitation coefficient, kernel choice, initial observations
Short-Probeinitial number of steps, probe budget, early-stopping threshold, confidence threshold
Active Learningexploration/uncertainty/diversity coefficients
Meta-learningembedding dimension, history size, meta-predictor learning rate
+

10.3 Principle of conditional functions

+

PRECOG never learns a universal constant, only conditional relationships:

+

$$ +LR^* = f(\text{Model}, \text{Dataset}, \text{Initialization}, \text{BatchSize}, \text{Optimizer}) +$$ +$$ +\text{Initialization}^* = f(\text{Architecture}, \text{Dataset}) +$$ +$$ +\text{BatchSize}^* = f(\text{ModelSize}, \text{DatasetSize}, \text{LR}, \text{Hardware}) +$$ +$$ +\text{Optimizer}^* = f(\text{Model}, \text{Dataset}, \text{LR}, \text{BatchSize}) +$$

+

and more generally a joint distribution $P(\theta^* \mid M, D, H)$, with an explicit interaction graph between variables (e.g. LR ↔ BatchSize ↔ gradient noise; Architecture ↔ Initialization ↔ signal propagation).

+
+

11. The Central Concept: Trainability

+

11.1 Operational definition

+

$$ +\text{Trainability} = f(\text{Gradient}, \text{Jacobian}, \text{Activation}, \text{Curvature}, \text{Conditioning}, \text{Initialization}, \text{Architecture}, \text{Data}) +$$

+

11.2 Exploitable signals

+
    +
  • Gradient norm and distribution $\|\nabla_\theta L\|$
  • +
  • Gradient variance $Var(\nabla_\theta L)$
  • +
  • Jacobian $J$, its singular values $\sigma_1, ..., \sigma_n$
  • +
  • Conditioning $\kappa(J) = \sigma_{max}/\sigma_{min}$
  • +
  • Activation statistics $E[a], Var(a)$
  • +
  • Local curvature $H = \nabla^2_\theta L$ (approximated, when cost allows)
  • +
  • Initialization properties and their link to dynamical isometry
  • +
+

11.3 Central research question

+
+

Which signals, observable on an untrained model, actually predict the future speed and quality of learning — and which are merely artifacts correlated with the architecture?

+
+

This question must be addressed both predictively (the meta-predictor) and causally (the causal discovery module, §9.12).

+
+

12. The Meta-Dataset: PRECOG's Scientific Memory

+

Every experiment — including every failure — must be recorded with, at minimum:

+
Experiment
+├── Model        (architecture, depth, width, params, FLOPs, activation, norm.)
+├── Dataset      (size, dimension, entropy, noise, imbalance, diversity)
+├── Hardware     (GPU/CPU, memory, precision, bandwidth)
+├── Initialization
+├── Optimizer, LR, batch size, weight decay, scheduler, warmup
+├── Zero-cost descriptors (SynFlow, SNIP, GraSP, Jacobian, NEAR…)
+├── Training dynamics (gradient norms, loss slope, activation statistics)
+├── Full learning curve
+├── Steps, compute (GPU-hours), memory, time, amount of data, seed
+└── Ground truth (final performance, convergence, real cost)
+
+

Prediction failures are kept and labeled (see Failure Analysis, §9.14): they constitute a learning signal at least as valuable as successes.

+

Strict separation: the meta-dataset is partitioned into TRAIN / VALIDATION / TEST, with the TEST set explicitly locked (never used to improve PRECOG), to avoid benchmark overfitting.

+
+

13. Experience Transfer and Task Embedding

+
                 New Task
+                    │
+                    ▼
+              Task Encoder
+                    │
+                    ▼
+              Task Embedding
+                    │
+          ┌─────────┴─────────┐
+          ▼                   ▼
+    Similar Tasks       Meta-Dataset
+          │                   │
+          └─────────┬─────────┘
+                    ▼
+              Prior Knowledge
+                    │
+                    ▼
+              Optimization
+
+

PRECOG must be able to recognize that a new problem "resembles" a problem already encountered and exploit that similarity as a prior, rather than starting from an uninformed search — this is one of the main expected levers for moving from a merely analytical system to a genuinely intelligent one.

+
+

14. End-to-End Experimental Pipeline

+
                    ┌──────────────────┐
+                    │  BENCHMARK TASKS │
+                    └────────┬─────────┘
+                             ▼
+                    ┌──────────────────┐
+                    │  PRECOG ANALYSIS │  (PURE mode)
+                    └────────┬─────────┘
+                             ▼
+                    ┌──────────────────┐
+                    │ META-PREDICTOR   │
+                    │ prediction +     │
+                    │ uncertainty      │
+                    └────────┬─────────┘
+                             ▼
+                    ┌──────────────────┐
+                    │ SEARCH ENGINE    │  (BO / Active Learning / Pareto)
+                    └────────┬─────────┘
+                             ▼
+                      TOP CANDIDATES
+                             ▼
+                    ┌──────────────────┐
+                    │ SHORT PROBES     │  (PROBE mode, optional)
+                    └────────┬─────────┘
+                     ┌───────┴────────┐
+                     ▼                ▼
+                 PROMISING          POOR
+                     │                │
+                     ▼                ▼
+               FULL TRAINING     STOP / LEARN
+                     ▼
+                 GROUND TRUTH
+                     ▼
+              META-DATASET UPDATE
+                     ▼
+              FAILURE ANALYSIS + RETRAIN
+                     ▼
+               PRECOG v(n+1)
+
+

This loop never stops after a single iteration: every PRECOG generation must be compared to the previous one under a strictly identical protocol.

+
+

15. Test Protocols

+ + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + +
ProtocolQuestionMain metric
P1 — RankingDoes PRECOG rank configurations correctly?Spearman ρ, Kendall τ
P2 — Top-KDoes it retrieve the best configurations?Recall@K
P3 — ConvergenceDoes the chosen configuration converge faster?Steps/Time-to-Target
P4 — ComputeHow much compute is saved?GPU-hours / FLOPs
P5 — Data efficiencySame quality with less data?Samples-to-Target
P6 — GeneralizationDoes it work on a never-seen model/dataset?Out-of-distribution performance
+

15.1 TRAIN/VALIDATION/TEST separation

+
PRECOG TRAIN        → known datasets and architectures, experiment history
+PRECOG VALIDATION   → different datasets, partially new architectures
+PRECOG TEST (locked) → never seen, never used to improve PRECOG
+
+

15.2 Reference benchmarks for the initial phase

+
    +
  • NATS-Bench (successor to the now-deprecated NAS-Bench-201): a reference architecture space with pre-computed performance (CIFAR-10, CIFAR-100, ImageNet16-120) — useful for testing ranking without having to train every architecture oneself.
  • +
  • NAS-Bench-Suite-Zero / JAHS-Bench / HPO-B: actively maintained benchmarks, the first specifically designed to evaluate zero-cost proxies (see stack.md §4 for the rationale behind these choices over the now poorly-maintained HPOBench).
  • +
  • Synthetic laboratory (generated in-house): fully controlled datasets and models (noise, entropy, dimensionality, depth, width), enabling candidate causal variables to be isolated before moving to real benchmarks.
  • +
+

15.3 Multi-seed and statistical tests

+

Every important experiment is repeated over several seeds, with mean, standard deviation, and confidence interval (95% CI) computed. Comparisons between methods (PRECOG vs. Random, vs. BO, vs. Hyperband, vs. Vizier) use appropriate statistical tests (e.g. a Wilcoxon signed-rank test rather than a t-test when parametric assumptions aren't guaranteed), to avoid declaring superiority based on a lucky seed.

+
+

16. Metrics and Objectives (to be demonstrated, not guaranteed)

+ + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + +
MetricDefinitionExperimental target
Ranking correlationSpearman ρ / Kendall τ between PRECOG's ranking and the real rankingρ ≥ 0.80 then ≥ 0.90
Top-K recall$Recall@K = \|\text{PredictedTopK} \cap \text{TrueTopK}\| / K$Recall@10 ≥ 80% then ≥ 90%
Compute reduction$1 - C_{PRECOG}/C_{baseline}$≥ 50% then ≥ 70%
Performance retention$Performance_{PRECOG}/Performance_{oracle}$≥ 99% (or a tolerance defined a priori)
Data efficiency$Samples_{baseline}/Samples_{PRECOG}$ for equal target performance≥ 30–50% reduction, to be refined
Time/Steps-to-TargetReduction in time/number of steps to reach a target≥ 50% reduction
Prediction error (learning curve)$\lvert \text{Prediction} - \text{Actual} \rvert$≈ 5–10% depending on the metric
GeneralizationRecall@K on never-seen tasks/architectures/datasetssame order of magnitude as on known data
+

These targets are progression hypotheses, formalized as successive gates (§17), never presented as already achieved.

+
+

17. Progression Gates

+
                PRECOG
+                   │
+             GATE 1: ρ ≥ 0.70 ?
+                   │
+             GATE 2: Recall@10 ≥ 80% ?
+                   │
+             GATE 3: Compute reduction ≥ 50% ?
+                   │
+             GATE 4: Generalization maintained (never-seen data)?
+                   │
+             GATE 5: Recall@10 ≥ 90% ?
+                   │
+             GATE 6: Compute reduction ≥ 70% ?
+                   │
+             PRECOG "advanced level"
+
+

Each gate is validated by independent metrics, on locked datasets, before considering the next generation.

+
+

18. Comparison Baselines

+

PRECOG must be systematically compared, at equal budget, against:

+
Random Search       Grid Search
+Bayesian Optimization   Hyperband
+ASHA                 BOHB
+Population Based Training
+Google Vizier         Optuna
+Zero-Cost NAS (proxy alone)
+Meta-learning HPO (without PRECOG's additional layers)
+
+

along the axes: final performance, compute, convergence speed, data needed, generalization.

+
+

19. Ablation Strategy

+

19.1 Pipeline component ablation

+
PRECOG-A = Zero-Cost only
+PRECOG-B = + NEAR
+PRECOG-C = + Initialization analysis
+PRECOG-D = + Meta-Learning
+PRECOG-E = + Bayesian Optimization
+PRECOG-F = + Adaptive Short Probe
+PRECOG-G = + Active Learning / Uncertainty
+PRECOG-H = + Causal Discovery / OOD detection
+
+

Expected table example (a template, not real results):

+ + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + +
SystemSpearmanRecall@10Compute used
Random0.1010%100%
ZC0.6055%10%
ZC+NEAR0.6864%12%
+Init0.7370%14%
+Meta0.7977%16%
+BO0.8282%20%
+Adaptive Probe0.8890%30%
+

19.2 Individual proxy ablation

+

SynFlow, SNIP, GraSP, Jacobian, NASWOT, Jacob-Cov, Gradient Norm, NEAR — tested individually then in combination, since the literature shows no proxy is universally dominant.

+

19.3 Robustness testing

+

Deliberate perturbations: dataset noise and imbalance, distribution shift, model depth/width, activation, seed, batch size, hardware — to verify that PRECOG's performance doesn't collapse outside the meta-predictor's training conditions.

+
+

20. Uncertainty Management

+

In addition to a prediction, PRECOG must systematically produce:

+
    +
  • a calibrated uncertainty (via predictor ensembles, quantile regression, or a Bayesian method),
  • +
  • a distinction between model uncertainty (lack of knowledge), data uncertainty (intrinsic ambiguity of the problem), and training stochasticity (variance across seeds).
  • +
+

Example output:

+
Configuration A: prediction = 95%, confidence = 91%
+Configuration B: prediction = 94%, confidence = 52%
+
+

Uncertainty directly feeds the acquisition function (§9.8) and the decision policy (§9.11): an uncertain but potentially informative configuration can be tested with priority to reduce the system's overall uncertainty (active learning).

+
+

21. Causation vs. Correlation

+

A correlation observed between a pre-training signal (e.g. gradient variance) and final performance can be confounded by a third variable (typically the architecture). PRECOG must therefore:

+
    +
  1. Identify candidate relationships from the meta-dataset's correlations.
  2. +
  3. Formulate explicit hypotheses.
  4. +
  5. Design controlled experiments where only the candidate variable changes (architecture, dataset, and optimizer fixed).
  6. +
  7. Only promote a relationship to "knowledge exploitable in production" after causal validation, or otherwise explicitly mark it as "correlation not causally validated".
  8. +
+
+

22. Generalization and Distribution-Shift Detection

+

The generalization test (P6, §15) is considered the most scientifically important. It requires:

+
    +
  • training/validating the meta-predictor on a subset of architectures and datasets, then
  • +
  • testing on architectures and datasets structurally absent from the training set (e.g. train on CNN/MLP/ResNet, test on Transformer).
  • +
+

The OOD module (§9.13) must estimate $P(\text{known task})$ and automatically trigger an increase in the validation budget (PROBE mode) when a task is judged far from the meta-dataset, rather than producing an overconfident PURE prediction out of distribution.

+
+

23. Methodological Risks and Mitigations

+ + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + +
RiskDescriptionMitigation
Data leakageReal dataset information leaking into the PURE analysisStrict Zero-Training Contract (§5), audit of allowed features
Benchmark overfittingPRECOG optimized in a loop on the same benchmarks (NAS-Bench-201, HPOBench…)Locked TEST set, not revealed before final evaluation
Meta-dataset biasOver-representation of certain architectures/domainsDiversification curriculum, explicit tracking of meta-dataset coverage
Undetected distribution shiftPRECOG applied outside its domain of validity without warningOOD module (§9.13) + adaptive validation budget
Training stochasticityConfusing seed variance with a configuration's real effectMandatory multi-seed runs, confidence intervals (§15.3)
Poorly calibrated uncertaintyDisplayed confidence not reflecting the real errorRegular calibration, calibration tests (e.g. reliability diagrams)
PRECOG's own excessive costAnalysis cost exceeds the savings achievedSystematic measurement of $Cost_{PRECOG} + Cost_{PROBE}$ vs. $Cost_{classic\ HPO}$ (§24)
Misleading correlationA relationship exploited in production isn't causalCausal discovery module (§21)
Dependence on one architecture familyGood performance only on the meta-dataset's architecturesProgressive curriculum (MLP → CNN → Transformer → unknown), strict generalization tests
+
+

24. System Economics

+

PRECOG only has practical value if:

+

$$ +Cost_{PRECOG} + Cost_{PROBE\ if\ any} \; \ll \; Cost_{classic\ HPO\ or\ multiple\ FULL\ TRAININGs} +$$

+

This constraint must be measured at every evaluation, not merely assumed. A system that is theoretically accurate but whose inference is too costly (e.g. a meta-predictor that itself requires enormous compute) must be considered an economic failure, even with a good ranking score.

+
+

25. Development Roadmap

+
V1 — Foundations
+  Learning Rate, Batch Size, Optimizer, Initialization
+  (basic zero-cost analysis, no meta-learning)
 
-    
-

03The number we got wrong, in public

-

- Our first controlled experiment (n=36: 12 tasks × 3 init methods) - found gradient_norm_variance correlating with real convergence - speed at ρ=+0.670. We used that number in downstream decisions. -

-
It didn't hold. Rechecked at n=936 (312 tasks, identical design, zero new training runs — the data already existed): ρ=0.395. The corrected best individual proxy is gradient_norm (ρ=0.540).
-

- We saw the same shrinkage independently for learning-rate prediction: - ρ=−0.726 at n=12 dropped to ρ=−0.343 at n=40. We don't - think this is unusual for the field — we just don't see it checked often. -

- Full correction → -
+V2 — Full configuration + Weight Decay, Warmup, Scheduler, Gradient Accumulation -
-

04A bug we found but can't fix

-

- jacob_cov never once recommends "He init" across 60 test - tasks — including the 10 where it's genuinely fastest. The cause is - structural: Xavier and He (zero-biased networks) draw from the same - underlying Gaussian values, differing only by a positive scale — and - jacob_cov only reads activation sign, which scaling - cannot flip. -

-
jacob_cov(He-init network)  ==  jacob_cov(Xavier-init network)
-

- exactly, for the same seed — checked across all 312 tasks, max difference - is 0.0. Two other proxies (effective_rank, - jacobian_condition_mean) share the property for the same - reason. Two attempted fixes (raw and population-normalized tie-breaking - with gradient_norm) both failed. -

- Full audit of all 11 proxies → -
+V3 — Architecture + Dropout, Architecture (depth/width/activation/normalization) -
-

05Reproduce it

-
pip install -r requirements.txt
-python scripts/gate1_ranking.py              # the original (overestimated) result
-python scripts/gate1_ranking_at_scale.py     # the correction, no retraining needed
-python scripts/explore_scale_invariance_blindspot.py
-python scripts/compare_meta_predictors.py    # full method comparison + regret
-

- data/meta_dataset.db ships with the repo (312 tasks, 936 - controlled experiments) so every scale-corrected result reproduces - without rerunning any training. CI reruns all of the above on every push. -

-
+V4 — Intelligence + Meta-learning, Task Embeddings, NEAR, combined Zero-Cost proxies -
-

06References

-

- Every method named above, linked to the actual paper — not a survey - mention, the source: -

- - - - - - - - - - - - - -
MethodPaperUsed for
NASWOT / jacob_covMellor et al., 2021the best-evidenced method (§2)
SynFlowTanaka et al., 2020zero-cost proxy, weak on this benchmark
SNIPLee et al., 2019zero-cost proxy
GraSPWang et al., 2020zero-cost proxy
NEAR / effective rankHusistein et al., 2025expressivity proxy
ZiCoLi et al., 2023tested, rho=0.328, not promoted
AZ-NAS (rank aggregation)Lee et al., CVPR 2024combination method, §2 above
LSUVMishkin & Matas, 2015data-aware init, tested, underperformed
Zero-Cost Proxies surveyAbdelfattah et al., ICLR 2021basis for the proxy set in precog/trainability.py
-

- Full bibliography, including hyperparameter optimization, meta-learning, - dynamical isometry, and active-learning literature not directly - implemented yet: - source.md. -

-
+V5 — Adaptive search + Active Learning, Bayesian Optimization, Adaptive Short-Probe -
-

07In the same spirit

-

- We're not affiliated with it, but - karpathy/autoresearch - is the closest thing we've seen to the same underlying instinct: fix a - tight, cheap experimental loop (there, a 5-minute training budget and one - metric; here, controlled synthetic tasks and a locked test split), then - let the loop run fast enough, and honestly enough, that what survives - it is worth trusting. Different problem — that project is about an - autonomous agent iterating on training code overnight, this one is - about predicting hyperparameters before training starts — but the same - bet: rigor and speed aren't in tension if the loop is small enough to - run a lot, and log every result, wins and failures alike, without - editorializing. -

-
- -
-

- Not a finished system — no success criterion in the full spec is met yet. - This is an honest, ongoing research snapshot, not a solved problem. -

-

- Repository · - Contributing · - Citation · - MIT License -

-
+V6 — Science + Causal Discovery, OOD Detection, Failure Analysis, + Scientific Discovery Engine +
+

Scientific progression by phase (indicative)

+
Phase A: analytical foundations (Zero-Cost, NEAR, Initialization)
+Phase B: meta-learning + Bayesian Optimization
+Phase C: uncertainty + active learning + adaptive acquisition
+Phase D: learning-curve prediction + adaptive probe + failure analysis
+Phase E: validation — never-seen tasks, multi-seed, statistical tests, reproducibility
+
+

Experimental curriculum

+
Level 1: MLP on synthetic datasets
+Level 2: CNN on vision
+Level 3: ResNet / modern architectures
+Level 4: Transformers
+Level 5: LLM fine-tuning
+Level 6: never-seen models and datasets (ultimate generalization test)
+
+
+

26. Success Criteria

+

A PRECOG milestone is only considered reached if, simultaneously, on a locked test set never used for training:

+
    +
  1. the ranking (Spearman ρ) reaches the target threshold for the level considered,
  2. +
  3. Recall@K reaches the target threshold,
  4. +
  5. the measured compute reduction reaches the target threshold,
  6. +
  7. the retained final performance stays within the tolerance for loss defined a priori,
  8. +
  9. results hold on never-seen tasks/architectures/datasets (generalization),
  10. +
  11. results are reproducible (multi-seed, confidence intervals, documented environment).
  12. +
+

A system that reaches only part of these criteria (e.g. good ranking but poor generalization) is not considered to have reached the milestone.

+
+

27. Known Limitations

+
    +
  • Generalization to radically new architecture families (beyond those represented in the meta-dataset) is not guaranteed and must be treated as a hypothesis to test, not as a given.
  • +
  • Current zero-cost signals from the literature are not universally reliable; combining them reduces but does not eliminate the risk.
  • +
  • The meta-dataset's quality intrinsically bounds the meta-predictor's quality: a poorly diversified meta-dataset will produce overly optimistic predictions outside its real coverage.
  • +
  • PROBE mode introduces a real cost, even if minimal; any claimed gain must be net of this cost.
  • +
  • The causation/correlation distinction remains partial: some exploited relationships will in practice remain robust correlations rather than demonstrated causes, and must be presented as such.
  • +
+
+

28. Outlook

+

In the longer term, PRECOG's scientific ambition goes beyond HPO: the goal is to build an operational theory of predictable learning dynamics, i.e. a function

+

$$ +F : (\text{Model}, \text{Data}, \text{Initialization}, \text{Hyperparameters}) \rightarrow \text{Training trajectory} +$$

+

able to anticipate the loss trajectory $L(t)$ before full training. If this direction succeeds, PRECOG would stop being just a hyperparameter optimizer and become a predictive model of learning dynamics, with potential for its own scientific contribution (beyond integrating existing tools).

+
+

29. Production Architecture (long-term target)

+

PRECOG, as a platform, must be able to:

+
    +
  1. Receive an untrained model, the dataset's allowed metadata/statistics, and a description of the hardware environment.
  2. +
  3. Run an analysis in PURE mode (no weight update on real data).
  4. +
  5. Produce a hyperparameter distribution with justification and confidence level, as well as a Pareto-optimal set of configurations according to constraints (performance/compute/data/time).
  6. +
  7. On request, validate the best hypotheses via a minimal budget in PROBE mode.
  8. +
  9. Systematically log the experiment (including production usage) into the meta-dataset, for continuous improvement.
  10. +
+
Untrained model + Dataset (stats) + Hardware
+                    │
+                    ▼
+              PRECOG (PURE mode)
+                    │
+                    ▼
+     Hyperparameter distribution + confidence
+                    │
+                    ▼
+        Pareto-optimal set of configurations
+                    │
+             (optional) PROBE
+                    │
+                    ▼
+        Recommended configuration + justification
+
+
+

30. Synthesis — the idea that distinguishes PRECOG from classic HPO

+
+

PRECOG does not simply search for the best hyperparameters after training many configurations; it seeks to learn the relationship between a model's initial state, the properties of the problem, and the learning conditions, in order to predict — before any training on real data — which configurations have the highest probability of leading to fast, efficient convergence.

+
+

Every evaluation, every benchmark, and every scientific communication about PRECOG must come back to this test: does the system provide information exploitable before training, that is measurable, generalizable, and economically justified — or does it merely reproduce classic HPO dressed up differently?

+
+ + From 0078795f9d68bb524f1666ed7f03d66012878d51 Mon Sep 17 00:00:00 2001 From: rustnew Date: Thu, 3 Sep 2026 22:02:28 +0100 Subject: [PATCH 2/2] Polish the spec page for a public launch (Reddit etc.) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Content is unchanged (still the full docs.md); this is a design pass requested ahead of sharing the link widely, so it reads as a finished project page rather than a raw markdown dump: - Section numbers pulled out of h2/h3 text into accent-colored chips (the same design language the original curated landing page used), instead of plain "3. Scientific Problem" text. - Proper type scale (Inter for body copy, JetBrains Mono for labels/ code -- both Google Fonts), tighter line length on the content column, clearer spacing rhythm between sections. - Hero rewritten: subtle background treatment, primary/secondary link buttons (repo / verified results / raw source), and an explicit methodological-note callout up front so the spec-vs-results distinction is visible before a reader scrolls past it in §0. - Sticky TOC now highlights the section currently in view (IntersectionObserver, no dependency). - Wide tables wrap in a horizontally-scrollable container instead of overflowing the fixed-width column on narrow screens. - Open Graph / Twitter Card meta tags, so a shared link renders a proper title and description instead of nothing. Co-Authored-By: Claude Sonnet 5 Claude-Session: https://claude.ai/code/session_01WSc9sb1otU6ssfxDBzeNHG --- docs/index.html | 509 +++++++++++++++++++++++++++++------------------- 1 file changed, 312 insertions(+), 197 deletions(-) diff --git a/docs/index.html b/docs/index.html index ead1517..9540784 100644 --- a/docs/index.html +++ b/docs/index.html @@ -4,7 +4,17 @@ PRECOG — Predictive Configuration & Trainability Engine - + + + + + + + + + + + @@ -12,117 +22,198 @@
-
-

Full specification · docs.md

-

PRECOG

-

Predictive Configuration & Trainability Engine — vision document, scientific specification, and research roadmap.

-
+

Full specification · docs.md

+

PRECOG: Predictive Configuration & Trainability Engine

+

Can signals computed on an untrained network — before a single optimizer step — predict which hyperparameters will actually train well? The complete vision document, scientific specification, and research roadmap.

+

Methodological note: this is a research specification, not a results report. Section 0 below states plainly which numbers are experimental targets versus what has actually been measured — see the README for verified findings, including a corrected result and a documented bug.

@@ -212,7 +303,7 @@

PRECOG

-