Repository navigation
Conversation
A registered metric is now five fields: name, function, family, cost and direction. The first line of its docstring is the question the column answers, written so it reads without the metric's name. Direction None means a diagnostic -- read beside the other columns, never ranked on -- which is what `dpp` becomes: three monotone transforms of one determinant that collapses toward underflow before the models stop differing. Where a metric is computed and how per-design values collapse to one number are no longer properties of the metric. `Board.evaluate` takes `space` (pixel, or a PCA of the reference designs) and `aggregation` (mean or median), so every metric is written once; the spec records the aggregation in each row and the `*_median` columns are retired. The board also scores half of the reference designs against the other half, which is the only honest measurement of what real designs score on a column. Port the math layer from the workshop branch (vendi, geometric-mean DPP, median-sigma, PRDC), fix `register_metric` silently targeting the global registry when handed an empty one, and replace "brief" with "conditions" throughout, including three PR #75 docstrings. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Ten code cells, one call each: what metrics exist, score five models, read the board, rank, change the aggregation, change the space, register a metric. Two of the models are constructions with known right answers; the one that returns the correct withheld designs under permuted conditions scores a perfect MMD, which is the lesson. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Twenty metrics, each one question: the set-level ones (mmd, coverage, vendi, dpp), the per-condition ones (per_condition_distance, volume_error), feasibility, memorization (train_distance, copy_rate), cond_sens, the optimality gaps, four call-budget metrics that price a defect in optimizer calls (calls_to_settle, gap_after_calls, reaches_reference_rate, first_call_gain), and three cost columns. Names say what they measure. `novelty` split into `train_distance` (nearest training design) and `copy_rate` (the submission gate, against everything the model could have copied); `cond_err` became `volume_error`, since it was never about conditions in general; `recovery` became `gap_after_calls`, and `parity_rate` became `reaches_reference_rate`. `calls_to_parity` is gone: a copier reaches parity in zero calls. `dpp` keeps its name and becomes the n-th-root form, since the raw determinant reads 1e-20 on every real board. The context now carries each design's re-optimization trajectory, the generator's parameter count and training time, so the new metrics have what they read. Specs move to v2 with the full list and the aggregation policy; v1 stays committed so published rows can be read under the protocol that produced them, and only the current spec must name live metrics. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
`Board.evaluate(space=...)` looks the space up in a registry: `pixel` and `pca` are built in, and a learned latent space registers one projection function. Metrics that need the actual designs declare `pixel_only` and are skipped elsewhere. The kernel bandwidth defaults to the median pairwise distance of the reference designs in the chosen space, so the diversity columns can see anything on a problem the spec never pinned a value for. `Board.from_generators` scores loaded models through an `Evaluator`, which is what cond_sens, the cost columns and the physics need, and `Board.load` pulls published checkpoints by name -- the workshop in one call. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The docs named `novelty`, listed the old default pass, and pointed at spec v1. They now name the columns that exist, say what --list-metrics prints, and send readers to the notebook for the rest. `Board` is exported from `engiopt.evaluation` as the front door. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
`ruff format --check` covers .ipynb files and the notebook had never been through the formatter. Re-executed after formatting so the stored outputs match the source. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
CI runs `ruff check .` over notebooks too. The first cell set pandas display options before importing the registry, two cells had imports out of the repo's order, and the toy problem's helpers lacked docstrings. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The first docstring line is the only text a reader gets from explain(), and several read as yes-or-no questions when the column is a quantity with a direction. Each line now names the quantity, says how it is computed in a phrase, and ends with what the two ends of the scale mean. Per-design metrics say "for each design" rather than "average", since the board's aggregation decides which that becomes. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The notebook scored a toy problem, which left cond_sens blank, made every copy rate zero or one, and let a reference-copying oracle pass for a model. It now loads four beams2d checkpoints from the Hub, builds three constructions from the dataset, and scores all seven under one spec with the physics columns included. Twelve conditions rather than the published fifty and medians rather than means, both declared in the notebook, so it runs in about seven minutes on a laptop. To score checkpoints and saved designs together, the evaluator gains `context_for_designs`, and `Board.from_evaluator` scores both under the evaluator's spec and keeps the sampled designs so they can be re-scored in another space. `Board.load` takes a config fingerprint per model, since two of the four families have no canonical checkpoint. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
`copy_rate` counted designs within a per-element tolerance of a corpus that held the training split plus the reference optima. No model can see the reference optima, so that half of the corpus caught nothing real, and the tolerance was an arbitrary number in design units deciding who counts as a copier. Both go. `train_distance` now emits two readings of one measurement against the full training split: the per-element distance to the nearest training design, and that distance divided by what the withheld reference designs score, so one means as far from the training data as real unseen designs are and zero means retrieval. The reference designs supply the scale and nothing else. The spec loses its three copy knobs, including the 512-design subsample that let a lookup table through half the time, and the board no longer flags rows for memorization: the columns are read, not enforced. The v1 specs lose the same three keys; their condition digests are untouched. The notebook is re-run on beams2d with the new columns. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The memorization change removed the spec from the publishing path, since no flag reads it any more, and this call site was missed. mypy caught it. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Left over from the previous fix; ruff flagged it as unused. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
SoheylM
left a comment
There was a problem hiding this comment.
Thanks, Matthew. I rechecked the current head. The default-spec mismatch and the saved-design path look like issues to address before merging; I left the smaller API and comparison points inline as well.
|
|
||
| problem_id: str | ||
| version: str = "v1" | ||
| version: str = "v2" |
There was a problem hiding this comment.
Could we align the v2 default across the public entry points? EvalSpec.load("beams2d") selects v2 here, but Evaluator.for_problem() and the CLI still select v1 when no spec is given. That v1 spec includes the removed novelty metric, so the default Board.load() and CLI paths fail when they score. EvalSpec(problem_id=...) also defaults to v2 while retaining novelty in its metric list.
There was a problem hiding this comment.
Fixed in cc5f16e. Evaluator.for_problem, the CLI and EvalSpec.load now all resolve to v2 when given no spec, and the dataclass default metric list names only registered metrics. Test: test_the_default_spec_metrics_are_all_registered.
There was a problem hiding this comment.
Thanks, the EvalSpec and Evaluator.for_problem defaults are fixed. The CLI still has spec = args.spec or f"{args.problem_id}/v1" in engiopt/evaluate.py:344, so running it without --spec still selects v1. I confirmed that selecting the metrics in beams2d/v1 raises KeyError: Unknown metric 'novelty'. Could the CLI use the same v2 default, with a regression test for its no---spec path?
There was a problem hiding this comment.
Reproduced first: the v1 spec still lists novelty, which the registry no longer serves, so the CLI's no-flag path raised the KeyError you saw. Fixed in 481fd39. The CLI no longer builds its own reference. --spec passes straight through and EvalSpec.load owns the only default, so the CLI cannot drift from the library again. tests/test_evaluate_default_spec.py pins both halves: main without --spec must hand spec=None to Evaluator.for_problem, and a bare problem reference must load the dataclass default version. The two docstrings that still said v1 are swept too.
| sigma = self.sigma if self.sigma is not None else metrics_mod.compute_median_sigma(reference_codes) | ||
|
|
||
| def score(generated: Any, against: Any, chosen: list[Any]) -> dict[str, float]: | ||
| ctx = EvaluationContext( |
There was a problem hiding this comment.
Could this path take the conditions that belong to the saved designs, or skip metrics that require them? On beams2d, the default viol check calls check_constraints without the required volfrac and raises; volume_error has no target and returns NaN. As written, the documented Board(problem, reference=...).evaluate(...) path does not work with its default metrics on that problem.
There was a problem hiding this comment.
Fixed in cc5f16e. Board takes conditions= and passes them into the context for the model rows, so the documented path works on beams2d when the conditions are given. Without them, viol now returns blank instead of raising, and its docstring says so; volume_error already returned NaN. Test: test_viol_is_blank_without_the_designs_conditions.
There was a problem hiding this comment.
Thanks, this prevents the viol crash. volume_error still cannot produce a value on the saved-design path: Board has no volume_condition field, and its new EvaluationContext(...) does not set one. I reproduced Board(..., conditions={"volfrac": ...}).evaluate(..., metrics=["volume_error"]) returning NaN even though the targets were supplied. Could Board pass an explicit volume-condition name, and carry it over from evaluator.spec in from_evaluator?
There was a problem hiding this comment.
Reproduced: your call returned NaN even with targets supplied, because the board never told its contexts which condition is the budget. Fixed in 481fd39. Board gains an explicit volume_condition field, declared rather than guessed, the same contract as EvalSpec, and passes it into every context it builds; from_evaluator carries it over from evaluator.spec, so re-scoring kept designs in another space still sees the budget. On the test problem, designs at their own budgets now score 0.0 and designs shifted by 0.05 score 0.048. test_volume_error_reaches_saved_designs pins both directions: a board told the budget scores it, and a board not told leaves the column blank rather than guessing a column name.
| f"{column!r} is a diagnostic: it says whether the other columns mean what they look like. " | ||
| f"Read it beside them; do not rank on it." | ||
| ) | ||
| return self.frame.sort_values(column, ascending=not spec.higher_is_better) |
There was a problem hiding this comment.
Could rank() exclude REFERENCE_ROW, as explain() already does? Right now the split-half comparison is sorted alongside the models and can appear as the winner, although it is not a model.
There was a problem hiding this comment.
Fixed in cc5f16e. rank drops the split-half row the same way explain does. Covered by test_the_reference_row_is_measured_and_never_picked.
| # Two halves of the reference set answer different conditions, so the | ||
| # per-condition metrics have no meaning there and are left blank. | ||
| unpaired = [spec for spec in specs if spec.family != "conditions"] | ||
| rows[REFERENCE_ROW] = score(reference[order[:half]], reference[order[half:]], unpaired) |
There was a problem hiding this comment.
Is this meant as a diagnostic baseline rather than a directly comparable score? Model rows use the full generated batch against the full reference set, while this row uses half against half. MMD and coverage depend on sample count, so I would avoid calling its values "what real designs score" without making that difference clear.
There was a problem hiding this comment.
Agreed. It is a scale reference at half the sample size, not a comparable row. The explain column is now named split-half reference, and the row's docstring and the notebook say it shows the order of magnitude a good model should reach, is not a number to beat, and is never ranked.
| ctx = evaluator.context_for_designs(batch) | ||
| rows[label] = evaluator.score_context(ctx, only=metrics, include_expensive=expensive) | ||
| sampled[label] = ctx.gen_designs | ||
| board = cls(problem=evaluator.problem, reference=evaluator.resolved.ref_designs, sigma=evaluator.spec.sigma) |
There was a problem hiding this comment.
Could the returned board keep evaluator.registry? The evaluator can score a custom metric, but this Board defaults to the global registry. Its custom column is then omitted by explain(), and rank() cannot recognize it.
There was a problem hiding this comment.
Fixed in cc5f16e. from_evaluator builds the board with evaluator.registry, so a custom metric the evaluator scored is readable in explain and rankable. Test: test_from_evaluator_keeps_the_evaluators_registry.
|
|
||
|
|
||
| @functools.cache | ||
| def _parameter_count(generator: Any) -> int | None: |
There was a problem hiding this comment.
Could we count parameters without caching by generator object? functools.cache keeps every scored generator, including its network, alive for the lifetime of the process. A sweep over many checkpoints could retain a lot of CPU or GPU memory for a calculation that is cheap to repeat.
There was a problem hiding this comment.
Fixed in cc5f16e. The functools.cache decorator is gone; the parameter count is recomputed per score and holds no reference to the generator.
…aining data `calls_to_settle` used a band of five percent of each design's own reach, so a start of 1e9 that dropped to 1e3 in one call read as settled after one call while a VQGAN start of 2.4 took 32. Matthew caught it. The column is now `calls_to_near_optimum`: calls until the gap stays within five percent of the reference optimum's objective for that design's conditions, with one more than the budget for a design that never arrives. The physics pass keeps the reference objective per design to supply the scale. `first_call_gain` is gone; it was a fraction of the guess's own improvement and had the same flaw with no optimum-relative form. The kernel bandwidth behind mmd, vendi and dpp is no longer a pinned 10.0. It defaults to the median pairwise distance of the training designs, in whatever space the context holds, resolved once per context and recorded in every row as `kernel_sigma`. A spec or a Board may still pin a number; the v2 specs pin none. On beams2d the pinned value had every model near ten effective designs; the data-sized kernel separates the collapsed GAN at 1.03 from the rest at five to seven. The two changes share files, so they ship together. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…tric The register-a-metric cell used attributes nobody had introduced. It now names what the context holds, with shapes, before the example, and the example reads each step in a comment. Re-run with the new settle column and the data-sized kernel; the interpretation reflects both. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…gns, and a board that keeps its registry Six findings from Soheyl on the current head, each reproduced from the code before changing it. The default spec disagreed with itself: `Evaluator.for_problem` and the CLI loaded v1 when given no spec, `EvalSpec.load` chose v2, and a bare `EvalSpec(problem_id=...)` still listed `novelty`. All three now resolve to v2, and the dataclass default names only registered metrics. The saved-designs path never carried the designs' conditions, so on beams2d `viol` reached the problem's constraint check without the volume fraction it reads and raised. `Board` now takes the conditions the reference designs answer and passes them through; `viol` reports blank rather than raising when none were given, and a test fixture that relied on the raise-free fake now supplies conditions. `rank` no longer lists the split-half reference row, which `explain` already excluded, and the row's docstring and `explain` column say what it is: a scale reference at half the sample size, not a number to beat. A board built from an evaluator keeps that evaluator's metric registry, so a custom metric it scored is readable and rankable. The parameter count is no longer memoized by generator object, which had kept every scored network alive for the life of the process. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
|
Thanks, Soheyl. All six are addressed in cc5f16e, each reproduced from the code first. CI on that commit: 309 passed, ruff and mypy clean. The notebook is being re-run with the renamed explain column and lands in the next push. |
The stored outputs still showed the old column name for the split-half row. Same numbers as before; the physics is unchanged. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
| gen_designs=project(np.asarray(generated)), | ||
| ref_designs=project(np.asarray(against)), | ||
| conditions=self.conditions if against is reference else None, | ||
| sigma=self.sigma, |
There was a problem hiding this comment.
When train and sigma are both omitted, kernel_sigma is inferred separately from each context's ref_designs. Model rows use the full reference set, but the split-half row uses only the other half, so its MMD/DPP/Vendi values use a different bandwidth as well as a different sample size. In a simple two-cluster, eight-design example, the bandwidths are 14.14 and 1e-6. Could we resolve the bandwidth once from the full reference set (or training set) and pass that same value to every row?
There was a problem hiding this comment.
Reproduced: with no training split, a context whose reference half lands in one cluster infers the 1e-6 floor while the full set gives 6.3 on the test problem, the same divergence as your example. Fixed in 481fd39. Board.evaluate now resolves the bandwidth once, from the training designs when given and otherwise from the full reference set, in whatever space is being scored, and hands the same value to every context, the split-half row included. One scope note: when a training split is supplied, every row already sized its kernel from that split, so this changes only boards built without train; the notebook's PCA board is unaffected and its stored outputs stand. test_every_row_shares_one_inferred_bandwidth asserts the median heuristic runs exactly once per evaluate, on the full reference set.
… knows, one bandwidth per board Three fixes, one per comment on the second review round. The CLI still built its own default spec as problem_id/v1 after the library default moved to v2, so a bare run selected a metric list the registry no longer serves. The CLI now passes --spec through and EvalSpec.load owns the only default, with a regression test pinning the no-flag path. Board never told its contexts which condition is the volume budget, so volume_error returned NaN on the saved-design path even when targets were supplied. Board gains an explicit volume_condition field, declared rather than guessed, and from_evaluator carries it over from the spec. When no bandwidth was pinned, each row inferred kernel_sigma from its own reference set, so the split-half row was scored under a different bandwidth as well as a different sample size; a half landing in one cluster hit the 1e-6 floor. evaluate now resolves the bandwidth once, from the training designs or the full reference set in the scored space, and every row shares it. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
Thanks, Soheyl. All three are addressed in 481fd39, each reproduced from the code first. CI on that commit collected 313 tests and passed 313, with ruff, mypy and pre-commit clean. The four new tests pin the no-flag CLI path, the volume budget on the saved-design path, and the single shared bandwidth. |
| def test_volume_error_reaches_saved_designs(fake_problem: Any, designs: Any) -> None: | ||
| """A board told which condition is the volume budget scores it; one not told leaves it blank.""" | ||
| models, ref, _ = designs | ||
| conditions = {"volfrac": [float(design.mean()) for design in ref]} |
There was a problem hiding this comment.
This column dictionary works for volume_error alone, but the same Board fails on a normal evaluate({"exact": ref.copy()}): default metrics include viol, and EvaluationContext.condition_at(0) calls conditions[0], raising KeyError: 0. I reproduced that path. Could we normalize column mappings to row-indexable conditions, or make the accepted conditions type explicit and use that type in this test?
There was a problem hiding this comment.
Reproduced: viol reads conditions row by row (conditions[0]), which a plain dict of columns answers with KeyError: 0, so the input my test blessed crashed a default evaluate. Fixed in 876e094, taking your first suggestion and one step past it. ConditionTable normalizes a dict of columns or a DataFrame at the board boundary so both access patterns work; test_dict_conditions_survive_a_full_evaluate is your reproduction as a test. A conditions object whose length disagrees with the reference is now refused with a message saying conditions are an answer key, not a filter, since a short one silently grades design i against the wrong brief. And because the briefs already sit beside the optima in the dataset, Board(reference=rows) now accepts the test rows directly and splits them itself, so the two cannot misalign; Evaluator.for_rows extends the same idea to the sampling path for a user-chosen slice, with the digest cleared since a custom draw is not the frozen contract. CI on 876e094: 319 passed, ruff, mypy and pre-commit clean.
…en slice The reviewer showed that a plain dict of condition columns feeds volume_error and crashes viol, because the two metrics read conditions differently: one by column name, one by row. The deeper problem is that the briefs already sit beside the optimal designs in the dataset, and slicing the designs into a bare array is what loses them. Three pieces, smallest surface first. ConditionTable wraps a dict of columns or a DataFrame so both access patterns work, and a conditions object whose length disagrees with the reference is refused, since a short answer key silently grades design i against the wrong brief. Board(reference=rows) accepts a dataset slice directly, splitting it into designs and briefs so the two cannot fall out of alignment. Evaluator.for_rows swaps the spec's seeded draw for rows the user chose, keeping the rest of the spec in force; generators are sampled under exactly those conditions, the old evaluate-script flow. The digest is cleared on that path, because it certifies the frozen draw and a custom slice is not leaderboard material. Board.load grows a rows argument that routes through it. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
What this is
After #75 the evaluation layer on
mainregistered eight metrics. There was no way to score several models at once without assembling anEvaluationContextby hand. The full suite lived onfeat/dataset-sensitivity-fieldsas about forty separately maintained functions with prefixed names such aspca_mmd,lv_vendiandiog_median. This PR brings the suite tomainon a contract that fits in one sentence, and adds the one object you use it through.Nothing here touches the LVAE, the baselines, or the pool tooling. Those are separate PRs.
How you use it
Two of the four beams2d families were never trained at the script's default configuration. They load by the fingerprint of the configuration that was trained. Diffusion and VQGAN load by name.
example_metrics_suite.ipynbat the repo root walks all of this on beams2d. It loads four published checkpoints, builds three constructions from the dataset, and scores all seven under one spec with the physics columns included. It uses twelve conditions rather than the published fifty and medians rather than means. Both are declared in the notebook, and it runs in about seven minutes on a laptop.The contract
A metric is a function of an
EvaluationContextplus four declarations: name, family, cost and direction. The few metrics that need the actual designs rather than codes in some space say so withpixel_only. The first line of the docstring is the sentenceexplain()and--list-metricsprint.higher_is_better=Nonemarks a diagnostic, read beside the other columns and never ranked on.Two things that used to be baked into metric names are arguments to
Board.evaluate.pixelandpcaare built in. A learned latent space registers one projection function withregister_space. This replaces thepca_*,lv_*,lvoff_*andpixel_*prefixes.*_mediancolumns.Metrics call
ctx.reduce(values)and never hard-code a mean.The metrics
Nineteen, down from about forty. Each name says what is measured.
mmd,coveragevendi,dppcond_sens,per_condition_distance,volume_errortrain_distance,train_distance_ratio(one metric, two outputs)violiog,cog,fog,calls_to_near_optimum,gap_after_calls,reaches_reference_rategeneration_seconds,n_parameters,train_minutesFour definitions changed on the way. Each change has a reason a reviewer should know.
Memorization is a distance, not a rate. The old
copy_ratecounted designs within a per-element tolerance of a corpus that included the reference optima. No model can see the reference optima, and the tolerance was an arbitrary number in design units.train_distancereports the per-element distance from each design to the nearest design in the full training split.train_distance_ratiodivides that by what the withheld reference designs score against the same split, so one reads as data-like and zero as retrieval. There is no tolerance anywhere and no automatic flag. The columns are read by people. The spec dropscopy_tol,max_copy_rateandcopy_corpus_size. The last of those was a 512-design subsample that had let a lookup table through half the time.Settling is measured against the optimum, not the guess.
calls_to_near_optimumcounts optimizer calls until a design's gap stays within five percent of the reference optimum's objective for its conditions. A design that never gets there counts one more than the budget. An earlier form set the band from the design's own starting gap, which read a start of 1e9 as settled after one call. That form is gone, and so isfirst_call_gain, which had the same flaw.calls_to_parityis gone too, because a copier satisfies it in zero calls.The kernel bandwidth comes from the data.
mmd,vendianddppdefault to the median pairwise distance of the training designs. The value is resolved per evaluation and recorded in every row askernel_sigma. The v2 specs pin no value. The old pinned 10.0 put every beams2d model near ten effective designs. The data-sized kernel separates the collapsed GAN at 1.03 from the rest at five to seven.dppis readable. It keeps its name and becomes the n-th root of the kernel determinant, bounded in (0, 1]. The raw determinant that v1 rows hold reads around 1e-20 on every real board and cannot be compared with this column.vendiis the diversity column to prefer, because a determinant still rises when a collapsed set is jittered.The board from saved designs also scores one random half of the reference designs against the other half. That row is labeled
reference (split-half). It is a scale reference at half the sample size, not a competitor. It shows where real designs land on each column, it is never ranked, andexplain()reports it undersplit-half reference.Specs
All four problems get a
v2.jsonwith the full metric list and the aggregation policy. v2 is the default everywhere a spec is resolved.v1.jsonstays committed so rows published under it can be read. Evaluating under v1 raisesUnknown metric 'novelty'at metric selection, because that metric no longer exists. The leaderboard dataset holds no v1 rows.Also in here
OptimizationResults.trajectoriesandreference_objectives, one gap path and one reference objective per design.EvaluationContext.reduce,kernel_sigma,model_params,train_minutesandtrain_designs.Evaluator.context_for_designs, so constructions built from the dataset score under the same spec as loaded models.Board.from_evaluatorscores both and keeps the sampled designs and the evaluator's metric registry.register_metricsilently registered into the global registry when handed an empty custom one, becauseMetricRegistryis aMappingand an empty mapping is falsy. Fixed.LEADERBOARD.md,README.mdandCONTRIBUTING_A_MODEL.mddescribe the suite that exists now.Review
Soheyl's six inline comments on
59c198dwere each reproduced from the code and fixed incc5f16e. Replies are inline, each naming its test.Evaluator.for_problem, the CLI andEvalSpec.load.Boardtakes the designs' conditions, andviolreports blank rather than raising without them.rankexcludes the split-half row.What was checked
python -m engiopt.evaluate --list-metricsruns as documented.What was not checked
Nothing has been run at the published fifty conditions. The notebook's twelve is the only real-problem run.
Board.loadby bare name has been exercised only for the two families that have a canonical checkpoint. The trajectory metrics assume a gap path that falls. For the diffusion model it rises after the first call while the optimizer restores feasibility. The columns report that faithfully but do not explain it.Follow-ups
knn_retrieval,deconv_regressionand the planted constructions, as a separate PR.space="lv", registered by the LVAE PR once the Lipschitz fixes land.🤖 Generated with Claude Code