Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
18 commits
Select commit Hold shift + click to select a range
efb8d42
Make a metric one decorator and a board one object
mkeeler43 Sep 25, 2026
2c21d7b
Add a notebook that walks the metrics suite on a toy problem
mkeeler43 Sep 25, 2026
5a5a317
Port the metric suite onto the one-decorator contract
mkeeler43 Sep 25, 2026
5224615
Let a board choose its space, and load published checkpoints by name
mkeeler43 Sep 25, 2026
1babd6a
Describe the current metric suite where contributors read about it
mkeeler43 Sep 25, 2026
3c6e6f5
Format the notebook's code cells the way CI does
mkeeler43 Sep 25, 2026
32afcc7
Put the notebook's imports where the linter expects them
mkeeler43 Sep 25, 2026
3cc4150
Say what each metric measures and how to read its scale
mkeeler43 Sep 28, 2026
62812ca
Run the metrics notebook on beams2d with published checkpoints
mkeeler43 Sep 28, 2026
1126e7d
Measure memorization as a distance, without a tolerance
mkeeler43 Sep 28, 2026
0af3a34
Stop passing the spec to push_to_hub, which no longer reads it
mkeeler43 Sep 28, 2026
ab2ef94
Drop the evaluator argument _publish stopped using
mkeeler43 Sep 28, 2026
68b32b9
Measure settling against the optimum, and size the kernel from the tr…
mkeeler43 Sep 29, 2026
59c198d
Explain the evaluation context before asking the reader to write a me…
mkeeler43 Sep 29, 2026
cc5f16e
Address the first review: one default spec, conditions for saved desi…
mkeeler43 Sep 29, 2026
6a1ad77
Re-run the notebook with the renamed explain column
mkeeler43 Sep 29, 2026
481fd39
Address the second review: one CLI default, a volume budget the board…
mkeeler43 Sep 30, 2026
876e094
Let the board take the test rows themselves, and the evaluator a chos…
mkeeler43 Oct 1, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
40 changes: 25 additions & 15 deletions CONTRIBUTING_A_MODEL.md
Original file line number Diff line number Diff line change
Expand Up @@ -175,25 +175,29 @@ python -m engiopt.evaluate --list-generators # your model should ap
python -m engiopt.evaluate --problem-id beams2d --generators my_model --hf-entity my-hf-username
```

The default pass runs the cheap metrics: `mmd`, `dpp`, `viol`, and the two
integrity checks (`novelty`, `cond_sens`). Feasibility is among them because it
describes the design as generated, so it is a constraint check rather than a
solver run. `--include-expensive` adds the optimality gaps (`iog`, `cog`,
`fog`), which do run the optimizer and are slow — leave them off while
iterating.
The default pass runs every metric that needs no simulator: the set-level
ones (`mmd`, `coverage`, `vendi`, `dpp`), the per-condition ones
(`per_condition_distance`, `volume_error`), feasibility (`viol`), the two
integrity checks (`train_distance`, `cond_sens`), and the cost
columns. `--include-expensive` adds the columns that re-optimize from each
generated design — the optimality gaps `iog`, `cog`, `fog` and the
call-budget metrics `calls_to_near_optimum`, `gap_after_calls` and
`reaches_reference_rate` — which run the optimizer and are
slow; leave them off while iterating. `python -m engiopt.evaluate
--list-metrics` prints the question each column answers.

Watch two columns while you develop:

- **`cond_sens`** should be greater than zero if you declared `conditional =
True`. Exactly zero means your conditions are not reaching the network, which
is a wiring bug far more often than a modeling choice.
- **`copy_rate`** should be near zero. High means your model is reproducing
training designs rather than generating, and the board will publish it without
ranking it.
- **`train_distance_ratio`** should be near one. Near zero means your model is
reproducing training designs rather than generating; the board shows it and
leaves the reading to people.

`cond_sens` draws the batch a second time, under shuffled conditions, so it
doubles sampling cost. That is nothing for a GAN and noticeable for a diffusion
model — pass `--metrics mmd dpp viol` while iterating if it slows you down, then
model — pass `--metrics mmd viol` while iterating if it slows you down, then
run the full list before publishing.

## 5. Publish it
Expand All @@ -215,13 +219,19 @@ from engiopt.evaluation import register_metric

@register_metric("my_metric", family="diversity", cost="cheap", higher_is_better=True)
def my_metric(ctx) -> float:
"""One line, shown by --list-metrics."""
return float(...) # ctx.gen_flat, ctx.ref_flat, ctx.conditions, ...
"""What question does the number answer? This line is what readers see."""
return ctx.reduce(...) # one value per design -> the caller's aggregation (mean or median)
```

Declare `cost="expensive"` if it touches `ctx.optimization` (the simulator or
optimizer). The registry enforces the split, so a cheap run can never
accidentally launch a simulation.
Five declarations: the name, the family (which question it belongs to), the
cost, the direction — `None` for a diagnostic that is read but never ranked on —
and, when the metric needs the actual designs rather than codes in some space
(a constraint check, a copy corpus), `pixel_only=True`. Declare
`cost="expensive"` if it touches `ctx.optimization` (the simulator or
optimizer); the registry enforces the split, so a cheap run can never
accidentally launch a simulation. Where the metric is computed (`space=`) and
how per-design values collapse to one number (`aggregation=`) are chosen when a
board is evaluated, not by the metric.

Importing the module is what registers it — which means you can define a metric
in a notebook cell and it will appear in the next leaderboard you build.
Expand Down
25 changes: 16 additions & 9 deletions LEADERBOARD.md
Original file line number Diff line number Diff line change
Expand Up @@ -132,7 +132,6 @@ to an entry covering every required seed.
| Flag | Trips when |
|---|---|
| `unverified` | No runner has reproduced it yet. |
| `memorized` | `copy_rate` exceeds the spec's `max_copy_rate`. |
| `ignores_conditions` | `cond_sens` is exactly zero for a model registered as conditional. |

## The copying problem
Expand All @@ -152,21 +151,29 @@ scored: the whole dataset is public, so a memorizer memorizes all of it.

So the board measures retrieval instead:

- **`novelty`** — mean per-element RMS distance from each generated design to the
nearest design the model could have copied: the training split it was fitted
on, plus the reference designs the protocol names.
- **`copy_rate`** — the share of the batch closer than `copy_tol`. This is the
one that flags an entry.
- **`train_distance`** — how far each generated design sits from the nearest
design in the training split, per element. Zero means the model reproduces
what it was trained on.
- **`train_distance_ratio`** — the same distance divided by what the withheld
reference designs score against the training split. One means the model's
designs are as far from its training data as real unseen optima are; near zero
means retrieval. There is no tolerance to set: the reference designs supply
the scale and nothing else.

Both are diagnostic, with no ranking direction, and that is not an oversight.
Ranking on novelty would put pure noise in first place — zero means retrieval,
No row is flagged for memorization; the columns are read, not enforced.
Ranking on `train_distance` would put pure noise in first place — zero means retrieval,
but large means only "unlike the data", which a broken model also achieves.
Closing one gaming vector by opening another is not progress. The same reasoning
applies to `cond_sens`: an unconditional model is a legitimate thing to build,
and responding to conditions *wrongly* also moves the output, so a large value is
not by itself a good one.

Read them next to `mmd` and `viol`, never on their own.
Read them next to `mmd` and `viol`, never on their own. Every column's question,
direction and cost is one call away — `METRICS.explain()` — and
[`example_metrics_suite.ipynb`](example_metrics_suite.ipynb) walks the whole suite,
including the construction that scores a perfect `mmd` by returning the correct
designs for the wrong conditions.

### What would actually close it

Expand All @@ -190,7 +197,7 @@ from engiopt.evaluation.leaderboard import disagreement, load_from_hub, rank
from engiopt.evaluation.spec import EvalSpec

board = load_from_hub("IDEALLab/engiopt-leaderboard")
spec = EvalSpec.load("beams2d/v1")
spec = EvalSpec.load("beams2d/v2")

rank(board, "fog", eval_spec=spec) # the public ordering
rank(board, "fog", eval_spec=spec, eligible_only=False) # everything, including claims
Expand Down
2 changes: 1 addition & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -166,7 +166,7 @@ python -m engiopt.verify --board ... --verifier ideallab-ci --publish # the o

Because the row carries the full address, that audit is not privileged — anyone can run the same command and get the same answer.

**Two integrity metrics decide whether a score means what it looks like.** The evaluation protocol is public, so a lookup table keyed on the condition vector returns the dataset-optimal designs and posts a perfect `mmd` and a zero `viol` — measured on beams2d, it beats a trained cGAN on every headline metric. `novelty` / `copy_rate` catch it (`copy_rate=1.00`), and `cond_sens` catches a model that ignores the conditions it claims to use. Both are diagnostic rather than ranked, because ranking on them would just reward the opposite extreme. Flagged rows are published and left out of the ordering.
**Two integrity metrics decide whether a score means what it looks like.** The evaluation protocol is public, so a lookup table keyed on the condition vector returns the dataset-optimal designs and posts a perfect `mmd` and a zero `viol` — measured on beams2d, it beats a trained cGAN on every headline metric. `train_distance` reads zero for it, and `cond_sens` catches a model that ignores the conditions it claims to use. Both are diagnostic rather than ranked, because ranking on them would just reward the opposite extreme. Flagged rows are published and left out of the ordering. The whole suite — what each column measures, which way is better, how to change the space or the aggregation, and how to add a metric — is in [`example_metrics_suite.ipynb`](example_metrics_suite.ipynb).

See **[LEADERBOARD.md](LEADERBOARD.md)** for the submission path, the flags, and what would actually close the copying hole.

Expand Down
25 changes: 7 additions & 18 deletions engiopt/evaluate.py
Original file line number Diff line number Diff line change
Expand Up @@ -32,7 +32,6 @@
from engiopt.evaluation.leaderboard import push_to_hub
from engiopt.evaluation.registry import METRICS
from engiopt.evaluation.submission import FLAG_IGNORES_CONDITIONS
from engiopt.evaluation.submission import FLAG_MEMORIZED
from engiopt.evaluation.submission import FLAG_UNVERIFIED
from engiopt.evaluation.submission import integrity_flags
from engiopt.utils.all_generators import BUILTIN_GENERATORS
Expand Down Expand Up @@ -65,7 +64,8 @@ class Args:
cgan_cnn_2d:023dd1fb gan_cnn_2d:06d9a9a1` asks for exactly the two that
exist."""
spec: str | None = None
"""Eval spec reference, e.g. `beams2d/v1`. Defaults to `<problem_id>/v1`."""
"""Eval spec reference, e.g. `beams2d/v2`. Omitted, the problem's current spec loads;
`EvalSpec.load` owns that default, so the CLI cannot drift from the library."""
metrics: tuple[str, ...] = ()
"""Metric names; defaults to the spec's list."""
include_expensive: bool = False
Expand Down Expand Up @@ -197,11 +197,7 @@ def _availability_label(count: int | None) -> str:
def _print_metrics() -> None:
"""Print every registered metric grouped by the question it answers."""
print(f"{len(METRICS)} metrics registered:\n")
for family in sorted({spec.family for spec in METRICS.values()}):
print(f" [{family}]")
for spec in METRICS.select(family=family):
direction = {True: "higher better", False: "lower better", None: "diagnostic"}[spec.higher_is_better]
print(f" {spec.name:<10} {spec.cost:<10} {direction:<14} {spec.description}")
print(METRICS.explain().sort_values(["family", "cost"]).to_string())


def _resolve_generator_names(requested: tuple[str, ...], problem_id: str) -> list[str]:
Expand Down Expand Up @@ -346,8 +342,7 @@ def main(args: Args) -> int:
if args.list_generators or args.list_metrics:
return 0

spec = args.spec or f"{args.problem_id}/v1"
evaluator = Evaluator.for_problem(args.problem_id, spec=spec)
evaluator = Evaluator.for_problem(args.problem_id, spec=args.spec)
print(f"Problem {args.problem_id} | spec {evaluator.spec.version} | n={evaluator.spec.n_samples}")

generators = _load_generators(args, evaluator)
Expand Down Expand Up @@ -375,14 +370,14 @@ def main(args: Args) -> int:
print(f"\n{board.to_string(index=False)}\n")
print(f"Wrote {len(board)} rows to {destination}")

_publish(args, evaluator, board)
_publish(args, board)
return 0


def _publish(args: Args, evaluator: Evaluator, board: pd.DataFrame) -> None:
def _publish(args: Args, board: pd.DataFrame) -> None:
"""Push the results wherever the flags asked, then print the ranking comparison."""
if args.push_to:
merged = push_to_hub(board, args.push_to, eval_spec=evaluator.spec)
merged = push_to_hub(board, args.push_to)
print(f"Published {len(board)} row(s) to {args.push_to}; board now holds {len(merged)} rows.")
print(
"These rows are unverified. They will not be ranked until a runner re-fetches the "
Expand Down Expand Up @@ -416,12 +411,6 @@ def _print_integrity_warnings(board: pd.DataFrame) -> None:
if not flags:
continue
label = f"{row.get('algo_id')} seed {row.get('seed')}"
if FLAG_MEMORIZED in flags:
print(
f"\n [{label}] copy_rate={row.get('copy_rate'):.2f}: most of this batch reproduces designs "
"from the dataset rather than generating them. Its distribution and performance scores "
"measure retrieval, and it will be published but not ranked."
)
if FLAG_IGNORES_CONDITIONS in flags:
print(
f"\n [{label}] cond_sens=0: output did not change at all when the conditions were shuffled, "
Expand Down
4 changes: 3 additions & 1 deletion engiopt/evaluation/__init__.py
Original file line number Diff line number Diff line change
Expand Up @@ -3,7 +3,7 @@
from engiopt.evaluation import Evaluator
from engiopt.utils.all_generators import BUILTIN_GENERATORS

ev = Evaluator.for_problem("beams2d", spec="beams2d/v1")
ev = Evaluator.for_problem("beams2d", spec="beams2d/v2")
gen = BUILTIN_GENERATORS["cgan_cnn_2d"].from_pretrained(ev.problem, problem_id="beams2d", seed=1)
ev.score(gen)

Expand All @@ -12,6 +12,7 @@
property of the comparison, not of the model.
"""

from engiopt.evaluation.board import Board
from engiopt.evaluation.context import EvaluationContext
from engiopt.evaluation.context import OptimizationResults
from engiopt.evaluation.evaluator import Evaluator
Expand All @@ -31,6 +32,7 @@

__all__ = [
"METRICS",
"Board",
"EvalSpec",
"EvaluationContext",
"Evaluator",
Expand Down
Loading
Loading