Summary
The controlled solver's catch-all (fle/eval/inspect/integration/solver.py:677-692) turns any exception — including pure infrastructure failures like "No Factorio containers available" — into a completed sample with production_score = 0.0. The eval finishes with status: success, sample.error = None, and the task scores 0.
This poisons benchmark data: a 0 that means "the environment was never created" is indistinguishable in results from a 0 that means "the model failed the task". I hit this via the server-pool allocation bug (#392) — a 3-of-8-task run reported clean zeros for five tasks whose model was never called once (0 assistant messages, 0 API calls). Given #392's default-32 pool, anyone benchmarking on a small cluster may have published such zeros without knowing.
Reproduction
Run any multi-task fle inspect-eval against fewer containers than the pool default (see #392). Tasks after the first complete "successfully" with score 0 in ~1 second, with the real error visible only as a logger.error event inside the .eval transcript.
Proposed fix
Distinguish environment-setup failures from in-episode errors: exceptions raised before the gym env is created (pool allocation, gym.make) should re-raise / mark the sample as errored so Inspect retries or reports it, rather than scoring it. In-episode step errors can keep the current recovery behavior. Happy to submit a PR.
🤖 Generated with Claude Code
Summary
The controlled solver's catch-all (
fle/eval/inspect/integration/solver.py:677-692) turns any exception — including pure infrastructure failures like "No Factorio containers available" — into a completed sample withproduction_score = 0.0. The eval finishes withstatus: success,sample.error = None, and the task scores 0.This poisons benchmark data: a 0 that means "the environment was never created" is indistinguishable in results from a 0 that means "the model failed the task". I hit this via the server-pool allocation bug (#392) — a 3-of-8-task run reported clean zeros for five tasks whose model was never called once (0 assistant messages, 0 API calls). Given #392's default-32 pool, anyone benchmarking on a small cluster may have published such zeros without knowing.
Reproduction
Run any multi-task
fle inspect-evalagainst fewer containers than the pool default (see #392). Tasks after the first complete "successfully" with score 0 in ~1 second, with the real error visible only as alogger.errorevent inside the .eval transcript.Proposed fix
Distinguish environment-setup failures from in-episode errors: exceptions raised before the gym env is created (pool allocation,
gym.make) should re-raise / mark the sample as errored so Inspect retries or reports it, rather than scoring it. In-episode step errors can keep the current recovery behavior. Happy to submit a PR.🤖 Generated with Claude Code