Skip to content

Controlled solver converts infrastructure failures into silent 0-score "successes" #393

Description

@phaestos2501

Summary

The controlled solver's catch-all (fle/eval/inspect/integration/solver.py:677-692) turns any exception — including pure infrastructure failures like "No Factorio containers available" — into a completed sample with production_score = 0.0. The eval finishes with status: success, sample.error = None, and the task scores 0.

This poisons benchmark data: a 0 that means "the environment was never created" is indistinguishable in results from a 0 that means "the model failed the task". I hit this via the server-pool allocation bug (#392) — a 3-of-8-task run reported clean zeros for five tasks whose model was never called once (0 assistant messages, 0 API calls). Given #392's default-32 pool, anyone benchmarking on a small cluster may have published such zeros without knowing.

Reproduction

Run any multi-task fle inspect-eval against fewer containers than the pool default (see #392). Tasks after the first complete "successfully" with score 0 in ~1 second, with the real error visible only as a logger.error event inside the .eval transcript.

Proposed fix

Distinguish environment-setup failures from in-episode errors: exceptions raised before the gym env is created (pool allocation, gym.make) should re-raise / mark the sample as errored so Inspect retries or reports it, rather than scoring it. In-episode step errors can keep the current recovery behavior. Happy to submit a PR.

🤖 Generated with Claude Code

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions