Skip to content

harbor_env: forward purpose, provider and eval_sampling from HarborSessionFactory - #1285

Open
jayzuccarelli wants to merge 2 commits into
huggingface:mainfrom
jayzuccarelli:harbor-session-factory-purpose-provider
Open

jayzuccarelli wants to merge 2 commits into
huggingface:mainfrom
jayzuccarelli:harbor-session-factory-purpose-provider

Conversation

@jayzuccarelli

@jayzuccarelli jayzuccarelli commented Sep 30, 2026 •

Copy link
Copy Markdown

HarborSessionFactory (where #1276 points opencode_env users) can't pass purpose, provider or eval_sampling, although HarborEnv.run_rollout and the server's run_rollout tool accept all three. So a TRL loop-owning run can't ask for purpose="train", which the provider qualification guide recommends.

That matters most for training: a vLLM started without --return-tokens-as-token-ids still runs every rollout, as eval, and fetch_proxy_trace() hands the trainer an empty list per rollout with only a warning. With purpose="train" the server refuses the rollout before creating a sandbox.

The factory now takes the three kwargs and passes them through. It validates them at construction, like sampling, so a bad value fails once rather than on every rollout. Defaults are unchanged.

Tests: new cases in tests/envs/test_harbor_session_factory.py fail on main (TypeError on the new kwargs) and pass with this change; the harbor and capture suites pass (606 passed, 3 skipped). ruff and usort clean on the touched files.

Prepared with AI assistance (Claude Code); I reviewed the change and ran the tests.


Note

Low Risk
Additive API with constructor validation and passthrough to existing rollout parameters; no changes to auth or core training logic beyond enabling explicit train/eval intent.

Overview
HarborSessionFactory and HarborSession now accept provider, purpose, and eval_sampling and pass them through to HarborEnv.run_rollout, matching what the harbor server already supports. Defaults stay openai, auto, and no eval sampling.

Factory construction validates these up front (like sampling): purpose must be auto, eval, or train; eval cannot combine with training sampling; eval_sampling requires purpose="eval" and is checked via evaluation_sampling.

This lets TRL loop-owning runs request purpose="train" so rollouts fail immediately when the engine cannot return token ids, instead of running as eval and yielding empty traces.

Tests cover forwarding purpose, provider, and eval_sampling, plus parametrized rejection of invalid combinations at factory init.

Reviewed by Cursor Bugbot for commit f567900. Bugbot is set up for automated code reviews on this repo. Configure here.

jayzuccarelli and others added 2 commits September 30, 2026 17:45
…ssionFactory

HarborSessionFactory is where the opencode_env deprecation (huggingface#1276) sends people, but it can't pass
`purpose`, `provider` or `eval_sampling`, even though `HarborEnv.run_rollout` and the server's
`run_rollout` tool both take them. So a TRL loop-owning run can't ask for `purpose="train"` (which
the provider qualification guide recommends), and can't point at a native Anthropic endpoint.

The train case is the one that bites: a vLLM started without `--return-tokens-as-token-ids`
still runs every rollout, just as eval, and `fetch_proxy_trace()` hands the trainer an empty list
per rollout with a warning. With `purpose="train"` the server refuses the rollout before a sandbox
is created ("training requires exact engine token capture").

Fix: add the three kwargs to the factory and session and pass them through. The factory checks
them at construction the same way it already checks `sampling`, mirroring `SessionRegistry.create`,
so a bad value fails once instead of on every rollout. Defaults are unchanged, and the client
already leaves defaults off the wire for older servers.

Tests: 6 new cases in tests/envs/test_harbor_session_factory.py fail on main (TypeError on the new
kwargs) and pass with this change. tests/envs/test_harbor_*.py + test_capture_*.py: 606 passed,
3 skipped. usort, ruff format and ruff check clean on the touched files.

Prepared with AI assistance (Claude Code); I reviewed the change and ran the tests.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ssionFactory

`HarborSessionFactory` (where huggingface#1276 points opencode_env users) can't pass `purpose`, `provider` or `eval_sampling`, although `HarborEnv.run_rollout` and the server's `run_rollout` tool accept all three. So a TRL loop-owning run can't ask for `purpose="train"`, which the provider qualification guide recommends.

That matters most for training: a vLLM started without `--return-tokens-as-token-ids` still runs every rollout, as eval, and `fetch_proxy_trace()` hands the trainer an empty list per rollout with only a warning. With `purpose="train"` the server refuses the rollout before creating a sandbox.

The factory now takes the three kwargs and passes them through. It validates them at construction, like `sampling`, so a bad value fails once rather than on every rollout. Defaults are unchanged.

Tests: new cases in `tests/envs/test_harbor_session_factory.py` fail on main (TypeError on the new kwargs) and pass with this change; the harbor and capture suites pass (606 passed, 3 skipped). ruff and usort clean on the touched files.

Prepared with AI assistance (Claude Code); I reviewed the change and ran the tests.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant