From fdfc05d684aedb7ef0b2cce542527e05e84d2f29 Mon Sep 17 00:00:00 2001 From: sergiopaniego Date: Thu, 1 Oct 2026 14:34:33 +0200 Subject: [PATCH 1/2] Link the TRL example and the multi-harness article from harbor_env Add a short Training with TRL section after the Quick Start, pointing to TRL's async_grpo_harbor example, its OpenEnv guide, and the multi-harness RL article with its collection. --- docs/source/environments/harbor.md | 14 ++++++++++++++ envs/harbor_env/README.md | 14 ++++++++++++++ 2 files changed, 28 insertions(+) diff --git a/docs/source/environments/harbor.md b/docs/source/environments/harbor.md index 349ef4d86..4bdb58a1d 100644 --- a/docs/source/environments/harbor.md +++ b/docs/source/environments/harbor.md @@ -259,6 +259,20 @@ with HarborEnv(base_url="http://localhost:8000") as env: `harness` and `sandbox` are per call, so consecutive rollouts against the same server can use different agents and different backends. +## Training with TRL + +TRL trains a policy on these rollouts with `AsyncGRPOTrainer` (on TRL `main` for now): a +`HarnessRolloutWorker` builds a `HarborSessionFactory` with its sampling policy, opens one session +per rollout and trains on each session's validated `TrainingTrace`. + +- [`examples/async_grpo_harbor`](https://github.com/huggingface/trl/tree/main/examples/async_grpo_harbor): + a complete training script, with local setup and a Hugging Face Jobs launcher. +- [TRL's OpenEnv guide](https://huggingface.co/docs/trl/openenv): how the worker, the reward and the + capture fit together. +- [The ultimate guide to multi-harness RL](https://huggingface.co/spaces/AdithyaSK/multi-harness-rl): + one policy trained across OpenCode, Claude Code, Codex and Mini-SWE-Agent through this + environment, with its [models, datasets and environments](https://huggingface.co/collections/FineEnvs/smoldataenv-multi-harness-rl). + ## The web UI `serve` (and a Space made with `push`) serves a UI at `/web`, in four tabs. diff --git a/envs/harbor_env/README.md b/envs/harbor_env/README.md index e1fc374d8..9fdbc0f34 100644 --- a/envs/harbor_env/README.md +++ b/envs/harbor_env/README.md @@ -267,6 +267,20 @@ with HarborEnv(base_url="http://localhost:8000") as env: `harness` and `sandbox` are per call, so consecutive rollouts against the same server can use different agents and different backends. +## Training with TRL + +TRL trains a policy on these rollouts with `AsyncGRPOTrainer` (on TRL `main` for now): a +`HarnessRolloutWorker` builds a `HarborSessionFactory` with its sampling policy, opens one session +per rollout and trains on each session's validated `TrainingTrace`. + +- [`examples/async_grpo_harbor`](https://github.com/huggingface/trl/tree/main/examples/async_grpo_harbor): + a complete training script, with local setup and a Hugging Face Jobs launcher. +- [TRL's OpenEnv guide](https://huggingface.co/docs/trl/openenv): how the worker, the reward and the + capture fit together. +- [The ultimate guide to multi-harness RL](https://huggingface.co/spaces/AdithyaSK/multi-harness-rl): + one policy trained across OpenCode, Claude Code, Codex and Mini-SWE-Agent through this + environment, with its [models, datasets and environments](https://huggingface.co/collections/FineEnvs/smoldataenv-multi-harness-rl). + ## The web UI `serve` (and a Space made with `push`) serves a UI at `/web`, in four tabs. From 06c4af3cf53173f529edaa3d18a72d4e070853a4 Mon Sep 17 00:00:00 2001 From: sergiopaniego Date: Thu, 1 Oct 2026 14:43:33 +0200 Subject: [PATCH 2/2] Link the FineEnvs multi-harness tutorial from harbor_env --- docs/source/environments/harbor.md | 5 ++++- envs/harbor_env/README.md | 5 ++++- 2 files changed, 8 insertions(+), 2 deletions(-) diff --git a/docs/source/environments/harbor.md b/docs/source/environments/harbor.md index 4bdb58a1d..9f51fca02 100644 --- a/docs/source/environments/harbor.md +++ b/docs/source/environments/harbor.md @@ -271,7 +271,10 @@ per rollout and trains on each session's validated `TrainingTrace`. capture fit together. - [The ultimate guide to multi-harness RL](https://huggingface.co/spaces/AdithyaSK/multi-harness-rl): one policy trained across OpenCode, Claude Code, Codex and Mini-SWE-Agent through this - environment, with its [models, datasets and environments](https://huggingface.co/collections/FineEnvs/smoldataenv-multi-harness-rl). + environment. Its runnable code is the + [FineEnvs multi-harness tutorial](https://github.com/adithya-s-k/FineEnvs/tree/main/05-multi-harness-rl), + and its [models, datasets and environments](https://huggingface.co/collections/FineEnvs/smoldataenv-multi-harness-rl) + are in one collection. ## The web UI diff --git a/envs/harbor_env/README.md b/envs/harbor_env/README.md index 9fdbc0f34..388a77fbe 100644 --- a/envs/harbor_env/README.md +++ b/envs/harbor_env/README.md @@ -279,7 +279,10 @@ per rollout and trains on each session's validated `TrainingTrace`. capture fit together. - [The ultimate guide to multi-harness RL](https://huggingface.co/spaces/AdithyaSK/multi-harness-rl): one policy trained across OpenCode, Claude Code, Codex and Mini-SWE-Agent through this - environment, with its [models, datasets and environments](https://huggingface.co/collections/FineEnvs/smoldataenv-multi-harness-rl). + environment. Its runnable code is the + [FineEnvs multi-harness tutorial](https://github.com/adithya-s-k/FineEnvs/tree/main/05-multi-harness-rl), + and its [models, datasets and environments](https://huggingface.co/collections/FineEnvs/smoldataenv-multi-harness-rl) + are in one collection. ## The web UI