From d34604392f6dba40b2e4fbaa9bdd71b2f6fea07d Mon Sep 17 00:00:00 2001 From: sergiopaniego Date: Wed, 30 Sep 2026 12:31:56 +0200 Subject: [PATCH 1/2] Deprecate opencode_env and pi_env in favour of harbor_env OpenCode and Pi are validated Harbor harnesses (`harness="opencode"`, `harness="pi"`), with the same token-level capture for training, so the per-harness envs are now a second path to the same agents. Both are removed in OpenEnv 0.8.0. - Importing either package emits a FutureWarning with the removal version and the HarborSessionFactory call to use instead. - The env READMEs (and their generated doc pages) open with the same notice and a migration snippet; both tutorials and the sidebar and catalog mark them deprecated. - The OpenCode tutorial's TRL links are pinned to v1.14.1, the last release with that recipe, so they keep working once TRL drops the example. - The Pi tutorial linked a TRL script that was never merged; it now says so and points to Harbor. --- docs/source/_toctree.yml | 8 ++++---- docs/source/environments.md | 4 ++-- docs/source/environments/opencode.md | 20 ++++++++++++++++++++ docs/source/environments/pi.md | 20 ++++++++++++++++++++ docs/source/tutorials/opencode-agent-grpo.md | 15 +++++++++++---- docs/source/tutorials/pi-agent-grpo.md | 10 ++++++++-- envs/opencode_env/README.md | 20 ++++++++++++++++++++ envs/opencode_env/__init__.py | 12 ++++++++++++ envs/pi_env/README.md | 20 ++++++++++++++++++++ envs/pi_env/__init__.py | 12 ++++++++++++ 10 files changed, 129 insertions(+), 12 deletions(-) diff --git a/docs/source/_toctree.yml b/docs/source/_toctree.yml index 707920b2fc..ac3475bec4 100644 --- a/docs/source/_toctree.yml +++ b/docs/source/_toctree.yml @@ -64,9 +64,9 @@ - local: tutorials/browsergym-harness title: Agentic Harness - local: tutorials/opencode-agent-grpo - title: OpenCode + title: OpenCode (deprecated) - local: tutorials/pi-agent-grpo - title: Pi + title: Pi (deprecated) title: Harnesses - sections: - local: tutorials/evaluation-inspect @@ -150,11 +150,11 @@ - local: environments/agent_world_model title: Agent World Model - local: environments/opencode - title: OpenCode + title: OpenCode (deprecated) - local: environments/pelican_svg title: Pelican SVG - local: environments/pi - title: Pi + title: Pi (deprecated) - local: environments/sophistry_bench_sprint title: Sophistry Bench Sprint title: Environments diff --git a/docs/source/environments.md b/docs/source/environments.md index 12ee6085a8..c23faaaf9e 100644 --- a/docs/source/environments.md +++ b/docs/source/environments.md @@ -259,7 +259,7 @@ The OpenEnv community has built a catalog of ready-to-run environments that cove
-
OpenCode
+
OpenCode (deprecated)

opencode_env runs the OpenCode coding agent inside an isolated E2B sandbox against any OpenAI-compatible LLM endpoint, optionally capturing per-token logprobs.

📄 Docs @@ -274,7 +274,7 @@ The OpenEnv community has built a catalog of ready-to-run environments that cove
-
Pi
+
Pi (deprecated)

pi_env runs the Pi coding agent inside an isolated Hugging Face sandbox against any OpenAI-compatible LLM endpoint, optionally capturing per-token logprobs.

📄 Docs diff --git a/docs/source/environments/opencode.md b/docs/source/environments/opencode.md index 8804eee01d..938952167d 100644 --- a/docs/source/environments/opencode.md +++ b/docs/source/environments/opencode.md @@ -1,6 +1,26 @@ # OpenCode Environment for OpenEnv +> [!WARNING] +> **Deprecated, will be removed in OpenEnv 0.8.0.** OpenCode now runs through +> [`harbor_env`](https://huggingface.co/docs/openenv/environments/harbor) as one of Harbor's harnesses, with the same token-level +> capture for training. Serve a Harbor task dataset and pick `opencode` as the harness: +> +> ```bash +> openenv harbor serve --dataset +> ``` +> +> ```python +> from harbor_env.harness import HarborSessionFactory +> +> factory = HarborSessionFactory( +> "http://localhost:8000", split="", harness="opencode", llm_url=vllm_url, model=model +> ) +> ``` +> +> Tasks become Harbor task directories (instruction, environment and verifier) +> instead of `OpenCodeTask`. + `opencode_env` runs the [OpenCode](https://opencode.ai) coding agent inside an isolated [E2B](https://e2b.dev) sandbox against any OpenAI-compatible LLM endpoint, optionally capturing per-token logprobs for GRPO training. diff --git a/docs/source/environments/pi.md b/docs/source/environments/pi.md index a69179844a..76331365f6 100644 --- a/docs/source/environments/pi.md +++ b/docs/source/environments/pi.md @@ -1,6 +1,26 @@ # Pi Environment for OpenEnv +> [!WARNING] +> **Deprecated, will be removed in OpenEnv 0.8.0.** Pi now runs through +> [`harbor_env`](https://huggingface.co/docs/openenv/environments/harbor) as one of Harbor's harnesses, with the same token-level +> capture for training. Serve a Harbor task dataset and pick `pi` as the harness: +> +> ```bash +> openenv harbor serve --dataset +> ``` +> +> ```python +> from harbor_env.harness import HarborSessionFactory +> +> factory = HarborSessionFactory( +> "http://localhost:8000", split="", harness="pi", llm_url=vllm_url, model=model +> ) +> ``` +> +> Tasks become Harbor task directories (instruction, environment and verifier) +> instead of `PiTask`. + `pi_env` runs the [Pi](https://github.com/badlogic/pi-mono) coding agent inside an isolated [Hugging Face sandbox](https://huggingface.co/docs/huggingface_hub/package_reference/sandbox) against any OpenAI-compatible LLM endpoint, optionally capturing per-token diff --git a/docs/source/tutorials/opencode-agent-grpo.md b/docs/source/tutorials/opencode-agent-grpo.md index bc497185ed..a900f22963 100644 --- a/docs/source/tutorials/opencode-agent-grpo.md +++ b/docs/source/tutorials/opencode-agent-grpo.md @@ -1,5 +1,12 @@ # Coding Agent Training with TRL (OpenCode) +> [!WARNING] +> **Deprecated.** This tutorial uses `opencode_env`, which will be removed in +> OpenEnv 0.8.0. OpenCode now runs through [`harbor_env`](https://huggingface.co/docs/openenv/environments/harbor) as one of Harbor's +> harnesses (`--harness opencode`), with the same token-level capture for +> training. The TRL links below are pinned to v1.14.1, the last release +> with this recipe. + This tutorial covers the black-box training path: training the actual [`opencode`](https://opencode.ai) coding agent, with its own planner, tools, context management, and stop condition, using TRL's experimental @@ -38,7 +45,7 @@ agent process. Three small functions adapt the recipe to your task: receive gradient), and `agent_turn_fn` (which trace entries are real agent turns rather than auxiliary calls like title generation). All three are documented in -[TRL's harness training guide](https://huggingface.co/docs/trl/openenv#training-on-harnesses-training-a-real-coding-agent-opencode). +[TRL's harness training guide](https://huggingface.co/docs/trl/v1.14.1/en/openenv#training-on-harnesses-training-a-real-coding-agent-opencode). ## Full Recipe @@ -53,12 +60,12 @@ remote Hugging Face sandbox instead of a local subprocess. Installation, the exact vLLM serving flags, and the run commands live next to the recipe in TRL: -- [Training on harnesses](https://huggingface.co/docs/trl/openenv#training-on-harnesses-training-a-real-coding-agent-opencode) +- [Training on harnesses](https://huggingface.co/docs/trl/v1.14.1/en/openenv#training-on-harnesses-training-a-real-coding-agent-opencode) in TRL's OpenEnv docs: rollout semantics, the reward path, turn selection, and the trace contract. -- [`examples/async_grpo_opencode/async_grpo_opencode.py`](https://github.com/huggingface/trl/blob/main/examples/async_grpo_opencode/async_grpo_opencode.py) +- [`examples/async_grpo_opencode/async_grpo_opencode.py`](https://github.com/huggingface/trl/blob/v1.14.1/examples/async_grpo_opencode/async_grpo_opencode.py) in TRL: the complete, runnable script. -- [`examples/async_grpo_opencode/opencode_hf_sandbox.py`](https://github.com/huggingface/trl/blob/main/examples/async_grpo_opencode/opencode_hf_sandbox.py) +- [`examples/async_grpo_opencode/opencode_hf_sandbox.py`](https://github.com/huggingface/trl/blob/v1.14.1/examples/async_grpo_opencode/opencode_hf_sandbox.py) in TRL: the same recipe, but each rollout runs in its own remote Hugging Face sandbox, so rollouts scale out beyond one node. - [`envs/opencode_env`](https://github.com/huggingface/OpenEnv/tree/main/envs/opencode_env): diff --git a/docs/source/tutorials/pi-agent-grpo.md b/docs/source/tutorials/pi-agent-grpo.md index 2f449425dd..c0d7c890ad 100644 --- a/docs/source/tutorials/pi-agent-grpo.md +++ b/docs/source/tutorials/pi-agent-grpo.md @@ -1,5 +1,11 @@ # Coding Agent Training with TRL (Pi) +> [!WARNING] +> **Deprecated.** This tutorial uses `pi_env`, which will be removed in +> OpenEnv 0.8.0. Pi now runs through [`harbor_env`](https://huggingface.co/docs/openenv/environments/harbor) as one of Harbor's +> harnesses (`--harness pi`), with the same token-level capture for +> training. + This tutorial covers the black-box training path: training the actual [`pi`](https://github.com/badlogic/pi-mono) coding agent, with its own planner, tools, context management, and stop condition, using TRL's experimental @@ -47,8 +53,8 @@ is self-contained, runs the agent in a local subprocess sandbox (no container setup needed), needs two GPUs (one serving the policy with vLLM, one training). -- [`examples/scripts/openenv/pi.py`](https://github.com/huggingface/trl/blob/main/examples/scripts/openenv/pi.py) - in TRL: the complete, runnable script. +- The TRL script for this recipe was never merged. To train Pi, use + [`harbor_env`](https://huggingface.co/docs/openenv/environments/harbor) with `--harness pi`. - [`envs/pi_env`](https://github.com/huggingface/OpenEnv/tree/main/envs/pi_env): the OpenEnv side, including the session factory, sandbox backends, and the transparent interception proxy. diff --git a/envs/opencode_env/README.md b/envs/opencode_env/README.md index ade85b182f..00c57fc89e 100644 --- a/envs/opencode_env/README.md +++ b/envs/opencode_env/README.md @@ -14,6 +14,26 @@ short_description: OpenCode coding agent in an E2B sandbox with logprob capture # OpenCode Environment for OpenEnv +> [!WARNING] +> **Deprecated, will be removed in OpenEnv 0.8.0.** OpenCode now runs through +> [`harbor_env`](https://huggingface.co/docs/openenv/environments/harbor) as one of Harbor's harnesses, with the same token-level +> capture for training. Serve a Harbor task dataset and pick `opencode` as the harness: +> +> ```bash +> openenv harbor serve --dataset +> ``` +> +> ```python +> from harbor_env.harness import HarborSessionFactory +> +> factory = HarborSessionFactory( +> "http://localhost:8000", split="", harness="opencode", llm_url=vllm_url, model=model +> ) +> ``` +> +> Tasks become Harbor task directories (instruction, environment and verifier) +> instead of `OpenCodeTask`. + `opencode_env` runs the [OpenCode](https://opencode.ai) coding agent inside an isolated [E2B](https://e2b.dev) sandbox against any OpenAI-compatible LLM endpoint, optionally capturing per-token logprobs for GRPO training. diff --git a/envs/opencode_env/__init__.py b/envs/opencode_env/__init__.py index 223be6f7b2..4f7666a7c5 100644 --- a/envs/opencode_env/__init__.py +++ b/envs/opencode_env/__init__.py @@ -19,6 +19,8 @@ See ``client.py`` and ``server/``. """ +import warnings + from openenv.core.env_server.mcp_types import CallToolAction, ListToolsAction from .client import OpenCodeEnv @@ -33,6 +35,16 @@ from .sandbox import E2BSandboxBackend, SandboxBackend, SandboxHandle from .task import OpenCodeTask +warnings.warn( + "opencode_env is deprecated and will be removed in OpenEnv 0.8.0. " + "Run opencode through harbor_env instead: " + "`HarborSessionFactory(server_url, harness='opencode')` against " + "`openenv harbor serve`. " + "See https://huggingface.co/docs/openenv/environments/harbor", + FutureWarning, + stacklevel=2, +) + __all__ = [ # Deployed-env client "OpenCodeEnv", diff --git a/envs/pi_env/README.md b/envs/pi_env/README.md index d4f06f9fce..3c5ea0a244 100644 --- a/envs/pi_env/README.md +++ b/envs/pi_env/README.md @@ -14,6 +14,26 @@ short_description: Pi coding agent in a Hugging Face sandbox with logprob captur # Pi Environment for OpenEnv +> [!WARNING] +> **Deprecated, will be removed in OpenEnv 0.8.0.** Pi now runs through +> [`harbor_env`](https://huggingface.co/docs/openenv/environments/harbor) as one of Harbor's harnesses, with the same token-level +> capture for training. Serve a Harbor task dataset and pick `pi` as the harness: +> +> ```bash +> openenv harbor serve --dataset +> ``` +> +> ```python +> from harbor_env.harness import HarborSessionFactory +> +> factory = HarborSessionFactory( +> "http://localhost:8000", split="", harness="pi", llm_url=vllm_url, model=model +> ) +> ``` +> +> Tasks become Harbor task directories (instruction, environment and verifier) +> instead of `PiTask`. + `pi_env` runs the [Pi](https://github.com/badlogic/pi-mono) coding agent inside an isolated [Hugging Face sandbox](https://huggingface.co/docs/huggingface_hub/package_reference/sandbox) against any OpenAI-compatible LLM endpoint, optionally capturing per-token diff --git a/envs/pi_env/__init__.py b/envs/pi_env/__init__.py index 6232949470..a08c611ba9 100644 --- a/envs/pi_env/__init__.py +++ b/envs/pi_env/__init__.py @@ -21,6 +21,8 @@ (to be consolidated into ``openenv.core``). """ +import warnings + from opencode_env.sandbox import ( HFSandboxBackend, SandboxBackend, @@ -38,6 +40,16 @@ ) from .task import PiTask +warnings.warn( + "pi_env is deprecated and will be removed in OpenEnv 0.8.0. " + "Run pi through harbor_env instead: " + "`HarborSessionFactory(server_url, harness='pi')` against " + "`openenv harbor serve`. " + "See https://huggingface.co/docs/openenv/environments/harbor", + FutureWarning, + stacklevel=2, +) + __all__ = [ # Deployed-env client "PiEnv", From a51d67cd772a2ee507b6ddacad61afb8d834c8e5 Mon Sep 17 00:00:00 2001 From: sergiopaniego Date: Wed, 30 Sep 2026 13:43:55 +0200 Subject: [PATCH 2/2] Mark the OpenCode and Pi tutorials deprecated in the tutorials index --- docs/source/tutorials/index.md | 6 +++--- 1 file changed, 3 insertions(+), 3 deletions(-) diff --git a/docs/source/tutorials/index.md b/docs/source/tutorials/index.md index c434e29a9a..f7009f5b76 100644 --- a/docs/source/tutorials/index.md +++ b/docs/source/tutorials/index.md @@ -4,7 +4,7 @@ Choose a learning path from the sidebar: - **Basics:** start with [Hello World](openenv-tutorial) and build [Your First Environment](../guides/first-environment). - **Training:** train a [reasoning model](end-to-end-walkthrough), use [TRL](wordle-grpo) or [Unsloth](rl-training-2048), or collect rollouts for [SFT](sft-warmup). -- **Harnesses:** run an [agentic harness](browsergym-harness) or train coding agents with [OpenCode](opencode-agent-grpo) and [Pi](pi-agent-grpo). +- **Harnesses:** run an [agentic harness](browsergym-harness) or train coding agents through [Harbor](https://huggingface.co/docs/openenv/environments/harbor). The [OpenCode](opencode-agent-grpo) and [Pi](pi-agent-grpo) tutorials are deprecated. - **Evals:** follow [Evaluating with Environments](evaluation-inspect). For [MCP environments](mcp-environment) and [rubrics](rubrics), see Concepts in the sidebar. @@ -35,6 +35,6 @@ Already familiar with the basics? These tutorials cover specific workflows in de | [RL Training with 2048](rl-training-2048) | Train a language model to play 2048 using GRPO. Covers game-state representation and reward shaping. | Yes | — | | [Evaluating with Environments](evaluation-inspect) | Wrap an OpenEnv environment in an Inspect AI `Task`, run it via `InspectAIHarness`, and get a structured `EvalResult`. | No | [![Open In Colab](https://colab.research.google.com/assets/colab-badge.svg)](https://colab.research.google.com/github/huggingface/OpenEnv/blob/main/examples/evaluation_inspect.ipynb) | | [BrowserGym Harness Rollouts](browsergym-harness) | Drive BrowserGym through the OpenEnv harness runtime when a trainer needs token sampling, logprobs, and reward assignment inside the training loop. | Yes | — | -| [Training a Real Coding Agent](opencode-agent-grpo) | Train the actual OpenCode agent (black-box, loop-owning) with TRL's `AsyncGRPOTrainer`: a transparent proxy captures each turn's token ids and logprobs while the agent runs its own tool loop. | Yes | — | +| [Training a Real Coding Agent (deprecated)](opencode-agent-grpo) | Deprecated, removed in OpenEnv 0.8.0: use [Harbor](https://huggingface.co/docs/openenv/environments/harbor) with `harness="opencode"`. Train the actual OpenCode agent (black-box, loop-owning) with TRL's `AsyncGRPOTrainer`: a transparent proxy captures each turn's token ids and logprobs while the agent runs its own tool loop. | Yes | — | | [Collecting rollouts for supervised training](sft-warmup) | Run a teacher model to collect reward-labeled rollouts, filter them, and fine-tune a student with TRL's `SFTTrainer` as a warm-start for GRPO. | Yes | [![Open In Colab](https://colab.research.google.com/assets/colab-badge.svg)](https://colab.research.google.com/github/huggingface/OpenEnv/blob/main/examples/sft_warmup.ipynb) | -| [Training a Real Coding Agent (Pi)](pi-agent-grpo) | Train the actual Pi agent (black-box, loop-owning) with TRL's AsyncGRPOTrainer: a transparent proxy captures each turn's token ids and logprobs while the agent runs its own tool loop. | Yes | — | +| [Training a Real Coding Agent (Pi, deprecated)](pi-agent-grpo) | Deprecated, removed in OpenEnv 0.8.0: use [Harbor](https://huggingface.co/docs/openenv/environments/harbor) with `harness="pi"`. Train the actual Pi agent (black-box, loop-owning) with TRL's AsyncGRPOTrainer: a transparent proxy captures each turn's token ids and logprobs while the agent runs its own tool loop. | Yes | — |