diff --git a/docs/source/_toctree.yml b/docs/source/_toctree.yml
index 707920b2f..ac3475bec 100644
--- a/docs/source/_toctree.yml
+++ b/docs/source/_toctree.yml
@@ -64,9 +64,9 @@
- local: tutorials/browsergym-harness
title: Agentic Harness
- local: tutorials/opencode-agent-grpo
- title: OpenCode
+ title: OpenCode (deprecated)
- local: tutorials/pi-agent-grpo
- title: Pi
+ title: Pi (deprecated)
title: Harnesses
- sections:
- local: tutorials/evaluation-inspect
@@ -150,11 +150,11 @@
- local: environments/agent_world_model
title: Agent World Model
- local: environments/opencode
- title: OpenCode
+ title: OpenCode (deprecated)
- local: environments/pelican_svg
title: Pelican SVG
- local: environments/pi
- title: Pi
+ title: Pi (deprecated)
- local: environments/sophistry_bench_sprint
title: Sophistry Bench Sprint
title: Environments
diff --git a/docs/source/environments.md b/docs/source/environments.md
index 12ee6085a..c23faaaf9 100644
--- a/docs/source/environments.md
+++ b/docs/source/environments.md
@@ -259,7 +259,7 @@ The OpenEnv community has built a catalog of ready-to-run environments that cove
-
OpenCode
+
OpenCode (deprecated)
opencode_env runs the OpenCode coding agent inside an isolated E2B sandbox against any OpenAI-compatible LLM endpoint, optionally capturing per-token logprobs.
📄 Docs
@@ -274,7 +274,7 @@ The OpenEnv community has built a catalog of ready-to-run environments that cove
-
Pi
+
Pi (deprecated)
pi_env runs the Pi coding agent inside an isolated Hugging Face sandbox against any OpenAI-compatible LLM endpoint, optionally capturing per-token logprobs.
📄 Docs
diff --git a/docs/source/environments/opencode.md b/docs/source/environments/opencode.md
index 4f299d5aa..0e9bad2f9 100644
--- a/docs/source/environments/opencode.md
+++ b/docs/source/environments/opencode.md
@@ -1,6 +1,26 @@
# OpenCode Environment for OpenEnv
+> [!WARNING]
+> **Deprecated, will be removed in OpenEnv 0.8.0.** OpenCode now runs through
+> [`harbor_env`](https://huggingface.co/docs/openenv/environments/harbor) as one of Harbor's harnesses, with the same token-level
+> capture for training. Serve a Harbor task dataset and pick `opencode` as the harness:
+>
+> ```bash
+> openenv harbor serve --dataset
+> ```
+>
+> ```python
+> from harbor_env.harness import HarborSessionFactory
+>
+> factory = HarborSessionFactory(
+> "http://localhost:8000", split="", harness="opencode", llm_url=vllm_url, model=model
+> )
+> ```
+>
+> Tasks become Harbor task directories (instruction, environment and verifier)
+> instead of `OpenCodeTask`.
+
`opencode_env` runs the [OpenCode](https://opencode.ai) coding agent inside
an isolated [E2B](https://e2b.dev) sandbox against any OpenAI-compatible
LLM endpoint, optionally capturing per-token logprobs for GRPO training.
diff --git a/docs/source/environments/pi.md b/docs/source/environments/pi.md
index 0717e66f6..03fdf6bc6 100644
--- a/docs/source/environments/pi.md
+++ b/docs/source/environments/pi.md
@@ -1,6 +1,26 @@
# Pi Environment for OpenEnv
+> [!WARNING]
+> **Deprecated, will be removed in OpenEnv 0.8.0.** Pi now runs through
+> [`harbor_env`](https://huggingface.co/docs/openenv/environments/harbor) as one of Harbor's harnesses, with the same token-level
+> capture for training. Serve a Harbor task dataset and pick `pi` as the harness:
+>
+> ```bash
+> openenv harbor serve --dataset
+> ```
+>
+> ```python
+> from harbor_env.harness import HarborSessionFactory
+>
+> factory = HarborSessionFactory(
+> "http://localhost:8000", split="", harness="pi", llm_url=vllm_url, model=model
+> )
+> ```
+>
+> Tasks become Harbor task directories (instruction, environment and verifier)
+> instead of `PiTask`.
+
`pi_env` runs the [Pi](https://github.com/badlogic/pi-mono) coding agent
inside an isolated [Hugging Face sandbox](https://huggingface.co/docs/huggingface_hub/package_reference/sandbox)
against any OpenAI-compatible LLM endpoint, optionally capturing per-token
diff --git a/docs/source/tutorials/index.md b/docs/source/tutorials/index.md
index c434e29a9..f7009f5b7 100644
--- a/docs/source/tutorials/index.md
+++ b/docs/source/tutorials/index.md
@@ -4,7 +4,7 @@ Choose a learning path from the sidebar:
- **Basics:** start with [Hello World](openenv-tutorial) and build [Your First Environment](../guides/first-environment).
- **Training:** train a [reasoning model](end-to-end-walkthrough), use [TRL](wordle-grpo) or [Unsloth](rl-training-2048), or collect rollouts for [SFT](sft-warmup).
-- **Harnesses:** run an [agentic harness](browsergym-harness) or train coding agents with [OpenCode](opencode-agent-grpo) and [Pi](pi-agent-grpo).
+- **Harnesses:** run an [agentic harness](browsergym-harness) or train coding agents through [Harbor](https://huggingface.co/docs/openenv/environments/harbor). The [OpenCode](opencode-agent-grpo) and [Pi](pi-agent-grpo) tutorials are deprecated.
- **Evals:** follow [Evaluating with Environments](evaluation-inspect).
For [MCP environments](mcp-environment) and [rubrics](rubrics), see Concepts in the sidebar.
@@ -35,6 +35,6 @@ Already familiar with the basics? These tutorials cover specific workflows in de
| [RL Training with 2048](rl-training-2048) | Train a language model to play 2048 using GRPO. Covers game-state representation and reward shaping. | Yes | — |
| [Evaluating with Environments](evaluation-inspect) | Wrap an OpenEnv environment in an Inspect AI `Task`, run it via `InspectAIHarness`, and get a structured `EvalResult`. | No | [](https://colab.research.google.com/github/huggingface/OpenEnv/blob/main/examples/evaluation_inspect.ipynb) |
| [BrowserGym Harness Rollouts](browsergym-harness) | Drive BrowserGym through the OpenEnv harness runtime when a trainer needs token sampling, logprobs, and reward assignment inside the training loop. | Yes | — |
-| [Training a Real Coding Agent](opencode-agent-grpo) | Train the actual OpenCode agent (black-box, loop-owning) with TRL's `AsyncGRPOTrainer`: a transparent proxy captures each turn's token ids and logprobs while the agent runs its own tool loop. | Yes | — |
+| [Training a Real Coding Agent (deprecated)](opencode-agent-grpo) | Deprecated, removed in OpenEnv 0.8.0: use [Harbor](https://huggingface.co/docs/openenv/environments/harbor) with `harness="opencode"`. Train the actual OpenCode agent (black-box, loop-owning) with TRL's `AsyncGRPOTrainer`: a transparent proxy captures each turn's token ids and logprobs while the agent runs its own tool loop. | Yes | — |
| [Collecting rollouts for supervised training](sft-warmup) | Run a teacher model to collect reward-labeled rollouts, filter them, and fine-tune a student with TRL's `SFTTrainer` as a warm-start for GRPO. | Yes | [](https://colab.research.google.com/github/huggingface/OpenEnv/blob/main/examples/sft_warmup.ipynb) |
-| [Training a Real Coding Agent (Pi)](pi-agent-grpo) | Train the actual Pi agent (black-box, loop-owning) with TRL's AsyncGRPOTrainer: a transparent proxy captures each turn's token ids and logprobs while the agent runs its own tool loop. | Yes | — |
+| [Training a Real Coding Agent (Pi, deprecated)](pi-agent-grpo) | Deprecated, removed in OpenEnv 0.8.0: use [Harbor](https://huggingface.co/docs/openenv/environments/harbor) with `harness="pi"`. Train the actual Pi agent (black-box, loop-owning) with TRL's AsyncGRPOTrainer: a transparent proxy captures each turn's token ids and logprobs while the agent runs its own tool loop. | Yes | — |
diff --git a/docs/source/tutorials/opencode-agent-grpo.md b/docs/source/tutorials/opencode-agent-grpo.md
index bc497185e..a900f2296 100644
--- a/docs/source/tutorials/opencode-agent-grpo.md
+++ b/docs/source/tutorials/opencode-agent-grpo.md
@@ -1,5 +1,12 @@
# Coding Agent Training with TRL (OpenCode)
+> [!WARNING]
+> **Deprecated.** This tutorial uses `opencode_env`, which will be removed in
+> OpenEnv 0.8.0. OpenCode now runs through [`harbor_env`](https://huggingface.co/docs/openenv/environments/harbor) as one of Harbor's
+> harnesses (`--harness opencode`), with the same token-level capture for
+> training. The TRL links below are pinned to v1.14.1, the last release
+> with this recipe.
+
This tutorial covers the black-box training path: training the actual
[`opencode`](https://opencode.ai) coding agent, with its own planner, tools,
context management, and stop condition, using TRL's experimental
@@ -38,7 +45,7 @@ agent process. Three small functions adapt the recipe to your task:
receive gradient), and `agent_turn_fn` (which trace entries are real agent
turns rather than auxiliary calls like title generation). All three are
documented in
-[TRL's harness training guide](https://huggingface.co/docs/trl/openenv#training-on-harnesses-training-a-real-coding-agent-opencode).
+[TRL's harness training guide](https://huggingface.co/docs/trl/v1.14.1/en/openenv#training-on-harnesses-training-a-real-coding-agent-opencode).
## Full Recipe
@@ -53,12 +60,12 @@ remote Hugging Face sandbox instead of a local subprocess.
Installation, the exact vLLM serving flags, and the run commands live next to
the recipe in TRL:
-- [Training on harnesses](https://huggingface.co/docs/trl/openenv#training-on-harnesses-training-a-real-coding-agent-opencode)
+- [Training on harnesses](https://huggingface.co/docs/trl/v1.14.1/en/openenv#training-on-harnesses-training-a-real-coding-agent-opencode)
in TRL's OpenEnv docs: rollout semantics, the reward path, turn selection,
and the trace contract.
-- [`examples/async_grpo_opencode/async_grpo_opencode.py`](https://github.com/huggingface/trl/blob/main/examples/async_grpo_opencode/async_grpo_opencode.py)
+- [`examples/async_grpo_opencode/async_grpo_opencode.py`](https://github.com/huggingface/trl/blob/v1.14.1/examples/async_grpo_opencode/async_grpo_opencode.py)
in TRL: the complete, runnable script.
-- [`examples/async_grpo_opencode/opencode_hf_sandbox.py`](https://github.com/huggingface/trl/blob/main/examples/async_grpo_opencode/opencode_hf_sandbox.py)
+- [`examples/async_grpo_opencode/opencode_hf_sandbox.py`](https://github.com/huggingface/trl/blob/v1.14.1/examples/async_grpo_opencode/opencode_hf_sandbox.py)
in TRL: the same recipe, but each rollout runs in its own remote Hugging Face
sandbox, so rollouts scale out beyond one node.
- [`envs/opencode_env`](https://github.com/huggingface/OpenEnv/tree/main/envs/opencode_env):
diff --git a/docs/source/tutorials/pi-agent-grpo.md b/docs/source/tutorials/pi-agent-grpo.md
index 2f449425d..c0d7c890a 100644
--- a/docs/source/tutorials/pi-agent-grpo.md
+++ b/docs/source/tutorials/pi-agent-grpo.md
@@ -1,5 +1,11 @@
# Coding Agent Training with TRL (Pi)
+> [!WARNING]
+> **Deprecated.** This tutorial uses `pi_env`, which will be removed in
+> OpenEnv 0.8.0. Pi now runs through [`harbor_env`](https://huggingface.co/docs/openenv/environments/harbor) as one of Harbor's
+> harnesses (`--harness pi`), with the same token-level capture for
+> training.
+
This tutorial covers the black-box training path: training the actual
[`pi`](https://github.com/badlogic/pi-mono) coding agent, with its own planner,
tools, context management, and stop condition, using TRL's experimental
@@ -47,8 +53,8 @@ is self-contained, runs the agent in a local subprocess sandbox (no container
setup needed), needs two GPUs (one serving the policy with vLLM, one
training).
-- [`examples/scripts/openenv/pi.py`](https://github.com/huggingface/trl/blob/main/examples/scripts/openenv/pi.py)
- in TRL: the complete, runnable script.
+- The TRL script for this recipe was never merged. To train Pi, use
+ [`harbor_env`](https://huggingface.co/docs/openenv/environments/harbor) with `--harness pi`.
- [`envs/pi_env`](https://github.com/huggingface/OpenEnv/tree/main/envs/pi_env):
the OpenEnv side, including the session factory, sandbox backends, and the
transparent interception proxy.
diff --git a/envs/opencode_env/README.md b/envs/opencode_env/README.md
index 0bbbddd30..b49f5dfa3 100644
--- a/envs/opencode_env/README.md
+++ b/envs/opencode_env/README.md
@@ -14,6 +14,26 @@ short_description: OpenCode coding agent in an E2B sandbox with logprob capture
# OpenCode Environment for OpenEnv
+> [!WARNING]
+> **Deprecated, will be removed in OpenEnv 0.8.0.** OpenCode now runs through
+> [`harbor_env`](https://huggingface.co/docs/openenv/environments/harbor) as one of Harbor's harnesses, with the same token-level
+> capture for training. Serve a Harbor task dataset and pick `opencode` as the harness:
+>
+> ```bash
+> openenv harbor serve --dataset
+> ```
+>
+> ```python
+> from harbor_env.harness import HarborSessionFactory
+>
+> factory = HarborSessionFactory(
+> "http://localhost:8000", split="", harness="opencode", llm_url=vllm_url, model=model
+> )
+> ```
+>
+> Tasks become Harbor task directories (instruction, environment and verifier)
+> instead of `OpenCodeTask`.
+
`opencode_env` runs the [OpenCode](https://opencode.ai) coding agent inside
an isolated [E2B](https://e2b.dev) sandbox against any OpenAI-compatible
LLM endpoint, optionally capturing per-token logprobs for GRPO training.
diff --git a/envs/opencode_env/__init__.py b/envs/opencode_env/__init__.py
index 223be6f7b..4f7666a7c 100644
--- a/envs/opencode_env/__init__.py
+++ b/envs/opencode_env/__init__.py
@@ -19,6 +19,8 @@
See ``client.py`` and ``server/``.
"""
+import warnings
+
from openenv.core.env_server.mcp_types import CallToolAction, ListToolsAction
from .client import OpenCodeEnv
@@ -33,6 +35,16 @@
from .sandbox import E2BSandboxBackend, SandboxBackend, SandboxHandle
from .task import OpenCodeTask
+warnings.warn(
+ "opencode_env is deprecated and will be removed in OpenEnv 0.8.0. "
+ "Run opencode through harbor_env instead: "
+ "`HarborSessionFactory(server_url, harness='opencode')` against "
+ "`openenv harbor serve`. "
+ "See https://huggingface.co/docs/openenv/environments/harbor",
+ FutureWarning,
+ stacklevel=2,
+)
+
__all__ = [
# Deployed-env client
"OpenCodeEnv",
diff --git a/envs/pi_env/README.md b/envs/pi_env/README.md
index 9cb7c7012..994338cdd 100644
--- a/envs/pi_env/README.md
+++ b/envs/pi_env/README.md
@@ -14,6 +14,26 @@ short_description: Pi coding agent in a Hugging Face sandbox with logprob captur
# Pi Environment for OpenEnv
+> [!WARNING]
+> **Deprecated, will be removed in OpenEnv 0.8.0.** Pi now runs through
+> [`harbor_env`](https://huggingface.co/docs/openenv/environments/harbor) as one of Harbor's harnesses, with the same token-level
+> capture for training. Serve a Harbor task dataset and pick `pi` as the harness:
+>
+> ```bash
+> openenv harbor serve --dataset
+> ```
+>
+> ```python
+> from harbor_env.harness import HarborSessionFactory
+>
+> factory = HarborSessionFactory(
+> "http://localhost:8000", split="", harness="pi", llm_url=vllm_url, model=model
+> )
+> ```
+>
+> Tasks become Harbor task directories (instruction, environment and verifier)
+> instead of `PiTask`.
+
`pi_env` runs the [Pi](https://github.com/badlogic/pi-mono) coding agent
inside an isolated [Hugging Face sandbox](https://huggingface.co/docs/huggingface_hub/package_reference/sandbox)
against any OpenAI-compatible LLM endpoint, optionally capturing per-token
diff --git a/envs/pi_env/__init__.py b/envs/pi_env/__init__.py
index 623294947..a08c611ba 100644
--- a/envs/pi_env/__init__.py
+++ b/envs/pi_env/__init__.py
@@ -21,6 +21,8 @@
(to be consolidated into ``openenv.core``).
"""
+import warnings
+
from opencode_env.sandbox import (
HFSandboxBackend,
SandboxBackend,
@@ -38,6 +40,16 @@
)
from .task import PiTask
+warnings.warn(
+ "pi_env is deprecated and will be removed in OpenEnv 0.8.0. "
+ "Run pi through harbor_env instead: "
+ "`HarborSessionFactory(server_url, harness='pi')` against "
+ "`openenv harbor serve`. "
+ "See https://huggingface.co/docs/openenv/environments/harbor",
+ FutureWarning,
+ stacklevel=2,
+)
+
__all__ = [
# Deployed-env client
"PiEnv",