Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
8 changes: 4 additions & 4 deletions docs/source/_toctree.yml
Original file line number Diff line number Diff line change
Expand Up @@ -64,9 +64,9 @@
- local: tutorials/browsergym-harness
title: Agentic Harness
- local: tutorials/opencode-agent-grpo
title: OpenCode
title: OpenCode (deprecated)
- local: tutorials/pi-agent-grpo
title: Pi
title: Pi (deprecated)
title: Harnesses
- sections:
- local: tutorials/evaluation-inspect
Expand Down Expand Up @@ -150,11 +150,11 @@
- local: environments/agent_world_model
title: Agent World Model
- local: environments/opencode
title: OpenCode
title: OpenCode (deprecated)
- local: environments/pelican_svg
title: Pelican SVG
- local: environments/pi
title: Pi
title: Pi (deprecated)
- local: environments/sophistry_bench_sprint
title: Sophistry Bench Sprint
title: Environments
Expand Down
4 changes: 2 additions & 2 deletions docs/source/environments.md
Original file line number Diff line number Diff line change
Expand Up @@ -259,7 +259,7 @@ The OpenEnv community has built a catalog of ready-to-run environments that cove
</div>
</div>
<div class="border dark:border-gray-700 p-5 rounded-lg shadow">
<div class="font-bold mb-2">OpenCode</div>
<div class="font-bold mb-2">OpenCode (deprecated)</div>
<p class="text-sm"><code>opencode_env</code> runs the OpenCode coding agent inside an isolated E2B sandbox against any OpenAI-compatible LLM endpoint, optionally capturing per-token logprobs.</p>
<div class="flex gap-2 mt-3">
<a href="environments/opencode" class="!no-underline border dark:border-gray-700 px-3 py-1 rounded text-sm hover:shadow">📄 Docs</a>
Expand All @@ -274,7 +274,7 @@ The OpenEnv community has built a catalog of ready-to-run environments that cove
</div>
</div>
<div class="border dark:border-gray-700 p-5 rounded-lg shadow">
<div class="font-bold mb-2">Pi</div>
<div class="font-bold mb-2">Pi (deprecated)</div>
<p class="text-sm"><code>pi_env</code> runs the Pi coding agent inside an isolated Hugging Face sandbox against any OpenAI-compatible LLM endpoint, optionally capturing per-token logprobs.</p>
<div class="flex gap-2 mt-3">
<a href="environments/pi" class="!no-underline border dark:border-gray-700 px-3 py-1 rounded text-sm hover:shadow">📄 Docs</a>
Expand Down
20 changes: 20 additions & 0 deletions docs/source/environments/opencode.md
Original file line number Diff line number Diff line change
@@ -1,6 +1,26 @@
<!-- openenv-source: opencode_env -->
# OpenCode Environment for OpenEnv

> [!WARNING]
> **Deprecated, will be removed in OpenEnv 0.8.0.** OpenCode now runs through
> [`harbor_env`](https://huggingface.co/docs/openenv/environments/harbor) as one of Harbor's harnesses, with the same token-level
> capture for training. Serve a Harbor task dataset and pick `opencode` as the harness:
>
> ```bash
> openenv harbor serve --dataset <hf-dataset>
> ```
>
> ```python
> from harbor_env.harness import HarborSessionFactory
>
> factory = HarborSessionFactory(
> "http://localhost:8000", split="<hf-dataset>", harness="opencode", llm_url=vllm_url, model=model
> )
> ```
>
> Tasks become Harbor task directories (instruction, environment and verifier)
> instead of `OpenCodeTask`.

`opencode_env` runs the [OpenCode](https://opencode.ai) coding agent inside
an isolated [E2B](https://e2b.dev) sandbox against any OpenAI-compatible
LLM endpoint, optionally capturing per-token logprobs for GRPO training.
Expand Down
20 changes: 20 additions & 0 deletions docs/source/environments/pi.md
Original file line number Diff line number Diff line change
@@ -1,6 +1,26 @@
<!-- openenv-source: pi_env -->
# Pi Environment for OpenEnv

> [!WARNING]
> **Deprecated, will be removed in OpenEnv 0.8.0.** Pi now runs through
> [`harbor_env`](https://huggingface.co/docs/openenv/environments/harbor) as one of Harbor's harnesses, with the same token-level
> capture for training. Serve a Harbor task dataset and pick `pi` as the harness:
>
> ```bash
> openenv harbor serve --dataset <hf-dataset>
> ```
>
> ```python
> from harbor_env.harness import HarborSessionFactory
>
> factory = HarborSessionFactory(
> "http://localhost:8000", split="<hf-dataset>", harness="pi", llm_url=vllm_url, model=model
> )
> ```
>
> Tasks become Harbor task directories (instruction, environment and verifier)
> instead of `PiTask`.

`pi_env` runs the [Pi](https://github.com/badlogic/pi-mono) coding agent
inside an isolated [Hugging Face sandbox](https://huggingface.co/docs/huggingface_hub/package_reference/sandbox)
against any OpenAI-compatible LLM endpoint, optionally capturing per-token
Expand Down
6 changes: 3 additions & 3 deletions docs/source/tutorials/index.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,7 +4,7 @@ Choose a learning path from the sidebar:

- **Basics:** start with [Hello World](openenv-tutorial) and build [Your First Environment](../guides/first-environment).
- **Training:** train a [reasoning model](end-to-end-walkthrough), use [TRL](wordle-grpo) or [Unsloth](rl-training-2048), or collect rollouts for [SFT](sft-warmup).
- **Harnesses:** run an [agentic harness](browsergym-harness) or train coding agents with [OpenCode](opencode-agent-grpo) and [Pi](pi-agent-grpo).
- **Harnesses:** run an [agentic harness](browsergym-harness) or train coding agents through [Harbor](https://huggingface.co/docs/openenv/environments/harbor). The [OpenCode](opencode-agent-grpo) and [Pi](pi-agent-grpo) tutorials are deprecated.
- **Evals:** follow [Evaluating with Environments](evaluation-inspect).

For [MCP environments](mcp-environment) and [rubrics](rubrics), see Concepts in the sidebar.
Expand Down Expand Up @@ -35,6 +35,6 @@ Already familiar with the basics? These tutorials cover specific workflows in de
| [RL Training with 2048](rl-training-2048) | Train a language model to play 2048 using GRPO. Covers game-state representation and reward shaping. | Yes | — |
| [Evaluating with Environments](evaluation-inspect) | Wrap an OpenEnv environment in an Inspect AI `Task`, run it via `InspectAIHarness`, and get a structured `EvalResult`. | No | [![Open In Colab](https://colab.research.google.com/assets/colab-badge.svg)](https://colab.research.google.com/github/huggingface/OpenEnv/blob/main/examples/evaluation_inspect.ipynb) |
| [BrowserGym Harness Rollouts](browsergym-harness) | Drive BrowserGym through the OpenEnv harness runtime when a trainer needs token sampling, logprobs, and reward assignment inside the training loop. | Yes | — |
| [Training a Real Coding Agent](opencode-agent-grpo) | Train the actual OpenCode agent (black-box, loop-owning) with TRL's `AsyncGRPOTrainer`: a transparent proxy captures each turn's token ids and logprobs while the agent runs its own tool loop. | Yes | — |
| [Training a Real Coding Agent (deprecated)](opencode-agent-grpo) | Deprecated, removed in OpenEnv 0.8.0: use [Harbor](https://huggingface.co/docs/openenv/environments/harbor) with `harness="opencode"`. Train the actual OpenCode agent (black-box, loop-owning) with TRL's `AsyncGRPOTrainer`: a transparent proxy captures each turn's token ids and logprobs while the agent runs its own tool loop. | Yes | — |
| [Collecting rollouts for supervised training](sft-warmup) | Run a teacher model to collect reward-labeled rollouts, filter them, and fine-tune a student with TRL's `SFTTrainer` as a warm-start for GRPO. | Yes | [![Open In Colab](https://colab.research.google.com/assets/colab-badge.svg)](https://colab.research.google.com/github/huggingface/OpenEnv/blob/main/examples/sft_warmup.ipynb) |
| [Training a Real Coding Agent (Pi)](pi-agent-grpo) | Train the actual Pi agent (black-box, loop-owning) with TRL's AsyncGRPOTrainer: a transparent proxy captures each turn's token ids and logprobs while the agent runs its own tool loop. | Yes | — |
| [Training a Real Coding Agent (Pi, deprecated)](pi-agent-grpo) | Deprecated, removed in OpenEnv 0.8.0: use [Harbor](https://huggingface.co/docs/openenv/environments/harbor) with `harness="pi"`. Train the actual Pi agent (black-box, loop-owning) with TRL's AsyncGRPOTrainer: a transparent proxy captures each turn's token ids and logprobs while the agent runs its own tool loop. | Yes | — |
15 changes: 11 additions & 4 deletions docs/source/tutorials/opencode-agent-grpo.md
Original file line number Diff line number Diff line change
@@ -1,5 +1,12 @@
# Coding Agent Training with TRL (OpenCode)

> [!WARNING]
> **Deprecated.** This tutorial uses `opencode_env`, which will be removed in
> OpenEnv 0.8.0. OpenCode now runs through [`harbor_env`](https://huggingface.co/docs/openenv/environments/harbor) as one of Harbor's
> harnesses (`--harness opencode`), with the same token-level capture for
> training. The TRL links below are pinned to v1.14.1, the last release
> with this recipe.

This tutorial covers the black-box training path: training the actual
[`opencode`](https://opencode.ai) coding agent, with its own planner, tools,
context management, and stop condition, using TRL's experimental
Expand Down Expand Up @@ -38,7 +45,7 @@ agent process. Three small functions adapt the recipe to your task:
receive gradient), and `agent_turn_fn` (which trace entries are real agent
turns rather than auxiliary calls like title generation). All three are
documented in
[TRL's harness training guide](https://huggingface.co/docs/trl/openenv#training-on-harnesses-training-a-real-coding-agent-opencode).
[TRL's harness training guide](https://huggingface.co/docs/trl/v1.14.1/en/openenv#training-on-harnesses-training-a-real-coding-agent-opencode).

## Full Recipe

Expand All @@ -53,12 +60,12 @@ remote Hugging Face sandbox instead of a local subprocess.
Installation, the exact vLLM serving flags, and the run commands live next to
the recipe in TRL:

- [Training on harnesses](https://huggingface.co/docs/trl/openenv#training-on-harnesses-training-a-real-coding-agent-opencode)
- [Training on harnesses](https://huggingface.co/docs/trl/v1.14.1/en/openenv#training-on-harnesses-training-a-real-coding-agent-opencode)
in TRL's OpenEnv docs: rollout semantics, the reward path, turn selection,
and the trace contract.
- [`examples/async_grpo_opencode/async_grpo_opencode.py`](https://github.com/huggingface/trl/blob/main/examples/async_grpo_opencode/async_grpo_opencode.py)
- [`examples/async_grpo_opencode/async_grpo_opencode.py`](https://github.com/huggingface/trl/blob/v1.14.1/examples/async_grpo_opencode/async_grpo_opencode.py)
in TRL: the complete, runnable script.
- [`examples/async_grpo_opencode/opencode_hf_sandbox.py`](https://github.com/huggingface/trl/blob/main/examples/async_grpo_opencode/opencode_hf_sandbox.py)
- [`examples/async_grpo_opencode/opencode_hf_sandbox.py`](https://github.com/huggingface/trl/blob/v1.14.1/examples/async_grpo_opencode/opencode_hf_sandbox.py)
in TRL: the same recipe, but each rollout runs in its own remote Hugging Face
sandbox, so rollouts scale out beyond one node.
- [`envs/opencode_env`](https://github.com/huggingface/OpenEnv/tree/main/envs/opencode_env):
Expand Down
10 changes: 8 additions & 2 deletions docs/source/tutorials/pi-agent-grpo.md
Original file line number Diff line number Diff line change
@@ -1,5 +1,11 @@
# Coding Agent Training with TRL (Pi)

> [!WARNING]
> **Deprecated.** This tutorial uses `pi_env`, which will be removed in
> OpenEnv 0.8.0. Pi now runs through [`harbor_env`](https://huggingface.co/docs/openenv/environments/harbor) as one of Harbor's
> harnesses (`--harness pi`), with the same token-level capture for
> training.

This tutorial covers the black-box training path: training the actual
[`pi`](https://github.com/badlogic/pi-mono) coding agent, with its own planner,
tools, context management, and stop condition, using TRL's experimental
Expand Down Expand Up @@ -47,8 +53,8 @@ is self-contained, runs the agent in a local subprocess sandbox (no container
setup needed), needs two GPUs (one serving the policy with vLLM, one
training).

- [`examples/scripts/openenv/pi.py`](https://github.com/huggingface/trl/blob/main/examples/scripts/openenv/pi.py)
in TRL: the complete, runnable script.
- The TRL script for this recipe was never merged. To train Pi, use
[`harbor_env`](https://huggingface.co/docs/openenv/environments/harbor) with `--harness pi`.
- [`envs/pi_env`](https://github.com/huggingface/OpenEnv/tree/main/envs/pi_env):
the OpenEnv side, including the session factory, sandbox backends, and the
transparent interception proxy.
20 changes: 20 additions & 0 deletions envs/opencode_env/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -14,6 +14,26 @@ short_description: OpenCode coding agent in an E2B sandbox with logprob capture

# OpenCode Environment for OpenEnv

> [!WARNING]
> **Deprecated, will be removed in OpenEnv 0.8.0.** OpenCode now runs through
> [`harbor_env`](https://huggingface.co/docs/openenv/environments/harbor) as one of Harbor's harnesses, with the same token-level
> capture for training. Serve a Harbor task dataset and pick `opencode` as the harness:
>
> ```bash
> openenv harbor serve --dataset <hf-dataset>
> ```
>
> ```python
> from harbor_env.harness import HarborSessionFactory
>
> factory = HarborSessionFactory(
> "http://localhost:8000", split="<hf-dataset>", harness="opencode", llm_url=vllm_url, model=model
> )
> ```
>
> Tasks become Harbor task directories (instruction, environment and verifier)
> instead of `OpenCodeTask`.

`opencode_env` runs the [OpenCode](https://opencode.ai) coding agent inside
an isolated [E2B](https://e2b.dev) sandbox against any OpenAI-compatible
LLM endpoint, optionally capturing per-token logprobs for GRPO training.
Expand Down
12 changes: 12 additions & 0 deletions envs/opencode_env/__init__.py
Original file line number Diff line number Diff line change
Expand Up @@ -19,6 +19,8 @@
See ``client.py`` and ``server/``.
"""

import warnings

from openenv.core.env_server.mcp_types import CallToolAction, ListToolsAction

from .client import OpenCodeEnv
Expand All @@ -33,6 +35,16 @@
from .sandbox import E2BSandboxBackend, SandboxBackend, SandboxHandle
from .task import OpenCodeTask

warnings.warn(
"opencode_env is deprecated and will be removed in OpenEnv 0.8.0. "
"Run opencode through harbor_env instead: "
"`HarborSessionFactory(server_url, harness='opencode')` against "
"`openenv harbor serve`. "
"See https://huggingface.co/docs/openenv/environments/harbor",
FutureWarning,
stacklevel=2,
)

__all__ = [
# Deployed-env client
"OpenCodeEnv",
Expand Down
20 changes: 20 additions & 0 deletions envs/pi_env/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -14,6 +14,26 @@ short_description: Pi coding agent in a Hugging Face sandbox with logprob captur

# Pi Environment for OpenEnv

> [!WARNING]
> **Deprecated, will be removed in OpenEnv 0.8.0.** Pi now runs through
> [`harbor_env`](https://huggingface.co/docs/openenv/environments/harbor) as one of Harbor's harnesses, with the same token-level
> capture for training. Serve a Harbor task dataset and pick `pi` as the harness:
>
> ```bash
> openenv harbor serve --dataset <hf-dataset>
> ```
>
> ```python
> from harbor_env.harness import HarborSessionFactory
>
> factory = HarborSessionFactory(
> "http://localhost:8000", split="<hf-dataset>", harness="pi", llm_url=vllm_url, model=model
> )
> ```
>
> Tasks become Harbor task directories (instruction, environment and verifier)
> instead of `PiTask`.

`pi_env` runs the [Pi](https://github.com/badlogic/pi-mono) coding agent
inside an isolated [Hugging Face sandbox](https://huggingface.co/docs/huggingface_hub/package_reference/sandbox)
against any OpenAI-compatible LLM endpoint, optionally capturing per-token
Expand Down
12 changes: 12 additions & 0 deletions envs/pi_env/__init__.py
Original file line number Diff line number Diff line change
Expand Up @@ -21,6 +21,8 @@
(to be consolidated into ``openenv.core``).
"""

import warnings

from opencode_env.sandbox import (
HFSandboxBackend,
SandboxBackend,
Expand All @@ -38,6 +40,16 @@
)
from .task import PiTask

warnings.warn(
"pi_env is deprecated and will be removed in OpenEnv 0.8.0. "
"Run pi through harbor_env instead: "
"`HarborSessionFactory(server_url, harness='pi')` against "
"`openenv harbor serve`. "
"See https://huggingface.co/docs/openenv/environments/harbor",
FutureWarning,
stacklevel=2,
)

__all__ = [
# Deployed-env client
"PiEnv",
Expand Down
Loading