Skip to content

Add RL environments post - #3541

Merged
burtenshaw merged 9 commits into
mainfrom
ben/rl-environments-blog
Oct 5, 2026
Merged

burtenshaw merged 9 commits into
mainfrom
ben/rl-environments-blog

Conversation

@burtenshaw

@burtenshaw burtenshaw commented Sep 28, 2026 •

Copy link
Copy Markdown
Collaborator

This PR adds the RL Environments blog post, metadata, an index entry, and a Harbor Space embed while preserving the prose. The search screenshot and NeMo example remain placeholders; the thumbnail reuses the dataset-filter image.

@burtenshaw
burtenshaw marked this pull request as draft September 28, 2026 09:15
Comment thread rl-environments.md Outdated
Co-authored-by: Sylendran95 <sarunagiri@nvidia.com>
@burtenshaw
burtenshaw marked this pull request as ready for review October 1, 2026 12:35

@merveenoyan merveenoyan left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

you cooked as always! @burtenshaw <3
imo most important piece, at the end of the blog, for people who want to use them, maybe we can have some more e2e getting started snippets in a bucket that pull an env, use HF sandboxes or locally to run an agent, or we add it to env loading part for all of them (and not only simply loading an env), this way people get started easily. we could also do a call for contributions of more env types from different players (in case there is)

Comment thread rl-environments.md Outdated

# Welcome RL Environments to the hub

Reinforcement Learning environments give new capabilities to agentic AI systems, and they’re a great way to measure and improve performance in your agents. Therefore, the hugging face hub now has a special place for RL Environments.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
Reinforcement Learning environments give new capabilities to agentic AI systems, and they’re a great way to measure and improve performance in your agents. Therefore, the hugging face hub now has a special place for RL Environments.
Reinforcement Learning environments give new capabilities to agentic AI systems, and they’re a great way to measure and improve performance in your agents. Therefore, Hugging Face Hub now has a special place for RL Environments.

maybe add a small note with some educational content or add it as link

Comment thread rl-environments.md Outdated

Every RL paper or framework ships its own way to find environments. A hub here, a registry there, a GitHub list of tasks with a custom loader. Each one is a small walled garden. If you publish an environment for one framework, users of the other three can’t load it. If you want to train on an environment from another framework or a new paper, you’ll need to port it by hand.

We think this is the wrong shape. An environment is tasks, tests, containers, and a reward rule. That is just data with a runtime attached. The Hub already stores data, versions it, gates it, previews it, and serves it to millions of people. It does not need a second system to hold environments. It needs a way to say "this data is an environment, and here is how you run it."

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

it sounds a bit claude-y so let me rephrase if ok

Suggested change
We think this is the wrong shape. An environment is tasks, tests, containers, and a reward rule. That is just data with a runtime attached. The Hub already stores data, versions it, gates it, previews it, and serves it to millions of people. It does not need a second system to hold environments. It needs a way to say "this data is an environment, and here is how you run it."
We think this is the wrong shape. An environment is tasks, tests, containers, and a reward rule, which are data with a runtime on top. The Hub already stores data, versions it, gates it, previews it, and serves it to millions of people. It does not need a second system to hold environments. It needs a way to say "this data is an environment, and here is how you run it."

Comment thread rl-environments.md Outdated

The frameworks keep doing what they are good at. The Hub does what it is good at, which is hosting, discovery, and versioning. Nobody has to own the catalogue. In fact, catalogues can run on other platforms too, powered by the hub.

Rl Environments are not a new repository type. An environment is a dataset repo, so it gets everything a dataset repo gets: gating, versioning, the viewer, discussions, and PRs.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

we are repeating a bit here

Comment thread rl-environments.md Outdated

Rl Environments are not a new repository type. An environment is a dataset repo, so it gets everything a dataset repo gets: gating, versioning, the viewer, discussions, and PRs.

The Hub does not have to run your environment, but you can as jobs, if you need. The tags just describe compatibility and generate loading commands. Execution stays in the framework, on your hardware or your sandbox provider.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
The Hub does not have to run your environment, but you can as jobs, if you need. The tags just describe compatibility and generate loading commands. Execution stays in the framework, on your hardware or your sandbox provider.
The dataset repository hosts your environment data, and you can use [Hugging Face Jobs](https://huggingface.co/docs/hub/en/jobs) to run the environment if you need. The tags just describe compatibility and generate loading commands. Execution stays in the framework, on your hardware or your sandbox provider.

HF sandboxes?

Comment thread rl-environments.md
Every RL paper or framework ships its own way to find environments. A hub here, a registry there, a GitHub list of tasks with a custom loader. Each one is a small walled garden. If you publish an environment for one framework, users of the other three can’t load it. If you want to train on an environment from another framework or a new paper, you’ll need to port it by hand.

We think this is the wrong shape. An environment is tasks, tests, containers, and a reward rule. That is just data with a runtime attached. The Hub already stores data, versions it, gates it, previews it, and serves it to millions of people. It does not need a second system to hold environments. It needs a way to say "this data is an environment, and here is how you run it."

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

maybe you could add a small excalidraw illustration of an env and what happens inside for illustration (verifier, task, agent etc and what part is held in a dataset)

Comment thread rl-environments.md Outdated

taskset = vf.HarborTaskset(
config=vf.HarborTasksetConfig(
dataset="hf://datasets/<org>/<dataset>",

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I wonder why they don't directly take repo id for convenience

Comment thread rl-environments.md Outdated
env = AutoEnv.from_env("<org>/<dataset>", trust_remote_code=False)
```

And NeMo Gym pulls the data down and evaluates against it:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

this is a bit unclear, is it only used for evals?

@merveenoyan merveenoyan left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

very nice!

Comment thread rl-environments.md Outdated

### Verifiers: run a model on the same task

The [Verifiers v1 API](https://github.com/PrimeIntellect-ai/verifiers/blob/f2382d3c285ecb85578caf2948f24d0eeed84561/docs/v1/harbor.md) can read the same Harbor task directories. With Docker running and a tool-calling model served at `http://localhost:8000/v1`, replace `local-model` below with the served model ID. This example uses uv and Python 3.13, and pins the v1 source revision. The `bash` harness lets the model act in the task's container; the task verifier scores the completed run.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The Harbor integration of [verifiers v1](https://www.primeintellect.ai/blog/verifiers-v1) can run the same task directories in different runtimes, such as Docker. It also supports different harnesses, including a minimal bash harness.

uvx --python 3.13 --from 'verifiers[harbor]' eval harbor \
    --env.taskset.repo https://huggingface.co/datasets/harborframework/terminal-bench-2.1 \
    --env.taskset.dataset terminal-bench-2.1@2.1.0 \
    --env.taskset.tasks '["regex-log"]' \
    --env.agent.runtime.type docker \
    --env.agent.harness.id bash \
    --model "$MODEL" \
    --client.base-url "$LLM_URL"

Here repo is the full Hugging Face Git URL, while dataset is the name and version in that repo's registry.json. This loader uses Harbor's registry conventions, so a bare Hub repo ID cannot replace both values.

@burtenshaw
burtenshaw merged commit 9c89cb1 into main Oct 5, 2026
2 checks passed
@burtenshaw
burtenshaw deleted the ben/rl-environments-blog branch October 5, 2026 14:23
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants