Add RL environments post - #3541
Conversation
Co-authored-by: Sylendran95 <sarunagiri@nvidia.com>
…face/blog into ben/rl-environments-blog
There was a problem hiding this comment.
you cooked as always! @burtenshaw <3
imo most important piece, at the end of the blog, for people who want to use them, maybe we can have some more e2e getting started snippets in a bucket that pull an env, use HF sandboxes or locally to run an agent, or we add it to env loading part for all of them (and not only simply loading an env), this way people get started easily. we could also do a call for contributions of more env types from different players (in case there is)
|
|
||
| # Welcome RL Environments to the hub | ||
|
|
||
| Reinforcement Learning environments give new capabilities to agentic AI systems, and they’re a great way to measure and improve performance in your agents. Therefore, the hugging face hub now has a special place for RL Environments. |
There was a problem hiding this comment.
| Reinforcement Learning environments give new capabilities to agentic AI systems, and they’re a great way to measure and improve performance in your agents. Therefore, the hugging face hub now has a special place for RL Environments. | |
| Reinforcement Learning environments give new capabilities to agentic AI systems, and they’re a great way to measure and improve performance in your agents. Therefore, Hugging Face Hub now has a special place for RL Environments. |
maybe add a small note with some educational content or add it as link
|
|
||
| Every RL paper or framework ships its own way to find environments. A hub here, a registry there, a GitHub list of tasks with a custom loader. Each one is a small walled garden. If you publish an environment for one framework, users of the other three can’t load it. If you want to train on an environment from another framework or a new paper, you’ll need to port it by hand. | ||
|
|
||
| We think this is the wrong shape. An environment is tasks, tests, containers, and a reward rule. That is just data with a runtime attached. The Hub already stores data, versions it, gates it, previews it, and serves it to millions of people. It does not need a second system to hold environments. It needs a way to say "this data is an environment, and here is how you run it." |
There was a problem hiding this comment.
it sounds a bit claude-y so let me rephrase if ok
| We think this is the wrong shape. An environment is tasks, tests, containers, and a reward rule. That is just data with a runtime attached. The Hub already stores data, versions it, gates it, previews it, and serves it to millions of people. It does not need a second system to hold environments. It needs a way to say "this data is an environment, and here is how you run it." | |
| We think this is the wrong shape. An environment is tasks, tests, containers, and a reward rule, which are data with a runtime on top. The Hub already stores data, versions it, gates it, previews it, and serves it to millions of people. It does not need a second system to hold environments. It needs a way to say "this data is an environment, and here is how you run it." |
|
|
||
| The frameworks keep doing what they are good at. The Hub does what it is good at, which is hosting, discovery, and versioning. Nobody has to own the catalogue. In fact, catalogues can run on other platforms too, powered by the hub. | ||
|
|
||
| Rl Environments are not a new repository type. An environment is a dataset repo, so it gets everything a dataset repo gets: gating, versioning, the viewer, discussions, and PRs. |
There was a problem hiding this comment.
we are repeating a bit here
|
|
||
| Rl Environments are not a new repository type. An environment is a dataset repo, so it gets everything a dataset repo gets: gating, versioning, the viewer, discussions, and PRs. | ||
|
|
||
| The Hub does not have to run your environment, but you can as jobs, if you need. The tags just describe compatibility and generate loading commands. Execution stays in the framework, on your hardware or your sandbox provider. |
There was a problem hiding this comment.
| The Hub does not have to run your environment, but you can as jobs, if you need. The tags just describe compatibility and generate loading commands. Execution stays in the framework, on your hardware or your sandbox provider. | |
| The dataset repository hosts your environment data, and you can use [Hugging Face Jobs](https://huggingface.co/docs/hub/en/jobs) to run the environment if you need. The tags just describe compatibility and generate loading commands. Execution stays in the framework, on your hardware or your sandbox provider. |
HF sandboxes?
| Every RL paper or framework ships its own way to find environments. A hub here, a registry there, a GitHub list of tasks with a custom loader. Each one is a small walled garden. If you publish an environment for one framework, users of the other three can’t load it. If you want to train on an environment from another framework or a new paper, you’ll need to port it by hand. | ||
|
|
||
| We think this is the wrong shape. An environment is tasks, tests, containers, and a reward rule. That is just data with a runtime attached. The Hub already stores data, versions it, gates it, previews it, and serves it to millions of people. It does not need a second system to hold environments. It needs a way to say "this data is an environment, and here is how you run it." | ||
|
|
There was a problem hiding this comment.
maybe you could add a small excalidraw illustration of an env and what happens inside for illustration (verifier, task, agent etc and what part is held in a dataset)
|
|
||
| taskset = vf.HarborTaskset( | ||
| config=vf.HarborTasksetConfig( | ||
| dataset="hf://datasets/<org>/<dataset>", |
There was a problem hiding this comment.
I wonder why they don't directly take repo id for convenience
| env = AutoEnv.from_env("<org>/<dataset>", trust_remote_code=False) | ||
| ``` | ||
|
|
||
| And NeMo Gym pulls the data down and evaluates against it: |
There was a problem hiding this comment.
this is a bit unclear, is it only used for evals?
|
|
||
| ### Verifiers: run a model on the same task | ||
|
|
||
| The [Verifiers v1 API](https://github.com/PrimeIntellect-ai/verifiers/blob/f2382d3c285ecb85578caf2948f24d0eeed84561/docs/v1/harbor.md) can read the same Harbor task directories. With Docker running and a tool-calling model served at `http://localhost:8000/v1`, replace `local-model` below with the served model ID. This example uses uv and Python 3.13, and pins the v1 source revision. The `bash` harness lets the model act in the task's container; the task verifier scores the completed run. |
There was a problem hiding this comment.
The Harbor integration of [verifiers v1](https://www.primeintellect.ai/blog/verifiers-v1) can run the same task directories in different runtimes, such as Docker. It also supports different harnesses, including a minimal bash harness.
uvx --python 3.13 --from 'verifiers[harbor]' eval harbor \
--env.taskset.repo https://huggingface.co/datasets/harborframework/terminal-bench-2.1 \
--env.taskset.dataset terminal-bench-2.1@2.1.0 \
--env.taskset.tasks '["regex-log"]' \
--env.agent.runtime.type docker \
--env.agent.harness.id bash \
--model "$MODEL" \
--client.base-url "$LLM_URL"Here repo is the full Hugging Face Git URL, while dataset is the name and version in that repo's registry.json. This loader uses Harbor's registry conventions, so a bare Hub repo ID cannot replace both values.
This PR adds the RL Environments blog post, metadata, an index entry, and a Harbor Space embed while preserving the prose. The search screenshot and NeMo example remain placeholders; the thumbnail reuses the dataset-filter image.