From 0a13dbe9c60a44ba673e707b62051b37f7f3bc1f Mon Sep 17 00:00:00 2001 From: Matheus F Scatolin Date: Sun, 27 Sep 2026 18:44:39 -0300 Subject: [PATCH] docs(rfcs): add RFC 013, task sampling and in-training evaluation Proposes a TaskSampler protocol (sequential, uniform, cost-aware curriculum), a generic TaskAPIClient, an EnvEvalHarness built on the collect pipeline, EvalSchedule, and an `openenv eval` command. Includes preliminary results from a single-seed pilot (Llama-3.2-1B, GRPO + LoRA, reasoning_gym_env) and lists the index entry in rfcs/README.md. Co-Authored-By: Tiago Perrupato Antunes Co-Authored-By: Claude Opus 5.5 --- rfcs/013-task-sampling-and-training-evals.md | 467 +++++++++++++++++++ rfcs/README.md | 1 + 2 files changed, 468 insertions(+) create mode 100644 rfcs/013-task-sampling-and-training-evals.md diff --git a/rfcs/013-task-sampling-and-training-evals.md b/rfcs/013-task-sampling-and-training-evals.md new file mode 100644 index 000000000..9170624a5 --- /dev/null +++ b/rfcs/013-task-sampling-and-training-evals.md @@ -0,0 +1,467 @@ +# RFC: Task Sampling and In-Training Evaluation + +**Status**: In Review +**Created**: 2026-09-26 +**Updated**: 2026-09-27 +**Authors**: @Matheus-F-Scatolin, @tiagoperrupato +**RFC ID**: 013 + +## Summary + +This RFC adds two trainer-side capabilities on top of the existing Task API: **task sampling**, which decides which task each episode runs, and **in-training evaluation**, which measures a policy checkpoint on a held-out split in a reproducible way. Together they turn "train on an environment" into a closed loop: evaluate a baseline, train, and evaluate the same fixed test set at regular steps, with the same environment and the same tool interface throughout. + +It introduces a small `TaskSampler` protocol with three implementations (`SequentialSampler`, `UniformSampler`, `CurriculumSampler`), a generic Task API client helper, an `EnvEvalHarness` that implements the existing `EvalHarness` base class by reusing the `openenv collect` rollout pipeline, and an `openenv eval` CLI command. The environment boundary does not change: environments keep computing rewards, and samplers only consume them. + +The protocol and the evaluation harness are the core of the proposal and are useful with any sampler. The curriculum is one pluggable sampler. A single-seed pilot on a laptop (see [Preliminary Results](#preliminary-results)) shows why it has to be cost-aware: a curriculum that targets GRPO signal reached higher test accuracy per environment episode than uniform sampling, but lower accuracy per generated token, because the harder tasks it chose produced answers 2.4 times longer. The `CurriculumSampler` therefore weights tasks by expected signal per unit of cost, and the trainer chooses the cost unit. + +## Motivation + +### Problem Statement + +**1. Task selection has no shared mechanism.** The Task API (`list_splits`, `list_tasks`, `num_tasks`, `get_task`, `get_task_range`, and `reset(split=..., index=...)`) lets a dataset-backed environment publish its tasks, but it is discovery only. Nothing in OpenEnv decides which task an episode should use. Harbor-backed environments and the wrappers generated by `openenv import` implement it; the hand-written dataset-backed environments in `envs/`, such as `finqa_env` and `reasoning_gym_env`, do not. The core clients deliberately ship no Task API methods, because task specs are environment-specific (Task API guide), so each environment client writes its own HTTP helpers; `HarborEnv` has its own, and the guide shows the same pattern for a LaTeX OCR client. + +Today, selecting tasks well is manual. The Harbor harness takes an explicit `indices` list, and its code explains why: a task that every generation solves and one that none solves both give `reward_std` 0, and "the band that splits has to be chosen per model" (`envs/harbor_env/harness.py`). This RFC turns that manual choice into a sampler that measures the band as training runs. + +**2. Small models waste most of their GRPO rollouts.** GRPO computes each completion's advantage relative to the other completions in its group. When all `G` completions for a prompt receive the same reward, every advantage is zero and the group contributes no gradient. For a binary reward with success probability `p`, the chance that a group of size `G` carries any signal is + +``` +P(signal) = 1 - p^G - (1 - p)^G +``` + +which is maximal at `p = 0.5` and collapses as `p` approaches 0 or 1. In our pilot, 80% of GRPO groups carried no gradient under uniform sampling (Llama-3.2-1B, `G = 8`). DAPO's dynamic sampling filters these groups after generation; a sampler that targets tasks near `p = 0.5` avoids generating them in the first place. What a wasted group costs depends on the environment: in a single-turn reasoning task it is mostly generated tokens, while in a sandboxed or tool-using environment it is mostly the episode itself (container time, tool calls, API calls). A sampler has to account for that cost, not only for the signal. + +**3. There is no native way to evaluate a checkpoint on an environment.** `openenv.core.evals` provides the abstract `EvalHarness` and an Inspect AI wrapper. There is no OpenEnv-native way to run a checkpoint over an environment's `test` split with fixed seeds, aggregate metrics, and compare two checkpoints. Reproducibility is also fragile today: at least two environments accept `seed` in `reset()` and discard it (`textarena_env` in #1102 and `finqa_env` in #1250), so two evaluation runs of the same checkpoint can see different tasks. + +RFC 000 schedules evals for phase three (0.6 to 0.9), starting with "simple evals that can be scored by a script". This RFC proposes that starting point, built so that training and evaluation share one interface (design principle 1: minimize lifecycle deltas). + +### Goals + +1. A trainer can choose tasks for each episode through one small protocol, independent of the training framework. +2. A reproducible evaluation of any policy reachable through an OpenAI-compatible endpoint, on any environment that implements the Task API, returning the existing `EvalResult` type. +3. A way to run that evaluation every N training steps and compare checkpoints task by task. +4. A curriculum sampler that raises the number of GRPO groups with non-zero advantage per unit of cost, where the trainer chooses the unit (environment episodes or generated tokens), with no environment changes beyond publishing tasks. +5. A reference experiment that runs on a single laptop-class machine and reports results per episode and per generated token. + +### Non-goals + +- Owning GRPO or any trainer implementation (consistent with the non-goals of the draft recipes RFC, #924). +- LLM-judge or multi-agent evals; this RFC covers evals scored by the environment's own reward. +- Changing `reset`, `step`, or `state` signatures, or the wire types. +- Exposing task selection to the agent. Samplers run in orchestration code only. +- Claiming that a curriculum beats uniform sampling in general. The samplers make the comparison easy to run; the answer depends on the environment and the cost unit. + +## Design + +### Architecture Overview + +``` + trainer process (orchestration side) environment container + ┌─────────────────────────────────────────────────────┐ ┌───────────────────────┐ + │ │ HTTP │ Task API routes │ + │ TaskAPIClient ──── list_splits / get_task_range ───┼──────▶│ (metadata only) │ + │ │ │ │ │ + │ ▼ │ │ │ + │ TaskSampler ── sample(n) ──▶ task specs │ │ │ + │ ▲ │ │ WS │ │ + │ │ ▼ │ │ reset(split, index, │ + │ update(task, rewards, rollout (harness / ├──────▶│ seed) │ + │ costs) CollectRunner) │ │ step(action) │ + │ └──── rewards, costs ◀─────┘ │◀──────┤ reward (computed │ + │ │ │ inside the env) │ + │ EnvEvalHarness = CollectRunner + SequentialSampler │ └───────────────────────┘ + │ + metric aggregation → EvalResult │ + └─────────────────────────────────────────────────────┘ +``` + +Training uses a `UniformSampler` or `CurriculumSampler` over the `train` split. Evaluation uses a `SequentialSampler` over the `test` split, with fixed seeds. Both feed the same rollout path, so the only difference between a training episode and an evaluation episode is which task is chosen and whether a gradient step follows. + +### Core Abstractions + +#### 1. Task spec conventions + +Task specs stay untyped (`Any`), as in the Task API guide. This RFC reserves a few optional keys so samplers and evaluators can work across environments: + +| Key | Type | Meaning | +|-----|------|---------| +| `id` | `str` | Stable task identifier. Already the convention. | +| `split` | `str` | Split name. Already the convention. | +| `index` | `int` | Position in the split. Already the convention. | +| `difficulty` | `int` or `float` | Optional. Higher means harder. Tiers (`0..K`) or a continuous score. | +| `group` | `str` | Optional. A family of related tasks (for example a reasoning_gym dataset name). | + +An environment that publishes neither `difficulty` nor `group` still works with every sampler; the curriculum then keys its statistics by task `id`. + +Tiers work best when they are close together. The curriculum can only concentrate on a tier whose success rate is near 0.5; in the pilot, once the middle tier reached about 95%, the next tier was still below 40% and no tier sat in between. The Task API guide should recommend enough tiers that neighbouring tiers differ by a modest step in difficulty. These keys are an addition to the guide, not a requirement: the Task API checklist and existing environments are unaffected. + +#### 2. `TaskAPIClient` + +A generic helper in `openenv.core` that calls the standard Task API routes for any environment, deriving the HTTP base URL from a `ws://` client URL: + +```python +class TaskAPIClient: + def __init__(self, base_url: str, env_name: str): ... + @classmethod + def from_env_client(cls, env: EnvClient, env_name: str | None = None) -> "TaskAPIClient": ... + def list_splits(self) -> list[dict[str, str]]: ... + def list_tasks(self, split: str) -> list[Any]: ... + def num_tasks(self, split: str) -> int: ... + def get_task(self, split: str, index: int) -> Any: ... + def get_task_range(self, split: str, start: int | None = None, + stop: int | None = None) -> list[Any]: ... + def iter_tasks(self, split: str, page_size: int = 256) -> Iterator[Any]: ... +``` + +It returns task specs as untyped JSON, exactly as the routes serve them, and maps `501` to `NotImplementedError` and `400` to `IndexError`, mirroring the server. `iter_tasks` pages with `get_task_range` and relies on `num_tasks` for the true total, since `list_tasks` may return a bounded preview. Environment clients can keep their own typed helpers (for example `HarborEnv.get_task` returning a `HarborTaskRef`) and build them on this client. See D6 for why this belongs in core although core clients currently leave Task API methods out. + +#### 3. `TaskSampler` + +```python +class TaskSampler(Protocol): + def sample(self, n: int) -> list[Any]: + """Return n task specs for the next n prompts (GRPO groups).""" + + def update(self, task: Any, rewards: Sequence[float], + costs: Sequence[float] | None = None) -> None: + """Report the rewards of every rollout of one task (one GRPO group). + + costs, if given, has one entry per rollout in the trainer's cost unit + (for example the number of generated tokens). + """ + + def state_dict(self) -> dict[str, Any]: ... + def load_state_dict(self, state: dict[str, Any]) -> None: ... +``` + +`state_dict` and `load_state_dict` let a trainer checkpoint and resume the sampler together with the model, which a long run on one machine needs. (The pilot lost its intermediate checkpoints because they were not saved separately; resuming requires the model, the optimizer and the sampler state together.) + +Every sampler keeps the same counters, so every training log can report efficiency the same way: + +- `groups`, `signal_groups` (groups with non-zero reward variance), `episodes`, and `cost` (the sum of reported costs) +- derived: `signal_groups_per_episode`, `signal_groups_per_cost`, and `zero_signal_fraction` + +The zero-signal fraction on its own is misleading: in the pilot the curriculum had the lower zero-signal fraction and the worse result per token. Logs should report signal groups per episode and per unit of cost. + +Three implementations ship in core: + +- **`SequentialSampler(tasks, seeds)`**: deterministic enumeration of a split, optionally crossed with several seeds. Used by evaluation. `update` only updates the counters. +- **`UniformSampler(tasks, rng_seed)`**: shuffled sampling without replacement per epoch. The baseline for training. +- **`CurriculumSampler(tasks, group_size, key="difficulty", cost="episodes", ...)`**: described below. + +`SequentialSampler` and `UniformSampler` are also valid `tasks` iterables for the existing `CollectRunner`, so `openenv collect` and the evaluation harness gain task selection without changes to `CollectRunner`. An adaptive sampler also needs each episode's reward, which `CollectRunner` does not report back to its `tasks` iterable today; trainers call `update` from their own loop (see the example), and open question 7 asks whether `CollectRunner` should take a per-episode callback. + +#### 4. `CurriculumSampler` + +The sampler keeps, for each key `k` (a difficulty tier, a group, or a task id), decayed success counts and an estimate `p̂_k`: + +``` +s_k ← γ · s_k + successes +f_k ← γ · f_k + failures +p̂_k = (s_k + 1) / (s_k + f_k + 2) # Beta(1, 1) prior +``` + +A rollout counts as a success when its reward is at least `success_threshold` (default `1.0` for binary rewards; configurable for graded rewards). The decay `γ` (default `0.9` per update) lets the estimate follow a policy that is improving. + +It also keeps a decayed mean cost per episode, `ĉ_k`, from the `costs` passed to `update`: + +``` +ĉ_k ← γ · ĉ_k + (1 - γ) · mean(costs) +``` + +Each key is sampled with probability proportional to its expected signal per unit of cost: + +``` +w_k = (1 - p̂_k^G - (1 - p̂_k)^G + ε) / ĉ_k +``` + +where `G` is the trainer's group size and `ε` (default `0.05`) keeps every key reachable so the sampler can notice when a hard tier becomes learnable. Unseen keys start at `p̂ = 0.5`, the maximum signal, and at the mean cost of the keys seen so far, so every tier is tried early. Within a key, tasks are drawn uniformly without replacement. + +The `cost` argument selects the unit: + +- `"episodes"` (default): `ĉ_k = 1`, and the weight is the probability that a group carries signal. This is the right unit when episodes are expensive and completions are short, as in sandboxed, tool-using, or API-backed environments. +- `"tokens"`: the trainer passes the number of generated tokens per rollout. This is the right unit when generation dominates the cost, as in single-turn reasoning tasks on local hardware. +- A callable `cost(task, rollout) -> float` for anything else, such as wall-clock seconds. + +Weighting by `P(signal) / ĉ_k` is a heuristic, not an optimum: it moves probability toward keys with more signal per unit of cost without abandoning the others. Whether it closes the gap to uniform sampling in the per-token setting is the main open question of the reference experiment. + +#### 5. `EnvEvalHarness` + +An implementation of the existing `EvalHarness` base class: + +```python +class EnvEvalHarness(EvalHarness): + def __init__(self, session_factory: ResourceSessionFactory, + harness_adapter: HarnessAdapter, + task_client: TaskAPIClient, + output_dir: Path | None = None): ... + + def run(self, harness_version, library_versions, dataset, eval_parameters) -> dict: + ... +``` + +`eval_parameters` accepts: + +| Key | Default | Meaning | +|-----|---------|---------| +| `split` | `"test"` | Split to evaluate. | +| `limit` | `None` | Evaluate only the first `limit` tasks (for cheap in-training evals). | +| `seeds` | `[0]` | Seeds crossed with each task. | +| `samples_per_task` | `1` | `n` completions per task, for pass@k. | +| `temperature` | `0.0` | Sampling temperature; must be > 0 when `samples_per_task > 1`. | +| `success_threshold` | `1.0` | Reward at or above which an episode counts as solved. | +| `model`, `llm_endpoint`, `max_tokens`, `max_turns` | as in `openenv collect` | Policy and rollout limits. | + +It builds a `SequentialSampler` from the Task API, drives it through `CollectRunner`, writes every episode as JSONL through the existing `RolloutSerializer`, and returns scores: + +- `success_rate` with its standard error, `mean_reward`, `mean_turns`, `mean_completion_tokens` +- `pass@k` for each `k ≤ samples_per_task`, with the unbiased estimator `1 - C(n - c, k) / C(n, k)` +- `error_rate`: episodes that ended in a harness or parse failure, reported separately so a formatting failure is not mistaken for a wrong answer +- `by_difficulty` and `by_group` breakdowns when task specs carry those keys +- `per_task`: task id → outcome, which is what checkpoint comparison needs + +A helper compares two results: + +```python +diff = compare_eval_results(before, after) +diff.fixed # task ids that went from fail to pass +diff.regressed # task ids that went from pass to fail +``` + +#### 6. `openenv eval` + +A CLI command with the same flags as `openenv collect` for environment, endpoint and model, plus the `eval_parameters` above. It writes episodes and an `EvalResult` JSON to `--output-dir`, and `--compare ` prints the fixed and regressed tasks. + +#### 7. Trainer integration + +Core provides a framework-neutral schedule object; framework adapters live in examples, following the recipes RFC: + +```python +class EvalSchedule: + def __init__(self, harness: EnvEvalHarness, config: EvalConfig, + every_n_steps: int, on_result: Callable[[int, EvalResult], None]): ... + def maybe_run(self, step: int) -> EvalResult | None: ... +``` + +A TRL callback, an MLX loop, or a plain Python loop calls `maybe_run(step)` after each optimizer step. The policy under evaluation is served through an OpenAI-compatible endpoint (vLLM, `mlx_lm.server`, or any other server), so the harness does not depend on how the trainer holds the weights. `on_result` receives the step, and the trainer should log the sampler's `episodes` and `cost` counters next to it, so learning curves can be drawn against episodes and against cost, not only against steps or wall-clock time. + +### Key Design Decisions + +#### D1. Sampling lives on the trainer side, not in the environment + +- **Chosen approach**: Samplers run in orchestration code and select tasks through `reset(split=..., index=...)`. Environments only publish task metadata. +- **Rationale**: This keeps the invariants intact. Rewards are still computed inside the environment; the sampler only reads them. The agent has no path to influence which task it gets, because selection goes through `reset()`, which stays on the infrastructure side of the dual API boundary. The environment server stays free of cross-episode state, which also keeps "one env = one trajectory" and stacking of environment instances simple. It is also where task selection already happens today: the Harbor harness picks `indices` on the trainer side. +- **Trade-offs**: Each trainer process holds its own statistics. A distributed trainer with many rollout workers must aggregate them; `state_dict` gives a merge point, but a shared sampler service is out of scope for this RFC. + +#### D2. The curriculum targets GRPO signal per unit of cost + +- **Chosen approach**: Weight tasks by the probability that a group of size `G` has non-zero reward variance, divided by the expected cost of an episode in a unit the trainer chooses. +- **Rationale**: The probability of signal is the quantity that decides whether a GRPO group produces a gradient, and it needs one parameter the trainer already knows (`G`). Signal alone is not enough: in the pilot, the signal-only curriculum produced 39% more gradient-bearing groups per episode than uniform sampling, but 42% fewer per generated token, because the tiers it favoured had completions 2.4 times longer. Which of the two matters depends on where the environment's cost is, so the unit is a parameter rather than a fixed choice. Learning-progress curricula (weighting by the change in success rate) remain a reasonable alternative and can be added as another `TaskSampler`. +- **Trade-offs**: Gradient-bearing groups are a proxy for learning, not the objective. In the pilot, uniform sampling trained mostly on an easier tier and still improved the two hardest tiers, from 7% and 0% to 46% and 24%. The evaluation harness, not the sampler's counters, is what decides whether a sampler helps. The estimates also need repeated observations per key: for large datasets without `difficulty` or `group`, per-task statistics are sparse and the sampler behaves close to uniform until tasks repeat. Publishing closely spaced difficulty tiers is the intended remedy. + +#### D3. Evaluation reuses the collect pipeline + +- **Chosen approach**: `EnvEvalHarness` is `CollectRunner` plus a `SequentialSampler` plus aggregation. +- **Rationale**: Training rollouts, collected datasets, and evaluation episodes then go through one code path and one serialization format (design principle 1). Evaluation episodes can be pushed to the Hub with the existing `push_to_hf_hub` for inspection. +- **Trade-offs**: Evaluation inherits the limits of the collect pipeline. The `openenv collect` CLI resolves only `openspiel:` and `reasoning_gym:` environment specs today, and `--llm-endpoint` rejects URLs with a port or path (#1188). `openenv eval` would start with the same environment coverage and grow with `collect`; the Python `EnvEvalHarness` accepts any `ResourceSessionFactory` from the start. + +#### D4. Evaluation checks reproducibility instead of assuming it + +- **Chosen approach**: Before a run, `EnvEvalHarness` resets two fresh sessions with the same `(split, index, seed)` and compares a hash of the initial observations. If they differ, it warns and records `reproducible: false` in the result. +- **Rationale**: An eval that silently sees different tasks across checkpoints produces curves that cannot be trusted. Environments that ignore `seed` exist today (#1102, #1250), and the RFC 008 repeatability probes (#1247) address the same concern at validation time. +- **Trade-offs**: Two extra resets per run. The check covers the initial observation only, not later nondeterminism in the environment. + +#### D5. Task metadata stays on HTTP for now + +- **Chosen approach**: `TaskAPIClient` uses the existing HTTP Task API routes. This RFC adds no routes and no wire types. +- **Rationale**: The Task API guide documents these routes as HTTP-only, with no WebSocket message types, so that the `/ws` session protocol stays focused on `reset`, `step`, `state` and `close`. The routes serve metadata, never the step loop. +- **Trade-offs**: This sits in tension with the communication-patterns invariant, which lists WebSocket for all environment communication, "Gym-like API + metadata", while HTTP is being deprecated (#252, #1156). This RFC does not resolve that tension; it builds on the Task API as documented. Because every sampler and the evaluation harness go through `TaskAPIClient`, moving the Task API to WebSocket later changes only that class. See open question 1. + +#### D6. A generic Task API client in core + +- **Chosen approach**: Add `TaskAPIClient` to `openenv.core`, separate from `EnvClient`, returning untyped task specs. +- **Rationale**: The Task API guide leaves Task API methods out of the core clients because task specs are environment-specific. `TaskAPIClient` keeps specs untyped and adds no methods to `EnvClient`, so that reasoning still holds; what it removes is the duplicated plumbing that is not environment-specific (route names, HTTP base derivation from a `ws://` URL, paging, error mapping). Samplers and the evaluation harness need to list tasks for any environment, which is impossible if each client names its helpers differently. +- **Trade-offs**: One more public class in core. Typed task specs remain each environment client's job. + +## Examples + +### Server side: publishing difficulty tiers + +A procedurally generated environment exposes tiers as a finite split. For `reasoning_gym_env`, a task maps to a dataset name, a difficulty configuration, and a seed. The tiers below are the ones used in the pilot: + +```python +TIERS = [ + {"min_terms": 2, "max_terms": 2, "min_digits": 1, "max_digits": 1}, + {"min_terms": 3, "max_terms": 3, "min_digits": 2, "max_digits": 2}, + {"min_terms": 4, "max_terms": 4, "min_digits": 3, "max_digits": 3}, + {"min_terms": 5, "max_terms": 5, "min_digits": 4, "max_digits": 4, "allow_negation": True}, + {"min_terms": 6, "max_terms": 6, "min_digits": 5, "max_digits": 5, "allow_negation": True}, +] +TASKS_PER_TIER = {"train": 2000, "test": 100} +SPLIT_SEED_OFFSET = {"train": 0, "test": 500_000_000} # train and test tasks never collide + +class ReasoningGymEnvironment(Environment): + def list_splits(self) -> list[str]: + return ["train", "test"] + + def num_tasks(self, split: str) -> int: + return len(TIERS) * TASKS_PER_TIER[split] + + def get_task(self, split: str, index: int) -> dict: + if not 0 <= index < self.num_tasks(split): + raise IndexError(index) + tier = index // TASKS_PER_TIER[split] + return {"id": f"chain_sum-{split}-{index}", "split": split, "index": index, + "group": "chain_sum", "difficulty": tier} + + def reset(self, split: str | None = None, index: int | None = None, + seed: int | None = None, episode_id: str | None = None, **kwargs): + if index is None: + # No selection: keep today's seed-driven behaviour, as the Task API checklist asks. + return self._reset_from_kwargs(seed=seed, episode_id=episode_id, **kwargs) + split = split or "train" + task = self.get_task(split, index) + return self._reset_dataset("chain_sum", TIERS[task["difficulty"]], + seed=SPLIT_SEED_OFFSET[split] + index) +``` + +`_reset_from_kwargs` and `_reset_dataset` stand for the environment's existing reset paths. + +The pilot showed these five tiers are too coarse for a 1B model (see D2 and the task spec conventions); the reference experiment will use finer steps between tiers 2 and 4. + +### Client side: a plain training loop + +```python +tasks = TaskAPIClient.from_env_client(env).get_task_range("train") +sampler = CurriculumSampler(tasks, group_size=8, key="difficulty", cost="tokens") + +evaluator = EnvEvalHarness(session_factory, harness_adapter, task_client) +schedule = EvalSchedule( + evaluator, + EvalConfig(harness_name="EnvEvalHarness", harness_version="0.1", + library_versions={}, dataset="reasoning_gym:chain_sum", + eval_parameters={"split": "test", "llm_endpoint": POLICY_URL, + "model": "policy", "temperature": 0.0}), + every_n_steps=50, + on_result=lambda step, r: logger.log({"step": step, "episodes": sampler.episodes, + "cost": sampler.cost, **r.scores}), +) + +schedule.maybe_run(step=0) # baseline before any training +for step in range(1, num_steps + 1): + batch = sampler.sample(n=prompts_per_step) + groups = [rollout_group(task, G=8) for task in batch] # trainer-specific + for task, group in zip(batch, groups): + sampler.update(task, [ep.reward for ep in group], + costs=[ep.completion_tokens for ep in group]) + trainer_update(groups) # trainer-specific + logger.log({"step": step, + "signal_groups_per_episode": sampler.signal_groups_per_episode, + "signal_groups_per_cost": sampler.signal_groups_per_cost}) + schedule.maybe_run(step) +``` + +### CLI: comparing checkpoints + +The environment spec is positional, as in `openenv collect`: + +```bash +openenv eval reasoning_gym:chain_sum --split test \ + --llm-endpoint http://localhost:8080/v1 --model base \ + --output-dir runs/eval-step0000 + +openenv eval reasoning_gym:chain_sum --split test \ + --llm-endpoint http://localhost:8080/v1 --model ckpt-1000 \ + --output-dir runs/eval-step1000 --compare runs/eval-step0000/result.json +``` + +## Validation Plan + +The claims in this RFC are measured, not assumed. The reference experiment is designed to fit one laptop-class machine. + +### Preliminary Results + +We ran a pilot of the samplers before proposing the API. The code is a standalone script, not the proposed API: a GRPO loop in MLX with the `CurriculumSampler` logic above (signal-only weights, that is `cost="episodes"`). The code, the per-step logs and a figure of both learning curves are at [Matheus-F-Scatolin/OpenEnv@rfc-013-pilot](https://github.com/Matheus-F-Scatolin/OpenEnv/tree/rfc-013-pilot/pilot); the figure is also in the PR description. + +- **Model**: Llama-3.2-1B-Instruct, LoRA rank 16 on every linear layer (11.3M trainable parameters). +- **Environment**: `reasoning_gym_env` vendored from commit `9e0fd9b`, run in-process, with the five `chain_sum` tiers in the example above. 2,000 train tasks and 100 test tasks per tier, with disjoint seeds. The reward is whatever the environment's `step()` returns. +- **Training**: 8 prompts × 8 completions per step (64 episodes), temperature 1.0, on-policy GRPO with one update per batch, no KL term, token-level loss normalization. Groups with no reward variance are skipped. +- **Evaluation**: greedy decoding on all 500 test tasks every 100 steps. The standard error of one evaluation is about 2 points. +- **Budget**: 150 training minutes per sampler on a MacBook Pro (M5 Pro, 24 GB). One seed per arm. + +At the same number of environment episodes, the curriculum was ahead at 9 of 10 evaluation points: + +| Training episodes (steps) | Uniform | Curriculum | Difference | +|---|---|---|---| +| 6,400 (100) | 50.6% | 58.0% | +7.4 | +| 19,200 (300) | 62.2% | 65.2% | +3.0 | +| 32,000 (500) | 65.4% | 68.4% | +3.0 | +| 44,800 (700) | 62.8% | 63.6% | +0.8 | +| 51,200 (800) | 62.4% | 69.0% | +6.6 | +| 64,000 (1,000) | 67.4% | 71.4% | +4.0 | + +Over the whole run: + +| | Uniform | Curriculum | +|---|---|---| +| Steps in 150 minutes | 2,196 | 1,081 | +| Tokens per completion | 44 | 106 | +| Groups with no gradient | 79.5% | 71.5% | +| Gradient-bearing groups per 1,000 episodes | 25.7 | **35.6** | +| Gradient-bearing groups per million generated tokens | **578** | 337 | +| Test success, start → end | 52.0% → 73.2% | 52.0% → 72.4% | +| Test success at ~6 M generated tokens | **73.6%** | 67.8% | +| Tiers 3 and 4, start → end | 7%, 0% → 46%, 24% | 7%, 0% → 43%, 22% | + +What we take from it: + +1. **The mechanism works as designed.** The curriculum concentrated on the tier nearest `p = 0.5` and cut the share of groups without gradient. +2. **The winner depends on the cost unit.** Per episode, the curriculum learned faster: it reached 72.4% after 69,184 episodes, which uniform sampling reached only after about 100,000. Per generated token, uniform sampling was ahead from about 3.3 M tokens on, because the curriculum's harder tasks produced much longer completions. This is why the `CurriculumSampler` weights by signal per unit of cost and lets the trainer pick the unit. +3. **Tiers were too coarse.** After tier 2 reached about 95%, no tier was near 50%, which limited what any curriculum could do. + +Caveats: one seed per arm, so a difference of 2 to 4 points between single evaluations is within noise, although the curriculum was ahead at 9 of 10 matched points. The two arms ran one after the other on the same machine. The curriculum arm was slowed for its last 40 minutes by an underpowered charger. This affects the comparison by wall-clock time but not the per-episode or per-token comparisons, which do not depend on machine speed. + +### Reference experiment + +- **Arms**: `UniformSampler`, `CurriculumSampler(cost="episodes")`, and `CurriculumSampler(cost="tokens")`, with the same group size and three seeds each. +- **Environment**: `reasoning_gym_env` through the Task API (rollout step 3), with finer tiers between the pilot's tiers 2 and 4. +- **Hardware**: the same laptop class; the runs are sized to fit one overnight session, about 14 hours by our estimate from the pilot. A CUDA machine with vLLM reproduces the same experiment faster. +- **Measurements**: test success rate with standard error, overall and per tier, against training episodes and against generated tokens; gradient-bearing groups per episode and per token; wall-clock time per step, reported but not used for the comparison. +- **What would support the proposal**: the episode-weighted curriculum at least as good as uniform per episode, and the token-weighted curriculum at least as good as uniform per token, across seeds. If a curriculum arm does not beat uniform in its own unit, we report that as is; the sampler protocol and the evaluation harness do not depend on the curriculum winning. + +## Rollout Plan + +Each step is a separate PR: + +1. This RFC. +2. Reserved task spec keys in the Task API guide, including guidance on tier spacing, and `TaskAPIClient` with tests against a stub server. +3. Task API support in `reasoning_gym_env` with difficulty tiers; then `finqa_env`, after its `reset(seed)` fix (#1252). +4. `SequentialSampler`, `UniformSampler`, `CurriculumSampler`, with unit tests on synthetic reward and cost streams (for example, a simulated policy whose success rate per tier improves over time and whose completion length grows with difficulty). +5. `EnvEvalHarness`, `compare_eval_results`, and `openenv eval`. +6. `EvalSchedule`, the reference experiment as an example (an MLX loop for Apple Silicon and a TRL callback) in `examples/`. + +## Alternatives Considered + +- **Curriculum inside the environment.** The server would track success rates and pick the next task itself. Rejected: it makes the environment stateful across episodes, mixes orchestration into the environment boundary, and breaks when many trainer workers share or stack environment instances. +- **Only filtering zero-signal groups after generation (DAPO-style).** This is complementary, not a replacement: it still pays for generating the discarded groups. A trainer can use both. +- **Signal-only curriculum weights.** Simpler, and equivalent to `cost="episodes"`. Rejected as the only option because the pilot shows it loses to uniform sampling when generated tokens are the cost. +- **Relying on the Inspect AI harness for evaluation.** Inspect AI remains supported, but it runs its own task and solver definitions. An evaluation that goes through the same environment, tools and rollout path as training is what design principle 1 asks for. + +## Open Questions + +1. Should the Task API gain WebSocket message types, given #1156 and the WebSocket invariant, or remain an HTTP metadata surface? +2. Should `difficulty` be standardized further (for example, normalized to `[0, 1]`), or remain environment-defined and only compared within one environment? +3. Where should shared sampler state live for distributed trainers with many rollout workers? +4. Should `openenv eval` results be pushable to the Hub in a schema that the environments datasets RFC (#795) can reference? +5. What is the right default for `success_threshold` on environments whose rewards are graded rather than binary? +6. Should `cost` default to `"episodes"`, or should environments declare their dominant cost (for example in their manifest) so trainers can pick a sensible default? +7. Should `CollectRunner` accept a per-episode callback (task, `EpisodeRecord`) so adaptive samplers can be driven by `openenv collect` and not only by a trainer loop? + +## Related RFCs and Work + +- RFC 000: phase three (evals) of the roadmap. +- RFC 004 (Rubrics): the source of the rewards that samplers consume. +- RFC 005 (Agentic harnesses) and the collect pipeline: the rollout path reused by evaluation. +- RFC 008 (Environment auto-validation) and #1247: repeatability checks at validation time; D4 applies the same idea at evaluation time. +- Draft recipes RFC (#924): its open question 2 asks which environment is simple enough for smoke validation. `reasoning_gym_env` answers the "simple and cheap" half, but it is single-turn and does not exercise harness-mediated tool use. +- Environments datasets RFC (#795): possible home for published evaluation results. +- Task API guide (`docs/source/guides/task-api.md`) and the Harbor harness (`envs/harbor_env/harness.py`): the current Task API contract and the manual task selection this RFC generalizes. +- #1250 and its fix #1252 (by one of the authors): the `finqa_env` seed issue named in rollout step 3. diff --git a/rfcs/README.md b/rfcs/README.md index c60eec4a8..9f8e62525 100644 --- a/rfcs/README.md +++ b/rfcs/README.md @@ -89,6 +89,7 @@ Each RFC should include the following sections: ### Reward & Evaluation - [004-rubrics.md](./004-rubrics.md) - Rubric System for Reward Computation +- [013-task-sampling-and-training-evals.md](./013-task-sampling-and-training-evals.md) - Task Sampling and In-Training Evaluation: `TaskSampler` (uniform, cost-aware curriculum), `EnvEvalHarness` and `openenv eval` ### Agentic Harnesses - [005-agentic-harnesses.md](./005-agentic-harnesses.md) - Agentic Harness Integration (OpenClaw, Claude Code, etc.)