Skip to content

RFC 013: Task sampling and in-training evaluation - #1256

Open
Matheus-F-Scatolin wants to merge 1 commit into
huggingface:mainfrom
Matheus-F-Scatolin:rfc/013-task-sampling
Open

Matheus-F-Scatolin wants to merge 1 commit into
huggingface:mainfrom
Matheus-F-Scatolin:rfc/013-task-sampling

Conversation

@Matheus-F-Scatolin

@Matheus-F-Scatolin Matheus-F-Scatolin commented Sep 27, 2026 •

Copy link
Copy Markdown

Summary

Adds RFC 013, which proposes trainer-side task sampling and in-training evaluation on top of the existing Task API: a TaskSampler protocol (sequential, uniform and a cost-aware curriculum), a generic TaskAPIClient, an EnvEvalHarness built on the openenv collect pipeline, an EvalSchedule, and an openenv eval command. It targets the "simple evals that can be scored by a script" step of RFC 000's phase three. This PR only adds the RFC and its index entry; no code changes.

The RFC includes preliminary results from a single-seed pilot on a laptop (Llama-3.2-1B, GRPO + LoRA in MLX, reasoning_gym_env chain_sum in five difficulty tiers). Per environment episode, a signal-targeting curriculum learned faster than uniform sampling; per generated token, uniform sampling won, because the curriculum's harder tasks produced answers 2.4 times longer. That result is why the proposed CurriculumSampler weights tasks by expected signal per unit of cost, with the cost unit chosen by the trainer.

Test success per environment episode and per generated token

Pilot code and logs: Matheus-F-Scatolin/OpenEnv@rfc-013-pilot.

Type of Change

  • Bug fix
  • New feature
  • Breaking change
  • Documentation
  • New environment
  • Refactoring

Alignment Checklist

Before submitting, verify:

  • I have read .claude/docs/PRINCIPLES.md and this PR aligns with our principles
  • I have checked .claude/docs/INVARIANTS.md and no invariants are violated (one tension is flagged below)
  • I have run /pre-submit-pr (or bash .claude/hooks/lint.sh and tests) and addressed all issues. Not applicable: the PR changes only Markdown in rfcs/, which the lint and tests do not cover.

RFC Status

  • Not required (bug fix, docs, minor refactoring)
  • RFC exists: this PR is the RFC
  • RFC needed (will create before merge)

Test Plan

Docs only. To reproduce the preliminary results (Apple Silicon, MLX needs the Metal GPU):

git clone -b rfc-013-pilot https://github.com/Matheus-F-Scatolin/OpenEnv.git rfc-013-pilot
cd rfc-013-pilot/pilot && uv sync
caffeinate -i uv run python smoke.py train --model llama3.2-1b --sampler curriculum --steps 100000 --max-minutes 150 --eval-every 100 --eval-at-start --eval-per-tier 100 --out runs/pilot
caffeinate -i uv run python smoke.py train --model llama3.2-1b --sampler uniform --steps 100000 --max-minutes 150 --eval-every 100 --eval-at-start --eval-per-tier 100 --out runs/pilot
uv run --no-project --with matplotlib python rfc_figure.py runs/pilot ../figures/pilot-episodes-vs-tokens.png

Claude Code Review

Alignment review done with Claude Code, following .claude/skills/alignment-review by hand against this diff (not the /alignment-review command itself).

Automated Checks

  • Lint: N/A (Markdown only)
  • Debug code: N/A

Open RFCs Context

  • RFC: GRPO recipes #924 (GRPO recipes, draft): this RFC keeps trainers out of core, consistent with its non-goals, and offers reasoning_gym_env as a partial answer to its open question 2 (cheap, but single-turn with no tool use).
  • docs: add hf rl environment datasets rfc #795 (HF RL environment datasets): possible home for published openenv eval results (open question 4).
  • RFC 008 and Validate runtime repeatability #1247 (runtime repeatability): D4 applies the same repeatability idea at evaluation time instead of duplicating the validation probe.

Tier 2: Alignment Discussion

ALIGNMENT FLAG: Task metadata over HTTP

  • Principle/RFC at stake: INVARIANTS.md, Architectural Invariant 4 (WebSocket for all environment communication, "Gym-like API + metadata"; HTTP deprecation in [PATCH] Deprecate HTTP in websockets #252, Deprecate HTTP /reset, /step, /state #1156)
  • The concern: TaskAPIClient consumes the existing HTTP-only Task API routes, as documented in the Task API guide. The RFC adds no routes, but it adds a core consumer of them. D5 isolates this in one class so a later WebSocket move changes only TaskAPIClient.
  • Suggested reviewer: @Darktex, @adithya-s-k

ALIGNMENT FLAG: Task API helpers in core

  • Principle/RFC at stake: Task API guide ("The core clients (EnvClient, SyncEnvClient) do not ship Task API methods, because task specs are environment-specific")
  • The concern: The RFC adds a generic TaskAPIClient to openenv.core. D6 argues it keeps specs untyped and adds nothing to EnvClient, so the guide's reasoning still holds, but this reverses a documented choice and should be agreed explicitly.
  • Suggested reviewer: @adithya-s-k

ALIGNMENT FLAG: Evaluation coverage follows openenv collect

  • Principle/RFC at stake: Principle 1 (minimize lifecycle deltas)
  • The concern: Reusing the collect pipeline keeps training and evaluation on one path, but openenv collect resolves only openspiel: and reasoning_gym: specs today, so openenv eval would start with the same coverage (D3).
  • Suggested reviewer: @burtenshaw

No conflicts found with the dual API boundary, rewards-in-environment, client-server separation, or "one env = one trajectory": samplers run on the orchestration side and only read rewards the environment computes.

Disclosure

#1250 and #1252, cited in the RFC, are by one of the authors. The RFC and this description were drafted with Claude Code and reviewed by the authors.

🤖 Generated with Claude Code


Note

Low Risk
Markdown-only RFC and README index update; no production code or behavior changes.

Overview
Adds RFC 013 (rfcs/013-task-sampling-and-training-evals.md) and lists it under Reward & Evaluation in rfcs/README.md. There is no runtime or API code in this PR—only the design doc and index entry.

The RFC proposes trainer-side task sampling (TaskSampler with sequential, uniform, and cost-aware curriculum samplers), a shared TaskAPIClient, and in-training evaluation via EnvEvalHarness (reusing the collect rollout path), openenv eval, and EvalSchedule. It documents motivation (no shared task selection, wasted GRPO groups, missing reproducible env evals), design decisions (sampling on the orchestration side, curriculum weighted by GRPO signal per chosen cost unit), pilot results, a six-step rollout plan, and open questions (HTTP Task API vs WebSocket, distributed sampler state).

Reviewed by Cursor Bugbot for commit 0a13dbe. Bugbot is set up for automated code reviews on this repo. Configure here.

Proposes a TaskSampler protocol (sequential, uniform, cost-aware
curriculum), a generic TaskAPIClient, an EnvEvalHarness built on the
collect pipeline, EvalSchedule, and an `openenv eval` command. Includes
preliminary results from a single-seed pilot (Llama-3.2-1B, GRPO + LoRA,
reasoning_gym_env) and lists the index entry in rfcs/README.md.

Co-Authored-By: Tiago Perrupato Antunes <tiagoperrupato@gmail.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@cursor cursor Bot mentioned this pull request Sep 28, 2026
4 of 16 tasks

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant