RFC 013: Task sampling and in-training evaluation - #1256
Open
Matheus-F-Scatolin wants to merge 1 commit into
Open
Matheus-F-Scatolin wants to merge 1 commit into
Matheus-F-Scatolin wants to merge 1 commit into
Conversation
Proposes a TaskSampler protocol (sequential, uniform, cost-aware curriculum), a generic TaskAPIClient, an EnvEvalHarness built on the collect pipeline, EvalSchedule, and an `openenv eval` command. Includes preliminary results from a single-seed pilot (Llama-3.2-1B, GRPO + LoRA, reasoning_gym_env) and lists the index entry in rfcs/README.md. Co-Authored-By: Tiago Perrupato Antunes <tiagoperrupato@gmail.com> Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Adds RFC 013, which proposes trainer-side task sampling and in-training evaluation on top of the existing Task API: a
TaskSamplerprotocol (sequential, uniform and a cost-aware curriculum), a genericTaskAPIClient, anEnvEvalHarnessbuilt on theopenenv collectpipeline, anEvalSchedule, and anopenenv evalcommand. It targets the "simple evals that can be scored by a script" step of RFC 000's phase three. This PR only adds the RFC and its index entry; no code changes.The RFC includes preliminary results from a single-seed pilot on a laptop (Llama-3.2-1B, GRPO + LoRA in MLX,
reasoning_gym_envchain_sumin five difficulty tiers). Per environment episode, a signal-targeting curriculum learned faster than uniform sampling; per generated token, uniform sampling won, because the curriculum's harder tasks produced answers 2.4 times longer. That result is why the proposedCurriculumSamplerweights tasks by expected signal per unit of cost, with the cost unit chosen by the trainer.Pilot code and logs: Matheus-F-Scatolin/OpenEnv@rfc-013-pilot.
Type of Change
Alignment Checklist
Before submitting, verify:
.claude/docs/PRINCIPLES.mdand this PR aligns with our principles.claude/docs/INVARIANTS.mdand no invariants are violated (one tension is flagged below)/pre-submit-pr(orbash .claude/hooks/lint.shand tests) and addressed all issues. Not applicable: the PR changes only Markdown inrfcs/, which the lint and tests do not cover.RFC Status
Test Plan
Docs only. To reproduce the preliminary results (Apple Silicon, MLX needs the Metal GPU):
Claude Code Review
Alignment review done with Claude Code, following
.claude/skills/alignment-reviewby hand against this diff (not the/alignment-reviewcommand itself).Automated Checks
Open RFCs Context
reasoning_gym_envas a partial answer to its open question 2 (cheap, but single-turn with no tool use).openenv evalresults (open question 4).Tier 2: Alignment Discussion
ALIGNMENT FLAG: Task metadata over HTTP
/reset,/step,/state#1156)TaskAPIClientconsumes the existing HTTP-only Task API routes, as documented in the Task API guide. The RFC adds no routes, but it adds a core consumer of them. D5 isolates this in one class so a later WebSocket move changes onlyTaskAPIClient.ALIGNMENT FLAG: Task API helpers in core
EnvClient,SyncEnvClient) do not ship Task API methods, because task specs are environment-specific")TaskAPIClienttoopenenv.core. D6 argues it keeps specs untyped and adds nothing toEnvClient, so the guide's reasoning still holds, but this reverses a documented choice and should be agreed explicitly.ALIGNMENT FLAG: Evaluation coverage follows
openenv collectopenenv collectresolves onlyopenspiel:andreasoning_gym:specs today, soopenenv evalwould start with the same coverage (D3).No conflicts found with the dual API boundary, rewards-in-environment, client-server separation, or "one env = one trajectory": samplers run on the orchestration side and only read rewards the environment computes.
Disclosure
#1250 and #1252, cited in the RFC, are by one of the authors. The RFC and this description were drafted with Claude Code and reviewed by the authors.
🤖 Generated with Claude Code
Note
Low Risk
Markdown-only RFC and README index update; no production code or behavior changes.
Overview
Adds RFC 013 (
rfcs/013-task-sampling-and-training-evals.md) and lists it under Reward & Evaluation inrfcs/README.md. There is no runtime or API code in this PR—only the design doc and index entry.The RFC proposes trainer-side task sampling (
TaskSamplerwith sequential, uniform, and cost-aware curriculum samplers), a sharedTaskAPIClient, and in-training evaluation viaEnvEvalHarness(reusing the collect rollout path),openenv eval, andEvalSchedule. It documents motivation (no shared task selection, wasted GRPO groups, missing reproducible env evals), design decisions (sampling on the orchestration side, curriculum weighted by GRPO signal per chosen cost unit), pilot results, a six-step rollout plan, and open questions (HTTP Task API vs WebSocket, distributed sampler state).Reviewed by Cursor Bugbot for commit 0a13dbe. Bugbot is set up for automated code reviews on this repo. Configure here.