Repository navigation
Add single-node Qwen3 DAPO math example - #3951
Conversation
78cab15 to
4b14a75
Compare
4b14a75 to
bb8e1d3
Compare
| # Rule-based verification for math RL examples | ||
| "math-verify==0.9.0", |
There was a problem hiding this comment.
This should be recipe / example level of requirements, rather than repo-wise? Some one who works on pre-training / diffusion models may not need this lib.
If you need to put this in CI, we can do something similar to https://github.com/pytorch/torchtitan/blob/main/.ci/docker/requirements-vlm.txt
There was a problem hiding this comment.
Do we need this? How should I appreciate this vs. https://github.com/felipemello1/torchtitan/blob/main/torchtitan/experiments/rl/train.py
| ## 150-step result | ||
|
|
||
| The single-node TP=2 trial completed optimizer steps 0 through 149 in 1 hour 46 minutes. This trial used the earlier 7,168-token response and 9,216-token packing limits; the runnable configuration above restores the full 8,192-token response budget. | ||
|
|
||
| ```text | ||
| metric first 10 steps last 10 steps | ||
| rollout reward, mean 0.067 0.401 | ||
| rollout total length, mean 988 tokens 3,284 tokens | ||
| rollout truncation rate, mean 1.75% 14.55% | ||
| ``` |
There was a problem hiding this comment.
Could you share a wandb log on the real job?
Summary: Test Plan: Reviewers: Subscribers: Tasks: Tags:
|  | ||
|
|
||
|  |
There was a problem hiding this comment.
Curves in these pictures would be easy to get outdated and not reproducible on later commit. Maybe only add them in PR summary, or add a commit pin to the results.
|
|
||
| Training uses the 12,643-row [filtered DAPO-Math dataset](https://huggingface.co/datasets/hamishivi/DAPO-Math-17k-Processed_filtered). Each row contains one user prompt and its verifiable final answer. | ||
|
|
||
| Validation uses all 30 problems from [AIME 2025](https://huggingface.co/datasets/opencompass/AIME2025). The same single-turn environment and Math-Verify reward are used for training and validation. |
There was a problem hiding this comment.
can we also put a validation set accuracy figure in the readme so people can be more convinced
There was a problem hiding this comment.
yes, i ahve to rerun it. I left as a todo at the bottom of the readme
|
LGTM, very clean PR |
## Summary - Add a single-node Qwen3-4B-Base DAPO-Math example with a TP=2 trainer and six independent TP=1 generators. - Train on the 12,643-row filtered DAPO-Math dataset and validate on AIME 2025 with Math-Verify rewards. - Document the recipe and the measured 150-step training result. - Smoke run: https://meta.wandb.io/felipemello/titan_rl/runs/3xhecewt <img width="524" height="313" alt="image" src="https://github.com/user-attachments/assets/b45f68a1-f3f4-4a1f-be45-f0c4ad6515f7" /> <img width="517" height="315" alt="image" src="https://github.com/user-attachments/assets/252b3da4-30d8-4b5d-b4df-ec518efb9385" /> --------- Co-authored-by: Felipe Mello <felipemello@fb.com>
Summary