Skip to content

Add single-node Qwen3 DAPO math example - #3951

Merged
felipemello1 merged 10 commits into
pytorch:mainfrom
felipemello1:feature/qwen3-4b-dapo-8k
Jul 21, 2026
Merged

felipemello1 merged 10 commits into
pytorch:mainfrom
felipemello1:feature/qwen3-4b-dapo-8k

Conversation

@felipemello1

@felipemello1 felipemello1 commented Jul 20, 2026 •

Copy link
Copy Markdown
Contributor

Summary

  • Add a single-node Qwen3-4B-Base DAPO-Math example with a TP=2 trainer and six independent TP=1 generators.
  • Train on the 12,643-row filtered DAPO-Math dataset and validate on AIME 2025 with Math-Verify rewards.
  • Document the recipe and the measured 150-step training result.
  • Smoke run: https://meta.wandb.io/felipemello/titan_rl/runs/3xhecewt
image image

@meta-cla meta-cla Bot added the CLA Signed This label is managed by the Meta Open Source bot. label Jul 20, 2026
@felipemello1
felipemello1 force-pushed the feature/qwen3-4b-dapo-8k branch from 78cab15 to 4b14a75 Compare July 21, 2026 01:22
@felipemello1
felipemello1 force-pushed the feature/qwen3-4b-dapo-8k branch from 4b14a75 to bb8e1d3 Compare July 21, 2026 03:00
Comment thread pyproject.toml Outdated
Comment on lines +21 to +22
# Rule-based verification for math RL examples
"math-verify==0.9.0",

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This should be recipe / example level of requirements, rather than repo-wise? Some one who works on pre-training / diffusion models may not need this lib.

If you need to put this in CI, we can do something similar to https://github.com/pytorch/torchtitan/blob/main/.ci/docker/requirements-vlm.txt

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Comment on lines +58 to +67
## 150-step result

The single-node TP=2 trial completed optimizer steps 0 through 149 in 1 hour 46 minutes. This trial used the earlier 7,168-token response and 9,216-token packing limits; the runnable configuration above restores the full 8,192-token response budget.

```text
metric first 10 steps last 10 steps
rollout reward, mean 0.067 0.401
rollout total length, mean 988 tokens 3,284 tokens
rollout truncation rate, mean 1.75% 14.55%
```

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Could you share a wandb log on the real job?

Comment on lines +96 to +98
![Qwen3-4B DAPO Math-Verify reward](./assets/qwen3_4b_7k_reward.png)

![Qwen3-4B DAPO mean response length](./assets/qwen3_4b_7k_response_length.png)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Curves in these pictures would be easy to get outdated and not reproducible on later commit. Maybe only add them in PR summary, or add a commit pin to the results.


Training uses the 12,643-row [filtered DAPO-Math dataset](https://huggingface.co/datasets/hamishivi/DAPO-Math-17k-Processed_filtered). Each row contains one user prompt and its verifiable final answer.

Validation uses all 30 problems from [AIME 2025](https://huggingface.co/datasets/opencompass/AIME2025). The same single-turn environment and Math-Verify reward are used for training and validation.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

can we also put a validation set accuracy figure in the readme so people can be more convinced

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

yes, i ahve to rerun it. I left as a todo at the bottom of the readme

@yichuan-w

Copy link
Copy Markdown
Member

LGTM, very clean PR

@felipemello1
felipemello1 merged commit e00b27d into pytorch:main Jul 21, 2026
7 of 10 checks passed
yichuan-w pushed a commit to yichuan-w/torchtitan that referenced this pull request Jul 28, 2026
## Summary

- Add a single-node Qwen3-4B-Base DAPO-Math example with a TP=2 trainer
and six independent TP=1 generators.
- Train on the 12,643-row filtered DAPO-Math dataset and validate on
AIME 2025 with Math-Verify rewards.
- Document the recipe and the measured 150-step training result.
- Smoke run: https://meta.wandb.io/felipemello/titan_rl/runs/3xhecewt

<img width="524" height="313" alt="image"
src="https://github.com/user-attachments/assets/b45f68a1-f3f4-4a1f-be45-f0c4ad6515f7"
/>
<img width="517" height="315" alt="image"
src="https://github.com/user-attachments/assets/252b3da4-30d8-4b5d-b4df-ec518efb9385"
/>

---------

Co-authored-by: Felipe Mello <felipemello@fb.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ciflow/rl ciflow/8gpu CLA Signed This label is managed by the Meta Open Source bot.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants