diff --git a/README.md b/README.md index 85d65f89a..d149b37ee 100644 --- a/README.md +++ b/README.md @@ -10,6 +10,13 @@ > Agent Lightning was completely refactored in v1.0. For legacy releases earlier than v1.0, see [this branch](https://github.com/microsoft/agent-lightning/tree/v0.x). +## ⚡ News + +- [2026/09] We release [a new coding agent example based on an MoE model](https://microsoft.github.io/agent-lightning/stable/76-example-coding-agent-moe/). Pure RL improves Qwen3.5-35B-A3B on SWE-bench Verified from 47.8% to 61.6% after training it on only 1.8K training examples! +- [2026/08] [Agent Lightning Skill](https://github.com/microsoft/agent-lightning/tree/main/skills#agent-lightning-skill) is released. It helps a coding agent improve another AI agent against a benchmark. +- [2026/08] [Agent Lightning v1.0 technical report](https://arxiv.org/abs/2608.17528) is released! +- [2026/08] We open source Agent Lightning v1.0! + ## ⚡ Key Features - 🪶 **~3,500 lines of code:** We treat simplicity as the first principle. @@ -74,6 +81,7 @@ We evaluate Agent Lightning v1.0 across several practical training domains, incl | [Search-R1](https://microsoft.github.io/agent-lightning/stable/65-example-search-r1/) | Multi-turn retrieval and reasoning agent. | | [LLM-in-Sandbox](https://microsoft.github.io/agent-lightning/stable/70-example-llm-in-sandbox/) | General agent with computer and code execution tools. | | [Coding Agent](https://microsoft.github.io/agent-lightning/stable/75-example-coding-agent/) | Coding agent trained with repository tests. | +| [Coding Agent: MoE](https://microsoft.github.io/agent-lightning/stable/76-example-coding-agent-moe/) | Train Qwen3.5-35B-A3B with Megatron and R3. | ## ⚡ Articles diff --git a/docs/76-example-coding-agent-moe.md b/docs/76-example-coding-agent-moe.md index 4b895e4a2..b9ca4f381 100644 --- a/docs/76-example-coding-agent-moe.md +++ b/docs/76-example-coding-agent-moe.md @@ -8,6 +8,20 @@ This is the MoE variant of the [Coding Agent](75-example-coding-agent.md) exampl SWE-smith data, Kubernetes controller, repository images, agent, and reward. The existing `Qwen/Qwen3.5-9B` FSDP path remains available through `examples/swe_smith/run.sh`. +## Results + +With pure RL, Qwen3.5-35B-A3B improves from 47.8% to 61.6% on SWE-bench Verified after +1,792 training examples (about 1.8K), a gain of 13.8 percentage points. Results use the +official SWE-bench Verified harness on all 500 instances; trained example counts assume +16 examples per step. + +| Step | Trained examples | Result on SWE-bench Verified | +|---|---:|---:| +| 0 (base) | 0 | 47.8% (239/500) | +| 64 | 1,024 | 58.8% (294/500) | +| 112 | 1,792 | 61.6% (308/500) | +| 176 | 2,816 | 60.6% (303/500) | + ## Why R3 An MoE token can select different experts during rollout and training. R3 records vLLM's rollout