Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
8 changes: 8 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -10,6 +10,13 @@

> Agent Lightning was completely refactored in v1.0. For legacy releases earlier than v1.0, see [this branch](https://github.com/microsoft/agent-lightning/tree/v0.x).

## ⚡ News

- [2026/09] We release [a new coding agent example based on an MoE model](https://microsoft.github.io/agent-lightning/stable/76-example-coding-agent-moe/). Pure RL improves Qwen3.5-35B-A3B on SWE-bench Verified from 47.8% to 61.6% after training it on only 1.8K training examples!
- [2026/08] [Agent Lightning Skill](https://github.com/microsoft/agent-lightning/tree/main/skills#agent-lightning-skill) is released. It helps a coding agent improve another AI agent against a benchmark.
- [2026/08] [Agent Lightning v1.0 technical report](https://arxiv.org/abs/2608.17528) is released!
- [2026/08] We open source Agent Lightning v1.0!

## ⚡ Key Features

- 🪶 **~3,500 lines of code:** We treat simplicity as the first principle.
Expand Down Expand Up @@ -74,6 +81,7 @@ We evaluate Agent Lightning v1.0 across several practical training domains, incl
| [Search-R1](https://microsoft.github.io/agent-lightning/stable/65-example-search-r1/) | Multi-turn retrieval and reasoning agent. |
| [LLM-in-Sandbox](https://microsoft.github.io/agent-lightning/stable/70-example-llm-in-sandbox/) | General agent with computer and code execution tools. |
| [Coding Agent](https://microsoft.github.io/agent-lightning/stable/75-example-coding-agent/) | Coding agent trained with repository tests. |
| [Coding Agent: MoE](https://microsoft.github.io/agent-lightning/stable/76-example-coding-agent-moe/) | Train Qwen3.5-35B-A3B with Megatron and R3. |

## ⚡ Articles

Expand Down
14 changes: 14 additions & 0 deletions docs/76-example-coding-agent-moe.md
Original file line number Diff line number Diff line change
Expand Up @@ -8,6 +8,20 @@ This is the MoE variant of the [Coding Agent](75-example-coding-agent.md) exampl
SWE-smith data, Kubernetes controller, repository images, agent, and reward. The existing
`Qwen/Qwen3.5-9B` FSDP path remains available through `examples/swe_smith/run.sh`.

## Results

With pure RL, Qwen3.5-35B-A3B improves from 47.8% to 61.6% on SWE-bench Verified after
1,792 training examples (about 1.8K), a gain of 13.8 percentage points. Results use the
official SWE-bench Verified harness on all 500 instances; trained example counts assume
16 examples per step.

| Step | Trained examples | Result on SWE-bench Verified |
|---|---:|---:|
| 0 (base) | 0 | 47.8% (239/500) |
| 64 | 1,024 | 58.8% (294/500) |
| 112 | 1,792 | 61.6% (308/500) |
| 176 | 2,816 | 60.6% (303/500) |

## Why R3

An MoE token can select different experts during rollout and training. R3 records vLLM's rollout
Expand Down
Loading