From 1a756573567627628eabe0c56ed5cec6ba9ee477 Mon Sep 17 00:00:00 2001 From: GitHub Actions Date: Tue, 29 Sep 2026 15:58:37 +0800 Subject: [PATCH 1/3] docs: highlight MoE coding agent results in key features Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> --- README.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/README.md b/README.md index d149b37ee..25fa19712 100644 --- a/README.md +++ b/README.md @@ -22,7 +22,7 @@ - 🪶 **~3,500 lines of code:** We treat simplicity as the first principle. - 🧩 **Train with real agent harnesses:** Agents interact with the model through the Agent Lightning v1.0 proxy with **ZERO changes**, while keeping tools, context, control flow, and environments in the loop. - ☸️ **Native Kubernetes support:** Run agents directly as Kubernetes Jobs without relying on external sandbox services. -- 💻 **Full coding agent training example:** Using only **6K training samples**, an end-to-end Qwen3.5-9B workflow improves SWE-bench Verified from **41.8% to 56.4%**, a gain of **14.6 percentage points**. We release the full pipeline, including data cleaning, reward-hacking prevention, and training scripts. +- 💻 **Full coding agent training example:** Using only **6K training samples**, an end-to-end Qwen3.5-9B workflow improves SWE-bench Verified from **41.8% to 56.4%**, a gain of **14.6 percentage points**. We release the full pipeline, including data cleaning, reward-hacking prevention, and training scripts. **Update:** Our [Qwen3.5-35B-A3B MoE coding agent](https://microsoft.github.io/agent-lightning/stable/76-example-coding-agent-moe/) improves SWE-bench Verified from **47.8% to 61.6%** with pure RL after only **1.8K training examples**, a gain of **13.8 percentage points**. ## ⚡ Installation From 6c9ce33dbc0c7ecc72b1052b357c499c82332ac9 Mon Sep 17 00:00:00 2001 From: GitHub Actions Date: Tue, 29 Sep 2026 16:00:33 +0800 Subject: [PATCH 2/3] docs: clarify MoE example announcement Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> --- README.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/README.md b/README.md index 25fa19712..1e00d4958 100644 --- a/README.md +++ b/README.md @@ -22,7 +22,7 @@ - 🪶 **~3,500 lines of code:** We treat simplicity as the first principle. - 🧩 **Train with real agent harnesses:** Agents interact with the model through the Agent Lightning v1.0 proxy with **ZERO changes**, while keeping tools, context, control flow, and environments in the loop. - ☸️ **Native Kubernetes support:** Run agents directly as Kubernetes Jobs without relying on external sandbox services. -- 💻 **Full coding agent training example:** Using only **6K training samples**, an end-to-end Qwen3.5-9B workflow improves SWE-bench Verified from **41.8% to 56.4%**, a gain of **14.6 percentage points**. We release the full pipeline, including data cleaning, reward-hacking prevention, and training scripts. **Update:** Our [Qwen3.5-35B-A3B MoE coding agent](https://microsoft.github.io/agent-lightning/stable/76-example-coding-agent-moe/) improves SWE-bench Verified from **47.8% to 61.6%** with pure RL after only **1.8K training examples**, a gain of **13.8 percentage points**. +- 💻 **Full coding agent training example:** Using only **6K training samples**, an end-to-end Qwen3.5-9B workflow improves SWE-bench Verified from **41.8% to 56.4%**, a gain of **14.6 percentage points**. We release the full pipeline, including data cleaning, reward-hacking prevention, and training scripts. **Update:** We release a [new coding agent training example](https://microsoft.github.io/agent-lightning/stable/76-example-coding-agent-moe/) based on Qwen3.5-35B-A3B. With pure RL, the model improves SWE-bench Verified from **47.8% to 61.6%** after only **1.8K training examples**, a gain of **13.8 percentage points**. ## ⚡ Installation From 97ed6eee49abe0aff75c88f039fd9d147a923ca4 Mon Sep 17 00:00:00 2001 From: GitHub Actions Date: Tue, 29 Sep 2026 16:01:06 +0800 Subject: [PATCH 3/3] docs: make pure RL the subject of MoE result Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> --- README.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/README.md b/README.md index 1e00d4958..19c3c7ea4 100644 --- a/README.md +++ b/README.md @@ -22,7 +22,7 @@ - 🪶 **~3,500 lines of code:** We treat simplicity as the first principle. - 🧩 **Train with real agent harnesses:** Agents interact with the model through the Agent Lightning v1.0 proxy with **ZERO changes**, while keeping tools, context, control flow, and environments in the loop. - ☸️ **Native Kubernetes support:** Run agents directly as Kubernetes Jobs without relying on external sandbox services. -- 💻 **Full coding agent training example:** Using only **6K training samples**, an end-to-end Qwen3.5-9B workflow improves SWE-bench Verified from **41.8% to 56.4%**, a gain of **14.6 percentage points**. We release the full pipeline, including data cleaning, reward-hacking prevention, and training scripts. **Update:** We release a [new coding agent training example](https://microsoft.github.io/agent-lightning/stable/76-example-coding-agent-moe/) based on Qwen3.5-35B-A3B. With pure RL, the model improves SWE-bench Verified from **47.8% to 61.6%** after only **1.8K training examples**, a gain of **13.8 percentage points**. +- 💻 **Full coding agent training example:** Using only **6K training samples**, an end-to-end Qwen3.5-9B workflow improves SWE-bench Verified from **41.8% to 56.4%**, a gain of **14.6 percentage points**. We release the full pipeline, including data cleaning, reward-hacking prevention, and training scripts. **Update:** We release a [new coding agent training example](https://microsoft.github.io/agent-lightning/stable/76-example-coding-agent-moe/) based on Qwen3.5-35B-A3B. Pure RL improves Qwen3.5-35B-A3B on SWE-bench Verified from **47.8% to 61.6%** after only **1.8K training examples**, a gain of **13.8 percentage points**. ## ⚡ Installation