From a24456725e6196c9b1723c2a73f27c05b6eb2144 Mon Sep 17 00:00:00 2001 From: cheng874 Date: Wed, 30 Sep 2026 13:51:01 +0800 Subject: [PATCH 1/2] Draft update --- .../getting_started/flagtree-cpu.md | 59 +++++++++++++++++++ .../getting_started/getting-started.md | 1 + .../getting_started/flagtree-cpu.md | 59 +++++++++++++++++++ .../getting_started/getting-started.md | 1 + 4 files changed, 120 insertions(+) create mode 100644 docs/flagtree_en/getting_started/flagtree-cpu.md create mode 100644 docs/flagtree_zh/getting_started/flagtree-cpu.md diff --git a/docs/flagtree_en/getting_started/flagtree-cpu.md b/docs/flagtree_en/getting_started/flagtree-cpu.md new file mode 100644 index 0000000000..25937d905b --- /dev/null +++ b/docs/flagtree_en/getting_started/flagtree-cpu.md @@ -0,0 +1,59 @@ +[中文版|English] + +# Enable FlagTree CPU on Arm64 + +[FlagTree CPU](https://github.com/flagos-ai/flagtree-cpu/tree/triton_v3.7.x) provides a CPU backend for Triton, allowing kernels written with Triton's Python API to compile and run on CPUs. Its compiler lowers Triton kernels through MLIR/LLVM into native CPU code, while its runtime handles CPU target selection, kernel launches, and compilation caching. This guide covers the Triton 3.7.2 implementation on Linux Arm64 and validates it with a vector-add kernel and FlagGems W4A8 operator tests. + +Build and validate the Triton 3.7.2 CPU backend using the scripts in [flagos-ai/community](https://github.com/flagos-ai/community/tree/main/fep/sig-edge/scripts/flagtree-cpu37). Follow these steps in the same shell. + +## Step 1 Prepare the host + +Use **Debian 13 on Linux Arm64 (`aarch64`)**, with Git installed, sudo access (or root), and at least **25 GiB free disk space**. Ensure the host can reach the source repositories and dependency download services. + +## Step 2 Get the scripts + +```bash +git clone --single-branch --branch main \ + https://github.com/flagos-ai/community.git community-flagtree +cd community-flagtree + +RUN=fep/sig-edge/scripts/flagtree-cpu37/run.sh +export WORK_DIR="$HOME/flagtree-cpu-3.7-test" +export BUILD_JOBS=4 +``` + +The dedicated work directory separates the compiler, Python environment, and caches from other installations. Keep the full script directory; `run.sh` depends on its companion files. + +## Step 3 Build and install + +```bash +bash "$RUN" setup +``` + +Setup installs system dependencies, creates a Python 3.11 environment, fetches pinned FlagTree CPU and FlagGems revisions, initializes submodules, and builds Triton 3.7.2. Pinned LLVM and SLEEF versions make the build reproducible. The installed package imports as `triton`. + +## Step 4 Validate CPU execution + +```bash +bash "$RUN" test +``` + +Tests check source/toolchain revisions and the CPU target, execute a real vector-add kernel with fresh and reused caches, and compare FlagGems W4A8 results against PyTorch. Acceptance requires **7 passed, 0 skipped**. Logs and test reports are saved under `$WORK_DIR/logs/`. + +## Step 5 Enable the backend for your application + +```bash +source "$WORK_DIR/.venv/bin/activate" +export TRITON_CPU_BACKEND=1 +export FLAGGEMS_VENDOR=arm +export TRITON_CACHE_DIR="$WORK_DIR/triton-runtime-cache" +mkdir -p "$TRITON_CACHE_DIR" + +python your_program.py +``` + +Replace `your_program.py` with your application. `TRITON_CPU_BACKEND=1` selects CPU execution; `FLAGGEMS_VENDOR=arm` selects the FlagGems Arm implementation. A persistent cache lets compatible kernel specializations reuse compiled code. These variables must be exported in your application shell because the setup/test scripts run in separate Bash processes. + +Passing these checks validates the compiler and tested operators; model integration and model cold-start performance require separate validation. + +Reference: [FEP-0082](https://github.com/flagos-ai/community/blob/main/fep/sig-edge/0082-flagtree-cpu-bump-to-triton-3_7.md). diff --git a/docs/flagtree_en/getting_started/getting-started.md b/docs/flagtree_en/getting_started/getting-started.md index 245bc3bfd7..9c9f7b259d 100644 --- a/docs/flagtree_en/getting_started/getting-started.md +++ b/docs/flagtree_en/getting_started/getting-started.md @@ -8,6 +8,7 @@ This section covers the requirements for installing and running FlagTree and gui requirements.md install.md install-arm64-cpu.md +flagtree-cpu.md multi-backend-prebuilt-docker-image-install/install-nv.md multi-backend-prebuilt-docker-image-install/install-tileir.md multi-backend-prebuilt-docker-image-install/install-amd.md diff --git a/docs/flagtree_zh/getting_started/flagtree-cpu.md b/docs/flagtree_zh/getting_started/flagtree-cpu.md new file mode 100644 index 0000000000..936a358943 --- /dev/null +++ b/docs/flagtree_zh/getting_started/flagtree-cpu.md @@ -0,0 +1,59 @@ +[英文版|中文版] + +# 在 Arm64 上启用 FlagTree CPU + +[FlagTree CPU](https://github.com/flagos-ai/flagtree-cpu/tree/triton_v3.7.x) 为 Triton 提供 CPU 后端,使使用 Triton Python API 编写的 kernel 能够在 CPU 上编译和运行。它的编译器通过 MLIR/LLVM 将 Triton kernel 下降为原生 CPU 代码,运行时负责 CPU 目标选择、kernel 启动和编译缓存。本指南覆盖 Linux Arm64 上的 Triton 3.7.2 实现,并通过一个 vector-add kernel 和 FlagGems W4A8 算子测试进行验证。 + +使用 [flagos-ai/community](https://github.com/flagos-ai/community/tree/main/fep/sig-edge/scripts/flagtree-cpu37) 中的脚本构建并验证 Triton 3.7.2 CPU 后端。请在同一个 shell 中按以下步骤操作。 + +## 步骤 1 准备主机 + +使用 **Debian 13(Linux Arm64,`aarch64`)**,已安装 Git、具备 sudo 权限(或 root),并且至少有 **25 GiB 可用磁盘空间**。确保主机能够访问源码仓库和依赖下载服务。 + +## 步骤 2 获取脚本 + +```bash +git clone --single-branch --branch main \ + https://github.com/flagos-ai/community.git community-flagtree +cd community-flagtree + +RUN=fep/sig-edge/scripts/flagtree-cpu37/run.sh +export WORK_DIR="$HOME/flagtree-cpu-3.7-test" +export BUILD_JOBS=4 +``` + +专用工作目录将编译器、Python 环境和缓存与其他安装隔离开来。请保留完整的脚本目录;`run.sh` 依赖其配套文件。 + +## 步骤 3 构建并安装 + +```bash +bash "$RUN" setup +``` + +setup 会安装系统依赖、创建 Python 3.11 环境、拉取固定版本的 FlagTree CPU 与 FlagGems 修订、初始化子模块,并构建 Triton 3.7.2。固定的 LLVM 与 SLEEF 版本保证构建可复现。安装后的包以 `triton` 名称导入。 + +## 步骤 4 验证 CPU 执行 + +```bash +bash "$RUN" test +``` + +测试会检查源码/工具链修订与 CPU 目标,分别使用全新缓存和复用缓存执行真实的 vector-add kernel,并将 FlagGems W4A8 结果与 PyTorch 对比。验收要求 **7 passed, 0 skipped**。日志与测试报告保存在 `$WORK_DIR/logs/` 下。 + +## 步骤 5 为你的应用启用该后端 + +```bash +source "$WORK_DIR/.venv/bin/activate" +export TRITON_CPU_BACKEND=1 +export FLAGGEMS_VENDOR=arm +export TRITON_CACHE_DIR="$WORK_DIR/triton-runtime-cache" +mkdir -p "$TRITON_CACHE_DIR" + +python your_program.py +``` + +将 `your_program.py` 替换为你的应用。`TRITON_CPU_BACKEND=1` 选择 CPU 执行;`FLAGGEMS_VENDOR=arm` 选择 FlagGems 的 Arm 实现。持久化缓存使与之兼容的 kernel 特化可以复用已编译的代码。这些变量必须在你的应用 shell 中导出,因为 setup/test 脚本运行在独立的 Bash 进程中。 + +通过上述检查可以验证编译器和已测算子;模型集成与模型冷启动性能需要单独验证。 + +参考:[FEP-0082](https://github.com/flagos-ai/community/blob/main/fep/sig-edge/0082-flagtree-cpu-bump-to-triton-3_7.md)。 diff --git a/docs/flagtree_zh/getting_started/getting-started.md b/docs/flagtree_zh/getting_started/getting-started.md index 44953df07f..164e2d5620 100644 --- a/docs/flagtree_zh/getting_started/getting-started.md +++ b/docs/flagtree_zh/getting_started/getting-started.md @@ -8,6 +8,7 @@ requirements.md install.md install-arm64-cpu.md +flagtree-cpu.md multi-backend-prebuilt-docker-image-install/install-nv.md multi-backend-prebuilt-docker-image-install/install-tileir.md multi-backend-prebuilt-docker-image-install/install-amd.md From 115ae4445f1b4d19d063eaa343cf6b662bf36417 Mon Sep 17 00:00:00 2001 From: cheng874 Date: Wed, 30 Sep 2026 14:30:26 +0800 Subject: [PATCH 2/2] Draft update --- docs/flagtree_en/getting_started/install.md | 1 + docs/flagtree_en/getting_started/requirements.md | 1 + docs/flagtree_en/release_notes/release_notes_v070.md | 1 + docs/flagtree_zh/getting_started/install.md | 1 + docs/flagtree_zh/getting_started/requirements.md | 1 + docs/flagtree_zh/release_notes/release_notes_v070.md | 1 + 6 files changed, 6 insertions(+) diff --git a/docs/flagtree_en/getting_started/install.md b/docs/flagtree_en/getting_started/install.md index 077d0037bb..9a050696ae 100644 --- a/docs/flagtree_en/getting_started/install.md +++ b/docs/flagtree_en/getting_started/install.md @@ -88,6 +88,7 @@ For installing FlagTree on different backends from source, see the following lis - [Tsingmicro](/getting_started/multi-backend-prebuilt-docker-image-install/install-tsingmicro.md) - [KLX](/getting_started/multi-backend-prebuilt-docker-image-install/install-xpu.md) - [ARM64 CPU](/getting_started/install-arm64-cpu.md) +- [ARM64 CPU (Triton 3.7.2)](/getting_started/flagtree-cpu.md) ## Option 3: Install wheel package diff --git a/docs/flagtree_en/getting_started/requirements.md b/docs/flagtree_en/getting_started/requirements.md index b13ffc499c..2a687eb7cb 100644 --- a/docs/flagtree_en/getting_started/requirements.md +++ b/docs/flagtree_en/getting_started/requirements.md @@ -21,6 +21,7 @@ Each backend is based on different versions of Triton, and therefore resides in |[triton_v3.4.x](https://github.com/flagos-ai/flagtree/tree/triton_v3.4.x)|NVIDIA
AMD
Sunrise|[nvidia](https://github.com/flagos-ai/FlagTree/tree/triton_v3.4.x/third_party/nvidia/)
[amd](https://github.com/flagos-ai/FlagTree/tree/triton_v3.4.x/third_party/amd/)
[sunrise](https://github.com/flagos-ai/FlagTree/tree/triton_v3.4.x/third_party/sunrise/)|3.4| |[triton_v3.5.x](https://github.com/flagos-ai/flagtree/tree/triton_v3.5.x)|NVIDIA
AMD
Enflame|[nvidia](https://github.com/flagos-ai/FlagTree/tree/triton_v3.5.x/third_party/nvidia/)
[amd](https://github.com/flagos-ai/FlagTree/tree/triton_v3.5.x/third_party/amd/)
[enflame](https://github.com/flagos-ai/FlagTree/tree/triton_v3.5.x/third_party/enflame/)|3.5| |[triton_v3.6.x](https://github.com/flagos-ai/flagtree/tree/triton_v3.6.x)|NVIDIA
AMD
Enflame
HYGON
Moore Threads
Thrive|[nvidia](https://github.com/flagos-ai/FlagTree/tree/triton_v3.6.x/third_party/nvidia/)
[amd](https://github.com/flagos-ai/FlagTree/tree/triton_v3.6.x/third_party/amd/)
[enflame](https://github.com/flagos-ai/FlagTree/tree/triton_v3.6.x/third_party/enflame/)
[hcu](https://github.com/flagos-ai/FlagTree/tree/triton_v3.6.x/third_party/hcu/)
[mthreads](https://github.com/flagos-ai/FlagTree/tree/triton_v3.6.x/third_party/mthreads/)
[damoacademy](https://github.com/flagos-ai/FlagTree/tree/triton_v3.6.x/third_party/thrive/)|3.6| +|[triton_v3.7.x](https://github.com/flagos-ai/flagtree/tree/triton_v3.7.x)|ARM64 cpu|[cpu](https://github.com/flagos-ai/FlagTree/tree/triton_v3.7.x/third_party/cpu/)|3.7| ## Features on different branches diff --git a/docs/flagtree_en/release_notes/release_notes_v070.md b/docs/flagtree_en/release_notes/release_notes_v070.md index 9d9d7d2cbd..8adb292370 100644 --- a/docs/flagtree_en/release_notes/release_notes_v070.md +++ b/docs/flagtree_en/release_notes/release_notes_v070.md @@ -22,6 +22,7 @@ - Upgraded the following backends to Triton 3.6 and added CI/CD: [sunrise](/getting_started/multi-backend-prebuilt-docker-image-install/install-sunrise.md), [xpu](/getting_started/multi-backend-prebuilt-docker-image-install/install-xpu.md), [iluvatar](/getting_started/multi-backend-prebuilt-docker-image-install/install-iluvatar.md), and [tsingmicro](/getting_started/multi-backend-prebuilt-docker-image-install/install-tsingmicro.md). - Added TLE support for the [amd](/getting_started/multi-backend-prebuilt-docker-image-install/install-amd.md) backend and added CI/CD. - [rpu](/getting_started/multi-backend-prebuilt-docker-image-install/install-rpu.md) (Huixi Intelligence, Triton 3.6) is also supported. On the 3.3.x branch, [ARM64 CPU](/getting_started/install-arm64-cpu.md) provides [TLE-CPU](/user_guide/use-tle-cpu.md). + - [ARM64 CPU](/getting_started/flagtree-cpu.md) provides a CPU backend based on Triton 3.7.2, validated on Linux Arm64 with a vector-add kernel and FlagGems W4A8 operator tests. - **DevTools (Debugger & Profiler)** - FlagPrism ([flagos-ai/FlagPrism](https://github.com/flagos-ai/FlagPrism)) provides debugging and performance-analysis tools for Triton programs, containing `flagtree.debugger` and `flagtree.profiler`, and is integrated into FlagTree as the `third_party/FlagPrism` submodule. It initially supports a subset of backends: Huawei Ascend, Iluvatar, and Moore Threads. diff --git a/docs/flagtree_zh/getting_started/install.md b/docs/flagtree_zh/getting_started/install.md index d2d5bdd363..2dd3b87aff 100644 --- a/docs/flagtree_zh/getting_started/install.md +++ b/docs/flagtree_zh/getting_started/install.md @@ -88,6 +88,7 @@ - [清微智能](/getting_started/multi-backend-prebuilt-docker-image-install/install-tsingmicro.md) - [KLX](/getting_started/multi-backend-prebuilt-docker-image-install/install-xpu.md) - [ARM64 CPU](/getting_started/install-arm64-cpu.md) +- [ARM64 CPU (Triton 3.7.2)](/getting_started/flagtree-cpu.md) ## 方式三:安装 wheel 包 diff --git a/docs/flagtree_zh/getting_started/requirements.md b/docs/flagtree_zh/getting_started/requirements.md index 1f68981f6c..3317fee9b4 100644 --- a/docs/flagtree_zh/getting_started/requirements.md +++ b/docs/flagtree_zh/getting_started/requirements.md @@ -21,6 +21,7 @@ |[triton_v3.4.x](https://github.com/flagos-ai/flagtree/tree/triton_v3.4.x)|NVIDIA
AMD
曦望芯科|[nvidia](https://github.com/flagos-ai/FlagTree/tree/triton_v3.4.x/third_party/nvidia/)
[amd](https://github.com/flagos-ai/FlagTree/tree/triton_v3.4.x/third_party/amd/)
[sunrise](https://github.com/flagos-ai/FlagTree/tree/triton_v3.4.x/third_party/sunrise/)|3.4| |[triton_v3.5.x](https://github.com/flagos-ai/flagtree/tree/triton_v3.5.x)|NVIDIA
AMD
燧原|[nvidia](https://github.com/flagos-ai/FlagTree/tree/triton_v3.5.x/third_party/nvidia/)
[amd](https://github.com/flagos-ai/FlagTree/tree/triton_v3.5.x/third_party/amd/)
[enflame](https://github.com/flagos-ai/FlagTree/tree/triton_v3.5.x/third_party/enflame/)|3.5| |[triton_v3.6.x](https://github.com/flagos-ai/flagtree/tree/triton_v3.6.x)|NVIDIA
AMD
燧原
海光信息
摩尔线程
Thrive|[nvidia](https://github.com/flagos-ai/FlagTree/tree/triton_v3.6.x/third_party/nvidia/)
[amd](https://github.com/flagos-ai/FlagTree/tree/triton_v3.6.x/third_party/amd/)
[enflame](https://github.com/flagos-ai/FlagTree/tree/triton_v3.6.x/third_party/enflame/)
[hcu](https://github.com/flagos-ai/FlagTree/tree/triton_v3.6.x/third_party/hcu/)
[mthreads](https://github.com/flagos-ai/FlagTree/tree/triton_v3.6.x/third_party/mthreads/)
[damoacademy](https://github.com/flagos-ai/FlagTree/tree/triton_v3.6.x/third_party/thrive/)|3.6| +|[triton_v3.7.x](https://github.com/flagos-ai/flagtree/tree/triton_v3.7.x)|ARM64 cpu|[cpu](https://github.com/flagos-ai/FlagTree/tree/triton_v3.7.x/third_party/cpu/)|3.7| ## 不同分支上的功能特性 diff --git a/docs/flagtree_zh/release_notes/release_notes_v070.md b/docs/flagtree_zh/release_notes/release_notes_v070.md index 1f0a30fec6..a8288acd3f 100644 --- a/docs/flagtree_zh/release_notes/release_notes_v070.md +++ b/docs/flagtree_zh/release_notes/release_notes_v070.md @@ -22,6 +22,7 @@ - 将以下后端升级至 Triton 3.6 并新增 CI/CD:[sunrise](/getting_started/multi-backend-prebuilt-docker-image-install/install-sunrise.md)、[xpu](/getting_started/multi-backend-prebuilt-docker-image-install/install-xpu.md)、[iluvatar](/getting_started/multi-backend-prebuilt-docker-image-install/install-iluvatar.md) 和 [tsingmicro](/getting_started/multi-backend-prebuilt-docker-image-install/install-tsingmicro.md)。 - 为 [amd](/getting_started/multi-backend-prebuilt-docker-image-install/install-amd.md) 后端新增 TLE 支持并新增 CI/CD。 - 另外还支持 [rpu](/getting_started/multi-backend-prebuilt-docker-image-install/install-rpu.md)(辉羲智能,Triton 3.6)。在 3.3.x 分支上,[ARM64 CPU](/getting_started/install-arm64-cpu.md) 提供 [TLE-CPU](/user_guide/use-tle-cpu.md)。 + - [ARM64 CPU](/getting_started/flagtree-cpu.md) 基于 Triton 3.7.2 提供 CPU 后端,已在 Linux Arm64 上通过 vector-add kernel 与 FlagGems W4A8 算子测试验证。 - **DevTools(调试器与性能分析器)** - FlagPrism([flagos-ai/FlagPrism](https://github.com/flagos-ai/FlagPrism))为 Triton 程序提供调试与性能分析工具,包含 `flagtree.debugger` 与 `flagtree.profiler`,以 `third_party/FlagPrism` 子模块集成在 FlagTree 中。先支持部分后端:华为昇腾、天数智芯、摩尔线程。