Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
59 changes: 59 additions & 0 deletions docs/flagtree_en/getting_started/flagtree-cpu.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,59 @@
[<a href="../../../flagtree_zh/getting_started/flagtree-cpu.html">中文版</a>|English]

# Enable FlagTree CPU on Arm64

[FlagTree CPU](https://github.com/flagos-ai/flagtree-cpu/tree/triton_v3.7.x) provides a CPU backend for Triton, allowing kernels written with Triton's Python API to compile and run on CPUs. Its compiler lowers Triton kernels through MLIR/LLVM into native CPU code, while its runtime handles CPU target selection, kernel launches, and compilation caching. This guide covers the Triton 3.7.2 implementation on Linux Arm64 and validates it with a vector-add kernel and FlagGems W4A8 operator tests.

Build and validate the Triton 3.7.2 CPU backend using the scripts in [flagos-ai/community](https://github.com/flagos-ai/community/tree/main/fep/sig-edge/scripts/flagtree-cpu37). Follow these steps in the same shell.

## Step 1 Prepare the host

Use **Debian 13 on Linux Arm64 (`aarch64`)**, with Git installed, sudo access (or root), and at least **25 GiB free disk space**. Ensure the host can reach the source repositories and dependency download services.

## Step 2 Get the scripts

```bash
git clone --single-branch --branch main \
https://github.com/flagos-ai/community.git community-flagtree
cd community-flagtree

RUN=fep/sig-edge/scripts/flagtree-cpu37/run.sh
export WORK_DIR="$HOME/flagtree-cpu-3.7-test"
export BUILD_JOBS=4
```

The dedicated work directory separates the compiler, Python environment, and caches from other installations. Keep the full script directory; `run.sh` depends on its companion files.

## Step 3 Build and install

```bash
bash "$RUN" setup
```

Setup installs system dependencies, creates a Python 3.11 environment, fetches pinned FlagTree CPU and FlagGems revisions, initializes submodules, and builds Triton 3.7.2. Pinned LLVM and SLEEF versions make the build reproducible. The installed package imports as `triton`.

## Step 4 Validate CPU execution

```bash
bash "$RUN" test
```

Tests check source/toolchain revisions and the CPU target, execute a real vector-add kernel with fresh and reused caches, and compare FlagGems W4A8 results against PyTorch. Acceptance requires **7 passed, 0 skipped**. Logs and test reports are saved under `$WORK_DIR/logs/`.

## Step 5 Enable the backend for your application

```bash
source "$WORK_DIR/.venv/bin/activate"
export TRITON_CPU_BACKEND=1
export FLAGGEMS_VENDOR=arm
export TRITON_CACHE_DIR="$WORK_DIR/triton-runtime-cache"
mkdir -p "$TRITON_CACHE_DIR"

python your_program.py
```

Replace `your_program.py` with your application. `TRITON_CPU_BACKEND=1` selects CPU execution; `FLAGGEMS_VENDOR=arm` selects the FlagGems Arm implementation. A persistent cache lets compatible kernel specializations reuse compiled code. These variables must be exported in your application shell because the setup/test scripts run in separate Bash processes.

Passing these checks validates the compiler and tested operators; model integration and model cold-start performance require separate validation.

Reference: [FEP-0082](https://github.com/flagos-ai/community/blob/main/fep/sig-edge/0082-flagtree-cpu-bump-to-triton-3_7.md).
1 change: 1 addition & 0 deletions docs/flagtree_en/getting_started/getting-started.md
Original file line number Diff line number Diff line change
Expand Up @@ -8,6 +8,7 @@ This section covers the requirements for installing and running FlagTree and gui
requirements.md
install.md
install-arm64-cpu.md
flagtree-cpu.md
multi-backend-prebuilt-docker-image-install/install-nv.md
multi-backend-prebuilt-docker-image-install/install-tileir.md
multi-backend-prebuilt-docker-image-install/install-amd.md
Expand Down
1 change: 1 addition & 0 deletions docs/flagtree_en/getting_started/install.md
Original file line number Diff line number Diff line change
Expand Up @@ -88,6 +88,7 @@ For installing FlagTree on different backends from source, see the following lis
- [Tsingmicro](/getting_started/multi-backend-prebuilt-docker-image-install/install-tsingmicro.md)
- [KLX](/getting_started/multi-backend-prebuilt-docker-image-install/install-xpu.md)
- [ARM64 CPU](/getting_started/install-arm64-cpu.md)
- [ARM64 CPU (Triton 3.7.2)](/getting_started/flagtree-cpu.md)

## Option 3: Install wheel package

Expand Down
1 change: 1 addition & 0 deletions docs/flagtree_en/getting_started/requirements.md
Original file line number Diff line number Diff line change
Expand Up @@ -21,6 +21,7 @@ Each backend is based on different versions of Triton, and therefore resides in
|[triton_v3.4.x](https://github.com/flagos-ai/flagtree/tree/triton_v3.4.x)|NVIDIA<br>AMD<br>Sunrise|[nvidia](https://github.com/flagos-ai/FlagTree/tree/triton_v3.4.x/third_party/nvidia/)<br>[amd](https://github.com/flagos-ai/FlagTree/tree/triton_v3.4.x/third_party/amd/)<br>[sunrise](https://github.com/flagos-ai/FlagTree/tree/triton_v3.4.x/third_party/sunrise/)|3.4|
|[triton_v3.5.x](https://github.com/flagos-ai/flagtree/tree/triton_v3.5.x)|NVIDIA<br>AMD<br>Enflame|[nvidia](https://github.com/flagos-ai/FlagTree/tree/triton_v3.5.x/third_party/nvidia/)<br>[amd](https://github.com/flagos-ai/FlagTree/tree/triton_v3.5.x/third_party/amd/)<br>[enflame](https://github.com/flagos-ai/FlagTree/tree/triton_v3.5.x/third_party/enflame/)|3.5|
|[triton_v3.6.x](https://github.com/flagos-ai/flagtree/tree/triton_v3.6.x)|NVIDIA<br>AMD<br>Enflame<br>HYGON<br>Moore Threads<br>Thrive|[nvidia](https://github.com/flagos-ai/FlagTree/tree/triton_v3.6.x/third_party/nvidia/)<br>[amd](https://github.com/flagos-ai/FlagTree/tree/triton_v3.6.x/third_party/amd/)<br>[enflame](https://github.com/flagos-ai/FlagTree/tree/triton_v3.6.x/third_party/enflame/)<br>[hcu](https://github.com/flagos-ai/FlagTree/tree/triton_v3.6.x/third_party/hcu/)<br>[mthreads](https://github.com/flagos-ai/FlagTree/tree/triton_v3.6.x/third_party/mthreads/)<br>[damoacademy](https://github.com/flagos-ai/FlagTree/tree/triton_v3.6.x/third_party/thrive/)|3.6|
|[triton_v3.7.x](https://github.com/flagos-ai/flagtree/tree/triton_v3.7.x)|ARM64 cpu|[cpu](https://github.com/flagos-ai/FlagTree/tree/triton_v3.7.x/third_party/cpu/)|3.7|

## Features on different branches

Expand Down
1 change: 1 addition & 0 deletions docs/flagtree_en/release_notes/release_notes_v070.md
Original file line number Diff line number Diff line change
Expand Up @@ -22,6 +22,7 @@
- Upgraded the following backends to Triton 3.6 and added CI/CD: [sunrise](/getting_started/multi-backend-prebuilt-docker-image-install/install-sunrise.md), [xpu](/getting_started/multi-backend-prebuilt-docker-image-install/install-xpu.md), [iluvatar](/getting_started/multi-backend-prebuilt-docker-image-install/install-iluvatar.md), and [tsingmicro](/getting_started/multi-backend-prebuilt-docker-image-install/install-tsingmicro.md).
- Added TLE support for the [amd](/getting_started/multi-backend-prebuilt-docker-image-install/install-amd.md) backend and added CI/CD.
- [rpu](/getting_started/multi-backend-prebuilt-docker-image-install/install-rpu.md) (Huixi Intelligence, Triton 3.6) is also supported. On the 3.3.x branch, [ARM64 CPU](/getting_started/install-arm64-cpu.md) provides [TLE-CPU](/user_guide/use-tle-cpu.md).
- [ARM64 CPU](/getting_started/flagtree-cpu.md) provides a CPU backend based on Triton 3.7.2, validated on Linux Arm64 with a vector-add kernel and FlagGems W4A8 operator tests.

- **DevTools (Debugger & Profiler)**
- FlagPrism ([flagos-ai/FlagPrism](https://github.com/flagos-ai/FlagPrism)) provides debugging and performance-analysis tools for Triton programs, containing `flagtree.debugger` and `flagtree.profiler`, and is integrated into FlagTree as the `third_party/FlagPrism` submodule. It initially supports a subset of backends: Huawei Ascend, Iluvatar, and Moore Threads.
59 changes: 59 additions & 0 deletions docs/flagtree_zh/getting_started/flagtree-cpu.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,59 @@
[<a href="../../../flagtree_en/getting_started/flagtree-cpu.html">英文版</a>|中文版]

# 在 Arm64 上启用 FlagTree CPU

[FlagTree CPU](https://github.com/flagos-ai/flagtree-cpu/tree/triton_v3.7.x) 为 Triton 提供 CPU 后端,使使用 Triton Python API 编写的 kernel 能够在 CPU 上编译和运行。它的编译器通过 MLIR/LLVM 将 Triton kernel 下降为原生 CPU 代码,运行时负责 CPU 目标选择、kernel 启动和编译缓存。本指南覆盖 Linux Arm64 上的 Triton 3.7.2 实现,并通过一个 vector-add kernel 和 FlagGems W4A8 算子测试进行验证。

使用 [flagos-ai/community](https://github.com/flagos-ai/community/tree/main/fep/sig-edge/scripts/flagtree-cpu37) 中的脚本构建并验证 Triton 3.7.2 CPU 后端。请在同一个 shell 中按以下步骤操作。

## 步骤 1 准备主机

使用 **Debian 13(Linux Arm64,`aarch64`)**,已安装 Git、具备 sudo 权限(或 root),并且至少有 **25 GiB 可用磁盘空间**。确保主机能够访问源码仓库和依赖下载服务。

## 步骤 2 获取脚本

```bash
git clone --single-branch --branch main \
https://github.com/flagos-ai/community.git community-flagtree
cd community-flagtree

RUN=fep/sig-edge/scripts/flagtree-cpu37/run.sh
export WORK_DIR="$HOME/flagtree-cpu-3.7-test"
export BUILD_JOBS=4
```

专用工作目录将编译器、Python 环境和缓存与其他安装隔离开来。请保留完整的脚本目录;`run.sh` 依赖其配套文件。

## 步骤 3 构建并安装

```bash
bash "$RUN" setup
```

setup 会安装系统依赖、创建 Python 3.11 环境、拉取固定版本的 FlagTree CPU 与 FlagGems 修订、初始化子模块,并构建 Triton 3.7.2。固定的 LLVM 与 SLEEF 版本保证构建可复现。安装后的包以 `triton` 名称导入。

## 步骤 4 验证 CPU 执行

```bash
bash "$RUN" test
```

测试会检查源码/工具链修订与 CPU 目标,分别使用全新缓存和复用缓存执行真实的 vector-add kernel,并将 FlagGems W4A8 结果与 PyTorch 对比。验收要求 **7 passed, 0 skipped**。日志与测试报告保存在 `$WORK_DIR/logs/` 下。

## 步骤 5 为你的应用启用该后端

```bash
source "$WORK_DIR/.venv/bin/activate"
export TRITON_CPU_BACKEND=1
export FLAGGEMS_VENDOR=arm
export TRITON_CACHE_DIR="$WORK_DIR/triton-runtime-cache"
mkdir -p "$TRITON_CACHE_DIR"

python your_program.py
```

将 `your_program.py` 替换为你的应用。`TRITON_CPU_BACKEND=1` 选择 CPU 执行;`FLAGGEMS_VENDOR=arm` 选择 FlagGems 的 Arm 实现。持久化缓存使与之兼容的 kernel 特化可以复用已编译的代码。这些变量必须在你的应用 shell 中导出,因为 setup/test 脚本运行在独立的 Bash 进程中。

通过上述检查可以验证编译器和已测算子;模型集成与模型冷启动性能需要单独验证。

参考:[FEP-0082](https://github.com/flagos-ai/community/blob/main/fep/sig-edge/0082-flagtree-cpu-bump-to-triton-3_7.md)。
1 change: 1 addition & 0 deletions docs/flagtree_zh/getting_started/getting-started.md
Original file line number Diff line number Diff line change
Expand Up @@ -8,6 +8,7 @@
requirements.md
install.md
install-arm64-cpu.md
flagtree-cpu.md
multi-backend-prebuilt-docker-image-install/install-nv.md
multi-backend-prebuilt-docker-image-install/install-tileir.md
multi-backend-prebuilt-docker-image-install/install-amd.md
Expand Down
1 change: 1 addition & 0 deletions docs/flagtree_zh/getting_started/install.md
Original file line number Diff line number Diff line change
Expand Up @@ -88,6 +88,7 @@
- [清微智能](/getting_started/multi-backend-prebuilt-docker-image-install/install-tsingmicro.md)
- [KLX](/getting_started/multi-backend-prebuilt-docker-image-install/install-xpu.md)
- [ARM64 CPU](/getting_started/install-arm64-cpu.md)
- [ARM64 CPU (Triton 3.7.2)](/getting_started/flagtree-cpu.md)

## 方式三:安装 wheel 包

Expand Down
1 change: 1 addition & 0 deletions docs/flagtree_zh/getting_started/requirements.md
Original file line number Diff line number Diff line change
Expand Up @@ -21,6 +21,7 @@
|[triton_v3.4.x](https://github.com/flagos-ai/flagtree/tree/triton_v3.4.x)|NVIDIA<br>AMD<br>曦望芯科|[nvidia](https://github.com/flagos-ai/FlagTree/tree/triton_v3.4.x/third_party/nvidia/)<br>[amd](https://github.com/flagos-ai/FlagTree/tree/triton_v3.4.x/third_party/amd/)<br>[sunrise](https://github.com/flagos-ai/FlagTree/tree/triton_v3.4.x/third_party/sunrise/)|3.4|
|[triton_v3.5.x](https://github.com/flagos-ai/flagtree/tree/triton_v3.5.x)|NVIDIA<br>AMD<br>燧原|[nvidia](https://github.com/flagos-ai/FlagTree/tree/triton_v3.5.x/third_party/nvidia/)<br>[amd](https://github.com/flagos-ai/FlagTree/tree/triton_v3.5.x/third_party/amd/)<br>[enflame](https://github.com/flagos-ai/FlagTree/tree/triton_v3.5.x/third_party/enflame/)|3.5|
|[triton_v3.6.x](https://github.com/flagos-ai/flagtree/tree/triton_v3.6.x)|NVIDIA<br>AMD<br>燧原<br>海光信息<br>摩尔线程<br>Thrive|[nvidia](https://github.com/flagos-ai/FlagTree/tree/triton_v3.6.x/third_party/nvidia/)<br>[amd](https://github.com/flagos-ai/FlagTree/tree/triton_v3.6.x/third_party/amd/)<br>[enflame](https://github.com/flagos-ai/FlagTree/tree/triton_v3.6.x/third_party/enflame/)<br>[hcu](https://github.com/flagos-ai/FlagTree/tree/triton_v3.6.x/third_party/hcu/)<br>[mthreads](https://github.com/flagos-ai/FlagTree/tree/triton_v3.6.x/third_party/mthreads/)<br>[damoacademy](https://github.com/flagos-ai/FlagTree/tree/triton_v3.6.x/third_party/thrive/)|3.6|
|[triton_v3.7.x](https://github.com/flagos-ai/flagtree/tree/triton_v3.7.x)|ARM64 cpu|[cpu](https://github.com/flagos-ai/FlagTree/tree/triton_v3.7.x/third_party/cpu/)|3.7|

## 不同分支上的功能特性

Expand Down
1 change: 1 addition & 0 deletions docs/flagtree_zh/release_notes/release_notes_v070.md
Original file line number Diff line number Diff line change
Expand Up @@ -22,6 +22,7 @@
- 将以下后端升级至 Triton 3.6 并新增 CI/CD:[sunrise](/getting_started/multi-backend-prebuilt-docker-image-install/install-sunrise.md)、[xpu](/getting_started/multi-backend-prebuilt-docker-image-install/install-xpu.md)、[iluvatar](/getting_started/multi-backend-prebuilt-docker-image-install/install-iluvatar.md) 和 [tsingmicro](/getting_started/multi-backend-prebuilt-docker-image-install/install-tsingmicro.md)。
- 为 [amd](/getting_started/multi-backend-prebuilt-docker-image-install/install-amd.md) 后端新增 TLE 支持并新增 CI/CD。
- 另外还支持 [rpu](/getting_started/multi-backend-prebuilt-docker-image-install/install-rpu.md)(辉羲智能,Triton 3.6)。在 3.3.x 分支上,[ARM64 CPU](/getting_started/install-arm64-cpu.md) 提供 [TLE-CPU](/user_guide/use-tle-cpu.md)。
- [ARM64 CPU](/getting_started/flagtree-cpu.md) 基于 Triton 3.7.2 提供 CPU 后端,已在 Linux Arm64 上通过 vector-add kernel 与 FlagGems W4A8 算子测试验证。

- **DevTools(调试器与性能分析器)**
- FlagPrism([flagos-ai/FlagPrism](https://github.com/flagos-ai/FlagPrism))为 Triton 程序提供调试与性能分析工具,包含 `flagtree.debugger` 与 `flagtree.profiler`,以 `third_party/FlagPrism` 子模块集成在 FlagTree 中。先支持部分后端:华为昇腾、天数智芯、摩尔线程。
Loading