From a417bad3f8a8b4ca095776fd644089b6c1f60309 Mon Sep 17 00:00:00 2001 From: cheng874 Date: Tue, 29 Sep 2026 15:58:47 +0800 Subject: [PATCH 1/3] 4th update of FlagQutumn --- .../FlagQuantum-overview.md | 51 ++++++ .../FlagQuantum_overview/architecture.md | 70 ++++++++ .../FlagQuantum_overview/features.md | 44 +++++ .../getting_started/getting-started.md | 39 +++- .../flagquantum_zh/getting_started/install.md | 74 ++++---- .../getting_started/requirements.md | 70 ++++---- docs/flagquantum_zh/index.md | 50 ++---- docs/flagquantum_zh/reference/capabilities.md | 63 +++++++ docs/flagquantum_zh/reference/reference.md | 83 +++++++-- .../release_notes/release-notes.md | 55 +++--- docs/flagquantum_zh/user_guide/algorithms.md | 105 +++++++---- docs/flagquantum_zh/user_guide/basic-usage.md | 52 ++++++ .../user_guide/circuits-and-ir.md | 71 ++++++++ .../user_guide/distributed-execution.md | 46 +++-- .../user_guide/hardware-and-remote.md | 91 ++++++++++ .../user_guide/measurement-and-noise.md | 86 +++++++++ docs/flagquantum_zh/user_guide/run-tests.md | 52 +++--- .../user_guide/simulation-representations.md | 56 ++++++ .../user_guide/training-with-pytorch.md | 74 ++++++++ docs/flagquantum_zh/user_guide/tutorials.md | 44 +++++ docs/flagquantum_zh/user_guide/user-guide.md | 170 +++++------------- 21 files changed, 1093 insertions(+), 353 deletions(-) create mode 100644 docs/flagquantum_zh/FlagQuantum_overview/FlagQuantum-overview.md create mode 100644 docs/flagquantum_zh/FlagQuantum_overview/architecture.md create mode 100644 docs/flagquantum_zh/FlagQuantum_overview/features.md create mode 100644 docs/flagquantum_zh/reference/capabilities.md create mode 100644 docs/flagquantum_zh/user_guide/basic-usage.md create mode 100644 docs/flagquantum_zh/user_guide/circuits-and-ir.md create mode 100644 docs/flagquantum_zh/user_guide/hardware-and-remote.md create mode 100644 docs/flagquantum_zh/user_guide/measurement-and-noise.md create mode 100644 docs/flagquantum_zh/user_guide/simulation-representations.md create mode 100644 docs/flagquantum_zh/user_guide/training-with-pytorch.md create mode 100644 docs/flagquantum_zh/user_guide/tutorials.md diff --git a/docs/flagquantum_zh/FlagQuantum_overview/FlagQuantum-overview.md b/docs/flagquantum_zh/FlagQuantum_overview/FlagQuantum-overview.md new file mode 100644 index 0000000000..8e7d7ada48 --- /dev/null +++ b/docs/flagquantum_zh/FlagQuantum_overview/FlagQuantum-overview.md @@ -0,0 +1,51 @@ +# FlagQuantum 概览 + +FlagQuantum 是构建在 PyTorch 之上的分片、可微量子模拟与训练框架。它把量子线路变成可训练模型,让同一份程序在多种模拟表示之间切换,并通过 FlagOS 触达国产加速器。FlagQuantum 是 FlagOS 生态的一部分——FlagOS 是一套统一的开源 AI 系统软件栈,通过无缝整合各类模型、系统与芯片来构建开放技术生态。 + +## 为什么需要 FlagQuantum? + +随着量子线路的比特数与深度增长,精确模拟的成本会变得难以为继;而把已经跑通的模型搬到真实硬件上,通常意味着重写。FlagQuantum 用同一套公开模型同时解决这两个问题: + +- 线路只构建一次,用标准的 PyTorch 自动微分与优化器训练; +- 按负载而不是按框架来选择模拟表示; +- 在本地、分布式与硬件执行目标之间迁移时,线路与所请求的可观测量始终保持显式。 + +## 一份程序,多个执行目标 + +```text +fq.Circuit / fq.Module + │ + ▼ + FlagQuantum IR + │ + ├── 编译与导出 + ├── 本地的态向量、MPS 与张量网络运行时 + ├── 位于 PyTorch 接口之后的可选 JAX 内核 + ├── 跨卡切分的态向量与 MPS 执行 + └── 面向提供方与硬件目标的部署包 +``` + +架构上的不变式很简单:后端选择可以改变执行方式,但不得改变程序的含义或结果契约。 + +## 适用人群 + +- **量子机器学习与 VQE 用户**:希望把可训练线路放进常规 PyTorch 训练循环。 +- **算法研究者**:需要在不同表示之间迁移,并在同一份程序上做比较。 +- **加速器用户**:需要一条国产加速器路径,并且要求证据显式、可审计,而不是隐藏回退。 + +## 关键概念 + +| 概念 | 含义 | +| --- | --- | +| `fq.Circuit` | 线路构建与面向线路的便捷方法,底层是 FlagQuantum IR | +| `fq.Module` | 可训练量子参数的 PyTorch 原生持有者 | +| `fq.plan` | 可解释的运行时规划:预期表示、策略与阻塞项 | +| `fq.run` | 唯一推荐的执行入口,返回 `fq.ExecutionResult` | +| `fq.train` | 由调用方掌握的最小 PyTorch 优化循环,返回 `fq.TrainingResult` | +| FlagQuantum IR | 编译、执行与部署共享的版本化算子、测量与元数据表示 | + +规划不是执行证据。规划描述意图与估计;运行时记录描述实际执行了什么。 + +## 支持边界 + +实现了某个特性,不等于在生产环境中支持它。FlagQuantum 会公开每个能力所处的等级——发布认证、生产支持、开发证据或实验性——见[能力参考](../reference/capabilities.md);已验证的工作流与仍属研究目标的部分都在那里列出,而不是靠示例去暗示。 diff --git a/docs/flagquantum_zh/FlagQuantum_overview/architecture.md b/docs/flagquantum_zh/FlagQuantum_overview/architecture.md new file mode 100644 index 0000000000..9be88b34d1 --- /dev/null +++ b/docs/flagquantum_zh/FlagQuantum_overview/architecture.md @@ -0,0 +1,70 @@ +# 架构 + +FlagQuantum 为量子 AI 程序在本地开发、加速内核、分布式模拟到部署之间提供同一套公开模型。它的组织方式保证后端变化不会改变程序含义。 + +## 核心层次 + +| 层次 | 职责 | 入口 | +| --- | --- | --- | +| 用户 API | 线路构建、PyTorch 模块、规划、执行、训练与部署 | `import flagquantum as fq` | +| 编译 | 变换线路并合法化目标输出,但不执行线路 | `fq.compile`、`flagquantum.compiler` | +| FlagQuantum IR | 版本化算子、测量、元数据、序列化与校验 | `fq.CircuitIR` | +| 规划 | 选择表示与执行策略,解释阻塞项与回退 | `fq.plan`、`Circuit.runtime_plan` | +| 运行时 | 本地或跨 rank 执行,并返回带类型的证据 | `fq.run`、`fq.ExecutionResult` | +| 训练 | 在各受支持运行时上保持 PyTorch 自动微分与优化器语义 | `fq.Module`、`fq.train` | +| 部署 | 绑定训练好的参数、面向目标编译并封装可审计的包 | `flagquantum.deployment.create_deployment_package` | + +## 源码结构 + +```text +flagquantum/ +├── _api.py # 根层 compile、plan、run 的组合 +├── circuit.py # 线路构建 +├── core/ # 后端中立的 IR 与共享语义 +├── compiler/ # 校验、优化、降级与代码生成 +├── runtime/ # 规划、执行生命周期、结果与协调 +├── simulation/ # 数值方法与内核 +├── noise/ # 后端中立的噪声模型与信道 +├── observables/ # 面向用户的测量构建 +├── qec/ # 纠错工作流与领域模型 +├── twin/ # 硬件数字孪生模型 +├── compute/ # 当前进程所控制的资源 +├── remote/ # 外部任务系统与结果获取 +├── ecosystem/ # 框架与格式适配器 +├── deployment/ # 已封装的、目标中立的执行包 +├── services/ # 可复用的多步应用工作流 +├── algorithms/ # 面向用户的算法组合 +├── benchmarking/ # 可复现的测量与证据生成 +├── drawer/ # 线路可视化 +└── experimental/ # 明确不稳定的 API +``` + +## 依赖方向 + +依赖指向内部,因此可选的集成始终位于必需的本地 PyTorch 路径之外: + +```text +用户门面 → 编译 / 运行时 / 应用工作流 → Core + │ + ├── Simulation 数值方法 + ├── Compute 本地资源适配器 + └── Remote 外部任务系统适配器 +``` + +Core 既不引入编排、也不引入数值引擎或厂商集成。编译器负责变换程序,但不执行程序;运行时负责组织执行,但不实现数值内核;Simulation 使用 Core 语义,但不选择资源;Compute 与 Remote 把硬件与外部系统的细节与其他领域隔离开来。 + +## 执行与训练契约 + +`fq.run` 是规范的执行入口,对受支持的本地与分布式模式返回 `fq.ExecutionResult`。专用的原生函数属于高级接口,可能暴露后端特有的对象。 + +`fq.train` 负责常规的 PyTorch 优化循环。按 rank 拥有的态向量与 MPS 训练有各自独立的分布式入口,它们并不会因为调用 `fq.train` 而自动生效。只有当分布式执行的前向、梯度、优化器更新与检查点归属都保持所声明的分布式语义时,分布式训练才算完成。 + +要主张分布式可扩展性,必须把同一份逻辑负载切分到多个 rank 上。复制的数据并行、rank 本地内核与手工张量切片会按各自的语义报告,绝不会被改称为容量扩展。 + +## 公开接口与内部接口 + +- 公开示例统一使用 `import flagquantum as fq`。 +- 稳定名称由受检清单圈定,并由测试验证。 +- `fq.experimental` 不提供兼容性保证。 +- 兼容模块用于迁移,它们不定义新的稳定 API。 +- 基准与研究工具永远不会成为运行时依赖。 diff --git a/docs/flagquantum_zh/FlagQuantum_overview/features.md b/docs/flagquantum_zh/FlagQuantum_overview/features.md new file mode 100644 index 0000000000..4dd6ffef26 --- /dev/null +++ b/docs/flagquantum_zh/FlagQuantum_overview/features.md @@ -0,0 +1,44 @@ +# 特性 + +## PyTorch 原生的量子训练 + +- **线路即 PyTorch 模块**:`fq.Module` 把量子模型暴露给 PyTorch,因此 `loss.backward()` 与常规优化器可以在同一个循环里训练量子层与经典层。 +- **平铺或命名参数分组**:既支持按位置书写的参数,也支持 `{"encoder": (4,), "readout": ()}` 这样的命名分组,并提供可复现的初始化与模块级随机种子。 +- **稳定的训练结果**:`fq.train` 返回带版本化摘要的 `fq.TrainingResult`,检查点与恢复由 `fq.Module` 负责。 + +## 多套模拟表示,一份程序 + +- **态向量**:CPU 或单卡 GPU 上的精确本地路径,小规模系统会自动给出稠密参考值。 +- **矩阵乘积态(MPS)**:面向低纠缠系统,可处理的比特数远超稠密态向量,并包含受限 TEBD。 +- **张量网络**:面向稠密态向量与 MPS 都不合适的线路结构,支持切片与收缩。 +- **可选 JAX 内核**:位于同一套 PyTorch 接口之后,跨接口支持一阶梯度。 +- **一条命令切换**:`python examples/quick_start.py --mode sv|mps|tn` 不改代码即可让同一混合模型跑在三种表示上。 + +## 编译与导出 + +- **与目标无关的优化**:`flagquantum.compiler.optimize` 迭代应用规范化重写直至收敛,并返回新的 IR。 +- **面向目标的编译**:`fq.compile(..., coupling_map=..., routing_strategy=...)` 只输出符合拓扑的双比特操作,并记录路由决策。 +- **编译器插件**:独立安装的包可以通过扩展注册表注册编译器,例如用于 Quafu 提交的 QSteed 插件。 +- **OpenQASM 与框架导出**:支持导出 OpenQASM 3.0,并提供 Qiskit、PennyLane、Cirq、Braket 与 CUDA-Q 适配器。 + +## 测量、噪声与纠错 + +- **可观测量与输出**:`fq.X`、`fq.Y`、`fq.Z` 用 `@` 组成泡利乘积、用普通算术组成哈密顿量;`fq.expectation`、`fq.probabilities`、`fq.samples`、`fq.counts` 请求结果。 +- **后端中立的噪声**:同一个 `NoiseModel` 驱动精确密度矩阵演化、批态向量轨迹与 MPS 量子轨迹,覆盖热弛豫、去极化、比特/相位翻转与读出错误。 +- **硬件泡利测量规划**:按量子位逐位对易分组,并为每组生成一个已封装的部署包。 +- **动态线路**:支持线路中测量与经典反馈,并提供只读的后端预检。 +- **量子纠错**:提供连接症状提取、译码与纠正的重复码存储实验。 + +## 分布式与加速器执行 + +- **分片态向量训练**:一份逻辑态向量跨 rank 切分,包含分布式前向、反向与优化器状态。 +- **按 rank 拥有的 MPS**:针对单卡放不下的负载,提供分布式 MPS 前向、反向与优化器状态。 +- **进程组下的 PyTorch 原生路径**:初始化多 rank 进程组后,同一模块与结果面会自动使用分片运行时。 +- **经 Torch-FL 触达 FlagOS 加速器**:逻辑设备 `flagos:0` 通过 Torch-FL 访问,厂商探测与派发由 Torch-FL 负责。 + +## 部署、互操作与生态 + +- **已封装的部署包**:提交前把训练好的参数绑定、面向目标编译,并生成可审计的包。 +- **QPU 数字孪生**:以标定数据为条件的设备模型,可预测测量分布,并附带明确的证据边界。 +- **算法单元**:Grover 搜索、振幅估计、量子 PCA、量子 k-medians、量子核估计与核岭分类、以 QUBO 形式表述的特征选择,以及 QUBO 到 Ising 的映射。 +- **线路可视化**:面向终端的文本绘制,以及用于论文插图的 Matplotlib 绘制。 diff --git a/docs/flagquantum_zh/getting_started/getting-started.md b/docs/flagquantum_zh/getting_started/getting-started.md index fdc06eae97..a3228017d4 100644 --- a/docs/flagquantum_zh/getting_started/getting-started.md +++ b/docs/flagquantum_zh/getting_started/getting-started.md @@ -1,11 +1,44 @@ -# FlagQuantum 快速入门 +# 快速入门 -本节介绍 FlagQuantum 的安装要求、安装过程,以及第一个可训练的量子模型。 +本节介绍安装 FlagQuantum 的环境要求,并引导你完成安装过程。 ```{toctree} :maxdepth: 2 +:hidden: requirements.md install.md -quick-start.md ``` + +## 你的第一个量子模型 + +FlagQuantum 是 PyTorch 优先的框架,因此最短的上手路径就是一个可训练的小线路。下面的示例构建一个双比特线路,通过最小化测量得到的期望值来学习其旋转角,最后打印训练后的测量结果。它不需要 GPU、凭据或任何可选后端。 + +```python +import torch +import flagquantum as fq + + +def circuit(parameters): + return fq.Circuit(2).ry(0, parameters[0]).cx(0, 1) + + +model = fq.Module(circuit, n_parameters=1, init=torch.tensor([0.25])) +training = fq.train( + model, + optimizer=torch.optim.Adam(model.parameters(), lr=0.05), + objective=lambda z: z.mean(), + steps=10, +) + +trained_circuit = circuit(next(model.parameters()).detach()) +measurement = fq.expectation(fq.Z(0)) +result = fq.run(trained_circuit, outputs=measurement) +print(result.expectation()) +``` + +接下来可以阅读: + +- [环境要求](requirements.md):受支持的 Python、PyTorch 与硬件平台,以及可选依赖分组; +- [安装 FlagQuantum](install.md):安装正式版本或开发版并验证; +- [用户指南](../user_guide/user-guide.md):线路、训练、模拟表示、分布式执行与硬件目标。 diff --git a/docs/flagquantum_zh/getting_started/install.md b/docs/flagquantum_zh/getting_started/install.md index 88055a9a4a..785c118a97 100644 --- a/docs/flagquantum_zh/getting_started/install.md +++ b/docs/flagquantum_zh/getting_started/install.md @@ -1,60 +1,56 @@ # 安装 FlagQuantum -开始前请先阅读[环境要求](requirements.md)。 +请先阅读[环境要求](requirements.md)。 -## 安装正式发布的包 +## 步骤 -```bash -python -m pip install flagquantum -``` +1. 安装 FlagQuantum -## 安装开发版本 + - 安装正式发布的包 -```bash -git clone https://github.com/flagos-ai/FlagQuantum.git -cd FlagQuantum -python -m pip install -e ".[dev]" -``` + ```{code-block} shell + python -m pip install flagquantum + ``` -## 验证安装 + - 安装带开发工具的开发版检出 -```bash -python -c "import flagquantum as fq; print(fq.__version__)" -``` + ```{code-block} shell + git clone https://github.com/flagos-ai/FlagQuantum.git + cd FlagQuantum + python -m pip install -e ".[dev]" + ``` -包通过 `import flagquantum as fq` 引入;导入它不会导入可选依赖、不会发现扩展, -也不会激活厂商适配器。 +2. 按需增加可选依赖 -## 运行一个本地示例 + ```{code-block} shell + python -m pip install -e ".[jax,cuda,viz]" + ``` -受维护的本地路径不需要凭据、不需要远程资源,也不需要可选后端: + 请使用[环境要求](requirements.md)中的分组名,而不是手工安装外部框架,这样才会沿用已锁定的兼容区间。 -```bash -python -m examples.local.simulate -python -m examples.local.measure -python -m examples.local.train -``` +3. 验证安装 -开发版安装还可以直接运行仓库中的示例,例如: + ```{code-block} python + import flagquantum as fq + print(fq.__version__) + ``` -```bash -python examples/quick_start.py --mode sv --steps 40 -``` +4. 验证本地执行路径 -## 可选依赖组 + ```{code-block} shell + python -m examples.local.simulate + python -m examples.local.measure + python -m examples.local.train + ``` -只在需要时安装对应能力: + 这三个示例覆盖本地态向量模拟、测量与 PyTorch 原生训练,不需要凭据、远程资源或可选后端。 -```bash -python -m pip install "flagquantum[jax]" -python -m pip install "flagquantum[viz]" -python -m pip install "flagquantum[qiskit]" -``` +## 开发容器 -互操作桥接把外部框架留在边界的另一侧:导入 FlagQuantum 永远不会导入 Qiskit、 -PennyLane、Cirq、CUDA-Q 或任何厂商 SDK。 +仓库同时提供面向 CPU 与 CUDA 环境的开发容器定义,并区分是否包含 JAX 与 QSteed 编译器。如果你更希望使用准备好的环境,请按仓库中的容器指南操作;镜像从本地检出安装 FlagQuantum 并内置 JupyterLab,且都不包含提供方凭据。 ## 下一步 -继续阅读[快速开始](quick-start.md),或直接查看[模拟模式](../user_guide/simulation-modes.md) -以选择表示形式,以及[远程执行](../user_guide/remote-execution.md)了解厂商目标。 +- [构建并运行你的第一个程序](../user_guide/basic-usage.md) +- [使用 PyTorch 训练](../user_guide/training-with-pytorch.md) +- [选择模拟表示](../user_guide/simulation-representations.md) diff --git a/docs/flagquantum_zh/getting_started/requirements.md b/docs/flagquantum_zh/getting_started/requirements.md index 5212aaa727..ccdfe9ebe6 100644 --- a/docs/flagquantum_zh/getting_started/requirements.md +++ b/docs/flagquantum_zh/getting_started/requirements.md @@ -1,48 +1,44 @@ # 环境要求 -FlagQuantum 支持 Python 3.10 至 3.12,并要求 PyTorch 2.5 或更高版本。正式发布的 -包只依赖 PyTorch,其余能力都是可选扩展。 +本节介绍 FlagQuantum 的硬件平台与软件要求。 ## 软件要求 -| 项目 | 要求 | -| --- | --- | -| Python | 3.10、3.11 或 3.12(`>=3.10,<3.13`) | -| PyTorch | `>=2.5,<2.14` | -| 操作系统 | Linux、macOS 或 Windows,可在 CPU 或单张加速卡上运行 | +- Python 3.10、3.11 或 3.12 +- PyTorch 2.5 及以上,低于 2.14 -可选扩展在 `pyproject.toml` 中声明,用 `pip install "flagquantum[]"` 安装: +正式发布的包只依赖 PyTorch,其余都是可选项。 -| 扩展 | 新增内容 | -| --- | --- | -| `dev` | pytest、覆盖率、xdist、ruff、black、mypy、pre-commit | -| `jax` | 运行在 PyTorch 接口之后的 JAX 内核 | -| `cotengra` | 张量网络收缩路径搜索 | -| `cuda` | 受支持的本地门路径所用 Triton 内核 | -| `viz` | 线路绘制所需的 Matplotlib | -| `qiskit` | Qiskit 与 Aer 的转换与执行桥接 | -| `pennylane` | PennyLane 的转换与 Lightning 执行桥接 | -| `cirq` | Cirq 的转换与模拟器执行桥接 | -| `cudaq` | CUDA-Q 内核导出(仅 Linux) | -| `braket`、`azure`、`quafu` | 其他厂商适配器 | -| `examples` | 较大示例所用的数据集与 transformer 辅助依赖 | -| `all` | 开发、绘图、JAX、Triton 与示例依赖 | - -## 硬件平台 - -| 平台 | 状态 | +## 支持的硬件平台 + +| 平台 | 触达方式 | | --- | --- | -| CPU | 默认本地路径;精确态向量模拟与训练 | -| 单张 CUDA GPU | 在 `ExecutionOptions` 中显式选择的单设备执行 | -| 多 rank | 基于 Gloo 或 NCCL 的跨卡态向量执行与按 rank 拥有的 MPS 执行 | -| FlagOS 加速器 | 通过 FlagOS 统一多芯片层及其逻辑设备 `flagos:0` 接入 | +| CPU | 默认的本地态向量、MPS 与张量网络执行 | +| NVIDIA GPU | 通过 `fq.ExecutionOptions(device="cuda:0")` 选择的 CUDA 设备 | +| FlagOS 支持的加速器 | 逻辑设备 `flagos:0`,经 Torch-FL 触达 | -FlagQuantum 不做厂商探测,也不含厂商分支。物理设备探测、厂商运行时以及 -`flagos:0` 到物理卡的映射,都属于 FlagOS 厂商集成;FlagQuantum 只记录它拿到 -的运行时标识与路由证据。因此,仅存在一条集成路径并不等于某款国产加速器已获得认证。 +FlagQuantum 不会按设备名去探测或派发某个具体国产加速器。厂商探测、运行时激活与兼容路由都由 Torch-FL 负责,并通过 `flagos` 契约暴露出来。 -## 多机环境预期 +## 可选依赖分组 -分布式支持依赖具体环境。在已复核的配置中,双机 A800、每机一张卡的部署可以基于 -NCCL 与 TCP 运行前向、梯度与训练/恢复负载;多机的发布认证是另一个需要证据支撑的 -独立步骤。 +| 分组 | 增加的内容 | +| --- | --- | +| `dev` | pytest、覆盖率、xdist、ruff、black、mypy、build、pre-commit | +| `jax` | JAX 支撑的内核 | +| `cuda` | Triton | +| `qiskit` | Qiskit 与 Aer 互操作 | +| `pennylane` | PennyLane 互操作 | +| `cirq` | Cirq 互操作 | +| `braket` | Amazon Braket 互操作 | +| `cudaq` | CUDA-Q 内核导出 | +| `quafu` | Quafu 提交支持 | +| `azure` | Azure Quantum 提交支持 | +| `viz` | Matplotlib 线路绘制 | +| `examples` | 示例工作流所需的 `datasets` 与 `transformers` | +| `all` | 以上全部可选依赖 | + +互操作适配器属于可选的控制面边界:`import flagquantum` 永远不会连带导入 Qiskit、PennyLane、Cirq 或其他外部框架。 + +## 选择目标前应了解的支持边界 + +每个能力的成熟度,以及各目标实际执行过什么,都公布在[能力参考](../reference/capabilities.md)中。实现了某个 API,并不等于自动获得生产支持。 diff --git a/docs/flagquantum_zh/index.md b/docs/flagquantum_zh/index.md index 87e71cac8a..6ffe0a0727 100644 --- a/docs/flagquantum_zh/index.md +++ b/docs/flagquantum_zh/index.md @@ -11,41 +11,41 @@ ::::{grid} 1 2 2 3 :gutter: 1 1 1 2 -::{grid-item-card} {octicon}`browser;1.5em;sd-mr-1` 概览 -:link: overview/overview +:::{grid-item-card} {octicon}`browser;1.5em;sd-mr-1` 概览 +:link: FlagQuantum_overview/FlagQuantum-overview :link-type: doc -FlagQuantum 是什么、如何组织,以及能力等级如何划分。 +FlagQuantum 是什么,以及 PyTorch 优先的量子工作流背后的基本概念。 +++ -[了解更多 »](overview/overview.md) +[了解更多 »](FlagQuantum_overview/FlagQuantum-overview.md) ::: -::{grid-item-card} {octicon}`book;1.5em;sd-mr-1` 快速入门 +:::{grid-item-card} {octicon}`book;1.5em;sd-mr-1` 快速入门 :link: getting_started/getting-started :link-type: doc -环境要求、安装步骤,以及第一个可训练的量子模型。 +查看环境要求,逐步安装 FlagQuantum。 +++ [了解更多 »](getting_started/getting-started.md) ::: -::{grid-item-card} {octicon}`broadcast;1.5em;sd-mr-1` 用户指南 +:::{grid-item-card} {octicon}`broadcast;1.5em;sd-mr-1` 用户指南 :link: user_guide/user-guide :link-type: doc -线路、规划、训练、模拟表示、噪声、硬件目标、数字孪生与纠错。 +构建、训练、模拟、分布式运行并部署量子程序。 +++ [了解更多 »](user_guide/user-guide.md) ::: -::{grid-item-card} {octicon}`bookmark;1.5em;sd-mr-1` 参考资料 +:::{grid-item-card} {octicon}`list-unordered;1.5em;sd-mr-1` 参考资料 :link: reference/reference :link-type: doc -稳定 API 清单、能力目录、扩展 SDK 与当前支持边界。 +稳定 API 清单、能力等级与支持边界。 +++ [了解更多 »](reference/reference.md) @@ -68,9 +68,9 @@ release_notes/release-notes.md :maxdepth: 2 :hidden: -overview/overview.md -overview/features.md -overview/architecture.md +FlagQuantum_overview/FlagQuantum-overview.md +FlagQuantum_overview/features.md +FlagQuantum_overview/architecture.md ``` ```{toctree} @@ -81,7 +81,6 @@ overview/architecture.md getting_started/getting-started.md getting_started/requirements.md getting_started/install.md -getting_started/quick-start.md ``` ```{toctree} @@ -90,24 +89,15 @@ getting_started/quick-start.md :hidden: user_guide/user-guide.md -user_guide/first-quantum-model.md -user_guide/build-and-run.md -user_guide/runtime-planning.md -user_guide/local-workflows.md -user_guide/choose-a-simulator.md +user_guide/basic-usage.md +user_guide/circuits-and-ir.md user_guide/training-with-pytorch.md -user_guide/custom-operations.md -user_guide/compile-and-target.md -user_guide/compiler-and-remote.md -user_guide/noisy-simulation.md -user_guide/digital-twin-and-qec.md -user_guide/qpu-digital-twin.md -user_guide/qec.md -user_guide/run-on-hardware.md +user_guide/measurement-and-noise.md +user_guide/simulation-representations.md user_guide/distributed-execution.md -user_guide/dynamic-circuits.md +user_guide/hardware-and-remote.md user_guide/algorithms.md -user_guide/examples-and-tutorials.md +user_guide/tutorials.md user_guide/run-tests.md ``` @@ -117,7 +107,5 @@ user_guide/run-tests.md :hidden: reference/reference.md -reference/api.md reference/capabilities.md -reference/extensions.md ``` diff --git a/docs/flagquantum_zh/reference/capabilities.md b/docs/flagquantum_zh/reference/capabilities.md new file mode 100644 index 0000000000..585bbabc85 --- /dev/null +++ b/docs/flagquantum_zh/reference/capabilities.md @@ -0,0 +1,63 @@ +# 能力参考 + +FlagQuantum 公开每个能力的成熟度,而不是靠示例去暗示。本页是摘要;仓库中经机器校验的能力矩阵才是权威来源。 + +## 如何理解成熟度 + +| 等级 | 含义 | +| --- | --- | +| 发布认证 | 有受发布门禁约束、经审计且可复现的证据,且无未解决的发布阻塞项 | +| 生产支持 | 具备兼容性、运维指引与目标硬件证据的受支持路径 | +| 开发证据 | 可执行且经过测试的开发结果,不构成生产或通用可扩展性主张 | +| 实验性 | 研究性接口,不提供兼容性或生产保证 | + +成熟度只对每个能力所声明的范围成立。本地、复制、切片或仅停留在规划阶段的执行路径都不构成分布式可扩展性证据。稳定的公开 API 也不会提升某个实验性后端的等级。 + +## 构建、模拟与训练 + +| 能力 | 成熟度 | 边界简述 | +| --- | --- | --- | +| 统一线路 API 与 FlagQuantum IR | 发布认证 | IR v1;不兼容的 schema 变更需要显式迁移 | +| 本地态向量模拟与训练 | 生产支持 | 容量受单设备限制 | +| 分片态向量训练 | 生产支持 | 多机发布认证取决于已提升的审计硬件证据 | +| 可微与分片 MPS 训练 | 开发证据 | 已有单机与双机证据;层并行收缩与容量长跑仍不完整 | +| 张量网络执行与训练 | 实验性 | 通用反向收缩与生产级分布式传输尚未认证 | +| 受限本地 MPS TEBD | 实验性 | 仅支持静态实系数的一体与相邻二体泡利项、批大小为 1、二阶虚时间 | +| 精确与轨迹式含噪模拟 | 实验性 | 脉冲重叠、串扰、泄漏、提供方标定适配器与含噪梯度均不支持 | +| 连续时间 Lindblad 演化 | 生产支持 | 仅限时不变稠密哈密顿量的 CPU complex64/complex128 | +| Double-Single FP32 数值原语 | 实验性 | 面向 FP32 硬件的软件扩展精度;不等价于 FP64 | +| 分离实部/虚部 FP32 态向量(P0–P5) | 实验性 | 默认运行时永不选择的显式前向、期望值与 Double-Single 路径 | +| 量子态制备 | 实验性 | 演示规模;输入本身已是指数级大小 | +| Oracle 构建块与真值表合成 | 实验性 | 可逆经典逻辑;辅助位必须初始化为零态 | +| Grover 搜索与振幅估计 | 实验性 | 优势体现在查询复杂度,且前提是 oracle 或态制备免费 | +| 量子 PCA、k-medians、核估计 | 实验性 | 优势前提是数据访问模型,而这些单元并不提供 | +| 特征选择与 QUBO 到 Ising 映射 | 实验性 | 多项式级的经典变换;本身没有求解器,也没有优势 | +| 相位估计求奇异值 | 实验性 | 演示规模;经典计算部分是显式付出的 | + +## 分布式与 FlagOS 执行 + +| 能力 | 成熟度 | 边界简述 | +| --- | --- | --- | +| FlagOS 本地态向量 CUDA 参考 | 开发证据 | 基于 CUDA 的参考路径,未认证任何国产加速器 | +| FlagOS 分布式态向量负载 | 开发证据 | 单台 A800 节点上 2、4、8 卡;多机行为尚未确立 | +| FlagOS 态向量容量扩展 | 开发证据 | 八卡 A800 节点上一个确切的 32 比特 complex128 前向负载 | +| FlagOS 分布式传输可观测性 | 开发证据 | 2、4、8 rank 的集合通信;设备活动捕获不完整 | +| 国产单卡认证试验台 | 开发证据 | 采集由供方独占证明的候选结果,本身不提升任何能力等级 | + +## 部署、互操作与研究性接口 + +| 能力 | 成熟度 | 边界简述 | +| --- | --- | --- | +| 线路打包与云端部署 | 开发证据 | 提供方支持与凭据行为各不相同;没有提供方获得发布认证 | +| 带证据约束的 QPU 数字孪生 | 开发证据 | 一致性指测量分布的全变差一致性,且只对所声明的线路成立 | +| 互操作适配器契约 | 实验性 | 协议处于候选稳定级,尚待 API 负责人批准 | +| Qiskit、PennyLane、Cirq 与 CUDA-Q 互操作 | 实验性 | 在版本化边界上做静态转换;外部对象不会进入运行时或加速层 | +| PennyLane Lightning、Cirq Simulator 与 Qiskit Aer 桥接 | 实验性 | 仅支持一个完全绑定、单批次的线路,不支持梯度、噪声、动态线路、路由或回退 | +| 基于证据的模拟器顾问 | 实验性 | 建议绑定到已入库的线路与环境,绝不按规模推断性能 | +| 动态线路与后端评估 | 实验性 | 本地动态噪声仅限一比特比特翻转信道与读出混淆 | +| 扩展 SDK | 实验性 | 协议已批准但未冻结;编译器插件只交换 `CircuitIR` | +| 重复码存储实验 | 开发证据 | 仅一个固定的三数据比特配置;不主张逻辑抑制或阈值结论 | + +## 已公开的性能主张 + +FlagQuantum 公布实测性能主张时,会同时给出原始产物、其摘要、记录环境与确切范围,并继承所属能力的成熟度。没有经过审计的产物时,本页不做任何性能主张。 diff --git a/docs/flagquantum_zh/reference/reference.md b/docs/flagquantum_zh/reference/reference.md index 6341a6e9e9..9aceb21bf1 100644 --- a/docs/flagquantum_zh/reference/reference.md +++ b/docs/flagquantum_zh/reference/reference.md @@ -1,28 +1,79 @@ # 参考资料 -稳定接口、能力成熟度、运行时契约与当前支持边界。 +稳定接口、配置与支持边界。 -## 稳定 API 与接口 +## 稳定 API -- [接口参考](api.md) —— 经过整理的 `import flagquantum as fq` 接口面,附可运行示例。 -- [能力目录](capabilities.md) —— 哪些属于发布认证、生产可用、开发证据或实验性。 +FlagQuantum 只暴露一套经过整理的 Python 接口,即 `import flagquantum as fq`。稳定名称由受检清单圈定,并由可执行的契约测试验证。 -## 契约与策略 +| API | 稳定性 | +| --- | --- | +| `fq.Circuit`、`fq.CircuitIR`、`fq.Instruction` | 稳定 | +| `fq.ExecutionOptions`、`fq.ExecutionPlan`、`fq.ExecutionResult` | 稳定 | +| `fq.MeasurementResult`、`fq.Observable`、`fq.OutputRequest` | 稳定 | +| `fq.Parameter`、`fq.ParameterExpression`、`fq.RuntimePolicy` | 稳定 | +| `fq.Module`、`fq.TrainingResult` | 稳定 | +| `fq.I`、`fq.X`、`fq.Y`、`fq.Z` | 稳定 | +| `fq.IR_VERSION`、`fq.IRSerializationError`、`fq.IRValidationError` | 稳定 | +| `fq.compile`、`fq.plan`、`fq.run`、`fq.train` | 稳定 | +| `fq.counts`、`fq.expectation`、`fq.probabilities`、`fq.samples` | 稳定 | +| `fq.submit`、`fq.restore_job` | 稳定 | +| `fq.experimental`、`fq.twin`、`fq.__version__` | 稳定 | + +## API 对照 + +| 任务 | 主要接口 | 结果 | +| --- | --- | --- | +| 构建程序 | `fq.Circuit` | 由 FlagQuantum IR 支撑的线路 | +| 优化程序 | `flagquantum.compiler.optimize` | `fq.CircuitIR` | +| 面向所选工具与目标编译 | `fq.compile` | `fq.CircuitIR` | +| 查看执行决策 | `fq.plan`、`Circuit.runtime_plan` | 可解释的运行时规划 | +| 本地或远程执行 | `fq.run` | `fq.ExecutionResult` | +| 定义可训练量子层 | `fq.Module` | PyTorch 模块 | +| 训练 | `fq.train` | `fq.TrainingResult` | +| 面向目标打包 | `flagquantum.deployment.create_deployment_package` | 已封装的部署包 | + +## 运行时配置 + +`RuntimeConfig` 是 FlagQuantum 的不可变执行策略:后端、设备、实数与复数精度、JAX 精度、矩阵乘法策略与绘制风格。`Circuit` 在构造时捕获配置,并把版本化清单嵌入 IR 与执行规划中,因此分布式工作进程可以重建同一套策略。 + +```{code-block} python +import flagquantum as fq +from flagquantum.runtime.configuration import RuntimeConfig, runtime_config + +config = RuntimeConfig(device="cuda") +circuit = fq.Circuit(4, config=config) + +with runtime_config(complex_dtype="complex128"): + circuit = fq.Circuit(2) # 捕获 complex128 +``` + +每次执行只解析一次复数精度:`complex64` 意味着 float32 参数与实部组件,`complex128` 意味着 float64。解析结果支配规划字节数、态分配、参数张量、门矩阵与分布式模拟,这些阶段不会各自去查询进程默认值。dtype 冲突会在应用任何门之前失败。 + +## 算子 + +线路操作与后端降级能力来自同一份带类型的算子注册表。注册了某个后端的降级,只意味着该算子可以为这个后端降级——它不是发布证据。仓库中生成的算子表是注册集合的权威来源,覆盖泡利门、Clifford 门、旋转门、受控门、对称双比特门与相位门,以及稳定噪声模型所用的噪声信道。 + +## 错误 -| 主题 | 定义位置 | +| 类别 | 含义 | | --- | --- | -| 稳定名称 | 受校验的 API 清单,在接口参考中呈现 | -| 算子与后端下沉 | 单一类型化算子注册表,以自动生成的算子表呈现 | -| 运行时证据字段 | 带版本的运行时契约;这套词汇是契约,不是运行发生过的证明 | -| 实测基准证据 | 带来源信息的审计产物;计划与夹具不算证据 | -| 支持的 Python 与依赖 | 包元数据:Python 3.10–3.12,PyTorch 2.5–2.13 | +| `ValidationError` | 语义输入非法 | +| `PlanningError` | 规划过期、被篡改或不兼容 | +| `CapabilityError` | 请求的能力不可用 | +| `ExecutionError` | 执行或训练失败 | -## 结论边界 +所有类别都继承 `FlagQuantumError` 以及一个兼容的 Python 内置异常。类型错误与未知关键字参数抛出 `TypeError`。 -本地执行、已注册的下沉支持、分布式开发探针与发布级可扩展性是彼此独立的结论。注册了下沉只证明某算子可在某路径上执行,并不证明该路径在生产规模上可用。分布式与加速卡结论还额外要求由运行时生成的证据,并通过仓库的基准审计与发布策略。 +## 稳定性边界 -规划中的能力属于路线图文档;在可执行清单与测试出现之前,不会列为受支持。 +- `fq.experimental` 不提供兼容性保证。 +- 兼容导入属于迁移辅助,并不意味着稳定。 +- 规划结果描述意图与估计,永远不是运行时或基准证据。 +- 运行时证据必须满足带类型、带版本的运行时契约。 +- 新的公开名称必须首先可导入、通过快照测试,并被加入稳定 API 清单。 -## 上游参考 +## 延伸阅读 -FlagQuantum 仓库保有完整的参考集:接口参考、运行时架构与配置参考、已知限制目录、测试手册与路线图。本页发布的取值来自该权威来源;当支持边界变化时,上游参考与本文档集同步更新。 +- [能力参考](capabilities.md):各能力对应的成熟度与支持边界。 +- [架构](../FlagQuantum_overview/architecture.md):层次划分、依赖方向与各项契约。 diff --git a/docs/flagquantum_zh/release_notes/release-notes.md b/docs/flagquantum_zh/release_notes/release-notes.md index 3e62009028..b760d3f11e 100644 --- a/docs/flagquantum_zh/release_notes/release-notes.md +++ b/docs/flagquantum_zh/release_notes/release-notes.md @@ -1,41 +1,50 @@ # 发布说明 -发布说明记录用户可见的行为变化与支持边界变化。性能数字必须有审计产物作为依据, -不能由本页推断。 - ## v0.2.0 **发布日期**:2026-09-11 - **新增特性** - - **统一程序模型**——`fq.Circuit`、`fq.Module`、`fq.run`、`fq.train` 与 `fq.plan` 构成一个覆盖带版本 FlagQuantum IR 的公开接口,具备可查看的执行计划、带类型的结果与失败即拒的校验。 - - **每次运行可选表示**——本地态向量、MPS 与张量网络执行都从同一个程序出发选择,可选 JAX 内核运行在 PyTorch 接口之后。 - - **分布式执行**——按 rank 拥有的态向量与 MPS 前向、梯度与训练路径把同一个逻辑负载切分到多个 rank,并显式声明分布式语义。 - - **噪声与精度选项**——一个与后端无关的噪声模型同时驱动精确密度矩阵演化、批量态向量轨迹与 MPS 轨迹;在仅有 FP32 的设备上可用 Double-Single 软件扩展精度。 - - **硬件执行路径**——显式选择编译器与厂商,完成编译、打包、提交,并通过同一个 `fq.ExecutionResult` 返回结果,其中包含面向采样硬件的分组泡利测量。 - - **生态适配器**——与框架无关的 PennyLane、Qiskit、Cirq、CUDA-Q 互转,并提供本地 PennyLane Lightning、Qiskit Aer 与 Cirq 模拟器的显式执行桥接。 - - **部署与扩展接口**——密封部署包把训练好的参数绑定到目标;扩展 SDK 支持后端、内核、算子、编译 pass、设备、厂商、测量收集器与规划器插件,并具备能力协商。 - - **证据纪律**——每项能力都有成熟度等级,对外性能结论必须有带摘要与环境记录的入库产物。 + - 统一线路 API:`fq.Circuit` 采用 `n_qubits` 写法,支持语义化量子比特关键字,并兼容旧的别名。 + + - FlagQuantum IR:一份版本化、可序列化并经过校验的表示,由编译、执行与部署共享。 + + - 运行时规划:`fq.plan` 与 `Circuit.runtime_plan` 在执行之前解释所选的表示、执行策略与阻塞项。 + + - 唯一执行入口:`fq.run` 返回 `fq.ExecutionResult`,提供稳定的值、态、采样、规划、精度、指标、溯源、运行时与兼容性接口面,并对未知关键字失败即拒绝。 + + - PyTorch 原生训练:`fq.Module` 返回可自动微分的张量,`fq.train` 运行由调用方掌握的优化循环并返回 `fq.TrainingResult`,模块检查点支持保存与恢复。 + + - 一份程序背后的三套模拟表示:态向量、矩阵乘积态与张量网络,并可在同一套 PyTorch 接口之后使用可选的 JAX 内核。 -- **变更** + - 编译:用 `compiler.optimize` 做与目标无关的优化,用显式耦合图做面向目标的编译并记录路由决策,编译器插件通过扩展注册表发现。 - - 稳定根 API 由一个受校验的清单限定;实验性与兼容性导入不再静默扩张受支持的接口面。 - - 运行时配置变为不可变、任务作用域,并贯穿编译、规划与执行边界,而不再依赖进程全局状态。 + - 测量与噪声:泡利可观测量配合 `expectation`、`probabilities`、`samples`、`counts` 输出;同一个后端中立的 `NoiseModel` 驱动精确密度矩阵演化、批态向量轨迹与 MPS 量子轨迹。 -- **移除** + - 分布式执行:分片态向量的前向、反向、训练与优化器状态,以及按 rank 拥有的分布式 MPS 训练。 - - 预发布 v0.1 的面向设备接口——`DistributedQuantumDevice`、`GeneralEncoder`、可逆酉辅助、DTensor 交换辅助,以及面向门的设备接口与面向设备的测量路径——不属于 v0.2 产品,已在不提供兼容层的情况下移除。 + - 硬件执行路径:通过显式的 `compiler` 与 `target` 选择完成远程提交、有序的逻辑到物理映射、哈密顿量测量的量子位逐位对易分组,以及已封装的部署包。 + + - 生态与研究接口:Qiskit、PennyLane、Cirq、Braket 与 CUDA-Q 的互操作适配器;QPU 数字孪生;重复码存储实验;线路绘制;以及算法包。 + + - `flagquantum.errors` 中的稳定错误类别,并与对应的 Python 内置异常保持兼容。 + +- **移除与替换** + + - 预发布阶段面向设备的 API——`DistributedQuantumDevice`、`GeneralEncoder`、可逆酉模式、DTensor 交换辅助工具,以及面向设备的门与测量——都不属于 v0.2.0 的产品范围,已在不提供兼容层的情况下移除。 ## v0.1.0 **发布日期**:2026-06-24 -- **新增特性** +- FlagQuantum 首个版本,作为构建在 PyTorch 之上的分布式量子态向量模拟器发布。 +- 使用 `DTensor` 进行多卡分布式态向量模拟,并在门操作期间自动重分片。 +- 支持泡利门、Clifford 门、旋转门与受控门及参数化能力,并支持自定义门注册。 +- 支持角度编码、振幅编码与基编码,以及用户自定义的通用编码器。 +- 通过可逆反向传播实现内存高效的梯度计算,并提供测量后选择与去极化噪声模型。 +- 支持文本与 Matplotlib 线路绘制,以及 OpenQASM 2.0/3.0 导出。 + +--- - - FlagQuantum 首个版本:基于 PyTorch 的分布式量子态向量模拟器。 - - 多 GPU 态向量模拟,门操作期间自动重分片。 - - 覆盖泡利、Clifford、旋转与受控门的门集合,并支持可训练参数。 - - 内存高效的可逆反向传播与自定义门注册。 - - 后选择与去极化噪声模型,以及角度、振幅与基编码。 - - 文本与 Matplotlib 两种线路可视化模式,以及面向真实硬件平台的 OpenQASM 导出。 +仓库中的开发版本还额外提供 QPU 数字孪生研究、带回执的非阻塞远程任务提交、动态线路以及更多模拟表示。如需使用请从源码安装。 diff --git a/docs/flagquantum_zh/user_guide/algorithms.md b/docs/flagquantum_zh/user_guide/algorithms.md index 52b1e09549..879dd73f97 100644 --- a/docs/flagquantum_zh/user_guide/algorithms.md +++ b/docs/flagquantum_zh/user_guide/algorithms.md @@ -1,36 +1,75 @@ -# 算法 - -FlagQuantum 以可直接运行的示例形式提供算法单元,规模为演示级别。每个单元都会说明 -自己的优势前提,而诚实的解读通常是:量子例程所要求的输入模型,本单元并未提供。 - -## 单元列表 - -| 单元 | 接口 | 优势前提 | -| --- | --- | --- | -| QUBO 到 Ising 映射 | `flagquantum.algorithms` | 只是多项式级的经典变换,本身没有优势;任何收益都属于消费该哈密顿量的求解器 | -| 量子态制备 | `state_preparation` | 在已知振幅的前提下高效制备态;其经典输入本身已是指数规模 | -| 预言机构建块 | `oracle` | O(n) 个 Toffoli 级别的可逆经典逻辑,本身没有优势前提 | -| 真值表预言机合成 | `oracle` | 用经典方式枚举真值表,成本随寄存器宽度指数增长 | -| Grover 搜索 | `grover` | 改进体现在对合成预言机的查询复杂度上,而非端到端 | -| 振幅估计 | `amplitude_estimation` | 仅当态制备酉算子免费时才有二次加速;本单元不提供 QRAM | -| 量子主成分分析 | `pca` | 密度矩阵在经典侧物化;仅低有效秩时才有意义 | -| 量子 k-medians | `kmedians` | 优势前提是免费的 Grover 预言机,而距离表是经典计算的 | -| 量子核估计 | `quantum_kernel` | 未满足数据访问模型前提,且每个核矩阵元素都是采样估计 | -| 特征选择作为 QUBO | `feature_selection` | 只构建并求值目标函数,不做求解;仓库中没有退火机 | -| 频繁项比例 | `qarm` | 假设存在连贯的逐项数据库访问,而这里的访问代价是实付的 | -| 相位估计求奇异值 | `svd` | 输入态由读数所要估计的那次分解本身制备而来 | - -## 运行某个单元 - -每个单元都从各自的模块导出,而不是从根命名空间导出,因此示例会直接导入它。可运行 -入口位于 `examples/algorithms/`,每个示例都会打印它与经典参考值对比的结果。 - -```bash -python examples/algorithms/grover.py +# 算法、纠错与数字孪生 + +## 算法单元 + +`flagquantum.algorithms` 在稳定的线路与运行时 API 之上组合面向用户的算法单元。它们从子包接口导出,而不是通过 `fq` 别名导出。 + +| 单元 | 作用 | +| --- | --- | +| Grover 搜索 | 由真值表合成 oracle 并执行搜索 | +| 振幅估计 | 在寄存器自身的网格上估计振幅 | +| 量子 PCA | 构建数据集的密度矩阵并读出其主要特征值 | +| 量子 k-medians | 通过 Grover 搜索采样质心归属 | +| 量子核估计 | 估计核矩阵元素并训练核岭分类器 | +| 以 QUBO 形式表述的特征选择 | 构建特征选择实例的目标函数 | +| QUBO 到 Ising 映射 | 在 QUBO 与 Ising 哈密顿量之间双向转换 | +| 哈密顿量辅助工具 | 泡利项、精确基态参考与 VQE 辅助 | + +这些单元都是演示规模。每个单元都会显式记录其优势前提:其中若干需要 qRAM 或免费的 oracle,而单元本身并不提供,构建问题的经典代价是实际付出的,而不是被假定消失。引用某个单元的加速比之前,请先阅读它的优势前提。VQE 与 ADAPT-VQE 单元的优化器选择使用 `optimizer_factory` 协议,默认是 Adam。 + +## 本地哈密顿量梯度 + +对于批大小为 1、系数为实常数的 Z 与 ZZ 项态向量线路,`Hamiltonian.expectation` 提供了一条内存受限的伴随路径: + +```{code-block} python +import torch +import flagquantum as fq +from flagquantum import algorithms as fqa + +theta = torch.tensor(0.2, dtype=torch.float64, requires_grad=True) +circuit = fq.Circuit(3, dtype=torch.complex128).ry(0, theta).cx(0, 1) +hamiltonian = fqa.Hamiltonian(( + fqa.pauli_term(0.7, "ZZ", (0, 1)), + fqa.pauli_term(0.2, "Z", (2,)), +)) + +energy = hamiltonian.expectation(circuit, differentiation="adjoint") +energy.backward() +``` + +默认仍是 `differentiation="autograd"`。伴随模式会拒绝 X/Y 项、可训练或复数系数以及批量线路,而不是悄悄切换算法。 + +## 量子纠错 + +`flagquantum.qec` 连接症状提取、译码、纠正与逻辑结果分析。参考实验是一个注入错误的三数据比特重复码存储实验: + +```{code-block} python +from flagquantum.qec import ErrorEvent, ErrorSchedule, run_repetition_memory_experiment + +result = run_repetition_memory_experiment( + error_schedule=ErrorSchedule((ErrorEvent(round_index=0, wire=1),)), + rounds=3, + shots=16, + seed=0, +) +print(result.logical_error_rate) ``` -## 边界 +扫描只报告有限采样下的观测结果,不给出逻辑抑制或阈值结论,其时序规则也不是最大似然译码。通用码、相关噪声与硬实时硬件反馈仍属研究目标。 + +## QPU 数字孪生 + +数字孪生是一台设备的标定条件模型,由携带设备配置的 `NoiseModel` 构建。执行目标与有序物理映射会成为孪生不可变身份的一部分: + +```{code-block} python +import flagquantum as fq + +twin = fq.twin.from_noise_model( + device_noise_model, + target="your-provider:your-qpu", + qubits=(12, 13), +) +prediction = twin.predict(fq.Circuit(2).h(0).cx(0, 1)) +``` -这些单元是演示规模的研究接口:不对它们附加性能、收敛或硬件结论,默认运行时也不会 -选择它们。其中若干单元的寄存器宽度被有意限制,因为更宽的线路需要本单元并未建模的 -辅助位或态制备成本。 +孪生的预测是经典测量分布之间的全变差一致性,不是量子态保真度。证据只对所声明的线路、操作、映射、物理耦合器、深度、标定快照与置信边界成立:`TwinCircuitSupport` 把证据包络收窄到真正验证过的有向耦合器与深度,组合也不会推断跨单元相关噪声,更不会把局部边界合并成区域级精度主张。模型、验证历史与提交记录都可以在不含凭据的情况下持久化与恢复,且加载永不提交或轮询任务。 diff --git a/docs/flagquantum_zh/user_guide/basic-usage.md b/docs/flagquantum_zh/user_guide/basic-usage.md new file mode 100644 index 0000000000..7235d49db6 --- /dev/null +++ b/docs/flagquantum_zh/user_guide/basic-usage.md @@ -0,0 +1,52 @@ +# 基本用法 + +构建线路、查看规划、执行并读取结果。这是贯穿稳定 API 的最短完整路径。 + +```{code-block} python +import flagquantum as fq + +circuit = fq.Circuit(n_qubits=2).h(0).cx(0, 1) +options = fq.ExecutionOptions(mode="auto", precision="complex64") +plan = fq.plan(circuit, options=options) +result = fq.run(plan) + +print(plan.identity) +print(plan.summary()["recommended_mode"]) +print(result.plan.identity) +print(result.state) +``` + +## 先规划,再执行同一个规划 + +规划解释将要执行什么以及为什么。把规划交给 `fq.run` 会执行这份确切的规划,不重新规划、不重新编译,在同一进程内满足 `result.plan is plan`。规划身份覆盖规范化 IR、解析后的执行语义、编译流水线、所需环境与所选决策,并且可以经受 JSON 往返: + +```{code-block} python +text = plan.to_json() +restored = fq.ExecutionPlan.from_json(text) +result = fq.run(restored) +assert result.plan.identity == plan.identity +``` + +已有规划对语义覆盖是封闭的:同时传入 `options`、`measurements` 或 `noise_model` 会抛出 `TypeError`。环境或 world size 不兼容会在内核启动前失败,而不是悄悄重新规划或回退。 + +## 只查看决策而不执行 + +如果只需要决策,可以使用线路级规划器: + +```{code-block} python +circuit = ( + fq.Circuit(n_qubits=4) + .h(0) + .cx(0, 1) + .rzz(1, 2, theta=0.2) +) + +plan = circuit.runtime_plan(prefer_jax=True, require_gradients=True) +print(plan.summary()) +``` + +规划是对预期执行的解释,不是基准证据。 + +## 本地选项与远程目标 + +`fq.ExecutionOptions` 描述当前进程所控制的资源,例如 `device="cuda:0"` 或模拟 `mode`。`target` 参数则保留给外部执行目标,例如九鼎工作区或 Quafu 后端。详见[硬件与远程目标](hardware-and-remote.md)。 diff --git a/docs/flagquantum_zh/user_guide/circuits-and-ir.md b/docs/flagquantum_zh/user_guide/circuits-and-ir.md new file mode 100644 index 0000000000..12feafaee0 --- /dev/null +++ b/docs/flagquantum_zh/user_guide/circuits-and-ir.md @@ -0,0 +1,71 @@ +# 线路与 FlagQuantum IR + +## 线路构建 + +`fq.Circuit(n_qubits=...)` 是表述线路规模的首选写法。按位置构造、`n_wires=` 与旧的 `nqubits=` 写法仍然兼容;别名冲突会在构造时失败。运行时、编译器与 IR 内部依旧使用 *wire* 表示逻辑映射。 + +```{code-block} python +import flagquantum as fq + +circuit = fq.Circuit(2).h(0).cx(0, 1).ry(1, theta=0.3) +``` + +生成的门的方​​法既保留简洁的位置写法,也接受语义化的量子比特关键字:`h(0)` 与 `h(qubit=0)` 等价,`cx(0, 1)` 与 `cx(control=0, target=1)` 等价,对称双比特门使用 `qubit1=` 与 `qubit2=`。重复、冲突或缺失的量子比特参数会在添加指令之前失败。 + +## 与目标无关的优化 + +编译器优化是面向专家的、与目标无关的变换,返回新的 IR 且不修改输入: + +```{code-block} python +import flagquantum as fq +import flagquantum.compiler as compiler + +circuit = fq.Circuit(2).h(0).h(0).cx(0, 1) +optimized_ir = compiler.optimize(circuit) +result = fq.run(optimized_ir) +``` + +## 面向目标的编译与路由 + +当具体拓扑很重要时,请提供显式耦合图: + +```{code-block} python +import flagquantum.compiler as compiler + +coupling = compiler.CouplingMap.line(circuit.n_qubits) +compiled_ir = compiler.compile( + circuit, + coupling_map=coupling, + routing_strategy="auto", +) +``` + +编译器只输出符合拓扑的双比特操作,并把路由决策记录在 `compiled_ir.metadata["routing"]` 中。它不选择也不调用执行后端。 + +`fq.compile(circuit, compiler="qsteed", target="quafu:")` 是直接选择编译器的路径:产出的 IR 保留逻辑线号,并把有序的物理映射带入部署阶段。 + +## IR 序列化与校验 + +FlagQuantum IR 是版本化、可序列化并经过校验的。`fq.CircuitIR` 属于稳定接口面,`fq.IR_VERSION`、`fq.IRSerializationError` 与 `fq.IRValidationError` 描述了这一边界。不兼容的 schema 变更需要显式迁移。 + +## 错误 + +可以从 `flagquantum.errors` 捕获稳定的生命周期类别: + +```{code-block} python +import flagquantum as fq +import flagquantum.errors as fqe + +try: + result = fq.run(fq.plan(circuit)) +except fqe.ValidationError: + ... # 语义输入非法 +except fqe.PlanningError: + ... # 规划过期、被篡改或不兼容 +except fqe.CapabilityError: + ... # 请求的能力不可用 +except fqe.ExecutionError: + ... # 执行或训练失败 +``` + +所有类别都继承 `FlagQuantumError` 以及与之兼容的 Python 内置异常。类型错误与未知关键字参数仍然抛出 `TypeError`。 diff --git a/docs/flagquantum_zh/user_guide/distributed-execution.md b/docs/flagquantum_zh/user_guide/distributed-execution.md index 155e445b58..811edfdbcc 100644 --- a/docs/flagquantum_zh/user_guide/distributed-execution.md +++ b/docs/flagquantum_zh/user_guide/distributed-execution.md @@ -1,47 +1,43 @@ # 分布式执行 -分布式执行保持一个逻辑负载,并把它切分到多个 rank 上。 +分布式执行把同一套编程模型扩展到真正切分的负载上。同一份逻辑负载被切到多个 rank,复制式执行永远不会被当作容量扩展来呈现。 -## 跨卡切分的态向量 +## 分片态向量 -已初始化的多 rank 进程组改变的是运行时,而不是程序: +在已初始化的多 rank 进程组下,同一态向量模块与结果面会自动使用原生分片态向量运行时。分布式拓扑来自执行环境,而不是另一套模式词汇。 -```bash +```{code-block} shell torchrun --nproc_per_node=4 your_script.py ``` -```python +```{code-block} python import flagquantum as fq -module = fq.Module(build_circuit, n_parameters=2) +def circuit(parameters): + return fq.Circuit(20).ry(0, parameters[0]).cx(0, 1) + +module = fq.Module(circuit, n_parameters=1) +result = module.execute() ``` -此后,同一个模块与结果接口会使用原生跨卡态向量运行时,振幅归属保持在本地 rank。 -该路径禁止物化完整态,仅前向训练的阻塞项会被显式报告,分布式拓扑取自执行环境。 +分片态向量训练让优化器状态与检查点都归所声明的分布所有,因此前向执行、梯度、优化器更新与重启都保持同一套语义。 ## 按 rank 拥有的 MPS -MPS 执行同样可以把一个态分布到多个 rank: +分布式 MPS 训练让前向、反向与优化器状态都由 rank 拥有。它是单设备放不下、纠缠又较低的大系统的路径,并支持语义匹配的检查点与恢复。面向有键结构的系统,还可以使用带 owner 本地有限内存历史的实验性 `adam_lbfgs` 调度。 -```bash +```{code-block} shell python examples/distributed_mps/variable_bond_capacity_8gpu.py +python examples/distributed_statevector_topologies/run.sh ``` -在已复核的证据中,针对已入库的 all-rank 与 all-boundary 负载,一次 batch-one、complex64 -的 MPS 训练步在键维数 768、16 个 rank 上达到 131,072 个格点,单 rank 峰值分配内存 -最大 72.41 GiB,累计丢弃权重 8.39e-06。这是一次确切负载的结果,不是固定计划的强 -扩展性结论。 - -## 分布式训练 +## 这里的分布式证据意味着什么 -按 rank 拥有的态向量与 MPS 训练,可通过 `flagquantum.experimental.distributed` 下显式 -的实验性入口使用,具备按 rank 拥有的优化器状态、检查点/恢复、取消与进度上报。它们 -不会因调用 `fq.train` 而被隐式启用;并且只有当前向执行、梯度、优化器更新与检查点 -归属都保持声明的分布式语义时,分布式训练才算完成。 +- 主张分布式可扩展性,必须把同一份逻辑负载切分到多个 rank。 +- CPU 分布式分层只证明语义与失败即拒绝的行为,不是容量证据。 +- 运行时记录会报告自身的 `distribution_semantics`,从而把分片容量与复制吞吐区分开。 +- 在受支持的分片态向量与 MPS 配置之外,分布式训练仍明确属于实验性。 -## 必须说明的语义 +## FlagOS 加速器 -- 数据并行复制属于吞吐,而不是容量扩展。 -- rank 本地内核与手工张量切片按各自语义报告。 -- CPU 分布式运行只能证明语义与失败即拒行为,绝不能替代真实加速卡的容量证据。 -- 多机发布认证需要获得提升的、经审计的硬件证据。 +在 FlagOS 支持的加速器上,同一份程序通过 Torch-FL 运行在逻辑设备 `flagos:0` 上,因此分布式传输与集合通信都走 FlagOS 路由。详见[硬件与远程目标](hardware-and-remote.md)。 diff --git a/docs/flagquantum_zh/user_guide/hardware-and-remote.md b/docs/flagquantum_zh/user_guide/hardware-and-remote.md new file mode 100644 index 0000000000..f71fe5f10d --- /dev/null +++ b/docs/flagquantum_zh/user_guide/hardware-and-remote.md @@ -0,0 +1,91 @@ +# 硬件与远程目标 + +FlagQuantum 既能在当前进程可控的资源上运行同一份程序,也能把它送到外部执行目标。两者刻意采用不同的命名:`ExecutionOptions` 描述当前进程的设备,而 `target` 命名外部系统。 + +## 经 Torch-FL 使用 FlagOS 加速器 + +`flagquantum` 依赖 PyTorch,而不是 Torch-FL。导入 FlagQuantum、查询后端或运行 CPU 与 CUDA 时都不会导入 `torch_fl`;只有在显式选择 `flagos` 时才会激活这个可选提供方,而 Torch-FL 缺失或不兼容会在激活时报出诊断错误,而不会禁用 CPU 或 CUDA。 + +```{code-block} python +import flagquantum as fq + +result = fq.run( + circuit, + options=fq.ExecutionOptions(device="flagos:0"), +) +``` + +国产加速器的相关职责全部留在 Torch-FL:由它识别厂商与运行时,并通过 `flagos` 契约暴露路由。FlagQuantum 记录 Torch-FL 的运行时身份与路由证据,而不重复实现厂商分支。CUDA 仍是唯一会被自动选择的加速器;`flagos` 需要显式选择。 + +在 `flagos` 上执行本地态向量之前,会强制检查随包发布的算子配置:设备名与环境提示不被接受为已验证证据。单卡认证试验台用于采集由供方独占证明的执行候选结果,即使运行成功也仍然只是待评审候选,而不是自动获得硬件认证。 + +## 远程任务 + +面向 Notebook 的提交最看重可用性:`fq.run()` 会一直等到结果,而 `fq.submit()` 在准备与提供方确认完成后就返回,因此在任务排队期间内核仍然可用。 + +```{code-block} python +import flagquantum as fq + +job = fq.submit(fq.Circuit(2).x(0), target="quafu:Baihua", shots=1024) +job.save("quafu-job.json") +print(job.id) +``` + +```{code-block} python +state = job.status() +if state == "succeeded": + result = job.result() + print(result.counts) +``` + +`status()` 只查询一次,且从不把未知或缺失状态当作成功。`result()` 不做轮询:需要阻塞时请显式使用 `job.wait(timeout=...)`。回执是不含凭据的 JSON,恢复回执不会重新提交任务。九鼎工作区为同一套可观测量接口提供低延迟远程算力: + +```{code-block} python +result = fq.run( + circuit, + target="jiuding:gpu", + outputs=(fq.samples(wires=(0, 1)), fq.counts(wires=(0, 1))), + shots=1024, +) +``` + +采样与计数归约无需返回完整态向量即可执行;可通过 `result.runtime` 与 `result.provenance` 查看所选设备、结果传输、计数聚合以及任何 CPU 回退证据。 + +## 真实量子硬件 + +显式指定编译器与提供方目标可以让这条路径保持失败即拒绝——该路径绝不会隐式选择或替换编译器、提供方: + +```{code-block} python +result = fq.run( + circuit, + compiler="qsteed", + target="quafu:Baihua", + shots=1024, + name="bell calibration", +) +counts = result.measurement("counts").value[0] +``` + +线路会被编译、封装、提交并等待,而稳定的 `fq.ExecutionResult` 返回类型不变。`target_qubits` 可选地按序把逻辑线号映射到物理比特;一旦提供,除非当前目标快照能证明该选择有效且连通,否则编译会失败。同一个远程入口也可以测量单个泡利期望值,量子位逐位对易的各项会在同一物理映射下分到不同的封装任务中测量。 + +## 部署包 + +`flagquantum.deployment.create_deployment_package` 会绑定训练好的参数、面向目标编译,并封装出一个可审计的包,便于之后持久化、签名或提交。预检会把程序与目标做校验,并对确切的包做身份核验,而不联系提供方: + +```{code-block} python +import flagquantum.deployment as deployment +from flagquantum.services import preflight_deployment + +backend = deployment.CloudBackendProfile.simulator(2) +report = preflight_deployment(circuit, backend=backend, shots=1024) +if report.approved_for_submission: + package = report.package +``` + +提供方的支持范围与凭据行为各不相同,能力目录中没有任何提供方获得发布认证。远程路径需要真实的提供方访问权限,本地检查并不覆盖它们。 + +## 边界 + +- 凭据只保留在宿主环境中;不要把原始令牌、密码或 API key 放进扩展清单、错误、测量或序列化载荷。 +- 回执只是本地上下文,不是硬件执行内容的密码学证明。 +- 实验性远程适配器的范围有限:没有自动上传、镜像构建、多卡或多机任务、日志流式输出,也不会自动恢复部分提交的工作。 diff --git a/docs/flagquantum_zh/user_guide/measurement-and-noise.md b/docs/flagquantum_zh/user_guide/measurement-and-noise.md new file mode 100644 index 0000000000..a646b2d0de --- /dev/null +++ b/docs/flagquantum_zh/user_guide/measurement-and-noise.md @@ -0,0 +1,86 @@ +# 测量与噪声 + +## 可观测量与输出 + +用 `fq.X`、`fq.Y`、`fq.Z` 描述数学上的可观测量,再从 `fq.plan` 或 `fq.run` 请求具名输出。泡利乘积使用 `@`;哈密顿量求和与实数系数使用普通算术。 + +```{code-block} python +import flagquantum as fq + +circuit = fq.Circuit(2).h(0).cx(0, 1) +outputs = ( + fq.expectation(fq.Z(0) + fq.Z(1), name="magnetization"), + fq.expectation(fq.X(0) @ fq.Z(1), name="correlation"), + fq.samples(wires=(0, 1)), +) +plan = fq.plan(circuit, outputs=outputs, options=fq.ExecutionOptions(shots=1024, seed=7)) +result = fq.run(plan) + +print(result.expectation("magnetization")) +print(result.expectation("correlation")) +print(result.require_samples()) +``` + +公开的输出工厂是 `expectation`、`probabilities`、`samples` 与 `counts`。采样与计数接受计算基下的线号或一个未加权的泡利乘积,并要求正的采样次数。用 `result.measurement(index_or_name)` 获取指定请求,需要态向量时用 `result.statevector()`。后端原生属性不会隐式透传:`result.native()` 是显式的应急出口。 + +概率与期望值默认精确;计数与采样需要显式给出采样次数。本地执行保留批维度,因此 `counts` 会为每个批次元素返回一个字典。 + +## 噪声模型 + +稳定的含噪执行在规划阶段接受 `flagquantum.noise.NoiseModel`: + +```{code-block} python +import flagquantum as fq +import flagquantum.noise as fqn + +circuit = fq.Circuit(2).h(0).cx(0, 1) +noise = ( + fqn.NoiseModel() + .add("h", fqn.thermal_relaxation_channel(t1=50_000, t2=70_000, duration=35)) + .add("cx", fqn.depolarizing_channel(0.01)) + .add_readout(0, fqn.ReadoutError(((0.98, 0.02), (0.07, 0.93)))) +) + +result = fq.run( + circuit, + noise_model=noise, + options=fq.ExecutionOptions(mode="density_matrix"), + outputs=fq.expectation(fq.Z(0) + fq.Z(1)), +) +print(result.expectation()) +``` + +版本化的模型载荷及其 SHA-256 身份会作为规划的一部分被校验,身份可以通过 `noise.identity` 获取。精确密度矩阵演化是小规模系统的正确性基准;批态向量轨迹与 MPS 量子轨迹会报告采样统计量,MPS 路径还会额外报告截断数据。当精确密度矩阵超出显式的内存预算时,调用方必须明确选择近似路径——否则规划会失败,而不是悄悄改变语义。 + +## 硬件泡利测量规划 + +在基于采样的硬件上测量含 X、Y、Z 项的哈密顿量时,可使用 `create_pauli_measurement_plan`。它会贪心地对量子位逐位对易的项分组、追加所需的基变换旋转,并为每组生成一个已封装的部署包: + +```{code-block} python +import flagquantum.deployment as fqd + +plan = fqd.create_pauli_measurement_plan(circuit, hamiltonian, backend=backend, shots=4096) +results = tuple(provider.run(package) for package in plan.packages) +energy = plan.expectation(tuple(result.counts for result in results)) +``` + +每个包都会记录自己的分组序号、项索引、基、路由证据与封装后的部署身份。 + +## 动态线路 + +候选稳定级的构造函数与实验性执行是隔离的: + +```{code-block} python +import flagquantum as fq +from flagquantum.dynamic import DynamicCircuit + +circuit = DynamicCircuit(2) +circuit.h(0) +circuit.measure(0, classical_bit=0) +circuit.conditional("x", 1, classical_bit=0) + +result = fq.experimental.dynamic.run_dynamic(circuit, shots=128, seed=7) +stable_result = result.to_execution_result() +``` + +本地动态噪声仅限匹配已执行门之后的一比特比特翻转信道,外加独立的读出混淆;其他信道会直接失败。执行前可用只读预检 `fq.experimental.dynamic.assess_dynamic_backend(circuit, backend)` 检查后端兼容性。 diff --git a/docs/flagquantum_zh/user_guide/run-tests.md b/docs/flagquantum_zh/user_guide/run-tests.md index f1ae0ec987..5daae6a0e9 100644 --- a/docs/flagquantum_zh/user_guide/run-tests.md +++ b/docs/flagquantum_zh/user_guide/run-tests.md @@ -1,35 +1,39 @@ # 运行测试 -FlagQuantum 的测试按层级组织。先运行最小且有意义的层级,再根据影响范围逐步扩大。 +先安装开发依赖,然后从最小的有效分层开始,按影响范围逐步扩展。 -## 安装测试依赖 +```{code-block} shell +python -m pip install -e ".[dev]" -```bash -python -m pip install "flagquantum[dev]" -``` +# 日常开发基线 +python tools/ci_tier.py pr-default + +# 本地 API、运行时、规划器与编译器改动 +python tools/ci_tier.py pr-runtime -## 快速开始 +# CPU 上的分布式规划、审计与基准契约 +python tools/ci_tier.py pr-distributed -| 场景 | 命令 | -| --- | --- | -| 日常开发或 issue 基线 | `python tools/ci_tier.py pr-default` | -| 本地 API、运行时、规划器或编译器改动 | `python tools/ci_tier.py pr-runtime` | -| 分布式规划器、审计或基准契约改动 | `python tools/ci_tier.py pr-distributed` | -| 真实多卡或加速卡测试 | `python tools/ci_tier.py gpu-scheduled` | +# 加速器与设备绑定的 Triton 测试 +python tools/ci_tier.py gpu-scheduled +``` -空的选择不是验证:没有收集到任何算例的层级什么也没有证明。 +也可以直接使用 pytest: -## 各层级证明什么 +```{code-block} shell +python -m pytest tests/unit -q +python -m pytest tests/qec -q +python -m pytest -m qiskit +python -m pytest -m pennylane +``` -- 默认层覆盖导入、最小线路、自动求导以及纯规划器/审计辅助。 -- 运行时层覆盖带种子的本地运行时与 API 行为。 -- 分布式 CPU 层覆盖多进程语义、失败即拒门禁、基准 JSON 契约与发布门禁校验;它绝不 - 能替代真实多卡容量扩展。 -- 设备层覆盖加速卡支撑的测试与绑定设备的内核;一张卡覆盖本地一致性,两张卡覆盖必需 - 的跨卡语义,更宽度的定时任务覆盖 rank 数、拓扑、划分与集合通信行为。 +## 各分层证明了什么 -## 发布边界 +| 分层 | 能证明 | 不能证明 | +| --- | --- | --- | +| 推送门禁 | 导入、最小线路、自动微分、纯规划与审计辅助 | 运行时集成、分布式行为、性能、发布就绪 | +| 本地运行时 | 本地运行时与 API 行为的固定种子集成覆盖 | 多进程传输、GPU 执行、可扩展性 | +| 分布式 CPU | CPU 分布式语义、失败即拒绝门禁、基准契约、发布门禁校验 | 真实多卡或多机容量扩展 | +| GPU 定时任务 | 加速器与设备绑定内核测试 | 多机传输或单独构成发布级可扩展性 | -长时间运行的任务会报告阶段、最后操作、已完成工作、内存、集合通信状态与 rank, -无进展看门狗会对停滞进行分类,而不是任其静默挂起。发布级可扩展性证据需要已提升的 -基准载荷通过发布策略与审计命令;正确性套件、覆盖率数字与设备冒烟运行都不是发布证据。 +空的标记选择不算验证。CPU 分布式测试只证明语义,永远不会被用作可扩展性的发布证据。 diff --git a/docs/flagquantum_zh/user_guide/simulation-representations.md b/docs/flagquantum_zh/user_guide/simulation-representations.md new file mode 100644 index 0000000000..c23e256f6a --- /dev/null +++ b/docs/flagquantum_zh/user_guide/simulation-representations.md @@ -0,0 +1,56 @@ +# 模拟表示 + +FlagQuantum 在同一份程序之后提供多套模拟表示,因此负载变化时无需重写模型。 + +| 表示 | 适用场景 | 边界 | +| --- | --- | --- | +| 态向量 | 小到中等规模线路的精确本地模拟与训练 | 容量受单设备限制;分布式容量主张使用分片路径 | +| 矩阵乘积态(MPS) | 比特数很大的低纠缠系统,包含受限 TEBD | 近似质量取决于键维 | +| 张量网络 | 稠密态向量与 MPS 都不合适的线路结构 | 通用反向收缩与生产级分布式传输尚未认证 | +| JAX 内核 | 位于 PyTorch 接口之后的加速内核 | 仅一阶梯度,二阶反向会直接报错 | + +## 不改代码切换表示 + +```{code-block} shell +python examples/quick_start.py --mode sv --steps 40 +python examples/quick_start.py --mode mps --steps 40 +python examples/quick_start.py --mode tn --steps 40 +``` + +同一个混合模型——`torch.nn.Linear` 编码器接到 `fq.Module` 量子层——在每种表示上都用同一个 PyTorch 优化循环训练。 + +## 显式请求某种表示 + +```{code-block} python +import flagquantum as fq + +circuit = fq.Circuit(4).h(0).cx(0, 1).rzz(1, 2, theta=0.2) +result = fq.run( + circuit, + options=fq.ExecutionOptions(mode="mps"), +) +``` + +本地态向量执行是默认路径。需要时可显式选择当前进程可控的一张 GPU: + +```{code-block} python +result = fq.run(circuit, options=fq.ExecutionOptions(device="cuda:0")) +``` + +## 交给规划器判断 + +如果不确定哪种表示合适,先做规划:运行时规划器会报告所选的表示、梯度支持与任何阻塞项,而不是在执行时才失败。 + +```{code-block} python +plan = circuit.runtime_plan(require_gradients=True) +print(plan.summary()) +``` + +## MPS 与张量网络工作流 + +```{code-block} shell +python examples/single_machine_quantum_ai/03_mps_training.py --steps 100 --n-qubits 8 +python examples/vqe_switch_sv_mps_tn.py +``` + +1000 比特的 dimer 示例是面向 PyTorch 侧 JAX/MPS 路径的结构化 MPS 基准,并不是对任意 1000 比特线路的主张。每种表示的准确适用范围见[能力参考](../reference/capabilities.md)。 diff --git a/docs/flagquantum_zh/user_guide/training-with-pytorch.md b/docs/flagquantum_zh/user_guide/training-with-pytorch.md new file mode 100644 index 0000000000..f2096aba73 --- /dev/null +++ b/docs/flagquantum_zh/user_guide/training-with-pytorch.md @@ -0,0 +1,74 @@ +# 使用 PyTorch 训练 + +执行与训练是有意分开的。可训练程序就是放在常规 PyTorch 训练循环里的 `fq.Module`。 + +```{code-block} python +import torch +import flagquantum as fq + + +def build_circuit(parameters, inputs=None): + return ( + fq.Circuit(n_qubits=2) + .ry(0, theta=parameters[0]) + .cx(0, 1) + .ry(1, theta=parameters[1]) + ) + + +module = fq.Module( + build_circuit, + n_parameters=2, + policy=fq.RuntimePolicy(observable_wires=(1,)), +) +optimizer = torch.optim.Adam(module.parameters(), lr=0.01) + +training = fq.train( + module, + optimizer=optimizer, + objective=lambda value: value.mean(), + steps=100, +) + +print(training.losses[-1]) +``` + +## 模块行为 + +- `module(inputs)` 与 `module.forward(inputs)` 返回兼容自动微分的张量,因此同一次反向传播就能把梯度传给量子参数与经典参数。 +- 当调用方需要溯源、运行时诊断或明确的后端信息时,使用 `module.execute(inputs)`,它返回 `fq.ExecutionResult`。 +- `fq.run` 接受线路、IR 或执行规划——不接受模块——并且从不更新参数。 + +`fq.train` 刻意做成由调用方掌握的最小优化循环:它执行 `zero_grad`、`backward`、`step`,然后返回 `fq.TrainingResult`。当量子模块只是更大经典模型的一部分时,请使用常规 PyTorch 循环。 + +## 精度 + +`fq.Module` 通过 `PrecisionPolicy` 统一持有端到端的精度选择,它决定实数参数 dtype 与复数线路/执行 dtype。未指定 `dtype` 的线路构造函数会继承该选择,而显式的 `ExecutionOptions.precision` 或线路 dtype 必须与之一致。模块会在执行前失败,而不是悄悄做类型转换。 + +## 命名参数分组 + +命名分组免去了较大线路中按位置索引的记账工作: + +```{code-block} python +def named_circuit(parameters): + return (fq.Circuit(2) + .ry(0, parameters["encoder"][0]) + .rx(1, parameters["readout"])) + +model = fq.Module( + named_circuit, + parameters={"encoder": (4,), "readout": ()}, + init={"encoder": "uniform", "readout": 0.1}, + seed=42, +) +``` + +`init="uniform"` 从 `[0, 2π)` 采样角度,`init="normal"` 则从均值为零、标准差为 `0.01` 的正态分布采样。`seed` 使用模块内部的生成器,不会重置 PyTorch 的全局随机状态。 + +## 检查点 + +模块的参数与策略都参与 `state_dict()` 的保存与加载,检查点与恢复由 `Module.save_checkpoint()` 与 `Module.load_checkpoint()` 负责,而不是变成 `fq.train` 的隐藏选项。线路构造函数仍属于应用代码,重建模块时必须提供,这与常规 PyTorch 模块的构造方式一致。 + +## 分布式训练 + +在已初始化的多 rank 进程组下,同一模块与结果面会自动使用原生分片态向量运行时。按 rank 拥有的态向量与 MPS 训练也各有独立的分布式入口,目前仍属实验性。详见[分布式执行](distributed-execution.md)。 diff --git a/docs/flagquantum_zh/user_guide/tutorials.md b/docs/flagquantum_zh/user_guide/tutorials.md new file mode 100644 index 0000000000..a116f4352b --- /dev/null +++ b/docs/flagquantum_zh/user_guide/tutorials.md @@ -0,0 +1,44 @@ +# 教程 + +教程系列讲解可运行示例背后的概念。新用户建议按顺序阅读,但每个笔记本同样可以独立使用。 + +| 序号 | 笔记本 | 学习目标 | +| --- | --- | --- | +| 00 | 理解态 | 态张量、振幅、概率与量子比特顺序 | +| 01 | 基本操作 | 单比特与双比特操作 | +| 02 | 测量 | 测量态并解释期望值 | +| 03 | 参数化门 | 可训练门与梯度 | +| 04 | 线路构建器 | 构建并查看可复用线路 | +| 05 | 量子机器学习 | 端到端训练一个小型 QML 模型 | +| 06 | 态向量 VQE | 用本地态向量模拟训练小型 VQE 模型 | +| 07 | 运行时选择 | 比较态向量、MPS 与张量网络的摘要结果 | +| 08 | PyTorch 与 JAX 层 | 训练由 JAX 量子内核支撑的 `fq.Module` | +| 09 | 梯度精度与速度 | 比较各运行时的梯度精度与「值+梯度」速度 | + +笔记本以空输出入库,请在全新的内核中运行。JAX 教程需要可选的 `jax` 依赖。 + +## 教程之外的示例 + +| 目标 | 起点 | +| --- | --- | +| 学习线路、测量、梯度与 QML | 教程 00–05 | +| 验证本地 CPU 或单卡 GPU 路径 | `examples/single_machine_quantum_ai/` | +| 训练本地态向量 VQE | `examples/single_machine_quantum_ai/01_vqe_statevector.py` | +| 使用 MPS 训练 | `examples/single_machine_quantum_ai/03_mps_training.py` | +| 在 PyTorch 中使用 JAX 内核 | `examples/single_machine_quantum_ai/04_jax_kernel_torch_layer.py` | +| 查看分片态向量的归属 | `examples/distributed_statevector_topologies/` | +| 查看按 rank 拥有的 MPS 执行 | `examples/distributed_mps/` | +| 训练并打包一个线路 | `examples/train_parameterized_circuit_then_deploy.py` | +| 构建扩展 | `examples/extensions/`(实验性 API) | +| 端到端运行一个算法单元 | `examples/algorithms/` | + +## 推荐的冒烟运行 + +```{code-block} shell +python examples/single_machine_quantum_ai/00_local_fast_path_check.py +python examples/single_machine_quantum_ai/01_vqe_statevector.py --backend torch --steps 2 --n-qubits 3 +python examples/single_machine_quantum_ai/02_quantum_classifier.py --steps 2 +python examples/single_machine_quantum_ai/03_mps_training.py --steps 2 --n-qubits 4 --max-bond 8 +``` + +示例会先报告正确性参考,再谈速度:VQE 与 MPS 示例会打印小哈密顿量的精确稠密基态能量以及最终能量差,分类器则与教师线路生成的数据对比。 diff --git a/docs/flagquantum_zh/user_guide/user-guide.md b/docs/flagquantum_zh/user_guide/user-guide.md index c21488b60b..adbda2c073 100644 --- a/docs/flagquantum_zh/user_guide/user-guide.md +++ b/docs/flagquantum_zh/user_guide/user-guide.md @@ -1,198 +1,124 @@ # 用户指南 -本指南介绍如何用 FlagQuantum 构建、规划、训练并运行量子程序。 +本指南介绍如何用 FlagQuantum 做量子线路的模拟与训练:构建程序、规划与运行、使用 PyTorch 训练、选择模拟表示、加入噪声与测量、跨 rank 扩展,并把同一份程序迁移到硬件上。 ::::{grid} 1 2 2 3 :gutter: 1 1 1 2 -:::{grid-item-card} {octicon}`play;1.5em;sd-mr-1` 第一个量子模型 -:link: first-quantum-model +:::{grid-item-card} {octicon}`play;1.5em;sd-mr-1` 基本用法 +:link: basic-usage :link-type: doc -只依赖基础安装,端到端训练一个小模型。 +构建线路、规划、执行并读取结果。 +++ -[了解更多 »](first-quantum-model.md) +[了解更多 »](basic-usage.md) ::: -:::{grid-item-card} {octicon}`tools;1.5em;sd-mr-1` 构建与运行 -:link: build-and-run +:::{grid-item-card} {octicon}`code;1.5em;sd-mr-1` 线路与 IR +:link: circuits-and-ir :link-type: doc -构建线路、规划、请求测量、处理错误与回放。 +线路构建、FlagQuantum IR、编译与路由。 +++ -[了解更多 »](build-and-run.md) -::: - -:::{grid-item-card} {octicon}`graph;1.5em;sd-mr-1` 运行时规划 -:link: runtime-planning -:link-type: doc - -执行前查看表示、策略、阻塞项与回退。 - -+++ -[了解更多 »](runtime-planning.md) -::: - -:::{grid-item-card} {octicon}`cpu;1.5em;sd-mr-1` 本地工作流 -:link: local-workflows -:link-type: doc - -零配置的本地模拟、测量、绘制与精度配置。 - -+++ -[了解更多 »](local-workflows.md) -::: - -:::{grid-item-card} {octicon}`server;1.5em;sd-mr-1` 选择模拟器 -:link: choose-a-simulator -:link-type: doc - -态向量、密度矩阵、MPS 与张量网络的取舍与边界。 - -+++ -[了解更多 »](choose-a-simulator.md) +[了解更多 »](circuits-and-ir.md) ::: :::{grid-item-card} {octicon}`gear;1.5em;sd-mr-1` 使用 PyTorch 训练 :link: training-with-pytorch :link-type: doc -模块、命名参数分组、可观测量、检查点与混合模型。 +可训练线路、优化器、命名参数与检查点。 +++ [了解更多 »](training-with-pytorch.md) ::: -:::{grid-item-card} {octicon}`pencil;1.5em;sd-mr-1` 自定义操作 -:link: custom-operations -:link-type: doc - -自定义酉矩阵、算子注册表与扩展。 - -+++ -[了解更多 »](custom-operations.md) -::: - -:::{grid-item-card} {octicon}`tools;1.5em;sd-mr-1` 编译与目标 -:link: compile-and-target +:::{grid-item-card} {octicon}`graph;1.5em;sd-mr-1` 测量与噪声 +:link: measurement-and-noise :link-type: doc -与目标无关的优化、面向拓扑的编译与导出格式。 +可观测量、结果访问、噪声模型与动态线路。 +++ -[了解更多 »](compile-and-target.md) +[了解更多 »](measurement-and-noise.md) ::: -:::{grid-item-card} {octicon}`globe;1.5em;sd-mr-1` 编译器与远程目标 -:link: compiler-and-remote +:::{grid-item-card} {octicon}`cpu;1.5em;sd-mr-1` 模拟表示 +:link: simulation-representations :link-type: doc -面向提供方编译、硬件上测量哈密顿量、非阻塞投递与打包。 +态向量、MPS、张量网络与 JAX 内核。 +++ -[了解更多 »](compiler-and-remote.md) +[了解更多 »](simulation-representations.md) ::: -:::{grid-item-card} {octicon}`zap;1.5em;sd-mr-1` 含噪模拟 -:link: noisy-simulation -:link-type: doc - -一个后端中立的噪声模型,覆盖精确与轨迹执行。 - -+++ -[了解更多 »](noisy-simulation.md) -::: - -:::{grid-item-card} {octicon}`telescope;1.5em;sd-mr-1` 数字孪生与纠错 -:link: digital-twin-and-qec -:link-type: doc - -标定条件下的 QPU 模型,以及重复码存储实验。 - -+++ -[了解更多 »](digital-twin-and-qec.md) -::: - -:::{grid-item-card} {octicon}`telescope;1.5em;sd-mr-1` QPU 数字孪生 -:link: qpu-digital-twin -:link-type: doc - -预测、证据包络、硬件验证与漂移跟踪。 - -+++ -[了解更多 »](qpu-digital-twin.md) -::: - -:::{grid-item-card} {octicon}`shield;1.5em;sd-mr-1` 量子纠错 -:link: qec -:link-type: doc - -症状提取、译码、纠正与逻辑结果分析。 - -+++ -[了解更多 »](qec.md) -::: - -:::{grid-item-card} {octicon}`cloud;1.5em;sd-mr-1` 在硬件上运行 -:link: run-on-hardware -:link-type: doc - -编译、提交、恢复任务,以及面向提供方的打包。 - -+++ -[了解更多 »](run-on-hardware.md) -::: - -:::{grid-item-card} {octicon}`stack;1.5em;sd-mr-1` 分布式执行 +:::{grid-item-card} {octicon}`server;1.5em;sd-mr-1` 分布式执行 :link: distributed-execution :link-type: doc -跨卡切分的态向量与按 rank 拥有的 MPS 负载。 +分片态向量与按 rank 拥有的 MPS 训练。 +++ [了解更多 »](distributed-execution.md) ::: -:::{grid-item-card} {octicon}`git-branch;1.5em;sd-mr-1` 动态线路 -:link: dynamic-circuits +:::{grid-item-card} {octicon}`plug;1.5em;sd-mr-1` 硬件与远程目标 +:link: hardware-and-remote :link-type: doc -线路中测量、条件操作与后端评估。 +FlagOS 加速器、远程任务与部署包。 +++ -[了解更多 »](dynamic-circuits.md) +[了解更多 »](hardware-and-remote.md) ::: -:::{grid-item-card} {octicon}`stack;1.5em;sd-mr-1` 算法 +:::{grid-item-card} {octicon}`beaker;1.5em;sd-mr-1` 算法与纠错 :link: algorithms :link-type: doc -演示规模的算法单元及其优势前提。 +算法单元、纠错实验与数字孪生。 +++ [了解更多 »](algorithms.md) ::: -:::{grid-item-card} {octicon}`file-code;1.5em;sd-mr-1` 示例与教程 -:link: examples-and-tutorials +:::{grid-item-card} {octicon}`book;1.5em;sd-mr-1` 教程 +:link: tutorials :link-type: doc -教程笔记本、脚本示例与冒烟运行。 +教程系列与示例目录。 +++ -[了解更多 »](examples-and-tutorials.md) +[了解更多 »](tutorials.md) ::: -:::{grid-item-card} {octicon}`checklist;1.5em;sd-mr-1` 运行测试 +:::{grid-item-card} {octicon}`flame;1.5em;sd-mr-1` 运行测试 :link: run-tests :link-type: doc -正确性分层、设备通道与发布边界。 +测试分层与标记命令。 +++ [了解更多 »](run-tests.md) ::: :::: + +```{toctree} +:maxdepth: 2 +:hidden: + +basic-usage.md +circuits-and-ir.md +training-with-pytorch.md +measurement-and-noise.md +simulation-representations.md +distributed-execution.md +hardware-and-remote.md +algorithms.md +tutorials.md +run-tests.md +``` From 0d4c5563bf4201601d0a14793beb2699b34c36ee Mon Sep 17 00:00:00 2001 From: cheng874 Date: Tue, 29 Sep 2026 16:18:17 +0800 Subject: [PATCH 2/3] 5th update of FlagQutumn doc --- .../getting_started/quick-start.md | 75 ----------- docs/flagquantum_en/overview/architecture.md | 85 ------------- docs/flagquantum_en/overview/features.md | 53 -------- docs/flagquantum_en/overview/overview.md | 51 -------- docs/flagquantum_en/reference/api.md | 111 ----------------- docs/flagquantum_en/reference/extensions.md | 54 -------- .../reference/operator-capabilities.md | 63 ---------- .../user_guide/build-and-run.md | 72 ----------- .../user_guide/choose-a-simulator.md | 42 ------- docs/flagquantum_en/user_guide/compilation.md | 73 ----------- .../user_guide/compile-and-target.md | 60 --------- .../user_guide/compiler-and-remote.md | 97 --------------- .../user_guide/custom-operations.md | 51 -------- docs/flagquantum_en/user_guide/deployment.md | 117 ------------------ .../user_guide/digital-twin-and-qec.md | 42 ------- .../user_guide/dynamic-circuits.md | 48 ------- .../user_guide/examples-and-tutorials.md | 48 ------- docs/flagquantum_en/user_guide/examples.md | 76 ------------ docs/flagquantum_en/user_guide/extensions.md | 62 ---------- .../user_guide/first-quantum-model.md | 51 -------- .../user_guide/interoperability.md | 94 -------------- .../user_guide/local-workflows.md | 88 ------------- docs/flagquantum_en/user_guide/measurement.md | 106 ---------------- .../user_guide/noisy-simulation.md | 70 ----------- docs/flagquantum_en/user_guide/qec.md | 44 ------- .../user_guide/qpu-digital-twin.md | 57 --------- .../user_guide/run-on-hardware.md | 73 ----------- .../user_guide/runtime-planning.md | 55 -------- .../user_guide/simulation-modes.md | 99 --------------- docs/flagquantum_en/user_guide/training.md | 108 ---------------- .../getting_started/quick-start.md | 69 ----------- docs/flagquantum_zh/overview/architecture.md | 85 ------------- docs/flagquantum_zh/overview/features.md | 52 -------- docs/flagquantum_zh/overview/overview.md | 51 -------- docs/flagquantum_zh/reference/api.md | 46 ------- docs/flagquantum_zh/reference/extensions.md | 54 -------- .../user_guide/compiler-and-remote.md | 97 --------------- .../user_guide/dynamic-circuits.md | 41 ------ .../user_guide/local-workflows.md | 88 ------------- .../user_guide/noisy-simulation.md | 83 ------------- docs/flagquantum_zh/user_guide/qec.md | 44 ------- .../user_guide/qpu-digital-twin.md | 57 --------- .../user_guide/runtime-planning.md | 55 -------- 43 files changed, 2947 deletions(-) delete mode 100644 docs/flagquantum_en/getting_started/quick-start.md delete mode 100644 docs/flagquantum_en/overview/architecture.md delete mode 100644 docs/flagquantum_en/overview/features.md delete mode 100644 docs/flagquantum_en/overview/overview.md delete mode 100644 docs/flagquantum_en/reference/api.md delete mode 100644 docs/flagquantum_en/reference/extensions.md delete mode 100644 docs/flagquantum_en/reference/operator-capabilities.md delete mode 100644 docs/flagquantum_en/user_guide/build-and-run.md delete mode 100644 docs/flagquantum_en/user_guide/choose-a-simulator.md delete mode 100644 docs/flagquantum_en/user_guide/compilation.md delete mode 100644 docs/flagquantum_en/user_guide/compile-and-target.md delete mode 100644 docs/flagquantum_en/user_guide/compiler-and-remote.md delete mode 100644 docs/flagquantum_en/user_guide/custom-operations.md delete mode 100644 docs/flagquantum_en/user_guide/deployment.md delete mode 100644 docs/flagquantum_en/user_guide/digital-twin-and-qec.md delete mode 100644 docs/flagquantum_en/user_guide/dynamic-circuits.md delete mode 100644 docs/flagquantum_en/user_guide/examples-and-tutorials.md delete mode 100644 docs/flagquantum_en/user_guide/examples.md delete mode 100644 docs/flagquantum_en/user_guide/extensions.md delete mode 100644 docs/flagquantum_en/user_guide/first-quantum-model.md delete mode 100644 docs/flagquantum_en/user_guide/interoperability.md delete mode 100644 docs/flagquantum_en/user_guide/local-workflows.md delete mode 100644 docs/flagquantum_en/user_guide/measurement.md delete mode 100644 docs/flagquantum_en/user_guide/noisy-simulation.md delete mode 100644 docs/flagquantum_en/user_guide/qec.md delete mode 100644 docs/flagquantum_en/user_guide/qpu-digital-twin.md delete mode 100644 docs/flagquantum_en/user_guide/run-on-hardware.md delete mode 100644 docs/flagquantum_en/user_guide/runtime-planning.md delete mode 100644 docs/flagquantum_en/user_guide/simulation-modes.md delete mode 100644 docs/flagquantum_en/user_guide/training.md delete mode 100644 docs/flagquantum_zh/getting_started/quick-start.md delete mode 100644 docs/flagquantum_zh/overview/architecture.md delete mode 100644 docs/flagquantum_zh/overview/features.md delete mode 100644 docs/flagquantum_zh/overview/overview.md delete mode 100644 docs/flagquantum_zh/reference/api.md delete mode 100644 docs/flagquantum_zh/reference/extensions.md delete mode 100644 docs/flagquantum_zh/user_guide/compiler-and-remote.md delete mode 100644 docs/flagquantum_zh/user_guide/dynamic-circuits.md delete mode 100644 docs/flagquantum_zh/user_guide/local-workflows.md delete mode 100644 docs/flagquantum_zh/user_guide/noisy-simulation.md delete mode 100644 docs/flagquantum_zh/user_guide/qec.md delete mode 100644 docs/flagquantum_zh/user_guide/qpu-digital-twin.md delete mode 100644 docs/flagquantum_zh/user_guide/runtime-planning.md diff --git a/docs/flagquantum_en/getting_started/quick-start.md b/docs/flagquantum_en/getting_started/quick-start.md deleted file mode 100644 index 0ef3fc7ef3..0000000000 --- a/docs/flagquantum_en/getting_started/quick-start.md +++ /dev/null @@ -1,75 +0,0 @@ -# Quick start - -This page trains a two-qubit model, then shows how to switch the simulation -representation without editing the model. - -## Train your first quantum model - -Build a two-qubit circuit and learn its rotation angle by minimising the -expectation value of `Z` on wire 0: - -```python -import torch -import flagquantum as fq - - -def circuit(parameters): - return fq.Circuit(2).ry(0, parameters[0]).cx(0, 1) - - -model = fq.Module(circuit, n_parameters=1, init=torch.tensor([0.25])) -training = fq.train( - model, - optimizer=torch.optim.Adam(model.parameters(), lr=0.05), - objective=lambda z: z.mean(), - steps=10, -) - -trained_circuit = circuit(next(model.parameters()).detach()) -measurement = fq.expectation(fq.Z(0)) -result = fq.run(trained_circuit, outputs=measurement) -print(result.expectation()) -``` - -`fq.Module` exposes the quantum model to PyTorch and owns its trainable -parameters; `fq.train` performs `zero_grad`, `backward`, and `step`, and returns -a training result whose `final_loss` and `losses` are stable accessors. - -## Inspect a circuit and its plan before running it - -```python -import flagquantum as fq - -circuit = fq.Circuit(n_qubits=2).h(0).cx(0, 1) -options = fq.ExecutionOptions(mode="auto", precision="complex64") -plan = fq.plan(circuit, options=options) -result = fq.run(plan) - -print(plan.identity) -print(plan.summary()["recommended_mode"]) -print(result.state) -``` - -The plan can be serialised and restored. Passing the restored plan to `fq.run` -executes exactly that plan: it is not replanned, and it is not silently replaced -by another backend. - -## Run the same model on another representation - -`examples/quick_start.py` trains a hybrid classical-quantum model whose exact -solution is known, and switches the simulation representation from the command -line: - -```bash -python examples/quick_start.py --mode sv --steps 40 -python examples/quick_start.py --mode mps --steps 40 -python examples/quick_start.py --mode tn --steps 40 -``` - -Statevector is the recommended first run; the MPS and tensor-network support -boundaries are listed in [Simulation modes](../user_guide/simulation-modes.md). - -## Next steps - -- [User guide](../user_guide/user-guide.md) — circuits, training, measurement, noise, twins, error correction, deployment, and remote execution. -- [Reference](../reference.md) — the stable API inventory and the executable operator lowering table. diff --git a/docs/flagquantum_en/overview/architecture.md b/docs/flagquantum_en/overview/architecture.md deleted file mode 100644 index 98f5764d61..0000000000 --- a/docs/flagquantum_en/overview/architecture.md +++ /dev/null @@ -1,85 +0,0 @@ -# Architecture - -FlagQuantum gives a quantum AI program one public model across local development, accelerated kernels, distributed simulation, and deployment: - -```text -fq.Circuit / fq.Module - | - v - FlagQuantum IR - | - +-- compile and export - +-- local statevector, MPS, and tensor-network runtimes - +-- optional JAX kernels behind the PyTorch interface - +-- sharded statevector and MPS execution - +-- deployment packages for provider and hardware targets -``` - -The architectural invariant is simple: backend selection may change execution, but it must not change the meaning of the program or the result contract. - -## Core layers - -| Layer | Responsibility | Entry point | -| --- | --- | --- | -| User API | Circuit construction, PyTorch modules, planning, execution, training, deployment | `import flagquantum as fq` | -| Compilation | Transform circuits and legalize target output without executing them | `fq.compile`, `flagquantum.compiler` | -| FlagQuantum IR | Versioned operators, measurements, metadata, serialization, validation | `fq.CircuitIR` | -| Planning | Select a representation and execution policy; explain blockers and fallbacks | `fq.plan`, `Circuit.runtime_plan` | -| Runtime | Execute locally or across ranks and return typed evidence | `fq.run`, `fq.ExecutionResult` | -| Training | Preserve PyTorch autograd and optimizer semantics across supported runtimes | `fq.Module`, `fq.train` | -| Deployment | Bind trained parameters, compile for a target, seal an auditable package | `flagquantum.deployment.create_deployment_package` | - -A plan describes intent and estimates; it is never execution or benchmark evidence. Runtime records describe what actually ran. - -## Source map - -```text -flagquantum/ -+-- _api.py # root compile, plan, and run composition -+-- circuit.py # circuit construction -+-- core/ # backend-neutral IR and shared semantics -+-- compiler/ # validation, optimization, lowering, code generation -+-- runtime/ # planning, execution lifecycle, results, coordination -+-- simulation/ # numerical methods and kernels -+-- noise/ # backend-neutral noise models and channels -+-- observables/ # user-facing measurement construction -+-- qec/ # error-correction workflows and domain models -+-- twin/ # hardware digital-twin models -+-- compute/ # resources controlled by the current process -+-- remote/ # external task systems and result retrieval -+-- ecosystem/ # framework and format adapters -+-- deployment/ # sealed target-neutral execution packages -+-- services/ # reusable multi-step application workflows -+-- algorithms/ # user-facing algorithm composition -+-- benchmarking/ # reproducible measurement and evidence generation -+-- drawer/ # circuit visualization -+-- testing/ # reusable correctness and conformance helpers -+-- experimental/ # explicitly unstable APIs -``` - -## Dependency direction - -Dependencies point inward: the user facade calls the compiler, runtime, and application workflows; those call `core`. Numerical methods, local compute adapters, and remote adapters sit beside them and are not consulted by `core`. - -- Core does not import orchestration, numerical engines, or vendor integrations. -- The compiler transforms programs but does not execute them. -- Runtime organizes execution but does not implement numerical kernels. -- `simulation` is a numerical-method domain, not a second public runtime. -- Compute and remote isolate hardware and external-system details from other domains. -- Optional integrations stay outside the mandatory local PyTorch path. - -## Execution and training contracts - -`fq.run` is the canonical execution entry point and returns `fq.ExecutionResult` for supported local and distributed modes. Specialized native functions are advanced interfaces and may expose backend-specific objects. - -`fq.train` owns the ordinary PyTorch optimization loop. Owner-sharded statevector and MPS training have separate experimental distributed entry points and are not implied by calling `fq.train`. Distributed training counts as complete only when forward execution, gradients, optimizer updates, and checkpoint ownership all preserve the declared distribution semantics. - -For a claim of distributed scalability, one logical workload must be sharded across ranks. Replicated data parallelism, rank-local kernels, and manual tensor slicing are reported under their own semantics and are never relabelled as capacity expansion. - -## Public versus internal interfaces - -- Public examples use `import flagquantum as fq`. -- Stable names are listed in the [stable API inventory](../reference.md). -- `fq.experimental` carries no compatibility guarantee. -- Compatibility modules support migration; they do not define new stable API. -- Benchmark and research utilities must not become runtime dependencies. diff --git a/docs/flagquantum_en/overview/features.md b/docs/flagquantum_en/overview/features.md deleted file mode 100644 index f0c73485f8..0000000000 --- a/docs/flagquantum_en/overview/features.md +++ /dev/null @@ -1,53 +0,0 @@ -# Features - -This page summarises what FlagQuantum provides. Each area links to the guide that -covers its interface and its evidence boundary. - -## PyTorch-native quantum training - -- `fq.Module` owns trainable quantum parameters and behaves as an ordinary PyTorch module: `forward()` returns an autograd tensor, so existing optimizers, losses, and training loops work unchanged. -- `fq.train` runs a minimal `zero_grad` / `backward` / `step` loop for a single module and returns a versioned training result. -- Classical and quantum layers compose in one model and receive gradients from the same `loss.backward()` call. -- Parameters may be flat or named groups, and checkpoints belong to the module. - -## One program, several representations - -- Local statevector simulation is the default exact path. -- Matrix product state (MPS) simulation targets large, low-entanglement systems. -- Tensor-network execution supports contraction-based workflows. -- Sharded statevector and rank-owned MPS execution extend one logical workload across ranks. -- Optional JAX kernels run behind the same PyTorch interface. - -## Circuits, compilation, and planning - -- `fq.Circuit` builds programs with concise gate methods, and FlagQuantum IR is the versioned representation shared by execution, compilation, and deployment. -- `flagquantum.compiler.optimize` applies target-independent canonical rewrites to a fixed point. -- `flagquantum.compiler.compile` targets an explicit coupling map and emits only topology-valid two-qubit operations, recording its routing decision. -- `fq.plan` returns an explainable runtime plan with explicit blockers and no silent fallback, and an existing plan executes exactly without replanning. - -## Measurement, noise, and precision - -- Named outputs: Pauli expectation values, exact probabilities, samples, and counts. -- Gradients include native autograd and a memory-bounded local adjoint path for Z/ZZ Hamiltonians. -- A backend-neutral noise model covers exact density-matrix evolution, batched statevector trajectories, and MPS quantum trajectories. -- Precision is an explicit per-execution decision, including software-expanded Double-Single arithmetic on FP32-only devices. - -## Deployment and hardware execution - -- Sealed deployment packages bind trained parameters and compile for a target. -- Grouped Pauli measurement plans create one package per qubit-wise-commuting group for shot-based hardware. -- Remote compute and provider targets are addressed by name, with credential-free job receipts that survive a process restart. - -## Digital twins and error correction - -- Calibration-conditioned QPU digital twins predict measurement distributions, then connect those predictions to traceable hardware evidence, drift histories, and holdout validation. -- A repetition-code memory experiment connects syndrome extraction, decoding, correction, and logical-result analysis for fault-tolerance research. - -## Ecosystem and extension - -- Framework-neutral adapters convert to and from PennyLane, Qiskit, Cirq, and CUDA-Q, and explicit bridges execute supported circuits on local PennyLane Lightning, Qiskit Aer, and Cirq simulators. -- An extension SDK supports backend, kernel, operator, compiler-pass, device, provider, measurement-collector, and planner extensions with manifest identity, capability negotiation, and lifecycle containment. - -## Evidence discipline - -Implemented APIs, passing correctness checks, historical measurements, and research goals are three different things. Each capability is graded as release certified, production supported, development evidence, or experimental, and the grade is bound to the exact workload and environment that was validated. diff --git a/docs/flagquantum_en/overview/overview.md b/docs/flagquantum_en/overview/overview.md deleted file mode 100644 index dcf60e1d72..0000000000 --- a/docs/flagquantum_en/overview/overview.md +++ /dev/null @@ -1,51 +0,0 @@ -# Overview - -FlagQuantum is a distributed, differentiable quantum computing framework built on PyTorch. It turns quantum circuits into trainable models: the same program is trained with ordinary PyTorch optimizers, simulated with different representations, scaled across ranks when a workload needs it, and evaluated on remote compute or quantum hardware. It is part of the FlagOS ecosystem — a unified, open-source AI system software stack that integrates diverse models, systems, and chips. - -## Why FlagQuantum? - -A quantum program has three separable concerns: the circuit, the measurement request, and the execution target. FlagQuantum keeps them explicit. - -- A circuit is built once with `fq.Circuit` and stays representation-neutral in FlagQuantum IR. -- Execution is selected at run time — local statevector, MPS, tensor-network, sharded, or a remote target — without rewriting the model. -- Training keeps PyTorch semantics: `fq.Module` returns autograd tensors, and `fq.train` is an ordinary caller-owned optimizer loop. - -The architectural invariant is that backend selection may change execution, but it must not change the meaning of the program or the result contract it returns. - -## Entry points - -| Goal | Primary interface | -| --- | --- | -| Build a program | `fq.Circuit` | -| Inspect its stable representation | `fq.CircuitIR` | -| Plan before executing | `fq.plan`, `Circuit.runtime_plan` | -| Execute locally | `fq.run` | -| Define a trainable quantum layer | `fq.Module` | -| Train | `fq.train` | -| Measure | `fq.expectation`, `fq.probabilities`, `fq.samples`, `fq.counts` | -| Compile or optimize | `fq.compile`, `flagquantum.compiler.optimize` | -| Package for a target | `flagquantum.deployment.create_deployment_package` | -| Run remotely | `fq.run(target=...)`, `fq.submit`, `fq.restore_job` | - -## Capability maturity - -Support is specific to each backend and workload. Every capability is graded, and the grade applies only to the scope that was actually verified: - -| Level | Meaning | -| --- | --- | -| Release certified | Release-gated with audited, reproducible evidence and no unresolved release blocker. | -| Production supported | Supported path with compatibility, operational guidance, and target-hardware evidence. | -| Development evidence | Executable and tested development result; not a production or general scalability claim. | -| Experimental | Research surface without compatibility or production guarantees. | - -A stable public API does not promote an experimental backend, and CPU semantic evidence does not promote a distributed capability. The [Reference](../reference.md) page lists the certified stable names; advertised performance claims require a checked-in audited artifact. - -## How it fits into FlagOS - -FlagQuantum depends on PyTorch, not on a vendor runtime. Domestic accelators are reached through the FlagOS unified multi-chip layer, which owns physical-device detection, vendor runtimes, and the logical `flagos:0` device; FlagQuantum itself contains no vendor dispatch and records the routing evidence it receives. See [Architecture](architecture.md) for the layer boundaries and [Remote execution](../user_guide/remote-execution.md) for provider paths. - -## Where to start - -- [Install](../getting_started/install.md) FlagQuantum and train a two-qubit model in the [quick start](../getting_started/quick-start.md). -- Choose a representation in [Simulation modes](../user_guide/simulation-modes.md) without touching the model. -- Check what is verified before relying on a path: [Capabilities](../reference.md) and the upstream validation scope. diff --git a/docs/flagquantum_en/reference/api.md b/docs/flagquantum_en/reference/api.md deleted file mode 100644 index 188f6ad9a1..0000000000 --- a/docs/flagquantum_en/reference/api.md +++ /dev/null @@ -1,111 +0,0 @@ -# API Reference - -FlagQuantum exposes one curated Python interface: `import flagquantum as fq`. -Build a circuit, inspect its runtime plan, execute it through a stable result -contract, and train parameterized programs with PyTorch. - -Exact stable names are defined by the repository's `public_api_v1.json`, verified -by executable contract tests, and rendered in the stable API inventory below. - -## API map - -| Task | Primary interface | Result | -| --- | --- | --- | -| Build a program | `fq.Circuit` | Circuit backed by FlagQuantum IR | -| Optimize a program | `flagquantum.compiler.optimize` | `fq.CircuitIR` | -| Compile for a selected tool and target | `fq.compile` | `fq.CircuitIR` | -| Inspect execution | `fq.plan`, `Circuit.runtime_plan` | Explainable runtime plan | -| Execute locally or remotely | `fq.run` | `fq.ExecutionResult` | -| Submit a detached job | `fq.submit`, `fq.restore_job` | Job handle with credential-free receipt | -| Define a trainable quantum layer | `fq.Module` | PyTorch module | -| Train | `fq.train` | `fq.TrainingResult` | -| Package for a target | `flagquantum.deployment.create_deployment_package` | Sealed deployment package | - -## Stable API inventory - -| API | Stability | Verification | -| --- | --- | --- | -| `fq.Circuit` | Stable | executable contract | -| `fq.CircuitIR` | Stable | executable contract | -| `fq.ExecutionOptions` | Stable | executable contract | -| `fq.ExecutionPlan` | Stable | executable contract | -| `fq.ExecutionResult` | Stable | executable contract | -| `fq.I` | Stable | executable contract | -| `fq.IRSerializationError` | Stable | executable contract | -| `fq.IRValidationError` | Stable | executable contract | -| `fq.IR_VERSION` | Stable | executable contract | -| `fq.Instruction` | Stable | executable contract | -| `fq.MeasurementResult` | Stable | executable contract | -| `fq.Module` | Stable | executable contract | -| `fq.Observable` | Stable | executable contract | -| `fq.OutputRequest` | Stable | executable contract | -| `fq.Parameter` | Stable | executable contract | -| `fq.ParameterExpression` | Stable | executable contract | -| `fq.RuntimePolicy` | Stable | executable contract | -| `fq.TrainingResult` | Stable | executable contract | -| `fq.X` | Stable | executable contract | -| `fq.Y` | Stable | executable contract | -| `fq.Z` | Stable | executable contract | -| `fq.__version__` | Stable | executable contract | -| `fq.compile` | Stable | executable contract | -| `fq.counts` | Stable | executable contract | -| `fq.expectation` | Stable | executable contract | -| `fq.experimental` | Stable | executable contract | -| `fq.plan` | Stable | executable contract | -| `fq.probabilities` | Stable | executable contract | -| `fq.restore_job` | Stable | executable contract | -| `fq.run` | Stable | executable contract | -| `fq.samples` | Stable | executable contract | -| `fq.submit` | Stable | executable contract | -| `fq.train` | Stable | executable contract | -| `fq.twin` | Stable | executable contract | - -`fq.experimental` is a stable import path, but its contents carry no -compatibility guarantee. Compatibility imports are migration aids and are not -implied stable. - -## Errors - -Catch stable lifecycle categories from `flagquantum.errors`: - -```python -import flagquantum.errors as fqe - -try: - result = fq.run(fq.plan(circuit, options=options)) -except fqe.ValidationError: - ... # invalid semantic input -except fqe.PlanningError: - ... # stale, tampered, or incompatible plan -except fqe.CapabilityError: - ... # requested capability is unavailable -except fqe.ExecutionError: - ... # execution or training failure -``` - -All categories inherit `FlagQuantumError` and the compatible Python built-in -exception (`ValueError`, `RuntimeError`, or `NotImplementedError`). Wrong Python -types and unknown keyword arguments raise `TypeError`, and specific errors such -as `IRValidationError`, `IRSerializationError`, and -`flagquantum.training.TrainingStateError` remain available inside the -corresponding category. - -## Result contract - -`ExecutionResult.diagnostics()` returns a versioned envelope with `metrics`, -`provenance`, `runtime`, and `compatibility` sections whose keys may grow -compatibly. `TrainingResult.final_loss` and its versioned `summary()` provide -stable training output access. Result summaries carry schema and version fields, -and backend-native attributes are not forwarded implicitly: use -`result.native()` when intentionally depending on one. - -## Stability boundaries - -- A planner result describes intent and estimates; it is never runtime or - benchmark evidence. -- Operator and backend support comes from the executable lowering registry in - [Operator Capabilities](operator-capabilities.md). -- Runtime evidence must satisfy the repository's typed runtime contracts. -- Examples in stable documentation are executed by documentation contract tests, - and new public names must first be importable, snapshot-tested, and added to - the stable API manifest. diff --git a/docs/flagquantum_en/reference/extensions.md b/docs/flagquantum_en/reference/extensions.md deleted file mode 100644 index 1c893464b8..0000000000 --- a/docs/flagquantum_en/reference/extensions.md +++ /dev/null @@ -1,54 +0,0 @@ -# Extension SDK - -The extension SDK lets a separately installed package contribute an execution backend, a circuit compiler, a compiler pass, a kernel, an operator, a device, a provider, a measurement collector, or a planner, without becoming a runtime dependency of FlagQuantum. - -## Where it lives - -The approved, pre-freeze SDK contract lives under `flagquantum.ecosystem.extensions` and adds no root exports. Extensions declare a versioned manifest, negotiate capabilities before activation, and are installed into a task-local immutable registry. The namespace moved to its current location without a compatibility layer, and import paths for the submodule remain equivalent. - -## How an extension is written - -A package registers one zero-argument factory in its entry-point group and returns a manifest with a matching identity: - -```{code-block} python -from flagquantum.ecosystem.extensions import ExtensionManifest - -ExtensionManifest( - name="qsteed", - version="0.1.0", - kind="compiler", - capabilities=frozenset({"circuit_ir"}), -) -``` - -A circuit compiler implements negotiation, start, compile, and close. `compile` accepts a FlagQuantum `CircuitIR`, an optional target mapping, and returns a FlagQuantum `CircuitIR`; third-party compiler objects stay inside the plugin. A compiler pass is the smaller transformation hook and exchanges the same IR boundary. - -## Discovery and activation - -Installed packages are discovered only through an explicit, kind-specific discovery call. Discovery validates the entry-point identity against the manifest and adds the result to the existing immutable registry rather than introducing a second plugin registry. Importing FlagQuantum never discovers, imports, or activates an extension. - -A user-facing integration names a compiler or provider explicitly and never relies on implicit selection: - -```{code-block} python -import flagquantum as fq - -result = fq.run(circuit, compiler="qsteed", target="quafu:Baihua", shots=1024) -``` - -## Compatibility lifecycle - -- SDK API mismatches fail during registration with upgrade guidance. -- Individual extensions are experimental by default; stabilisation requires conformance, security review, documentation, and a declared compatibility window. -- Deprecations declare a removal version in the manifest and must retain the previous contract for that window. -- Capability negotiation is fail-closed: a missing dtype, device, gradient, or semantic capability produces blockers before activation. - -## Isolation and security - -- Registration returns a new immutable registry, and scopes are task-local. An extension never mutates root exports, core operator tables, or another task's registry. -- Lifecycle wrappers translate extension exceptions and attempt cleanup after a failed start; cleanup also runs at normal scope exit. -- Raw tokens, passwords, API keys, secrets, and credentials are rejected from extension configuration. Providers must obtain credentials through a host-owned resolver and must not place them in manifests, errors, measurements, or serialized payloads. -- Extensions execute with the authority of the Python process. Discovery is not a sandbox: install only trusted packages. - -## Conformance - -The SDK supplies reusable backend, provider, and circuit-compiler checks that cover manifest and payload serialization, capability honesty, PyTorch gradients, dtype and device preservation, IR ownership, determinism, isolated errors, and cleanup. Passing them is a prerequisite for stabilising an extension, not a substitute for independent qualification of the extension itself. diff --git a/docs/flagquantum_en/reference/operator-capabilities.md b/docs/flagquantum_en/reference/operator-capabilities.md deleted file mode 100644 index b29a54ce1c..0000000000 --- a/docs/flagquantum_en/reference/operator-capabilities.md +++ /dev/null @@ -1,63 +0,0 @@ -# Operator Capabilities - -An executable registered lowering for an operator on a given backend is not a -release or scalability claim. `yes` means the operator can be lowered on that -backend today; the support level of the execution path itself is published in -[Capabilities](capabilities.md). - -| Operator | jax | mps | provider | pytorch | qasm | qcis | tensor_network | -| --- | --- | --- | --- | --- | --- | --- | --- | -| `amplitude_damping` | no | no | no | yes | no | no | no | -| `bit_flip` | no | no | no | yes | no | no | no | -| `ccx` | yes | yes | yes | yes | yes | yes | yes | -| `cphase` | yes | yes | yes | yes | yes | no | yes | -| `crx` | yes | yes | yes | yes | yes | no | yes | -| `cry` | yes | yes | yes | yes | yes | no | yes | -| `crz` | yes | yes | yes | yes | yes | no | yes | -| `cswap` | yes | yes | yes | yes | yes | no | yes | -| `cx` | yes | yes | yes | yes | yes | yes | yes | -| `cy` | yes | yes | yes | yes | yes | yes | yes | -| `cz` | yes | yes | yes | yes | yes | yes | yes | -| `depolarizing` | no | no | no | yes | no | no | no | -| `h` | yes | yes | yes | yes | yes | yes | yes | -| `i` | yes | yes | yes | yes | yes | yes | yes | -| `phase` | yes | yes | yes | yes | yes | yes | yes | -| `phase_flip` | no | no | no | yes | no | no | no | -| `rx` | yes | yes | yes | yes | yes | yes | yes | -| `rxx` | yes | yes | yes | yes | yes | yes | yes | -| `ry` | yes | yes | yes | yes | yes | yes | yes | -| `ryy` | yes | yes | yes | yes | yes | yes | yes | -| `rz` | yes | yes | yes | yes | yes | yes | yes | -| `rzz` | yes | yes | yes | yes | yes | yes | yes | -| `s` | yes | yes | yes | yes | yes | yes | yes | -| `sdg` | yes | yes | yes | yes | yes | yes | yes | -| `swap` | yes | yes | yes | yes | yes | yes | yes | -| `sx` | yes | yes | yes | yes | yes | yes | yes | -| `sxdg` | yes | yes | yes | yes | yes | yes | yes | -| `t` | yes | yes | yes | yes | yes | yes | yes | -| `tdg` | yes | yes | yes | yes | yes | yes | yes | -| `u1` | yes | yes | yes | yes | yes | yes | yes | -| `u2` | yes | yes | yes | yes | yes | yes | yes | -| `u3` | yes | yes | yes | yes | yes | yes | yes | -| `x` | yes | yes | yes | yes | yes | yes | yes | -| `y` | yes | yes | yes | yes | yes | yes | yes | -| `z` | yes | yes | yes | yes | yes | yes | yes | - -Noise channels (`bit_flip`, `phase_flip`, `depolarizing`, `amplitude_damping`) -are registered on the PyTorch path only. They are consumed by the noise model -rather than executed as circuit instructions. - -## Reading the columns - -| Column | Backend | -| --- | --- | -| `pytorch` | Local PyTorch execution, the reference path | -| `jax` | Optional JAX kernels behind the PyTorch interface | -| `mps` | Matrix product state execution | -| `tensor_network` | Tensor-network execution | -| `qasm` | OpenQASM export | -| `qcis` | QCIS export | -| `provider` | Sealed provider and hardware deployment packages | - -The table is generated from the repository's `operator_manifest.json`; treat the -manifest as authoritative and this page as its rendering. diff --git a/docs/flagquantum_en/user_guide/build-and-run.md b/docs/flagquantum_en/user_guide/build-and-run.md deleted file mode 100644 index 35184bd9f9..0000000000 --- a/docs/flagquantum_en/user_guide/build-and-run.md +++ /dev/null @@ -1,72 +0,0 @@ -# Build and Run - -## Build a circuit - -```{code-block} python -import flagquantum as fq - -circuit = fq.Circuit(n_qubits=2).h(0).cx(0, 1) -``` - -`n_qubits` is the preferred public spelling for circuit size. Positional construction, `n_wires=`, and Qiskit-compatible `num_qubits` remain available, and conflicting aliases fail during construction instead of being resolved silently. - -Generated gate methods keep their concise positional form and also accept semantic qubit keywords: `h(0)` and `h(qubit=0)` are equivalent, `cx(0, 1)` and `cx(control=0, target=1)` are equivalent, and symmetric two-qubit gates take `qubit1=` and `qubit2=`. - -The built-in circuit surface covers Pauli gates (`i`, `x`, `y`, `z`), Clifford gates (`h`, `s`, `sdg`, `sx`, `sxdg`, `t`, `tdg`, `cx`, `cy`, `cz`, `swap`), rotations (`rx`, `ry`, `rz`, `rxx`, `ryy`, `rzz`), phase and generic single-qubit gates (`phase`, `u1`, `u2`, `u3`), controlled rotations (`crx`, `cry`, `crz`, `cphase`), Toffoli and Fredkin (`ccx`, `cswap`), and custom matrix operations. - -## Plan, then run - -```{code-block} python -circuit = fq.Circuit(n_qubits=2).h(0).cx(0, 1) -options = fq.ExecutionOptions(mode="auto", precision="complex64") -plan = fq.plan(circuit, options=options) -result = fq.run(plan) - -print(plan.identity) -print(plan.summary()["recommended_mode"]) -print(result.state) -``` - -A plan records the canonical IR, resolved execution semantics, compiler pipeline, required environment, and the selected decision under one SHA-256 identity. Passing a plan to `fq.run` executes exactly that plan: there is no replanning, recompilation, or silent fallback, and a JSON round trip verifies every fingerprint before execution. - -An existing plan is closed to semantic overrides. Passing `options`, `measurements`, or `noise_model` together with a plan raises `TypeError`, and an incompatible environment or world size fails before kernel launch. - -## Request measurements - -```{code-block} python -outputs = ( - fq.expectation(fq.Z(0) + fq.Z(1), name="magnetization"), - fq.expectation(fq.X(0) @ fq.Z(1), name="correlation"), - fq.samples(wires=(0, 1)), -) -plan = fq.plan(circuit, outputs=outputs, options=fq.ExecutionOptions(shots=1024, seed=7)) -result = fq.run(plan) - -print(result.expectation("magnetization")) -print(result.require_samples()) -``` - -Pauli products use `@`; Hamiltonian sums and real coefficients use ordinary arithmetic. The public output factories are `expectation`, `probabilities`, `samples`, and `counts`. Probabilities and expectations are exact by default; `samples` and `counts` require a positive shot count. Use `result.measurement(index_or_name)` for a specific request, `result.statevector()` when a statevector is required, and `result.native()` only when intentionally depending on an unstable backend-native object. - -## Handle errors - -```{code-block} python -import flagquantum.errors as fqe - -try: - result = fq.run(fq.plan(circuit, options=options)) -except fqe.ValidationError: # invalid semantic input - ... -except fqe.PlanningError: # stale, tampered, or incompatible plan - ... -except fqe.CapabilityError: # requested capability is unavailable - ... -except fqe.ExecutionError: # execution or training failure - ... -``` - -All categories inherit `FlagQuantumError` and a compatible Python built-in exception. Wrong Python types and unknown keyword arguments continue to raise `TypeError`. - -## Record and replay - -`Circuit(..., record_op=True)` keeps the instruction list that produced a state, and the same circuit can be exported for tools outside FlagQuantum (see [Compile and Target](compile-and-target.md)). A recorded program is the input to deployment packaging, so a circuit that was trained, compiled, or submitted always has an inspectable source. diff --git a/docs/flagquantum_en/user_guide/choose-a-simulator.md b/docs/flagquantum_en/user_guide/choose-a-simulator.md deleted file mode 100644 index 9be84da93f..0000000000 --- a/docs/flagquantum_en/user_guide/choose-a-simulator.md +++ /dev/null @@ -1,42 +0,0 @@ -# Choose a Simulator - -FlagQuantum selects a simulation representation from the same program. The choice changes resource use and numerical behaviour, not the meaning of the circuit or the result contract. - -| Representation | Select with | Strength | Boundary | -| --- | --- | --- | --- | -| Statevector | `mode="statevector"` or `mode="auto"` | Exact amplitudes, exact expectations, full measurement set | Memory grows with 2ⁿ; a shared state must fit on one device unless sharded | -| Density matrix | `mode="density_matrix"` | Exact noisy evolution | Memory grows with 4ⁿ; small-system correctness oracle for noise work | -| MPS | `mode="mps"` | Low-entanglement circuits with far larger qubit counts | Approximate; truncation is reported | -| Tensor network | `mode="tensor_network"` | Contraction-path execution with native slicing | Experimental; unsupported instructions fail closed instead of falling back | - -## Inspect the decision before running - -```{code-block} python -circuit = fq.Circuit(n_qubits=4).h(0).cx(0, 1).rzz(1, 2, theta=0.2) -plan = circuit.runtime_plan(prefer_jax=True, require_gradients=True) -print(plan.summary()) -``` - -The planner reports the selected representation, gradient support, and the blockers that would prevent a request from running. Use `fq.plan(circuit, options=...)` when the execution options are part of the decision. A plan is an explanation of intent; it is not benchmark evidence. - -## Precision - -Each execution resolves complex precision once. `complex64` implies float32 parameters and real components; `complex128` implies float64. An explicit `ExecutionOptions.precision` or circuit dtype must agree with the module's own precision choice, otherwise the module fails before execution instead of silently casting. - -`RuntimeConfig` is the immutable execution policy for long-lived or distributed work: - -```{code-block} python -from flagquantum.runtime.configuration import RuntimeConfig, runtime_config - -config = RuntimeConfig(device="cuda") -circuit = fq.Circuit(4, config=config) - -with runtime_config(complex_dtype="complex128"): - precise = fq.Circuit(2) # captures complex128 -``` - -Overrides are context-local: nested contexts restore exactly, asyncio tasks inherit a snapshot, threads start from their own default, and processes rebuild the policy from the plan's manifest rather than inheriting mutable state. - -## Software-extended precision - -On devices without native double precision, FlagQuantum can represent the requested logical precision with paired FP32 values ("Double-Single"). This is a backend representation of the requested precision, not a second user-facing precision. The split real/imag and Double-Single paths are explicitly experimental, are never selected by the default runtime, and expose their own acceptance evidence; their supported scope is recorded in the capability catalog. diff --git a/docs/flagquantum_en/user_guide/compilation.md b/docs/flagquantum_en/user_guide/compilation.md deleted file mode 100644 index 5d6cdefd6d..0000000000 --- a/docs/flagquantum_en/user_guide/compilation.md +++ /dev/null @@ -1,73 +0,0 @@ -# Compilation - -Compilation transforms a program without executing it. FlagQuantum separates -target-independent optimization from target-aware lowering. - -## Optimize - -```{code-block} python -import flagquantum as fq -import flagquantum.compiler as compiler - -circuit = fq.Circuit(2).h(0).h(0).cx(0, 1) -optimized_ir = compiler.optimize(circuit) -result = fq.run(optimized_ir) -``` - -`optimize` returns a new `CircuitIR`, leaves the input unchanged, and applies -canonical rewrites to a fixed point, removing redundant gates while preserving -the numerical result. - -## Compile for a target topology - -```{code-block} python -coupling = compiler.CouplingMap.line(circuit.n_qubits) -compiled_ir = compiler.compile( - circuit, - coupling_map=coupling, - routing_strategy="auto", -) -print(compiled_ir.metadata["routing"]) -``` - -The compiler emits only topology-valid two-qubit operations, records its routing -decision, and does not select or invoke an execution backend. Use it when a -concrete coupling map or target-aware lowering is required. - -## Select a compiler for a provider journey - -```{code-block} python -import flagquantum as fq - -compiled = fq.compile(fq.Circuit(2).h(0).cx(0, 1), compiler="qsteed", target="quafu:Baihua") -result = fq.run( - fq.Circuit(2).h(0).cx(0, 1), - compiler="qsteed", - target="quafu:Baihua", - shots=1024, -) -``` - -`fq.compile` is the direct compiler-selection journey. Logical wire numbers are -preserved and an ordered physical `target_qubits` mapping travels through -packaging and submission; omitting `compiler` means provider-side compilation. -The selected compiler, provider, and any fallback are always explicit. - -## Compiler plugins - -A compiler plugin is an independently installed package that owns its compiler -dependency and translation code, for example -`flagquantum-compiler-qsteed`: - -```{code-block} toml -[project.entry-points."flagquantum.extensions"] -"compiler.qsteed" = "flagquantum_compiler_qsteed:create_extension" -``` - -The plugin receives a `CircuitIR`, the target snapshot, and an optional mapping, -and returns a `CircuitIR`; compiler-native objects stay inside the plugin. -Installing a plugin never imports or activates it during `import flagquantum`, -and independently installed plugins are discovered only through an explicit -discovery call with entry-point identity, capability negotiation, determinism, -lifecycle cleanup, and fail-closed errors covered by conformance tests. See -[Extensions](extensions.md). diff --git a/docs/flagquantum_en/user_guide/compile-and-target.md b/docs/flagquantum_en/user_guide/compile-and-target.md deleted file mode 100644 index 5f3cda2343..0000000000 --- a/docs/flagquantum_en/user_guide/compile-and-target.md +++ /dev/null @@ -1,60 +0,0 @@ -# Compile and Target - -## Target-independent optimisation - -```{code-block} python -import flagquantum as fq -import flagquantum.compiler as compiler - -circuit = fq.Circuit(2).h(0).h(0).cx(0, 1) -optimized_ir = compiler.optimize(circuit) -result = fq.run(optimized_ir, options=fq.ExecutionOptions(precision="complex128")) -``` - -`optimize` returns a new `CircuitIR`, leaves the input unchanged, and applies canonical rewrites to a fixed point. It never selects or invokes an execution backend. A runnable example is `python -m examples.compiler_optimize`. - -## Target-aware compilation - -```{code-block} python -coupling = compiler.CouplingMap.line(circuit.n_qubits) -compiled_ir = compiler.compile(circuit, coupling_map=coupling, routing_strategy="auto") -print(compiled_ir.metadata["routing"]) -``` - -The compiler emits only topology-valid two-qubit operations and records its routing decision in the metadata. Logical wire numbers are preserved. A runnable example is `python -m examples.target_aware_compilation`. - -## Select a compiler for a hardware target - -```{code-block} python -result = fq.run( - circuit, - compiler="qsteed", - target="quafu:Baihua", - shots=1024, -) -counts = result.measurement("counts").value[0] -``` - -The named compiler receives the current chip snapshot, selects a physical subgraph, and returns logical IR with an ordered physical mapping that FlagQuantum carries through packaging and submission. When `target_qubits` is provided, its order maps logical wires to physical qubits and compilation fails unless the target snapshot proves the selection is valid and connected; an explicit mapping is never silently replaced. - -## Export - -FlagQuantum programs can leave the framework for other toolchains: - -| Format | Use | -| --- | --- | -| OpenQASM | Circuit exchange with Qiskit, PennyLane, Cirq, CUDA-Q, Azure Quantum, and OpenQASM-based devices | -| QCIS | Submission to QCIS-based providers | - -Conversion at these boundaries is versioned and fail-closed: an operation, parameter expression, or control-flow construct that cannot be represented losslessly is rejected with machine-readable diagnostics rather than approximated. CUDA-Q export is one-way by design; external framework objects never enter the compiler or runtime layers. - -## Compiler plugins - -A compiler plugin registers one zero-argument factory in the `flagquantum.extensions` entry-point group and returns a manifest with the matching identity: - -```{code-block} toml -[project.entry-points."flagquantum.extensions"] -"compiler.qsteed" = "flagquantum_compiler_qsteed:create_extension" -``` - -The extension implements `negotiate`, `start`, `compile`, and `close`; `compile` accepts FlagQuantum `CircuitIR`, an optional target mapping, and returns FlagQuantum `CircuitIR`. Compiler-specific objects stay inside the plugin, so installing a plugin never makes its dependency a FlagQuantum requirement and `import flagquantum` never activates it. diff --git a/docs/flagquantum_en/user_guide/compiler-and-remote.md b/docs/flagquantum_en/user_guide/compiler-and-remote.md deleted file mode 100644 index 3b91b4cd9b..0000000000 --- a/docs/flagquantum_en/user_guide/compiler-and-remote.md +++ /dev/null @@ -1,97 +0,0 @@ -# Compiler and Remote Targets - -Compilation transforms a program; execution runs it. FlagQuantum keeps the two separate so that a compiled artifact can be inspected, sealed, and submitted without ambiguity about what will run. - -## Target-independent optimization - -```{code-block} python -import flagquantum as fq -import flagquantum.compiler as compiler - -circuit = fq.Circuit(2).h(0).h(0).cx(0, 1) -optimized_ir = compiler.optimize(circuit) - -result = fq.run(optimized_ir, options=fq.ExecutionOptions(mode="auto")) -``` - -`optimize` returns a new `CircuitIR`, leaves the input unchanged, and applies canonical rewrites to a fixed point. - -## Target-aware compilation - -Provide an explicit coupling map when the emitted program must respect a topology: - -```{code-block} python -coupling = compiler.CouplingMap.line(circuit.n_qubits) -compiled_ir = compiler.compile( - circuit, - coupling_map=coupling, - routing_strategy="auto", -) -``` - -The compiler emits only topology-valid two-qubit operations and records its routing decision in the compiled IR metadata. It does not select or invoke an execution backend. - -## Compile for a provider - -```{code-block} python -result = fq.run( - circuit, - compiler="qsteed", - target="quafu:Baihua", - shots=1024, -) -counts = result.measurement("counts").value[0] -``` - -This path compiles, packages, submits, and waits for the remote result without changing the `fq.ExecutionResult` return type. It never selects or substitutes a compiler or provider implicitly. - -Optional arguments keep the journey inspectable: - -- `name` labels the deployment; omitting it uses the deployment default, and the provider-assigned task ID remains independent of that display name. -- `target_qubits` gives an optional ordered logical-to-physical mapping. When provided, compilation fails unless the target snapshot proves the mapping is valid and connected; an explicit mapping is never silently replaced. - -## Measure a Hamiltonian on hardware - -A Pauli expectation can use the same entry point. The circuit is compiled once, then qubit-wise-commuting terms are measured in separate sealed jobs without changing the selected physical-qubit mapping: - -```{code-block} python -energy = fq.run( - circuit, - outputs=fq.expectation(0.5 * (fq.X(0) @ fq.X(1)) + fq.Z(0)), - compiler="qsteed", - target="quafu:Baihua", - shots=4096, -).expectation() -``` - -Shots apply to each measurement group. The statistics report the estimator standard error, the number of groups, per-group shots, and total shots; provenance records every provider task and deployment identity. Mixed outputs and unsupported remote outputs fail before compilation or submission. - -## Detached submission and restore - -Long-running remote work does not have to block a notebook: - -```{code-block} python -job = fq.submit(fq.Circuit(2).x(0), target="quafu:Baihua", shots=1024) -job.save("quafu-job.json") -print(job.status()) -``` - -Normalized states are `queued`, `running`, `succeeded`, `failed`, `cancelled`, and `unknown`; a provider state that only means "compilation finished" maps to `queued`, and unknown states never count as success. `job.result()` does not poll, `job.wait(timeout=...)` blocks deliberately, and `job.cancel()` requests cancellation that should be confirmed by a later status query. - -Receipts are credential-free JSON, saving never overwrites an existing file, and restoring a job never resubmits. Configure credentials again after a restart, and reconcile with the provider before retrying a submission whose network response was lost. - -## Packaging a trained program - -A deployment package binds trained parameters, target compilation, and auditable identity so that a program can be saved, signed, or submitted later: - -```{code-block} python -import flagquantum.deployment as deployment - -package = deployment.create_deployment_package( - circuit=trained_circuit, - backend=deployment.CloudBackendProfile.simulator(4), - shots=1024, -) -``` - -Preflight helpers validate a program against a target and identity-check the exact package that may later be submitted, without contacting a provider or consuming remote capacity. diff --git a/docs/flagquantum_en/user_guide/custom-operations.md b/docs/flagquantum_en/user_guide/custom-operations.md deleted file mode 100644 index bd0919cedd..0000000000 --- a/docs/flagquantum_en/user_guide/custom-operations.md +++ /dev/null @@ -1,51 +0,0 @@ -# Custom Operations - -## Custom matrix operations in a circuit - -A custom operation applies a caller-supplied unitary matrix to one or more wires: - -```{code-block} python -import torch -import flagquantum as fq - -matrix = torch.tensor([[0, 1], [1, 0]], dtype=torch.complex64) -circuit = fq.Circuit(2).h(0).matrix(matrix, wires=(1,)).cx(0, 1) -result = fq.run(circuit) -``` - -The matrix shape determines the wire arity, and shape, finite-value, and unitarity checks fail closed before the instruction enters the IR. Custom matrices are stored in the versioned IR, so a circuit that uses one serializes, plans, and compiles like any other circuit. - -## Discover the registered operators - -Every gate and lowering available to the runtime comes from one typed operator registry, and the generated capability table reports which execution paths have an executable lowering: - -```{code-block} python -from flagquantum import operators - -print(operators.gate_info("ry")) -``` - -Gate matrices and their inverses are exposed through the same registry, which is what makes the compiler's canonical rewrites and the capability table consistent with what the runtime can actually execute. - -## Extend FlagQuantum with an extension - -Backends, kernels, operators, compilers, compiler passes, devices, providers, measurement collectors, and planners are extension kinds. An extension is an independently installed Python package that declares a versioned manifest and negotiates capabilities before activation: - -```{code-block} python -from flagquantum.ecosystem.extensions import ( - ExtensionManifest, discover_extensions, -) - -manifest = ExtensionManifest( - name="my-kernel", - version="0.1.0", - kind="kernel", - capabilities=frozenset({"complex64"}), -) -``` - -Installed extensions are discovered only through an explicit, kind-specific `discover_extensions(...)` call; importing FlagQuantum never discovers, imports, or activates one. Registration returns a new immutable registry, lifecycle cleanup runs on normal exit and after a failed start, and credential material is rejected from extension configuration. Individual extensions stay experimental until separately qualified, and discovery is not a sandbox: install only trusted packages. - -## Compiler plugins - -Circuit compilers are a specific extension kind that accepts and returns FlagQuantum `CircuitIR`. See [Compile and Target](compile-and-target.md) for the plugin lifecycle and a working QSteed journey. diff --git a/docs/flagquantum_en/user_guide/deployment.md b/docs/flagquantum_en/user_guide/deployment.md deleted file mode 100644 index 2a77a39669..0000000000 --- a/docs/flagquantum_en/user_guide/deployment.md +++ /dev/null @@ -1,117 +0,0 @@ -# Deployment and Hardware - -## Sealed deployment packages - -`flagquantum.deployment` binds trained parameters, compiles for a target, and -seals an auditable package whose identity can be verified before submission: - -```{code-block} python -import flagquantum.deployment as fqd - -backend = fqd.CloudBackendProfile.simulator(2) -package = fqd.create_deployment_package( - trained_circuit, - backend=backend, - shots=1024, -) -``` - -A deployment package is the artifact a provider journey transports; local -`ExecutionPlan` objects and provider-facing packages stay separate concerns. - -## Preflight before submission - -```{code-block} python -from flagquantum.services import preflight_deployment, preflight_execution - -report = preflight_execution(trained_circuit, target="expectation") -if not report.executable: - print(report.blockers) - -deployment_report = preflight_deployment(trained_circuit, backend=backend, shots=1024) -if deployment_report.approved_for_submission: - package = deployment_report.package -``` - -Preflight converts expected failures into structured blockers and validates the -exact package that may later be submitted, without contacting a provider. -Authentication, approval, tenant state, job persistence, and paid-resource -submission stay with the consuming application. - -## Remote execution - -```{code-block} python -import flagquantum as fq - -# Jiuding: GPU simulation in a running workspace -jiuding_result = fq.run(trained_circuit, target="jiuding:gpu", outputs=fq.expectation(fq.Z(0))) - -# Quafu: compilation, submission, and estimation from measured shots -quafu_result = fq.run( - trained_circuit, target="quafu:Baihua", outputs=fq.expectation(fq.Z(0)), shots=1024, -) -``` - -Both return the canonical `fq.ExecutionResult`. Jiuding computes a simulated -expectation; Quafu estimates one from hardware measurements. Live provider -access is required and is not certified by the local checks. - -Detached jobs keep a notebook responsive while a task is queued or running: - -```{code-block} python -job = fq.submit(fq.Circuit(2).x(0), target="quafu:Baihua", shots=1024) -job.save("quafu-job.json") - -state = job.status() # queued, running, succeeded, failed, cancelled, unknown -result = job.result() # raises unless the job succeeded; job.wait(timeout=...) blocks - -restored = fq.restore_job("quafu-job.json") # a later process; restoration never resubmits -``` - -Receipts are credential-free JSON files that preserve job identity and decoding -context; they never overwrite an existing file and never act as proof of what -hardware executed. - -## QPU digital twins - -An experimental, provider-neutral model predicts a QPU's measurement -distribution from a frozen calibration snapshot: - -```{code-block} python -twin = fq.twin.from_noise_model( - device_noise_model, - target="your-provider:your-qpu", - qubits=(12, 13), -) -prediction = twin.predict(fq.Circuit(2).h(0).cx(0, 1)) -``` - -Twin evidence is qualified by circuit, operations, mapping, physical couplers, -depth, calibration snapshot, and confidence bound. Agreement is total-variation -agreement for classical measurement distributions, not quantum-state fidelity, -and the framework performs no automatic calibration collection, scheduling, -model promotion, or trust decision. Native Quafu support supplies calibration -and execution data; any provider can construct the same model from its own -calibration through the provider-neutral noise and device-profile types. - -## Error correction - -The repetition-code memory experiment connects syndrome extraction, decoding, -and correction in one local reference circuit: - -```{code-block} python -from flagquantum.qec import ErrorEvent, ErrorSchedule, run_repetition_memory_experiment - -result = run_repetition_memory_experiment( - error_schedule=ErrorSchedule((ErrorEvent(round_index=0, wire=1),)), - rounds=3, - shots=16, - seed=0, -) -print(result.logical_error_rate) -``` - -This is development evidence for one fixed three-data-qubit profile: sweeps -report finite-shot observations only, not logical suppression or thresholds, and -general codes, correlated noise, hard-real-time hardware feedback, gradients, -and distributed execution are unsupported. diff --git a/docs/flagquantum_en/user_guide/digital-twin-and-qec.md b/docs/flagquantum_en/user_guide/digital-twin-and-qec.md deleted file mode 100644 index 74094a4f60..0000000000 --- a/docs/flagquantum_en/user_guide/digital-twin-and-qec.md +++ /dev/null @@ -1,42 +0,0 @@ -# Digital Twins and Error Correction - -## QPU digital twins - -A digital twin is a calibration-conditioned model of a physical QPU. It predicts the measurement distribution of a circuit offline, and its predictions are compared with traceable hardware evidence rather than with an ideal simulator. - -```{code-block} python -import flagquantum as fq - -twin = fq.twin.from_noise_model( - device_noise_model, - target="your-provider:your-qpu", - qubits=(12, 13), -) -prediction = twin.predict(fq.Circuit(2).h(0).cx(0, 1)) -``` - -The execution target and the ordered physical mapping are part of the twin's immutable identity, so a prediction always names the device, the calibration snapshot, and the mapping it used. A twin can be saved and restored, compared across calibration snapshots, and validated against recorded hardware tasks. - -Validation evidence is narrowed to the directed physical couplers and circuit depth that were actually exercised: `TwinCircuitSupport` refuses to extend a bound to an untested interaction, a reversed direction, an excessive depth, or an operation of arity greater than two. Overlapping support cells from one device can be composed into a structural region, but local bounds are never combined into a region-level accuracy claim, and composition never infers cross-cell correlated noise. - -Agreement is agreement of classical measurement distributions (total-variation), not quantum-state fidelity. A twin does not route or submit work; submissions stay explicit. - -## Quantum error correction - -The error-correction namespace connects syndrome extraction, decoding, correction, and logical-result analysis. A memory experiment returns a logical error rate together with the syndrome and correction records that produced it: - -```{code-block} python -from flagquantum.qec import ErrorEvent, ErrorSchedule, run_repetition_memory_experiment - -result = run_repetition_memory_experiment( - error_schedule=ErrorSchedule((ErrorEvent(round_index=0, wire=1),)), - rounds=3, - shots=16, - seed=0, -) -print(result.logical_error_rate) -``` - -Supported today: one fixed three-data-qubit repetition-code profile, bounded deterministic X-error schedules, replaceable per-round trajectory decoding with physical-X or Pauli-frame-X actions, a two-round temporal rule that rejects an isolated readout excursion, independent bit flips after parity-check CNOTs, and independent syndrome and final-readout confusion. Feedback traces separate true and observed bits, actions, and frame evolution. - -The temporal rule is not maximum-likelihood decoding, and repeated readout faults can mimic data errors. Sweeps report finite-shot observations only. General codes and channels, correlated or timing noise, batched decoder feedback, hard-real-time or provider control, gradients, and distributed execution are not supported, and logical-suppression or threshold claims require separate statistical and scaling evidence. diff --git a/docs/flagquantum_en/user_guide/dynamic-circuits.md b/docs/flagquantum_en/user_guide/dynamic-circuits.md deleted file mode 100644 index f928453130..0000000000 --- a/docs/flagquantum_en/user_guide/dynamic-circuits.md +++ /dev/null @@ -1,48 +0,0 @@ -# Dynamic circuits - -Dynamic circuits add mid-circuit measurement and classical control. - -## Build and run - -```python -import flagquantum as fq -from flagquantum.dynamic import DynamicCircuit - -circuit = DynamicCircuit(2) -circuit.h(0) -circuit.measure(0, classical_bit=0) -circuit.conditional("x", 1, classical_bit=0) - -result = fq.experimental.dynamic.run_dynamic(circuit, shots=128, seed=7) -stable_result = result.to_execution_result() -``` - -`DynamicCircuit` and its IR encoding are candidate-stable pending API-owner -approval; execution and backend assessment remain experimental. The stable -dynamic path will keep returning the canonical `fq.ExecutionResult`, and -provider-native state will not be frozen into that contract. - -## Execution strategies - -`run_dynamic(..., strategy="auto")` uses batched statevector trajectories for -eligible workloads of at least 32 shots and falls back to the reference -trajectory path when batching would exceed `max_batched_bytes` (256 MiB by -default) or the input is already batched. Callers may request -`strategy="trajectory"` or `"batched"` explicitly, and -`statistics["gate_execution_strategy"]` records the selected path for benchmark -attribution. - -## Backend assessment - -```python -report = fq.experimental.dynamic.assess_dynamic_backend(circuit, backend) -assert report.compatible, report.blockers -``` - -The preflight is read-only and makes no task submission. Local dynamic noise is -limited to one-wire bit-flip channels matched to executed gates plus independent -readout confusion on explicit measurements and final sampling; other channels, -correlated readout, device-profile timing noise, and provider-noise execution -fail closed rather than approximating silently. Provider-neutral conformance -passes locally and on Qiskit Aer, and a real dynamic QPU execution is not -claimed. diff --git a/docs/flagquantum_en/user_guide/examples-and-tutorials.md b/docs/flagquantum_en/user_guide/examples-and-tutorials.md deleted file mode 100644 index 3054356cca..0000000000 --- a/docs/flagquantum_en/user_guide/examples-and-tutorials.md +++ /dev/null @@ -1,48 +0,0 @@ -# Examples and Tutorials - -## Tutorials - -Ten notebooks in `examples/tutorials` teach the concepts behind the runnable examples and are meant to be read in order, though each stands alone: - -| # | Notebook | Learning goal | -| --- | --- | --- | -| 00 | `00_understanding_states.ipynb` | State tensors, amplitudes, probabilities, and qubit order | -| 01 | `01_basic_operations.ipynb` | Single-qubit and two-qubit operations | -| 02 | `02_measurement.ipynb` | Measurement and expectation values | -| 03 | `03_parameterized_gates.ipynb` | Trainable gates and gradients | -| 04 | `04_quantum_circuit_builder.ipynb` | Build and inspect reusable circuits | -| 05 | `05_quantum_machine_learning.ipynb` | Train a small QML model end to end | -| 06 | `06_vqe_statevector.ipynb` | Train a small VQE model with local statevector simulation | -| 07 | `07_runtime_selection_statevector_mps_tn.ipynb` | Compare statevector, MPS, and tensor-network summaries | -| 08 | `08_pytorch_jax_qml_layer.ipynb` | Train an `fq.Module` backed by a JAX quantum kernel | -| 09 | `09_gradient_precision_speed_benchmark.ipynb` | Compare gradient precision and value-plus-gradient speed across runtimes | - -Tutorial 08 needs the optional JAX dependency. Stored notebook outputs are intentionally empty so a reader always runs the current code. - -## Script examples - -| Goal | Example | -| --- | --- | -| Learn circuits, measurements, and gradients | `examples/tutorials` | -| Run one algorithm unit end to end | `examples/algorithms` | -| Verify the local CPU or one-GPU path | `examples/single_machine_quantum_ai` | -| Train a local statevector VQE | `examples/single_machine_quantum_ai/01_vqe_statevector.py` | -| Train with MPS, up to structured thousand-qubit systems | `examples/single_machine_quantum_ai/03_mps_training.py`, `05_mps_1000q_dimer_training.py` | -| Use a JAX kernel through PyTorch | `examples/single_machine_quantum_ai/04_jax_kernel_torch_layer.py` | -| Inspect sharded statevector ownership | `examples/distributed_statevector_topologies` | -| Inspect rank-owned MPS execution | `examples/distributed_mps` | -| Train and package a circuit | `examples/train_parameterized_circuit_then_deploy.py` | -| Build an extension | `examples/extensions/reference_extensions.py` | -| Run on a provider | `examples/remote/quafu_bell.py`, `examples/remote/jiuding_workspace_bell.py` | - -## Suggested smoke runs - -```{code-block} shell -python -m examples.local.simulate -python -m examples.local.measure -python -m examples.local.train -python -m examples.cpu_statevector -python -m examples.quick_start --mode sv --steps 40 -``` - -These commands need neither credentials nor optional backends. The curated single-machine examples do not initialise distributed backends and make no distributed scalability claim. diff --git a/docs/flagquantum_en/user_guide/examples.md b/docs/flagquantum_en/user_guide/examples.md deleted file mode 100644 index 262a87dd33..0000000000 --- a/docs/flagquantum_en/user_guide/examples.md +++ /dev/null @@ -1,76 +0,0 @@ -# Examples and Tutorials - -## Learning path - -Tutorials teach the concepts behind the runnable examples and should be read in -order, though each notebook stands alone. - -| Order | Notebook | Learning goal | -| --- | --- | --- | -| 00 | Understanding states | State tensors, amplitudes, probabilities, and qubit order | -| 01 | Basic operations | Single-qubit and two-qubit operations | -| 02 | Measurement | Measuring states and interpreting expectation values | -| 03 | Parameterized gates | Trainable gates and gradients | -| 04 | Quantum circuit builder | Building and inspecting reusable circuits | -| 05 | Quantum machine learning | Training a small QML model end to end | -| 06 | VQE statevector | Training a small VQE model with local statevector simulation | -| 07 | Runtime selection | Comparing statevector, MPS, and tensor-network summaries | -| 08 | PyTorch and JAX layer | Training an `fq.Module` backed by a JAX quantum kernel | -| 09 | Gradient precision and speed | Comparing gradient precision and value+gradient speed across runtimes | - -Run notebooks in a fresh kernel and keep stored outputs empty; the JAX tutorial -requires the optional `jax` extra. Support boundaries belong in -[Capabilities](../reference/capabilities.md). - -## Recommended entry points - -| Goal | Recommended entry | -| --- | --- | -| Learn circuits, measurements, gradients, and QML | Tutorial notebooks | -| Verify the local CPU or one-GPU path | `single_machine_quantum_ai/00_local_fast_path_check.py` | -| Train a local statevector VQE | `single_machine_quantum_ai/01_vqe_statevector.py` | -| Train with MPS | `single_machine_quantum_ai/03_mps_training.py` | -| Use a JAX kernel through PyTorch | `single_machine_quantum_ai/04_jax_kernel_torch_layer.py` | -| Inspect sharded statevector ownership | `distributed_statevector_topologies/` | -| Inspect rank-owned MPS execution | `distributed_mps/variable_bond_capacity_8gpu.py` | -| Train and package a circuit | `train_parameterized_circuit_then_deploy.py` | -| Build an extension | `extensions/reference_extensions.py` | - -```{code-block} shell -python -m examples.local.simulate -python -m examples.local.measure -python -m examples.local.train -python -m examples.cpu_statevector -python examples/quick_start.py --mode sv --steps 40 -python examples/quick_start.py --mode mps --steps 40 -``` - -The curated single-machine examples initialise no distributed backend and make -no distributed scalability claim. The 1000-qubit dimer example is a -structure-aware MPS benchmark, not a claim about arbitrary 1000-qubit circuits. - -## Remote examples - -```{code-block} shell -python examples/remote/quafu_bell.py -python examples/remote/jiuding_workspace_bell.py -``` - -The Quafu path compiles and validates a circuit before submitting it to -hardware; the Jiuding path reuses a running workspace for low-latency remote -compute. Both require provider credentials and configured remote resources, and -a single visible Jiuding workspace is selected automatically. - -## Plan before executing - -```{code-block} python -import flagquantum as fq - -circuit = fq.Circuit(n_qubits=4).h(0).cx(0, 1).rzz(1, 2, theta=0.2) -plan = circuit.runtime_plan(prefer_jax=True, require_gradients=True) -print(plan.summary()) -``` - -A plan explains intended execution; it is not benchmark evidence. Performance -and scalability statements must use runtime-generated records and report their -`distribution_semantics`. diff --git a/docs/flagquantum_en/user_guide/extensions.md b/docs/flagquantum_en/user_guide/extensions.md deleted file mode 100644 index 26e48af148..0000000000 --- a/docs/flagquantum_en/user_guide/extensions.md +++ /dev/null @@ -1,62 +0,0 @@ -# Extensions - -The approved, pre-freeze SDK contract lives under -`flagquantum.ecosystem.extensions` and adds no root exports. Extensions declare -a versioned manifest, negotiate capabilities before activation, and are -installed into a task-local immutable registry. Individual extensions remain -experimental by default and require independent qualification. - -## Extension kinds - -Execution backends, circuit compilers, compiler passes, kernels, operators, -devices, providers, measurement collectors, and planners. A circuit compiler -accepts and returns FlagQuantum `CircuitIR`; compiler passes are the smaller -transformation hook. - -```{code-block} python -from flagquantum.ecosystem.extensions import ( - ExtensionManifest, discover_extensions, extension_scope, -) - -manifest = ExtensionManifest( - name="my-backend", - version="0.1.0", - kind="backend", - capabilities=frozenset({"circuit_ir"}), -) -``` - -## Discovery and lifecycle - -Installed packages are discovered only through an explicit, kind-specific -`discover_extensions(...)` call; they register zero-argument factories in the -`flagquantum.extensions` entry-point group. Discovery validates the entry-point -identity against the manifest and adds the result to the existing immutable -registry instead of introducing a second plugin registry. Importing FlagQuantum -does not discover, import, or activate extensions. - -- SDK API mismatches fail during registration with upgrade guidance. -- Deprecations declare a removal version in the manifest and must retain the - previous contract for that window. -- Capability negotiation is fail-closed: missing dtype, device, gradient, or - semantic capabilities produce blockers before activation. -- Compatibility failures are `flagquantum.errors.CapabilityError`; lifecycle - failures are `flagquantum.errors.ExecutionError`, so applications handle - extensions with the same stable error categories as core execution. - -## Isolation and credentials - -Registration returns a new immutable registry and `extension_scope` uses -task-local context, so an extension never mutates root exports, core operator -tables, or another task's registry. Lifecycle wrappers translate extension -exceptions and run cleanup after failed startup and at normal scope exit. Raw -tokens, passwords, API keys, secrets, and credentials are rejected from -`ExtensionConfig`; providers obtain credentials through a host-owned resolver -and must not place them in manifests, errors, measurements, or serialized -payloads. - -Extensions execute with the Python process's authority: discovery is not a -sandbox, and only trusted packages should be installed. Conformance helpers -cover manifest and payload serialization, capability honesty, PyTorch -gradients, dtype and device preservation, `CircuitIR` ownership, determinism, -isolated errors, and cleanup. diff --git a/docs/flagquantum_en/user_guide/first-quantum-model.md b/docs/flagquantum_en/user_guide/first-quantum-model.md deleted file mode 100644 index b2a6e90caf..0000000000 --- a/docs/flagquantum_en/user_guide/first-quantum-model.md +++ /dev/null @@ -1,51 +0,0 @@ -# First Quantum Model - -This page trains a small quantum model end to end. It needs only the base installation. - -## Build, train, and measure - -```{code-block} python -import torch -import flagquantum as fq - - -def circuit(parameters): - return fq.Circuit(2).ry(0, parameters[0]).cx(0, 1) - - -model = fq.Module(circuit, n_parameters=1, init=torch.tensor([0.25])) -training = fq.train( - model, - optimizer=torch.optim.Adam(model.parameters(), lr=0.05), - objective=lambda z: z.mean(), - steps=10, -) - -trained_circuit = circuit(next(model.parameters()).detach()) -measurement = fq.expectation(fq.Z(0)) -result = fq.run(trained_circuit, outputs=measurement) -print(result.expectation()) -``` - -`fq.Module` exposes the quantum model to PyTorch, `fq.train` owns the optimizer loop, and `outputs` selects what to measure after training. - -## A hybrid classical-quantum model - -Quantum layers compose with ordinary PyTorch layers. For a complete workflow, run the repository example: - -```{code-block} shell -python examples/quick_start.py --mode sv --steps 40 -``` - -The same model can switch representation without a code change: - -```{code-block} shell -python examples/quick_start.py --mode mps --steps 40 -python examples/quick_start.py --mode tn --steps 40 -``` - -## Next steps - -- [Build and Run](build-and-run.md) explains planning, execution, and measurement in detail. -- [Train with PyTorch](training-with-pytorch.md) covers parameter groups, checkpoints, and precision. -- [Run on Hardware](run-on-hardware.md) takes a trained circuit to a provider or QPU. diff --git a/docs/flagquantum_en/user_guide/interoperability.md b/docs/flagquantum_en/user_guide/interoperability.md deleted file mode 100644 index bdb4107879..0000000000 --- a/docs/flagquantum_en/user_guide/interoperability.md +++ /dev/null @@ -1,94 +0,0 @@ -# Interoperability - -External frameworks connect through one candidate-stable, framework-neutral -adapter protocol under `flagquantum.ecosystem`. The default registry stores -import-safe descriptors and loads an implementation only when requested, so -importing an adapter never imports its dependency. - -```{code-block} python -from flagquantum.ecosystem import available_adapters, get_adapter - -assert available_adapters() == ("braket", "cirq", "cudaq", "pennylane", "qiskit") -adapter = get_adapter("qiskit") -result = adapter.import_program(external_circuit) -flagquantum_ir = result.ir -``` - -Registries are immutable: adding a descriptor returns a new registry and cannot -alter the process-wide default. Every adapter must pass the same identity, -lossless round-trip, fail-closed, and explicit-loss checks through -`run_adapter_conformance()`, which emits a versioned -`flagquantum_interop_conformance_v1` payload. - -## Qiskit - -```{code-block} python -from flagquantum.ecosystem.qiskit import ( - from_qiskit, to_qiskit, import_qiskit, export_qiskit, run_qiskit_conformance, -) -``` - -Certified against Qiskit 2.0.x/2.5.x with Aer 0.17.x for operations, wire order, -statevectors, classical bits, arithmetic parameter expressions, custom unitaries -up to three qubits, and round trips. Functions, powers, other symbolic -operations, and control flow fail closed; named or multiple registers require -explicit lossy flattening. - -The Qiskit Aer bridge executes one fully bound, single-batch circuit on local CPU -Aer and returns an owned `ExecutionResult` with explicit wire order, seed, and -CPU thread controls. - -## PennyLane - -```{code-block} python -from flagquantum.ecosystem.pennylane import from_pennylane, to_pennylane, run -``` - -PennyLane 0.44.1/0.45.1 on Python 3.11 or newer supports static `QuantumScript` -conversion with complex128 semantics. QNodes, device execution, shots, -measurements, trainable parameters, and arbitrary wire labels without explicit -flattening are out of scope for v1. The Lightning bridge executes one fully -bound circuit on local CPU `lightning.qubit` and returns an owned -`ExecutionResult`; it does not support gradients, noise models, dynamic -circuits, automatic routing, GPU, or fallback. - -## Cirq and CUDA-Q - -```{code-block} python -from flagquantum.ecosystem.cirq import from_cirq, to_cirq, run -from flagquantum.ecosystem.cudaq import export_cudaq, to_cudaq -``` - -Cirq support covers the contract gate subset, contiguous `LineQubit` indices, -bound real parameters, and an explicit qubit order; the Simulator bridge -executes one fully bound circuit locally. CUDA-Q export is deliberately -one-way, does not execute the kernel, and reverses bit axes explicitly because -CUDA-Q treats wire zero as least-significant. - -## Dynamic circuits - -```{code-block} python -from flagquantum.dynamic import DynamicCircuit - -circuit = DynamicCircuit(2) -circuit.h(0) -circuit.measure(0, classical_bit=0) -circuit.conditional("x", 1, classical_bit=0) - -report = fq.experimental.dynamic.assess_dynamic_backend(circuit, backend) -result = fq.experimental.dynamic.run_dynamic(circuit, shots=128, seed=7) -stable_result = result.to_execution_result() -``` - -Builders are candidate-stable pending API-owner approval, while dynamic -execution and backend assessment remain experimental. Local dynamic noise is -limited to one-wire bit-flip channels and independent readout confusion, and -provider-noise execution fails closed. No real IQM QPU task has been used; the -Braket IQM integration was validated through the real SDK serializer and a -mocked task contract only. - -## Boundary - -Conversion does not install dependencies, sandbox third-party Python, or -certify provider hardware. External framework objects never enter the -compiler, PyTorch runtime, Torch-FL, accelerator, or QPU layers. diff --git a/docs/flagquantum_en/user_guide/local-workflows.md b/docs/flagquantum_en/user_guide/local-workflows.md deleted file mode 100644 index d043b4771d..0000000000 --- a/docs/flagquantum_en/user_guide/local-workflows.md +++ /dev/null @@ -1,88 +0,0 @@ -# Local Workflows - -Local execution is the zero-configuration path. It needs no provider account, compiler plugin, task scheduler, or network connection. - -## Simulate - -```{code-block} python -import flagquantum as fq - -circuit = fq.Circuit(2).h(0).cx(0, 1) -result = fq.run(circuit) -state = result.to_statevector() -``` - -CPU statevector execution is the default. Select one locally controlled GPU explicitly when needed: - -```{code-block} python -result = fq.run( - circuit, - options=fq.ExecutionOptions(device="cuda:0"), -) -``` - -`ExecutionOptions.device` describes a device controlled by the current process. The `target` argument is reserved for external execution destinations such as a Jiuding GPU workspace or Quafu hardware. - -## Measure - -Request the scientific result you need instead of manually inspecting the statevector: - -```{code-block} python -probabilities = fq.run( - circuit, - outputs=fq.probabilities(), -).probabilities - -correlation = fq.run( - circuit, - outputs=fq.expectation(fq.Z(0) @ fq.Z(1)), -).expectation() - -counts = fq.run( - circuit, - outputs=fq.counts(), - shots=1024, -).counts[0] -``` - -Probabilities and expectations are exact by default. Counts and samples require an explicit shot count. Local execution preserves its batch dimension, so `counts` returns one dictionary per batch item and `[0]` selects the default single-circuit batch. - -## Choose a simulation representation - -The same program can run under different representations without being rewritten: - -```{code-block} python -statevector_result = fq.run(circuit, options=fq.ExecutionOptions(mode="statevector")) -mps_result = fq.run(circuit, options=fq.ExecutionOptions(mode="mps")) -tensor_result = fq.run(circuit, options=fq.ExecutionOptions(mode="tensor_network")) -``` - -Statevector simulation is the recommended first run. MPS suits large low-entanglement systems, and tensor networks suit structured contraction workloads; each representation carries its own support boundary. - -## Precision - -Precision is resolved once per execution. `complex64` implies float32 parameters, and `complex128` implies float64. Use an explicit runtime configuration for long-lived or distributed work so that circuits, plans, and workers agree: - -```{code-block} python -from flagquantum.runtime.configuration import RuntimeConfig - -config = RuntimeConfig(device="cuda", complex_dtype="complex128") -circuit = fq.Circuit(4, config=config) -``` - -## Draw a circuit - -```{code-block} python -print(circuit.draw()) # terminal text form -figure, axes = circuit.draw(format="mpl") # publication-quality figure -``` - -## Run the maintained local paths - -```{code-block} shell -python -m examples.local.simulate -python -m examples.local.measure -python -m examples.local.train -``` - -Continue with the single-machine examples for explicit simulator selection, larger models, and configurable CPU/GPU runs. diff --git a/docs/flagquantum_en/user_guide/measurement.md b/docs/flagquantum_en/user_guide/measurement.md deleted file mode 100644 index 9e8805de29..0000000000 --- a/docs/flagquantum_en/user_guide/measurement.md +++ /dev/null @@ -1,106 +0,0 @@ -# Measurement - -Describe mathematical observables with `fq.X`, `fq.Y`, and `fq.Z`, then request -named outputs from `fq.plan` or `fq.run`. Pauli products use `@`; Hamiltonian -sums and real coefficients use ordinary arithmetic. - -```{code-block} python -import flagquantum as fq - -outputs = ( - fq.expectation(fq.Z(0) + fq.Z(1), name="magnetization"), - fq.expectation(fq.X(0) @ fq.Z(1), name="correlation"), - fq.samples(wires=(0, 1)), -) -plan = fq.plan( - circuit, - outputs=outputs, - options=fq.ExecutionOptions(shots=1024, seed=7), -) -result = fq.run(plan) - -z_sum = result.expectation("magnetization") -xz_value = result.expectation("correlation") -bit_samples = result.require_samples() -``` - -## Output kinds - -The public output factories are `expectation`, `probabilities`, `samples`, and -`counts`: - -| Output | Meaning | Shots | -| --- | --- | --- | -| `fq.expectation(...)` | Exact or estimated Pauli expectation value | Not required | -| `fq.probabilities()` | Exact joint marginal over the requested wires | Not required | -| `fq.samples(wires=...)` | Computational-basis samples | Required | -| `fq.counts(wires=...)` | Aggregated bitstring counts | Required | - -Sampling and counts accept computational-basis wires or one unweighted Pauli -product, and require a positive shot count. Read results through -`result.expectation()`, `result.expectations`, `result.probabilities`, -`result.samples`, and `result.counts`. Unsupported output kinds and missing shot -counts fail before execution. - -`probabilities` reduces the distribution a statevector or density-matrix result -already carries over the complement of the requested wires, and normalizes a -marginal whose total is not one. An MPS or tensor-network result keeps the -parity route instead of materializing a dense state, and the default limit of -eight wires bounds the marginal width; set `max_marginal_wires` explicitly to -request a wider one. - -## Local adjoint Hamiltonian gradients - -For batch-size-one statevector circuits with real, constant-coefficient Z and ZZ -terms, `Hamiltonian.expectation` exposes a memory-bounded adjoint path: - -```{code-block} python -import torch -import flagquantum as fq -from flagquantum import algorithms as fqa - -theta = torch.tensor(0.2, dtype=torch.float64, requires_grad=True) -circuit = fq.Circuit(3, dtype=torch.complex128).ry(0, theta).cx(0, 1) -hamiltonian = fqa.Hamiltonian(( - fqa.pauli_term(0.7, "ZZ", (0, 1)), - fqa.pauli_term(0.2, "Z", (2,)), -)) - -energy = hamiltonian.expectation(circuit, differentiation="adjoint") -energy.backward() -``` - -The default remains `differentiation="autograd"`. Adjoint mode rejects X/Y -terms, trainable or complex coefficients, and batched circuits instead of -silently switching to a different algorithm. - -## Measurement on hardware - -Use `create_pauli_measurement_plan` to measure a Hamiltonian containing X, Y, -and Z terms on shot-based hardware. The plan greedily groups -qubit-wise-commuting terms, appends the required basis rotations, and creates one -sealed deployment package per group: - -```{code-block} python -import flagquantum.deployment as fqd - -plan = fqd.create_pauli_measurement_plan( - circuit, - hamiltonian, - backend=backend, - shots=4096, -) - -results = tuple(provider.run(package) for package in plan.packages) -energy = plan.expectation(tuple(result.counts for result in results)) -``` - -X measurements append H, Y measurements append RZ(-pi/2) followed by H, and -aggregation validates group count, shot totals, bitstring widths, and real -coefficients before producing an expectation value. Each package records its -group index, term indices, basis, routing evidence, and sealed deployment -identity. - -Core IR measurement nodes stay available to runtime implementers for advanced -features such as bounded postselection, but they are intentionally absent from -the root user API. diff --git a/docs/flagquantum_en/user_guide/noisy-simulation.md b/docs/flagquantum_en/user_guide/noisy-simulation.md deleted file mode 100644 index b6734eda3f..0000000000 --- a/docs/flagquantum_en/user_guide/noisy-simulation.md +++ /dev/null @@ -1,70 +0,0 @@ -# Noisy Simulation - -FlagQuantum uses one backend-neutral `NoiseModel` for exact density-matrix -evolution, batched statevector trajectories, and MPS quantum trajectories. The -exact path is the small-system correctness oracle; trajectory paths report -sampling statistics, and the MPS path additionally reports truncation data. - -```{code-block} python -import flagquantum as fq -import flagquantum.noise as fqn -import flagquantum.runtime as fqr - -circuit = fq.Circuit(2).h(0).cx(0, 1) -noise = ( - fqn.NoiseModel() - .add("h", fqn.thermal_relaxation_channel(t1=50_000, t2=70_000, duration=35)) - .add("cx", fqn.depolarizing_channel(0.01)) - .add_readout(0, fqn.ReadoutError(((0.98, 0.02), (0.07, 0.93)))) -) - -exact = fq.run( - circuit, - noise_model=noise, - options=fq.ExecutionOptions(mode="density_matrix"), - outputs=fq.expectation(fq.Z(0) + fq.Z(1)), -) -print(exact.expectation()) - -sampled = fqr.run_noisy_mps( - circuit, - noise, - trajectories=4096, - min_trajectories=128, - target_standard_error=1e-3, - seed=42, -) -print(noise.identity) -print(sampled.expectation_z_mean, sampled.statistics.standard_error) -print(sampled.converged, sampled.stopped_early) -``` - -## Built-in channels - -`bit_flip`, `phase_flip`, `depolarizing`, and `amplitude_damping` channels, -thermal relaxation with gate timing, readout confusion matrices, and -calibration-conditioned device profiles with gate and idle noise lowering. - -## Selection under a memory budget - -The stable `fq.run(...)` entry point can select noisy MPS trajectories when the -exact density matrix would exceed an explicit memory budget. Approximation is -opt-in: without that budget, planning fails rather than silently changing the -semantics of the request. - -## Reproducibility - -A noise model has an identity that participates in plan verification, so a -versioned model survives a plan JSON round trip and a changed model is detected -instead of being applied silently. Trajectory runs accept an explicit seed and -report their sampling statistics. - -## Support boundary - -Validated Markovian Kraus channels, timestamped device profiles, ASAP gate and -idle thermal lowering, classical readout confusion, exact density execution, and -reproducible MPS trajectories with single-rank adaptive stopping are available. -Pulse overlap, crosstalk, leakage, provider calibration adapters, distributed -adaptive stopping, batched statevector trajectories, production multi-GPU -scheduling, and noisy gradients are unsupported. Multi-wire MPS channels use an -explicit dense correctness fallback. diff --git a/docs/flagquantum_en/user_guide/qec.md b/docs/flagquantum_en/user_guide/qec.md deleted file mode 100644 index 59838ee116..0000000000 --- a/docs/flagquantum_en/user_guide/qec.md +++ /dev/null @@ -1,44 +0,0 @@ -# Quantum Error Correction - -QEC connects syndrome extraction, decoding, correction, and logical-result analysis. The long-term direction is a complete workflow for fault-tolerant quantum computing research, including logical operations and hardware feedback. - -QEC owns codes, decoder semantics, detection events, and Pauli frames. It composes compiler control flow, runtime feedback, simulation kernels, noise models, and remote hardware interfaces. - -## Start with a memory experiment - -A three-data-qubit repetition-code memory experiment injects an error and follows it through syndrome extraction, decoding, and correction: - -```{code-block} python -from flagquantum.qec import ErrorEvent, ErrorSchedule, run_repetition_memory_experiment - -result = run_repetition_memory_experiment( - error_schedule=ErrorSchedule((ErrorEvent(round_index=0, wire=1),)), - rounds=3, - shots=16, - seed=0, -) -print(result.logical_error_rate) -``` - -The experiment reports syndrome histories, correction actions, and final logical outcomes, and feedback traces separate true and observed bits, actions, and frame evolution. - -## What the reference implementation covers - -- A fixed repetition-code profile with bounded deterministic error schedules. -- Replaceable per-round trajectory decoding with physical-X or Pauli-frame-X actions. -- A temporal rule that rejects an isolated readout excursion, requiring a following round to confirm a data error. -- A circuit-location stochastic profile with independent bit flips after parity-check operations and independent readout confusion. - -## Run and verify - -```{code-block} shell -python -m pytest tests/qec -q -``` - -Check syndrome histories, correction actions, and final logical outcomes for known injected errors. Decoder changes must also cover readout faults and errors near the final round. - -## Boundaries - -- Sweeps report finite-shot observations only; logical suppression and threshold claims require separate statistical and scaling evidence. -- The temporal rule is not maximum-likelihood decoding, and repeated readout faults can mimic data errors. -- Batched decoder feedback, general channels and codes, correlated or timing noise, hard-real-time provider control, gradients, and distributed execution remain outside the current capability. diff --git a/docs/flagquantum_en/user_guide/qpu-digital-twin.md b/docs/flagquantum_en/user_guide/qpu-digital-twin.md deleted file mode 100644 index ae45c899ba..0000000000 --- a/docs/flagquantum_en/user_guide/qpu-digital-twin.md +++ /dev/null @@ -1,57 +0,0 @@ -# QPU Digital Twin - -A digital twin is a calibration-conditioned model of one QPU. It predicts the measurement distribution you would see on hardware, and it keeps that prediction tied to evidence: which circuits were validated, on which physical couplers, at which depth, at which calibration snapshot, and with which confidence bound. - -## Build a Twin for any QPU - -Start from a `NoiseModel` carrying a device profile. The execution target and the ordered physical mapping become part of the immutable Twin identity: - -```{code-block} python -import flagquantum as fq - -twin = fq.twin.from_noise_model( - device_noise_model, - target="your-provider:your-qpu", - qubits=(12, 13), -) -prediction = twin.predict(fq.Circuit(2).h(0).cx(0, 1)) -``` - -The framework side is provider-neutral: any QPU integration can build the same model after converting its calibration into FlagQuantum's noise and device-profile types. A native Quafu calibration path is included. - -## What a prediction is - -- Predictions cover the full computational-basis measurement distribution. Agreement is a total-variation agreement between measured distributions, not a quantum-state fidelity. -- `evidence_report()` is offline and never touches hardware. -- A validation run binds the corresponding all-qubit measurement when it emits the program for hardware submission; submission itself is always explicit. - -## Qualify evidence by topology and depth - -`TwinCircuitSupport` narrows a statistical envelope to the directed physical couplers and maximum circuit depth that were actually validated. It does not infer support merely because a coupler exists on the device, and it fails closed for untested interactions, reversed directions, excessive depth, or operations of arity greater than two. - -Overlapping contemporaneous cells from one QPU can be composed into a connected-region coverage report for circuit-mapping checks, preserving directed edges and conservative limits. Local statistical bounds are never combined into a region-level accuracy claim. - -## Validate against hardware - -The validation workflow is deliberately explicit and resumable: - -- freeze the circuits you intend to validate and the comparator they will be compared against; -- submit each task explicitly and checkpoint each receipt; -- assemble the ordered results into one simultaneous confidence statement; -- keep Twin↔QPU agreement, noiseless reference agreement, and QPU repeatability as separate quantities. - -Receipts are credential-free JSON and restoration never resubmits, polls, or retries: those remain explicit provider calls. - -## Track drift over time - -Calibration drift and prediction accuracy can be aligned across chronological snapshots for the same fixed circuits, so a chart-ready series of Twin↔QPU agreement, reference agreement, repeatability, and verified bounds is available without claiming causality or choosing a promotion policy. Histories persist offline with private permissions and refuse destructive replacement. - -## Compare candidates prospectively - -An incumbent and a candidate model can be compared on the exact same prospectively collected hardware counts. The result reports a conservative finite-shot interval and returns `improved`, `degraded`, or `inconclusive`; it never promotes a model automatically, and no model is promoted without an explicit application decision. - -## Boundaries - -- Evidence remains specific to declared circuits, operations, mappings, physical couplers, depth, calibration snapshots, and confidence bounds. -- Holdout evidence applies only to the prospectively declared circuits; it does not establish arbitrary-circuit accuracy or training generalization. -- Automatic calibration collection, scheduling, region-wide statistical inference, trust policy, model promotion, and release-certified provider support are outside the framework. diff --git a/docs/flagquantum_en/user_guide/run-on-hardware.md b/docs/flagquantum_en/user_guide/run-on-hardware.md deleted file mode 100644 index 16e59fba59..0000000000 --- a/docs/flagquantum_en/user_guide/run-on-hardware.md +++ /dev/null @@ -1,73 +0,0 @@ -# Run on Hardware - -FlagQuantum keeps one execution contract across local, remote compute, and quantum-hardware targets. The circuit and the requested observable stay explicit; the target changes where the program runs, not what the result means. - -## Compile, package, submit - -```{code-block} python -import flagquantum as fq - -circuit = fq.Circuit(2).h(0).cx(0, 1) -result = fq.run( - circuit, - compiler="qsteed", - target="quafu:Baihua", - shots=1024, - name="bell calibration", -) -counts = result.measurement("counts").value[0] -``` - -This path compiles, packages, submits, and waits for the remote result without changing the `fq.ExecutionResult` return type. It never selects or substitutes a compiler or provider implicitly, and an unsupported remote output fails before compilation or submission. - -One Pauli expectation can use the same entry point. The circuit is compiled once, qubit-wise-commuting terms are measured in separate sealed jobs without changing the selected physical mapping, and the measurement statistics report the estimator standard error, the group count, per-group shots, and total shots: - -```{code-block} python -energy = fq.run( - circuit, - outputs=fq.expectation(0.5 * (fq.X(0) @ fq.X(1)) + fq.Z(0)), - compiler="qsteed", - target="quafu:Baihua", - shots=4096, -).expectation() -``` - -## Submit without blocking - -`fq.run()` waits for a result. When a notebook should stay available while a task is queued, submit it and query later: - -```{code-block} python -job = fq.submit(fq.Circuit(2).x(0), target="quafu:Baihua", shots=1024) -job.save("quafu-job.json") -print(job.id) -``` - -Submission waits only for preparation and the provider's acknowledgement. `status()` queries once and normalises the provider state to `queued`, `running`, `succeeded`, `failed`, `cancelled`, or `unknown`; provider states that mean "compilation finished" are reported as queued, not as completed execution. `result()` does not poll and raises unless the job has succeeded; use `job.wait(timeout=...)` to block deliberately. `job.cancel()` requests cancellation and the state must be queried afterwards to confirm it. Interrupting a cell or closing a notebook is not cancellation. - -## Restore after a restart - -```{code-block} python -job = fq.restore_job("quafu-job.json") -print(job.status()) -``` - -The receipt is JSON, contains job identity and decoding context but no credentials, is written once, and never overwrites an existing file. Restoration never resubmits. If a submission loses its network response, reconcile with the provider before retrying: the server may already have accepted the job. - -## Managed compute targets - -A running Jiuding workspace can execute a circuit as a remote compute target. Sampling and count reduction execute without returning a full statevector: - -```{code-block} python -result = fq.run( - circuit, - target="jiuding:gpu", - outputs=(fq.samples(wires=(0, 1)), fq.counts(wires=(0, 1))), - shots=1024, -) -``` - -Inspect `result.runtime` and `result.provenance` for the selected device, result transfer, counts aggregation, and any CPU-fallback evidence. The Jiuding path and the Quafu provider path both require live provider access and credentials, and their evidence scope is recorded with their guides in the upstream repository. - -## Packaging for a provider - -`flagquantum.deployment.create_deployment_package` binds trained parameters, compiles for a target, and seals an auditable package. `flagquantum.services.preflight_deployment` validates a program against a target and identity-checks the exact package that may later be submitted, without contacting the provider or consuming remote capacity. diff --git a/docs/flagquantum_en/user_guide/runtime-planning.md b/docs/flagquantum_en/user_guide/runtime-planning.md deleted file mode 100644 index 27eb25a12b..0000000000 --- a/docs/flagquantum_en/user_guide/runtime-planning.md +++ /dev/null @@ -1,55 +0,0 @@ -# Runtime Planning - -A plan explains what the runtime intends to do before anything runs. It selects a representation and an execution policy, and it reports blockers and fallbacks instead of silently changing the program's semantics. - -## Inspect a plan - -```{code-block} python -import flagquantum as fq - -circuit = ( - fq.Circuit(n_qubits=4) - .h(0) - .cx(0, 1) - .rzz(1, 2, theta=0.2) -) - -plan = circuit.runtime_plan(prefer_jax=True, require_gradients=True) -print(plan.summary()) -``` - -The same planning entry point is available directly: - -```{code-block} python -plan = fq.plan(circuit, options=fq.ExecutionOptions(mode="auto")) -print(plan.summary()["recommended_mode"]) -``` - -## Plans are executable and reproducible - -Pass the result of `fq.plan` directly to `fq.run` for an inspectable and reproducible execution. The supplied plan is validated and executed without replanning or recompiling, and plan identity covers the canonical IR, the resolved execution semantics, the compiler pipeline, the required environment, and the selected decision: - -```{code-block} python -plan = fq.plan(circuit, options=fq.ExecutionOptions(mode="auto", precision="complex64")) -result = fq.run(plan) -assert result.plan.identity == plan.identity -``` - -JSON round trips verify all fingerprints and the final identity before execution, and an existing plan is closed to semantic overrides: passing `options`, `measurements`, or `noise_model` alongside it raises `TypeError`. An environment or world-size incompatibility fails before kernel launch instead of silently replanning or falling back. - -## Planning is not evidence - -A planner result describes intent and estimates. It is never runtime evidence, benchmark evidence, or a scalability claim. Performance and capacity statements must come from runtime-generated records that state their distribution semantics. - -## Fail-closed behavior - -- Unknown execution options fail before planning rather than being ignored. -- Unsupported output kinds and missing shot counts fail before execution. -- A request that cannot be honored on the selected path raises a capability or planning error instead of switching representation behind the caller's back. -- Distributed execution is described by the execution environment; a conflicting world size or precision fails up front. - -## Related - -- [Local Workflows](local-workflows.md) — executing a planned program locally. -- [Compiler and Remote Targets](compiler-and-remote.md) — compiling for a topology or a provider. -- [Training](training.md) — how the module and the runtime policy interact. diff --git a/docs/flagquantum_en/user_guide/simulation-modes.md b/docs/flagquantum_en/user_guide/simulation-modes.md deleted file mode 100644 index cf1f6884b6..0000000000 --- a/docs/flagquantum_en/user_guide/simulation-modes.md +++ /dev/null @@ -1,99 +0,0 @@ -# Simulation Modes - -One program, several execution representations. `ExecutionOptions(mode=...)` -selects one explicitly, and `Circuit.runtime_plan(...)` explains what the -planner would select and why. - -## Statevector - -The default and the reference implementation. Probabilities and expectations are -exact, small and medium circuits run with zero configuration, and the local -PyTorch path supports training with vector-valued Z observables. - -```{code-block} python -import flagquantum as fq - -result = fq.run( - fq.Circuit(4).h(0).cx(0, 1), - outputs=fq.expectation(fq.Z(0) @ fq.Z(1)), -) -print(result.expectation()) -``` - -Capacity is bounded by one device; for larger states use the sharded runtime or -a low-entanglement representation. On a CUDA-backed FlagOS logical device, the -local statevector path is gated by a packaged operator profile and attached -evidence. - -## Matrix product states - -MPS is the low-entanglement representation: cost follows bond dimension rather -than state size, and rank-owned MPS training distributes one state across -several devices. - -```{code-block} python -from flagquantum.simulation.mps import run_mps - -result = run_mps(circuit, max_bond=64) -``` - -MPS execution reports truncation data, so a caller can see the approximation the -representation introduced. Boundary instructions still execute serially by -owner, the only public capacity measurement is one exact-workload result, and -general scalability or release evidence is not claimed. A constrained MPS TEBD -path covers static real one-site and adjacent two-site Pauli terms on an open -chain with second-order imaginary-time evolution; real-time evolution, periodic -and nonlocal terms, gradients, and TDVP are unsupported. - -## Tensor networks - -Contraction-based execution for structured circuits, with slicing, contraction -ordering, and differentiable reverse contraction. - -```{code-block} python -from flagquantum.simulation.tensor_network import run_tensor_network - -result = run_tensor_network(circuit) -``` - -General reverse contraction and production distributed transport are not -certified, and noise channels are unsupported in this mode: they fail closed -instead of falling back to statevector. - -## Density matrix - -Exact noise evolution for small systems, used as the correctness oracle for the -trajectory paths. - -```{code-block} python -import flagquantum as fq -import flagquantum.noise as fqn - -noise = fqn.NoiseModel().add("x", fqn.bit_flip_channel(0.01)) -result = fq.run( - circuit, - noise_model=noise, - options=fq.ExecutionOptions(mode="density_matrix"), -) -``` - -Continuous-time Lindblad evolution adds dense time-independent Hamiltonians with -Markovian collapse operators on finite strictly increasing time grids, on CPU -complex64/complex128 only. - -## Choosing a mode - -```{code-block} python -plan = circuit.runtime_plan(require_gradients=True) -print(plan.summary()) -``` - -Rules that hold across modes: - -- The planner reports intent and estimates; runtime records report what ran. -- A representation is never silently substituted: an unsupported request fails - during planning instead of falling back. -- Precision is resolved once per execution: `complex64` implies float32 - parameters and `complex128` implies float64. -- Advanced modes are labelled by maturity, and a mode being reachable is not a - production or scalability claim. diff --git a/docs/flagquantum_en/user_guide/training.md b/docs/flagquantum_en/user_guide/training.md deleted file mode 100644 index d8d3c5db17..0000000000 --- a/docs/flagquantum_en/user_guide/training.md +++ /dev/null @@ -1,108 +0,0 @@ -# Training with PyTorch - -Execution and training are intentionally separate. `fq.Module` owns the -trainable quantum parameters, circuit builder, observable selection, runtime -policy, and precision, while `fq.train` is a minimal optimizer loop you can -replace with your own. - -## A trainable quantum layer - -```{code-block} python -import torch -import flagquantum as fq - - -def build_circuit(parameters, inputs=None): - return ( - fq.Circuit(n_qubits=2) - .ry(0, theta=parameters[0]) - .cx(0, 1) - .ry(1, theta=parameters[1]) - ) - - -module = fq.Module( - build_circuit, - n_parameters=2, - policy=fq.RuntimePolicy(observable_wires=(1,)), -) -optimizer = torch.optim.Adam(module.parameters(), lr=0.01) - -training = fq.train( - module, - optimizer=optimizer, - objective=lambda value: value.mean(), - steps=100, -) -print(training.final_loss) -``` - -`module(inputs)` and `module.forward(inputs)` always return an autograd tensor -and use a tensor-only local training path. Use `module.execute(inputs)` when you -need an `ExecutionResult` with provenance, runtime diagnostics, or explicit -backend compatibility information. `fq.run` accepts a circuit, IR, or plan — -never a module — and never updates parameters. - -## Parameter groups - -Flat, named, and symbolic parameters are supported: - -```{code-block} python -def named_circuit(parameters): - return ( - fq.Circuit(2) - .ry(0, parameters["encoder"][0]) - .rx(1, parameters["readout"]) - ) - - -model = fq.Module( - named_circuit, - parameters={"encoder": (4,), "readout": ()}, - init={"encoder": "uniform", "readout": 0.1}, - seed=42, -) -``` - -`init="uniform"` samples angles from `[0, 2*pi)`, `init="normal"` samples a -zero-mean normal distribution with standard deviation 0.01, and `seed` uses a -module-local generator without resetting PyTorch's global random state. A -symbolic circuit can be used directly: its `fq.Parameter` names are inferred, -registered as scalar parameters, and bound automatically on execution. - -## Observables and result shape - -`RuntimePolicy(observable="z", observable_wires=(0, 1))` keeps the observable -axis instead of reducing it, so the leading dimension is the circuit batch -dimension and unbatched circuits return `(1, observable_count)`. Use -`observable="z_sum"` when the observable axis should be summed. Vector-valued Z -execution is available on the local statevector, MPS, and tensor-network fast -paths; JAX and distributed statevector execution remain single-observable and -fail closed when fallback is disabled. - -## Checkpoints, precision, and hybrid models - -`Module.parameters()` and the runtime policy participate in `state_dict()` -save/load, with checkpoint and restore handled by -`Module.save_checkpoint()` / `Module.load_checkpoint()`. Checkpoint/resume -deliberately stays on the module rather than becoming a hidden option of -`fq.train`. - -`fq.Module` owns one end-to-end precision choice through `PrecisionPolicy`: -circuit builders that omit `dtype` inherit it, and an explicit -`ExecutionOptions.precision` or circuit dtype that disagrees fails before -execution instead of silently casting. - -A hybrid model is ordinary PyTorch: place a classical `nn.Linear` encoder -before the quantum layer and train both sides from one `loss.backward()` call. -The repository's `examples/quick_start.py` trains exactly that model with an -analytical target, so it reports correctness as well as loss. - -## Distributed training - -Under an initialized multi-rank process group, the same statevector module and -result surface use the native sharded statevector runtime; distribution topology -comes from the execution environment rather than a second mode vocabulary. -Owner-sharded statevector and MPS training have separate experimental entry -points under `flagquantum.experimental.distributed` and are not implied by -calling `fq.train`. See [Distributed Execution](distributed-execution.md). diff --git a/docs/flagquantum_zh/getting_started/quick-start.md b/docs/flagquantum_zh/getting_started/quick-start.md deleted file mode 100644 index 5f607c0f6c..0000000000 --- a/docs/flagquantum_zh/getting_started/quick-start.md +++ /dev/null @@ -1,69 +0,0 @@ -# 快速开始 - -本页训练一个双量子比特模型,然后演示如何在不改动模型的情况下切换模拟表示。 - -## 训练第一个量子模型 - -构建一个双量子比特线路,并通过最小化 0 号线上的 `Z` 期望值来学习它的旋转角: - -```python -import torch -import flagquantum as fq - - -def circuit(parameters): - return fq.Circuit(2).ry(0, parameters[0]).cx(0, 1) - - -model = fq.Module(circuit, n_parameters=1, init=torch.tensor([0.25])) -training = fq.train( - model, - optimizer=torch.optim.Adam(model.parameters(), lr=0.05), - objective=lambda z: z.mean(), - steps=10, -) - -trained_circuit = circuit(next(model.parameters()).detach()) -measurement = fq.expectation(fq.Z(0)) -result = fq.run(trained_circuit, outputs=measurement) -print(result.expectation()) -``` - -`fq.Module` 把量子模型暴露给 PyTorch,并拥有其可训练参数;`fq.train` 执行 -`zero_grad`、`backward`、`step`,返回的训练结果提供 `final_loss` 与 `losses` -等稳定访问器。 - -## 执行前先查看线路与计划 - -```python -import flagquantum as fq - -circuit = fq.Circuit(n_qubits=2).h(0).cx(0, 1) -options = fq.ExecutionOptions(mode="auto", precision="complex64") -plan = fq.plan(circuit, options=options) -result = fq.run(plan) - -print(plan.identity) -print(plan.summary()["recommended_mode"]) -print(result.state) -``` - -计划可以序列化与恢复。把恢复后的计划交给 `fq.run` 会精确执行该计划:不会重新规划, -也不会被静默替换成另一个后端。 - -## 让同一个模型换一种表示运行 - -`examples/quick_start.py` 训练一个解已知的经典—量子混合模型,并可在命令行切换模拟表示: - -```bash -python examples/quick_start.py --mode sv --steps 40 -python examples/quick_start.py --mode mps --steps 40 -python examples/quick_start.py --mode tn --steps 40 -``` - -建议先跑态向量模式;MPS 与张量网络的支持边界见[模拟模式](../user_guide/simulation-modes.md)。 - -## 下一步 - -- [用户指南](../user_guide/user-guide.md)——线路、训练、测量、噪声、数字孪生、纠错、部署与远程执行。 -- [参考资料](../reference.md)——稳定 API 清单与可执行算子下发表。 diff --git a/docs/flagquantum_zh/overview/architecture.md b/docs/flagquantum_zh/overview/architecture.md deleted file mode 100644 index e4878843ac..0000000000 --- a/docs/flagquantum_zh/overview/architecture.md +++ /dev/null @@ -1,85 +0,0 @@ -# 架构 - -FlagQuantum 为量子 AI 程序提供统一模型,覆盖本地开发、加速内核、分布式模拟与部署: - -```text -fq.Circuit / fq.Module - | - v - FlagQuantum IR - | - +-- 编译与导出 - +-- 本地态向量、MPS 与张量网络运行时 - +-- PyTorch 接口之后的可选 JAX 内核 - +-- 跨卡切分的态向量与 MPS 执行 - +-- 面向厂商与硬件目标的部署包 -``` - -架构不变式很简单:后端选择可以改变执行方式,但不能改变程序含义或结果契约。 - -## 核心分层 - -| 层 | 职责 | 入口 | -| --- | --- | --- | -| 用户 API | 线路构建、PyTorch 模块、规划、执行、训练、部署 | `import flagquantum as fq` | -| 编译 | 变换线路并合法化目标输出,但不执行 | `fq.compile`、`flagquantum.compiler` | -| FlagQuantum IR | 带版本的算子、测量、元数据、序列化与校验 | `fq.CircuitIR` | -| 规划 | 选择表示与执行策略,解释阻塞与回退 | `fq.plan`、`Circuit.runtime_plan` | -| 运行时 | 本地或跨 rank 执行,返回带类型的结果 | `fq.run`、`fq.ExecutionResult` | -| 训练 | 在受支持的运行时上保持 PyTorch 自动求导与优化器语义 | `fq.Module`、`fq.train` | -| 部署 | 绑定训练参数、面向目标编译、密封可审计的部署包 | `flagquantum.deployment.create_deployment_package` | - -计划描述的是意图与估算,它永远不是执行或基准证据;运行时记录描述的是实际运行了什么。 - -## 源码结构 - -```text -flagquantum/ -+-- _api.py # 根级 compile / plan / run 组合 -+-- circuit.py # 线路构建 -+-- core/ # 与后端无关的 IR 与共享语义 -+-- compiler/ # 校验、优化、下降与代码生成 -+-- runtime/ # 规划、执行生命周期、结果与协调 -+-- simulation/ # 数值方法与内核 -+-- noise/ # 与后端无关的噪声模型与信道 -+-- observables/ # 面向用户的测量构建 -+-- qec/ # 纠错工作流与领域模型 -+-- twin/ # 硬件数字孪生模型 -+-- compute/ # 当前进程可支配的资源 -+-- remote/ # 外部任务系统与结果获取 -+-- ecosystem/ # 框架与格式适配器 -+-- deployment/ # 密封的、与目标无关的执行包 -+-- services/ # 可复用的多步应用工作流 -+-- algorithms/ # 面向用户的算法组合 -+-- benchmarking/ # 可复现的测量与证据生成 -+-- drawer/ # 线路可视化 -+-- testing/ # 可复用的正确性与一致性辅助 -+-- experimental/ # 明确不稳定的 API -``` - -## 依赖方向 - -依赖朝内:用户门面调用编译器、运行时与应用工作流,它们再调用 `core`。数值方法、本地算力适配器与远程适配器并列其侧,`core` 不会调用它们。 - -- `core` 不导入编排、数值引擎或厂商集成。 -- 编译器变换程序,但不执行程序。 -- 运行时组织执行,但不实现数值内核。 -- `simulation` 是数值方法域,而不是第二个公开运行时。 -- `compute` 与 `remote` 把硬件与外部系统的细节与其他域隔离。 -- 可选集成不得进入本地 PyTorch 的必需路径。 - -## 执行与训练契约 - -`fq.run` 是规范的执行入口,对受支持的本地与分布式模式返回 `fq.ExecutionResult`。专用原生函数属于进阶接口,可能暴露后端特有的对象。 - -`fq.train` 承担常规的 PyTorch 优化循环。按 rank 拥有的态向量与 MPS 训练有各自独立的实验性分布式入口,调用 `fq.train` 不会隐式启用它们。只有当前向执行、梯度、优化器更新与检查点归属都保持声明的分布式语义时,分布式训练才算完成。 - -要给出分布式可扩展性结论,必须把同一个逻辑负载切分到多个 rank 上。数据并行复制、rank 本地内核以及手工张量切片,各自以自身语义报告,绝不被重新表述为容量扩展。 - -## 公开与内部接口 - -- 公开示例统一使用 `import flagquantum as fq`。 -- 稳定名称列在[稳定 API 清单](../reference.md)中。 -- `fq.experimental` 不提供兼容性保证。 -- 兼容模块用于迁移,不定义新的稳定 API。 -- 基准与研究工具不得成为运行时依赖。 diff --git a/docs/flagquantum_zh/overview/features.md b/docs/flagquantum_zh/overview/features.md deleted file mode 100644 index 6fd33cff5c..0000000000 --- a/docs/flagquantum_zh/overview/features.md +++ /dev/null @@ -1,52 +0,0 @@ -# 特性 - -本页概述 FlagQuantum 提供的能力,每项能力都指向介绍其接口与证据边界的指南页面。 - -## PyTorch 原生的量子训练 - -- `fq.Module` 拥有可训练量子参数,行为与普通 PyTorch 模块一致:`forward()` 返回可自动求导的张量,现有优化器、损失函数与训练循环无需改动即可使用。 -- `fq.train` 对单个模块执行最小的 `zero_grad` / `backward` / `step` 循环,并返回带版本信息的训练结果。 -- 经典层与量子层可以组合在同一个模型中,从同一次 `loss.backward()` 获得梯度。 -- 参数支持平铺或命名分组,检查点由模块本身负责。 - -## 一个程序,多种表示 - -- 本地态向量模拟是默认的精确路径。 -- 矩阵乘积态(MPS)模拟面向规模大、纠缠低的系统。 -- 张量网络执行支持基于收缩的工作流。 -- 跨卡切分的态向量执行与按 rank 拥有的 MPS 执行,把同一个逻辑负载扩展到手多张卡上。 -- 可选的 JAX 内核在同一个 PyTorch 接口之后运行。 - -## 线路、编译与规划 - -- `fq.Circuit` 用简洁的门方法构建程序,FlagQuantum IR 是执行、编译与部署共享的带版本表示。 -- `flagquantum.compiler.optimize` 把与目标无关的规范化改写应用到不动点。 -- `flagquantum.compiler.compile` 面向显式给定的耦合图,只输出满足拓扑的双量子比特操作,并记录其路由决策。 -- `fq.plan` 返回可解释的运行时计划,显式给出阻塞原因且不会静默降级;已有计划会被精确执行而不会重新规划。 - -## 测量、噪声与精度 - -- 命名输出:泡利期望值、精确概率、采样与计数。 -- 梯度既包含原生自动求导,也包含面向 Z/ZZ 哈密顿量的内存受限局部伴随路径。 -- 一个与后端无关的噪声模型同时覆盖精确密度矩阵演化、批量态向量轨迹与 MPS 量子轨迹。 -- 精度是每次执行的显式决定,其中包括在仅有 FP32 的设备上用 Double-Single 软件扩展精度。 - -## 部署与硬件执行 - -- 密封部署包把训练好的参数与目标绑定,并完成目标编译。 -- 分组泡利测量计划为每个逐比特对易分组各生成一个密封包,用于基于采样次数的硬件。 -- 远程算力与厂商目标按名称寻址,任务收据不含凭据,且进程重启后仍可恢复。 - -## 数字孪生与纠错 - -- 以标定为条件的 QPU 数字孪生可预测测量分布,并把预测与可追溯的硬件证据、漂移历史与留出验证连接起来。 -- 重复码内存实验把症状提取、解码、纠正与逻辑结果分析串联起来,服务于容错研究。 - -## 生态与扩展 - -- 与框架无关的适配器支持与 PennyLane、Qiskit、Cirq、CUDA-Q 互转,并可用显式桥接在本地 PennyLane Lightning、Qiskit Aer 与 Cirq 模拟器上执行受支持的线路。 -- 扩展 SDK 支持后端、内核、算子、编译 pass、设备、厂商、测量收集器与规划器扩展,具备清单标识、能力协商与生命周期隔离。 - -## 证据纪律 - -已实现的 API、通过的算例正确性检查、历史测量结果与研究目标,是三件不同的事。每项能力都有等级(发布认证/生产可用/开发证据/实验性),且等级绑定到实际验证过的负载与环境。 diff --git a/docs/flagquantum_zh/overview/overview.md b/docs/flagquantum_zh/overview/overview.md deleted file mode 100644 index aadc713da2..0000000000 --- a/docs/flagquantum_zh/overview/overview.md +++ /dev/null @@ -1,51 +0,0 @@ -# 概览 - -FlagQuantum 是一个基于 PyTorch 构建的分布式、可微量子计算框架。它把量子线路变成可训练模型:同一个程序既可以用 PyTorch 常规优化器训练,也可以用不同表示形式模拟,在负载需要时跨卡切分,并在远程算力或量子硬件上求值。它属于 FlagOS 生态——一个统一的开源 AI 系统软件栈,用来整合多样化的模型、系统与芯片。 - -## 为什么需要 FlagQuantum? - -一个量子程序有三个彼此独立的关注点:线路、测量请求、执行目标。FlagQuantum 让它们保持显式。 - -- 线路用 `fq.Circuit` 构建一次,并以与表示无关的 FlagQuantum IR 保存。 -- 执行在运行时选择——本地态向量、MPS、张量网络、跨卡切分或远程目标——无需改写模型。 -- 训练保持 PyTorch 语义:`fq.Module` 返回可自动求导的张量,`fq.train` 就是调用方自己拥有的优化循环。 - -架构上的不变式是:后端选择可以改变执行方式,但不能改变程序含义,也不能改变结果契约。 - -## 入口 - -| 目标 | 主要接口 | -| --- | --- | -| 构建程序 | `fq.Circuit` | -| 查看其稳定表示 | `fq.CircuitIR` | -| 执行前先规划 | `fq.plan`、`Circuit.runtime_plan` | -| 本地执行 | `fq.run` | -| 定义可训练量子层 | `fq.Module` | -| 训练 | `fq.train` | -| 测量 | `fq.expectation`、`fq.probabilities`、`fq.samples`、`fq.counts` | -| 编译或优化 | `fq.compile`、`flagquantum.compiler.optimize` | -| 为目标打包 | `flagquantum.deployment.create_deployment_package` | -| 远程运行 | `fq.run(target=...)`、`fq.submit`、`fq.restore_job` | - -## 能力成熟度 - -支持范围因后端与负载而异。每项能力都有等级,且等级只适用于实际验证过的范围: - -| 等级 | 含义 | -| --- | --- | -| 发布认证(Release certified) | 有经审计、可复现的证据并通过发布门禁,且没有未解决的发布阻塞。 | -| 生产可用(Production supported) | 有兼容性、运维指引和目标硬件证据的支持路径。 | -| 开发证据(Development evidence) | 可执行且已测试的开发结果,不是生产或通用扩展性结论。 | -| 实验性(Experimental) | 研究性接口,不提供兼容性或生产保证。 | - -稳定的公开 API 不会提升实验性后端的等级,CPU 语义证据也不会提升分布式能力的等级。[参考资料](../reference.md)列出已认证的稳定名称;对外性能数字必须有已入库的审计产物作为依据。 - -## 与 FlagOS 的配合方式 - -FlagQuantum 依赖 PyTorch,而不依赖特定厂商运行时。国产加速器通过 FlagOS 统一多芯片层接入,物理设备探测、厂商运行时以及逻辑设备 `flagos:0` 的映射都由该层负责;FlagQuantum 自身不含任何厂商分支,只记录它收到的路由证据。层边界见[架构](architecture.md),厂商接入路径见[远程执行](../user_guide/remote-execution.md)。 - -## 从哪开始 - -- 参考[安装](../getting_started/install.md),并按[快速开始](../getting_started/quick-start.md)训练一个双量子比特模型。 -- 在不改动模型的前提下,按[模拟模式](../user_guide/simulation-modes.md)选择表示形式。 -- 依赖某条路径之前,先确认它验证到什么程度:[参考资料](../reference.md)与上游的验证范围说明。 diff --git a/docs/flagquantum_zh/reference/api.md b/docs/flagquantum_zh/reference/api.md deleted file mode 100644 index 79563be34c..0000000000 --- a/docs/flagquantum_zh/reference/api.md +++ /dev/null @@ -1,46 +0,0 @@ -# 接口参考 - -FlagQuantum 对外只暴露一套经过整理的 Python 接口:`import flagquantum as fq`。 - -## 接口总览 - -| 任务 | 主要接口 | 结果 | -| --- | --- | --- | -| 构建程序 | `fq.Circuit` | 由 FlagQuantum IR 支撑的电路 | -| 优化程序 | `flagquantum.compiler.optimize` | `fq.CircuitIR` | -| 为选定工具与目标编译 | `fq.compile` | `fq.CircuitIR` | -| 检视执行 | `fq.plan`、`Circuit.runtime_plan` | 可解释的运行时计划 | -| 本地或远程执行 | `fq.run` | `fq.ExecutionResult` | -| 非阻塞投递 | `fq.submit`、`fq.restore_job` | 带状态、结果与取消的作业句柄 | -| 定义可训练量子层 | `fq.Module` | PyTorch 模块 | -| 训练 | `fq.train` | `fq.TrainingResult` | -| 测量 | `fq.expectation`、`fq.probabilities`、`fq.samples`、`fq.counts` | 输出请求 | -| 为目标打包 | `flagquantum.deployment.create_deployment_package` | 密封部署包 | - -## 错误 - -稳定的生命周期类别位于 `flagquantum.errors`:语义输入非法时抛 `ValidationError`,计划过期、被篡改或不兼容时抛 `PlanningError`,所请求能力不可用时抛 `CapabilityError`,执行或训练失败时抛 `ExecutionError`。它们都继承 `FlagQuantumError` 及其兼容的 Python 内建异常,因此窄范围与通用 `except` 都能工作。 - -## 运行时配置 - -`RuntimeConfig` 记录后端、设备、实数与复数精度、JAX 精度、矩阵乘法策略与绘制风格。电路在构建时捕获配置,并把带版本的清单嵌入其 IR 与计划,因此分布式工作进程会重建同一套策略,而不是继承可变进程状态。`runtime_config(...)` 提供上下文局部的临时覆盖。 - -## 测量 - -`fq.expectation`、`fq.probabilities`、`fq.samples`、`fq.counts` 描述要测量什么。Pauli 乘积用 `@`,哈密顿量求和与实数系数用普通算术;采样与计数接受计算基比特或一个不带权重的 Pauli 乘积。结果暴露 `expectation()`、`expectations`、`probabilities`、`samples`、`counts` 与 `measurement(index_or_name)`。 - -## 算子与扩展点 - -- `flagquantum.operators` 报告已注册的门集合与每个门的信息。 -- `flagquantum.ecosystem` 在一个候选稳定的协议与不可变注册表之后,收纳各框架适配器(Qiskit、PennyLane、Cirq、CUDA-Q、Amazon Braket)。 -- `flagquantum.ecosystem.extensions` 是面向后端、编译器、pass、内核、算子、设备、提供方、测量收集器与规划器的冻结前扩展 SDK。 -- `flagquantum.services` 收纳可复用的多步工作流,例如执行与部署预检。 -- `flagquantum.experimental` 不提供任何兼容性保证,分布式训练与动态电路执行目前都在这里。 - -## 稳定性边界 - -- 只有受校验的清单定义稳定接口面;实验性与兼容性导入不会静默扩张它。 -- 兼容导入用于迁移,不隐含稳定。 -- 规划结果描述意图与估算,永远不是运行时或基准证据。 -- 算子与后端支持来自可执行的下沉注册表,而不是文字描述。 -- 文档示例由文档契约测试执行;新增公开名称必须先可导入、经过快照测试并加入清单。 diff --git a/docs/flagquantum_zh/reference/extensions.md b/docs/flagquantum_zh/reference/extensions.md deleted file mode 100644 index 52f9b64737..0000000000 --- a/docs/flagquantum_zh/reference/extensions.md +++ /dev/null @@ -1,54 +0,0 @@ -# 扩展 SDK - -扩展 SDK 允许独立安装的包贡献执行后端、电路编译器、编译 pass、内核、算子、设备、提供方、测量收集器或规划器,同时不成为 FlagQuantum 的运行时依赖。 - -## 位置 - -这套已获批准、尚未冻结的 SDK 契约位于 `flagquantum.ecosystem.extensions`,不新增顶层导出。扩展声明带版本号的清单,在激活前协商能力,并被安装到任务级不可变注册表中。该命名空间迁移到当前位置时未提供兼容层,其子模块导入路径保持等价。 - -## 如何编写扩展 - -包在其入口点组中注册一个零参工厂,并返回身份一致的清单: - -```{code-block} python -from flagquantum.ecosystem.extensions import ExtensionManifest - -ExtensionManifest( - name="qsteed", - version="0.1.0", - kind="compiler", - capabilities=frozenset({"circuit_ir"}), -) -``` - -电路编译器实现协商、启动、编译与关闭。`compile` 接受 FlagQuantum 的 `CircuitIR`、可选的目标映射,并返回 FlagQuantum 的 `CircuitIR`;第三方编译器的对象始终留在插件内部。编译 pass 是更小的变换钩子,交换同样的 IR 边界。 - -## 发现与激活 - -已安装的包只通过显式的、按类型区分的发现调用被识别。发现过程会校验入口点身份与清单是否一致,并把结果加入既有的不可变注册表,而不是引入第二套插件注册表。导入 FlagQuantum 绝不会发现、导入或激活任何扩展。 - -面向用户的集成为显式指定编译器或提供方,绝不依赖隐式选择: - -```{code-block} python -import flagquantum as fq - -result = fq.run(circuit, compiler="qsteed", target="quafu:Baihua", shots=1024) -``` - -## 兼容性生命周期 - -- SDK API 不匹配会在注册阶段失败,并给出升级提示。 -- 单个扩展默认是实验性的;要稳定化需要一致性验证、安全审查、文档与明确的兼容窗口。 -- 弃用需在清单中声明移除版本,并在该窗口内保留上一版契约。 -- 能力协商失败即关闭:缺少 dtype、设备、梯度或语义能力会在激活前产生阻塞项。 - -## 隔离与安全 - -- 注册返回新的不可变注册表,作用域为任务级。扩展绝不会改动顶层导出、核心算子表或其他任务的注册表。 -- 生命周期包装器会转换扩展异常,并在启动失败后尝试清理;正常退出作用域时同样会清理。 -- 原始令牌、口令、API key、密钥与凭据会被拒绝进入扩展配置。提供方必须通过宿主提供的解析器获取凭据,且不得把它们放进清单、错误、测量或序列化负载。 -- 扩展以 Python 进程的权限运行。发现机制不是沙箱:请只安装可信的包。 - -## 一致性验证 - -SDK 提供可复用的后端、提供方与电路编译器检查,覆盖清单与负载序列化、能力声明真实性、PyTorch 梯度、dtype 与设备保持、IR 归属、确定性、错误隔离与清理。通过这些检查是扩展稳定化的前提,但不能替代对该扩展本身的独立资质评估。 diff --git a/docs/flagquantum_zh/user_guide/compiler-and-remote.md b/docs/flagquantum_zh/user_guide/compiler-and-remote.md deleted file mode 100644 index aaf46986ee..0000000000 --- a/docs/flagquantum_zh/user_guide/compiler-and-remote.md +++ /dev/null @@ -1,97 +0,0 @@ -# 编译器与远程目标 - -编译负责变换程序,执行负责运行程序。FlagQuantum 刻意把两者分开,使编译产物可以被检视、密封并投递,而不会对“究竟执行了什么”产生歧义。 - -## 与目标无关的优化 - -```{code-block} python -import flagquantum as fq -import flagquantum.compiler as compiler - -circuit = fq.Circuit(2).h(0).h(0).cx(0, 1) -optimized_ir = compiler.optimize(circuit) - -result = fq.run(optimized_ir, options=fq.ExecutionOptions(mode="auto")) -``` - -`optimize` 返回新的 `CircuitIR`,不修改输入,并把规范重写迭代到不动点。 - -## 面向目标的编译 - -当产出的程序必须遵守某个拓扑时,提供显式的耦合图: - -```{code-block} python -coupling = compiler.CouplingMap.line(circuit.n_qubits) -compiled_ir = compiler.compile( - circuit, - coupling_map=coupling, - routing_strategy="auto", -) -``` - -编译器只产出拓扑合法的两比特操作,并把路由决策记录在编译后 IR 的元数据中。它不选择、也不调用执行后端。 - -## 为提供方编译 - -```{code-block} python -result = fq.run( - circuit, - compiler="qsteed", - target="quafu:Baihua", - shots=1024, -) -counts = result.measurement("counts").value[0] -``` - -该路径完成编译、打包、投递并等待远程结果,同时不改变 `fq.ExecutionResult` 返回类型,也绝不会隐式选择或替换编译器与提供方。 - -可选参数让流程保持可检视: - -- `name` 为部署命名;省略则使用部署默认名,提供方分配的任务 ID 与该展示名彼此独立。 -- `target_qubits` 给出可选的有序“逻辑位到物理位”映射。一旦提供,除非目标快照能证明该映射有效且连通,否则编译会失败;显式映射绝不会被静默替换。 - -## 在硬件上测量哈密顿量 - -Pauli 期望值可以使用同一个入口。电路只编译一次,然后按量子比特逐位对易关系分组,在互不相同的密封任务中测量,且不改变已选定的物理比特映射: - -```{code-block} python -energy = fq.run( - circuit, - outputs=fq.expectation(0.5 * (fq.X(0) @ fq.X(1)) + fq.Z(0)), - compiler="qsteed", - target="quafu:Baihua", - shots=4096, -).expectation() -``` - -采样次数按测量组生效。统计量报告估计量的标准误、分组数、每组采样数与总采样数;来源记录会保存每个提供方任务与部署身份。混合输出与不支持的远程输出会在编译或投递之前失败。 - -## 非阻塞投递与恢复 - -长时远程任务不必阻塞笔记本: - -```{code-block} python -job = fq.submit(fq.Circuit(2).x(0), target="quafu:Baihua", shots=1024) -job.save("quafu-job.json") -print(job.status()) -``` - -规范化状态为 `queued`、`running`、`succeeded`、`failed`、`cancelled` 与 `unknown`;仅表示“编译完成”的提供方状态会映射为 `queued`,未知状态绝不视为成功。`job.result()` 不会轮询,`job.wait(timeout=...)` 才是刻意阻塞,`job.cancel()` 发出取消请求,需再用状态查询确认。 - -回执是不含凭据的 JSON,保存不会覆盖已有文件,恢复任务不会重新投递。重启后需重新配置凭据;若投递时网络响应丢失,请先与提供方核对再重试。 - -## 打包训练好的程序 - -部署包把训练参数、按目标编译的结果与可审计身份绑定,便于稍后保存、签名或投递: - -```{code-block} python -import flagquantum.deployment as deployment - -package = deployment.create_deployment_package( - circuit=trained_circuit, - backend=deployment.CloudBackendProfile.simulator(4), - shots=1024, -) -``` - -预检工具会针对目标校验程序,并对随后可能被投递的那个包做身份校验,过程中不联系提供方、不消耗远程资源。 diff --git a/docs/flagquantum_zh/user_guide/dynamic-circuits.md b/docs/flagquantum_zh/user_guide/dynamic-circuits.md deleted file mode 100644 index 6a4f26d994..0000000000 --- a/docs/flagquantum_zh/user_guide/dynamic-circuits.md +++ /dev/null @@ -1,41 +0,0 @@ -# 动态线路 - -动态线路引入线路中测量与经典条件控制。 - -## 构建并运行 - -```python -import flagquantum as fq -from flagquantum.dynamic import DynamicCircuit - -circuit = DynamicCircuit(2) -circuit.h(0) -circuit.measure(0, classical_bit=0) -circuit.conditional("x", 1, classical_bit=0) - -result = fq.experimental.dynamic.run_dynamic(circuit, shots=128, seed=7) -stable_result = result.to_execution_result() -``` - -`DynamicCircuit` 及其 IR 编码处于“候选稳定”状态,等待 API 负责人批准;执行与后端 -评估仍属实验性。稳定的动态路径将继续返回规范的 `fq.ExecutionResult`,厂商原生状态 -不会被固化进该契约。 - -## 执行策略 - -`run_dynamic(..., strategy="auto")` 对符合条件的、采样次数不少于 32 的负载使用批量 -态向量轨迹;当批处理会超过 `max_batched_bytes`(默认 256 MiB)或输入本身已是批量时, -回退到参考轨迹路径。调用方可以显式请求 `strategy="trajectory"` 或 `"batched"`, -`statistics["gate_execution_strategy"]` 会记录所选路径,便于基准归因。 - -## 后端评估 - -```python -report = fq.experimental.dynamic.assess_dynamic_backend(circuit, backend) -assert report.compatible, report.blockers -``` - -该预检只读,不提交任何任务。本地动态噪声仅限于与已执行门相匹配的单线路比特翻转 -信道,以及在显式测量与最终采样上的独立读出混淆;其他信道、关联读出、设备 profile -的时序噪声以及厂商噪声执行都会失败即拒,而不是静默近似。与厂商无关的一致性测试在 -本地与 Qiskit Aer 上通过,且不声称已在真实动态 QPU 上执行。 diff --git a/docs/flagquantum_zh/user_guide/local-workflows.md b/docs/flagquantum_zh/user_guide/local-workflows.md deleted file mode 100644 index 9aad7ba712..0000000000 --- a/docs/flagquantum_zh/user_guide/local-workflows.md +++ /dev/null @@ -1,88 +0,0 @@ -# 本地工作流 - -本地执行是零配置路径,不需要提供方账号、编译器插件、任务调度器或网络连接。 - -## 模拟 - -```{code-block} python -import flagquantum as fq - -circuit = fq.Circuit(2).h(0).cx(0, 1) -result = fq.run(circuit) -state = result.to_statevector() -``` - -CPU 态向量执行是默认路径。需要时可显式选择本机可控的一块 GPU: - -```{code-block} python -result = fq.run( - circuit, - options=fq.ExecutionOptions(device="cuda:0"), -) -``` - -`ExecutionOptions.device` 描述当前进程可控的设备。`target` 参数保留给外部执行目的地,例如九鼎 GPU 工作区或 Quafu 硬件。 - -## 测量 - -直接请求所需的科学结果,而不是手工检查态向量: - -```{code-block} python -probabilities = fq.run( - circuit, - outputs=fq.probabilities(), -).probabilities - -correlation = fq.run( - circuit, - outputs=fq.expectation(fq.Z(0) @ fq.Z(1)), -).expectation() - -counts = fq.run( - circuit, - outputs=fq.counts(), - shots=1024, -).counts[0] -``` - -概率与期望值默认是精确值。计数与采样需要显式指定采样次数。本地执行会保留批次维度,因此 `counts` 对每个批次项返回一个字典,`[0]` 选取默认的单电路批次。 - -## 选择模拟表示 - -同一份程序不必改写即可在不同表示下运行: - -```{code-block} python -statevector_result = fq.run(circuit, options=fq.ExecutionOptions(mode="statevector")) -mps_result = fq.run(circuit, options=fq.ExecutionOptions(mode="mps")) -tensor_result = fq.run(circuit, options=fq.ExecutionOptions(mode="tensor_network")) -``` - -首次运行建议使用态向量模拟。MPS 适合低纠缠的大规模系统,张量网络适合结构化的收缩负载;每种表示都有各自的支持边界。 - -## 精度 - -每次执行只解析一次精度:`complex64` 隐含 float32 参数,`complex128` 隐含 float64。长时运行或分布式任务请使用显式的运行时配置,使电路、计划与各工作进程保持一致: - -```{code-block} python -from flagquantum.runtime.configuration import RuntimeConfig - -config = RuntimeConfig(device="cuda", complex_dtype="complex128") -circuit = fq.Circuit(4, config=config) -``` - -## 绘制电路 - -```{code-block} python -print(circuit.draw()) # 终端文本形式 -figure, axes = circuit.draw(format="mpl") # 出版级图形 -``` - -## 运行维护中的本地示例 - -```{code-block} shell -python -m examples.local.simulate -python -m examples.local.measure -python -m examples.local.train -``` - -接下来可查看单机示例,了解显式模拟器选择、更大的模型,以及可配置的 CPU/GPU 运行方式。 diff --git a/docs/flagquantum_zh/user_guide/noisy-simulation.md b/docs/flagquantum_zh/user_guide/noisy-simulation.md deleted file mode 100644 index 2a8e6220f2..0000000000 --- a/docs/flagquantum_zh/user_guide/noisy-simulation.md +++ /dev/null @@ -1,83 +0,0 @@ -# 含噪模拟 - -FlagQuantum 用同一个后端中立的 `NoiseModel` 支撑精确密度矩阵演化、批量态向量轨迹与 MPS 量子轨迹。精确路径是小规模系统的正确性基准;轨迹路径报告采样统计量,MPS 路径还额外报告截断数据。 - -## 定义噪声模型 - -```{code-block} python -import flagquantum as fq -import flagquantum.runtime as fqr -import flagquantum.noise as fqn - -circuit = fq.Circuit(2).h(0).cx(0, 1) -noise = ( - fqn.NoiseModel() - .add("h", fqn.thermal_relaxation_channel(t1=50_000, t2=70_000, duration=35)) - .add("cx", fqn.depolarizing_channel(0.01)) - .add_readout(0, fqn.ReadoutError(((0.98, 0.02), (0.07, 0.93)))) -) - -exact = fq.run( - circuit, - noise_model=noise, - options=fq.ExecutionOptions(mode="density_matrix"), - outputs=fq.expectation(fq.Z(0) + fq.Z(1)), -) -print(exact.expectation()) -``` - -## 内置信道 - -| 信道 | 用途 | -| --- | --- | -| `depolarizing_channel` | 均匀去极化噪声 | -| `bit_flip_channel`、`phase_flip_channel` | 离散泡利错误 | -| `amplitude_damping_channel` | 能量弛豫 | -| `thermal_relaxation_channel` | 带门时长的 T1/T2 弛豫 | -| `ReadoutError` | 被测量子比特上的经典读出混淆 | - -器件画像还可以提供门时长与空闲时间噪声,运行时会把它们下沉到真正发生的门与空闲窗口上。面向硬件目标的标定推导噪声模型同样受支持。 - -## 轨迹模拟 - -```{code-block} python -sampled = fqr.run_noisy_mps( - circuit, - noise, - trajectories=4096, - min_trajectories=128, - target_standard_error=1e-3, - seed=42, - retain_trajectories=False, -) -print(sampled.expectation_z_mean) -print(sampled.statistics.standard_error) -print(sampled.converged, sampled.stopped_early) -``` - -自适应停止在单个 rank 上可用;当精确密度矩阵会超出显式内存预算时,稳定的 `fq.run` 入口可以选择含噪 MPS 轨迹,但必须由调用方主动选择近似——否则规划会失败,而不会静默改变程序语义。 - -| 负载 | 路径 | -| --- | --- | -| 小规模系统,要求精确值 | 通过 `fq.run` 使用密度矩阵模式 | -| 较大系统,接受采样统计 | 批量态向量轨迹 | -| 大规模低纠缠系统 | 带截断报告的 MPS 量子轨迹 | -| 超出稠密内存上限的 CPU 计数 | 通过稳定入口选择含噪 MPS 轨迹 | - -## 可复现性 - -`NoiseModel` 携带版本化身份,该身份是计划的一部分。计划往返会重新校验模型负载及其摘要,因此含噪结果可以从保存的计划复现,而不必重新敲一遍噪声定义。轨迹路径接受随机种子,其结果携带判断估计是否收敛所需的采样统计量——报告包含标准误、轨迹或采样数量,以及是否使用了自适应停止。 - -## 连续时间演化 - -时间无关的马尔可夫系统可以用稠密哈密顿量与 Lindblad 塌缩算子在固定时间网格上演化: - -```{code-block} shell -python examples/lindblad_evolution.py -``` - -该路径在 complex64 与 complex128 下均为仅 CPU,网格决定定步长四阶 Runge-Kutta 积分器的精度。 - -## 边界 - -脉冲重叠、串扰、泄漏、厂商标定适配器、分布式自适应停止、批量态向量轨迹与含噪梯度都不在支持范围内。多量子比特的 MPS 信道使用显式的稠密正确性回退,而不是静默近似;每条含噪路径的确切支持范围记录在能力目录中。 diff --git a/docs/flagquantum_zh/user_guide/qec.md b/docs/flagquantum_zh/user_guide/qec.md deleted file mode 100644 index 2619532e64..0000000000 --- a/docs/flagquantum_zh/user_guide/qec.md +++ /dev/null @@ -1,44 +0,0 @@ -# 量子纠错 - -量子纠错把症状提取、译码、纠正与逻辑结果分析连成一条流程。长期目标是面向容错量子计算研究的完整工作流,包括逻辑操作与硬件反馈。 - -QEC 领域持有码、译码语义、探测事件与 Pauli 帧。它组合编译器的控制流、运行时反馈、模拟内核、噪声模型与远程硬件接口。 - -## 从存储实验开始 - -三数据比特重复码存储实验会注入一个错误,并跟踪它经过症状提取、译码与纠正的全过程: - -```{code-block} python -from flagquantum.qec import ErrorEvent, ErrorSchedule, run_repetition_memory_experiment - -result = run_repetition_memory_experiment( - error_schedule=ErrorSchedule((ErrorEvent(round_index=0, wire=1),)), - rounds=3, - shots=16, - seed=0, -) -print(result.logical_error_rate) -``` - -实验会报告症状历史、纠正动作与最终逻辑结果;反馈轨迹把真实比特与观测比特、动作以及帧演化分开记录。 - -## 参考实现覆盖的内容 - -- 固定的重复码画像,支持有界的确定性错误计划。 -- 可按轮替换的轨迹译码策略,动作可以是物理 X 或 Pauli 帧 X。 -- 时间规则:拒绝孤立的读出偏移,数据错误需要后续轮次确认。 -- 按电路位置定义的随机噪声画像:奇偶校验操作后独立比特翻转,读出独立混淆。 - -## 运行与校验 - -```{code-block} shell -python -m pytest tests/qec -q -``` - -针对已知注入错误,检查症状历史、纠正动作与最终逻辑结果。修改译码器时还必须覆盖读出故障与末轮附近的错误。 - -## 边界 - -- 扫描只报告有限采样观测;逻辑抑制与阈值结论需要单独的统计与规模证据。 -- 该时间规则不是最大似然译码,重复的读出故障可能被误判为数据错误。 -- 批量译码反馈、通用信道与码、关联或时序噪声、硬实时提供方控制、梯度与分布式执行仍不在当前能力范围内。 diff --git a/docs/flagquantum_zh/user_guide/qpu-digital-twin.md b/docs/flagquantum_zh/user_guide/qpu-digital-twin.md deleted file mode 100644 index 3341e56441..0000000000 --- a/docs/flagquantum_zh/user_guide/qpu-digital-twin.md +++ /dev/null @@ -1,57 +0,0 @@ -# QPU 数字孪生 - -数字孪生是某个 QPU 的、以标定为条件的模型。它预测你在硬件上会看到的测量分布,并把该预测与证据绑定:验证过哪些电路、涉及哪些物理耦合、多深、哪次标定快照、置信界是多少。 - -## 为任意 QPU 构建孪生 - -从携带器件画像的 `NoiseModel` 出发。执行目标与有序物理映射会成为孪生不可变身份的一部分: - -```{code-block} python -import flagquantum as fq - -twin = fq.twin.from_noise_model( - device_noise_model, - target="your-provider:your-qpu", - qubits=(12, 13), -) -prediction = twin.predict(fq.Circuit(2).h(0).cx(0, 1)) -``` - -框架侧是提供方中立的:任何 QPU 集成只要把自身标定转换为 FlagQuantum 的噪声与器件画像类型即可构建同样的模型。仓库内已包含原生的 Quafu 标定路径。 - -## 预测是什么 - -- 预测覆盖完整的计算基测量分布。一致性度量的是测量分布之间的全变差一致性,而不是量子态保真度。 -- `evidence_report()` 完全离线,不会访问硬件。 -- 验证运行时会在向硬件投递的程序中绑定对应的全量子比特测量;投递本身始终是显式动作。 - -## 按拓扑与深度限定证据 - -`TwinCircuitSupport` 把统计包络收窄到真正验证过的有向物理耦合与最大电路深度。它不会仅因为器件上存在某个耦合就推断支持,并且对未测试的相互作用、反向方向、过大深度或阶数大于 2 的操作一律失败即关闭。 - -来自同一 QPU 且时间相近的重叠证据单元可以组合成连通区域覆盖报告,用于电路映射检查,同时保留有向边与保守上下限。局部统计界绝不会被合并成区域级精度结论。 - -## 与硬件做验证 - -验证流程刻意保持显式且可恢复: - -- 先冻结要验证的电路及其对照对象; -- 逐个显式投递任务并保存回执; -- 把有序结果汇总为一个同时置信的陈述; -- 把“孪生↔QPU 一致性”“无噪参考↔QPU 一致性”“QPU 可重复性”作为彼此独立的量保留。 - -回执是不含凭据的 JSON,恢复过程不会重新投递、轮询或重试;这些动作始终是显式的提供方调用。 - -## 跟踪随时间漂移 - -对同一组固定电路,标定漂移与预测精度可以在按时间排列的快照之间对齐,从而得到可直接绘图的孪生↔QPU 一致性、参考一致性、可重复性与已验证界序列,且不宣称因果、也不代为选择模型升级策略。历史记录以私有权限离线持久化,并拒绝破坏性覆盖。 - -## 前瞻式比较候选模型 - -可以在完全相同的、前瞻收集的硬件计数上比较现有模型与候选模型。结果给出保守的有限采样区间,并返回 `improved`、`degraded` 或 `inconclusive`;它绝不会自动提升模型,没有明确的应用决策就不会有模型被提升。 - -## 边界 - -- 证据始终只针对已声明的电路、操作、映射、物理耦合、深度、标定快照与置信界。 -- 留出集证据只适用于前瞻声明的电路,不构成任意电路的精度结论,也不构成训练泛化结论。 -- 自动标定采集、调度、区域级统计推断、信任策略、模型提升与经发布认证的提供方支持不在框架范围内。 diff --git a/docs/flagquantum_zh/user_guide/runtime-planning.md b/docs/flagquantum_zh/user_guide/runtime-planning.md deleted file mode 100644 index 85b19a188f..0000000000 --- a/docs/flagquantum_zh/user_guide/runtime-planning.md +++ /dev/null @@ -1,55 +0,0 @@ -# 运行时规划 - -计划在真正执行之前解释运行时的意图:它选择表示与执行策略,并显式报告阻塞项与回退,而不是静默改变程序语义。 - -## 查看计划 - -```{code-block} python -import flagquantum as fq - -circuit = ( - fq.Circuit(n_qubits=4) - .h(0) - .cx(0, 1) - .rzz(1, 2, theta=0.2) -) - -plan = circuit.runtime_plan(prefer_jax=True, require_gradients=True) -print(plan.summary()) -``` - -也可以直接使用同一个规划入口: - -```{code-block} python -plan = fq.plan(circuit, options=fq.ExecutionOptions(mode="auto")) -print(plan.summary()["recommended_mode"]) -``` - -## 计划可执行且可复现 - -把 `fq.plan` 的结果直接交给 `fq.run`,即可获得可检视、可复现的执行。给定的计划会被校验并执行,不会重新规划或重新编译;计划身份覆盖规范 IR、解析后的执行语义、编译流水线、所需环境与最终决策: - -```{code-block} python -plan = fq.plan(circuit, options=fq.ExecutionOptions(mode="auto", precision="complex64")) -result = fq.run(plan) -assert result.plan.identity == plan.identity -``` - -JSON 往返会在执行前校验所有指纹与最终身份;已有计划对语义覆盖是封闭的:若同时传入 `options`、`measurements` 或 `noise_model` 会抛出 `TypeError`。环境或 world size 不匹配会在内核启动前失败,而不是静默重新规划或降级。 - -## 规划不是证据 - -规划结果描述意图与估算,永远不是运行时证据、基准证据或可扩展性结论。性能与容量结论必须来自运行时生成的记录,并说明其分布语义。 - -## 失败即关闭 - -- 未知的执行选项在规划前即报错,而不是被忽略。 -- 不支持的结果类型与缺少采样次数在执行前即报错。 -- 在所选路径上无法满足的请求会抛出能力或规划错误,而不是在调用方不知情的情况下切换表示。 -- 分布式执行由执行环境描述;world size 或精度冲突会在最开始就失败。 - -## 相关页面 - -- [本地工作流](local-workflows.md) —— 在本地执行已规划的程序。 -- [编译器与远程目标](compiler-and-remote.md) —— 面向拓扑或提供方进行编译。 -- [使用 PyTorch 训练](training-with-pytorch.md) —— 模块与运行时策略如何配合。 From 3e3de59aeb714256be1dea02cdecc84e4475beeb Mon Sep 17 00:00:00 2001 From: cheng874 Date: Tue, 29 Sep 2026 22:20:35 +0800 Subject: [PATCH 3/3] Draft updates --- docs/conf.py | 14 +++ docs/flagquantum_en/reference/capabilities.md | 2 +- docs/flagquantum_zh/reference/capabilities.md | 2 +- .../getting_started/getting-started.md | 10 +++ .../getting_started/install.md | 87 +++++++++++++++++++ .../getting_started/requirements.md | 26 ++++++ docs/verl_hardware_plugin_en/index.md | 75 ++++++++++++++++ .../overview/features.md | 27 ++++++ .../overview/overview.md | 66 ++++++++++++++ .../references/reference.md | 23 +++++ .../release_notes/release-notes.md | 32 +++++++ .../user_guide/grpo-baseline.md | 61 +++++++++++++ .../user_guide/platforms.md | 84 ++++++++++++++++++ .../user_guide/user-guide.md | 10 +++ .../getting_started/getting-started.md | 10 +++ .../getting_started/install.md | 87 +++++++++++++++++++ .../getting_started/requirements.md | 26 ++++++ docs/verl_hardware_plugin_zh/index.md | 75 ++++++++++++++++ .../overview/features.md | 27 ++++++ .../overview/overview.md | 66 ++++++++++++++ .../references/reference.md | 23 +++++ .../release_notes/release-notes.md | 32 +++++++ .../user_guide/grpo-baseline.md | 61 +++++++++++++ .../user_guide/platforms.md | 84 ++++++++++++++++++ .../user_guide/user-guide.md | 10 +++ 25 files changed, 1018 insertions(+), 2 deletions(-) create mode 100644 docs/verl_hardware_plugin_en/getting_started/getting-started.md create mode 100644 docs/verl_hardware_plugin_en/getting_started/install.md create mode 100644 docs/verl_hardware_plugin_en/getting_started/requirements.md create mode 100644 docs/verl_hardware_plugin_en/index.md create mode 100644 docs/verl_hardware_plugin_en/overview/features.md create mode 100644 docs/verl_hardware_plugin_en/overview/overview.md create mode 100644 docs/verl_hardware_plugin_en/references/reference.md create mode 100644 docs/verl_hardware_plugin_en/release_notes/release-notes.md create mode 100644 docs/verl_hardware_plugin_en/user_guide/grpo-baseline.md create mode 100644 docs/verl_hardware_plugin_en/user_guide/platforms.md create mode 100644 docs/verl_hardware_plugin_en/user_guide/user-guide.md create mode 100644 docs/verl_hardware_plugin_zh/getting_started/getting-started.md create mode 100644 docs/verl_hardware_plugin_zh/getting_started/install.md create mode 100644 docs/verl_hardware_plugin_zh/getting_started/requirements.md create mode 100644 docs/verl_hardware_plugin_zh/index.md create mode 100644 docs/verl_hardware_plugin_zh/overview/features.md create mode 100644 docs/verl_hardware_plugin_zh/overview/overview.md create mode 100644 docs/verl_hardware_plugin_zh/references/reference.md create mode 100644 docs/verl_hardware_plugin_zh/release_notes/release-notes.md create mode 100644 docs/verl_hardware_plugin_zh/user_guide/grpo-baseline.md create mode 100644 docs/verl_hardware_plugin_zh/user_guide/platforms.md create mode 100644 docs/verl_hardware_plugin_zh/user_guide/user-guide.md diff --git a/docs/conf.py b/docs/conf.py index ba9e254a29..60d7294d3c 100644 --- a/docs/conf.py +++ b/docs/conf.py @@ -411,6 +411,13 @@ def get_project(projects): "html_title": "verl-FL Documentation", }, }, + "verl_hardware_plugin_en": { + "use_config_file": False, + "config": { + "project": "verl-hardware-plugin Documentation", + "html_title": "verl-hardware-plugin Documentation", + }, + }, "flagos_robo_en": { "use_config_file": False, "config": { @@ -556,6 +563,13 @@ def get_project(projects): "html_title": "verl-FL 文档中心", }, }, + "verl_hardware_plugin_zh": { + "use_config_file": False, + "config": { + "project": "verl-hardware-plugin 文档中心", + "html_title": "verl-hardware-plugin 文档中心", + }, + }, "flagos_robo_zh": { "use_config_file": False, "config": { diff --git a/docs/flagquantum_en/reference/capabilities.md b/docs/flagquantum_en/reference/capabilities.md index 09631bbc93..e122e33c07 100644 --- a/docs/flagquantum_en/reference/capabilities.md +++ b/docs/flagquantum_en/reference/capabilities.md @@ -54,7 +54,7 @@ evidence. A stable public API does not promote an experimental backend. | --- | --- | --- | | Circuit packaging and cloud deployment | Development evidence | Provider support and credential behaviour vary; no provider is release-certified | | Evidence-qualified QPU digital twins | Development evidence | Agreement is total-variation agreement of measurement distributions, for declared circuits only | -| Interoperability adapter contract | Experimental | Candidate-stable protocol pending API-owner approval | +| Interoperability adapter contract | Experimental | Candidate-stable protocol, not yet approved by the API owners | | Qiskit, PennyLane, Cirq and CUDA-Q interoperability | Experimental | Static conversion at versioned boundaries; external objects never enter runtime or accelerator layers | | PennyLane Lightning, Cirq Simulator and Qiskit Aer bridges | Experimental | One fully bound single-batch circuit, no gradients, noise, dynamic circuits, routing or fallback | | Evidence-based simulator advisor | Experimental | Advice is bound to a checked-in circuit and environment; it never infers performance from size | diff --git a/docs/flagquantum_zh/reference/capabilities.md b/docs/flagquantum_zh/reference/capabilities.md index 585bbabc85..48f072bd40 100644 --- a/docs/flagquantum_zh/reference/capabilities.md +++ b/docs/flagquantum_zh/reference/capabilities.md @@ -50,7 +50,7 @@ FlagQuantum 公开每个能力的成熟度,而不是靠示例去暗示。本 | --- | --- | --- | | 线路打包与云端部署 | 开发证据 | 提供方支持与凭据行为各不相同;没有提供方获得发布认证 | | 带证据约束的 QPU 数字孪生 | 开发证据 | 一致性指测量分布的全变差一致性,且只对所声明的线路成立 | -| 互操作适配器契约 | 实验性 | 协议处于候选稳定级,尚待 API 负责人批准 | +| 互操作适配器契约 | 实验性 | 协议处于候选稳定级,尚未获得 API 负责人批准 | | Qiskit、PennyLane、Cirq 与 CUDA-Q 互操作 | 实验性 | 在版本化边界上做静态转换;外部对象不会进入运行时或加速层 | | PennyLane Lightning、Cirq Simulator 与 Qiskit Aer 桥接 | 实验性 | 仅支持一个完全绑定、单批次的线路,不支持梯度、噪声、动态线路、路由或回退 | | 基于证据的模拟器顾问 | 实验性 | 建议绑定到已入库的线路与环境,绝不按规模推断性能 | diff --git a/docs/verl_hardware_plugin_en/getting_started/getting-started.md b/docs/verl_hardware_plugin_en/getting_started/getting-started.md new file mode 100644 index 0000000000..200621cc6d --- /dev/null +++ b/docs/verl_hardware_plugin_en/getting_started/getting-started.md @@ -0,0 +1,10 @@ +# Getting Started with verl-hardware-plugin + +This section covers the requirements for installing verl-hardware-plugin and guides you through installing it on different hardware platforms. + +```{toctree} +:maxdepth: 2 + +requirements.md +install.md +``` diff --git a/docs/verl_hardware_plugin_en/getting_started/install.md b/docs/verl_hardware_plugin_en/getting_started/install.md new file mode 100644 index 0000000000..c5977a66a6 --- /dev/null +++ b/docs/verl_hardware_plugin_en/getting_started/install.md @@ -0,0 +1,87 @@ +# Installation + +verl-hardware-plugin is installed as a Python package and discovered by verl automatically through the `verl.plugins` entry-points group. + +## Prerequisites + +- Linux +- Python >= 3.10 +- [verl](https://github.com/verl-project/verl) >= 0.7.0 (the plugin registry mechanism is provided by [verl#6086](https://github.com/verl-project/verl/pull/6086); for Iluvatar, verl > 0.8.0 is used) +- The vendor software stack for your target hardware (driver, `torch` extension such as `torch_mlu`, and the communication library such as CNCL/MCCL/IXCCL) +- (Optional, for the FlagOS engine) [FlagCX](https://github.com/flagos-ai/FlagCX) and [FlagGems](https://github.com/flagos-ai/FlagGems) + +## Install from Source + +```bash +# 1. Install verl (see https://verl.readthedocs.io/en/latest/start/install.html) +git clone https://github.com/verl-project/verl +cd verl +pip install --no-build-isolation -e . + +# 2. Install verl-hardware-plugin +git clone https://github.com/verl-project/verl-hardware-plugin.git +cd verl-hardware-plugin +pip install --no-build-isolation -e . +``` + +After installation, no additional configuration in verl is required. When verl starts, it imports all packages registered under the `verl.plugins` group, which triggers registration of all platforms and engines. + +To verify registration: + +```bash +python3 -c "from verl.plugin.platform import get_platform; p = get_platform(); print(f'device: {p.device_name}'); print(f'vendor: {p.vendor_name}'); print(f'available: {p.is_available()}')" +``` + +## Platform Selection + +The platform is auto-detected at startup. You can override it with the `VERL_PLATFORM` environment variable: + +```bash +export VERL_PLATFORM=metax # MetaX +export VERL_PLATFORM=intel # Intel XPU +export VERL_PLATFORM=cambricon # Cambricon MLU +``` + +For the FlagOS engine on NVIDIA, set the engine device and vendor: + +```bash +export VERL_ENGINE_DEVICE=cuda +export VERL_ENGINE_VENDOR=flagos +``` + +## Platform-Specific Setup + +Each hardware platform has its own base image, driver mounts, and environment requirements. Use the vendor-provided container images where available. + +| Platform | Base image / stack | Platform env | Communication | Guide | +|----------|--------------------|--------------|---------------|-------| +| NVIDIA (FlagOS) | `harbor.baai.ac.cn/flagscale/flagscale-rl:dev-cu128-py3.12-*` | `VERL_ENGINE_VENDOR=flagos` | NCCL / FlagCX | [nvidia guide](https://github.com/verl-project/verl-hardware-plugin/tree/main/docs/user_guide_flagos/nvidia) | +| MetaX | MetaX Docker Hub image, e.g. `verl:0.7.1-maca.ai3.5.3.3-torch2.8-py312-ubuntu22.04-amd64` | `VERL_PLATFORM=metax` | MCCL | [metax guide](https://github.com/verl-project/verl-hardware-plugin/tree/main/docs/user_guide_metax) | +| Iluvatar | `harbor.baai.ac.cn/flagos21-base/iluvatarcorex-4.4.0-ubuntu24-py312-base:*` | auto / `iluvatar` | IXCCL | [iluvatar guide](https://github.com/verl-project/verl-hardware-plugin/tree/main/docs/user_guide_iluvatar) | +| Cambricon MLU | Cambricon release image (contact Cambricon), `pytorch_infer` env | auto / `cambricon` | CNCL | [mlu guide](https://github.com/verl-project/verl-hardware-plugin/tree/main/docs/user_guide_mlu) | +| Enflame GCU | Vendor image | auto | ECCL / FlagCX | [enflame guide](https://github.com/verl-project/verl-hardware-plugin/tree/main/docs/user_guide_enflame) | +| Intel XPU | oneAPI environment (`source /opt/intel/oneapi/setvars.sh`) | `VERL_PLATFORM=intel` | xccl (oneCCL) | [xpu guide](https://github.com/verl-project/verl-hardware-plugin/tree/main/docs/user_guide_xpu) | + +For Cambricon MLU, install the additional packages `numpy<2` and `TransferQueue` inside the container. For MetaX and Iluvatar, `mx-smi` / the CoreX stack must be available inside the container for hardware auto-detection. + +## Preparing Data and Models + +The platform guides use Qwen3-0.6B and GSM8K as the reference end-to-end example: + +```bash +# Model +modelscope download --model Qwen/Qwen3-0.6B --local_dir ./Qwen3-0.6B + +# Dataset +mkdir gsm8k && cd gsm8k +wget "https://baai-flagscale.ks3-cn-beijing.ksyuncs.com/rl/datasets/gsm8k/train.parquet" +wget "https://baai-flagscale.ks3-cn-beijing.ksyuncs.com/rl/datasets/gsm8k/test.parquet" +``` + +Start a Ray cluster before launching training: + +```bash +ray start --head --dashboard-host=0.0.0.0 +``` + +See the [User Guide](../user_guide/user-guide.md) for running the GSM8K GRPO baseline and the [References](../references/reference.md) for the environment-variable reference and FAQ. diff --git a/docs/verl_hardware_plugin_en/getting_started/requirements.md b/docs/verl_hardware_plugin_en/getting_started/requirements.md new file mode 100644 index 0000000000..7c3d758f08 --- /dev/null +++ b/docs/verl_hardware_plugin_en/getting_started/requirements.md @@ -0,0 +1,26 @@ +# Requirements + +## Supported Hardware + +| Vendor | Device | Device type | Communication | Status | +|--------|--------|-------------|---------------|--------| +| NVIDIA | CUDA GPUs | `cuda` | NCCL / FlagCX | Supported (FlagOS engine verified) | +| MetaX | C500/C550 series (CUDA-compatible) | `cuda` | MCCL | Supported | +| Iluvatar | BI-V150 (CUDA-compatible) | `cuda` | IXCCL | Supported | +| Cambricon | MLU | `mlu` | CNCL | Reference implementation (tracked in [flagos-ai/community#73](https://github.com/flagos-ai/community/issues/73)) | +| Enflame | GCU | `enflame` | ECCL / FlagCX | Example (requires vendor support) | +| Intel | Data Center GPU Max / Arc | `xpu` | xccl (oneCCL) | Example (requires vendor support) | +| Huawei | Ascend 910B | `npu` | HCCL | Built-in (verl core) | + +## Operating System + +Linux (official). + +## Software + +- Python >= 3.10 +- [verl](https://github.com/verl-project/verl) >= 0.7.0 +- PyTorch (matching the target device and vendor stack) +- For the FlagOS engine: FlagCX (optional, enabled via `USE_FLAGCX`), FlagGems (optional operator acceleration) + +Per-platform software stacks (vendor driver, firmware, `torch` extension such as `torch_mlu`, and communication libraries such as CNCL/MCCL/IXCCL) are provided by the respective hardware vendors. See each platform's installation guide for details. diff --git a/docs/verl_hardware_plugin_en/index.md b/docs/verl_hardware_plugin_en/index.md new file mode 100644 index 0000000000..54a7ff29cc --- /dev/null +++ b/docs/verl_hardware_plugin_en/index.md @@ -0,0 +1,75 @@ +# verl-hardware-plugin Documentation + +verl-hardware-plugin provides multi-chip hardware platform and engine plugins for [verl](https://github.com/verl-project/verl). It is jointly developed by the ByteDance verl team and the [FlagOS](https://github.com/flagos-ai) community, and allows the same RL post-training code to run on NVIDIA, MetaX, Iluvatar, Cambricon MLU, Enflame, and Intel XPU hardware. + +```{button-ref} getting_started/getting-started +:ref-type: myst +:color: primary +:class: sd-btn-lg sd-px-4 sd-py-2 sd-fw-bold + +Getting Started +``` + +::::{grid} 1 2 2 3 +:gutter: 1 1 1 2 + +:::{grid-item-card} {octicon}`browser;1.5em;sd-mr-1` Overview +:link: overview/overview +:link-type: doc + +Understand what verl-hardware-plugin is, its architecture, and the platform/engine registry design. + ++++ +[Learn more »](overview/overview.md) +::: + +:::{grid-item-card} {octicon}`book;1.5em;sd-mr-1` Getting Started +:link: getting_started/getting-started +:link-type: doc + +Requirements and step-by-step installation for supported hardware platforms. + ++++ +[Learn more »](getting_started/getting-started.md) +::: + +:::{grid-item-card} {octicon}`broadcast;1.5em;sd-mr-1` User Guide +:link: user_guide/user-guide +:link-type: doc + +Platform-specific guides and the end-to-end GSM8K GRPO baseline. + ++++ +[Learn more »](user_guide/user-guide.md) +::: + +:::: + +--- + +```{toctree} +:caption: 📑 Release Notes +:maxdepth: 5 +:hidden: + +release_notes/release-notes.md +``` + +```{toctree} +:caption: 📚 Guides +:maxdepth: 5 +:hidden: + +overview/overview.md +overview/features.md +getting_started/getting-started.md +user_guide/user-guide.md +``` + +```{toctree} +:caption: 📖 References +:maxdepth: 5 +:hidden: + +references/reference.md +``` diff --git a/docs/verl_hardware_plugin_en/overview/features.md b/docs/verl_hardware_plugin_en/overview/features.md new file mode 100644 index 0000000000..7c8203ba95 --- /dev/null +++ b/docs/verl_hardware_plugin_en/overview/features.md @@ -0,0 +1,27 @@ +# Features + +## Hardware-Agnostic Platform Abstraction + +Through the unified `PlatformBase` interface, hardware-specific logic for device management, collective communication, memory management, profiling, and rollout environment variables is abstracted into standard methods. A vendor only needs to implement one platform class and register it with `@PlatformRegistry.register` to integrate with verl. + +For CUDA-compatible hardware such as MetaX and Iluvatar, `torch.cuda.is_available()` returns True on multiple chips. The platform layer introduces a `vendor_name` identifier and SMI-based hardware detection (for example `mx-smi` for MetaX and `ixsmi` for Iluvatar) to distinguish the actual hardware during first-time auto-detection and avoid mismatching the NVIDIA engine. + +## Stage-Scoped Environment Manager (FLEnvManager) + +RL post-training has different operator-acceleration and communication requirements in the training stage and the rollout stage. The environment manager manages FlagGems and FlagCX configuration per stage: + +- **FlagGems operator acceleration** — supports independent operator allowlists / blocklists per stage to control which operators take the FlagGems accelerated path, and records operator hits for tuning and troubleshooting. +- **FlagCX unified communication** — the `USE_FLAGCX` switch enables the FlagCX heterogeneous communication library in multi-chip environments; when disabled, it falls back to the device-native backend (such as NCCL). + +## Dedicated FlagOS Engines for FSDP and Megatron + +FlagOS-specific engines are derived from verl's native FSDP and Megatron engines, covering the key roles in RL training: + +- `FSDPFlagOSEngineWithLMHead` / `FSDPFlagOSEngineWithValueHead` — support fsdp / fsdp2 and cover the policy and value models. +- `MegatronFlagOSEngineWithLMHead` — supports large-scale Megatron parallel training. + +The engines automatically inject FlagGems operator acceleration during the `initialize` stage based on the environment configuration, fully transparent to the upper-layer RL algorithm. + +## Zero-Configuration Plugin Discovery + +The plugin is discovered by verl through Python's `entry_points` mechanism. After `pip install`, verl imports the package registered under the `verl.plugins` group, which triggers registration of all platforms and engines. No changes to the verl main framework are required. diff --git a/docs/verl_hardware_plugin_en/overview/overview.md b/docs/verl_hardware_plugin_en/overview/overview.md new file mode 100644 index 0000000000..55df503179 --- /dev/null +++ b/docs/verl_hardware_plugin_en/overview/overview.md @@ -0,0 +1,66 @@ +# verl-hardware-plugin Overview + +verl-hardware-plugin provides **reference implementations** of multi-chip hardware platforms and training engines for [verl](https://github.com/verl-project/verl), the RL post-training framework. It supplies platform abstractions and training engine extensions for non-CUDA accelerators, and serves as a template for hardware vendors to adapt verl to their own devices through a unified plugin interface. + +The repository is jointly developed by the ByteDance verl team and the [FlagOS](https://github.com/flagos-ai) community. + +```{note} +The platforms and engines in this repository are reference implementations. Full production support and maintenance require collaboration with the respective hardware vendors. +``` + +## Relationship to verl and verl-FL + +- **verl** — the upstream RL post-training framework (HybridFlow). The platform/engine registry mechanism is implemented in [verl#6086](https://github.com/verl-project/verl/pull/6086). +- **verl-hardware-plugin** — an out-of-tree plugin discovered by verl through the `verl.plugins` entry-points group. No manual configuration in verl is required after `pip install`. +- **verl-FL** — FlagOS's fork of verl, which uses its own in-tree platform abstraction layer. verl-hardware-plugin targets the plugin mechanism of upstream verl. + +## Architecture + +``` +verl (main framework) + | + +-- entry_points: verl.plugins -> verl_hardware_plugin + | + +-- platforms/ -> @PlatformRegistry.register(platform="vendor_name") + | +-- PlatformFlagOS (device=cuda, vendor=flagos) + | +-- PlatformMetaX (device=cuda, vendor=metax) + | +-- PlatformCUDAIluvatar (device=cuda, vendor=iluvatar) + | +-- PlatformMLU (device=mlu, vendor=cambricon) + | +-- PlatformXPU (device=xpu, vendor=intel) + | +-- PlatformENFLAME (device=enflame) + | + +-- engines/ -> @EngineRegistry.register(device=..., vendor=...) + | +-- fsdp_flagos.py / megatron_flagos.py + | +-- fsdp_metax.py / megatron_metax.py + | +-- fsdp_mlu.py / megatron_mlu.py + | +-- fsdp_iluvatar.py / megatron_iluvatar.py + | +-- fsdp_enflame.py / megatron_enflame.py + | +-- fsdp_xpu.py / megatron_xpu.py + | + +-- utils/ -> FLEnvManager (stage-scoped FlagGems / FlagCX config) +``` + +The plugin integrates with verl through two registries: + +1. **PlatformRegistry** — registers hardware platform abstractions (device management, communication, memory). +2. **EngineRegistry** — registers training engines (hardware-specific variants of FSDP/Megatron). + +Engine lookup uses a two-level key `(device, vendor)`: + +1. Exact match `(device, vendor)` — a vendor-specific engine; +2. Fallback to a device-only key — the base engine for that device type; +3. For CUDA-compatible devices, fallback to the base CUDA engine. + +See [Features](features.md) for details on the platform abstraction, the stage-scoped environment manager, and the dedicated FlagOS engines. + +## Supported Hardware + +| Platform | Device | Communication | Status | +|----------|--------|---------------|--------| +| FlagOS / NVIDIA | NVIDIA GPU (verified) | FlagCX / NCCL | Supported | +| MetaX | MetaX GPUs (CUDA-compatible) | MCCL | Supported | +| Iluvatar | BI-V150 (CUDA-compatible) | IXCCL | Supported | +| Cambricon MLU | MLU | CNCL | Reference implementation (tracked in [flagos-ai/community#73](https://github.com/flagos-ai/community/issues/73)) | +| Enflame GCU | GCU | ECCL / FlagCX | Example (requires vendor support) | +| Intel XPU | Data Center GPU Max / Arc | xccl (oneCCL) | Example (requires vendor support) | +| Huawei Ascend | Ascend 910B | HCCL | Built-in (verl core) | diff --git a/docs/verl_hardware_plugin_en/references/reference.md b/docs/verl_hardware_plugin_en/references/reference.md new file mode 100644 index 0000000000..34293ce8be --- /dev/null +++ b/docs/verl_hardware_plugin_en/references/reference.md @@ -0,0 +1,23 @@ +# References + +## Project Links + +- **Repository**: [verl-project/verl-hardware-plugin](https://github.com/verl-project/verl-hardware-plugin) +- **Upstream verl**: [verl-project/verl](https://github.com/verl-project/verl) +- **Plugin registry mechanism**: [verl#6086](https://github.com/verl-project/verl/pull/6086) +- **FlagOS community**: [flagos-ai/FlagOS](https://github.com/flagos-ai) + +## Related FlagOS Components + +- [FlagCX](https://github.com/flagos-ai/FlagCX) — unified cross-vendor communication library +- [FlagGems](https://github.com/flagos-ai/FlagGems) — Triton-based general-purpose operator library +- [vllm-plugin-FL](https://github.com/flagos-ai/vllm-plugin-FL) — vLLM rollout/inference backend +- [verl-FL](https://github.com/flagos-ai/verl-FL) — FlagOS fork of verl with an in-tree platform abstraction layer + +## Upstream verl Documentation + +For verl features and commands not specific to hardware plugins, refer to the [verl Documentation](https://verl.readthedocs.io/en/latest/index.html). + +## License + +Apache License 2.0. diff --git a/docs/verl_hardware_plugin_en/release_notes/release-notes.md b/docs/verl_hardware_plugin_en/release_notes/release-notes.md new file mode 100644 index 0000000000..174ea58674 --- /dev/null +++ b/docs/verl_hardware_plugin_en/release_notes/release-notes.md @@ -0,0 +1,32 @@ +# Release Notes + +This section includes the verl-hardware-plugin release information. + +## v0.1.0 + +- **Overview** + + Initial release of verl-hardware-plugin, providing multi-chip hardware platform and engine plugins for [verl](https://github.com/verl-project/verl). The package is jointly developed by the ByteDance verl team and the FlagOS community, and is auto-discovered by verl through the `verl.plugins` entry-points group. + +- **Added Features** + + - Hardware platform implementations registered via `@PlatformRegistry.register`: + - FlagOS engine platform on `cuda` (vendor `flagos`, NVIDIA verified). + - MetaX platform on `cuda` (vendor `metax`). + - Iluvatar platform on `cuda` (vendor `iluvatar`, BI-V150). + - Cambricon MLU platform on `mlu` (vendor `cambricon`). + - Enflame GCU platform (vendor `enflame`). + - Intel XPU platform on `xpu` (vendor `intel`). + - FSDP and Megatron engine variants for each supported device/vendor pair, registered via `@EngineRegistry.register`. + - Stage-scoped environment manager (`FLEnvManager`) for per-stage FlagGems operator allowlists/blocklists and FlagCX communication control. + - SMI-based hardware detection for CUDA-compatible devices (`nvidia-smi`, `mx-smi`) to disambiguate vendors during auto-detection. + - Dedicated FlagOS engines (`FSDPFlagOSEngineWithLMHead`/`WithValueHead`, `MegatronFlagOSEngineWithLMHead`) that transparently inject FlagGems acceleration. + - Cambricon CNCL and CNI XL checkpoint engines. + - Per-platform user guides and the GSM8K GRPO acceptance baseline script. + - Validated platforms — the end-to-end GRPO run on GSM8K with Qwen3-0.6B has been completed on MetaX and Iluvatar; the remaining platforms ship as reference implementations. + +- **Requirements** + + - Python >= 3.10 + - verl >= 0.7.0 (plugin registry from [verl#6086](https://github.com/verl-project/verl/pull/6086); Iluvatar uses verl > 0.8.0) + diff --git a/docs/verl_hardware_plugin_en/user_guide/grpo-baseline.md b/docs/verl_hardware_plugin_en/user_guide/grpo-baseline.md new file mode 100644 index 0000000000..940743a9a1 --- /dev/null +++ b/docs/verl_hardware_plugin_en/user_guide/grpo-baseline.md @@ -0,0 +1,61 @@ +# GRPO Acceptance Baseline (GSM8K) + +The standard acceptance test for a new hardware platform adaptation is GRPO training on GSM8K with Qwen3-0.6B. The reference implementation is `scripts/baseline_grpo_gsm8k.sh` in the repository. + +## What the Baseline Validates + +- End-to-end RL post-training (FSDP actor/critic + vLLM rollout) on the target hardware. +- The `critic/rewards/mean` curve should align with the [NVIDIA reference run](https://swanlab.cn/@heavyrain/verl_grpo_gsm8k_math/runs/8h196r8o/chart). +- Platform registration, engine lookup, and (when enabled) FlagCX communication and FlagGems operator acceleration. + +## Running the Baseline + +1. Complete the platform [installation](../getting_started/install.md) and prepare Qwen3-0.6B and the GSM8K `train.parquet` / `test.parquet`. +2. Start Ray: + + ```bash + ray start --head --dashboard-host=0.0.0.0 + ``` + +3. Set the platform environment variables for your hardware (for example `VERL_PLATFORM=metax`, or `VERL_ENGINE_DEVICE=cuda` + `VERL_ENGINE_VENDOR=flagos` for the FlagOS engine). +4. Run the baseline script, adjusting `DATA_DIR` and `MODEL_DIR` to your local paths: + + ```bash + bash scripts/baseline_grpo_gsm8k.sh + ``` + +The script uses these default hyperparameters (all overridable via environment variables): + +| Parameter | Default | +|-----------|---------| +| Model | Qwen3-0.6B | +| `train_batch_size` | 64 | +| `ppo_mini_batch_size` | 16 | +| `max_prompt_length` / `max_response_length` | 1024 / 1024 | +| `rollout_n` | 5 | +| `total_epochs` | 15 | +| Algorithm | GRPO (`algorithm.adv_estimator=grpo`, `use_kl_in_reward=False`) | + +When training starts successfully, logs show platform auto-detection and step-level progress: + +```text +INFO platform_manager.py: Auto-detected platform: metax +INFO platform_manager.py: verl platform initialised: cuda +step:1 - actor/entropy:... - perf/mfu/actor_infer:... - critic/rewards/mean:... +``` + +## FlagOS Engine Environment Variables + +When running with the FlagOS engine (vendor `flagos`), the following variables control stage-scoped acceleration: + +| Variable | Description | Default | +|----------|-------------|---------| +| `VERL_ENGINE_DEVICE` | Device type, e.g. `cuda` | - | +| `VERL_ENGINE_VENDOR` | Vendor identifier, set to `flagos` | - | +| `TRAINING_FL_FLAGGEMS_ENABLE` | Enable FlagGems for training | `0` | +| `TRAINING_FL_FLAGOS_WHITELIST` | Training operator whitelist | (none) | +| `TRAINING_FL_FLAGOS_BLACKLIST` | Training operator blacklist | (none) | +| `USE_FLAGCX` | Enable FlagCX communication | `0` | +| `RAY_ACCEL_ENV_VAR_OVERRIDE_ON_ZERO` | Ray GPU detection override | `1` (set to `0` in the baseline) | + +For vLLM rollout dispatch variables (`VLLM_FL_PREFER`, `VLLM_FL_STRICT`, allow/deny lists, etc.), see the [vllm-plugin-FL dispatch documentation](https://github.com/flagos-ai/vllm-plugin-FL/blob/main/vllm_fl/dispatch/README.md#environment-variables). diff --git a/docs/verl_hardware_plugin_en/user_guide/platforms.md b/docs/verl_hardware_plugin_en/user_guide/platforms.md new file mode 100644 index 0000000000..79095b5fb6 --- /dev/null +++ b/docs/verl_hardware_plugin_en/user_guide/platforms.md @@ -0,0 +1,84 @@ +# Platform Guide + +This page summarizes the hardware platforms verl-hardware-plugin ships implementations for. For complete, up-to-date installation and quick-start steps, use the platform guide in the repository (linked for each platform). + +## Platform Summary + +| Platform | Device type | Vendor | Communication | Visibility env var | Detection | Ray resource | IPC | Repository guide | +|----------|-------------|--------|---------------|---------------------|-----------|--------------|-----|------------------| +| NVIDIA (FlagOS) | `cuda` | `flagos` | NCCL / FlagCX | `CUDA_VISIBLE_DEVICES` | `nvidia-smi` | `GPU` | Yes | [user_guide_flagos/nvidia](https://github.com/verl-project/verl-hardware-plugin/tree/main/docs/user_guide_flagos/nvidia) | +| MetaX | `cuda` | `metax` | NCCL API / MCCL | `CUDA_VISIBLE_DEVICES` | `mx-smi` | `GPU` | Yes | [user_guide_metax](https://github.com/verl-project/verl-hardware-plugin/tree/main/docs/user_guide_metax) | +| Iluvatar | `cuda` | `iluvatar` | NCCL API / IXCCL | `CUDA_VISIBLE_DEVICES` | `ixsmi` | `GPU` | Yes | [user_guide_iluvatar](https://github.com/verl-project/verl-hardware-plugin/tree/main/docs/user_guide_iluvatar) | +| Cambricon MLU | `mlu` | `cambricon` | CNCL | `MLU_VISIBLE_DEVICES` | `torch.mlu` | `MLU` | No | [user_guide_mlu](https://github.com/verl-project/verl-hardware-plugin/tree/main/docs/user_guide_mlu) | +| Enflame GCU | `enflame` | `enflame` | ECCL / FlagCX | - | - | - | - | [user_guide_enflame](https://github.com/verl-project/verl-hardware-plugin/tree/main/docs/user_guide_enflame) | +| Intel XPU | `xpu` | `intel` | xccl (oneCCL) | `ZE_AFFINITY_MASK` | - | - | No | [user_guide_xpu](https://github.com/verl-project/verl-hardware-plugin/tree/main/docs/user_guide_xpu) | +| Huawei Ascend | `npu` | (built-in) | HCCL | - | - | - | - | [verl ascend_tutorial](https://github.com/verl-project/verl/tree/main/docs/ascend_tutorial) | + +For CUDA-compatible devices (MetaX, Iluvatar, NVIDIA), `torch.cuda.is_available()` returns True on all of them. The platform layer uses `vendor_name` and SMI-based detection (`mx-smi` for MetaX, `nvidia-smi` for NVIDIA) at first auto-detection to select the correct engine. + +## Platform-Specific Notes + +### NVIDIA (FlagOS engine) + +The FlagOS engine registers as an engine vendor (`flagos`) on top of the `cuda` device platform, rather than as a standalone platform. + +```bash +export VERL_ENGINE_DEVICE=cuda +export VERL_ENGINE_VENDOR=flagos +export TRAINING_FL_FLAGGEMS_ENABLE=1 +export RAY_ACCEL_ENV_VAR_OVERRIDE_ON_ZERO=0 +export FLAGCX_PATH=/path/to/FlagCX # when using FlagCX +``` + +### MetaX + +```bash +export VERL_PLATFORM=metax +export MACA_MPS_MODE=1 +export MCCL_MAX_NCHANNELS=16 +``` + +Requires `/dev/dri` and `/dev/mxcd` device mounts and `mx-smi` inside the container. Uses MetaX Docker Hub images (for example `verl:0.7.1-maca.ai3.5.3.3-torch2.8-py312-ubuntu22.04-amd64`). + +### Iluvatar (BI-V150) + +Uses the CoreX base image `iluvatarcorex-4.4.0-ubuntu24-py312-base`. Requires verl > 0.8.0 with the plugin registry from [verl#6086](https://github.com/verl-project/verl/pull/6086). + +### Cambricon MLU + +Use the Cambricon release Docker image in the `pytorch_infer` environment, and install `numpy<2` and `TransferQueue`. Start Ray and run verl examples with the recommended `runtime_env.yaml`: + +```yaml +working_dir: ./ +excludes: ["/.git/"] +env_vars: + TORCH_NCCL_AVOID_RECORD_STREAMS: "1" + RAY_ACCEL_ENV_VAR_OVERRIDE_ON_ZERO: "0" +``` + +MLU uses a custom Ray resource so that device assignment stays separate from CUDA GPU scheduling; Ray workers must advertise it: + +```bash +ray start --resources='{"MLU": 8}' +``` + +### Intel XPU + +```bash +export VERL_PLATFORM=intel +source /opt/intel/oneapi/setvars.sh +``` + +### Enflame GCU + +Provided as a reference example; full production support requires collaboration with the vendor. + +## Adding a New Platform + +Vendors can add a new hardware platform by: + +1. Creating a platform class under `verl_hardware_plugin/platforms/` decorated with `@PlatformRegistry.register(platform="vendor_name")`. +2. Creating FSDP and/or Megatron engines under `verl_hardware_plugin/engines/` decorated with `@EngineRegistry.register(device=..., vendor=...)`. +3. Adding the device/engine checks to the plugin's `__init__.py` registration. + +See the [Development Guide](https://github.com/verl-project/verl-hardware-plugin/blob/main/docs/development.md) in the repository for a fully annotated platform template. diff --git a/docs/verl_hardware_plugin_en/user_guide/user-guide.md b/docs/verl_hardware_plugin_en/user_guide/user-guide.md new file mode 100644 index 0000000000..d60d100487 --- /dev/null +++ b/docs/verl_hardware_plugin_en/user_guide/user-guide.md @@ -0,0 +1,10 @@ +# User Guide + +This section provides guidance on running verl RL post-training workloads across hardware platforms with verl-hardware-plugin. + +```{toctree} +:maxdepth: 2 + +grpo-baseline.md +platforms.md +``` diff --git a/docs/verl_hardware_plugin_zh/getting_started/getting-started.md b/docs/verl_hardware_plugin_zh/getting_started/getting-started.md new file mode 100644 index 0000000000..b27d497b02 --- /dev/null +++ b/docs/verl_hardware_plugin_zh/getting_started/getting-started.md @@ -0,0 +1,10 @@ +# verl-hardware-plugin 快速开始 + +本节介绍安装 verl-hardware-plugin 的要求,并指导你在不同硬件平台上完成安装。 + +```{toctree} +:maxdepth: 2 + +requirements.md +install.md +``` diff --git a/docs/verl_hardware_plugin_zh/getting_started/install.md b/docs/verl_hardware_plugin_zh/getting_started/install.md new file mode 100644 index 0000000000..4e0cc8937d --- /dev/null +++ b/docs/verl_hardware_plugin_zh/getting_started/install.md @@ -0,0 +1,87 @@ +# 安装 + +verl-hardware-plugin 以 Python 包形式安装,verl 通过 `verl.plugins` entry-points 组自动发现。 + +## 前置条件 + +- Linux +- Python >= 3.10 +- [verl](https://github.com/verl-project/verl) >= 0.7.0(插件注册机制由 [verl#6086](https://github.com/verl-project/verl/pull/6086) 提供;Iluvatar 使用 verl > 0.8.0) +- 目标硬件的厂商软件栈(驱动、`torch_mlu` 等 torch 扩展,以及 CNCL/MCCL/IXCCL 等通信库) +- (可选,FlagOS 引擎需要)[FlagCX](https://github.com/flagos-ai/FlagCX) 与 [FlagGems](https://github.com/flagos-ai/FlagGems) + +## 从源码安装 + +```bash +# 1. 安装 verl(见 https://verl.readthedocs.io/en/latest/start/install.html) +git clone https://github.com/verl-project/verl +cd verl +pip install --no-build-isolation -e . + +# 2. 安装 verl-hardware-plugin +git clone https://github.com/verl-project/verl-hardware-plugin.git +cd verl-hardware-plugin +pip install --no-build-isolation -e . +``` + +安装后无需在 verl 中做任何额外配置。verl 启动时会导入注册在 `verl.plugins` 组下的所有包,从而触发所有平台与引擎的注册。 + +验证注册: + +```bash +python3 -c "from verl.plugin.platform import get_platform; p = get_platform(); print(f'device: {p.device_name}'); print(f'vendor: {p.vendor_name}'); print(f'available: {p.is_available()}')" +``` + +## 平台选择 + +平台在启动时自动检测。可通过 `VERL_PLATFORM` 环境变量覆盖: + +```bash +export VERL_PLATFORM=metax # 沐曦 MetaX +export VERL_PLATFORM=intel # Intel XPU +export VERL_PLATFORM=cambricon # 寒武纪 MLU +``` + +在 NVIDIA 上使用 FlagOS 引擎时,设置引擎设备与厂商: + +```bash +export VERL_ENGINE_DEVICE=cuda +export VERL_ENGINE_VENDOR=flagos +``` + +## 各平台特定设置 + +每个硬件平台都有各自的基础镜像、驱动挂载与环境要求。尽可能使用厂商提供的容器镜像。 + +| 平台 | 基础镜像 / 软件栈 | 平台环境变量 | 通信 | 指南 | +|------|-------------------|--------------|------|------| +| NVIDIA(FlagOS) | `harbor.baai.ac.cn/flagscale/flagscale-rl:dev-cu128-py3.12-*` | `VERL_ENGINE_VENDOR=flagos` | NCCL / FlagCX | [nvidia 指南](https://github.com/verl-project/verl-hardware-plugin/tree/main/docs/user_guide_flagos/nvidia) | +| 沐曦 MetaX | MetaX Docker Hub 镜像,如 `verl:0.7.1-maca.ai3.5.3.3-torch2.8-py312-ubuntu22.04-amd64` | `VERL_PLATFORM=metax` | MCCL | [metax 指南](https://github.com/verl-project/verl-hardware-plugin/tree/main/docs/user_guide_metax) | +| 天数智芯 Iluvatar | `harbor.baai.ac.cn/flagos21-base/iluvatarcorex-4.4.0-ubuntu24-py312-base:*` | 自动 / `iluvatar` | IXCCL | [iluvatar 指南](https://github.com/verl-project/verl-hardware-plugin/tree/main/docs/user_guide_iluvatar) | +| 寒武纪 MLU | 寒武纪 release 镜像(联系寒武纪获取),`pytorch_infer` 环境 | 自动 / `cambricon` | CNCL | [mlu 指南](https://github.com/verl-project/verl-hardware-plugin/tree/main/docs/user_guide_mlu) | +| 燧原 Enflame GCU | 厂商镜像 | 自动 | ECCL / FlagCX | [enflame 指南](https://github.com/verl-project/verl-hardware-plugin/tree/main/docs/user_guide_enflame) | +| Intel XPU | oneAPI 环境(`source /opt/intel/oneapi/setvars.sh`) | `VERL_PLATFORM=intel` | xccl (oneCCL) | [xpu 指南](https://github.com/verl-project/verl-hardware-plugin/tree/main/docs/user_guide_xpu) | + +寒武纪 MLU 需在容器内额外安装 `numpy<2` 与 `TransferQueue`。沐曦与天数智芯容器内需有 `mx-smi` / CoreX 软件栈以供硬件自动检测。 + +## 准备数据与模型 + +各平台指南以 Qwen3-0.6B 与 GSM8K 作为端到端参考示例: + +```bash +# 模型 +modelscope download --model Qwen/Qwen3-0.6B --local_dir ./Qwen3-0.6B + +# 数据集 +mkdir gsm8k && cd gsm8k +wget "https://baai-flagscale.ks3-cn-beijing.ksyuncs.com/rl/datasets/gsm8k/train.parquet" +wget "https://baai-flagscale.ks3-cn-beijing.ksyuncs.com/rl/datasets/gsm8k/test.parquet" +``` + +启动训练前先启动 Ray 集群: + +```bash +ray start --head --dashboard-host=0.0.0.0 +``` + +运行 GSM8K GRPO 基线见[用户指南](../user_guide/user-guide.md),环境变量参考与 FAQ 见[参考](../references/reference.md)。 diff --git a/docs/verl_hardware_plugin_zh/getting_started/requirements.md b/docs/verl_hardware_plugin_zh/getting_started/requirements.md new file mode 100644 index 0000000000..bb904ae33e --- /dev/null +++ b/docs/verl_hardware_plugin_zh/getting_started/requirements.md @@ -0,0 +1,26 @@ +# 要求 + +## 支持的硬件 + +| 厂商 | 设备 | 设备类型 | 通信 | 状态 | +|------|------|----------|------|------| +| NVIDIA | CUDA GPU | `cuda` | NCCL / FlagCX | 支持(FlagOS 引擎已验证) | +| 沐曦 MetaX | C500/C550 系列(CUDA 兼容) | `cuda` | MCCL | 支持 | +| 天数智芯 Iluvatar | BI-V150(CUDA 兼容) | `cuda` | IXCCL | 支持 | +| 寒武纪 | MLU | `mlu` | CNCL | 参考实现(跟踪于 [flagos-ai/community#73](https://github.com/flagos-ai/community/issues/73)) | +| 燧原 Enflame | GCU | `enflame` | ECCL / FlagCX | 示例(需厂商支持) | +| Intel | Data Center GPU Max / Arc | `xpu` | xccl (oneCCL) | 示例(需厂商支持) | +| 华为 | Ascend 910B | `npu` | HCCL | 内建(verl 核心) | + +## 操作系统 + +Linux(官方支持)。 + +## 软件 + +- Python >= 3.10 +- [verl](https://github.com/verl-project/verl) >= 0.7.0 +- PyTorch(与目标设备及厂商栈匹配) +- 对于 FlagOS 引擎:FlagCX(可选,通过 `USE_FLAGCX` 启用)、FlagGems(可选算子加速) + +各平台的软件栈(厂商驱动、固件、`torch_mlu` 等 torch 扩展,以及 CNCL/MCCL/IXCCL 等通信库)由相应硬件厂商提供。细节见各平台安装指南。 diff --git a/docs/verl_hardware_plugin_zh/index.md b/docs/verl_hardware_plugin_zh/index.md new file mode 100644 index 0000000000..c41dee4d7f --- /dev/null +++ b/docs/verl_hardware_plugin_zh/index.md @@ -0,0 +1,75 @@ +# verl-hardware-plugin 文档 + +verl-hardware-plugin 为 [verl](https://github.com/verl-project/verl) 提供多芯片硬件平台与训练引擎插件。它由字节跳动 verl 团队与 [FlagOS](https://github.com/flagos-ai) 社区联合开发,使同一份 RL 后训练代码能够运行在 NVIDIA、沐曦 MetaX、天数智芯 Iluvatar、寒武纪 MLU、燧原 Enflame、Intel XPU 等硬件上。 + +```{button-ref} getting_started/getting-started +:ref-type: myst +:color: primary +:class: sd-btn-lg sd-px-4 sd-py-2 sd-fw-bold + +快速开始 +``` + +::::{grid} 1 2 2 3 +:gutter: 1 1 1 2 + +:::{grid-item-card} {octicon}`browser;1.5em;sd-mr-1` 概述 +:link: overview/overview +:link-type: doc + +了解 verl-hardware-plugin 是什么、其架构以及平台/引擎注册表设计。 + ++++ +[了解更多 »](overview/overview.md) +::: + +:::{grid-item-card} {octicon}`book;1.5em;sd-mr-1` 快速开始 +:link: getting_started/getting-started +:link-type: doc + +支持硬件平台的安装要求与分步说明。 + ++++ +[了解更多 »](getting_started/getting-started.md) +::: + +:::{grid-item-card} {octicon}`broadcast;1.5em;sd-mr-1` 用户指南 +:link: user_guide/user-guide +:link-type: doc + +各平台指南与端到端 GSM8K GRPO 基线。 + ++++ +[了解更多 »](user_guide/user-guide.md) +::: + +:::: + +--- + +```{toctree} +:caption: 📑 发布说明 +:maxdepth: 5 +:hidden: + +release_notes/release-notes.md +``` + +```{toctree} +:caption: 📚 指南 +:maxdepth: 5 +:hidden: + +overview/overview.md +overview/features.md +getting_started/getting-started.md +user_guide/user-guide.md +``` + +```{toctree} +:caption: 📖 参考 +:maxdepth: 5 +:hidden: + +references/reference.md +``` diff --git a/docs/verl_hardware_plugin_zh/overview/features.md b/docs/verl_hardware_plugin_zh/overview/features.md new file mode 100644 index 0000000000..4f60d43b01 --- /dev/null +++ b/docs/verl_hardware_plugin_zh/overview/features.md @@ -0,0 +1,27 @@ +# 特性 + +## 硬件无关的平台抽象层 + +通过统一的 `PlatformBase` 接口,将设备管理、集合通信、内存管理、profiler、rollout 环境变量等硬件相关逻辑抽象为标准方法。厂商只需实现一个平台类并用 `@PlatformRegistry.register` 注册,即可接入 verl。 + +对于沐曦、天数智芯等 CUDA 兼容硬件,`torch.cuda.is_available()` 在多种芯片上均返回 True。平台层引入 `vendor_name` 标识与基于 SMI 命令的硬件探测(如沐曦的 `mx-smi`、天数智芯的 `ixsmi`),在首次自动检测时区分实际硬件,避免误匹配到 NVIDIA 引擎。 + +## 按阶段配置的环境管理器(FLEnvManager) + +RL 后训练在训练阶段与推理(rollout)阶段对算子加速和通信后端的需求不同。环境管理器按阶段(training / rollout)分别管理 FlagGems 与 FlagCX 配置: + +- **FlagGems 算子加速** —— 支持按阶段独立配置算子**白名单 / 黑名单**,精细控制哪些算子走 FlagGems 加速路径,并支持算子命中记录,便于调优与问题定位。 +- **FlagCX 统一通信** —— 通过 `USE_FLAGCX` 开关启用 FlagCX 异构通信库,在多芯片环境下提供统一集合通信;未启用时回退到设备原生后端(如 NCCL)。 + +## 针对 FSDP 和 Megatron 的专用 FlagOS 引擎 + +在 verl 原生 FSDP 与 Megatron 引擎基础上派生 FlagOS 专用引擎,覆盖 RL 训练的关键角色: + +- `FSDPFlagOSEngineWithLMHead` / `FSDPFlagOSEngineWithValueHead` —— 支持 fsdp / fsdp2,覆盖策略模型与价值模型。 +- `MegatronFlagOSEngineWithLMHead` —— 支持 Megatron 大规模并行训练。 + +引擎在 `initialize` 阶段依据环境配置自动注入 FlagGems 算子加速,对上层 RL 算法完全透明。 + +## 零配置插件发现 + +插件通过 Python `entry_points` 机制被 verl 发现。`pip install` 后,verl 会导入注册在 `verl.plugins` 组下的包,从而触发所有平台与引擎的注册,无需改动 verl 主框架。 diff --git a/docs/verl_hardware_plugin_zh/overview/overview.md b/docs/verl_hardware_plugin_zh/overview/overview.md new file mode 100644 index 0000000000..50615bef7f --- /dev/null +++ b/docs/verl_hardware_plugin_zh/overview/overview.md @@ -0,0 +1,66 @@ +# verl-hardware-plugin 概述 + +verl-hardware-plugin 为 [verl](https://github.com/verl-project/verl) RL 后训练框架提供多芯片硬件平台与训练引擎的**参考实现**。它为非 CUDA 加速器提供平台抽象与训练引擎扩展,并作为硬件厂商通过统一插件接口将 verl 适配到自有设备的模板与示例。 + +本仓库由字节跳动 verl 团队与 [FlagOS](https://github.com/flagos-ai) 社区联合开发。 + +```{note} +本仓库中的平台与引擎均为参考实现。完整的生产级支持与维护需要与相应硬件厂商协作。 +``` + +## 与 verl、verl-FL 的关系 + +- **verl** —— 上游 RL 后训练框架(HybridFlow)。平台/引擎注册机制在 [verl#6086](https://github.com/verl-project/verl/pull/6086) 中实现。 +- **verl-hardware-plugin** —— 一个 out-of-tree 插件,verl 通过 `verl.plugins` entry-points 组自动发现。`pip install` 后无需在 verl 中做任何手动配置。 +- **verl-FL** —— FlagOS 的 verl 分支,使用自带的树内平台抽象层。verl-hardware-plugin 面向的是上游 verl 的插件机制。 + +## 架构 + +``` +verl(主框架) + | + +-- entry_points: verl.plugins -> verl_hardware_plugin + | + +-- platforms/ -> @PlatformRegistry.register(platform="vendor_name") + | +-- PlatformFlagOS (device=cuda, vendor=flagos) + | +-- PlatformMetaX (device=cuda, vendor=metax) + | +-- PlatformCUDAIluvatar (device=cuda, vendor=iluvatar) + | +-- PlatformMLU (device=mlu, vendor=cambricon) + | +-- PlatformXPU (device=xpu, vendor=intel) + | +-- PlatformENFLAME (device=enflame) + | + +-- engines/ -> @EngineRegistry.register(device=..., vendor=...) + | +-- fsdp_flagos.py / megatron_flagos.py + | +-- fsdp_metax.py / megatron_metax.py + | +-- fsdp_mlu.py / megatron_mlu.py + | +-- fsdp_iluvatar.py / megatron_iluvatar.py + | +-- fsdp_enflame.py / megatron_enflame.py + | +-- fsdp_xpu.py / megatron_xpu.py + | + +-- utils/ -> FLEnvManager(按阶段配置 FlagGems / FlagCX) +``` + +插件通过两个注册表与 verl 集成: + +1. **PlatformRegistry** —— 注册硬件平台抽象(设备管理、通信、内存)。 +2. **EngineRegistry** —— 注册训练引擎(FSDP/Megatron 的硬件特定变体)。 + +引擎查找采用 `(device, vendor)` 二级键: + +1. 精确匹配 `(device, vendor)` —— 厂商专用引擎; +2. 回退到仅设备键 —— 该设备类型的基础引擎; +3. 对于 CUDA 兼容设备,再回退到基础 CUDA 引擎。 + +平台抽象、按阶段环境管理器与专用 FlagOS 引擎的细节见[特性](features.md)。 + +## 支持的硬件 + +| 平台 | 设备 | 通信 | 状态 | +|------|------|------|------| +| FlagOS / NVIDIA | NVIDIA GPU(已验证) | FlagCX / NCCL | 支持 | +| 沐曦 MetaX | MetaX GPU(CUDA 兼容) | MCCL | 支持 | +| 天数智芯 Iluvatar | BI-V150(CUDA 兼容) | IXCCL | 支持 | +| 寒武纪 Cambricon MLU | MLU | CNCL | 参考实现(跟踪于 [flagos-ai/community#73](https://github.com/flagos-ai/community/issues/73)) | +| 燧原 Enflame GCU | GCU | ECCL / FlagCX | 示例(需厂商支持) | +| Intel XPU | Data Center GPU Max / Arc | xccl (oneCCL) | 示例(需厂商支持) | +| 华为昇腾 | Ascend 910B | HCCL | 内建(verl 核心) | diff --git a/docs/verl_hardware_plugin_zh/references/reference.md b/docs/verl_hardware_plugin_zh/references/reference.md new file mode 100644 index 0000000000..756e0f7f8b --- /dev/null +++ b/docs/verl_hardware_plugin_zh/references/reference.md @@ -0,0 +1,23 @@ +# 参考 + +## 项目链接 + +- **仓库**:[verl-project/verl-hardware-plugin](https://github.com/verl-project/verl-hardware-plugin) +- **上游 verl**:[verl-project/verl](https://github.com/verl-project/verl) +- **插件注册机制**:[verl#6086](https://github.com/verl-project/verl/pull/6086) +- **FlagOS 社区**:[flagos-ai/FlagOS](https://github.com/flagos-ai) + +## 相关 FlagOS 组件 + +- [FlagCX](https://github.com/flagos-ai/FlagCX) —— 统一跨厂商通信库 +- [FlagGems](https://github.com/flagos-ai/FlagGems) —— 基于 Triton 的通用算子库 +- [vllm-plugin-FL](https://github.com/flagos-ai/vllm-plugin-FL) —— vLLM rollout/推理后端 +- [verl-FL](https://github.com/flagos-ai/verl-FL) —— FlagOS 的 verl 分支,自带树内平台抽象层 + +## 上游 verl 文档 + +非硬件插件特有的 verl 特性与命令,请参考 [verl 文档](https://verl.readthedocs.io/en/latest/index.html)。 + +## 许可证 + +Apache License 2.0。 diff --git a/docs/verl_hardware_plugin_zh/release_notes/release-notes.md b/docs/verl_hardware_plugin_zh/release_notes/release-notes.md new file mode 100644 index 0000000000..c46145e267 --- /dev/null +++ b/docs/verl_hardware_plugin_zh/release_notes/release-notes.md @@ -0,0 +1,32 @@ +# 发布说明 + +本节包含 verl-hardware-plugin 的发布信息。 + +## v0.1.0 + +- **概述** + + verl-hardware-plugin 首次发布,为 [verl](https://github.com/verl-project/verl) 提供多芯片硬件平台与训练引擎插件。该包由字节跳动 verl 团队与 FlagOS 社区联合开发,verl 通过 `verl.plugins` entry-points 组自动发现。 + +- **新增功能** + + - 通过 `@PlatformRegistry.register` 注册的硬件平台实现: + - `cuda` 上的 FlagOS 引擎平台(vendor `flagos`,NVIDIA 已验证)。 + - `cuda` 上的沐曦 MetaX 平台(vendor `metax`)。 + - `cuda` 上的天数智芯 Iluvatar 平台(vendor `iluvatar`,BI-V150)。 + - `mlu` 上的寒武纪 MLU 平台(vendor `cambricon`)。 + - 燧原 Enflame GCU 平台(vendor `enflame`)。 + - `xpu` 上的 Intel XPU 平台(vendor `intel`)。 + - 为每个支持的 device/vendor 对提供 FSDP 与 Megatron 引擎变体,通过 `@EngineRegistry.register` 注册。 + - 按阶段环境管理器(`FLEnvManager`),支持分阶段的 FlagGems 算子白名单/黑名单与 FlagCX 通信控制。 + - CUDA 兼容设备的基于 SMI 的硬件探测(`nvidia-smi`、`mx-smi`),用于自动检测时区分厂商。 + - 专用 FlagOS 引擎(`FSDPFlagOSEngineWithLMHead`/`WithValueHead`、`MegatronFlagOSEngineWithLMHead`),透明注入 FlagGems 加速。 + - 寒武纪 CNCL 与 CNI XL checkpoint 引擎。 + - 各平台用户指南与 GSM8K GRPO 验收基线脚本。 + - 已验证平台 —— Qwen3-0.6B 在 GSM8K 上的端到端 GRPO 训练已在沐曦 MetaX 与天数智芯 Iluvatar 完成;其余平台以参考实现形式提供。 + +- **要求** + + - Python >= 3.10 + - verl >= 0.7.0(插件注册机制来自 [verl#6086](https://github.com/verl-project/verl/pull/6086);Iluvatar 使用 verl > 0.8.0) + diff --git a/docs/verl_hardware_plugin_zh/user_guide/grpo-baseline.md b/docs/verl_hardware_plugin_zh/user_guide/grpo-baseline.md new file mode 100644 index 0000000000..88da979880 --- /dev/null +++ b/docs/verl_hardware_plugin_zh/user_guide/grpo-baseline.md @@ -0,0 +1,61 @@ +# GRPO 验收基线(GSM8K) + +新硬件平台适配的标准验收测试是使用 Qwen3-0.6B 在 GSM8K 上进行 GRPO 训练。参考实现为仓库中的 `scripts/baseline_grpo_gsm8k.sh`。 + +## 基线验证内容 + +- 在目标硬件上的端到端 RL 后训练(FSDP actor/critic + vLLM rollout)。 +- `critic/rewards/mean` 曲线应与 [NVIDIA 参考运行](https://swanlab.cn/@heavyrain/verl_grpo_gsm8k_math/runs/8h196r8o/chart)对齐。 +- 平台注册、引擎查找,以及(启用时)FlagCX 通信与 FlagGems 算子加速。 + +## 运行基线 + +1. 完成平台[安装](../getting_started/install.md),并准备 Qwen3-0.6B 与 GSM8K 的 `train.parquet` / `test.parquet`。 +2. 启动 Ray: + + ```bash + ray start --head --dashboard-host=0.0.0.0 + ``` + +3. 为你的硬件设置平台环境变量(例如 `VERL_PLATFORM=metax`,FlagOS 引擎用 `VERL_ENGINE_DEVICE=cuda` + `VERL_ENGINE_VENDOR=flagos`)。 +4. 运行基线脚本,将 `DATA_DIR` 与 `MODEL_DIR` 调整为本地路径: + + ```bash + bash scripts/baseline_grpo_gsm8k.sh + ``` + +脚本使用以下默认超参数(均可通过环境变量覆盖): + +| 参数 | 默认值 | +|------|--------| +| 模型 | Qwen3-0.6B | +| `train_batch_size` | 64 | +| `ppo_mini_batch_size` | 16 | +| `max_prompt_length` / `max_response_length` | 1024 / 1024 | +| `rollout_n` | 5 | +| `total_epochs` | 15 | +| 算法 | GRPO(`algorithm.adv_estimator=grpo`,`use_kl_in_reward=False`) | + +训练成功启动时,日志会显示平台自动检测与 step 级进度: + +```text +INFO platform_manager.py: Auto-detected platform: metax +INFO platform_manager.py: verl platform initialised: cuda +step:1 - actor/entropy:... - perf/mfu/actor_infer:... - critic/rewards/mean:... +``` + +## FlagOS 引擎环境变量 + +使用 FlagOS 引擎(vendor `flagos`)时,以下变量控制按阶段加速: + +| 变量 | 说明 | 默认值 | +|------|------|--------| +| `VERL_ENGINE_DEVICE` | 设备类型,如 `cuda` | - | +| `VERL_ENGINE_VENDOR` | 厂商标识,设为 `flagos` | - | +| `TRAINING_FL_FLAGGEMS_ENABLE` | 训练阶段启用 FlagGems | `0` | +| `TRAINING_FL_FLAGOS_WHITELIST` | 训练算子白名单 | (无) | +| `TRAINING_FL_FLAGOS_BLACKLIST` | 训练算子黑名单 | (无) | +| `USE_FLAGCX` | 启用 FlagCX 通信 | `0` | +| `RAY_ACCEL_ENV_VAR_OVERRIDE_ON_ZERO` | Ray GPU 检测覆盖 | `1`(基线中设为 `0`) | + +vLLM rollout 调度变量(`VLLM_FL_PREFER`、`VLLM_FL_STRICT`、允许/拒绝列表等)见 [vllm-plugin-FL 调度文档](https://github.com/flagos-ai/vllm-plugin-FL/blob/main/vllm_fl/dispatch/README.md#environment-variables)。 diff --git a/docs/verl_hardware_plugin_zh/user_guide/platforms.md b/docs/verl_hardware_plugin_zh/user_guide/platforms.md new file mode 100644 index 0000000000..dca3543305 --- /dev/null +++ b/docs/verl_hardware_plugin_zh/user_guide/platforms.md @@ -0,0 +1,84 @@ +# 平台指南 + +本页汇总 verl-hardware-plugin 已提供实现的硬件平台。完整、最新的安装与快速开始步骤请使用各平台对应的仓库指南(每平台均附链接)。 + +## 平台一览 + +| 平台 | 设备类型 | 厂商标识 | 通信 | 设备可见性变量 | 硬件探测 | Ray 资源 | IPC | 仓库指南 | +|------|----------|----------|------|----------------|----------|----------|-----|----------| +| NVIDIA(FlagOS) | `cuda` | `flagos` | NCCL / FlagCX | `CUDA_VISIBLE_DEVICES` | `nvidia-smi` | `GPU` | 是 | [user_guide_flagos/nvidia](https://github.com/verl-project/verl-hardware-plugin/tree/main/docs/user_guide_flagos/nvidia) | +| 沐曦 MetaX | `cuda` | `metax` | NCCL API / MCCL | `CUDA_VISIBLE_DEVICES` | `mx-smi` | `GPU` | 是 | [user_guide_metax](https://github.com/verl-project/verl-hardware-plugin/tree/main/docs/user_guide_metax) | +| 天数智芯 Iluvatar | `cuda` | `iluvatar` | NCCL API / IXCCL | `CUDA_VISIBLE_DEVICES` | `ixsmi` | `GPU` | 是 | [user_guide_iluvatar](https://github.com/verl-project/verl-hardware-plugin/tree/main/docs/user_guide_iluvatar) | +| 寒武纪 MLU | `mlu` | `cambricon` | CNCL | `MLU_VISIBLE_DEVICES` | `torch.mlu` | `MLU` | 否 | [user_guide_mlu](https://github.com/verl-project/verl-hardware-plugin/tree/main/docs/user_guide_mlu) | +| 燧原 Enflame GCU | `enflame` | `enflame` | ECCL / FlagCX | - | - | - | - | [user_guide_enflame](https://github.com/verl-project/verl-hardware-plugin/tree/main/docs/user_guide_enflame) | +| Intel XPU | `xpu` | `intel` | xccl (oneCCL) | `ZE_AFFINITY_MASK` | - | - | 否 | [user_guide_xpu](https://github.com/verl-project/verl-hardware-plugin/tree/main/docs/user_guide_xpu) | +| 华为昇腾 | `npu` | (内建) | HCCL | - | - | - | - | [verl ascend_tutorial](https://github.com/verl-project/verl/tree/main/docs/ascend_tutorial) | + +对于 CUDA 兼容设备(沐曦、天数智芯、NVIDIA),`torch.cuda.is_available()` 在所有设备上均返回 True。平台层在首次自动检测时使用 `vendor_name` 与基于 SMI 的探测(沐曦用 `mx-smi`、NVIDIA 用 `nvidia-smi`)来选择正确引擎。 + +## 各平台注意事项 + +### NVIDIA(FlagOS 引擎) + +FlagOS 引擎注册为 `cuda` 设备平台之上的引擎厂商(`flagos`),而非独立平台。 + +```bash +export VERL_ENGINE_DEVICE=cuda +export VERL_ENGINE_VENDOR=flagos +export TRAINING_FL_FLAGGEMS_ENABLE=1 +export RAY_ACCEL_ENV_VAR_OVERRIDE_ON_ZERO=0 +export FLAGCX_PATH=/path/to/FlagCX # 使用 FlagCX 时 +``` + +### 沐曦 MetaX + +```bash +export VERL_PLATFORM=metax +export MACA_MPS_MODE=1 +export MCCL_MAX_NCHANNELS=16 +``` + +需要挂载 `/dev/dri` 与 `/dev/mxcd` 设备,容器内需要有 `mx-smi`。使用 MetaX Docker Hub 镜像(如 `verl:0.7.1-maca.ai3.5.3.3-torch2.8-py312-ubuntu22.04-amd64`)。 + +### 天数智芯 Iluvatar(BI-V150) + +使用 CoreX 基础镜像 `iluvatarcorex-4.4.0-ubuntu24-py312-base`。需要 verl > 0.8.0 以及 [verl#6086](https://github.com/verl-project/verl/pull/6086) 提供的插件注册机制。 + +### 寒武纪 MLU + +在 `pytorch_infer` 环境中使用寒武纪 release Docker 镜像,并安装 `numpy<2` 与 `TransferQueue`。启动 Ray 后使用推荐的 `runtime_env.yaml` 运行 verl 示例: + +```yaml +working_dir: ./ +excludes: ["/.git/"] +env_vars: + TORCH_NCCL_AVOID_RECORD_STREAMS: "1" + RAY_ACCEL_ENV_VAR_OVERRIDE_ON_ZERO: "0" +``` + +MLU 使用自定义 Ray 资源,以便设备分配与 CUDA GPU 调度相互独立,Ray worker 需显式声明: + +```bash +ray start --resources='{"MLU": 8}' +``` + +### Intel XPU + +```bash +export VERL_PLATFORM=intel +source /opt/intel/oneapi/setvars.sh +``` + +### 燧原 Enflame GCU + +作为参考示例提供;完整生产级支持需要与厂商协作。 + +## 新增平台 + +厂商可通过以下步骤新增硬件平台: + +1. 在 `verl_hardware_plugin/platforms/` 下创建平台类,用 `@PlatformRegistry.register(platform="vendor_name")` 装饰。 +2. 在 `verl_hardware_plugin/engines/` 下创建 FSDP 和/或 Megatron 引擎,用 `@EngineRegistry.register(device=..., vendor=...)` 装饰。 +3. 在插件的 `__init__.py` 注册逻辑中加入设备/引擎检查。 + +带完整注释的平台模板见仓库中的[开发指南](https://github.com/verl-project/verl-hardware-plugin/blob/main/docs/development.md)。 diff --git a/docs/verl_hardware_plugin_zh/user_guide/user-guide.md b/docs/verl_hardware_plugin_zh/user_guide/user-guide.md new file mode 100644 index 0000000000..cd96f15368 --- /dev/null +++ b/docs/verl_hardware_plugin_zh/user_guide/user-guide.md @@ -0,0 +1,10 @@ +# 用户指南 + +本节介绍如何使用 verl-hardware-plugin 在各硬件平台上运行 verl RL 后训练任务。 + +```{toctree} +:maxdepth: 2 + +grpo-baseline.md +platforms.md +```