From f059626ebf923c22f721d2112dd3357e2c061d7c Mon Sep 17 00:00:00 2001 From: cheng874 Date: Wed, 30 Sep 2026 19:45:57 +0800 Subject: [PATCH] Torch-fl update --- docs/conf.py | 1854 ++++++++--------- docs/flagos_homepage/index.md | 604 +++--- .../locale/zh_CN/LC_MESSAGES/index.po | 1098 +++++----- .../locale/zh_CN/LC_MESSAGES/overview.po | 902 ++++---- docs/flagos_homepage/overview.md | 270 +-- docs/flagtree_en/getting_started/install.md | 14 +- docs/flagtree_zh/getting_started/install.md | 14 +- .../release_notes/release_notes_v070.md | 2 +- docs/torch_fl_en/architecture/distributed.md | 140 +- docs/torch_fl_en/architecture/profiler.md | 146 +- .../torch_fl_en/architecture/torch-compile.md | 196 +- .../getting_started/installation.md | 478 +++-- .../torch_fl_en/getting_started/quickstart.md | 96 + docs/torch_fl_en/index.md | 202 +- docs/torch_fl_en/overview/architecture.md | 152 +- docs/torch_fl_en/overview/features.md | 196 +- docs/torch_fl_en/overview/overview.md | 102 +- docs/torch_fl_en/reference/compatibility.md | 117 +- docs/torch_fl_en/reference/dtype-support.md | 40 + .../reference/environment-variables.md | 204 +- .../reference/platform-capability.md | 51 + docs/torch_fl_en/reference/troubleshooting.md | 63 + .../release_notes/release-notes.md | 28 +- docs/torch_fl_zh/architecture/distributed.md | 140 +- docs/torch_fl_zh/architecture/profiler.md | 146 +- .../torch_fl_zh/architecture/torch-compile.md | 196 +- .../getting_started/installation.md | 478 +++-- .../torch_fl_zh/getting_started/quickstart.md | 96 + docs/torch_fl_zh/index.md | 202 +- docs/torch_fl_zh/overview/architecture.md | 152 +- docs/torch_fl_zh/overview/features.md | 196 +- docs/torch_fl_zh/overview/overview.md | 102 +- docs/torch_fl_zh/reference/compatibility.md | 117 +- docs/torch_fl_zh/reference/dtype-support.md | 40 + .../reference/environment-variables.md | 204 +- .../reference/platform-capability.md | 51 + docs/torch_fl_zh/reference/troubleshooting.md | 63 + .../release_notes/release-notes.md | 28 +- 38 files changed, 4880 insertions(+), 4300 deletions(-) create mode 100644 docs/torch_fl_en/getting_started/quickstart.md create mode 100644 docs/torch_fl_en/reference/dtype-support.md create mode 100644 docs/torch_fl_en/reference/platform-capability.md create mode 100644 docs/torch_fl_en/reference/troubleshooting.md create mode 100644 docs/torch_fl_zh/getting_started/quickstart.md create mode 100644 docs/torch_fl_zh/reference/dtype-support.md create mode 100644 docs/torch_fl_zh/reference/platform-capability.md create mode 100644 docs/torch_fl_zh/reference/troubleshooting.md diff --git a/docs/conf.py b/docs/conf.py index df4c66edcb..023026494d 100644 --- a/docs/conf.py +++ b/docs/conf.py @@ -1,928 +1,926 @@ -""" -Shared Sphinx configuration using sphinx-multiproject. - -To build each project, the ``PROJECT`` environment variable is used. - -.. code:: console - - $ make html # build default project - $ PROJECT=flagos_homepage make html # build the FlagOS homepage - $ PROJECT=flagcx_en make html # build the flagcx English project - $ PROJECT=flaggems_en make html # build the flaggems English project - $ PROJECT=flaggems_vllm_en make html # build the flaggems-vllm English project - $ PROJECT=flagtree_en make html # build the flagtree English project - $ PROJECT=flagscale_en make html # build the flagscale English project - $ PROJECT=flagrelease_en make html # build the flagrelease English project - $ PROJECT=flagperf_en make html # build the flagperf English project - $ PROJECT=megatron_lm_fl_en make html # build the megatron_lm_fl English project - $ PROJECT=vllm_plugin_fl_en make html # build the vllm_plugin_fl English project - $ PROJECT=transformer_engine_fl_en make html # build the transformer_engine_fl English project - $ PROJECT=verl_fl_en make html # build the verl_fl English project - $ PROJECT=flagos_robo_en make html # build the flagos_robo English project - $ PROJECT=onlinelaboratory_en make html # build the onlinelaboratory English project - $ PROJECT=flagcicd_en make html # build the flagcicd English project - $ PROJECT=flagdnn_en make html # build the flagdnn English project - $ PROJECT=flagblas_en make html # build the flagblas English project - $ PROJECT=flagfft_en make html # build the flagfft English project - $ PROJECT=flagsparse_en make html # build the flagsparse English project - $ PROJECT=flagtensor_en make html # build the flagtensor English project - $ PROJECT=flagaudio_en make html # build the flagaudio English project - $ PROJECT=flagattention_en make html # build the flagattention English project - $ PROJECT=pytorch_plugin_fl_en make html # build the pytorch_plugin_fl English project - $ PROJECT=sglang_plugin_fl_en make html # build the sglang_plugin_fl English project - $ PROJECT=flagquantum_en make html # build the flagquantum English project - $ PROJECT=kernelgenbench_en make html # build the kernelgenbench English project - $ PROJECT=flagprism_en make html # build the flagprism English project - $ PROJECT=pytorch_plugin_fl_en make html # build the pytorch_plugin_fl English project - - $ PROJECT=flagos_zh make html # build the Chinese project - $ PROJECT=flagcx_zh make html # build the flagcx Chinese project - $ PROJECT=flaggems_zh make html # build the flaggems Chinese project - $ PROJECT=flaggems_vllm_zh make html # build the flaggems-vllm Chinese project - $ PROJECT=flagtree_zh make html # build the flagtree Chinese project - $ PROJECT=flagscale_zh make html # build the flagscale Chinese project - $ PROJECT=flagrelease_zh make html # build the flagrelease Chinese project - $ PROJECT=flagperf_zh make html # build the flagperf Chinese project - $ PROJECT=megatron_lm_fl_zh make html # build the megatron_lm_fl Chinese project - $ PROJECT=vllm_plugin_fl_zh make html # build the vllm_plugin_fl Chinese project - $ PROJECT=transformer_engine_fl_zh make html # build the transformer_engine_fl Chinese project - $ PROJECT=verl_fl_zh make html # build the transformer_engine_fl Chinese project - $ PROJECT=flagos_robo_zh make html # build the flagos_robo Chinese project - $ PROJECT=onlinelaboratory_zh make html # build the onlinelaboratory Chinese project - $ PROJECT=flagcicd_zh make html # build the flagcicd Chinese project - $ PROJECT=flagdnn_zh make html # build the flagdnn Chinese project - $ PROJECT=flagblas_zh make html # build the flagblas Chinese project - $ PROJECT=flagfft_zh make html # build the flagfft Chinese project - $ PROJECT=flagsparse_zh make html # build the flagsparse Chinese project - $ PROJECT=flagtensor_zh make html # build the flagtensor Chinese project - $ PROJECT=flagaudio_zh make html # build the flagaudio Chinese project - $ PROJECT=flagattention_zh make html # build the flagattention Chinese project - $ PROJECT=pytorch_plugin_fl_zh make html # build the pytorch_plugin_fl Chinese project - $ PROJECT=sglang_plugin_fl_zh make html # build the sglang_plugin_fl Chinese project - $ PROJECT=flagquantum_zh make html # build the flagquantum Chinese project - $ PROJECT=kernelgenbench_zh make html # build the kernelgenbench Chinese project - $ PROJECT=flagprism_zh make html # build the flagprism Chinese project - $ PROJECT=pytorch_plugin_fl_zh make html # build the pytorch_plugin_fl Chinese project - -For more information read https://sphinx-multiproject.readthedocs.io/. -""" - -import os -import sys - -# Fix imports: Check different import methods -try: - # First try sphinx_multiproject - from sphinx_multiproject.utils import get_project - print("INFO: Using sphinx_multiproject") -except ImportError: - try: - # Then try multiproject - from multiproject.utils import get_project - print("INFO: Using multiproject") - except ImportError: - # If both fail, create a simple get_project function - print("WARNING: sphinx-multiproject not found. Using simple project selection.") - def get_project(projects): - return os.environ.get("PROJECT", "flagos_homepage") - -sys.path.append(os.path.abspath("_ext")) - -# Base extensions - only include actually installed ones -extensions = [ - "multiproject", # Sphinx extension name, not Python module name - "myst_parser", - "sphinx_copybutton", - "sphinx_design", - # Temporarily comment out potentially problematic extensions - # "sphinx_tabs.tabs", # Module name might be different - # "sphinx_prompt", - "sphinx.ext.autodoc", - "sphinx.ext.autosectionlabel", - "sphinx.ext.extlinks", - "sphinx.ext.intersphinx", - # Comment out uninstalled extensions - # "sphinxcontrib.httpdomain", - # "sphinxcontrib.video", - # "sphinxemoji.sphinxemoji", - "sphinxext.opengraph", - "sphinx_tippy", - "sphinxcontrib.lightbox2", # click-to-enlarge / lightbox for images - "sphinx_tippy", - "sphinx_togglebutton", - "flagos_page_tags" -] - -# Check and add actually installed extensions -try: - import sphinx_tabs - extensions.append("sphinx_tabs.tabs") - print("INFO: sphinx_tabs extension added") -except ImportError: - print("INFO: sphinx_tabs not available") - -try: - import sphinx_prompt - extensions.append("sphinx_prompt") - print("INFO: sphinx_prompt extension added") -except ImportError: - print("INFO: sphinx_prompt not available") - -# Define all projects with their configurations -multiproject_projects = { - "flagos_homepage": { - "use_config_file": False, - "config": { - "project": "FlagOS Documentation", - "html_title": "FlagOS Documentation", - "locale_dirs": ["locale/"], - }, - }, - "flagcx_en": { - "use_config_file": False, - "config": { - "project": "FlagCX Documentation", - "html_title": "FlagCX Documentation", - }, - }, - "flaggems_en": { - "use_config_file": False, - "config": { - "project": "FlagGems Documentation", - "html_title": "FlagGems Documentation", - # Custom config values for extensions (paths relative to docs root) - "operator_yaml_path": "shared/conf/operators.yaml", - "benchmark_data_path": "shared/benchmark", - "coverage_data_path": "shared/coverage", - # Static files - include shared _static (for logo) and coverage - "html_static_path": ["_static", "flaggems_en/_static", "shared/coverage"], - "html_css_files": [ - "custom.css", # 全局的,包含 logo 设置 - "css/custom.css", # 项目特有的 - "https://unpkg.com/tabulator-tables@5.5.2/dist/css/tabulator.min.css", - ], - # # Theme options - # "html_theme": "sphinx_book_theme", - # "html_theme_options": { - # "github_url": "https://github.com/flagos-ai/FlagGems", - # "use_edit_page_button": True, - # "show_nav_level": 2, - # "navigation_with_keys": True, - # "show_toc_level": 2, - # }, - # "html_context": { - # "github_user": "flagos-ai", - # "github_repo": "FlagGems", - # "github_version": "master", - # "doc_path": "docs/flaggems_en", - # }, - # MyST config - "myst_enable_extensions": [ - "colon_fence", - "deflist", - "html_admonition", - "html_image", - "replacements", - "smartquotes", - "substitution", - "tasklist", - ], - "myst_heading_anchors": 3, - # Language - "language": "en", - }, - }, - "flaggems_vllm_en": { - "use_config_file": False, - "config": { - "project": "FlagGems-vLLM Documentation", - "html_title": "FlagGems-vLLM Documentation", - }, - }, - "flaggems_sglang_en": { - "use_config_file": False, - "config": { - "project": "FlagGems-sglang Documentation", - "html_title": "FlagGems-sglang Documentation", - }, - }, - "flaggems_sglang_zh": { - "use_config_file": False, - "config": { - "project": "FlagGems-sglang 文档中心", - "html_title": "FlagGems-sglang 文档中心", - }, - }, - "flagdnn_en": { - "use_config_file": False, - "config": { - "project": "FlagDNN Documentation", - "html_title": "FlagDNN Documentation", - }, - }, - "flagdnn_zh": { - "use_config_file": False, - "config": { - "project": "FlagDNN 文档中心", - "html_title": "FlagDNN 文档中心", - }, - }, - "flagblas_en": { - "use_config_file": False, - "config": { - "project": "FlagBLAS Documentation", - "html_title": "FlagBLAS Documentation", - }, - }, - "flagblas_zh": { - "use_config_file": False, - "config": { - "project": "FlagBLAS 文档中心", - "html_title": "FlagBLAS 文档中心", - }, - }, - "flagfft_en": { - "use_config_file": False, - "config": { - "project": "FlagFFT Documentation", - "html_title": "FlagFFT Documentation", - }, - }, - "flagfft_zh": { - "use_config_file": False, - "config": { - "project": "FlagFFT 文档中心", - "html_title": "FlagFFT 文档中心", - }, - }, - "flagsparse_en": { - "use_config_file": False, - "config": { - "project": "FlagSparse Documentation", - "html_title": "FlagSparse Documentation", - }, - }, - "flagsparse_zh": { - "use_config_file": False, - "config": { - "project": "FlagSparse 文档中心", - "html_title": "FlagSparse 文档中心", - }, - }, - "flagtensor_en": { - "use_config_file": False, - "config": { - "project": "FlagTensor Documentation", - "html_title": "FlagTensor Documentation", - }, - }, - "flagtensor_zh": { - "use_config_file": False, - "config": { - "project": "FlagTensor 文档中心", - "html_title": "FlagTensor 文档中心", - }, - }, - "flagaudio_en": { - "use_config_file": False, - "config": { - "project": "FlagAudio Documentation", - "html_title": "FlagAudio Documentation", - }, - }, - "flagaudio_zh": { - "use_config_file": False, - "config": { - "project": "FlagAudio 文档中心", - "html_title": "FlagAudio 文档中心", - }, - }, - "flagattention_en": { - "use_config_file": False, - "config": { - "project": "FlagAttention Documentation", - "html_title": "FlagAttention Documentation", - }, - }, - "flagattention_zh": { - "use_config_file": False, - "config": { - "project": "FlagAttention 文档中心", - "html_title": "FlagAttention 文档中心", - }, - }, - "pytorch_plugin_fl_en": { - "use_config_file": False, - "config": { - "project": "PyTorch-Plugin-FL Documentation", - "html_title": "PyTorch-Plugin-FL Documentation", - }, - }, - "pytorch_plugin_fl_zh": { - "use_config_file": False, - "config": { - "project": "PyTorch-Plugin-FL 文档中心", - "html_title": "PyTorch-Plugin-FL 文档中心", - }, - }, - "sglang_plugin_fl_en": { - "use_config_file": False, - "config": { - "project": "sglang-Plugin-FL Documentation", - "html_title": "sglang-Plugin-FL Documentation", - }, - }, - "sglang_plugin_fl_zh": { - "use_config_file": False, - "config": { - "project": "sglang-Plugin-FL 文档中心", - "html_title": "sglang-Plugin-FL 文档中心", - }, - }, - "flagquantum_en": { - "use_config_file": False, - "config": { - "project": "FlagQuantum Documentation", - "html_title": "FlagQuantum Documentation", - }, - }, - "flagquantum_zh": { - "use_config_file": False, - "config": { - "project": "FlagQuantum 文档中心", - "html_title": "FlagQuantum 文档中心", - }, - }, - "kernelgenbench_en": { - "use_config_file": False, - "config": { - "project": "KernelGenBench Documentation", - "html_title": "KernelGenBench Documentation", - }, - }, - "kernelgenbench_zh": { - "use_config_file": False, - "config": { - "project": "KernelGenBench 文档中心", - "html_title": "KernelGenBench 文档中心", - }, - }, - "flagprism_en": { - "use_config_file": False, - "config": { - "project": "FlagPrism Documentation", - "html_title": "FlagPrism Documentation", - }, - }, - "flagprism_zh": { - "use_config_file": False, - "config": { - "project": "FlagPrism 文档中心", - "html_title": "FlagPrism 文档中心", - }, - }, - "flagtree_en": { - "use_config_file": False, - "config": { - "project": "FlagTree Documentation", - "html_title": "FlagTree Documentation", - }, - }, - "flagscale_en": { - "use_config_file": False, - "config": { - "project": "FlagScale Documentation", - "html_title": "FlagScale Documentation", - }, - }, - "flagrelease_en": { - "use_config_file": False, - "config": { - "project": "FlagRelease Documentation", - "html_title": "FlagRelease Documentation", - }, - }, - "flagperf_en": { - "use_config_file": False, - "config": { - "project": "FlagPerf Documentation", - "html_title": "FlagPerf Documentation", - }, - }, - "megatron_lm_fl_en": { - "use_config_file": False, - "config": { - "project": "Megatron-LM-FL Documentation", - "html_title": "Megatron-LM-FL Documentation", - }, - }, - "vllm_plugin_fl_en": { - "use_config_file": False, - "config": { - "project": "VLLM-Plugin-FL Documentation", - "html_title": "VLLM-Plugin-FL Documentation", - }, - }, - "transformer_engine_fl_en": { - "use_config_file": False, - "config": { - "project": "Transformer-Engine-FL Documentation", - "html_title": "Transformer-Engine-FL Documentation", - }, - }, - "verl_fl_en": { - "use_config_file": False, - "config": { - "project": "verl-FL Documentation", - "html_title": "verl-FL Documentation", - }, - }, - "verl_hardware_plugin_en": { - "use_config_file": False, - "config": { - "project": "verl-hardware-plugin Documentation", - "html_title": "verl-hardware-plugin Documentation", - }, - }, - "flagos_robo_en": { - "use_config_file": False, - "config": { - "project": "FlagOS-Robo Documentation", - "html_title": "FlagOS-Robo Documentation", - }, - }, - "onlinelaboratory_en": { - "use_config_file": False, - "config": { - "project": "Online Laboratory Documentation", - "html_title": "Online Laboratory Documentation", - }, - }, - "flagcicd_en": { - "use_config_file": False, - "config": { - "project": "FlagCICD Documentation", - "html_title": "FlagCICD Documentation", - }, - }, - # "flagos_zh": { - # "use_config_file": False, - # "config": { - # "project": "FlagOS 文档中心", - # "html_title": "FlagOS 文档中心", - # }, - # }, - "flagcx_zh": { - "use_config_file": False, - "config": { - "project": "FlagCX 文档中心", - "html_title": "FlagCX 文档中心", - }, - }, - "flaggems_zh": { - "use_config_file": False, - "config": { - "project": "FlagGems 文档中心", - "html_title": "FlagGems 文档中心", - # Custom config values for extensions (paths relative to docs root) - "operator_yaml_path": "shared/conf/operators.yaml", - "benchmark_data_path": "shared/benchmark", - "coverage_data_path": "shared/coverage", - # Static files - include shared _static (for logo) and coverage - "html_static_path": ["_static", "flaggems_zh/_static", "shared/coverage"], - "html_css_files": [ - "custom.css", # 全局的,包含 logo 设置 - "css/custom.css", # 项目特有的 - "https://unpkg.com/tabulator-tables@5.5.2/dist/css/tabulator.min.css", - ], - # # Theme options - # "html_theme": "sphinx_book_theme", - # "html_theme_options": { - # "github_url": "https://github.com/flagos-ai/FlagGems", - # "use_edit_page_button": True, - # "show_nav_level": 2, - # "navigation_with_keys": True, - # "show_toc_level": 2, - # }, - # "html_context": { - # "github_user": "flagos-ai", - # "github_repo": "FlagGems", - # "github_version": "master", - # "doc_path": "docs/flaggems_zh", - # }, - # MyST config - "myst_enable_extensions": [ - "colon_fence", - "deflist", - "html_admonition", - "html_image", - "replacements", - "smartquotes", - "substitution", - "tasklist", - ], - "myst_heading_anchors": 3, - # Language - "language": "zh", - }, - }, - "flaggems_vllm_zh": { - "use_config_file": False, - "config": { - "project": "FlagGems-vLLM 文档中心", - "html_title": "FlagGems-vLLM 文档中心", - }, - }, - "flagtree_zh": { - "use_config_file": False, - "config": { - "project": "FlagTree 文档中心", - "html_title": "FlagTree 文档中心", - }, - }, - "flagscale_zh": { - "use_config_file": False, - "config": { - "project": "FlagScale 文档中心", - "html_title": "FlagScale 文档中心", - }, - }, - "flagrelease_zh": { - "use_config_file": False, - "config": { - "project": "FlagRelease 文档中心", - "html_title": "FlagRelease 文档中心", - }, - }, - "flagperf_zh": { - "use_config_file": False, - "config": { - "project": "FlagPerf 文档中心", - "html_title": "FlagPerf 文档中心", - }, - }, - "megatron_lm_fl_zh": { - "use_config_file": False, - "config": { - "project": "Megatron-LM-FL 文档中心", - "html_title": "Megatron-LM-FL 文档中心", - }, - }, - "vllm_plugin_fl_zh": { - "use_config_file": False, - "config": { - "project": "VLLM-Plugin-FL 文档中心", - "html_title": "VLLM-Plugin-FL 文档中心", - }, - }, - "transformer_engine_fl_zh": { - "use_config_file": False, - "config": { - "project": "Transformer-Engine-FL 文档中心", - "html_title": "Transformer-Engine-FL 文档中心", - }, - }, - "verl_fl_zh": { - "use_config_file": False, - "config": { - "project": "verl-FL 文档中心", - "html_title": "verl-FL 文档中心", - }, - }, - "verl_hardware_plugin_zh": { - "use_config_file": False, - "config": { - "project": "verl-hardware-plugin 文档中心", - "html_title": "verl-hardware-plugin 文档中心", - }, - }, - "flagos_robo_zh": { - "use_config_file": False, - "config": { - "project": "FlagOS-Robo 文档中心", - "html_title": "FlagOS-Robo 文档中心", - }, - }, - "onlinelaboratory_zh": { - "use_config_file": False, - "config": { - "project": "线上实验室文档中心", - "html_title": "线上实验室文档中心", - }, - }, - "flagcicd_zh": { - "use_config_file": False, - "config": { - "project": "FlagCICD 文档中心", - "html_title": "FlagCICD 文档中心", - }, - }, -} - -docset = get_project(multiproject_projects) - -# Add project-specific _ext directory for custom extensions (FlagGems) -if docset in ["flaggems_en", "flaggems_zh"]: - project_ext_path = os.path.abspath(os.path.join(docset, "_ext")) - sys.path.insert(0, project_ext_path) - print(f"INFO: Added {project_ext_path} to sys.path for {docset}") - - # Add custom extensions for FlagGems (must be in global extensions list) - try: - import operator_list - import benchmark_table - import coverage_data - extensions.extend(["operator_list", "benchmark_table", "coverage_data"]) - print(f"INFO: Added FlagGems custom extensions for {docset}") - except ImportError as e: - print(f"WARNING: Could not import FlagGems extensions: {e}") - -ogp_site_name = "KernelGen Documentation" -ogp_use_first_image = True -ogp_image = "https://docs.readthedocs.io/en/latest/_static/img/logo-opengraph.png" -ogp_custom_meta_tags = ( - '', -) -ogp_enable_meta_description = True -ogp_description_length = 300 - -# templates_path = ["_templates"] -html_baseurl = os.environ.get("READTHEDOCS_CANONICAL_URL", "/") - -master_doc = "index" -copyright = '2026, FlagOS Community' -author = 'FlagOS Community' -release = '1.0.0' -# release = version - -# Exclude patterns - exclude all other project directories -exclude_patterns = [ - "_build", - "shared", - "_includes", - "chip_adaptation_guide/_shared", - "chip_adaptation_guide/TODO.md", - "chip_adaptation_guide_toctree_backup", -] -all_projects = list(multiproject_projects.keys()) -for project in all_projects: - if project != docset: - exclude_patterns.append(project) -if docset in ["flagrelease_en", "flagrelease_zh"]: - exclude_patterns.append("model_readmes") - -# flagcicd 项目:排除管理员专用页面(用户管理、模型管理) -if docset in ["flagcicd_en", "flagcicd_zh"]: - exclude_patterns.extend([ - "function-description/user-management.md", - "function-description/model-management.md", - "operation-guide/user-management.md", - "operation-guide/model-management.md", - ]) - -default_role = "obj" -intersphinx_cache_limit = 14 -intersphinx_timeout = 3 -intersphinx_mapping = { - "python": ("https://docs.python.org/3.10/", None), - "sphinx": ("https://www.sphinx-doc.org/en/master/", None), -} - -intersphinx_disabled_reftypes = ["*"] - -myst_frontmatter_process = "yaml" - -myst_enable_extensions = [ - "dollarmath", - "amsmath", - "deflist", - "fieldlist", - "html_admonition", - "html_image", - "colon_fence", - "smartquotes", - "replacements", - # "linkify", - "strikethrough", - "substitution", - "tasklist", - "attrs_inline", - "attrs_block", - # "substitution", -] - -htmlhelp_basename = "KernelGendoc" -latex_documents = [ - ( - "index", - "KernelGen.tex", - "KernelGen Documentation", - "KernelGen Team", - "manual", - ), -] -man_pages = [ - ( - "index", - "kernelgen", - "KernelGen Documentation", - ["KernelGen Team"], - 1, - ) -] - -# Set language based on project suffix or environment variable (sphinx-intl support) -# language = os.environ.get("READTHEDOCS_LANGUAGE", "en") if docset == "flagos_homepage" else ("en" if docset.endswith("_en") else "zh_CN") -language = "en" - -# Detect the actual build language from Read the Docs environment variable -# Falls back to the language config variable for local builds -CURRENT_LANGUAGE = os.getenv("READTHEDOCS_LANGUAGE", language) - -# if docset == "flagos_homepage": -# is_zh = CURRENT_LANGUAGE in ["zh_CN", "zh", "zh-cn"] -# else: -# is_zh = docset.endswith("_zh") -# lang_prefix = "zh-cn" if is_zh else "en" - -# # 定义 myst_substitutions -# myst_substitutions = { -# "lang_prefix": lang_prefix, -# } - -# locale_dirs = [ -# f"{docset}/locale/", -# ] -gettext_compact = False - -html_short_title = "" - -# ============================================================================ -# HTML THEME CONFIGURATION - DIFFERENT THEMES FOR DIFFERENT PROJECTS -# ============================================================================ - -# Only flagos_homepage uses pydata_sphinx_theme, all others use sphinx_book_theme -if docset == "flagos_homepage": - html_theme = "pydata_sphinx_theme" -else: - html_theme = "sphinx_book_theme" - -# Common static paths -html_static_path = ["_static", f"{docset}/_static"] -html_css_files = ["custom.css", "homepage.css"] -if docset == "flagos_homepage": - html_css_files.append("guide.css") -html_js_files = [] - -# html_logo = "img/logo.png" -html_favicon = "_static/favicon.svg" - -# Theme-specific configurations -if html_theme == "pydata_sphinx_theme": - # PyData Sphinx Theme configuration for flagos_homepage - - # Set logo based on language (sphinx-intl support) - # Read the Docs sets READTHEDOCS_LANGUAGE environment variable during builds - # ReadTheDocs uses lowercase codes (zh, zh-cn), while Sphinx uses zh_CN - - # Debug output - print(f"DEBUG: CURRENT_LANGUAGE = '{CURRENT_LANGUAGE}'") - print(f"DEBUG: language = '{language}'") - print(f"DEBUG: READTHEDOCS_LANGUAGE env = '{os.getenv('READTHEDOCS_LANGUAGE', 'NOT SET')}'") - print(f"DEBUG: Is Chinese? {CURRENT_LANGUAGE in ['zh_CN', 'zh', 'zh-cn']}") - - if CURRENT_LANGUAGE in ["zh_CN", "zh", "zh-cn"]: - logo_config = { - "text": "文档中心", - "image_light": "_static/logo-zh-light.svg", - "image_dark": "_static/logo-zh-dark.svg", - } - print("DEBUG: Using CHINESE logo config") - else: - # Default to English configuration for all other languages - logo_config = { - "text": "Documentation", - "image_light": "_static/logo-en-light.svg", - "image_dark": "_static/logo-en-dark.svg", - } - print("DEBUG: Using ENGLISH logo config") - - html_theme_options = { - "logo": logo_config, - "home_page_in_toc": True, - "use_download_button": False, - "repository_url": "https://github.com/flagos-ai/KernelGen", - "use_repository_button": True, - "secondary_sidebar_items": { - "**": ["page-toc"], - "flagos_homepage/index": [], - }, - "show_toc_level": 2, - "footer_start": ["copyright"], - "footer_end": [], - "show_sphinx": False, - "navbar_end": ["navbar-icon-links"] - } - - # Keep the FlagOS homepage clean while enabling page-local TOC elsewhere. - - # html_sidebars is only for PyData Sphinx Theme - html_sidebars = {} - for project in all_projects: - html_sidebars[f"{project}/index"] = [] - - # html_context is only applied to PyData Sphinx Theme - html_context = { - "default_mode": "light" - } - - # No additional JS files needed for portal homepage - html_js_files = [] - -else: - # Sphinx Book Theme configuration for all other projects - - # # repo URL per project - # repository_urls = { - # "flagcx_en": "https://github.com/flagos-ai/FlagCX", - # "flagcx_zh": "https://github.com/flagos-ai/FlagCX", - # "flaggems_en": "https://github.com/flagos-ai/FlagGems", - # "flaggems_zh": "https://github.com/flagos-ai/FlagGems", - # "flagtree_en": "https://github.com/flagos-ai/FlagTree", - # "flagtree_zh": "https://github.com/flagos-ai/FlagTree", - # "flagrelease_en": "https://github.com/flagos-ai/FlagRelease", - # "flagrelease_zh": "https://github.com/flagos-ai/FlagRelease", - # "flagperf_en": "https://github.com/flagos-ai/FlagPerf", - # "flagperf_zh": "https://github.com/flagos-ai/FlagPerf", - # } - - # # Obtain the current repo URL, if failed, set the default value - # current_repo_url = repository_urls.get(docset, "https://github.com/flagos-ai") - - if docset.endswith("_en"): - main_site_url = "https://docs.flagos.io/en/latest/" - main_site_text = "Back to FlagOS Documentation" - else: - main_site_url = "https://docs.flagos.io/zh-cn/latest/" - main_site_text = "返回 FlagOS 文档" - - templates_path = ["_templates"] - - # Logo configuration per project - if docset in ["flagcicd_en", "flagcicd_zh"]: - logo_config = { - "image_light": "_static/flagcicd-logo-light.svg", - "image_dark": "_static/flagcicd-logo-dark.svg", - } - else: - logo_config = { - "image_light": "_static/logo-en-light.svg", - "image_dark": "_static/logo-en-dark.svg", - } - - # Sphinx Book Theme configuration for all other projects - html_theme_options = { - "logo": logo_config, - "home_page_in_toc": True, - "use_download_button": False, - "repository_url": "https://github.com/flagos-ai/docs", - "use_edit_page_button": True, - "use_repository_button": True, - "navbar_center": ["back_to_main.html"], - # "default_mode": "light", - } - - html_context = { - "main_site_url": main_site_url, - "main_site_text": main_site_text, - "default_mode": "light" - } - - # No html_sidebars for Sphinx Book Theme - html_sidebars = {} - # No html_context for Sphinx Book Theme - html_last_updated_fmt = '%b %d, %Y' - -rst_epilog = """ -.. |org_brand| replace:: KernelGen Community -.. |com_brand| replace:: KernelGen for Business -.. |git_providers_and| replace:: GitHub, Bitbucket, and GitLab -.. |git_providers_or| replace:: GitHub, Bitbucket, or GitLab -""" - -autosectionlabel_prefix_document = True - -linkcheck_retries = 2 -linkcheck_timeout = 1 -linkcheck_workers = 10 -linkcheck_ignore = [ - r"http://127\.0\.0\.1", - r"http://localhost", - r"https://github\.com.+?#L\d+", -] - -extlinks = { - "issue": ("https://github.com/armstrongttwalker-alt/test-i18n-KernelGen/issues/%s", "#%s"), -} - -suppress_warnings = ["epub.unknown_project_files"] +""" +Shared Sphinx configuration using sphinx-multiproject. + +To build each project, the ``PROJECT`` environment variable is used. + +.. code:: console + + $ make html # build default project + $ PROJECT=flagos_homepage make html # build the FlagOS homepage + $ PROJECT=flagcx_en make html # build the flagcx English project + $ PROJECT=flaggems_en make html # build the flaggems English project + $ PROJECT=flaggems_vllm_en make html # build the flaggems-vllm English project + $ PROJECT=flagtree_en make html # build the flagtree English project + $ PROJECT=flagscale_en make html # build the flagscale English project + $ PROJECT=flagrelease_en make html # build the flagrelease English project + $ PROJECT=flagperf_en make html # build the flagperf English project + $ PROJECT=megatron_lm_fl_en make html # build the megatron_lm_fl English project + $ PROJECT=vllm_plugin_fl_en make html # build the vllm_plugin_fl English project + $ PROJECT=transformer_engine_fl_en make html # build the transformer_engine_fl English project + $ PROJECT=verl_fl_en make html # build the verl_fl English project + $ PROJECT=flagos_robo_en make html # build the flagos_robo English project + $ PROJECT=onlinelaboratory_en make html # build the onlinelaboratory English project + $ PROJECT=flagcicd_en make html # build the flagcicd English project + $ PROJECT=flagdnn_en make html # build the flagdnn English project + $ PROJECT=flagblas_en make html # build the flagblas English project + $ PROJECT=flagfft_en make html # build the flagfft English project + $ PROJECT=flagsparse_en make html # build the flagsparse English project + $ PROJECT=flagtensor_en make html # build the flagtensor English project + $ PROJECT=flagaudio_en make html # build the flagaudio English project + $ PROJECT=flagattention_en make html # build the flagattention English project + $ PROJECT=torch_fl_en make html # build the torch_fl English project + $ PROJECT=sglang_plugin_fl_en make html # build the sglang_plugin_fl English project + $ PROJECT=flagquantum_en make html # build the flagquantum English project + $ PROJECT=kernelgenbench_en make html # build the kernelgenbench English project + $ PROJECT=flagprism_en make html # build the flagprism English project + + $ PROJECT=flagos_zh make html # build the Chinese project + $ PROJECT=flagcx_zh make html # build the flagcx Chinese project + $ PROJECT=flaggems_zh make html # build the flaggems Chinese project + $ PROJECT=flaggems_vllm_zh make html # build the flaggems-vllm Chinese project + $ PROJECT=flagtree_zh make html # build the flagtree Chinese project + $ PROJECT=flagscale_zh make html # build the flagscale Chinese project + $ PROJECT=flagrelease_zh make html # build the flagrelease Chinese project + $ PROJECT=flagperf_zh make html # build the flagperf Chinese project + $ PROJECT=megatron_lm_fl_zh make html # build the megatron_lm_fl Chinese project + $ PROJECT=vllm_plugin_fl_zh make html # build the vllm_plugin_fl Chinese project + $ PROJECT=transformer_engine_fl_zh make html # build the transformer_engine_fl Chinese project + $ PROJECT=verl_fl_zh make html # build the transformer_engine_fl Chinese project + $ PROJECT=flagos_robo_zh make html # build the flagos_robo Chinese project + $ PROJECT=onlinelaboratory_zh make html # build the onlinelaboratory Chinese project + $ PROJECT=flagcicd_zh make html # build the flagcicd Chinese project + $ PROJECT=flagdnn_zh make html # build the flagdnn Chinese project + $ PROJECT=flagblas_zh make html # build the flagblas Chinese project + $ PROJECT=flagfft_zh make html # build the flagfft Chinese project + $ PROJECT=flagsparse_zh make html # build the flagsparse Chinese project + $ PROJECT=flagtensor_zh make html # build the flagtensor Chinese project + $ PROJECT=flagaudio_zh make html # build the flagaudio Chinese project + $ PROJECT=flagattention_zh make html # build the flagattention Chinese project + $ PROJECT=torch_fl_zh make html # build the torch_fl Chinese project + $ PROJECT=sglang_plugin_fl_zh make html # build the sglang_plugin_fl Chinese project + $ PROJECT=flagquantum_zh make html # build the flagquantum Chinese project + $ PROJECT=kernelgenbench_zh make html # build the kernelgenbench Chinese project + $ PROJECT=flagprism_zh make html # build the flagprism Chinese project + +For more information read https://sphinx-multiproject.readthedocs.io/. +""" + +import os +import sys + +# Fix imports: Check different import methods +try: + # First try sphinx_multiproject + from sphinx_multiproject.utils import get_project + print("INFO: Using sphinx_multiproject") +except ImportError: + try: + # Then try multiproject + from multiproject.utils import get_project + print("INFO: Using multiproject") + except ImportError: + # If both fail, create a simple get_project function + print("WARNING: sphinx-multiproject not found. Using simple project selection.") + def get_project(projects): + return os.environ.get("PROJECT", "flagos_homepage") + +sys.path.append(os.path.abspath("_ext")) + +# Base extensions - only include actually installed ones +extensions = [ + "multiproject", # Sphinx extension name, not Python module name + "myst_parser", + "sphinx_copybutton", + "sphinx_design", + # Temporarily comment out potentially problematic extensions + # "sphinx_tabs.tabs", # Module name might be different + # "sphinx_prompt", + "sphinx.ext.autodoc", + "sphinx.ext.autosectionlabel", + "sphinx.ext.extlinks", + "sphinx.ext.intersphinx", + # Comment out uninstalled extensions + # "sphinxcontrib.httpdomain", + # "sphinxcontrib.video", + # "sphinxemoji.sphinxemoji", + "sphinxext.opengraph", + "sphinx_tippy", + "sphinxcontrib.lightbox2", # click-to-enlarge / lightbox for images + "sphinx_tippy", + "sphinx_togglebutton", + "flagos_page_tags" +] + +# Check and add actually installed extensions +try: + import sphinx_tabs + extensions.append("sphinx_tabs.tabs") + print("INFO: sphinx_tabs extension added") +except ImportError: + print("INFO: sphinx_tabs not available") + +try: + import sphinx_prompt + extensions.append("sphinx_prompt") + print("INFO: sphinx_prompt extension added") +except ImportError: + print("INFO: sphinx_prompt not available") + +# Define all projects with their configurations +multiproject_projects = { + "flagos_homepage": { + "use_config_file": False, + "config": { + "project": "FlagOS Documentation", + "html_title": "FlagOS Documentation", + "locale_dirs": ["locale/"], + }, + }, + "flagcx_en": { + "use_config_file": False, + "config": { + "project": "FlagCX Documentation", + "html_title": "FlagCX Documentation", + }, + }, + "flaggems_en": { + "use_config_file": False, + "config": { + "project": "FlagGems Documentation", + "html_title": "FlagGems Documentation", + # Custom config values for extensions (paths relative to docs root) + "operator_yaml_path": "shared/conf/operators.yaml", + "benchmark_data_path": "shared/benchmark", + "coverage_data_path": "shared/coverage", + # Static files - include shared _static (for logo) and coverage + "html_static_path": ["_static", "flaggems_en/_static", "shared/coverage"], + "html_css_files": [ + "custom.css", # 全局的,包含 logo 设置 + "css/custom.css", # 项目特有的 + "https://unpkg.com/tabulator-tables@5.5.2/dist/css/tabulator.min.css", + ], + # # Theme options + # "html_theme": "sphinx_book_theme", + # "html_theme_options": { + # "github_url": "https://github.com/flagos-ai/FlagGems", + # "use_edit_page_button": True, + # "show_nav_level": 2, + # "navigation_with_keys": True, + # "show_toc_level": 2, + # }, + # "html_context": { + # "github_user": "flagos-ai", + # "github_repo": "FlagGems", + # "github_version": "master", + # "doc_path": "docs/flaggems_en", + # }, + # MyST config + "myst_enable_extensions": [ + "colon_fence", + "deflist", + "html_admonition", + "html_image", + "replacements", + "smartquotes", + "substitution", + "tasklist", + ], + "myst_heading_anchors": 3, + # Language + "language": "en", + }, + }, + "flaggems_vllm_en": { + "use_config_file": False, + "config": { + "project": "FlagGems-vLLM Documentation", + "html_title": "FlagGems-vLLM Documentation", + }, + }, + "flaggems_sglang_en": { + "use_config_file": False, + "config": { + "project": "FlagGems-sglang Documentation", + "html_title": "FlagGems-sglang Documentation", + }, + }, + "flaggems_sglang_zh": { + "use_config_file": False, + "config": { + "project": "FlagGems-sglang 文档中心", + "html_title": "FlagGems-sglang 文档中心", + }, + }, + "flagdnn_en": { + "use_config_file": False, + "config": { + "project": "FlagDNN Documentation", + "html_title": "FlagDNN Documentation", + }, + }, + "flagdnn_zh": { + "use_config_file": False, + "config": { + "project": "FlagDNN 文档中心", + "html_title": "FlagDNN 文档中心", + }, + }, + "flagblas_en": { + "use_config_file": False, + "config": { + "project": "FlagBLAS Documentation", + "html_title": "FlagBLAS Documentation", + }, + }, + "flagblas_zh": { + "use_config_file": False, + "config": { + "project": "FlagBLAS 文档中心", + "html_title": "FlagBLAS 文档中心", + }, + }, + "flagfft_en": { + "use_config_file": False, + "config": { + "project": "FlagFFT Documentation", + "html_title": "FlagFFT Documentation", + }, + }, + "flagfft_zh": { + "use_config_file": False, + "config": { + "project": "FlagFFT 文档中心", + "html_title": "FlagFFT 文档中心", + }, + }, + "flagsparse_en": { + "use_config_file": False, + "config": { + "project": "FlagSparse Documentation", + "html_title": "FlagSparse Documentation", + }, + }, + "flagsparse_zh": { + "use_config_file": False, + "config": { + "project": "FlagSparse 文档中心", + "html_title": "FlagSparse 文档中心", + }, + }, + "flagtensor_en": { + "use_config_file": False, + "config": { + "project": "FlagTensor Documentation", + "html_title": "FlagTensor Documentation", + }, + }, + "flagtensor_zh": { + "use_config_file": False, + "config": { + "project": "FlagTensor 文档中心", + "html_title": "FlagTensor 文档中心", + }, + }, + "flagaudio_en": { + "use_config_file": False, + "config": { + "project": "FlagAudio Documentation", + "html_title": "FlagAudio Documentation", + }, + }, + "flagaudio_zh": { + "use_config_file": False, + "config": { + "project": "FlagAudio 文档中心", + "html_title": "FlagAudio 文档中心", + }, + }, + "flagattention_en": { + "use_config_file": False, + "config": { + "project": "FlagAttention Documentation", + "html_title": "FlagAttention Documentation", + }, + }, + "flagattention_zh": { + "use_config_file": False, + "config": { + "project": "FlagAttention 文档中心", + "html_title": "FlagAttention 文档中心", + }, + }, + "torch_fl_en": { + "use_config_file": False, + "config": { + "project": "Torch-FL Documentation", + "html_title": "Torch-FL Documentation", + }, + }, + "torch_fl_zh": { + "use_config_file": False, + "config": { + "project": "Torch-FL 文档中心", + "html_title": "Torch-FL 文档中心", + }, + }, + "sglang_plugin_fl_en": { + "use_config_file": False, + "config": { + "project": "sglang-Plugin-FL Documentation", + "html_title": "sglang-Plugin-FL Documentation", + }, + }, + "sglang_plugin_fl_zh": { + "use_config_file": False, + "config": { + "project": "sglang-Plugin-FL 文档中心", + "html_title": "sglang-Plugin-FL 文档中心", + }, + }, + "flagquantum_en": { + "use_config_file": False, + "config": { + "project": "FlagQuantum Documentation", + "html_title": "FlagQuantum Documentation", + }, + }, + "flagquantum_zh": { + "use_config_file": False, + "config": { + "project": "FlagQuantum 文档中心", + "html_title": "FlagQuantum 文档中心", + }, + }, + "kernelgenbench_en": { + "use_config_file": False, + "config": { + "project": "KernelGenBench Documentation", + "html_title": "KernelGenBench Documentation", + }, + }, + "kernelgenbench_zh": { + "use_config_file": False, + "config": { + "project": "KernelGenBench 文档中心", + "html_title": "KernelGenBench 文档中心", + }, + }, + "flagprism_en": { + "use_config_file": False, + "config": { + "project": "FlagPrism Documentation", + "html_title": "FlagPrism Documentation", + }, + }, + "flagprism_zh": { + "use_config_file": False, + "config": { + "project": "FlagPrism 文档中心", + "html_title": "FlagPrism 文档中心", + }, + }, + "flagtree_en": { + "use_config_file": False, + "config": { + "project": "FlagTree Documentation", + "html_title": "FlagTree Documentation", + }, + }, + "flagscale_en": { + "use_config_file": False, + "config": { + "project": "FlagScale Documentation", + "html_title": "FlagScale Documentation", + }, + }, + "flagrelease_en": { + "use_config_file": False, + "config": { + "project": "FlagRelease Documentation", + "html_title": "FlagRelease Documentation", + }, + }, + "flagperf_en": { + "use_config_file": False, + "config": { + "project": "FlagPerf Documentation", + "html_title": "FlagPerf Documentation", + }, + }, + "megatron_lm_fl_en": { + "use_config_file": False, + "config": { + "project": "Megatron-LM-FL Documentation", + "html_title": "Megatron-LM-FL Documentation", + }, + }, + "vllm_plugin_fl_en": { + "use_config_file": False, + "config": { + "project": "VLLM-Plugin-FL Documentation", + "html_title": "VLLM-Plugin-FL Documentation", + }, + }, + "transformer_engine_fl_en": { + "use_config_file": False, + "config": { + "project": "Transformer-Engine-FL Documentation", + "html_title": "Transformer-Engine-FL Documentation", + }, + }, + "verl_fl_en": { + "use_config_file": False, + "config": { + "project": "verl-FL Documentation", + "html_title": "verl-FL Documentation", + }, + }, + "verl_hardware_plugin_en": { + "use_config_file": False, + "config": { + "project": "verl-hardware-plugin Documentation", + "html_title": "verl-hardware-plugin Documentation", + }, + }, + "flagos_robo_en": { + "use_config_file": False, + "config": { + "project": "FlagOS-Robo Documentation", + "html_title": "FlagOS-Robo Documentation", + }, + }, + "onlinelaboratory_en": { + "use_config_file": False, + "config": { + "project": "Online Laboratory Documentation", + "html_title": "Online Laboratory Documentation", + }, + }, + "flagcicd_en": { + "use_config_file": False, + "config": { + "project": "FlagCICD Documentation", + "html_title": "FlagCICD Documentation", + }, + }, + # "flagos_zh": { + # "use_config_file": False, + # "config": { + # "project": "FlagOS 文档中心", + # "html_title": "FlagOS 文档中心", + # }, + # }, + "flagcx_zh": { + "use_config_file": False, + "config": { + "project": "FlagCX 文档中心", + "html_title": "FlagCX 文档中心", + }, + }, + "flaggems_zh": { + "use_config_file": False, + "config": { + "project": "FlagGems 文档中心", + "html_title": "FlagGems 文档中心", + # Custom config values for extensions (paths relative to docs root) + "operator_yaml_path": "shared/conf/operators.yaml", + "benchmark_data_path": "shared/benchmark", + "coverage_data_path": "shared/coverage", + # Static files - include shared _static (for logo) and coverage + "html_static_path": ["_static", "flaggems_zh/_static", "shared/coverage"], + "html_css_files": [ + "custom.css", # 全局的,包含 logo 设置 + "css/custom.css", # 项目特有的 + "https://unpkg.com/tabulator-tables@5.5.2/dist/css/tabulator.min.css", + ], + # # Theme options + # "html_theme": "sphinx_book_theme", + # "html_theme_options": { + # "github_url": "https://github.com/flagos-ai/FlagGems", + # "use_edit_page_button": True, + # "show_nav_level": 2, + # "navigation_with_keys": True, + # "show_toc_level": 2, + # }, + # "html_context": { + # "github_user": "flagos-ai", + # "github_repo": "FlagGems", + # "github_version": "master", + # "doc_path": "docs/flaggems_zh", + # }, + # MyST config + "myst_enable_extensions": [ + "colon_fence", + "deflist", + "html_admonition", + "html_image", + "replacements", + "smartquotes", + "substitution", + "tasklist", + ], + "myst_heading_anchors": 3, + # Language + "language": "zh", + }, + }, + "flaggems_vllm_zh": { + "use_config_file": False, + "config": { + "project": "FlagGems-vLLM 文档中心", + "html_title": "FlagGems-vLLM 文档中心", + }, + }, + "flagtree_zh": { + "use_config_file": False, + "config": { + "project": "FlagTree 文档中心", + "html_title": "FlagTree 文档中心", + }, + }, + "flagscale_zh": { + "use_config_file": False, + "config": { + "project": "FlagScale 文档中心", + "html_title": "FlagScale 文档中心", + }, + }, + "flagrelease_zh": { + "use_config_file": False, + "config": { + "project": "FlagRelease 文档中心", + "html_title": "FlagRelease 文档中心", + }, + }, + "flagperf_zh": { + "use_config_file": False, + "config": { + "project": "FlagPerf 文档中心", + "html_title": "FlagPerf 文档中心", + }, + }, + "megatron_lm_fl_zh": { + "use_config_file": False, + "config": { + "project": "Megatron-LM-FL 文档中心", + "html_title": "Megatron-LM-FL 文档中心", + }, + }, + "vllm_plugin_fl_zh": { + "use_config_file": False, + "config": { + "project": "VLLM-Plugin-FL 文档中心", + "html_title": "VLLM-Plugin-FL 文档中心", + }, + }, + "transformer_engine_fl_zh": { + "use_config_file": False, + "config": { + "project": "Transformer-Engine-FL 文档中心", + "html_title": "Transformer-Engine-FL 文档中心", + }, + }, + "verl_fl_zh": { + "use_config_file": False, + "config": { + "project": "verl-FL 文档中心", + "html_title": "verl-FL 文档中心", + }, + }, + "verl_hardware_plugin_zh": { + "use_config_file": False, + "config": { + "project": "verl-hardware-plugin 文档中心", + "html_title": "verl-hardware-plugin 文档中心", + }, + }, + "flagos_robo_zh": { + "use_config_file": False, + "config": { + "project": "FlagOS-Robo 文档中心", + "html_title": "FlagOS-Robo 文档中心", + }, + }, + "onlinelaboratory_zh": { + "use_config_file": False, + "config": { + "project": "线上实验室文档中心", + "html_title": "线上实验室文档中心", + }, + }, + "flagcicd_zh": { + "use_config_file": False, + "config": { + "project": "FlagCICD 文档中心", + "html_title": "FlagCICD 文档中心", + }, + }, +} + +docset = get_project(multiproject_projects) + +# Add project-specific _ext directory for custom extensions (FlagGems) +if docset in ["flaggems_en", "flaggems_zh"]: + project_ext_path = os.path.abspath(os.path.join(docset, "_ext")) + sys.path.insert(0, project_ext_path) + print(f"INFO: Added {project_ext_path} to sys.path for {docset}") + + # Add custom extensions for FlagGems (must be in global extensions list) + try: + import operator_list + import benchmark_table + import coverage_data + extensions.extend(["operator_list", "benchmark_table", "coverage_data"]) + print(f"INFO: Added FlagGems custom extensions for {docset}") + except ImportError as e: + print(f"WARNING: Could not import FlagGems extensions: {e}") + +ogp_site_name = "KernelGen Documentation" +ogp_use_first_image = True +ogp_image = "https://docs.readthedocs.io/en/latest/_static/img/logo-opengraph.png" +ogp_custom_meta_tags = ( + '', +) +ogp_enable_meta_description = True +ogp_description_length = 300 + +# templates_path = ["_templates"] +html_baseurl = os.environ.get("READTHEDOCS_CANONICAL_URL", "/") + +master_doc = "index" +copyright = '2026, FlagOS Community' +author = 'FlagOS Community' +release = '1.0.0' +# release = version + +# Exclude patterns - exclude all other project directories +exclude_patterns = [ + "_build", + "shared", + "_includes", + "chip_adaptation_guide/_shared", + "chip_adaptation_guide/TODO.md", + "chip_adaptation_guide_toctree_backup", +] +all_projects = list(multiproject_projects.keys()) +for project in all_projects: + if project != docset: + exclude_patterns.append(project) +if docset in ["flagrelease_en", "flagrelease_zh"]: + exclude_patterns.append("model_readmes") + +# flagcicd 项目:排除管理员专用页面(用户管理、模型管理) +if docset in ["flagcicd_en", "flagcicd_zh"]: + exclude_patterns.extend([ + "function-description/user-management.md", + "function-description/model-management.md", + "operation-guide/user-management.md", + "operation-guide/model-management.md", + ]) + +default_role = "obj" +intersphinx_cache_limit = 14 +intersphinx_timeout = 3 +intersphinx_mapping = { + "python": ("https://docs.python.org/3.10/", None), + "sphinx": ("https://www.sphinx-doc.org/en/master/", None), +} + +intersphinx_disabled_reftypes = ["*"] + +myst_frontmatter_process = "yaml" + +myst_enable_extensions = [ + "dollarmath", + "amsmath", + "deflist", + "fieldlist", + "html_admonition", + "html_image", + "colon_fence", + "smartquotes", + "replacements", + # "linkify", + "strikethrough", + "substitution", + "tasklist", + "attrs_inline", + "attrs_block", + # "substitution", +] + +htmlhelp_basename = "KernelGendoc" +latex_documents = [ + ( + "index", + "KernelGen.tex", + "KernelGen Documentation", + "KernelGen Team", + "manual", + ), +] +man_pages = [ + ( + "index", + "kernelgen", + "KernelGen Documentation", + ["KernelGen Team"], + 1, + ) +] + +# Set language based on project suffix or environment variable (sphinx-intl support) +# language = os.environ.get("READTHEDOCS_LANGUAGE", "en") if docset == "flagos_homepage" else ("en" if docset.endswith("_en") else "zh_CN") +language = "en" + +# Detect the actual build language from Read the Docs environment variable +# Falls back to the language config variable for local builds +CURRENT_LANGUAGE = os.getenv("READTHEDOCS_LANGUAGE", language) + +# if docset == "flagos_homepage": +# is_zh = CURRENT_LANGUAGE in ["zh_CN", "zh", "zh-cn"] +# else: +# is_zh = docset.endswith("_zh") +# lang_prefix = "zh-cn" if is_zh else "en" + +# # 定义 myst_substitutions +# myst_substitutions = { +# "lang_prefix": lang_prefix, +# } + +# locale_dirs = [ +# f"{docset}/locale/", +# ] +gettext_compact = False + +html_short_title = "" + +# ============================================================================ +# HTML THEME CONFIGURATION - DIFFERENT THEMES FOR DIFFERENT PROJECTS +# ============================================================================ + +# Only flagos_homepage uses pydata_sphinx_theme, all others use sphinx_book_theme +if docset == "flagos_homepage": + html_theme = "pydata_sphinx_theme" +else: + html_theme = "sphinx_book_theme" + +# Common static paths +html_static_path = ["_static", f"{docset}/_static"] +html_css_files = ["custom.css", "homepage.css"] +if docset == "flagos_homepage": + html_css_files.append("guide.css") +html_js_files = [] + +# html_logo = "img/logo.png" +html_favicon = "_static/favicon.svg" + +# Theme-specific configurations +if html_theme == "pydata_sphinx_theme": + # PyData Sphinx Theme configuration for flagos_homepage + + # Set logo based on language (sphinx-intl support) + # Read the Docs sets READTHEDOCS_LANGUAGE environment variable during builds + # ReadTheDocs uses lowercase codes (zh, zh-cn), while Sphinx uses zh_CN + + # Debug output + print(f"DEBUG: CURRENT_LANGUAGE = '{CURRENT_LANGUAGE}'") + print(f"DEBUG: language = '{language}'") + print(f"DEBUG: READTHEDOCS_LANGUAGE env = '{os.getenv('READTHEDOCS_LANGUAGE', 'NOT SET')}'") + print(f"DEBUG: Is Chinese? {CURRENT_LANGUAGE in ['zh_CN', 'zh', 'zh-cn']}") + + if CURRENT_LANGUAGE in ["zh_CN", "zh", "zh-cn"]: + logo_config = { + "text": "文档中心", + "image_light": "_static/logo-zh-light.svg", + "image_dark": "_static/logo-zh-dark.svg", + } + print("DEBUG: Using CHINESE logo config") + else: + # Default to English configuration for all other languages + logo_config = { + "text": "Documentation", + "image_light": "_static/logo-en-light.svg", + "image_dark": "_static/logo-en-dark.svg", + } + print("DEBUG: Using ENGLISH logo config") + + html_theme_options = { + "logo": logo_config, + "home_page_in_toc": True, + "use_download_button": False, + "repository_url": "https://github.com/flagos-ai/KernelGen", + "use_repository_button": True, + "secondary_sidebar_items": { + "**": ["page-toc"], + "flagos_homepage/index": [], + }, + "show_toc_level": 2, + "footer_start": ["copyright"], + "footer_end": [], + "show_sphinx": False, + "navbar_end": ["navbar-icon-links"] + } + + # Keep the FlagOS homepage clean while enabling page-local TOC elsewhere. + + # html_sidebars is only for PyData Sphinx Theme + html_sidebars = {} + for project in all_projects: + html_sidebars[f"{project}/index"] = [] + + # html_context is only applied to PyData Sphinx Theme + html_context = { + "default_mode": "light" + } + + # No additional JS files needed for portal homepage + html_js_files = [] + +else: + # Sphinx Book Theme configuration for all other projects + + # # repo URL per project + # repository_urls = { + # "flagcx_en": "https://github.com/flagos-ai/FlagCX", + # "flagcx_zh": "https://github.com/flagos-ai/FlagCX", + # "flaggems_en": "https://github.com/flagos-ai/FlagGems", + # "flaggems_zh": "https://github.com/flagos-ai/FlagGems", + # "flagtree_en": "https://github.com/flagos-ai/FlagTree", + # "flagtree_zh": "https://github.com/flagos-ai/FlagTree", + # "flagrelease_en": "https://github.com/flagos-ai/FlagRelease", + # "flagrelease_zh": "https://github.com/flagos-ai/FlagRelease", + # "flagperf_en": "https://github.com/flagos-ai/FlagPerf", + # "flagperf_zh": "https://github.com/flagos-ai/FlagPerf", + # } + + # # Obtain the current repo URL, if failed, set the default value + # current_repo_url = repository_urls.get(docset, "https://github.com/flagos-ai") + + if docset.endswith("_en"): + main_site_url = "https://docs.flagos.io/en/latest/" + main_site_text = "Back to FlagOS Documentation" + else: + main_site_url = "https://docs.flagos.io/zh-cn/latest/" + main_site_text = "返回 FlagOS 文档" + + templates_path = ["_templates"] + + # Logo configuration per project + if docset in ["flagcicd_en", "flagcicd_zh"]: + logo_config = { + "image_light": "_static/flagcicd-logo-light.svg", + "image_dark": "_static/flagcicd-logo-dark.svg", + } + else: + logo_config = { + "image_light": "_static/logo-en-light.svg", + "image_dark": "_static/logo-en-dark.svg", + } + + # Sphinx Book Theme configuration for all other projects + html_theme_options = { + "logo": logo_config, + "home_page_in_toc": True, + "use_download_button": False, + "repository_url": "https://github.com/flagos-ai/docs", + "use_edit_page_button": True, + "use_repository_button": True, + "navbar_center": ["back_to_main.html"], + # "default_mode": "light", + } + + html_context = { + "main_site_url": main_site_url, + "main_site_text": main_site_text, + "default_mode": "light" + } + + # No html_sidebars for Sphinx Book Theme + html_sidebars = {} + # No html_context for Sphinx Book Theme + html_last_updated_fmt = '%b %d, %Y' + +rst_epilog = """ +.. |org_brand| replace:: KernelGen Community +.. |com_brand| replace:: KernelGen for Business +.. |git_providers_and| replace:: GitHub, Bitbucket, and GitLab +.. |git_providers_or| replace:: GitHub, Bitbucket, or GitLab +""" + +autosectionlabel_prefix_document = True + +linkcheck_retries = 2 +linkcheck_timeout = 1 +linkcheck_workers = 10 +linkcheck_ignore = [ + r"http://127\.0\.0\.1", + r"http://localhost", + r"https://github\.com.+?#L\d+", +] + +extlinks = { + "issue": ("https://github.com/armstrongttwalker-alt/test-i18n-KernelGen/issues/%s", "#%s"), +} + +suppress_warnings = ["epub.unknown_project_files"] diff --git a/docs/flagos_homepage/index.md b/docs/flagos_homepage/index.md index 1b61bf5eeb..ca680720cb 100644 --- a/docs/flagos_homepage/index.md +++ b/docs/flagos_homepage/index.md @@ -1,302 +1,302 @@ ---- -sd_hide_title: true ---- - -# Documentation - -:::{div} flagos-header -:align: center - -# FlagOS - -A unified, open-source system software stack designed for a variety of AI chips - -[FlagOS Overview](overview.md){ .flagos-outline-btn } -[Cloud Chip Adaptation Guide](chip_adaptation_guide/cloud_adaptation_guide_index.md){ .flagos-outline-btn } -[Edge Chip Adaptation Guide](chip_adaptation_guide/edge_adaptation_guide_index.md){ .flagos-outline-btn } -::: - -```{toctree} -:maxdepth: 1 -:class: flagos-guide-root-toctree - -chip_adaptation_guide/cloud_adaptation_guide_index.md -chip_adaptation_guide/edge_adaptation_guide_index.md -``` - -## FlagOS Core Libraries - -````{grid} 1 1 1 1 -:gutter: 3 -:class: flagos-grid-sd - -```{grid-item-card} Operator Libraries -:class-card: flagos-card-sd - -High-performance operator libraries optimized for diverse hardware backends. - -+++ -:::{div} operator-item -**General-Purpose Operator Library** - -**FlagGems** - -Triton-based general-purpose operator library. - -[View Documentation →](https://docs.flagos.io/projects/FlagGems/en/latest/) -::: - -:::{div} operator-item -**Fused Operator Libraries** - -**FlagGems-vllm** - -Optimized vLLM operators for multiple backends. - -[View Documentation →](https://docs.flagos.io/projects/FlagGems-vllm/en/latest/) -::: - -:::{div} operator-item -**Multi-Domain Operator Libraries** - -- **FlagDNN** — Deep learning operators. [View Documentation →](https://docs.flagos.io/projects/FlagDNN/en/latest/) -- **FlagBLAS** — BLAS numerical library. [View Documentation →](https://docs.flagos.io/projects/FlagBLAS/en/latest/) -- **FlagFFT** — GPU FFT library. [View Documentation →](https://docs.flagos.io/projects/FlagFFT/en/latest/) -- **FlagSparse** — Sparse computation. [View Documentation →](https://docs.flagos.io/projects/FlagSparse/en/latest/) -- **FlagTensor** — Tensor primitives. [View Documentation →](https://docs.flagos.io/projects/FlagTensor/en/latest/) -- **FlagAudio** — Audio processing. [View Documentation →](https://docs.flagos.io/projects/FlagAudio/en/latest/) -::: -``` -```` - -````{grid} 1 1 3 3 -:gutter: 3 -:class: flagos-grid-sd - -```{grid-item-card} Compiler -:class-card: flagos-card-sd - -**FlagTree** - -An open-source, unified compiler for multiple AI chips, advancing and expanding the Triton ecosystem across diverse hardware platforms. - -+++ -[View Documentation →](https://docs.flagos.io/projects/FlagTree/en/latest/){ .card-link-sd } -``` - -```{grid-item-card} Training & Inference Framework -:class-card: flagos-card-sd - -**FlagScale** - -A comprehensive toolkit designed to support the entire lifecycle of large models, from training to inference and deployment. - -+++ -[View Documentation →](https://docs.flagos.io/projects/FlagScale/en/latest/){ .card-link-sd } -``` - -```{grid-item-card} Communication Library -:class-card: flagos-card-sd - -**FlagCX** - -A scalable and adaptive unified communication library for cross-chip environments, delivering high-performance collective communication capabilities. - -+++ -[View Documentation →](https://docs.flagos.io/projects/FlagCX/en/latest/){ .card-link-sd } -``` -```` - ---- - -## FlagOS Plugins for Diverse Chips - -````{grid} 1 1 3 3 -:gutter: 3 -:class: flagos-grid-sd - -```{grid-item-card} vllm-plugin-FL -:class-card: flagos-card-sd - -A plugin for the vLLM inference/serving framework, built on FlagOS's unified multi-chip backend. - -+++ -[View Documentation →](https://docs.flagos.io/projects/vllm-plugin-FL/en/latest/){ .card-link-sd } -``` - -```{grid-item-card} Megatron-LM-FL -:class-card: flagos-card-sd - -A fork of Megatron-LM that introduces a plugin-based architecture for supporting diverse AI chips, built on top of FlagOS. - -+++ -[View Documentation →](https://docs.flagos.io/projects/Megatron-LM-FL/en/latest/){ .card-link-sd } -``` - -```{grid-item-card} TransformerEngine-FL -:class-card: flagos-card-sd - -A fork of TransformerEngine that introduces a plugin-based architecture for supporting diverse AI chips, built on top of FlagOS. - -+++ -[View Documentation →](https://docs.flagos.io/projects/TransformerEngine-FL/en/latest/){ .card-link-sd } -``` - -```{grid-item-card} verl-FL -:class-card: flagos-card-sd - -A fork of veRL that extends the upstream library with multi-chip/multi-hardware support via the FlagOS ecosystem. - -+++ -[View Documentation →](https://docs.flagos.io/projects/verl-FL/en/latest/){ .card-link-sd } -``` - -```{grid-item-card} PyTorch-Plugin-FL -:class-card: flagos-card-sd - -A custom PyTorch device plugin based on the PrivateUse1 extension mechanism, registering FlagGems high-performance Triton operators as the flagos device backend. - -+++ -[View Documentation →](https://docs.flagos.io/projects/PyTorch-Plugin-FL/en/latest/){ .card-link-sd } -``` - -```{grid-item-card} sglang-plugin-FL -:class-card: flagos-card-sd - -An out-of-tree (OOT) plugin for SGLang, built on FlagOS's unified multi-chip backend, extending SGLang's inference capabilities across diverse hardware platforms. - -+++ -[View Documentation →](https://docs.flagos.io/projects/sglang-plugin-FL/en/latest/){ .card-link-sd } -``` -```` - ---- - -## FlagOS Domain-Specific Projects - -````{grid} 1 1 3 3 -:gutter: 3 -:class: flagos-grid-sd - -```{grid-item-card} FlagOS-Robo -:class-card: flagos-card-sd - -An integrated training and inference framework for AI models used in robots, so-called Embodied Intelligence. - -+++ -[View Documentation →](https://docs.flagos.io/projects/FlagOS-Robo/en/latest/){ .card-link-sd } -``` - -```{grid-item-card} FlagQuantum -:class-card: flagos-card-sd - -A high-performance distributed quantum statevector simulator built on PyTorch, enabling quantum circuit simulation across multiple GPUs. - -+++ -[View Documentation →](https://docs.flagos.io/projects/FlagQuantum/en/latest/){ .card-link-sd } -``` -```` - ---- - -## FlagOS Developer Tools - -````{grid} 1 1 3 3 -:gutter: 3 -:class: flagos-grid-sd - -```{grid-item-card} KernelGen -:class-card: flagos-card-sd - -An operator auto-generation tool. - -+++ -[View Documentation →](https://docs.flagos.io/projects/kernelgen/en/latest/){ .card-link-sd } -``` - -```{grid-item-card} KernelGenBench -:class-card: flagos-card-sd - -A benchmark framework for evaluating LLM and agent-based Triton kernel generation across multiple hardware platforms. - -+++ -[View Documentation →](https://docs.flagos.io/projects/kernelgenbench/en/latest/){ .card-link-sd } -``` - -```{grid-item-card} FlagOS Skills -:class-card: flagos-card-sd -:link: https://github.com/flagos-ai/skills - -Compatible with Claude Code, Cursor, Codex, and any agent supporting the Agent Skills standard. - -+++ -[View Documentation →](https://github.com/flagos-ai/skills){ .card-link-sd } -``` - -```{grid-item-card} Online Laboratory -:class-card: flagos-card-sd - -An online laboratory providing cloud-based development environments. - -+++ -[View Documentation →](https://docs.flagos.io/projects/onlinelaboratory/en/latest/){ .card-link-sd } -``` - -```{grid-item-card} FlagPrism -:class-card: flagos-card-sd - -A multi-backend debugging and performance-analysis toolkit for Triton programs. - -+++ -[View Documentation →](https://docs.flagos.io/projects/FlagPrism/en/latest/){ .card-link-sd } -``` -```` - ---- - -## FlagOS Platform Services - -````{grid} 1 1 3 3 -:gutter: 3 -:class: flagos-grid-sd - -```{grid-item-card} FlagRelease -:class-card: flagos-card-sd - -An automated platform for the cross-chip migration and release of open-source large models. - -+++ -[View Documentation →](https://docs.flagos.io/projects/FlagRelease/en/latest/){ .card-link-sd } -``` - -```{grid-item-card} FlagPerf -:class-card: flagos-card-sd - -An integrated AI hardware evaluation engine. - -+++ -[View Documentation →](https://docs.flagos.io/projects/FlagPerf/en/latest/){ .card-link-sd } -``` - -```{grid-item-card} FlagCICD -:class-card: flagos-card-sd -:link: https://docs.flagos.io/projects/FlagCICD/zh-cn/latest/ - -A CI/CD toolchain that streamlines large-model development across diverse AI chips. - -+++ -[View Documentation →](https://docs.flagos.io/projects/FlagCICD/zh-cn/latest/){ .card-link-sd } -``` -```` - ---- - -:::{div} call-to-action -:align: center - -## Start to Use FlagOS - -Join us to co-build an open AI chip development ecosystem - -[FlagOS Homepage](https://flagos.io/){ .btn .btn-primary .btn-lg } -::: +--- +sd_hide_title: true +--- + +# Documentation + +:::{div} flagos-header +:align: center + +# FlagOS + +A unified, open-source system software stack designed for a variety of AI chips + +[FlagOS Overview](overview.md){ .flagos-outline-btn } +[Cloud Chip Adaptation Guide](chip_adaptation_guide/cloud_adaptation_guide_index.md){ .flagos-outline-btn } +[Edge Chip Adaptation Guide](chip_adaptation_guide/edge_adaptation_guide_index.md){ .flagos-outline-btn } +::: + +```{toctree} +:maxdepth: 1 +:class: flagos-guide-root-toctree + +chip_adaptation_guide/cloud_adaptation_guide_index.md +chip_adaptation_guide/edge_adaptation_guide_index.md +``` + +## FlagOS Core Libraries + +````{grid} 1 1 1 1 +:gutter: 3 +:class: flagos-grid-sd + +```{grid-item-card} Operator Libraries +:class-card: flagos-card-sd + +High-performance operator libraries optimized for diverse hardware backends. + ++++ +:::{div} operator-item +**General-Purpose Operator Library** + +**FlagGems** + +Triton-based general-purpose operator library. + +[View Documentation →](https://docs.flagos.io/projects/FlagGems/en/latest/) +::: + +:::{div} operator-item +**Fused Operator Libraries** + +**FlagGems-vllm** + +Optimized vLLM operators for multiple backends. + +[View Documentation →](https://docs.flagos.io/projects/FlagGems-vllm/en/latest/) +::: + +:::{div} operator-item +**Multi-Domain Operator Libraries** + +- **FlagDNN** — Deep learning operators. [View Documentation →](https://docs.flagos.io/projects/FlagDNN/en/latest/) +- **FlagBLAS** — BLAS numerical library. [View Documentation →](https://docs.flagos.io/projects/FlagBLAS/en/latest/) +- **FlagFFT** — GPU FFT library. [View Documentation →](https://docs.flagos.io/projects/FlagFFT/en/latest/) +- **FlagSparse** — Sparse computation. [View Documentation →](https://docs.flagos.io/projects/FlagSparse/en/latest/) +- **FlagTensor** — Tensor primitives. [View Documentation →](https://docs.flagos.io/projects/FlagTensor/en/latest/) +- **FlagAudio** — Audio processing. [View Documentation →](https://docs.flagos.io/projects/FlagAudio/en/latest/) +::: +``` +```` + +````{grid} 1 1 3 3 +:gutter: 3 +:class: flagos-grid-sd + +```{grid-item-card} Compiler +:class-card: flagos-card-sd + +**FlagTree** + +An open-source, unified compiler for multiple AI chips, advancing and expanding the Triton ecosystem across diverse hardware platforms. + ++++ +[View Documentation →](https://docs.flagos.io/projects/FlagTree/en/latest/){ .card-link-sd } +``` + +```{grid-item-card} Training & Inference Framework +:class-card: flagos-card-sd + +**FlagScale** + +A comprehensive toolkit designed to support the entire lifecycle of large models, from training to inference and deployment. + ++++ +[View Documentation →](https://docs.flagos.io/projects/FlagScale/en/latest/){ .card-link-sd } +``` + +```{grid-item-card} Communication Library +:class-card: flagos-card-sd + +**FlagCX** + +A scalable and adaptive unified communication library for cross-chip environments, delivering high-performance collective communication capabilities. + ++++ +[View Documentation →](https://docs.flagos.io/projects/FlagCX/en/latest/){ .card-link-sd } +``` +```` + +--- + +## FlagOS Plugins for Diverse Chips + +````{grid} 1 1 3 3 +:gutter: 3 +:class: flagos-grid-sd + +```{grid-item-card} vllm-plugin-FL +:class-card: flagos-card-sd + +A plugin for the vLLM inference/serving framework, built on FlagOS's unified multi-chip backend. + ++++ +[View Documentation →](https://docs.flagos.io/projects/vllm-plugin-FL/en/latest/){ .card-link-sd } +``` + +```{grid-item-card} Megatron-LM-FL +:class-card: flagos-card-sd + +A fork of Megatron-LM that introduces a plugin-based architecture for supporting diverse AI chips, built on top of FlagOS. + ++++ +[View Documentation →](https://docs.flagos.io/projects/Megatron-LM-FL/en/latest/){ .card-link-sd } +``` + +```{grid-item-card} TransformerEngine-FL +:class-card: flagos-card-sd + +A fork of TransformerEngine that introduces a plugin-based architecture for supporting diverse AI chips, built on top of FlagOS. + ++++ +[View Documentation →](https://docs.flagos.io/projects/TransformerEngine-FL/en/latest/){ .card-link-sd } +``` + +```{grid-item-card} verl-FL +:class-card: flagos-card-sd + +A fork of veRL that extends the upstream library with multi-chip/multi-hardware support via the FlagOS ecosystem. + ++++ +[View Documentation →](https://docs.flagos.io/projects/verl-FL/en/latest/){ .card-link-sd } +``` + +```{grid-item-card} Torch-FL +:class-card: flagos-card-sd + +A custom PyTorch device plugin based on the PrivateUse1 extension mechanism, registering FlagGems high-performance Triton operators as the flagos device backend. + ++++ +[View Documentation →](https://docs.flagos.io/projects/torch-FL/en/latest/){ .card-link-sd } +``` + +```{grid-item-card} sglang-plugin-FL +:class-card: flagos-card-sd + +An out-of-tree (OOT) plugin for SGLang, built on FlagOS's unified multi-chip backend, extending SGLang's inference capabilities across diverse hardware platforms. + ++++ +[View Documentation →](https://docs.flagos.io/projects/sglang-plugin-FL/en/latest/){ .card-link-sd } +``` +```` + +--- + +## FlagOS Domain-Specific Projects + +````{grid} 1 1 3 3 +:gutter: 3 +:class: flagos-grid-sd + +```{grid-item-card} FlagOS-Robo +:class-card: flagos-card-sd + +An integrated training and inference framework for AI models used in robots, so-called Embodied Intelligence. + ++++ +[View Documentation →](https://docs.flagos.io/projects/FlagOS-Robo/en/latest/){ .card-link-sd } +``` + +```{grid-item-card} FlagQuantum +:class-card: flagos-card-sd + +A high-performance distributed quantum statevector simulator built on PyTorch, enabling quantum circuit simulation across multiple GPUs. + ++++ +[View Documentation →](https://docs.flagos.io/projects/FlagQuantum/en/latest/){ .card-link-sd } +``` +```` + +--- + +## FlagOS Developer Tools + +````{grid} 1 1 3 3 +:gutter: 3 +:class: flagos-grid-sd + +```{grid-item-card} KernelGen +:class-card: flagos-card-sd + +An operator auto-generation tool. + ++++ +[View Documentation →](https://docs.flagos.io/projects/kernelgen/en/latest/){ .card-link-sd } +``` + +```{grid-item-card} KernelGenBench +:class-card: flagos-card-sd + +A benchmark framework for evaluating LLM and agent-based Triton kernel generation across multiple hardware platforms. + ++++ +[View Documentation →](https://docs.flagos.io/projects/kernelgenbench/en/latest/){ .card-link-sd } +``` + +```{grid-item-card} FlagOS Skills +:class-card: flagos-card-sd +:link: https://github.com/flagos-ai/skills + +Compatible with Claude Code, Cursor, Codex, and any agent supporting the Agent Skills standard. + ++++ +[View Documentation →](https://github.com/flagos-ai/skills){ .card-link-sd } +``` + +```{grid-item-card} Online Laboratory +:class-card: flagos-card-sd + +An online laboratory providing cloud-based development environments. + ++++ +[View Documentation →](https://docs.flagos.io/projects/onlinelaboratory/en/latest/){ .card-link-sd } +``` + +```{grid-item-card} FlagPrism +:class-card: flagos-card-sd + +A multi-backend debugging and performance-analysis toolkit for Triton programs. + ++++ +[View Documentation →](https://docs.flagos.io/projects/FlagPrism/en/latest/){ .card-link-sd } +``` +```` + +--- + +## FlagOS Platform Services + +````{grid} 1 1 3 3 +:gutter: 3 +:class: flagos-grid-sd + +```{grid-item-card} FlagRelease +:class-card: flagos-card-sd + +An automated platform for the cross-chip migration and release of open-source large models. + ++++ +[View Documentation →](https://docs.flagos.io/projects/FlagRelease/en/latest/){ .card-link-sd } +``` + +```{grid-item-card} FlagPerf +:class-card: flagos-card-sd + +An integrated AI hardware evaluation engine. + ++++ +[View Documentation →](https://docs.flagos.io/projects/FlagPerf/en/latest/){ .card-link-sd } +``` + +```{grid-item-card} FlagCICD +:class-card: flagos-card-sd +:link: https://docs.flagos.io/projects/FlagCICD/zh-cn/latest/ + +A CI/CD toolchain that streamlines large-model development across diverse AI chips. + ++++ +[View Documentation →](https://docs.flagos.io/projects/FlagCICD/zh-cn/latest/){ .card-link-sd } +``` +```` + +--- + +:::{div} call-to-action +:align: center + +## Start to Use FlagOS + +Join us to co-build an open AI chip development ecosystem + +[FlagOS Homepage](https://flagos.io/){ .btn .btn-primary .btn-lg } +::: diff --git a/docs/flagos_homepage/locale/zh_CN/LC_MESSAGES/index.po b/docs/flagos_homepage/locale/zh_CN/LC_MESSAGES/index.po index d99383e88d..a88882ff83 100644 --- a/docs/flagos_homepage/locale/zh_CN/LC_MESSAGES/index.po +++ b/docs/flagos_homepage/locale/zh_CN/LC_MESSAGES/index.po @@ -1,549 +1,549 @@ -# FlagOS Documentation Chinese Translation -# Copyright (C) 2025-2026, FlagOS Community -# This file is distributed under the same license as the FlagOS -# Documentation package. -# FlagOS Community , 2026. -# -msgid "" -msgstr "" -"Project-Id-Version: FlagOS Documentation \n" -"Report-Msgid-Bugs-To: \n" -"POT-Creation-Date: 2026-09-20 17:05+0800\n" -"PO-Revision-Date: 2026-06-23 15:10+0800\n" -"Last-Translator: FlagOS Community \n" -"Language: zh_CN\n" -"Language-Team: zh_CN \n" -"Plural-Forms: nplurals=1; plural=0;\n" -"MIME-Version: 1.0\n" -"Content-Type: text/plain; charset=utf-8\n" -"Content-Transfer-Encoding: 8bit\n" -"Generated-By: Babel 2.17.0\n" - -#: ../../index.md:5 -msgid "Documentation" -msgstr "文档中心" - -#: ../../index.md:10 -msgid "FlagOS" -msgstr "FlagOS" - -#: ../../index.md:12 -msgid "" -"A unified, open-source system software stack designed for a variety of AI" -" chips" -msgstr "一个统一的开源系统软件栈,专为多种 AI 芯片设计" - -#: ../../index.md:14 -#, python-brace-format -msgid "" -"[FlagOS Overview](overview.md){ .flagos-outline-btn } [Cloud Chip " -"Adaptation Guide](chip_adaptation_guide/cloud_adaptation_guide_index.md){" -" .flagos-outline-btn } [Edge Chip Adaptation " -"Guide](chip_adaptation_guide/edge_adaptation_guide_index.md){ .flagos-" -"outline-btn }" -msgstr "" -"[FlagOS 概览](overview.md){ .flagos-outline-btn } [云端芯片适配指南](chip_adaptation_guide/cloud_adaptation_guide_index.md){" -" .flagos-outline-btn } [端侧芯片适配指南](chip_adaptation_guide/edge_adaptation_guide_index.md){ .flagos-" -"outline-btn }" - -#: ../../index.md:27 -msgid "FlagOS Core Libraries" -msgstr "FlagOS 核心库" - -#: ../../index.md:33 -msgid "Operator Libraries" -msgstr "算子库" - -#: ../../index.md:36 -msgid "" -"High-performance operator libraries optimized for diverse hardware " -"backends." -msgstr "面向多种硬件后端优化的高性能算子库。" - -#: ../../index.md:40 -msgid "**General-Purpose Operator Library**" -msgstr "**通用算子库**" - -#: ../../index.md:42 -msgid "**FlagGems**" -msgstr "**FlagGems**" - -#: ../../index.md:44 -msgid "Triton-based general-purpose operator library." -msgstr "基于 Triton 的通用算子库。" - -#: ../../index.md:46 -msgid "" -"[View Documentation " -"→](https://docs.flagos.io/projects/FlagGems/en/latest/)" -msgstr "[查看文档 →](https://docs.flagos.io/projects/FlagGems/zh-cn/latest/)" - -#: ../../index.md:50 -msgid "**Fused Operator Libraries**" -msgstr "**融合算子库**" - -#: ../../index.md:52 -msgid "**FlagGems-vllm**" -msgstr "**FlagGems-vllm**" - -#: ../../index.md:54 -msgid "Optimized vLLM operators for multiple backends." -msgstr "面向多种后端优化的 vLLM 算子。" - -#: ../../index.md:56 -msgid "" -"[View Documentation →](https://docs.flagos.io/projects/FlagGems-" -"vllm/en/latest/)" -msgstr "[查看文档 →](https://docs.flagos.io/projects/FlagGems-vllm/zh-cn/latest/)" - -#: ../../index.md:60 -msgid "**Multi-Domain Operator Libraries**" -msgstr "**多领域算子库**" - -#: ../../index.md:62 -msgid "" -"**FlagDNN** — Deep learning operators. [View Documentation " -"→](https://docs.flagos.io/projects/FlagDNN/en/latest/)" -msgstr "" -"**FlagDNN** — 深度学习算子。[查看文档 →](https://docs.flagos.io/projects/FlagDNN/zh-" -"cn/latest/)" - -#: ../../index.md:63 -msgid "" -"**FlagBLAS** — BLAS numerical library. [View Documentation " -"→](https://docs.flagos.io/projects/FlagBLAS/en/latest/)" -msgstr "" -"**FlagBLAS** — BLAS 数值库。[查看文档 →](https://docs.flagos.io/projects/FlagBLAS" -"/zh-cn/latest/)" - -#: ../../index.md:64 -msgid "" -"**FlagFFT** — GPU FFT library. [View Documentation " -"→](https://docs.flagos.io/projects/FlagFFT/en/latest/)" -msgstr "" -"**FlagFFT** — GPU FFT 库。[查看文档 →](https://docs.flagos.io/projects/FlagFFT" -"/zh-cn/latest/)" - -#: ../../index.md:65 -msgid "" -"**FlagSparse** — Sparse computation. [View Documentation " -"→](https://docs.flagos.io/projects/FlagSparse/en/latest/)" -msgstr "" -"**FlagSparse** — 稀疏计算。[查看文档 →](https://docs.flagos.io/projects/FlagSparse" -"/zh-cn/latest/)" - -#: ../../index.md:66 -msgid "" -"**FlagTensor** — Tensor primitives. [View Documentation " -"→](https://docs.flagos.io/projects/FlagTensor/en/latest/)" -msgstr "" -"**FlagTensor** — 张量原语。[查看文档 →](https://docs.flagos.io/projects/FlagTensor" -"/zh-cn/latest/)" - -#: ../../index.md:67 -msgid "" -"**FlagAudio** — Audio processing. [View Documentation " -"→](https://docs.flagos.io/projects/FlagAudio/en/latest/)" -msgstr "" -"**FlagAudio** — 音频处理。[查看文档 →](https://docs.flagos.io/projects/FlagAudio" -"/zh-cn/latest/)" - -#: ../../index.md:76 -msgid "Compiler" -msgstr "编译器" - -#: ../../index.md:79 -msgid "**FlagTree**" -msgstr "**FlagTree**" - -#: ../../index.md:81 -msgid "" -"An open-source, unified compiler for multiple AI chips, advancing and " -"expanding the Triton ecosystem across diverse hardware platforms." -msgstr "面向多种 AI 芯片的开源统一编译器,在多种硬件平台上推进和扩展 Triton 生态系统。" - -#: ../../index.md:84 -#, python-brace-format -msgid "" -"[View Documentation " -"→](https://docs.flagos.io/projects/FlagTree/en/latest/){ .card-link-sd }" -msgstr "" -"[查看文档 →](https://docs.flagos.io/projects/FlagTree/zh-cn/latest/){ .card-" -"link-sd }" - -#: ../../index.md:87 -msgid "Training & Inference Framework" -msgstr "训练与推理框架" - -#: ../../index.md:90 -msgid "**FlagScale**" -msgstr "**FlagScale**" - -#: ../../index.md:92 -msgid "" -"A comprehensive toolkit designed to support the entire lifecycle of large" -" models, from training to inference and deployment." -msgstr "全面支持大模型全生命周期的工具集,涵盖从训练到推理和部署。" - -#: ../../index.md:95 -#, python-brace-format -msgid "" -"[View Documentation " -"→](https://docs.flagos.io/projects/FlagScale/en/latest/){ .card-link-sd }" -msgstr "" -"[查看文档 →](https://docs.flagos.io/projects/FlagScale/zh-cn/latest/){ .card-" -"link-sd }" - -#: ../../index.md:98 -msgid "Communication Library" -msgstr "通信库" - -#: ../../index.md:101 -msgid "**FlagCX**" -msgstr "**FlagCX**" - -#: ../../index.md:103 -msgid "" -"A scalable and adaptive unified communication library for cross-chip " -"environments, delivering high-performance collective communication " -"capabilities." -msgstr "面向跨芯片环境的可扩展自适应统一通信库,提供高性能集合通信能力。" - -#: ../../index.md:106 -#, python-brace-format -msgid "" -"[View Documentation " -"→](https://docs.flagos.io/projects/FlagCX/en/latest/){ .card-link-sd }" -msgstr "" -"[查看文档 →](https://docs.flagos.io/projects/FlagCX/zh-cn/latest/){ .card-" -"link-sd }" - -#: ../../index.md:112 -msgid "FlagOS Plugins for Diverse Chips" -msgstr "FlagOS 多芯片插件" - -#: ../../index.md:118 -msgid "vllm-plugin-FL" -msgstr "vllm-plugin-FL" - -#: ../../index.md:121 -msgid "" -"A plugin for the vLLM inference/serving framework, built on FlagOS's " -"unified multi-chip backend." -msgstr "vLLM 推理/服务框架插件,基于 FlagOS 统一多芯片后端构建。" - -#: ../../index.md:124 -#, python-brace-format -msgid "" -"[View Documentation →](https://docs.flagos.io/projects/vllm-plugin-" -"FL/en/latest/){ .card-link-sd }" -msgstr "" -"[查看文档 →](https://docs.flagos.io/projects/vllm-plugin-FL/zh-cn/latest/){ " -".card-link-sd }" - -#: ../../index.md:127 -msgid "Megatron-LM-FL" -msgstr "Megatron-LM-FL" - -#: ../../index.md:130 -msgid "" -"A fork of Megatron-LM that introduces a plugin-based architecture for " -"supporting diverse AI chips, built on top of FlagOS." -msgstr "Megatron-LM 的分支,引入插件化架构以支持多种 AI 芯片,基于 FlagOS 构建。" - -#: ../../index.md:133 -#, python-brace-format -msgid "" -"[View Documentation →](https://docs.flagos.io/projects/Megatron-LM-" -"FL/en/latest/){ .card-link-sd }" -msgstr "" -"[查看文档 →](https://docs.flagos.io/projects/Megatron-LM-FL/zh-cn/latest/){ " -".card-link-sd }" - -#: ../../index.md:136 -msgid "TransformerEngine-FL" -msgstr "TransformerEngine-FL" - -#: ../../index.md:139 -msgid "" -"A fork of TransformerEngine that introduces a plugin-based architecture " -"for supporting diverse AI chips, built on top of FlagOS." -msgstr "TransformerEngine 的分支,引入插件化架构以支持多种 AI 芯片,基于 FlagOS 构建。" - -#: ../../index.md:142 -#, python-brace-format -msgid "" -"[View Documentation →](https://docs.flagos.io/projects/TransformerEngine-" -"FL/en/latest/){ .card-link-sd }" -msgstr "" -"[查看文档 →](https://docs.flagos.io/projects/TransformerEngine-FL/zh-" -"cn/latest/){ .card-link-sd }" - -#: ../../index.md:145 -msgid "verl-FL" -msgstr "verl-FL" - -#: ../../index.md:148 -msgid "" -"A fork of veRL that extends the upstream library with multi-chip/multi-" -"hardware support via the FlagOS ecosystem." -msgstr "veRL 的分支,通过 FlagOS 生态系统扩展上游库的多芯片/多硬件支持。" - -#: ../../index.md:151 -#, python-brace-format -msgid "" -"[View Documentation →](https://docs.flagos.io/projects/verl-" -"FL/en/latest/){ .card-link-sd }" -msgstr "" -"[查看文档 →](https://docs.flagos.io/projects/verl-FL/zh-cn/latest/){ .card-" -"link-sd }" - -#: ../../index.md:154 -msgid "PyTorch-Plugin-FL" -msgstr "PyTorch-Plugin-FL" - -#: ../../index.md:157 -msgid "" -"A custom PyTorch device plugin based on the PrivateUse1 extension " -"mechanism, registering FlagGems high-performance Triton operators as the " -"flagos device backend." -msgstr "" -"基于 PrivateUse1 扩展机制的自定义 PyTorch 设备插件,将 FlagGems 高性能 Triton 算子注册为 flagos " -"设备后端。" - -#: ../../index.md:160 -#, python-brace-format -msgid "" -"[View Documentation →](https://docs.flagos.io/projects/PyTorch-Plugin-" -"FL/en/latest/){ .card-link-sd }" -msgstr "" -"[查看文档 →](https://docs.flagos.io/projects/PyTorch-Plugin-FL/zh-" -"cn/latest/){ .card-link-sd }" - -#: ../../index.md:163 -msgid "sglang-plugin-FL" -msgstr "sglang-plugin-FL" - -#: ../../index.md:166 -msgid "" -"An out-of-tree (OOT) plugin for SGLang, built on FlagOS's unified multi-" -"chip backend, extending SGLang's inference capabilities across diverse " -"hardware platforms." -msgstr "SGLang 的树外 (OOT) 插件,基于 FlagOS 统一多芯片后端构建,将 SGLang 的推理能力扩展到多种硬件平台。" - -#: ../../index.md:169 -#, python-brace-format -msgid "" -"[View Documentation →](https://docs.flagos.io/projects/sglang-plugin-" -"FL/en/latest/){ .card-link-sd }" -msgstr "" -"[查看文档 →](https://docs.flagos.io/projects/sglang-plugin-FL/zh-cn/latest/){" -" .card-link-sd }" - -#: ../../index.md:175 -msgid "FlagOS Domain-Specific Projects" -msgstr "FlagOS 领域专用项目" - -#: ../../index.md:181 -msgid "FlagOS-Robo" -msgstr "FlagOS-Robo" - -#: ../../index.md:184 -msgid "" -"An integrated training and inference framework for AI models used in " -"robots, so-called Embodied Intelligence." -msgstr "面向机器人 AI 模型的集成训练与推理框架,即具身智能。" - -#: ../../index.md:187 -#, python-brace-format -msgid "" -"[View Documentation →](https://docs.flagos.io/projects/FlagOS-" -"Robo/en/latest/){ .card-link-sd }" -msgstr "" -"[查看文档 →](https://docs.flagos.io/projects/FlagOS-Robo/zh-cn/latest/){ " -".card-link-sd }" - -#: ../../index.md:190 -msgid "FlagQuantum" -msgstr "FlagQuantum" - -#: ../../index.md:193 -msgid "" -"A high-performance distributed quantum statevector simulator built on " -"PyTorch, enabling quantum circuit simulation across multiple GPUs." -msgstr "基于 PyTorch 构建的高性能分布式量子态矢量模拟器,支持跨多 GPU 的量子电路模拟。" - -#: ../../index.md:196 -#, python-brace-format -msgid "" -"[View Documentation " -"→](https://docs.flagos.io/projects/FlagQuantum/en/latest/){ .card-link-sd" -" }" -msgstr "" -"[查看文档 →](https://docs.flagos.io/projects/FlagQuantum/zh-cn/latest/){ " -".card-link-sd }" - -#: ../../index.md:202 -msgid "FlagOS Developer Tools" -msgstr "FlagOS 开发者工具" - -#: ../../index.md:208 -msgid "KernelGen" -msgstr "KernelGen" - -#: ../../index.md:211 -msgid "An operator auto-generation tool." -msgstr "算子自动生成工具。" - -#: ../../index.md:214 -#, python-brace-format -msgid "" -"[View Documentation " -"→](https://docs.flagos.io/projects/kernelgen/en/latest/){ .card-link-sd }" -msgstr "" -"[查看文档 →](https://docs.flagos.io/projects/kernelgen/zh-cn/latest/){ .card-" -"link-sd }" - -#: ../../index.md:217 -msgid "KernelGenBench" -msgstr "KernelGenBench" - -#: ../../index.md:220 -msgid "" -"A benchmark framework for evaluating LLM and agent-based Triton kernel " -"generation across multiple hardware platforms." -msgstr "用于评估跨多硬件平台 LLM 和 Agent 驱动的 Triton kernel 生成的基准测试框架。" - -#: ../../index.md:223 -#, python-brace-format -msgid "" -"[View Documentation " -"→](https://docs.flagos.io/projects/kernelgenbench/en/latest/){ .card-" -"link-sd }" -msgstr "" -"[查看文档 →](https://docs.flagos.io/projects/kernelgenbench/zh-cn/latest/){ " -".card-link-sd }" - -#: ../../index.md:226 -msgid "FlagOS Skills" -msgstr "FlagOS Skills" - -#: ../../index.md:230 -msgid "" -"Compatible with Claude Code, Cursor, Codex, and any agent supporting the " -"Agent Skills standard." -msgstr "兼容 Claude Code、Cursor、Codex 以及任何支持 Agent Skills 标准的 Agent。" - -#: ../../index.md:233 -#, python-brace-format -msgid "" -"[View Documentation →](https://github.com/flagos-ai/skills){ .card-link-" -"sd }" -msgstr "[查看文档 →](https://github.com/flagos-ai/skills){ .card-link-sd }" - -#: ../../index.md:236 -msgid "Online Laboratory" -msgstr "线上实验室" - -#: ../../index.md:239 -msgid "An online laboratory providing cloud-based development environments." -msgstr "提供云端开发环境的线上实验室。" - -#: ../../index.md:242 -#, python-brace-format -msgid "" -"[View Documentation " -"→](https://docs.flagos.io/projects/onlinelaboratory/en/latest/){ .card-" -"link-sd }" -msgstr "" -"[查看文档 →](https://docs.flagos.io/projects/onlinelaboratory/zh-cn/latest/){" -" .card-link-sd }" - -msgid "FlagPrism" -msgstr "FlagPrism" - -msgid "" -"A multi-backend debugging and performance-analysis toolkit for Triton " -"programs." -msgstr "面向 Triton 程序的多后端调试与性能分析工具。" - -#, python-brace-format -msgid "" -"[View Documentation " -"→](https://docs.flagos.io/projects/FlagPrism/en/latest/){ .card-link-sd }" -msgstr "" -"[查看文档 →](https://docs.flagos.io/projects/FlagPrism/zh-cn/latest/){ " -".card-link-sd }" - -#: ../../flagos_homepage/index.md:238 -msgid "FlagOS Platform Services" -msgstr "FlagOS 平台服务" - -#: ../../index.md:254 -msgid "FlagRelease" -msgstr "FlagRelease" - -#: ../../index.md:257 -msgid "" -"An automated platform for the cross-chip migration and release of open-" -"source large models." -msgstr "开源大模型跨芯片迁移与发布的自动化平台。" - -#: ../../index.md:260 -#, python-brace-format -msgid "" -"[View Documentation " -"→](https://docs.flagos.io/projects/FlagRelease/en/latest/){ .card-link-sd" -" }" -msgstr "" -"[查看文档 →](https://docs.flagos.io/projects/FlagRelease/zh-cn/latest/){ " -".card-link-sd }" - -#: ../../index.md:263 -msgid "FlagPerf" -msgstr "FlagPerf" - -#: ../../index.md:266 -msgid "An integrated AI hardware evaluation engine." -msgstr "集成式 AI 硬件评测引擎。" - -#: ../../index.md:269 -#, python-brace-format -msgid "" -"[View Documentation " -"→](https://docs.flagos.io/projects/FlagPerf/en/latest/){ .card-link-sd }" -msgstr "" -"[查看文档 →](https://docs.flagos.io/projects/FlagPerf/zh-cn/latest/){ .card-" -"link-sd }" - -#: ../../index.md:272 -msgid "FlagCICD" -msgstr "FlagCICD" - -#: ../../index.md:276 -msgid "" -"A CI/CD toolchain that streamlines large-model development across diverse" -" AI chips." -msgstr "简化跨多种 AI 芯片大模型开发的 CI/CD 工具链。" - -#: ../../index.md:279 -#, python-brace-format -msgid "" -"[View Documentation →](https://docs.flagos.io/projects/FlagCICD/zh-" -"cn/latest/){ .card-link-sd }" -msgstr "" -"[查看文档 →](https://docs.flagos.io/projects/FlagCICD/zh-cn/latest/){ .card-" -"link-sd }" - -#: ../../index.md:288 -msgid "Start to Use FlagOS" -msgstr "开始使用 FlagOS" - -#: ../../index.md:290 -msgid "Join us to co-build an open AI chip development ecosystem" -msgstr "加入我们,共建开放 AI 芯片开发生态" - -#: ../../index.md:292 -#, python-brace-format -msgid "[FlagOS Homepage](https://flagos.io/){ .btn .btn-primary .btn-lg }" -msgstr "[FlagOS 主页](https://flagos.io/){ .btn .btn-primary .btn-lg }" +# FlagOS Documentation Chinese Translation +# Copyright (C) 2025-2026, FlagOS Community +# This file is distributed under the same license as the FlagOS +# Documentation package. +# FlagOS Community , 2026. +# +msgid "" +msgstr "" +"Project-Id-Version: FlagOS Documentation \n" +"Report-Msgid-Bugs-To: \n" +"POT-Creation-Date: 2026-09-20 17:05+0800\n" +"PO-Revision-Date: 2026-06-23 15:10+0800\n" +"Last-Translator: FlagOS Community \n" +"Language: zh_CN\n" +"Language-Team: zh_CN \n" +"Plural-Forms: nplurals=1; plural=0;\n" +"MIME-Version: 1.0\n" +"Content-Type: text/plain; charset=utf-8\n" +"Content-Transfer-Encoding: 8bit\n" +"Generated-By: Babel 2.17.0\n" + +#: ../../index.md:5 +msgid "Documentation" +msgstr "文档中心" + +#: ../../index.md:10 +msgid "FlagOS" +msgstr "FlagOS" + +#: ../../index.md:12 +msgid "" +"A unified, open-source system software stack designed for a variety of AI" +" chips" +msgstr "一个统一的开源系统软件栈,专为多种 AI 芯片设计" + +#: ../../index.md:14 +#, python-brace-format +msgid "" +"[FlagOS Overview](overview.md){ .flagos-outline-btn } [Cloud Chip " +"Adaptation Guide](chip_adaptation_guide/cloud_adaptation_guide_index.md){" +" .flagos-outline-btn } [Edge Chip Adaptation " +"Guide](chip_adaptation_guide/edge_adaptation_guide_index.md){ .flagos-" +"outline-btn }" +msgstr "" +"[FlagOS 概览](overview.md){ .flagos-outline-btn } [云端芯片适配指南](chip_adaptation_guide/cloud_adaptation_guide_index.md){" +" .flagos-outline-btn } [端侧芯片适配指南](chip_adaptation_guide/edge_adaptation_guide_index.md){ .flagos-" +"outline-btn }" + +#: ../../index.md:27 +msgid "FlagOS Core Libraries" +msgstr "FlagOS 核心库" + +#: ../../index.md:33 +msgid "Operator Libraries" +msgstr "算子库" + +#: ../../index.md:36 +msgid "" +"High-performance operator libraries optimized for diverse hardware " +"backends." +msgstr "面向多种硬件后端优化的高性能算子库。" + +#: ../../index.md:40 +msgid "**General-Purpose Operator Library**" +msgstr "**通用算子库**" + +#: ../../index.md:42 +msgid "**FlagGems**" +msgstr "**FlagGems**" + +#: ../../index.md:44 +msgid "Triton-based general-purpose operator library." +msgstr "基于 Triton 的通用算子库。" + +#: ../../index.md:46 +msgid "" +"[View Documentation " +"→](https://docs.flagos.io/projects/FlagGems/en/latest/)" +msgstr "[查看文档 →](https://docs.flagos.io/projects/FlagGems/zh-cn/latest/)" + +#: ../../index.md:50 +msgid "**Fused Operator Libraries**" +msgstr "**融合算子库**" + +#: ../../index.md:52 +msgid "**FlagGems-vllm**" +msgstr "**FlagGems-vllm**" + +#: ../../index.md:54 +msgid "Optimized vLLM operators for multiple backends." +msgstr "面向多种后端优化的 vLLM 算子。" + +#: ../../index.md:56 +msgid "" +"[View Documentation →](https://docs.flagos.io/projects/FlagGems-" +"vllm/en/latest/)" +msgstr "[查看文档 →](https://docs.flagos.io/projects/FlagGems-vllm/zh-cn/latest/)" + +#: ../../index.md:60 +msgid "**Multi-Domain Operator Libraries**" +msgstr "**多领域算子库**" + +#: ../../index.md:62 +msgid "" +"**FlagDNN** — Deep learning operators. [View Documentation " +"→](https://docs.flagos.io/projects/FlagDNN/en/latest/)" +msgstr "" +"**FlagDNN** — 深度学习算子。[查看文档 →](https://docs.flagos.io/projects/FlagDNN/zh-" +"cn/latest/)" + +#: ../../index.md:63 +msgid "" +"**FlagBLAS** — BLAS numerical library. [View Documentation " +"→](https://docs.flagos.io/projects/FlagBLAS/en/latest/)" +msgstr "" +"**FlagBLAS** — BLAS 数值库。[查看文档 →](https://docs.flagos.io/projects/FlagBLAS" +"/zh-cn/latest/)" + +#: ../../index.md:64 +msgid "" +"**FlagFFT** — GPU FFT library. [View Documentation " +"→](https://docs.flagos.io/projects/FlagFFT/en/latest/)" +msgstr "" +"**FlagFFT** — GPU FFT 库。[查看文档 →](https://docs.flagos.io/projects/FlagFFT" +"/zh-cn/latest/)" + +#: ../../index.md:65 +msgid "" +"**FlagSparse** — Sparse computation. [View Documentation " +"→](https://docs.flagos.io/projects/FlagSparse/en/latest/)" +msgstr "" +"**FlagSparse** — 稀疏计算。[查看文档 →](https://docs.flagos.io/projects/FlagSparse" +"/zh-cn/latest/)" + +#: ../../index.md:66 +msgid "" +"**FlagTensor** — Tensor primitives. [View Documentation " +"→](https://docs.flagos.io/projects/FlagTensor/en/latest/)" +msgstr "" +"**FlagTensor** — 张量原语。[查看文档 →](https://docs.flagos.io/projects/FlagTensor" +"/zh-cn/latest/)" + +#: ../../index.md:67 +msgid "" +"**FlagAudio** — Audio processing. [View Documentation " +"→](https://docs.flagos.io/projects/FlagAudio/en/latest/)" +msgstr "" +"**FlagAudio** — 音频处理。[查看文档 →](https://docs.flagos.io/projects/FlagAudio" +"/zh-cn/latest/)" + +#: ../../index.md:76 +msgid "Compiler" +msgstr "编译器" + +#: ../../index.md:79 +msgid "**FlagTree**" +msgstr "**FlagTree**" + +#: ../../index.md:81 +msgid "" +"An open-source, unified compiler for multiple AI chips, advancing and " +"expanding the Triton ecosystem across diverse hardware platforms." +msgstr "面向多种 AI 芯片的开源统一编译器,在多种硬件平台上推进和扩展 Triton 生态系统。" + +#: ../../index.md:84 +#, python-brace-format +msgid "" +"[View Documentation " +"→](https://docs.flagos.io/projects/FlagTree/en/latest/){ .card-link-sd }" +msgstr "" +"[查看文档 →](https://docs.flagos.io/projects/FlagTree/zh-cn/latest/){ .card-" +"link-sd }" + +#: ../../index.md:87 +msgid "Training & Inference Framework" +msgstr "训练与推理框架" + +#: ../../index.md:90 +msgid "**FlagScale**" +msgstr "**FlagScale**" + +#: ../../index.md:92 +msgid "" +"A comprehensive toolkit designed to support the entire lifecycle of large" +" models, from training to inference and deployment." +msgstr "全面支持大模型全生命周期的工具集,涵盖从训练到推理和部署。" + +#: ../../index.md:95 +#, python-brace-format +msgid "" +"[View Documentation " +"→](https://docs.flagos.io/projects/FlagScale/en/latest/){ .card-link-sd }" +msgstr "" +"[查看文档 →](https://docs.flagos.io/projects/FlagScale/zh-cn/latest/){ .card-" +"link-sd }" + +#: ../../index.md:98 +msgid "Communication Library" +msgstr "通信库" + +#: ../../index.md:101 +msgid "**FlagCX**" +msgstr "**FlagCX**" + +#: ../../index.md:103 +msgid "" +"A scalable and adaptive unified communication library for cross-chip " +"environments, delivering high-performance collective communication " +"capabilities." +msgstr "面向跨芯片环境的可扩展自适应统一通信库,提供高性能集合通信能力。" + +#: ../../index.md:106 +#, python-brace-format +msgid "" +"[View Documentation " +"→](https://docs.flagos.io/projects/FlagCX/en/latest/){ .card-link-sd }" +msgstr "" +"[查看文档 →](https://docs.flagos.io/projects/FlagCX/zh-cn/latest/){ .card-" +"link-sd }" + +#: ../../index.md:112 +msgid "FlagOS Plugins for Diverse Chips" +msgstr "FlagOS 多芯片插件" + +#: ../../index.md:118 +msgid "vllm-plugin-FL" +msgstr "vllm-plugin-FL" + +#: ../../index.md:121 +msgid "" +"A plugin for the vLLM inference/serving framework, built on FlagOS's " +"unified multi-chip backend." +msgstr "vLLM 推理/服务框架插件,基于 FlagOS 统一多芯片后端构建。" + +#: ../../index.md:124 +#, python-brace-format +msgid "" +"[View Documentation →](https://docs.flagos.io/projects/vllm-plugin-" +"FL/en/latest/){ .card-link-sd }" +msgstr "" +"[查看文档 →](https://docs.flagos.io/projects/vllm-plugin-FL/zh-cn/latest/){ " +".card-link-sd }" + +#: ../../index.md:127 +msgid "Megatron-LM-FL" +msgstr "Megatron-LM-FL" + +#: ../../index.md:130 +msgid "" +"A fork of Megatron-LM that introduces a plugin-based architecture for " +"supporting diverse AI chips, built on top of FlagOS." +msgstr "Megatron-LM 的分支,引入插件化架构以支持多种 AI 芯片,基于 FlagOS 构建。" + +#: ../../index.md:133 +#, python-brace-format +msgid "" +"[View Documentation →](https://docs.flagos.io/projects/Megatron-LM-" +"FL/en/latest/){ .card-link-sd }" +msgstr "" +"[查看文档 →](https://docs.flagos.io/projects/Megatron-LM-FL/zh-cn/latest/){ " +".card-link-sd }" + +#: ../../index.md:136 +msgid "TransformerEngine-FL" +msgstr "TransformerEngine-FL" + +#: ../../index.md:139 +msgid "" +"A fork of TransformerEngine that introduces a plugin-based architecture " +"for supporting diverse AI chips, built on top of FlagOS." +msgstr "TransformerEngine 的分支,引入插件化架构以支持多种 AI 芯片,基于 FlagOS 构建。" + +#: ../../index.md:142 +#, python-brace-format +msgid "" +"[View Documentation →](https://docs.flagos.io/projects/TransformerEngine-" +"FL/en/latest/){ .card-link-sd }" +msgstr "" +"[查看文档 →](https://docs.flagos.io/projects/TransformerEngine-FL/zh-" +"cn/latest/){ .card-link-sd }" + +#: ../../index.md:145 +msgid "verl-FL" +msgstr "verl-FL" + +#: ../../index.md:148 +msgid "" +"A fork of veRL that extends the upstream library with multi-chip/multi-" +"hardware support via the FlagOS ecosystem." +msgstr "veRL 的分支,通过 FlagOS 生态系统扩展上游库的多芯片/多硬件支持。" + +#: ../../index.md:151 +#, python-brace-format +msgid "" +"[View Documentation →](https://docs.flagos.io/projects/verl-" +"FL/en/latest/){ .card-link-sd }" +msgstr "" +"[查看文档 →](https://docs.flagos.io/projects/verl-FL/zh-cn/latest/){ .card-" +"link-sd }" + +#: ../../index.md:154 +msgid "Torch-FL" +msgstr "Torch-FL" + +#: ../../index.md:157 +msgid "" +"A custom PyTorch device plugin based on the PrivateUse1 extension " +"mechanism, registering FlagGems high-performance Triton operators as the " +"flagos device backend." +msgstr "" +"基于 PrivateUse1 扩展机制的自定义 PyTorch 设备插件,将 FlagGems 高性能 Triton 算子注册为 flagos " +"设备后端。" + +#: ../../index.md:160 +#, python-brace-format +msgid "" +"[View Documentation →](https://docs.flagos.io/projects/torch-" +"FL/en/latest/){ .card-link-sd }" +msgstr "" +"[查看文档 →](https://docs.flagos.io/projects/torch-FL/zh-" +"cn/latest/){ .card-link-sd }" + +#: ../../index.md:163 +msgid "sglang-plugin-FL" +msgstr "sglang-plugin-FL" + +#: ../../index.md:166 +msgid "" +"An out-of-tree (OOT) plugin for SGLang, built on FlagOS's unified multi-" +"chip backend, extending SGLang's inference capabilities across diverse " +"hardware platforms." +msgstr "SGLang 的树外 (OOT) 插件,基于 FlagOS 统一多芯片后端构建,将 SGLang 的推理能力扩展到多种硬件平台。" + +#: ../../index.md:169 +#, python-brace-format +msgid "" +"[View Documentation →](https://docs.flagos.io/projects/sglang-plugin-" +"FL/en/latest/){ .card-link-sd }" +msgstr "" +"[查看文档 →](https://docs.flagos.io/projects/sglang-plugin-FL/zh-cn/latest/){" +" .card-link-sd }" + +#: ../../index.md:175 +msgid "FlagOS Domain-Specific Projects" +msgstr "FlagOS 领域专用项目" + +#: ../../index.md:181 +msgid "FlagOS-Robo" +msgstr "FlagOS-Robo" + +#: ../../index.md:184 +msgid "" +"An integrated training and inference framework for AI models used in " +"robots, so-called Embodied Intelligence." +msgstr "面向机器人 AI 模型的集成训练与推理框架,即具身智能。" + +#: ../../index.md:187 +#, python-brace-format +msgid "" +"[View Documentation →](https://docs.flagos.io/projects/FlagOS-" +"Robo/en/latest/){ .card-link-sd }" +msgstr "" +"[查看文档 →](https://docs.flagos.io/projects/FlagOS-Robo/zh-cn/latest/){ " +".card-link-sd }" + +#: ../../index.md:190 +msgid "FlagQuantum" +msgstr "FlagQuantum" + +#: ../../index.md:193 +msgid "" +"A high-performance distributed quantum statevector simulator built on " +"PyTorch, enabling quantum circuit simulation across multiple GPUs." +msgstr "基于 PyTorch 构建的高性能分布式量子态矢量模拟器,支持跨多 GPU 的量子电路模拟。" + +#: ../../index.md:196 +#, python-brace-format +msgid "" +"[View Documentation " +"→](https://docs.flagos.io/projects/FlagQuantum/en/latest/){ .card-link-sd" +" }" +msgstr "" +"[查看文档 →](https://docs.flagos.io/projects/FlagQuantum/zh-cn/latest/){ " +".card-link-sd }" + +#: ../../index.md:202 +msgid "FlagOS Developer Tools" +msgstr "FlagOS 开发者工具" + +#: ../../index.md:208 +msgid "KernelGen" +msgstr "KernelGen" + +#: ../../index.md:211 +msgid "An operator auto-generation tool." +msgstr "算子自动生成工具。" + +#: ../../index.md:214 +#, python-brace-format +msgid "" +"[View Documentation " +"→](https://docs.flagos.io/projects/kernelgen/en/latest/){ .card-link-sd }" +msgstr "" +"[查看文档 →](https://docs.flagos.io/projects/kernelgen/zh-cn/latest/){ .card-" +"link-sd }" + +#: ../../index.md:217 +msgid "KernelGenBench" +msgstr "KernelGenBench" + +#: ../../index.md:220 +msgid "" +"A benchmark framework for evaluating LLM and agent-based Triton kernel " +"generation across multiple hardware platforms." +msgstr "用于评估跨多硬件平台 LLM 和 Agent 驱动的 Triton kernel 生成的基准测试框架。" + +#: ../../index.md:223 +#, python-brace-format +msgid "" +"[View Documentation " +"→](https://docs.flagos.io/projects/kernelgenbench/en/latest/){ .card-" +"link-sd }" +msgstr "" +"[查看文档 →](https://docs.flagos.io/projects/kernelgenbench/zh-cn/latest/){ " +".card-link-sd }" + +#: ../../index.md:226 +msgid "FlagOS Skills" +msgstr "FlagOS Skills" + +#: ../../index.md:230 +msgid "" +"Compatible with Claude Code, Cursor, Codex, and any agent supporting the " +"Agent Skills standard." +msgstr "兼容 Claude Code、Cursor、Codex 以及任何支持 Agent Skills 标准的 Agent。" + +#: ../../index.md:233 +#, python-brace-format +msgid "" +"[View Documentation →](https://github.com/flagos-ai/skills){ .card-link-" +"sd }" +msgstr "[查看文档 →](https://github.com/flagos-ai/skills){ .card-link-sd }" + +#: ../../index.md:236 +msgid "Online Laboratory" +msgstr "线上实验室" + +#: ../../index.md:239 +msgid "An online laboratory providing cloud-based development environments." +msgstr "提供云端开发环境的线上实验室。" + +#: ../../index.md:242 +#, python-brace-format +msgid "" +"[View Documentation " +"→](https://docs.flagos.io/projects/onlinelaboratory/en/latest/){ .card-" +"link-sd }" +msgstr "" +"[查看文档 →](https://docs.flagos.io/projects/onlinelaboratory/zh-cn/latest/){" +" .card-link-sd }" + +msgid "FlagPrism" +msgstr "FlagPrism" + +msgid "" +"A multi-backend debugging and performance-analysis toolkit for Triton " +"programs." +msgstr "面向 Triton 程序的多后端调试与性能分析工具。" + +#, python-brace-format +msgid "" +"[View Documentation " +"→](https://docs.flagos.io/projects/FlagPrism/en/latest/){ .card-link-sd }" +msgstr "" +"[查看文档 →](https://docs.flagos.io/projects/FlagPrism/zh-cn/latest/){ " +".card-link-sd }" + +#: ../../flagos_homepage/index.md:238 +msgid "FlagOS Platform Services" +msgstr "FlagOS 平台服务" + +#: ../../index.md:254 +msgid "FlagRelease" +msgstr "FlagRelease" + +#: ../../index.md:257 +msgid "" +"An automated platform for the cross-chip migration and release of open-" +"source large models." +msgstr "开源大模型跨芯片迁移与发布的自动化平台。" + +#: ../../index.md:260 +#, python-brace-format +msgid "" +"[View Documentation " +"→](https://docs.flagos.io/projects/FlagRelease/en/latest/){ .card-link-sd" +" }" +msgstr "" +"[查看文档 →](https://docs.flagos.io/projects/FlagRelease/zh-cn/latest/){ " +".card-link-sd }" + +#: ../../index.md:263 +msgid "FlagPerf" +msgstr "FlagPerf" + +#: ../../index.md:266 +msgid "An integrated AI hardware evaluation engine." +msgstr "集成式 AI 硬件评测引擎。" + +#: ../../index.md:269 +#, python-brace-format +msgid "" +"[View Documentation " +"→](https://docs.flagos.io/projects/FlagPerf/en/latest/){ .card-link-sd }" +msgstr "" +"[查看文档 →](https://docs.flagos.io/projects/FlagPerf/zh-cn/latest/){ .card-" +"link-sd }" + +#: ../../index.md:272 +msgid "FlagCICD" +msgstr "FlagCICD" + +#: ../../index.md:276 +msgid "" +"A CI/CD toolchain that streamlines large-model development across diverse" +" AI chips." +msgstr "简化跨多种 AI 芯片大模型开发的 CI/CD 工具链。" + +#: ../../index.md:279 +#, python-brace-format +msgid "" +"[View Documentation →](https://docs.flagos.io/projects/FlagCICD/zh-" +"cn/latest/){ .card-link-sd }" +msgstr "" +"[查看文档 →](https://docs.flagos.io/projects/FlagCICD/zh-cn/latest/){ .card-" +"link-sd }" + +#: ../../index.md:288 +msgid "Start to Use FlagOS" +msgstr "开始使用 FlagOS" + +#: ../../index.md:290 +msgid "Join us to co-build an open AI chip development ecosystem" +msgstr "加入我们,共建开放 AI 芯片开发生态" + +#: ../../index.md:292 +#, python-brace-format +msgid "[FlagOS Homepage](https://flagos.io/){ .btn .btn-primary .btn-lg }" +msgstr "[FlagOS 主页](https://flagos.io/){ .btn .btn-primary .btn-lg }" diff --git a/docs/flagos_homepage/locale/zh_CN/LC_MESSAGES/overview.po b/docs/flagos_homepage/locale/zh_CN/LC_MESSAGES/overview.po index 84b7cd496a..89e3a53b1e 100644 --- a/docs/flagos_homepage/locale/zh_CN/LC_MESSAGES/overview.po +++ b/docs/flagos_homepage/locale/zh_CN/LC_MESSAGES/overview.po @@ -1,451 +1,451 @@ -# FlagOS Documentation Chinese Translation -# Copyright (C) 2025-2026, FlagOS Community -# This file is distributed under the same license as the FlagOS -# Documentation package. -# FlagOS Community , 2026. -# -msgid "" -msgstr "" -"Project-Id-Version: FlagOS Documentation \n" -"Report-Msgid-Bugs-To: \n" -"POT-Creation-Date: 2026-09-20 17:05+0800\n" -"PO-Revision-Date: 2026-06-22 13:15+0800\n" -"Last-Translator: FlagOS Community \n" -"Language: zh_CN\n" -"Language-Team: zh_CN \n" -"Plural-Forms: nplurals=1; plural=0;\n" -"MIME-Version: 1.0\n" -"Content-Type: text/plain; charset=utf-8\n" -"Content-Transfer-Encoding: 8bit\n" -"Generated-By: Babel 2.17.0\n" - -#: ../../overview.md:1 -msgid "FlagOS Overview" -msgstr "FlagOS 概览" - -#: ../../overview.md:3 -msgid "" -"FlagOS is a fully open-source AI system software stack for heterogeneous " -"AI chips, allowing AI models to be developed once and seamlessly ported " -"to a wide range of AI hardware with minimal effort." -msgstr "FlagOS 是一个完全开源的异构 AI 芯片系统软件栈,允许 AI 模型一次开发即可无缝移植到广泛的 AI 硬件平台,实现最小化的适配成本。" - -#: ../../overview.md:6 ../../overview.md:10 -msgid "FlagOS architecture" -msgstr "FlagOS 架构" - -#: ../../overview.md:8 -msgid "" -"The figure below shows the position of FlagOS in the AI ecosystem and its" -" composition modules." -msgstr "下图展示了 FlagOS 在 AI 生态系统中的位置及其组成模块。" - -#: ../../overview.md:10 -msgid "![FlagOS architecture](images/flagos-architecture-en.png)" -msgstr "![FlagOS 架构](images/flagos-architecture-zh.png)" - -#: ../../overview.md:12 -msgid "" -"FlagOS 2.1 comprises the following core libraries, plugins, domain-" -"specific projects, developer tools, and platform services." -msgstr "FlagOS 2.1 包含以下核心库、插件、领域专用项目、开发者工具和平台服务。" - -#: ../../overview.md:14 -msgid "Open-source core libraries" -msgstr "开源核心库" - -#: ../../overview.md:16 -msgid "Operator libraries" -msgstr "算子库" - -#: ../../overview.md:17 -msgid "**General-purpose operator library**" -msgstr "**通用算子库**" - -#: ../../overview.md:18 -msgid "**FlagGems** (v5.3.0)" -msgstr "**FlagGems** (v5.3.0)" - -#: ../../overview.md:20 -msgid "" -"FlagGems is a high-performance general-purpose operator library " -"implemented with the Triton programming language and its extended " -"languages. FlagGems is designed to provide a suite of general-purpose " -"operators for large models, accelerating the inference and training of " -"models across multiple backend platforms." -msgstr "" -"FlagGems 是一个使用 Triton 编程语言及其扩展语言实现的高性能通用算子库。FlagGems " -"旨在为大模型提供一套通用算子,加速多后端平台上的模型推理和训练。" - -#: ../../overview.md:22 -msgid "**Fused operator libraries**" -msgstr "**融合算子库**" - -#: ../../overview.md:23 -msgid "**FlagGems-vllm** (v0.1.0)" -msgstr "**FlagGems-vllm** (v0.1.0)" - -#: ../../overview.md:25 -msgid "" -"A high-performance operator library designed for multiple hardware " -"backends. It provides optimized implementations of common vLLM operators " -"and supports high-performance inference and deployment for a variety of " -"widely used models." -msgstr "一个面向多硬件后端的高性能算子库。它提供常见 vLLM 算子的优化实现,支持多种广泛使用模型的高性能推理和部署。" - -#: ../../overview.md:27 -msgid "**Multi-domain operator libraries**" -msgstr "**多领域算子库**" - -#: ../../overview.md:29 -msgid "**FlagDNN** (v0.2.0)" -msgstr "**FlagDNN** (v0.2.0)" - -#: ../../overview.md:31 -msgid "" -"A deep neural network computing library oriented towards multiple chip " -"backends. It provides high-performance implementations of common deep " -"learning operators." -msgstr "一个面向多芯片后端的深度神经网络计算库。它提供常见深度学习算子的高性能实现。" - -#: ../../overview.md:33 -msgid "**FlagBLAS** (v0.2.0)" -msgstr "**FlagBLAS** (v0.2.0)" - -#: ../../overview.md:35 -msgid "" -"A computing library that follows the BLAS standard interface and is " -"oriented towards multiple chip backends. It defines core operations for " -"numerical calculations." -msgstr "一个遵循 BLAS 标准接口、面向多芯片后端的计算库。它定义数值计算的核心操作。" - -#: ../../overview.md:37 -msgid "**FlagFFT** (v0.1.0)" -msgstr "**FlagFFT** (v0.1.0)" - -#: ../../overview.md:39 -msgid "" -"A JIT-compiled GPU FFT library. It generates CUDA kernels at runtime via " -"Triton/TLE and libtriton_jit, targeting arbitrary-length transforms that " -"cuFFT does not optimally support." -msgstr "" -"一个 JIT 编译的 GPU FFT 库。它通过 Triton/TLE 和 libtriton_jit 在运行时生成 CUDA 内核,针对 " -"cuFFT 无法最优支持的任意长度变换。" - -#: ../../overview.md:41 -msgid "**FlagSparse** (v0.2.0)" -msgstr "**FlagSparse** (v0.2.0)" - -#: ../../overview.md:43 -msgid "" -"A domain-specific operator library that contains operators dedicated to " -"sparse computation scenarios." -msgstr "一个领域专用算子库,包含专门用于稀疏计算场景的算子。" - -#: ../../overview.md:45 -msgid "**FlagTensor** (v0.2.0)" -msgstr "**FlagTensor** (v0.2.0)" - -#: ../../overview.md:47 -msgid "" -"A high-performance tensor-primitive library implemented in Triton " -"language. It provides optimized implementations of common tensor " -"primitives (unary, binary, and tensor contraction operations) benchmarked" -" against cuTensor baselines." -msgstr "" -"一个使用 Triton 语言实现的高性能张量原语库。它提供常见张量原语(一元、二元和张量缩并操作)的优化实现,并以 cuTensor " -"为基准进行测试。" - -#: ../../overview.md:49 -msgid "**FlagAudio** (v0.2.0)" -msgstr "**FlagAudio** (v0.2.0)" - -#: ../../overview.md:51 -msgid "" -"A multi-backend computing library that adheres to Audio standard " -"interfaces. It delivers a high-performance computing solution designed " -"for audio signal processing and speech AI applications." -msgstr "一个遵循 Audio 标准接口的多后端计算库。它为音频信号处理和语音 AI 应用提供高性能计算解决方案。" - -#: ../../overview.md:53 -msgid "**FlagTree** (v0.6.0)" -msgstr "**FlagTree** (v0.6.0)" - -#: ../../overview.md:55 -msgid "" -"FlagTree is an open-source, unified compiler for multiple AI chips. " -"FlagTree is dedicated to building a compiler and associated tooling " -"platform for diverse AI chips, advancing and expanding the upstream and " -"downstream Triton ecosystem, with the goals of supporting existing " -"adaptation solutions, unifying code repositories, and enabling rapid " -"multi-backend support from a single repository. For upstream model users," -" FlagTree provides unified compilation support across multiple backends; " -"for downstream chip vendors, FlagTree offers reference implementations " -"for integration into the Triton ecosystem." -msgstr "" -"FlagTree 是一个开源的多 AI 芯片统一编译器。FlagTree 致力于为多样化的 AI 芯片构建编译器及相关工具平台,推进和扩展 " -"Triton 上下游生态系统,目标是支持现有适配方案、统一代码仓库,并实现从单一仓库快速支持多后端。对于上游模型用户,FlagTree " -"提供跨多后端的统一编译支持;对于下游芯片厂商,FlagTree 提供集成到 Triton 生态的参考实现。" - -#: ../../overview.md:57 -msgid "**FlagScale** (v2.0.0)" -msgstr "**FlagScale** (v2.0.0)" - -#: ../../overview.md:59 -msgid "" -"FlagScale is a comprehensive toolkit designed to support the entire " -"lifecycle of large models. FlagScale builds on the strengths of several " -"prominent open-source projects, including Megatron-LM and vLLM, to " -"provide a robust, end-to-end solution for managing and scaling large " -"models." -msgstr "" -"FlagScale 是一个全面的大模型全生命周期工具集。FlagScale 基于 Megatron-LM 和 vLLM " -"等多个知名开源项目的优势,为管理和扩展大模型提供稳健的端到端解决方案。" - -#: ../../overview.md:61 -msgid "**FlagCX** (v0.13.0)" -msgstr "**FlagCX** (v0.13.0)" - -#: ../../overview.md:63 -msgid "" -"FlagCX is a scalable and adaptive unified communication library for " -"cross-chip environments. FlagCX delivers high-performance point-to-point " -"and collective communication capabilities tailored for multi-chip, multi-" -"platform scenarios. By leveraging the native collective communication " -"capabilities of each platform, FlagCX incorporates technologies such as " -"device-buffer IPC and RDMA to enable highly efficient collective " -"communication in both cross-chip and single-chip scenarios, while also " -"providing adaptive tuning capabilities for communication optimization." -msgstr "" -"FlagCX 是一个可扩展、自适应的跨芯片环境统一通信库。FlagCX " -"为多芯片、多平台场景提供高性能的点对点和集合通信能力。通过利用每个平台的原生集合通信能力,FlagCX 采用设备缓冲区 IPC 和 RDMA " -"等技术,在跨芯片和单芯片场景中实现高效的集合通信,同时提供通信优化的自适应调优能力。" - -#: ../../overview.md:65 -msgid "Plugins" -msgstr "插件" - -#: ../../overview.md:67 -msgid "" -"The FlagOS ecosystem enablement layer adopts a plugin architecture " -"composed of the following modules. Each module bridges an upstream " -"library and its backend engine with the FlagOS core libraries." -msgstr "FlagOS 生态使能层采用插件架构,由以下模块组成。每个模块将上游库及其后端引擎与 FlagOS 核心库连接起来。" - -#: ../../overview.md:69 -msgid "**vllm-plugin-FL** (v0.2.0)" -msgstr "**vllm-plugin-FL** (v0.2.0)" - -#: ../../overview.md:71 -msgid "" -"vllm-plugin-FL extends the inference capabilities of vLLM to diverse AI " -"chips, enabling efficient model serving beyond the original supported " -"hardware. Built on FlagOS's unified multi-chip backend." -msgstr "" -"vllm-plugin-FL 将 vLLM 的推理能力扩展到多种 AI 芯片,实现超越原始支持硬件的高效模型服务。基于 FlagOS " -"的统一多芯片后端构建。" - -#: ../../overview.md:73 -msgid "**sglang-plugin-FL** (v0.1.0)" -msgstr "**sglang-plugin-FL** (v0.1.0)" - -#: ../../overview.md:75 -msgid "" -"sglang-plugin-FL is an out-of-tree (OOT) plugin for SGLang, built on " -"FlagOS's unified multi-chip backend. It extends SGLang's inference " -"capabilities across diverse hardware platforms." -msgstr "" -"sglang-plugin-FL 是 SGLang 的一个树外 (OOT) 插件,基于 FlagOS 的统一多芯片后端构建。它将 SGLang " -"的推理能力扩展到多种硬件平台。" - -#: ../../overview.md:77 -msgid "**PyTorch-Plugin-FL** (v0.1.0)" -msgstr "**PyTorch-Plugin-FL** (v0.1.0)" - -#: ../../overview.md:79 -msgid "" -"PyTorch-Plugin-FL is a custom PyTorch device plugin based on the " -"PrivateUse1 extension mechanism, registering FlagGems high-performance " -"Triton operators as the flagos device backend for unified multi-chip " -"support." -msgstr "" -"PyTorch-Plugin-FL 是一个基于 PrivateUse1 扩展机制的自定义 PyTorch 设备插件,将 FlagGems 高性能 " -"Triton 算子注册为 flagos 设备后端,实现统一的多芯片支持。" - -#: ../../overview.md:81 -msgid "**Megatron-LM-FL** (v0.2.0)" -msgstr "**Megatron-LM-FL** (v0.2.0)" - -#: ../../overview.md:83 -msgid "" -"Megatron-LM-FL extends the distributed training capabilities of Megatron-" -"LM to diverse AI chips, supporting scalable large-model training across " -"heterogeneous hardware." -msgstr "Megatron-LM-FL 将 Megatron-LM 的分布式训练能力扩展到多种 AI 芯片,支持跨异构硬件的可扩展大模型训练。" - -#: ../../overview.md:85 -msgid "**TransformerEngine-FL** (v0.2.0)" -msgstr "**TransformerEngine-FL** (v0.2.0)" - -#: ../../overview.md:87 -msgid "" -"TransformerEngine-FL extends the transformer acceleration capabilities of" -" Transformer Engine to diverse AI chips, enabling hardware-agnostic " -"training acceleration." -msgstr "" -"TransformerEngine-FL 将 Transformer Engine 的 transformer 加速能力扩展到多种 AI " -"芯片,实现硬件无关的训练加速。" - -#: ../../overview.md:89 -msgid "**verl-FL** (v0.2.0)" -msgstr "**verl-FL** (v0.2.0)" - -#: ../../overview.md:91 -msgid "" -"verl-FL extends the reinforcement learning capabilities of veRL to " -"diverse AI chips, broadening the hardware coverage for RL-based training " -"workflows." -msgstr "verl-FL 将 veRL 的强化学习能力扩展到多种 AI 芯片,拓宽基于强化学习训练工作流的硬件覆盖范围。" - -#: ../../overview.md:93 -msgid "" -"vllm-plugin-FL, Megatron-LM-FL, TransformerEngine-FL, and verl-FL can be " -"used together with FlagScale. When only one or two capabilities are " -"required — such as training, inference, or reinforcement learning — the " -"corresponding module can independently bridge its upstream library and " -"backend engine with the relevant FlagOS core library modules, offering " -"the flexibility to meet diverse user deployment scenarios." -msgstr "" -"vllm-plugin-FL、Megatron-LM-FL、TransformerEngine-FL 和 verl-FL 可以与 " -"FlagScale 配合使用。当只需要一两项能力(如训练、推理或强化学习)时,相应模块可以独立地将其上游库和后端引擎与相关 FlagOS " -"核心库模块连接,提供满足多样化用户部署场景的灵活性。" - -#: ../../overview.md:95 -msgid "Domain-specific projects" -msgstr "领域专用项目" - -#: ../../overview.md:97 -msgid "**FlagOS-Robo** (v0.1.0)" -msgstr "**FlagOS-Robo** (v0.1.0)" - -#: ../../overview.md:99 -msgid "" -"FlagOS-Robo is a chip-agnostic framework for training and deploying " -"Vision Language Models (VLMs) and Vision Language Action (VLA) models " -"across edge-to-cloud scenarios in Embodied Intelligence. It treats VLMs " -"as the \"brain\" for task planning and VLA models as the \"cerebellum\" " -"for generating robot control actions." -msgstr "" -"FlagOS-Robo 是一个芯片无关的框架,用于在具身智能的端到云场景中训练和部署视觉语言模型 (VLM) 和视觉语言动作模型 (VLA)。它将" -" VLM 视为任务规划的\"大脑\",将 VLA 模型视为生成机器人控制动作的\"小脑\"。" - -#: ../../overview.md:101 -msgid "**FlagQuantum** (v0.1.0)" -msgstr "**FlagQuantum** (v0.1.0)" - -#: ../../overview.md:103 -msgid "" -"FlagQuantum is a high-performance distributed quantum statevector " -"simulator built on PyTorch, enabling quantum circuit simulation across " -"multiple GPUs with automatic sharding and resharding." -msgstr "" -"FlagQuantum 是一个基于 PyTorch 构建的高性能分布式量子态矢量模拟器,支持跨多 GPU " -"的量子电路模拟,具有自动分片和重新分片功能。" - -#: ../../overview.md:105 -msgid "Developer tools" -msgstr "开发者工具" - -#: ../../overview.md:107 -msgid "**KernelGen** (v2.1.0)" -msgstr "**KernelGen** (v2.1.0)" - -#: ../../overview.md:109 -msgid "" -"KernelGen is an operator auto-generation tool. KernelGen is designed to " -"construct operator definitions through natural language prompts, retrieve" -" existing similar operator definitions, automatically execute operator " -"accuracy and performance testing, generate accuracy and performance test " -"results, and produce Triton Kernels." -msgstr "" -"KernelGen 是一个算子自动生成工具。KernelGen " -"旨在通过自然语言提示构建算子定义,检索现有相似算子定义,自动执行算子精度和性能测试,生成精度和性能测试结果,并产出 Triton Kernel。" - -#: ../../overview.md:111 -msgid "**FlagOS Skills** (v1.1.0)" -msgstr "**FlagOS Skills** (v1.1.0)" - -#: ../../overview.md:113 -msgid "" -"FlagOS Skills are agent-compatible capabilities designed to streamline " -"key FlagOS workflows, including deployment, operator development, " -"migration, adoption, and performance evaluation. Compatible with Claude " -"Code, Cursor, Codex, and any agent supporting the Agent Skills standard." -msgstr "" -"FlagOS Skills 是与 Agent 兼容的能力集,旨在简化关键 FlagOS 工作流程,包括部署、算子开发、迁移、适配和性能评估。兼容 " -"Claude Code、Cursor、Codex 以及任何支持 Agent Skills 标准的 Agent。" - -#: ../../overview.md:115 -msgid "**Online Laboratory**" -msgstr "**线上实验室**" - -#: ../../overview.md:117 -msgid "" -"An online laboratory providing cloud-based development environments for " -"FlagOS projects." -msgstr "一个为 FlagOS 项目提供云端开发环境的线上实验室。" - -#: ../../overview.md:119 -msgid "Platform services" -msgstr "平台服务" - -#: ../../overview.md:121 -msgid "**FlagRelease** (v0.1.0)" -msgstr "**FlagRelease** (v0.1.0)" - -#: ../../overview.md:123 -msgid "" -"FlagRelease is a platform dedicated to the automatic migration, " -"adaptation and release of large models for multi-architecture AI chips. " -"FlagRelease aims to enable mainstream large models to be migrated, " -"validated, and released on diverse domestic AI hardware with lower cost " -"and higher efficiency through automated, standardized, and intelligent " -"adaptation workflows." -msgstr "" -"FlagRelease 是一个致力于多架构 AI 芯片大模型自动迁移、适配和发布的平台。FlagRelease " -"旨在通过自动化、标准化和智能化的适配工作流程,使主流大模型能够以更低成本、更高效率在多样化国产 AI 硬件上完成迁移、验证和发布。" - -#: ../../overview.md:125 -msgid "**FlagPerf** (v1.2.0)" -msgstr "**FlagPerf** (v1.2.0)" - -#: ../../overview.md:127 -msgid "" -"FlagPerf is an integrated AI hardware evaluation engine. FlagPerf aims to" -" establish an industry practice-oriented indicator system and evaluate " -"the actual performance of AI hardware under combinations of software " -"stacks (model + framework + compiler)." -msgstr "" -"FlagPerf 是一个集成的 AI 硬件评测引擎。FlagPerf 旨在建立业界实践导向的指标体系,评估 AI 硬件在软件栈组合(模型 + 框架" -" + 编译器)下的实际性能。" - -#: ../../overview.md:129 -msgid "**FlagCICD** (v0.1.0)" -msgstr "**FlagCICD** (v0.1.0)" - -#: ../../overview.md:131 -msgid "" -"FlagCICD is a CI/CD toolchain that streamlines large-model development " -"across diverse AI chips, eliminating fragmentation and cutting adaptation" -" costs." -msgstr "FlagCICD 是一个 CI/CD 工具链,简化跨多种 AI 芯片的大模型开发,消除碎片化并降低适配成本。" - -#: ../../overview.md:133 -msgid "**KernelGenBench** (v0.1.0)" -msgstr "**KernelGenBench** (v0.1.0)" - -#: ../../overview.md:135 -msgid "" -"KernelGenBench is a benchmark framework for evaluating LLM and agent-" -"based Triton kernel generation across multiple hardware platforms." -msgstr "KernelGenBench 是一个基准测试框架,用于评估跨多硬件平台的 LLM 和 Agent 驱动的 Triton kernel 生成能力。" +# FlagOS Documentation Chinese Translation +# Copyright (C) 2025-2026, FlagOS Community +# This file is distributed under the same license as the FlagOS +# Documentation package. +# FlagOS Community , 2026. +# +msgid "" +msgstr "" +"Project-Id-Version: FlagOS Documentation \n" +"Report-Msgid-Bugs-To: \n" +"POT-Creation-Date: 2026-09-20 17:05+0800\n" +"PO-Revision-Date: 2026-06-22 13:15+0800\n" +"Last-Translator: FlagOS Community \n" +"Language: zh_CN\n" +"Language-Team: zh_CN \n" +"Plural-Forms: nplurals=1; plural=0;\n" +"MIME-Version: 1.0\n" +"Content-Type: text/plain; charset=utf-8\n" +"Content-Transfer-Encoding: 8bit\n" +"Generated-By: Babel 2.17.0\n" + +#: ../../overview.md:1 +msgid "FlagOS Overview" +msgstr "FlagOS 概览" + +#: ../../overview.md:3 +msgid "" +"FlagOS is a fully open-source AI system software stack for heterogeneous " +"AI chips, allowing AI models to be developed once and seamlessly ported " +"to a wide range of AI hardware with minimal effort." +msgstr "FlagOS 是一个完全开源的异构 AI 芯片系统软件栈,允许 AI 模型一次开发即可无缝移植到广泛的 AI 硬件平台,实现最小化的适配成本。" + +#: ../../overview.md:6 ../../overview.md:10 +msgid "FlagOS architecture" +msgstr "FlagOS 架构" + +#: ../../overview.md:8 +msgid "" +"The figure below shows the position of FlagOS in the AI ecosystem and its" +" composition modules." +msgstr "下图展示了 FlagOS 在 AI 生态系统中的位置及其组成模块。" + +#: ../../overview.md:10 +msgid "![FlagOS architecture](images/flagos-architecture-en.png)" +msgstr "![FlagOS 架构](images/flagos-architecture-zh.png)" + +#: ../../overview.md:12 +msgid "" +"FlagOS 2.1 comprises the following core libraries, plugins, domain-" +"specific projects, developer tools, and platform services." +msgstr "FlagOS 2.1 包含以下核心库、插件、领域专用项目、开发者工具和平台服务。" + +#: ../../overview.md:14 +msgid "Open-source core libraries" +msgstr "开源核心库" + +#: ../../overview.md:16 +msgid "Operator libraries" +msgstr "算子库" + +#: ../../overview.md:17 +msgid "**General-purpose operator library**" +msgstr "**通用算子库**" + +#: ../../overview.md:18 +msgid "**FlagGems** (v5.3.0)" +msgstr "**FlagGems** (v5.3.0)" + +#: ../../overview.md:20 +msgid "" +"FlagGems is a high-performance general-purpose operator library " +"implemented with the Triton programming language and its extended " +"languages. FlagGems is designed to provide a suite of general-purpose " +"operators for large models, accelerating the inference and training of " +"models across multiple backend platforms." +msgstr "" +"FlagGems 是一个使用 Triton 编程语言及其扩展语言实现的高性能通用算子库。FlagGems " +"旨在为大模型提供一套通用算子,加速多后端平台上的模型推理和训练。" + +#: ../../overview.md:22 +msgid "**Fused operator libraries**" +msgstr "**融合算子库**" + +#: ../../overview.md:23 +msgid "**FlagGems-vllm** (v0.1.0)" +msgstr "**FlagGems-vllm** (v0.1.0)" + +#: ../../overview.md:25 +msgid "" +"A high-performance operator library designed for multiple hardware " +"backends. It provides optimized implementations of common vLLM operators " +"and supports high-performance inference and deployment for a variety of " +"widely used models." +msgstr "一个面向多硬件后端的高性能算子库。它提供常见 vLLM 算子的优化实现,支持多种广泛使用模型的高性能推理和部署。" + +#: ../../overview.md:27 +msgid "**Multi-domain operator libraries**" +msgstr "**多领域算子库**" + +#: ../../overview.md:29 +msgid "**FlagDNN** (v0.2.0)" +msgstr "**FlagDNN** (v0.2.0)" + +#: ../../overview.md:31 +msgid "" +"A deep neural network computing library oriented towards multiple chip " +"backends. It provides high-performance implementations of common deep " +"learning operators." +msgstr "一个面向多芯片后端的深度神经网络计算库。它提供常见深度学习算子的高性能实现。" + +#: ../../overview.md:33 +msgid "**FlagBLAS** (v0.2.0)" +msgstr "**FlagBLAS** (v0.2.0)" + +#: ../../overview.md:35 +msgid "" +"A computing library that follows the BLAS standard interface and is " +"oriented towards multiple chip backends. It defines core operations for " +"numerical calculations." +msgstr "一个遵循 BLAS 标准接口、面向多芯片后端的计算库。它定义数值计算的核心操作。" + +#: ../../overview.md:37 +msgid "**FlagFFT** (v0.1.0)" +msgstr "**FlagFFT** (v0.1.0)" + +#: ../../overview.md:39 +msgid "" +"A JIT-compiled GPU FFT library. It generates CUDA kernels at runtime via " +"Triton/TLE and libtriton_jit, targeting arbitrary-length transforms that " +"cuFFT does not optimally support." +msgstr "" +"一个 JIT 编译的 GPU FFT 库。它通过 Triton/TLE 和 libtriton_jit 在运行时生成 CUDA 内核,针对 " +"cuFFT 无法最优支持的任意长度变换。" + +#: ../../overview.md:41 +msgid "**FlagSparse** (v0.2.0)" +msgstr "**FlagSparse** (v0.2.0)" + +#: ../../overview.md:43 +msgid "" +"A domain-specific operator library that contains operators dedicated to " +"sparse computation scenarios." +msgstr "一个领域专用算子库,包含专门用于稀疏计算场景的算子。" + +#: ../../overview.md:45 +msgid "**FlagTensor** (v0.2.0)" +msgstr "**FlagTensor** (v0.2.0)" + +#: ../../overview.md:47 +msgid "" +"A high-performance tensor-primitive library implemented in Triton " +"language. It provides optimized implementations of common tensor " +"primitives (unary, binary, and tensor contraction operations) benchmarked" +" against cuTensor baselines." +msgstr "" +"一个使用 Triton 语言实现的高性能张量原语库。它提供常见张量原语(一元、二元和张量缩并操作)的优化实现,并以 cuTensor " +"为基准进行测试。" + +#: ../../overview.md:49 +msgid "**FlagAudio** (v0.2.0)" +msgstr "**FlagAudio** (v0.2.0)" + +#: ../../overview.md:51 +msgid "" +"A multi-backend computing library that adheres to Audio standard " +"interfaces. It delivers a high-performance computing solution designed " +"for audio signal processing and speech AI applications." +msgstr "一个遵循 Audio 标准接口的多后端计算库。它为音频信号处理和语音 AI 应用提供高性能计算解决方案。" + +#: ../../overview.md:53 +msgid "**FlagTree** (v0.6.0)" +msgstr "**FlagTree** (v0.6.0)" + +#: ../../overview.md:55 +msgid "" +"FlagTree is an open-source, unified compiler for multiple AI chips. " +"FlagTree is dedicated to building a compiler and associated tooling " +"platform for diverse AI chips, advancing and expanding the upstream and " +"downstream Triton ecosystem, with the goals of supporting existing " +"adaptation solutions, unifying code repositories, and enabling rapid " +"multi-backend support from a single repository. For upstream model users," +" FlagTree provides unified compilation support across multiple backends; " +"for downstream chip vendors, FlagTree offers reference implementations " +"for integration into the Triton ecosystem." +msgstr "" +"FlagTree 是一个开源的多 AI 芯片统一编译器。FlagTree 致力于为多样化的 AI 芯片构建编译器及相关工具平台,推进和扩展 " +"Triton 上下游生态系统,目标是支持现有适配方案、统一代码仓库,并实现从单一仓库快速支持多后端。对于上游模型用户,FlagTree " +"提供跨多后端的统一编译支持;对于下游芯片厂商,FlagTree 提供集成到 Triton 生态的参考实现。" + +#: ../../overview.md:57 +msgid "**FlagScale** (v2.0.0)" +msgstr "**FlagScale** (v2.0.0)" + +#: ../../overview.md:59 +msgid "" +"FlagScale is a comprehensive toolkit designed to support the entire " +"lifecycle of large models. FlagScale builds on the strengths of several " +"prominent open-source projects, including Megatron-LM and vLLM, to " +"provide a robust, end-to-end solution for managing and scaling large " +"models." +msgstr "" +"FlagScale 是一个全面的大模型全生命周期工具集。FlagScale 基于 Megatron-LM 和 vLLM " +"等多个知名开源项目的优势,为管理和扩展大模型提供稳健的端到端解决方案。" + +#: ../../overview.md:61 +msgid "**FlagCX** (v0.13.0)" +msgstr "**FlagCX** (v0.13.0)" + +#: ../../overview.md:63 +msgid "" +"FlagCX is a scalable and adaptive unified communication library for " +"cross-chip environments. FlagCX delivers high-performance point-to-point " +"and collective communication capabilities tailored for multi-chip, multi-" +"platform scenarios. By leveraging the native collective communication " +"capabilities of each platform, FlagCX incorporates technologies such as " +"device-buffer IPC and RDMA to enable highly efficient collective " +"communication in both cross-chip and single-chip scenarios, while also " +"providing adaptive tuning capabilities for communication optimization." +msgstr "" +"FlagCX 是一个可扩展、自适应的跨芯片环境统一通信库。FlagCX " +"为多芯片、多平台场景提供高性能的点对点和集合通信能力。通过利用每个平台的原生集合通信能力,FlagCX 采用设备缓冲区 IPC 和 RDMA " +"等技术,在跨芯片和单芯片场景中实现高效的集合通信,同时提供通信优化的自适应调优能力。" + +#: ../../overview.md:65 +msgid "Plugins" +msgstr "插件" + +#: ../../overview.md:67 +msgid "" +"The FlagOS ecosystem enablement layer adopts a plugin architecture " +"composed of the following modules. Each module bridges an upstream " +"library and its backend engine with the FlagOS core libraries." +msgstr "FlagOS 生态使能层采用插件架构,由以下模块组成。每个模块将上游库及其后端引擎与 FlagOS 核心库连接起来。" + +#: ../../overview.md:69 +msgid "**vllm-plugin-FL** (v0.2.0)" +msgstr "**vllm-plugin-FL** (v0.2.0)" + +#: ../../overview.md:71 +msgid "" +"vllm-plugin-FL extends the inference capabilities of vLLM to diverse AI " +"chips, enabling efficient model serving beyond the original supported " +"hardware. Built on FlagOS's unified multi-chip backend." +msgstr "" +"vllm-plugin-FL 将 vLLM 的推理能力扩展到多种 AI 芯片,实现超越原始支持硬件的高效模型服务。基于 FlagOS " +"的统一多芯片后端构建。" + +#: ../../overview.md:73 +msgid "**sglang-plugin-FL** (v0.1.0)" +msgstr "**sglang-plugin-FL** (v0.1.0)" + +#: ../../overview.md:75 +msgid "" +"sglang-plugin-FL is an out-of-tree (OOT) plugin for SGLang, built on " +"FlagOS's unified multi-chip backend. It extends SGLang's inference " +"capabilities across diverse hardware platforms." +msgstr "" +"sglang-plugin-FL 是 SGLang 的一个树外 (OOT) 插件,基于 FlagOS 的统一多芯片后端构建。它将 SGLang " +"的推理能力扩展到多种硬件平台。" + +#: ../../overview.md:77 +msgid "**Torch-FL** (v0.1.0)" +msgstr "**Torch-FL** (v0.1.0)" + +#: ../../overview.md:79 +msgid "" +"Torch-FL is a custom PyTorch device plugin based on the " +"PrivateUse1 extension mechanism, registering FlagGems high-performance " +"Triton operators as the flagos device backend for unified multi-chip " +"support." +msgstr "" +"Torch-FL 是一个基于 PrivateUse1 扩展机制的自定义 PyTorch 设备插件,将 FlagGems 高性能 " +"Triton 算子注册为 flagos 设备后端,实现统一的多芯片支持。" + +#: ../../overview.md:81 +msgid "**Megatron-LM-FL** (v0.2.0)" +msgstr "**Megatron-LM-FL** (v0.2.0)" + +#: ../../overview.md:83 +msgid "" +"Megatron-LM-FL extends the distributed training capabilities of Megatron-" +"LM to diverse AI chips, supporting scalable large-model training across " +"heterogeneous hardware." +msgstr "Megatron-LM-FL 将 Megatron-LM 的分布式训练能力扩展到多种 AI 芯片,支持跨异构硬件的可扩展大模型训练。" + +#: ../../overview.md:85 +msgid "**TransformerEngine-FL** (v0.2.0)" +msgstr "**TransformerEngine-FL** (v0.2.0)" + +#: ../../overview.md:87 +msgid "" +"TransformerEngine-FL extends the transformer acceleration capabilities of" +" Transformer Engine to diverse AI chips, enabling hardware-agnostic " +"training acceleration." +msgstr "" +"TransformerEngine-FL 将 Transformer Engine 的 transformer 加速能力扩展到多种 AI " +"芯片,实现硬件无关的训练加速。" + +#: ../../overview.md:89 +msgid "**verl-FL** (v0.2.0)" +msgstr "**verl-FL** (v0.2.0)" + +#: ../../overview.md:91 +msgid "" +"verl-FL extends the reinforcement learning capabilities of veRL to " +"diverse AI chips, broadening the hardware coverage for RL-based training " +"workflows." +msgstr "verl-FL 将 veRL 的强化学习能力扩展到多种 AI 芯片,拓宽基于强化学习训练工作流的硬件覆盖范围。" + +#: ../../overview.md:93 +msgid "" +"vllm-plugin-FL, Megatron-LM-FL, TransformerEngine-FL, and verl-FL can be " +"used together with FlagScale. When only one or two capabilities are " +"required — such as training, inference, or reinforcement learning — the " +"corresponding module can independently bridge its upstream library and " +"backend engine with the relevant FlagOS core library modules, offering " +"the flexibility to meet diverse user deployment scenarios." +msgstr "" +"vllm-plugin-FL、Megatron-LM-FL、TransformerEngine-FL 和 verl-FL 可以与 " +"FlagScale 配合使用。当只需要一两项能力(如训练、推理或强化学习)时,相应模块可以独立地将其上游库和后端引擎与相关 FlagOS " +"核心库模块连接,提供满足多样化用户部署场景的灵活性。" + +#: ../../overview.md:95 +msgid "Domain-specific projects" +msgstr "领域专用项目" + +#: ../../overview.md:97 +msgid "**FlagOS-Robo** (v0.1.0)" +msgstr "**FlagOS-Robo** (v0.1.0)" + +#: ../../overview.md:99 +msgid "" +"FlagOS-Robo is a chip-agnostic framework for training and deploying " +"Vision Language Models (VLMs) and Vision Language Action (VLA) models " +"across edge-to-cloud scenarios in Embodied Intelligence. It treats VLMs " +"as the \"brain\" for task planning and VLA models as the \"cerebellum\" " +"for generating robot control actions." +msgstr "" +"FlagOS-Robo 是一个芯片无关的框架,用于在具身智能的端到云场景中训练和部署视觉语言模型 (VLM) 和视觉语言动作模型 (VLA)。它将" +" VLM 视为任务规划的\"大脑\",将 VLA 模型视为生成机器人控制动作的\"小脑\"。" + +#: ../../overview.md:101 +msgid "**FlagQuantum** (v0.1.0)" +msgstr "**FlagQuantum** (v0.1.0)" + +#: ../../overview.md:103 +msgid "" +"FlagQuantum is a high-performance distributed quantum statevector " +"simulator built on PyTorch, enabling quantum circuit simulation across " +"multiple GPUs with automatic sharding and resharding." +msgstr "" +"FlagQuantum 是一个基于 PyTorch 构建的高性能分布式量子态矢量模拟器,支持跨多 GPU " +"的量子电路模拟,具有自动分片和重新分片功能。" + +#: ../../overview.md:105 +msgid "Developer tools" +msgstr "开发者工具" + +#: ../../overview.md:107 +msgid "**KernelGen** (v2.1.0)" +msgstr "**KernelGen** (v2.1.0)" + +#: ../../overview.md:109 +msgid "" +"KernelGen is an operator auto-generation tool. KernelGen is designed to " +"construct operator definitions through natural language prompts, retrieve" +" existing similar operator definitions, automatically execute operator " +"accuracy and performance testing, generate accuracy and performance test " +"results, and produce Triton Kernels." +msgstr "" +"KernelGen 是一个算子自动生成工具。KernelGen " +"旨在通过自然语言提示构建算子定义,检索现有相似算子定义,自动执行算子精度和性能测试,生成精度和性能测试结果,并产出 Triton Kernel。" + +#: ../../overview.md:111 +msgid "**FlagOS Skills** (v1.1.0)" +msgstr "**FlagOS Skills** (v1.1.0)" + +#: ../../overview.md:113 +msgid "" +"FlagOS Skills are agent-compatible capabilities designed to streamline " +"key FlagOS workflows, including deployment, operator development, " +"migration, adoption, and performance evaluation. Compatible with Claude " +"Code, Cursor, Codex, and any agent supporting the Agent Skills standard." +msgstr "" +"FlagOS Skills 是与 Agent 兼容的能力集,旨在简化关键 FlagOS 工作流程,包括部署、算子开发、迁移、适配和性能评估。兼容 " +"Claude Code、Cursor、Codex 以及任何支持 Agent Skills 标准的 Agent。" + +#: ../../overview.md:115 +msgid "**Online Laboratory**" +msgstr "**线上实验室**" + +#: ../../overview.md:117 +msgid "" +"An online laboratory providing cloud-based development environments for " +"FlagOS projects." +msgstr "一个为 FlagOS 项目提供云端开发环境的线上实验室。" + +#: ../../overview.md:119 +msgid "Platform services" +msgstr "平台服务" + +#: ../../overview.md:121 +msgid "**FlagRelease** (v0.1.0)" +msgstr "**FlagRelease** (v0.1.0)" + +#: ../../overview.md:123 +msgid "" +"FlagRelease is a platform dedicated to the automatic migration, " +"adaptation and release of large models for multi-architecture AI chips. " +"FlagRelease aims to enable mainstream large models to be migrated, " +"validated, and released on diverse domestic AI hardware with lower cost " +"and higher efficiency through automated, standardized, and intelligent " +"adaptation workflows." +msgstr "" +"FlagRelease 是一个致力于多架构 AI 芯片大模型自动迁移、适配和发布的平台。FlagRelease " +"旨在通过自动化、标准化和智能化的适配工作流程,使主流大模型能够以更低成本、更高效率在多样化国产 AI 硬件上完成迁移、验证和发布。" + +#: ../../overview.md:125 +msgid "**FlagPerf** (v1.2.0)" +msgstr "**FlagPerf** (v1.2.0)" + +#: ../../overview.md:127 +msgid "" +"FlagPerf is an integrated AI hardware evaluation engine. FlagPerf aims to" +" establish an industry practice-oriented indicator system and evaluate " +"the actual performance of AI hardware under combinations of software " +"stacks (model + framework + compiler)." +msgstr "" +"FlagPerf 是一个集成的 AI 硬件评测引擎。FlagPerf 旨在建立业界实践导向的指标体系,评估 AI 硬件在软件栈组合(模型 + 框架" +" + 编译器)下的实际性能。" + +#: ../../overview.md:129 +msgid "**FlagCICD** (v0.1.0)" +msgstr "**FlagCICD** (v0.1.0)" + +#: ../../overview.md:131 +msgid "" +"FlagCICD is a CI/CD toolchain that streamlines large-model development " +"across diverse AI chips, eliminating fragmentation and cutting adaptation" +" costs." +msgstr "FlagCICD 是一个 CI/CD 工具链,简化跨多种 AI 芯片的大模型开发,消除碎片化并降低适配成本。" + +#: ../../overview.md:133 +msgid "**KernelGenBench** (v0.1.0)" +msgstr "**KernelGenBench** (v0.1.0)" + +#: ../../overview.md:135 +msgid "" +"KernelGenBench is a benchmark framework for evaluating LLM and agent-" +"based Triton kernel generation across multiple hardware platforms." +msgstr "KernelGenBench 是一个基准测试框架,用于评估跨多硬件平台的 LLM 和 Agent 驱动的 Triton kernel 生成能力。" diff --git a/docs/flagos_homepage/overview.md b/docs/flagos_homepage/overview.md index 5677a883a8..52f68e4667 100644 --- a/docs/flagos_homepage/overview.md +++ b/docs/flagos_homepage/overview.md @@ -1,135 +1,135 @@ -# FlagOS Overview - -FlagOS is a fully open-source AI system software stack for heterogeneous AI chips, allowing AI models to be developed once and seamlessly ported to a wide range of AI hardware with minimal effort. - - -## FlagOS architecture - -The figure below shows the position of FlagOS in the AI ecosystem and its composition modules. - -![FlagOS architecture](images/flagos-architecture-en.png) - -FlagOS 2.1 comprises the following core libraries, plugins, domain-specific projects, developer tools, and platform services. - -### Open-source core libraries - -- Operator libraries - - **General-purpose operator library** - - **FlagGems** (v5.3.0) - - FlagGems is a high-performance general-purpose operator library implemented with the Triton programming language and its extended languages. FlagGems is designed to provide a suite of general-purpose operators for large models, accelerating the inference and training of models across multiple backend platforms. - - - **Fused operator libraries** - - **FlagGems-vllm** (v0.1.0) - - A high-performance operator library designed for multiple hardware backends. It provides optimized implementations of common vLLM operators and supports high-performance inference and deployment for a variety of widely used models. - - - **Multi-domain operator libraries** - - - **FlagDNN** (v0.2.0) - - A deep neural network computing library oriented towards multiple chip backends. It provides high-performance implementations of common deep learning operators. - - - **FlagBLAS** (v0.2.0) - - A computing library that follows the BLAS standard interface and is oriented towards multiple chip backends. It defines core operations for numerical calculations. - - - **FlagFFT** (v0.1.0) - - A JIT-compiled GPU FFT library. It generates CUDA kernels at runtime via Triton/TLE and libtriton_jit, targeting arbitrary-length transforms that cuFFT does not optimally support. - - - **FlagSparse** (v0.2.0) - - A domain-specific operator library that contains operators dedicated to sparse computation scenarios. - - - **FlagTensor** (v0.2.0) - - A high-performance tensor-primitive library implemented in Triton language. It provides optimized implementations of common tensor primitives (unary, binary, and tensor contraction operations) benchmarked against cuTensor baselines. - - - **FlagAudio** (v0.2.0) - - A multi-backend computing library that adheres to Audio standard interfaces. It delivers a high-performance computing solution designed for audio signal processing and speech AI applications. - -- **FlagTree** (v0.6.0) - - FlagTree is an open-source, unified compiler for multiple AI chips. FlagTree is dedicated to building a compiler and associated tooling platform for diverse AI chips, advancing and expanding the upstream and downstream Triton ecosystem, with the goals of supporting existing adaptation solutions, unifying code repositories, and enabling rapid multi-backend support from a single repository. For upstream model users, FlagTree provides unified compilation support across multiple backends; for downstream chip vendors, FlagTree offers reference implementations for integration into the Triton ecosystem. - -- **FlagScale** (v2.0.0) - - FlagScale is a comprehensive toolkit designed to support the entire lifecycle of large models. FlagScale builds on the strengths of several prominent open-source projects, including Megatron-LM and vLLM, to provide a robust, end-to-end solution for managing and scaling large models. - -- **FlagCX** (v0.13.0) - - FlagCX is a scalable and adaptive unified communication library for cross-chip environments. FlagCX delivers high-performance point-to-point and collective communication capabilities tailored for multi-chip, multi-platform scenarios. By leveraging the native collective communication capabilities of each platform, FlagCX incorporates technologies such as device-buffer IPC and RDMA to enable highly efficient collective communication in both cross-chip and single-chip scenarios, while also providing adaptive tuning capabilities for communication optimization. - -### Plugins - -The FlagOS ecosystem enablement layer adopts a plugin architecture composed of the following modules. Each module bridges an upstream library and its backend engine with the FlagOS core libraries. - -- **vllm-plugin-FL** (v0.2.0) - - vllm-plugin-FL extends the inference capabilities of vLLM to diverse AI chips, enabling efficient model serving beyond the original supported hardware. Built on FlagOS's unified multi-chip backend. - -- **sglang-plugin-FL** (v0.1.0) - - sglang-plugin-FL is an out-of-tree (OOT) plugin for SGLang, built on FlagOS's unified multi-chip backend. It extends SGLang's inference capabilities across diverse hardware platforms. - -- **PyTorch-Plugin-FL** (v0.1.0) - - PyTorch-Plugin-FL is a custom PyTorch device plugin based on the PrivateUse1 extension mechanism, registering FlagGems high-performance Triton operators as the flagos device backend for unified multi-chip support. - -- **Megatron-LM-FL** (v0.2.0) - - Megatron-LM-FL extends the distributed training capabilities of Megatron-LM to diverse AI chips, supporting scalable large-model training across heterogeneous hardware. - -- **TransformerEngine-FL** (v0.2.0) - - TransformerEngine-FL extends the transformer acceleration capabilities of Transformer Engine to diverse AI chips, enabling hardware-agnostic training acceleration. - -- **verl-FL** (v0.2.0) - - verl-FL extends the reinforcement learning capabilities of veRL to diverse AI chips, broadening the hardware coverage for RL-based training workflows. - -vllm-plugin-FL, Megatron-LM-FL, TransformerEngine-FL, and verl-FL can be used together with FlagScale. When only one or two capabilities are required — such as training, inference, or reinforcement learning — the corresponding module can independently bridge its upstream library and backend engine with the relevant FlagOS core library modules, offering the flexibility to meet diverse user deployment scenarios. - -### Domain-specific projects - -- **FlagOS-Robo** (v0.1.0) - - FlagOS-Robo is a chip-agnostic framework for training and deploying Vision Language Models (VLMs) and Vision Language Action (VLA) models across edge-to-cloud scenarios in Embodied Intelligence. It treats VLMs as the "brain" for task planning and VLA models as the "cerebellum" for generating robot control actions. - -- **FlagQuantum** (v0.1.0) - - FlagQuantum is a high-performance distributed quantum statevector simulator built on PyTorch, enabling quantum circuit simulation across multiple GPUs with automatic sharding and resharding. - -### Developer tools - -- **KernelGen** (v2.1.0) - - KernelGen is an operator auto-generation tool. KernelGen is designed to construct operator definitions through natural language prompts, retrieve existing similar operator definitions, automatically execute operator accuracy and performance testing, generate accuracy and performance test results, and produce Triton Kernels. - -- **FlagOS Skills** (v1.1.0) - - FlagOS Skills are agent-compatible capabilities designed to streamline key FlagOS workflows, including deployment, operator development, migration, adoption, and performance evaluation. Compatible with Claude Code, Cursor, Codex, and any agent supporting the Agent Skills standard. - -- **Online Laboratory** - - An online laboratory providing cloud-based development environments for FlagOS projects. - -### Platform services - -- **FlagRelease** (v0.1.0) - - FlagRelease is a platform dedicated to the automatic migration, adaptation and release of large models for multi-architecture AI chips. FlagRelease aims to enable mainstream large models to be migrated, validated, and released on diverse domestic AI hardware with lower cost and higher efficiency through automated, standardized, and intelligent adaptation workflows. - -- **FlagPerf** (v1.2.0) - - FlagPerf is an integrated AI hardware evaluation engine. FlagPerf aims to establish an industry practice-oriented indicator system and evaluate the actual performance of AI hardware under combinations of software stacks (model + framework + compiler). - -- **FlagCICD** (v0.1.0) - - FlagCICD is a CI/CD toolchain that streamlines large-model development across diverse AI chips, eliminating fragmentation and cutting adaptation costs. - -- **KernelGenBench** (v0.1.0) - - KernelGenBench is a benchmark framework for evaluating LLM and agent-based Triton kernel generation across multiple hardware platforms. +# FlagOS Overview + +FlagOS is a fully open-source AI system software stack for heterogeneous AI chips, allowing AI models to be developed once and seamlessly ported to a wide range of AI hardware with minimal effort. + + +## FlagOS architecture + +The figure below shows the position of FlagOS in the AI ecosystem and its composition modules. + +![FlagOS architecture](images/flagos-architecture-en.png) + +FlagOS 2.1 comprises the following core libraries, plugins, domain-specific projects, developer tools, and platform services. + +### Open-source core libraries + +- Operator libraries + - **General-purpose operator library** + - **FlagGems** (v5.3.0) + + FlagGems is a high-performance general-purpose operator library implemented with the Triton programming language and its extended languages. FlagGems is designed to provide a suite of general-purpose operators for large models, accelerating the inference and training of models across multiple backend platforms. + + - **Fused operator libraries** + - **FlagGems-vllm** (v0.1.0) + + A high-performance operator library designed for multiple hardware backends. It provides optimized implementations of common vLLM operators and supports high-performance inference and deployment for a variety of widely used models. + + - **Multi-domain operator libraries** + + - **FlagDNN** (v0.2.0) + + A deep neural network computing library oriented towards multiple chip backends. It provides high-performance implementations of common deep learning operators. + + - **FlagBLAS** (v0.2.0) + + A computing library that follows the BLAS standard interface and is oriented towards multiple chip backends. It defines core operations for numerical calculations. + + - **FlagFFT** (v0.1.0) + + A JIT-compiled GPU FFT library. It generates CUDA kernels at runtime via Triton/TLE and libtriton_jit, targeting arbitrary-length transforms that cuFFT does not optimally support. + + - **FlagSparse** (v0.2.0) + + A domain-specific operator library that contains operators dedicated to sparse computation scenarios. + + - **FlagTensor** (v0.2.0) + + A high-performance tensor-primitive library implemented in Triton language. It provides optimized implementations of common tensor primitives (unary, binary, and tensor contraction operations) benchmarked against cuTensor baselines. + + - **FlagAudio** (v0.2.0) + + A multi-backend computing library that adheres to Audio standard interfaces. It delivers a high-performance computing solution designed for audio signal processing and speech AI applications. + +- **FlagTree** (v0.6.0) + + FlagTree is an open-source, unified compiler for multiple AI chips. FlagTree is dedicated to building a compiler and associated tooling platform for diverse AI chips, advancing and expanding the upstream and downstream Triton ecosystem, with the goals of supporting existing adaptation solutions, unifying code repositories, and enabling rapid multi-backend support from a single repository. For upstream model users, FlagTree provides unified compilation support across multiple backends; for downstream chip vendors, FlagTree offers reference implementations for integration into the Triton ecosystem. + +- **FlagScale** (v2.0.0) + + FlagScale is a comprehensive toolkit designed to support the entire lifecycle of large models. FlagScale builds on the strengths of several prominent open-source projects, including Megatron-LM and vLLM, to provide a robust, end-to-end solution for managing and scaling large models. + +- **FlagCX** (v0.13.0) + + FlagCX is a scalable and adaptive unified communication library for cross-chip environments. FlagCX delivers high-performance point-to-point and collective communication capabilities tailored for multi-chip, multi-platform scenarios. By leveraging the native collective communication capabilities of each platform, FlagCX incorporates technologies such as device-buffer IPC and RDMA to enable highly efficient collective communication in both cross-chip and single-chip scenarios, while also providing adaptive tuning capabilities for communication optimization. + +### Plugins + +The FlagOS ecosystem enablement layer adopts a plugin architecture composed of the following modules. Each module bridges an upstream library and its backend engine with the FlagOS core libraries. + +- **vllm-plugin-FL** (v0.2.0) + + vllm-plugin-FL extends the inference capabilities of vLLM to diverse AI chips, enabling efficient model serving beyond the original supported hardware. Built on FlagOS's unified multi-chip backend. + +- **sglang-plugin-FL** (v0.1.0) + + sglang-plugin-FL is an out-of-tree (OOT) plugin for SGLang, built on FlagOS's unified multi-chip backend. It extends SGLang's inference capabilities across diverse hardware platforms. + +- **Torch-FL** (v0.1.0) + + Torch-FL is a custom PyTorch device plugin based on the PrivateUse1 extension mechanism, registering FlagGems high-performance Triton operators as the flagos device backend for unified multi-chip support. + +- **Megatron-LM-FL** (v0.2.0) + + Megatron-LM-FL extends the distributed training capabilities of Megatron-LM to diverse AI chips, supporting scalable large-model training across heterogeneous hardware. + +- **TransformerEngine-FL** (v0.2.0) + + TransformerEngine-FL extends the transformer acceleration capabilities of Transformer Engine to diverse AI chips, enabling hardware-agnostic training acceleration. + +- **verl-FL** (v0.2.0) + + verl-FL extends the reinforcement learning capabilities of veRL to diverse AI chips, broadening the hardware coverage for RL-based training workflows. + +vllm-plugin-FL, Megatron-LM-FL, TransformerEngine-FL, and verl-FL can be used together with FlagScale. When only one or two capabilities are required — such as training, inference, or reinforcement learning — the corresponding module can independently bridge its upstream library and backend engine with the relevant FlagOS core library modules, offering the flexibility to meet diverse user deployment scenarios. + +### Domain-specific projects + +- **FlagOS-Robo** (v0.1.0) + + FlagOS-Robo is a chip-agnostic framework for training and deploying Vision Language Models (VLMs) and Vision Language Action (VLA) models across edge-to-cloud scenarios in Embodied Intelligence. It treats VLMs as the "brain" for task planning and VLA models as the "cerebellum" for generating robot control actions. + +- **FlagQuantum** (v0.1.0) + + FlagQuantum is a high-performance distributed quantum statevector simulator built on PyTorch, enabling quantum circuit simulation across multiple GPUs with automatic sharding and resharding. + +### Developer tools + +- **KernelGen** (v2.1.0) + + KernelGen is an operator auto-generation tool. KernelGen is designed to construct operator definitions through natural language prompts, retrieve existing similar operator definitions, automatically execute operator accuracy and performance testing, generate accuracy and performance test results, and produce Triton Kernels. + +- **FlagOS Skills** (v1.1.0) + + FlagOS Skills are agent-compatible capabilities designed to streamline key FlagOS workflows, including deployment, operator development, migration, adoption, and performance evaluation. Compatible with Claude Code, Cursor, Codex, and any agent supporting the Agent Skills standard. + +- **Online Laboratory** + + An online laboratory providing cloud-based development environments for FlagOS projects. + +### Platform services + +- **FlagRelease** (v0.1.0) + + FlagRelease is a platform dedicated to the automatic migration, adaptation and release of large models for multi-architecture AI chips. FlagRelease aims to enable mainstream large models to be migrated, validated, and released on diverse domestic AI hardware with lower cost and higher efficiency through automated, standardized, and intelligent adaptation workflows. + +- **FlagPerf** (v1.2.0) + + FlagPerf is an integrated AI hardware evaluation engine. FlagPerf aims to establish an industry practice-oriented indicator system and evaluate the actual performance of AI hardware under combinations of software stacks (model + framework + compiler). + +- **FlagCICD** (v0.1.0) + + FlagCICD is a CI/CD toolchain that streamlines large-model development across diverse AI chips, eliminating fragmentation and cutting adaptation costs. + +- **KernelGenBench** (v0.1.0) + + KernelGenBench is a benchmark framework for evaluating LLM and agent-based Triton kernel generation across multiple hardware platforms. diff --git a/docs/flagtree_en/getting_started/install.md b/docs/flagtree_en/getting_started/install.md index 9a050696ae..1574d9044a 100644 --- a/docs/flagtree_en/getting_started/install.md +++ b/docs/flagtree_en/getting_started/install.md @@ -109,12 +109,12 @@ If you do not wish to build from source, you can directly pull and install whl ( |Backend |Install command
(The version corresponds to the git tag)|Triton
ver.|libc.so &
libstdc++.so| |:---------|:---------|:---------|:---------| |nvidia |python3.12 -m pip install flagtree===0.7.0 $RES |3.6|GLIBC_2.39
GLIBCXX_3.4.33
CXXABI_1.3.15| - |tileir |python3.12 -m pip install flagtree===0.6.1+tileir3.6 $RES |3.6|GLIBC_2.39
GLIBCXX_3.4.33
CXXABI_1.3.15| - |amd |python3.12 -m pip install flagtree===0.7.0rc1+amd3.6 $RES |3.6|GLIBC_2.39
GLIBCXX_3.4.33
CXXABI_1.3.15| + |tileir |python3.12 -m pip install flagtree===0.7.0+tileir3.6 $RES |3.6|GLIBC_2.39
GLIBCXX_3.4.33
CXXABI_1.3.15| + |amd |python3.12 -m pip install flagtree===0.7.0+amd3.6 $RES |3.6|GLIBC_2.39
GLIBCXX_3.4.33
CXXABI_1.3.15| |aipu |python3.10 -m pip install flagtree===0.5.0+aipu3.3 $RES |3.3|GLIBC_2.35
GLIBCXX_3.4.30
CXXABI_1.3.13| |ascend |python3.11 -m pip install flagtree===0.7.0+ascend3.5 $RES |3.5|GLIBC_2.35
GLIBCXX_3.4.30
CXXABI_1.3.13| |ascend |python3.11 -m pip install flagtree===0.6.0+ascend3.2 $RES |3.2|GLIBC_2.35
GLIBCXX_3.4.30
CXXABI_1.3.13| - |enflame |python3.12 -m pip install flagtree===0.7.0rc2+enflame3.6 $RES |3.6|GLIBC_2.39
GLIBCXX_3.4.33
CXXABI_1.3.15| + |enflame |python3.12 -m pip install flagtree===0.7.0+enflame3.6 $RES |3.6|GLIBC_2.39
GLIBCXX_3.4.33
CXXABI_1.3.15| |enflame |python3.12 -m pip install flagtree===0.5.0+enflame3.5 $RES |3.5|GLIBC_2.39
GLIBCXX_3.4.33
CXXABI_1.3.15| |enflame |python3.10 -m pip install flagtree===0.4.0+enflame3.3 $RES |3.3|GLIBC_2.35
GLIBCXX_3.4.30
CXXABI_1.3.13| |hcu |python3.10 -m pip install flagtree===0.7.0+hcu3.6 $RES |3.6|GLIBC_2.35
GLIBCXX_3.4.30
CXXABI_1.3.13| @@ -122,17 +122,17 @@ If you do not wish to build from source, you can directly pull and install whl ( |iluvatar |python3.12 -m pip install flagtree===0.7.0+iluvatar3.6 $RES |3.6|GLIBC_2.39
GLIBCXX_3.4.33
CXXABI_1.3.15| |iluvatar |python3.12 -m pip install flagtree===0.5.1+iluvatar3.1 $RES |3.1|GLIBC_2.39
GLIBCXX_3.4.33
CXXABI_1.3.15| |iluvatar |python3.10 -m pip install flagtree===0.5.1+iluvatar3.1 $RES |3.1|GLIBC_2.35
GLIBCXX_3.4.30
CXXABI_1.3.13| - |metax |python3.12 -m pip install flagtree===0.7.0rc3+metax3.6 $RES |3.6|GLIBC_2.39
GLIBCXX_3.4.33
CXXABI_1.3.15| + |metax |python3.12 -m pip install flagtree===0.7.0+metax3.6 $RES |3.6|GLIBC_2.35
GLIBCXX_3.4.30
CXXABI_1.3.13| |metax |python3.12 -m pip install flagtree===0.5.1+metax3.0 $RES |3.0|GLIBC_2.35
GLIBCXX_3.4.30
CXXABI_1.3.13| |mthreads |python3.10 -m pip install flagtree===0.7.0+mthreads3.6 $RES |3.6|GLIBC_2.35
GLIBCXX_3.4.30
CXXABI_1.3.13| |mthreads |python3.10 -m pip install flagtree===0.5.1+mthreads3.2 $RES |3.2|GLIBC_2.35
GLIBCXX_3.4.30
CXXABI_1.3.13| |mthreads |python3.10 -m pip install flagtree===0.5.1+mthreads3.1 $RES |3.1|GLIBC_2.35
GLIBCXX_3.4.30
CXXABI_1.3.13| |ppu |python3.12 -m pip install flagtree===0.7.0+ppu3.6 $RES |3.6|GLIBC_2.39
GLIBCXX_3.4.33
CXXABI_1.3.15| - |sunrise |python3.10 -m pip install flagtree===0.6.0+sunrise3.6 $RES |3.6|GLIBC_2.39
GLIBCXX_3.4.33
CXXABI_1.3.15| + |sunrise |python3.10 -m pip install flagtree===0.7.0+sunrise3.6 $RES |3.6|GLIBC_2.39
GLIBCXX_3.4.33
CXXABI_1.3.15| |sunrise |python3.10 -m pip install flagtree===0.4.0+sunrise3.4 $RES |3.4|GLIBC_2.39
GLIBCXX_3.4.33
CXXABI_1.3.15| |tsingmicro|python3.10 -m pip install flagtree===0.7.0+tsingmicro3.6 $RES |3.6|GLIBC_2.30
GLIBCXX_3.4.28
CXXABI_1.3.12| - |tsingmicro|python3.10 -m pip install flagtree===0.6.0+tsingmicro3.3 $RES |3.3|GLIBC_2.30
GLIBCXX_3.4.28
CXXABI_1.3.12| - |xpu |python3.10 -m pip install flagtree===0.7.0rc3+xpu3.6 $RES |3.6|GLIBC_2.35
GLIBCXX_3.4.30
CXXABI_1.3.13| + |tsingmicro|python3.10 -m pip install flagtree===0.7.0+tsingmicro3.3 $RES |3.3|GLIBC_2.30
GLIBCXX_3.4.28
CXXABI_1.3.12| + |xpu |python3.10 -m pip install flagtree===0.7.0+xpu3.6 $RES |3.6|GLIBC_2.35
GLIBCXX_3.4.30
CXXABI_1.3.13| |xpu |python3.10 -m pip install flagtree===0.5.1+xpu3.0 $RES |3.0|GLIBC_2.31
GLIBCXX_3.4.28
CXXABI_1.3.12| Historical versions of flagtree can be found at https://resource.flagos.net/#browse/search/pypi/=repository_name%3Dflagos-pypi-hosted%20AND%20name.raw%3Dflagtree diff --git a/docs/flagtree_zh/getting_started/install.md b/docs/flagtree_zh/getting_started/install.md index 2dd3b87aff..4e705e9e2b 100644 --- a/docs/flagtree_zh/getting_started/install.md +++ b/docs/flagtree_zh/getting_started/install.md @@ -109,12 +109,12 @@ |后端 |安装命令
(版本对应 git 标签)|Triton
版本|libc.so &
libstdc++.so| |:---------|:---------|:---------|:---------| |nvidia |python3.12 -m pip install flagtree===0.7.0 $RES |3.6|GLIBC_2.39
GLIBCXX_3.4.33
CXXABI_1.3.15| - |tileir |python3.12 -m pip install flagtree===0.6.1+tileir3.6 $RES |3.6|GLIBC_2.39
GLIBCXX_3.4.33
CXXABI_1.3.15| - |amd |python3.12 -m pip install flagtree===0.7.0rc1+amd3.6 $RES |3.6|GLIBC_2.39
GLIBCXX_3.4.33
CXXABI_1.3.15| + |tileir |python3.12 -m pip install flagtree===0.7.0+tileir3.6 $RES |3.6|GLIBC_2.39
GLIBCXX_3.4.33
CXXABI_1.3.15| + |amd |python3.12 -m pip install flagtree===0.7.0+amd3.6 $RES |3.6|GLIBC_2.39
GLIBCXX_3.4.33
CXXABI_1.3.15| |aipu |python3.10 -m pip install flagtree===0.5.0+aipu3.3 $RES |3.3|GLIBC_2.35
GLIBCXX_3.4.30
CXXABI_1.3.13| |ascend |python3.11 -m pip install flagtree===0.7.0+ascend3.5 $RES |3.5|GLIBC_2.35
GLIBCXX_3.4.30
CXXABI_1.3.13| |ascend |python3.11 -m pip install flagtree===0.6.0+ascend3.2 $RES |3.2|GLIBC_2.35
GLIBCXX_3.4.30
CXXABI_1.3.13| - |enflame |python3.12 -m pip install flagtree===0.7.0rc2+enflame3.6 $RES |3.6|GLIBC_2.39
GLIBCXX_3.4.33
CXXABI_1.3.15| + |enflame |python3.12 -m pip install flagtree===0.7.0+enflame3.6 $RES |3.6|GLIBC_2.39
GLIBCXX_3.4.33
CXXABI_1.3.15| |enflame |python3.12 -m pip install flagtree===0.5.0+enflame3.5 $RES |3.5|GLIBC_2.39
GLIBCXX_3.4.33
CXXABI_1.3.15| |enflame |python3.10 -m pip install flagtree===0.4.0+enflame3.3 $RES |3.3|GLIBC_2.35
GLIBCXX_3.4.30
CXXABI_1.3.13| |hcu |python3.10 -m pip install flagtree===0.7.0+hcu3.6 $RES |3.6|GLIBC_2.35
GLIBCXX_3.4.30
CXXABI_1.3.13| @@ -122,17 +122,17 @@ |iluvatar |python3.12 -m pip install flagtree===0.7.0+iluvatar3.6 $RES |3.6|GLIBC_2.39
GLIBCXX_3.4.33
CXXABI_1.3.15| |iluvatar |python3.12 -m pip install flagtree===0.5.1+iluvatar3.1 $RES |3.1|GLIBC_2.39
GLIBCXX_3.4.33
CXXABI_1.3.15| |iluvatar |python3.10 -m pip install flagtree===0.5.1+iluvatar3.1 $RES |3.1|GLIBC_2.35
GLIBCXX_3.4.30
CXXABI_1.3.13| - |metax |python3.12 -m pip install flagtree===0.7.0rc3+metax3.6 $RES |3.6|GLIBC_2.39
GLIBCXX_3.4.33
CXXABI_1.3.15| + |metax |python3.12 -m pip install flagtree===0.7.0+metax3.6 $RES |3.6|GLIBC_2.35
GLIBCXX_3.4.30
CXXABI_1.3.13| |metax |python3.12 -m pip install flagtree===0.5.1+metax3.0 $RES |3.0|GLIBC_2.35
GLIBCXX_3.4.30
CXXABI_1.3.13| |mthreads |python3.10 -m pip install flagtree===0.7.0+mthreads3.6 $RES |3.6|GLIBC_2.35
GLIBCXX_3.4.30
CXXABI_1.3.13| |mthreads |python3.10 -m pip install flagtree===0.5.1+mthreads3.2 $RES |3.2|GLIBC_2.35
GLIBCXX_3.4.30
CXXABI_1.3.13| |mthreads |python3.10 -m pip install flagtree===0.5.1+mthreads3.1 $RES |3.1|GLIBC_2.35
GLIBCXX_3.4.30
CXXABI_1.3.13| |ppu |python3.12 -m pip install flagtree===0.7.0+ppu3.6 $RES |3.6|GLIBC_2.39
GLIBCXX_3.4.33
CXXABI_1.3.15| - |sunrise |python3.10 -m pip install flagtree===0.6.0+sunrise3.6 $RES |3.6|GLIBC_2.39
GLIBCXX_3.4.33
CXXABI_1.3.15| + |sunrise |python3.10 -m pip install flagtree===0.7.0+sunrise3.6 $RES |3.6|GLIBC_2.39
GLIBCXX_3.4.33
CXXABI_1.3.15| |sunrise |python3.10 -m pip install flagtree===0.4.0+sunrise3.4 $RES |3.4|GLIBC_2.39
GLIBCXX_3.4.33
CXXABI_1.3.15| |tsingmicro|python3.10 -m pip install flagtree===0.7.0+tsingmicro3.6 $RES |3.6|GLIBC_2.30
GLIBCXX_3.4.28
CXXABI_1.3.12| - |tsingmicro|python3.10 -m pip install flagtree===0.6.0+tsingmicro3.3 $RES |3.3|GLIBC_2.30
GLIBCXX_3.4.28
CXXABI_1.3.12| - |xpu |python3.10 -m pip install flagtree===0.7.0rc3+xpu3.6 $RES |3.6|GLIBC_2.35
GLIBCXX_3.4.30
CXXABI_1.3.13| + |tsingmicro|python3.10 -m pip install flagtree===0.7.0+tsingmicro3.3 $RES |3.3|GLIBC_2.30
GLIBCXX_3.4.28
CXXABI_1.3.12| + |xpu |python3.10 -m pip install flagtree===0.7.0+xpu3.6 $RES |3.6|GLIBC_2.35
GLIBCXX_3.4.30
CXXABI_1.3.13| |xpu |python3.10 -m pip install flagtree===0.5.1+xpu3.0 $RES |3.0|GLIBC_2.31
GLIBCXX_3.4.28
CXXABI_1.3.12| FlagTree 的历史版本可在 https://resource.flagos.net/#browse/search/pypi/=repository_name%3Dflagos-pypi-hosted%20AND%20name.raw%3Dflagtree 找到。 diff --git a/docs/flagtree_zh/release_notes/release_notes_v070.md b/docs/flagtree_zh/release_notes/release_notes_v070.md index 82b0abc291..bfc9d2429c 100644 --- a/docs/flagtree_zh/release_notes/release_notes_v070.md +++ b/docs/flagtree_zh/release_notes/release_notes_v070.md @@ -15,7 +15,7 @@ - 新增 `tle.gpu.buffered_tensor.slot` 和 `tle.gpu.buffered_tensor.reshape` 操作。在 NVIDIA 上支持。 - 为 `tle.gpu.alloc` 新增 `init_value` 和 `alias_offset_bytes` 参数。在 NVIDIA 上支持。 - TLE-Raw: - - 新增 `library` 与 `compiler` 参数,支持通过 `@dialect(..., library="nvshmem", compiler="clang")` 将 NVSHMEM 设备端接口内联进 TLE-Raw kernel。在 NVIDIA 上支持。 + - 新增 `library` 与 `compiler` 参数,支持通过 `@dialect(..., library="nvshmem")` 将 NVSHMEM 设备端接口内联进 TLE-Raw kernel。在 NVIDIA 上支持。 - 后端: - 新增以下后端集成(基于 Triton 3.6):[tileir](/getting_started/multi-backend-prebuilt-docker-image-install/install-tileir.md)(NVIDIA TileIR)、[ppu](/getting_started/multi-backend-prebuilt-docker-image-install/install-ppu.md)(平头哥)和 [spacemit](/getting_started/multi-backend-prebuilt-docker-image-install/install-spacemit.md)(进迭时空)。 diff --git a/docs/torch_fl_en/architecture/distributed.md b/docs/torch_fl_en/architecture/distributed.md index a0ad1bfa1e..9f59970738 100644 --- a/docs/torch_fl_en/architecture/distributed.md +++ b/docs/torch_fl_en/architecture/distributed.md @@ -1,70 +1,70 @@ -# Distributed Collectives - -Torch-FL provides distributed support for the `flagos` device through `ProcessGroupFlagOS`, a native `torch.distributed.ProcessGroup` subclass. It is registered at import, so `torch.distributed.init_process_group("flagos")` works with no `torch.distributed.*` monkeypatching. - -## How it works - -`flagos` tensors and the vendor's tensors share the same physical device memory, so a collective only needs a metadata conversion, not a copy: - -1. A collective virtual method is called with `privateuseone` tensors. -2. The tensors are converted into the device view the inner backend expects (a zero-copy view over the same `data_ptr`), where the inner backend requires one. -3. The call is delegated to the wrapped inner backend. -4. The inner backend's `Work` object is returned unchanged, so callers and the DDP reducer receive properly typed futures. - -`ProcessGroupFlagOS` overrides every collective virtual function — allreduce, allgather (list and into-tensor forms), reduce-scatter, all-to-all (plain and single), broadcast, reduce, gather, scatter, send/recv and their immediate variants, and barrier — rather than a handful of APIs, so a collective cannot silently miss the conversion. - -## Backend selection - -The inner communication backend is resolved at group-construction time, in priority order: - -1. **FlagCX** — the heterogeneous collective library, used when it is importable. FlagCX registers its own backend (`flagcx`) for its own device; `ProcessGroupFlagOS` builds its `ProcessGroupFlagCX` through the `extended_api=True` creator form. -2. **Vendor native** — `NCCL` for NVIDIA and MetaX, `HCCL` for Ascend, `MCCL` for Moore Threads MUSA. -3. **Host-staged gloo** — the last tier and the only one that needs no vendor library; it copies every operand device → host → device per collective. Set `FLAGOS_DIST_STAGED_GLOO=0` to decline this tier and fail loudly instead. One warning is emitted the first time a group is built on it. - -Requests that never name the `flagos` backend are handled as well: `torch.distributed` routes any device type it does not recognize to gloo, and a `ProcessGroupGloo` rejects flagos tensors outright. With `FLAGOS_DIST_REDIRECT_GLOO` (on by default) a plain `init_process_group(backend="gloo")` or `new_group` request is answered with the `flagos` backend when the process accelerator is the flagos device. - -## Usage - -```python -import torch -import torch_fl -import torch_fl.distributed as flagos_dist - -# "auto" (default): FlagCX first, vendor-native backend as fallback -# "flagcx": force flagos / FlagCX, falling back to the vendor backend -# "nccl": force NCCL (NVIDIA, MetaX) -# "hccl": force HCCL (Ascend) -flagos_dist.init_process_group(backend="auto") - -model = MyModel().to("flagos:0") -model = flagos_dist.DistributedDataParallel(model) -flagos_dist.move_buffers_to_device(model, "flagos:0") -``` - -`torch_fl.distributed` exposes `init_process_group`, `DistributedDataParallel` and `move_buffers_to_device`. With the native backend registered you can equally call `torch.distributed.init_process_group("flagos")` directly, or let it be selected automatically from `device_id=torch.device("privateuseone:0")`. - -### DDP - -At import, Torch-FL patches `torch.nn.parallel.DistributedDataParallel.__init__`. When the model lives on a `flagos` device, the patch: - -- forces the Python reducer, bypassing the C++ reducer's CUDA assertion, and -- replaces the default gradient-accumulation hook (which uses functional collectives that have no `privateuseone` dispatch) with a version that goes through `dist.all_reduce` and therefore through `ProcessGroupFlagOS`. - -`torch.nn.DataParallel` and the functional `torch.nn.parallel.data_parallel` are patched the same way, so their replicas are placed on flagos devices instead of failing on device-type detection. - -## Vendor status - -| Vendor | FlagCX path | Native fallback | View conversion | Notes | -|---|---|---|---|---| -| NVIDIA | Yes | NCCL | flagos → cuda view | Collectives and DDP gradient sync live-verified on 2x/8x A100 | -| MetaX | Reused | NCCL-shaped MCCL via MACA's libtorch | flagos → cuda view | Not covered by CI | -| Ascend | Recommended primary path | HCCL (custom backend type) | flagos → npu view | No CUDA compatibility layer exists on Ascend. Architectural routing only; no collective-level CI coverage | -| Hygon DCU | Reused | RCCL via DTK | flagos → cuda view | `all_reduce`/DDP measured on 2 cards; not in CI | -| Moore Threads MUSA | Reused | MCCL | flagos → cuda view | Host-staged gloo measured on MTT S5000 | -| Enflame GCU | Primary path | None (FlagCX only) | None needed | Measured on two S60 devices: collectives, barrier, DDP forward/backward and gradient sync, FSDP2 `fully_shard` training and sharded state-dict save/load | - -## Limitations - -- Collective coverage is validated per vendor; gaps are recorded rather than implied. On Enflame GCU, point-to-point operations, `gather`/`scatter` roots, all-to-all, multi-node rendezvous, process-failure recovery and deployments larger than two devices remain unvalidated. -- Ascend distributed support is architectural: the routing and view logic exist, but there is no collective-level CI coverage. -- The host-staged gloo tier is correctness-first and pays a device-to-host-to-device copy per collective. +# Distributed Collectives + +Torch-FL provides distributed support for the `flagos` device through `ProcessGroupFlagOS`, a native `torch.distributed.ProcessGroup` subclass. It is registered at import, so `torch.distributed.init_process_group("flagos")` works with no `torch.distributed.*` monkeypatching. + +## How it works + +`flagos` tensors and the vendor's tensors share the same physical device memory, so a collective only needs a metadata conversion, not a copy: + +1. A collective virtual method is called with `privateuseone` tensors. +2. The tensors are converted into the device view the inner backend expects (a zero-copy view over the same `data_ptr`), where the inner backend requires one. +3. The call is delegated to the wrapped inner backend. +4. The inner backend's `Work` object is returned unchanged, so callers and the DDP reducer receive properly typed futures. + +`ProcessGroupFlagOS` overrides every collective virtual function — allreduce, allgather (list and into-tensor forms), reduce-scatter, all-to-all (plain and single), broadcast, reduce, gather, scatter, send/recv and their immediate variants, and barrier — rather than a handful of APIs, so a collective cannot silently miss the conversion. + +## Backend selection + +The inner communication backend is resolved at group-construction time, in priority order: + +1. **FlagCX** — the heterogeneous collective library, used when it is importable. FlagCX registers its own backend (`flagcx`) for its own device; `ProcessGroupFlagOS` builds its `ProcessGroupFlagCX` through the `extended_api=True` creator form. +2. **Vendor native** — `NCCL` for NVIDIA and MetaX, `HCCL` for Ascend, `MCCL` for Moore Threads MUSA. +3. **Host-staged gloo** — the last tier and the only one that needs no vendor library; it copies every operand device → host → device per collective. Set `FLAGOS_DIST_STAGED_GLOO=0` to decline this tier and fail loudly instead. One warning is emitted the first time a group is built on it. + +Requests that never name the `flagos` backend are handled as well: `torch.distributed` routes any device type it does not recognize to gloo, and a `ProcessGroupGloo` rejects flagos tensors outright. With `FLAGOS_DIST_REDIRECT_GLOO` (on by default) a plain `init_process_group(backend="gloo")` or `new_group` request is answered with the `flagos` backend when the process accelerator is the flagos device. + +## Usage + +```python +import torch +import torch_fl +import torch_fl.distributed as flagos_dist + +# "auto" (default): FlagCX first, vendor-native backend as fallback +# "flagcx": force flagos / FlagCX, falling back to the vendor backend +# "nccl": force NCCL (NVIDIA, MetaX) +# "hccl": force HCCL (Ascend) +flagos_dist.init_process_group(backend="auto") + +model = MyModel().to("flagos:0") +model = flagos_dist.DistributedDataParallel(model) +flagos_dist.move_buffers_to_device(model, "flagos:0") +``` + +`torch_fl.distributed` exposes `init_process_group`, `DistributedDataParallel` and `move_buffers_to_device`. With the native backend registered you can equally call `torch.distributed.init_process_group("flagos")` directly, or let it be selected automatically from `device_id=torch.device("privateuseone:0")`. + +### DDP + +At import, Torch-FL patches `torch.nn.parallel.DistributedDataParallel.__init__`. When the model lives on a `flagos` device, the patch: + +- forces the Python reducer, bypassing the C++ reducer's CUDA assertion, and +- replaces the default gradient-accumulation hook (which uses functional collectives that have no `privateuseone` dispatch) with a version that goes through `dist.all_reduce` and therefore through `ProcessGroupFlagOS`. + +`torch.nn.DataParallel` and the functional `torch.nn.parallel.data_parallel` are patched the same way, so their replicas are placed on flagos devices instead of failing on device-type detection. + +## Vendor status + +| Vendor | FlagCX path | Native fallback | View conversion | Notes | +|---|---|---|---|---| +| NVIDIA | Yes | NCCL | flagos → cuda view | Collectives and DDP gradient sync verified on multi-node NVIDIA hardware | +| MetaX | Reused | NCCL-shaped MCCL via MACA's libtorch | flagos → cuda view | Not continuously validated | +| Ascend | Recommended primary path | HCCL (custom backend type) | flagos → npu view | No CUDA compatibility layer exists on Ascend. Architectural routing only; no collective-level validation | +| Hygon DCU | Reused | RCCL via DTK | flagos → cuda view | `all_reduce`/DDP measured on 2 cards; not continuously validated | +| Moore Threads MUSA | Reused | MCCL | flagos → cuda view | FlagCX-first routing and MCCL fallback are implemented; end-to-end multi-process collectives are not yet validated on this host | +| Enflame GCU | Primary path | None (FlagCX only) | None needed | Measured on two S60 devices: collectives, barrier, DDP forward/backward and gradient sync, FSDP2 `fully_shard` training and sharded state-dict save/load | + +## Limitations + +- Collective coverage is validated per vendor; gaps are recorded rather than implied. On Enflame GCU, point-to-point operations, `gather`/`scatter` roots, all-to-all, multi-node rendezvous, process-failure recovery and deployments larger than two devices remain unvalidated. +- Ascend distributed support is architectural: the routing and view logic exist, but there is no collective-level validation. +- The host-staged gloo tier is correctness-first and pays a device-to-host-to-device copy per collective. diff --git a/docs/torch_fl_en/architecture/profiler.md b/docs/torch_fl_en/architecture/profiler.md index 0b31361dbe..8bf830f2d9 100644 --- a/docs/torch_fl_en/architecture/profiler.md +++ b/docs/torch_fl_en/architecture/profiler.md @@ -1,73 +1,73 @@ -# Profiler Integration - -`torch.profiler` supports the `flagos` device through a device tracer compiled into the wheel. A trace from `torch.profiler.profile(activities=[CPU, PrivateUse1])` is structurally equivalent to the trace of the same workload on `torch.cuda`: - -- **Flow arrows** connect each CPU operator to the device kernel it launched. -- **Device time attribution**: `prof.key_averages()` reports a per-operator `self_device_time_total`. -- **Complete kernel metadata** (grid/block, occupancy, shared memory, register count) and demangled kernel names. -- **Runtime events** carry real API names decoded from the callback id rather than a placeholder. -- **memcpy and memset** activities are collected alongside kernels. - -## Three-layer architecture - -Adding a vendor means writing one file: a tracer that satisfies the vendor-agnostic interface. - -| Layer | Files | Responsibility | -|---|---|---| -| Vendor-agnostic interface | `csrc/profiler/device_tracer.h` | `DeviceTracer`, `DeviceEvent`, `EventKind` — the complete contract a vendor implements | -| Vendor tracers | `cupti_device_tracer.cc`, `cann_device_tracer.cc`, `musa_mupti_device_tracer.cc`, `roctracer_device_tracer.cc`, `gcu_topspti_device_tracer.cc`, `unavailable_device_tracer.cc` | One implementation per accelerator, plus an explicit no-device-activity fallback | -| Generic adaptor | `flagos_kineto_profiler.{h,cc}` | Kineto/PyTorch profiler adaptor with zero vendor coupling; `dlopen` shims (`cupti_shim.h`, `mupti_shim.h`, `topspti_shim.h`, `roctracer`) bind the vendor activity library | - -| Accelerator | Activity API | Status | -|---|---|---| -| NVIDIA CUDA | CUPTI | Stable, parity suite in CI | -| MetaX | MCPTI (CUDA-compatible activity API in MACA) | Experimental: all seven parity assertions passed on C550 + MACA 3.8.0 (local validation, no vendor runner in CI) | -| Ascend | MSPTI | Beta: kernel/runtime/flow/memcpy events plus device-time linkage, CI-covered by the shared contract; the parity suite itself is not in CI | -| Hygon DCU | ROCtracer | Beta: parity suite runs in CI | -| Moore Threads MUSA | MUPTI | Experimental: device timeline measured on MTT S5000; CPU-Kineto linkage is environment-dependent | -| Enflame GCU | TOPSPTI | Runtime only: TOPSPTI collects activities, but a CPU-only Kineto build supplies no PrivateUse1 resolver, so activities do not surface as device events | -| Other | `unavailable_device_tracer.cc` | Explicit no-device-activity fallback | - -## Correlation ids - -A trace carries two independent numbering schemes, both called "correlation". They look alike and mean different things: - -| | `correlation_id` | `external_correlation_id` | -|---|---|---| -| Whose id | The activity API's | PyTorch's | -| Pairs what | A runtime call with the device kernel it produced | A device/runtime activity with the CPU operator that issued it | -| Used for | Drawing flow arrows | Device time attribution | -| Trace field | `args["correlation"]` | `args["External id"]` | - -Passing the wrong one fails silently: the trace still renders, flow arrows disappear, and `self_device_time_total` reads 0. The device-time value must be resolved through the `getLinkedActivity` callback (external id), while flow arrows are keyed on the activity correlation id. - -## Debugging - -| Variable | Default | Purpose | -|---|---|---| -| `FLAGOS_TRACE` | `0` | Verbose logging in the device profiler shim: event draining, capture window, tracer binding and registration | -| `FLAGOS_TRACER_LIBRARY` | auto-discovered | Override the tracer library the profiler shim `dlopen`s when the default path does not match the installed driver | - -Two warnings are deliberately not gated by `FLAGOS_TRACE` — an empty linked-activity callback and an activity-record layout mismatch — because both silently zero device time. - -## Parity test and baseline - -`tests/integration/test_profiler_parity.py` compares a `flagos` trace against a baseline captured on native `torch+cuda`. All seven assertions check structure, never counts or durations: - -| # | Assertion | What it checks | -|---|---|---| -| 1 | `test_category_coverage` | The `flagos` trace emits every category the `torch.cuda` baseline does | -| 2 | `test_flow_arrows_are_paired` | Every flow-arrow start half has a matching finish half | -| 3 | `test_arg_key_supersets` | Each category's argument keys are a superset of the baseline's | -| 4 | `test_device_time_attribution` | An operator's `self_device_time_total` reconciles with the duration of the device events it owns | -| 5 | `test_kernel_names_are_demangled` | No bare mangled C++ symbols | -| 6 | `test_runtime_names_come_from_cbid` | Runtime event names are decoded from the callback id | -| 7 | `test_capture_window_containment` | No device or runtime event escapes the capture window | - -The baseline lives in `tests/data/profiler_cuda_baseline.json`. Assertion 2 is deliberately stricter than upstream `torch.cuda`: it is a Torch-FL invariant, kept because a regression into dangling flow halves is exactly the bug it guards against. - -## Known gaps - -- The `overhead` activity category is not collected: it measures the cost of profiling itself rather than the user's workload. It is recorded as a known gap in the baseline rather than omitted silently. -- MetaX MCPTI runtime callback ids are not NVIDIA CUPTI ids; the tracer uses the MetaX callback namespace and defers API-name resolution until after activity flushing, because calling the resolver from the buffer callback can deadlock the profiler. The scanner has not been validated across multiple MetaX SDK versions, devices, or non-default streams. -- Enflame GCU profiler support is runtime-only — see the table above. +# Profiler Integration + +`torch.profiler` supports the `flagos` device through a device tracer compiled into the wheel. A trace from `torch.profiler.profile(activities=[CPU, PrivateUse1])` is structurally equivalent to the trace of the same workload on `torch.cuda`: + +- **Flow arrows** connect each CPU operator to the device kernel it launched. +- **Device time attribution**: `prof.key_averages()` reports a per-operator `self_device_time_total`. +- **Complete kernel metadata** (grid/block, occupancy, shared memory, register count) and demangled kernel names. +- **Runtime events** carry real API names decoded from the callback id rather than a placeholder. +- **memcpy and memset** activities are collected alongside kernels. + +## Three-layer architecture + +Adding a vendor means writing one file: a tracer that satisfies the vendor-agnostic interface. + +| Layer | Files | Responsibility | +|---|---|---| +| Vendor-agnostic interface | `csrc/profiler/device_tracer.h` | `DeviceTracer`, `DeviceEvent`, `EventKind` — the complete contract a vendor implements | +| Vendor tracers | `cupti_device_tracer.cc`, `cann_device_tracer.cc`, `musa_mupti_device_tracer.cc`, `roctracer_device_tracer.cc`, `gcu_topspti_device_tracer.cc`, `unavailable_device_tracer.cc` | One implementation per accelerator, plus an explicit no-device-activity fallback | +| Generic adaptor | `flagos_kineto_profiler.{h,cc}` | Kineto/PyTorch profiler adaptor with zero vendor coupling; `dlopen` shims (`cupti_shim.h`, `mupti_shim.h`, `topspti_shim.h`, `roctracer`) bind the vendor activity library | + +| Accelerator | Activity API | Status | +|---|---|---| +| NVIDIA CUDA | CUPTI | Stable, parity suite included | +| MetaX | MCPTI (CUDA-compatible activity API in MACA) | Experimental: parity validated on C550 + MACA 3.8.0 (local validation, no vendor runner) | +| Ascend | MSPTI | Beta: kernel/runtime/flow/memcpy events plus device-time linkage, covered by the shared contract; the parity suite itself is not included | +| Hygon DCU | ROCtracer | Beta: parity suite included | +| Moore Threads MUSA | MUPTI | Experimental: device timeline measured on MTT S5000; CPU-Kineto linkage is environment-dependent | +| Enflame GCU | TOPSPTI | Runtime only: TOPSPTI collects activities, but a CPU-only Kineto build supplies no PrivateUse1 resolver, so activities do not surface as device events | +| Other | `unavailable_device_tracer.cc` | Explicit no-device-activity fallback | + +## Correlation ids + +A trace carries two independent numbering schemes, both called "correlation". They look alike and mean different things: + +| | `correlation_id` | `external_correlation_id` | +|---|---|---| +| Whose id | The activity API's | PyTorch's | +| Pairs what | A runtime call with the device kernel it produced | A device/runtime activity with the CPU operator that issued it | +| Used for | Drawing flow arrows | Device time attribution | +| Trace field | `args["correlation"]` | `args["External id"]` | + +Passing the wrong one fails silently: the trace still renders, flow arrows disappear, and `self_device_time_total` reads 0. The device-time value must be resolved through the `getLinkedActivity` callback (external id), while flow arrows are keyed on the activity correlation id. + +## Debugging + +| Variable | Default | Purpose | +|---|---|---| +| `FLAGOS_TRACE` | `0` | Verbose logging in the device profiler shim: event draining, capture window, tracer binding and registration | +| `FLAGOS_TRACER_LIBRARY` | auto-discovered | Override the tracer library the profiler shim `dlopen`s when the default path does not match the installed driver | + +Two warnings are deliberately not gated by `FLAGOS_TRACE` — an empty linked-activity callback and an activity-record layout mismatch — because both silently zero device time. + +## Parity test and baseline + +The parity test compares a `flagos` trace against a baseline captured on native `torch+cuda`. Every assertion checks structure, never counts or durations: + +| # | Assertion | What it checks | +|---|---|---| +| 1 | `test_category_coverage` | The `flagos` trace emits every category the `torch.cuda` baseline does | +| 2 | `test_flow_arrows_are_paired` | Every flow-arrow start half has a matching finish half | +| 3 | `test_arg_key_supersets` | Each category's argument keys are a superset of the baseline's | +| 4 | `test_device_time_attribution` | An operator's `self_device_time_total` reconciles with the duration of the device events it owns | +| 5 | `test_kernel_names_are_demangled` | No bare mangled C++ symbols | +| 6 | `test_runtime_names_come_from_cbid` | Runtime event names are decoded from the callback id | +| 7 | `test_capture_window_containment` | No device or runtime event escapes the capture window | + +The baseline is a trace captured on native `torch+cuda`. Assertion 2 is deliberately stricter than upstream `torch.cuda`: it is a Torch-FL invariant, kept because a regression into dangling flow halves is exactly the bug it guards against. + +## Known gaps + +- The `overhead` activity category is not collected: it measures the cost of profiling itself rather than the user's workload. It is recorded as a known gap in the baseline rather than omitted silently. +- MetaX MCPTI runtime callback ids are not NVIDIA CUPTI ids; the tracer uses the MetaX callback namespace and defers API-name resolution until after activity flushing, because calling the resolver from the buffer callback can deadlock the profiler. The scanner has not been validated across multiple MetaX SDK versions, devices, or non-default streams. +- Enflame GCU profiler support is runtime-only — see the table above. diff --git a/docs/torch_fl_en/architecture/torch-compile.md b/docs/torch_fl_en/architecture/torch-compile.md index 7b4bf7c6ef..4ca8308e19 100644 --- a/docs/torch_fl_en/architecture/torch-compile.md +++ b/docs/torch_fl_en/architecture/torch-compile.md @@ -1,98 +1,98 @@ -# torch.compile Integration - -The `flagos` device supports `torch.compile` for automatic kernel fusion and reduced dispatch overhead. The graph stays on the `flagos` device: there is no device round trip and no copy at the graph boundary. - -## Quick start - -```python -import torch_fl # Import first on MetaX and Ascend. -import torch - -model = torch.nn.Sequential( - torch.nn.Linear(512, 512), - torch.nn.ReLU(), - torch.nn.Linear(512, 512), -).to("flagos:0") - -model = torch.compile(model, backend="flagos") - -x = torch.randn(64, 512, device="flagos:0") -y = model(x) # Fused kernels -``` - -Compilation modes: - -```python -model = torch.compile(model, backend="flagos") # default -model = torch.compile(model, backend="flagos", mode="max-autotune") # longer compile, better runtime -model = torch.compile(model, backend="flagos", options={"max_autotune": True}) -``` - -`mode` and `options` are expanded into Inductor configuration patches scoped to that compile. CUDA graphs are always off for this backend, so `mode="reduce-overhead"` — whose main lever is CUDA graphs — has little effect. - -## FlagTree compilation - -[FlagTree](https://github.com/flagos-ai/FlagTree) is a Triton fork whose compiler targets many vendor backends. It integrates by **substitution at install time**, which is the whole thing to understand about it: - -- Its wheel is named `flagtree`, but the module it installs is `triton`. -- Installing it uninstalls the official `triton` and takes its place. -- Inductor's own `import triton` therefore already resolves to FlagTree once it is installed, and nothing in Torch-FL patches `sys.modules`. - -The backend compiler is selected at FlagTree build time through `FLAGTREE_BACKEND` (unset for NVIDIA and AMD), not at runtime: the same Triton kernel code compiles for a different vendor backend. `is_flagtree_active()` detects a FlagTree build, and `FLAGOS_USE_FLAGTREE=1` asserts that the active Triton is FlagTree — it errors rather than silently compiling with stock Triton. - -Because the FlagTree install removes the existing `triton`, build it in a separate virtualenv on a machine whose `triton` is in use by FlagGems. Wheels from FlagTree 0.6.2 on also install a real `flagtree` package (the FlagPrism debugger/profiler host); reaching FlagTree is still done through `triton`. - -## Backend internals - -| Component | Responsibility | -|---|---| -| `torch_fl/compile/inductor_backend.py` | Registers the `flagos` backend with `torch._dynamo` and wires Inductor's device interface | -| `torch_fl/compile/device_interface.py` | Inductor GPU device registration for the `flagos` device | -| `torch_fl/flagos/meta.py` | Meta kernels so tracing can infer output shapes for ops without a default meta implementation | -| `torch_fl/compile/flagtree_shim.py` | FlagTree detection (`is_flagtree_active`, `require_flagtree`); no import patching | -| `torch_fl/compile/platform_profile.py` | Per-platform codegen profiles and vendor workarounds | -| `torch_fl/compile/triton_*.py` | Triton integration guards: 64-bit indexing, byte loads, libdevice, resource limits | -| `torch_fl/compile/flagtree_ascend_policy.py` | Ascend backend policy for FlagTree's strategy registry, answering from `torch.flagos` instead of `torch_npu` | - -## Platform notes - -- **Ascend** — compilation through FlagTree's Ascend backend; the plugin's `flagos` policy answers FlagTree's strategy names from its own runtime and forces `TRITON_ENABLE_TASKQUEUE=false` (the task queue is `torch_npu`-only). Not covered in CI. -- **PPU** — FlagTree initializes CUDA while selecting compiler hints in an asynchronous Inductor worker, which can fail after the parent process initialized the PPU context, so PPU FlagTree defaults to serial compilation. Set `TORCHINDUCTOR_COMPILE_THREADS` explicitly only when testing an upstream fix or deliberately choosing another worker configuration. -- **MetaX** — `torch.compile` is validated with the vendor Triton and with FlagTree main in CUDA-boxing mode. -- **Enflame GCU** — the 64-bit codegen guard is what turns "no 64-bit support" failures into an actionable error naming the operator; see Troubleshooting. -- **D-Robotics BPU** — compilation is the *only* acceleration path: `torch.compile(backend="bpu")` traces a graph, compiles it to an `.hbm` artifact through hbdk4 and runs it on the BPU, with int8 quantization inserted by default. -- **CUDA** — `flagos` is registered as a first-class Inductor GPU device. There is no `torch.compile` step in the CUDA CI job, so the path is exercised by the integration test rather than by CI. - -## Performance - -Fusion gain is verified for correctness (`tests/integration/test_compile.py`); benchmarking the gain against stock Inductor on CUDA is still open work. Structurally the two should land close together — same fusion passes, same Triton codegen, no per-call copy — but that is an expectation, not a measurement. - -```bash -python tests/perf/bench_compile.py --model=mlp --batch-size=64 -python tests/perf/bench_compile.py --model=transformer --compare-cuda -FLAGOS_USE_FLAGTREE=1 python tests/perf/bench_compile.py -``` - -## Environment variables - -| Variable | Default | Purpose | -|---|---|---| -| `FLAGOS_USE_FLAGTREE` | `0` | Require the active Triton to be FlagTree (assert, not switch) | -| `FLAGOS_COMPILE_FALLBACK_EAGER` | `0` | Fall back to eager mode on compile errors | -| `FLAGOS_TILEOPS_*` | see reference | TileOps/TileLang L2 tier, instance-cache capacity and cache disabling | - -Routing variables such as `FLAGOS_BACKEND_CONFIG`, `FLAGOS_OP_` and `FLAGOS_FORCE_BACKEND` still apply to compiled kernels, because the compiled graph dispatches through the same routing table. - -## Troubleshooting - -| Symptom | Cause and fix | -|---|---| -| Compilation raises during graph capture or codegen | Set `FLAGOS_COMPILE_FALLBACK_EAGER=1` to fall back to eager; check for unsupported ops (dynamic shapes, custom ops) and missing meta implementations | -| `InductorError: ... has no 64-bit support` on GCU | The 64-bit codegen guard rejects kernel shapes that need 64-bit indexing on a backend without it; reduce the tensor/index sizes or keep the operator on its route | -| No speedup over eager | Verify the graph really compiled (`TORCH_LOGS=inductor`), check whether Triton autotuning is still running, and confirm the workload is not launch-bound | -| FlagTree not active | `FLAGOS_USE_FLAGTREE=1` raises when the active Triton is stock; check `is_flagtree_active()` and reinstall FlagTree, which must replace the `triton` module | -| `torch.compile` unavailable | The backend is registered when `torch._dynamo` is importable; check the PyTorch version and that `import torch_fl` ran before compiling | - -```bash -pytest tests/integration/test_compile.py -v --tb=short -``` +# torch.compile Integration + +The `flagos` device supports `torch.compile` for automatic kernel fusion and reduced dispatch overhead. The graph stays on the `flagos` device: there is no device round trip and no copy at the graph boundary. + +## Quick start + +```python +import torch_fl # Import first on MetaX and Ascend. +import torch + +model = torch.nn.Sequential( + torch.nn.Linear(512, 512), + torch.nn.ReLU(), + torch.nn.Linear(512, 512), +).to("flagos:0") + +model = torch.compile(model, backend="flagos") + +x = torch.randn(64, 512, device="flagos:0") +y = model(x) # Fused kernels +``` + +Compilation modes: + +```python +model = torch.compile(model, backend="flagos") # default +model = torch.compile(model, backend="flagos", mode="max-autotune") # longer compile, better runtime +model = torch.compile(model, backend="flagos", options={"max_autotune": True}) +``` + +`mode` and `options` are expanded into Inductor configuration patches scoped to that compile. CUDA graphs are always off for this backend, so `mode="reduce-overhead"` — whose main lever is CUDA graphs — has little effect. + +## FlagTree compilation + +[FlagTree](https://github.com/flagos-ai/FlagTree) is a Triton fork whose compiler targets many vendor backends. It integrates by **substitution at install time**, which is the whole thing to understand about it: + +- Its wheel is named `flagtree`, but the module it installs is `triton`. +- Installing it uninstalls the official `triton` and takes its place. +- Inductor's own `import triton` therefore already resolves to FlagTree once it is installed, and nothing in Torch-FL patches `sys.modules`. + +The backend compiler is selected at FlagTree build time through `FLAGTREE_BACKEND` (unset for NVIDIA and AMD), not at runtime: the same Triton kernel code compiles for a different vendor backend. `is_flagtree_active()` detects a FlagTree build, and `FLAGOS_USE_FLAGTREE=1` asserts that the active Triton is FlagTree — it errors rather than silently compiling with stock Triton. + +Because the FlagTree install removes the existing `triton`, build it in a separate virtualenv on a machine whose `triton` is in use by FlagGems. Recent FlagTree wheels also install a real `flagtree` package (the FlagPrism debugger/profiler host); reaching FlagTree is still done through `triton`. + +## Backend internals + +| Component | Responsibility | +|---|---| +| `torch_fl/compile/inductor_backend.py` | Registers the `flagos` backend with `torch._dynamo` and wires Inductor's device interface | +| `torch_fl/compile/device_interface.py` | Inductor GPU device registration for the `flagos` device | +| `torch_fl/flagos/meta.py` | Meta kernels so tracing can infer output shapes for ops without a default meta implementation | +| `torch_fl/compile/flagtree_shim.py` | FlagTree detection (`is_flagtree_active`, `require_flagtree`); no import patching | +| `torch_fl/compile/platform_profile.py` | Per-platform codegen profiles and vendor workarounds | +| `torch_fl/compile/triton_*.py` | Triton integration guards: 64-bit indexing, byte loads, libdevice, resource limits | +| `torch_fl/compile/flagtree_ascend_policy.py` | Ascend backend policy for FlagTree's strategy registry, answering from `torch.flagos` instead of `torch_npu` | + +## Platform notes + +- **Ascend** — compilation through FlagTree's Ascend backend; the plugin's `flagos` policy answers FlagTree's strategy names from its own runtime and forces `TRITON_ENABLE_TASKQUEUE=false` (the task queue is `torch_npu`-only). Exercised by the integration test only. +- **PPU** — FlagTree initializes CUDA while selecting compiler hints in an asynchronous Inductor worker, which can fail after the parent process initialized the PPU context, so PPU FlagTree defaults to serial compilation. Set `TORCHINDUCTOR_COMPILE_THREADS` explicitly only when testing an upstream fix or deliberately choosing another worker configuration. +- **MetaX** — `torch.compile` is validated with the vendor Triton and with FlagTree main in CUDA-boxing mode. +- **Enflame GCU** — the 64-bit codegen guard is what turns "no 64-bit support" failures into an actionable error naming the operator; see Troubleshooting. +- **D-Robotics BPU** — compilation is the *only* acceleration path: `torch.compile(backend="bpu")` traces a graph, compiles it to an `.hbm` artifact through hbdk4 and runs it on the BPU, with int8 quantization inserted by default. +- **CUDA** — `flagos` is registered as a first-class Inductor GPU device. The path is exercised by the integration test rather than by a dedicated job. + +## Performance + +Fusion gain is verified for correctness (`tests/integration/test_compile.py`); benchmarking the gain against stock Inductor on CUDA is still open work. Structurally the two should land close together — same fusion passes, same Triton codegen, no per-call copy — but that is an expectation, not a measurement. + +```bash +python tests/perf/bench_compile.py --model=mlp --batch-size=64 +python tests/perf/bench_compile.py --model=transformer --compare-cuda +FLAGOS_USE_FLAGTREE=1 python tests/perf/bench_compile.py +``` + +## Environment variables + +| Variable | Default | Purpose | +|---|---|---| +| `FLAGOS_USE_FLAGTREE` | `0` | Require the active Triton to be FlagTree (assert, not switch) | +| `FLAGOS_COMPILE_FALLBACK_EAGER` | `0` | Fall back to eager mode on compile errors | +| `FLAGOS_TILEOPS_*` | see reference | TileOps/TileLang L2 tier, instance-cache capacity and cache disabling | + +Routing variables such as `FLAGOS_BACKEND_CONFIG`, `FLAGOS_OP_` and `FLAGOS_FORCE_BACKEND` still apply to compiled kernels, because the compiled graph dispatches through the same routing table. + +## Troubleshooting + +| Symptom | Cause and fix | +|---|---| +| Compilation raises during graph capture or codegen | Set `FLAGOS_COMPILE_FALLBACK_EAGER=1` to fall back to eager; check for unsupported ops (dynamic shapes, custom ops) and missing meta implementations | +| `InductorError: ... has no 64-bit support` on GCU | The 64-bit codegen guard rejects kernel shapes that need 64-bit indexing on a backend without it; reduce the tensor/index sizes or keep the operator on its route | +| No speedup over eager | Verify the graph really compiled (`TORCH_LOGS=inductor`), check whether Triton autotuning is still running, and confirm the workload is not launch-bound | +| FlagTree not active | `FLAGOS_USE_FLAGTREE=1` raises when the active Triton is stock; check `is_flagtree_active()` and reinstall FlagTree, which must replace the `triton` module | +| `torch.compile` unavailable | The backend is registered when `torch._dynamo` is importable; check the PyTorch version and that `import torch_fl` ran before compiling | + +```bash +pytest tests/integration/test_compile.py -v --tb=short +``` diff --git a/docs/torch_fl_en/getting_started/installation.md b/docs/torch_fl_en/getting_started/installation.md index 59172ad955..c5e5012196 100644 --- a/docs/torch_fl_en/getting_started/installation.md +++ b/docs/torch_fl_en/getting_started/installation.md @@ -1,223 +1,255 @@ -# Installation - -## Choose a platform - -Each platform is selected by a single build variable, `FLAGOS_ACCELERATOR`, and each platform has its own execution path: - -| Platform | `FLAGOS_ACCELERATOR` | Execution path | Status | -|---|---|---|---| -| NVIDIA CUDA | `cuda` (default) | CUDA boxing over an external `libtorch_cuda.so` | Stable | -| MetaX | `metax` | CUDA boxing via `cu-bridge` against the vendor libtorch | Stable | -| Ascend | `ascend` | Native ACLNN operator backend, FlagGems via FlagTree (Triton 3.5) | Beta | -| PPU | `ppu` | Same CUDA-boxing path as NVIDIA CUDA, against the PPU CUDA-13-compatible SDK, bundling its own libtorch | Experimental | -| Hygon DCU | `dcu` | CUDA boxing over the hipified DTK torch build | Beta | -| Enflame GCU | `gcu` | Native `libtopsaten.so` operator backend, with CPU fallback for unrouted/int64/float64 ops | Beta | -| Moore Threads MUSA | `musa` | FlagGems-first Triton kernels, native `mudnn` backend as fallback, CPU fallback for unrouted ops | Experimental | -| D-Robotics BPU | `bpu` | No eager kernel sets are built; eager ops run on CPU, acceleration comes from the graph path | Runtime only | -| TsingMicro | `tsingmicro` | Runtime/build selector present; no per-operator kernel set documented | Runtime only | - -## Common requirements - -All platforms require: - -- **Python**: 3.8 or later (platform SDKs and available wheels may impose a narrower range) -- **PyTorch**: 2.10.x (`>=2.10,<2.11`) — the generated ATen bindings are tied to this minor line -- **CMake**: 3.18 or later -- **C++ toolchain**: a working C++17 compiler (GCC 7+, Clang 5+, or MSVC 2017+) -- **Platform SDK/runtime**: the vendor-specific SDK, compiler and runtime libraries for your accelerator - -Patch releases inside the same PyTorch minor line (for example 2.10.0 to 2.10.1) are compatible. A different minor line (for example 2.11.x) fails at build or run time, because the generated bindings are sensitive to C++ ABI and operator schema changes. - -## Source installation contract - -Every platform installation follows the same pattern: - -```bash -FLAGOS_ACCELERATOR= pip install --no-build-isolation -e . -``` - -`--no-build-isolation` is required: without it, pip creates an isolated build environment that cannot see the PyTorch installation and platform SDK in your current environment, and the generated native bindings then link against the wrong torch or fail to find the vendor SDK. - -Each platform directory in the upstream repository (`docs/vendors//installation.md`) defines the platform's SDK environment variables and any extra build flags. Platform summaries: - -### NVIDIA CUDA - -```bash -git clone https://github.com/flagos-ai/Torch-FL.git && cd Torch-FL - -pip install torch==2.10.0+cpu --index-url https://download.pytorch.org/whl/cpu - -FLAGOS_ACCELERATOR=cuda pip install --no-build-isolation -vvv -e . -``` - -This generates the CUDA-boxing kernels from PyTorch's ATen schema, bundles `libtorch_cuda.so` and the related CUDA dispatcher libraries into `torch_fl/lib/`, and pins matching `nvidia-*-cu12` runtime dependencies. Requirements: an NVIDIA GPU with compute capability 7.0 or later, driver 470 or later, a CUDA 12.x toolkit for the build, plus `cmake`, `ninja` and `patchelf`. - -Optional FlagGems C++ dispatch (lowest-overhead FlagGems route): - -```bash -FLAGOS_ACCELERATOR=cuda \ - FLAGOS_BUILD_FLAGGEMS_CPP=1 \ - FLAGGEMS_DIR=/lib/cmake/FlagGems \ - pip install --no-build-isolation -vvv -e . -``` - -### MetaX - -MetaX ships a self-contained boxing wheel: CUDA-boxing kernels compiled with the host `g++`, plus the MetaX-forked libtorch C++ runtime bundled inside the wheel. The target machine needs only the stock `torch==2.10.0+cpu` wheel, the `torch_fl` wheel, and the `/opt/maca` driver runtime. - -The wheel is built on a machine with the full MACA SDK and a `torch+metax` wheel (both from the MetaX developer portal), in three steps: build the boxing artifacts, bundle the forked libtorch with `scripts/vendor/bundle_maca_libtorch.sh`, then repackage the wheel. The result is a large wheel (the bundled libtorch puts it above PyPI's size limit), distributed through a private index or a direct transfer. - -On the target host: - -```bash -pip install torch==2.10.0+cpu --index-url https://download.pytorch.org/whl/cpu -pip install torch_fl-+.whl -``` - -`FLAGOS_WHEEL_LOCAL` records the target SDK in the wheel's local version label (for example `0.1.0+metax3.8.1`), which keeps two SDK-incompatible wheels from being indistinguishable by filename. - -**Import order matters on MetaX**: import `torch_fl` before `import torch`, because PyTorch's bundled CUDA 12.x runtime is ABI-incompatible with MACA's `cu-bridge` and `torch_fl` preloads a shim that supplies the required symbol versions. - -### Ascend - -```bash -pip install torch==2.10.0 --index-url https://download.pytorch.org/whl/cpu -source /usr/local/Ascend/ascend-toolkit/set_env.sh - -FLAGOS_ACCELERATOR=ascend pip install --no-build-isolation -v -e . -``` - -Requirements: an Ascend 910 with CANN 9.0.0 or compatible, accessible `/dev/davinci*` device nodes, and Python 3.11 if you use the FlagTree Ascend 3.5 wheel (it is cp311-only; Python 3.8+ works for an ACLNN-only build). - -The default configuration enables the FlagGems Python route for measured operators, with native ACLNN kernels as the fallback. `scripts/codegen/codegen_ascend.py` generates the ACLNN kernels; operators without an ACLNN mapping fall back to CPU. - -FlagGems on Ascend runs on **FlagTree** (the FlagOS Triton fork) on its Ascend 3.5 line. FlagTree installs the module named `triton`, so remove any stock or vendor Triton first, then install FlagTree and FlagGems from the FlagOS index. Nothing in this path imports or links `torch_npu`: the backend installs a lightweight `torch_npu` stub and registers its own `flagos` policy on FlagTree's strategy registry. - -**Import order matters on Ascend** too: import `torch_fl` before packages that might register a device backend. - -### PPU - -PPU presents itself as a CUDA-compatible device: its torch wheel is a full CUDA 13 build and it registers operators under the `CUDA` dispatch key, so no stock `+cpu` wheel and no external `libtorch_cuda.so` are needed. - -```bash -FLAGOS_ACCELERATOR=ppu \ - CUDA_HOME=/usr/local/PPU_SDK/CUDA_SDK \ - FLAGOS_BUILD_FLAGGEMS_CPP=OFF \ - FLAGOS_BUILD_FLAGGEMS=OFF \ - FLAGOS_SKIP_CUDA_ASSETS=1 \ - pip install --no-build-isolation -vvv -e . -``` - -At runtime, export `FLAGOS_DISABLE_CUDA_ASSETS=1` so the import-time preload of a bundled `libtorch_cuda.so` is a no-op (PPU torch provides the runtime itself). - -### Hygon DCU - -```bash -source /opt/dtk/env.sh - -FLAGOS_ACCELERATOR=dcu pip install --no-build-isolation -vvv -e . -``` - -The build is pure boxing: the DTK torch wheel registers HIP kernels under the `CUDA` dispatch key, generated boxing kernels dispatch into `libtorch_hip.so` unchanged, and runtime sources compile with plain host `g++` — no `nvcc`, `hipcc` or hipify pass. FlagGems Python is on by default; the FlagGems C++ path stays off because DTK ships no `liboperators.so`. - -### Enflame GCU - -```bash -pip install torch==2.10.0 --index-url https://download.pytorch.org/whl/cpu - -FLAGOS_ACCELERATOR=gcu pip install --no-build-isolation -v -e . -``` - -Requires the TopsRider SDK (`libtopsrt.so` runtime and `libtopsaten.so` operator library). The build runs `scripts/codegen/codegen_gcu.py`, validating each operator against the `topsaten` symbols actually present in the installed `libtopsaten.so`; operators missing from the SDK are skipped with a warning. The `torch-gcu` wheel cannot be used alongside Torch-FL, because it claims `PrivateUse1` for itself. - -### Moore Threads MUSA - -```bash -pip install torch==2.10.0 --index-url https://download.pytorch.org/whl/cpu - -FLAGOS_ACCELERATOR=musa pip install --no-build-isolation -v -e . -``` - -Requires the MUSA toolkit under `/usr/local/musa`: `musart` (runtime), `mudnn` (operator library) and `murand` (device RNG). The build runs `scripts/codegen/codegen_mudnn.py`; coverage is the generated operator set plus handwritten convolution kernels, with native RNG kernels and CPU fallback for everything else. `--no-build-isolation` is mandatory here: without it, pip's build overlay resolves its own torch and `import torch_fl` then fails with an undefined `c10` symbol. - -### D-Robotics BPU - -```bash -pip install torch==2.10.0+cpu --index-url https://download.pytorch.org/whl/cpu -FLAGOS_ACCELERATOR=bpu pip install --no-build-isolation -e . -``` - -The BPU platform provides runtime acceleration only: no per-operator BPU kernels exist, eager operators run on the CPU via fallback, and acceleration comes from whole-graph compilation (`torch.compile(backend="bpu")`) or from the prebuilt-HBM LLM runtime. Graph compilation needs `hbdk4`, which ships x86_64-only wheels. - -## Build-time switches - -`setup.py` forces a per-accelerator value for the kernel-set switches and rejects an explicit environment value that contradicts it: - -| Variable | Default | Purpose | -|---|---|---| -| `FLAGOS_ACCELERATOR` | `cuda` | Hardware platform the wheel is built for | -| `FLAGOS_BUILD_VENDOR` | `ON`, `OFF` on `metax` | Compile the vendor's native kernels (a no-op where the vendor ships none) | -| `FLAGOS_BUILD_FLAGGEMS` | `ON`, `OFF` on `bpu` | Compile the FlagGems Python kernel wrappers | -| `FLAGOS_BUILD_FLAGGEMS_CPP` | `ON` on `cuda`/`tsingmicro` | Compile the FlagGems C++ wrapper (`liboperators.so`) | -| `FLAGOS_BUILD_BOXING` | `ON`, `OFF` on `ascend`/`gcu`/`musa` | Compile the generated CUDA-boxing kernels | -| `FLAGOS_BUILD_TILEOPS` | `ON` on `cuda` | Compile the TileOps kernel wrappers (SM90 NVIDIA parts) | -| `FLAGOS_BUILD_JOBS` | CPU count | Parallel jobs for the CMake build | -| `FLAGOS_SKIP_CUDA_ASSETS` | `0` | Do not bundle an external `libtorch_cuda.so` | -| `FLAGOS_WHEEL_LOCAL` | SDK-derived | Local version label for the wheel | - -The wheel records the accelerator and kernel sets it was built with in `torch_fl/_build_config.py`, and every runtime reader consults that record — a stale exported variable cannot make a wheel describe itself wrongly. Full details of every variable are in the {doc}`environment variable reference <../reference/environment-variables>`. - -## Verification - -```bash -python -c " -import torch_fl -import torch - -print(f'PyTorch version: {torch.__version__}') -print(f'flagos devices: {torch.flagos.device_count()}') -print(f'flagos available: {torch.flagos.is_available()}') - -x = torch.randn(4, 4, device='flagos:0') -y = (x @ x).sum() -print(f'Sample result: {y.cpu().item():.4f}') -" -``` - -A working installation reports the device count of the current machine and a floating-point result. If the count is 0, check the vendor SDK installation, the driver, the device nodes, and the import order rules above. - -## Tests - -```bash -# Unit tests: no hardware dependency -pytest tests/unit -q - -# Operator correctness on the active backend -pytest tests/integration/ops/ -m main_ops -v --tb=short - -# Factory operators respect device placement -pytest tests/integration/test_factory_ops.py -v --tb=short -``` - -Operator tests are selected by markers registered in `tests/integration/ops/conftest.py`: - -| Marker | Meaning | -|---|---| -| `main_ops` | Representative operator in the CI smoke subset | -| `anyplatform` | Runs on any accelerator backend | -| `cuda`, `metax`, `ascend`, `musa` | Requires that backend's kernels or hardware | -| `flaggems` | Asserts the FlagGems route from `backends_.conf` | -| `flaggems_python` | Requires the FlagGems Python wrapper backend | -| `flaggems_cpp` | Requires a wheel built with `FLAGOS_BUILD_FLAGGEMS_CPP=ON` | - -Cross-backend contracts (profiler, AMP) are selected with the `profiler*` and `amp*` markers from `tests/integration/conftest.py`; a test needing a capability the active platform does not provide skips with a reason naming the platform. - -Test filtering is automatic: the conftest detects the platform from the wheel's build record, the installed platform marker, or the routing table name, and skips tests marked for unavailable backends. Unit tests, model tests and the manual suites are documented in the upstream `docs/development/testing.md`. - -## Next steps - -- {doc}`Compatibility matrix <../reference/compatibility>` — per-platform capability validation -- {doc}`Environment variables <../reference/environment-variables>` — build and runtime configuration -- {doc}`Distributed collectives <../architecture/distributed>` — `ProcessGroupFlagOS` and FlagCX -- {doc}`Profiler <../architecture/profiler>` — `torch.profiler` integration -- {doc}`torch.compile <../architecture/torch-compile>` — Inductor and FlagTree integration +# Installation + +## Choose a platform + +Each platform is selected by a single build variable, `FLAGOS_ACCELERATOR`, and each platform has its own execution path: + +| Platform | `FLAGOS_ACCELERATOR` | Execution path | Status | +|---|---|---|---| +| NVIDIA CUDA | `cuda` (default) | CUDA boxing over an external `libtorch_cuda.so` | Stable | +| MetaX | `metax` | CUDA boxing via `cu-bridge` against the vendor libtorch | Stable | +| Ascend | `ascend` | Native ACLNN operator backend, FlagGems via FlagTree (Triton 3.5) | Beta | +| PPU | `ppu` | Same CUDA-boxing path as NVIDIA CUDA, against the PPU CUDA-13-compatible SDK, bundling its own libtorch | Experimental | +| Hygon DCU | `dcu` | CUDA boxing over the hipified DTK torch build | Beta | +| Enflame GCU | `gcu` | Native `libtopsaten.so` operator backend, with CPU fallback for unrouted/int64/float64 ops | Beta | +| Moore Threads MUSA | `musa` | FlagGems-first Triton kernels, native `mudnn` backend as fallback, CPU fallback for unrouted ops | Experimental | +| D-Robotics BPU | `bpu` | No eager kernel sets are built; eager ops run on CPU, acceleration comes from the graph path | Runtime only | +| TsingMicro | `tsingmicro` | Runtime/build selector present; no per-operator kernel set documented | Runtime only | + +## Common requirements + +All platforms require: + +- **Python**: one version per platform, not a range — 3.12 on CUDA/GCU/MetaX/PPU, 3.10 on DCU/MUSA, 3.11 on Ascend. A FlagTree build exists for exactly one cp tag and the wheel links it, so `Requires-Python` names a single interpreter +- **PyTorch**: 2.10.x (`>=2.10,<2.11`) — the generated ATen bindings are tied to this minor line +- **CMake**: 3.18 or later +- **C++ toolchain**: a working C++17 compiler (GCC 7+, Clang 5+, or MSVC 2017+) +- **Platform SDK/runtime**: the vendor-specific SDK, compiler and runtime libraries for your accelerator + +Patch releases inside the same PyTorch minor line (for example 2.10.0 to 2.10.1) are compatible. A different minor line (for example 2.11.x) fails at build or run time, because the generated bindings are sensitive to C++ ABI and operator schema changes. + +## Source installation contract + +Every platform installation follows the same pattern: + +```bash +FLAGOS_ACCELERATOR= pip install --no-build-isolation -e . +``` + +`--no-build-isolation` is required: without it, pip creates an isolated build environment that cannot see the PyTorch installation and platform SDK in your current environment, and the generated native bindings then link against the wrong torch or fail to find the vendor SDK. + +Each platform directory in the upstream repository (`docs/vendors//installation.md`) defines the platform's SDK environment variables and any extra build flags. Platform summaries: + +### NVIDIA CUDA + +```bash +git clone https://github.com/flagos-ai/Torch-FL.git && cd Torch-FL + +pip install torch==2.10.0+cpu --index-url https://download.pytorch.org/whl/cpu + +FLAGOS_ACCELERATOR=cuda pip install --no-build-isolation -vvv -e . +``` + +This generates the CUDA-boxing kernels from PyTorch's ATen schema, bundles `libtorch_cuda.so` and the related CUDA dispatcher libraries into `torch_fl/lib/`, and pins matching `nvidia-*-cu12` runtime dependencies. Requirements: an NVIDIA GPU with compute capability 7.0 or later, a driver new enough for the runtime the wheel bundles (a cu12.8 `libtorch_cuda.so` plus matching `nvidia-*-cu12` wheels), a CUDA toolkit providing `nvcc` and headers for the build, plus `cmake`, `ninja` and `patchelf`. + +Optional FlagGems C++ dispatch (lowest-overhead FlagGems route): + +```bash +FLAGOS_ACCELERATOR=cuda \ + FLAGOS_BUILD_FLAGGEMS_CPP=1 \ + FLAGGEMS_DIR=/lib/cmake/FlagGems \ + pip install --no-build-isolation -vvv -e . +``` + +### MetaX + +MetaX ships a self-contained boxing wheel: CUDA-boxing kernels compiled with the host `g++`, plus the MetaX-forked libtorch C++ runtime bundled inside the wheel. The target machine needs only the stock `torch==2.10.0+cpu` wheel, the `torch_fl` wheel, and the `/opt/maca` driver runtime. + +The wheel is built on a machine with the full MACA SDK and a `torch+metax` wheel (both from the MetaX developer portal), in three steps: build the boxing artifacts, bundle the forked libtorch with `scripts/vendor/bundle_maca_libtorch.sh`, then repackage the wheel. The result is a large wheel (the bundled libtorch puts it above PyPI's size limit), distributed through a private index or a direct transfer. + +On the target host: + +```bash +pip install torch==2.10.0+cpu --index-url https://download.pytorch.org/whl/cpu +pip install torch_fl-+.whl +``` + +`FLAGOS_WHEEL_LOCAL` records the target SDK in the wheel's local version label (for example `2.10.0+maca3.8.1.3`), which keeps two SDK-incompatible wheels from being indistinguishable by filename. + +**Import order matters on MetaX**: import `torch_fl` before `import torch`, because PyTorch's bundled CUDA 12.x runtime is ABI-incompatible with MACA's `cu-bridge` and `torch_fl` preloads a shim that supplies the required symbol versions. + +### Ascend + +```bash +pip install torch==2.10.0 --index-url https://download.pytorch.org/whl/cpu +source /usr/local/Ascend/ascend-toolkit/set_env.sh + +FLAGOS_ACCELERATOR=ascend pip install --no-build-isolation -v -e . +``` + +Requirements: an Ascend 910 with CANN 9.0.0 or compatible, accessible `/dev/davinci*` device nodes, and Python 3.11 — the FlagTree Ascend wheel is cp311-only and the wheel's `Requires-Python` names that single interpreter. + +The default configuration enables the FlagGems Python route for measured operators, with native ACLNN kernels as the fallback. `scripts/codegen/codegen_ascend.py` generates the ACLNN kernels; operators without an ACLNN mapping fall back to CPU. + +FlagGems on Ascend runs on **FlagTree** (the FlagOS Triton fork) on its Ascend 3.5 line. FlagTree installs the module named `triton`, so remove any stock or vendor Triton first, then install FlagTree and FlagGems from the FlagOS index. Nothing in this path imports or links `torch_npu`: the backend installs a lightweight `torch_npu` stub and registers its own `flagos` policy on FlagTree's strategy registry. + +**Import order matters on Ascend** too: import `torch_fl` before packages that might register a device backend. + +### PPU + +PPU presents itself as a CUDA-compatible device: its torch wheel is a full CUDA 13 build and it registers operators under the `CUDA` dispatch key, so no stock `+cpu` wheel and no external `libtorch_cuda.so` are needed. + +```bash +FLAGOS_ACCELERATOR=ppu \ + CUDA_HOME=/usr/local/PPU_SDK/CUDA_SDK \ + FLAGOS_BUILD_FLAGGEMS_CPP=OFF \ + FLAGOS_BUILD_FLAGGEMS=OFF \ + FLAGOS_SKIP_CUDA_ASSETS=1 \ + pip install --no-build-isolation -vvv -e . +``` + +At runtime, export `FLAGOS_DISABLE_CUDA_ASSETS=1` so the import-time preload of a bundled `libtorch_cuda.so` is a no-op (PPU torch provides the runtime itself). + +### Hygon DCU + +```bash +source /opt/dtk/env.sh + +FLAGOS_ACCELERATOR=dcu pip install --no-build-isolation -vvv -e . +``` + +The build is pure boxing: the DTK torch wheel registers HIP kernels under the `CUDA` dispatch key, generated boxing kernels dispatch into `libtorch_hip.so` unchanged, and runtime sources compile with plain host `g++` — no `nvcc`, `hipcc` or hipify pass. FlagGems Python is on by default; the FlagGems C++ path stays off because DTK ships no `liboperators.so`. + +### Enflame GCU + +```bash +pip install torch==2.10.0 --index-url https://download.pytorch.org/whl/cpu + +FLAGOS_ACCELERATOR=gcu pip install --no-build-isolation -v -e . +``` + +Requires the TopsRider SDK (`libtopsrt.so` runtime and `libtopsaten.so` operator library). The build runs `scripts/codegen/codegen_gcu.py`, validating each operator against the `topsaten` symbols actually present in the installed `libtopsaten.so`; operators missing from the SDK are skipped with a warning. The `torch-gcu` wheel cannot be used alongside Torch-FL, because it claims `PrivateUse1` for itself. + +### Moore Threads MUSA + +```bash +pip install torch==2.10.0 --index-url https://download.pytorch.org/whl/cpu + +FLAGOS_ACCELERATOR=musa pip install --no-build-isolation -v -e . +``` + +Requires the MUSA toolkit under `/usr/local/musa`: `musart` (runtime), `mudnn` (operator library) and `murand` (device RNG). The build runs `scripts/codegen/codegen_mudnn.py`; coverage is the generated operator set plus handwritten convolution kernels, with native RNG kernels and CPU fallback for everything else. `--no-build-isolation` is mandatory here: without it, pip's build overlay resolves its own torch and `import torch_fl` then fails with an undefined `c10` symbol. + +### D-Robotics BPU + +```bash +pip install torch==2.10.0+cpu --index-url https://download.pytorch.org/whl/cpu +FLAGOS_ACCELERATOR=bpu pip install --no-build-isolation -e . +``` + +The BPU platform provides runtime acceleration only: no per-operator BPU kernels exist, eager operators run on the CPU via fallback, and acceleration comes from whole-graph compilation (`torch.compile(backend="bpu")`) or from the prebuilt-HBM LLM runtime. Graph compilation needs `hbdk4`, which ships x86_64-only wheels. + +## Runtime dependencies and the package index + +A wheel declares the three packages built outside this repository that it cannot work without — **FlagTree** (the Triton build carrying the vendor's backend), **FlagGems** (the operator source) and **FlagCX** (the distributed backend) — pinned to the exact versions it was built against, taken from `.github/version-pins.env`. Nothing about them is a range: `flagtree` and `flagcx` are not on PyPI at all, a FlagTree build is per-platform (its package name carries the vendor's Triton backend), and the `flag_gems` on PyPI is an older cohort than the one the per-op routing tables in `torch_fl/configs/backends_*.conf` were generated against. + +That means the index has to carry more than one location: + +| Requirement | Where it is published | +|---|---| +| `torch_fl` | `flagos-pypi-` — the lane named by the wheel's local version (`2.10.0+hygon` → `flagos-pypi-hygon`) | +| `flag_gems`, `flagcx` | the same vendor lane | +| `flagtree` | `flagos-pypi-hosted`, for every platform | +| `torch==2.10.0+cpu` | `https://download.pytorch.org/whl/cpu` | +| everything else (`packaging`, `PyYAML`, `numpy`, …) | PyPI (or a mirror) | + +A single `--index-url` therefore has to name a **group repository** that contains all of those. Where one is not configured, list them instead — this is the DCU case: + +```bash +BASE=https://resource.flagos.net/repository +pip install \ + --index-url "$BASE/flagos-pypi-hygon/simple/" \ + --extra-index-url "$BASE/flagos-pypi-hosted/simple/" \ + --extra-index-url "$BASE/pypi-proxy/simple/" \ + --extra-index-url "https://download.pytorch.org/whl/cpu" \ + torch_fl==2.10.0+hygon +``` + +Two things are easy to get wrong here: the vendor lane alone is not enough even where it already carries all three FlagOS packages (`flag_gems` itself declares `packaging>=26.0` and `PyYAML==6.0.1`, which the lanes do not serve), and `flagtree` is not in most lanes — it comes from `flagos-pypi-hosted` for every platform. + +## Build-time switches + +`setup.py` forces a per-accelerator value for the kernel-set switches and rejects an explicit environment value that contradicts it: + +| Variable | Default | Purpose | +|---|---|---| +| `FLAGOS_ACCELERATOR` | `cuda` | Hardware platform the wheel is built for | +| `FLAGOS_BUILD_VENDOR` | `ON`, `OFF` on `metax` | Compile the vendor's native kernels (a no-op where the vendor ships none) | +| `FLAGOS_BUILD_FLAGGEMS` | `ON`, `OFF` on `bpu` | Compile the FlagGems Python kernel wrappers | +| `FLAGOS_BUILD_FLAGGEMS_CPP` | `ON` on `cuda`/`tsingmicro` | Compile the FlagGems C++ wrapper (`liboperators.so`) | +| `FLAGOS_BUILD_BOXING` | `ON`, `OFF` on `ascend`/`gcu`/`musa` | Compile the generated CUDA-boxing kernels | +| `FLAGOS_BUILD_TILEOPS` | `ON` on `cuda` | Compile the TileOps kernel wrappers (SM90 NVIDIA parts) | +| `FLAGOS_BUILD_JOBS` | CPU count | Parallel jobs for the CMake build | +| `FLAGOS_SKIP_CUDA_ASSETS` | `0` | Do not bundle an external `libtorch_cuda.so` | +| `FLAGOS_WHEEL_LOCAL` | SDK-derived | Local version label for the wheel | + +The wheel records the accelerator and kernel sets it was built with in `torch_fl/_build_config.py`, and every runtime reader consults that record — a stale exported variable cannot make a wheel describe itself wrongly. Full details of every variable are in the {doc}`environment variable reference <../reference/environment-variables>`. + +## Verification + +```bash +python -c " +import torch_fl +import torch + +print(f'PyTorch version: {torch.__version__}') +print(f'flagos devices: {torch.flagos.device_count()}') +print(f'flagos available: {torch.flagos.is_available()}') + +x = torch.randn(4, 4, device='flagos:0') +y = (x @ x).sum() +print(f'Sample result: {y.cpu().item():.4f}') +" +``` + +A working installation reports the device count of the current machine and a floating-point result. If the count is 0, check the vendor SDK installation, the driver, the device nodes, and the import order rules above. + +## Tests + +```bash +# Unit tests: no hardware dependency +pytest tests/unit -q + +# Operator correctness on the active backend +pytest tests/integration/ops/ -m main_ops -v --tb=short + +# Factory operators respect device placement +pytest tests/integration/test_factory_ops.py -v --tb=short +``` + +Operator tests are selected by markers registered in `tests/integration/ops/conftest.py`: + +| Marker | Meaning | +|---|---| +| `main_ops` | Representative operator in the smoke subset | +| `anyplatform` | Runs on any accelerator backend | +| `cuda`, `metax`, `ascend`, `musa` | Requires that backend's kernels or hardware | +| `flaggems` | Asserts the FlagGems route from `backends_.conf` | +| `flaggems_python` | Requires the FlagGems Python wrapper backend | +| `flaggems_cpp` | Requires a wheel built with `FLAGOS_BUILD_FLAGGEMS_CPP=ON` | + +Cross-backend contracts (profiler, AMP) are selected with the `profiler*` and `amp*` markers from `tests/integration/conftest.py`; a test needing a capability the active platform does not provide skips with a reason naming the platform. + +Test filtering is automatic: the conftest detects the platform from the wheel's build record, the installed platform marker, or the routing table name, and skips tests marked for unavailable backends. Unit tests, model tests and the manual suites are documented in the upstream `docs/development/testing.md`. + +## Next steps + +- {doc}`Quick start ` — platform-independent usage patterns +- {doc}`Compatibility matrix <../reference/compatibility>` — per-platform capability validation +- {doc}`Platform capability matrix <../reference/platform-capability>` — what each accelerator builds and routes +- {doc}`Dtype support <../reference/dtype-support>` — storage, AMP targets and fallback boundaries +- {doc}`Troubleshooting <../reference/troubleshooting>` — device count 0, missing libraries, compiler errors +- {doc}`Environment variables <../reference/environment-variables>` — build and runtime configuration +- {doc}`Distributed collectives <../architecture/distributed>` — `ProcessGroupFlagOS` and FlagCX +- {doc}`Profiler <../architecture/profiler>` — `torch.profiler` integration +- {doc}`torch.compile <../architecture/torch-compile>` — Inductor and FlagTree integration diff --git a/docs/torch_fl_en/getting_started/quickstart.md b/docs/torch_fl_en/getting_started/quickstart.md new file mode 100644 index 0000000000..8e6de50f5c --- /dev/null +++ b/docs/torch_fl_en/getting_started/quickstart.md @@ -0,0 +1,96 @@ +# Quick start + +This page shows platform-independent usage. After installing for your platform (see {doc}`Installation `), the same code runs on every supported accelerator. + +## Basic usage + +```python +import torch +import torch_fl + +x = torch.randn(4, 4, device="flagos:0") +y = torch.relu(x @ x) +print(y.cpu()) +``` + +The tensors are created on the first `flagos` device, the matrix multiply and activation are routed to platform-appropriate kernels, and the result is copied back to the CPU for printing. + +## Moving tensors between devices + +```python +import torch +import torch_fl + +x = torch.randn(4, 4) # CPU tensor +x_flagos = x.to("flagos") # device 0 +x_flagos_1 = x.to("flagos:1") # device 1 + +y = torch.randn(4, 4, device="flagos") +y_cpu = y.cpu() # back to CPU +``` + +## Selecting a device + +`torch.flagos.device()` sets the current-device context, so a later `device="flagos"` tensor lands on it: + +```python +import torch +import torch_fl + +with torch.flagos.device(0): + x = torch.randn(4, 4, device="flagos") + +with torch.flagos.device(1): + y = torch.randn(4, 4, device="flagos") +``` + +## Synchronization + +Kernel launches are asynchronous, as on any other PyTorch device: + +```python +import torch +import torch_fl + +x = torch.randn(1000, 1000, device="flagos") +y = x @ x # enqueued, not necessarily finished +torch.flagos.synchronize() # wait for the device to drain +``` + +## Device queries + +```python +import torch +import torch_fl + +if torch.flagos.is_available(): + print(f"Found {torch.flagos.device_count()} device(s)") + print(f"Current device: {torch.flagos.current_device()}") + + props = torch.flagos.get_device_properties(0) + print(f"Device name: {props.name}") + print(f"Total memory: {props.total_memory / 1024**3:.2f} GB") +else: + print("No flagos devices available") +``` + +If `device_count()` reports 0 on a machine with a working driver, see {doc}`Troubleshooting <../reference/troubleshooting>`. + +## Operator routing + +Operations on `flagos` tensors are dispatched per operator, not per device or per model: + +- **Portable compiler kernels** — FlagGems Triton kernels, where the platform's route enables them +- **Native vendor kernels** — the vendor operator library (ACLNN, topsaten, mudnn, …) +- **Compatibility boxing** — generated kernels delegating to an external vendor `libtorch` (CUDA, MetaX, PPU, DCU) +- **CPU fallback** — a correctness-first CPU implementation for operators with no device kernel, copied back to the device + +Routing is transparent: no code change is needed when moving between platforms or kernel sources. To see which backend served a given call, set `FLAGOS_LOG=dispatch` (see the {doc}`environment variable reference <../reference/environment-variables>`). + +## Next steps + +- {doc}`Installation ` — per-platform build and verification +- {doc}`Platform capability matrix <../reference/platform-capability>` — what each accelerator supports +- {doc}`Dtype support <../reference/dtype-support>` — storage, AMP targets and fallback boundaries +- {doc}`Distributed collectives <../architecture/distributed>` — `ProcessGroupFlagOS`, DDP and FSDP2 +- {doc}`Profiler <../architecture/profiler>` — `torch.profiler` integration diff --git a/docs/torch_fl_en/index.md b/docs/torch_fl_en/index.md index 02b3da20ed..8ee0b1e0bd 100644 --- a/docs/torch_fl_en/index.md +++ b/docs/torch_fl_en/index.md @@ -1,102 +1,100 @@ -# Torch-FL - -Torch-FL (`torch_fl`) is a PyTorch device plugin for the FlagOS software stack. It exposes a single `flagos` device that routes operators among reusable native kernels, portable compiler kernels, vendor-native implementations, and explicit CPU fallback, so the same PyTorch program runs across accelerators without workload changes. - -![Torch-FL architecture](assets/images/torch-fl.png) - -::::{grid} 1 2 2 3 -:gutter: 1 1 1 2 - -:::{grid-item-card} {octicon}`browser;1.5em;sd-mr-1` Overview -:link: overview/overview -:link-type: doc - -What Torch-FL is, its design principles, capabilities, and component architecture. - -+++ -[Learn more »](overview/overview.md) -::: - -:::{grid-item-card} {octicon}`book;1.5em;sd-mr-1` Getting Started -:link: getting_started/installation -:link-type: doc - -Platform selection, build requirements, environment setup, verification, and test markers. - -+++ -[Learn more »](getting_started/installation.md) -::: - -:::{grid-item-card} {octicon}`broadcast;1.5em;sd-mr-1` Distributed -:link: architecture/distributed -:link-type: doc - -`ProcessGroupFlagOS`, FlagCX, vendor fallbacks, DDP and DataParallel support. - -+++ -[Learn more »](architecture/distributed.md) -::: - -:::{grid-item-card} {octicon}`pulse;1.5em;sd-mr-1` Profiler -:link: architecture/profiler -:link-type: doc - -`torch.profiler` integration, vendor tracers, correlation ids, and parity with `torch.cuda`. - -+++ -[Learn more »](architecture/profiler.md) -::: - -:::{grid-item-card} {octicon}`zap;1.5em;sd-mr-1` torch.compile -:link: architecture/torch-compile -:link-type: doc - -Inductor integration for the `flagos` device, FlagTree compilation, and platform notes. - -+++ -[Learn more »](architecture/torch-compile.md) -::: - -:::{grid-item-card} {octicon}`gear;1.5em;sd-mr-1` Reference -:link: reference/compatibility -:link-type: doc - -Per-platform capability validation, environment variables, and operation-routing configuration. - -+++ -[Learn more »](reference/compatibility.md) -::: - -:::: - -- **Repository**: [flagos-ai/Torch-FL](https://github.com/flagos-ai/Torch-FL) -- **FlagGems**: [flagos-ai/FlagGems](https://github.com/flagos-ai/FlagGems) -- **FlagTree**: [flagos-ai/FlagTree](https://github.com/flagos-ai/FlagTree) -- **FlagCX**: [flagos-ai/FlagCX](https://github.com/flagos-ai/FlagCX) -- **License**: Apache License 2.0 - ---- - -```{toctree} -:caption: 📑 Release Notes -:maxdepth: 5 -:hidden: - -release_notes/release-notes.md -``` - -```{toctree} -:caption: 📚 Guides -:maxdepth: 5 -:hidden: - -overview/overview.md -overview/features.md -overview/architecture.md -getting_started/installation.md -architecture/distributed.md -architecture/profiler.md -architecture/torch-compile.md -reference/compatibility.md -reference/environment-variables.md -``` +# Torch-FL + +Torch-FL (`torch_fl`) is a PyTorch device plugin for the FlagOS software stack. It exposes a single `flagos` device that routes operators among reusable native kernels, portable compiler kernels, vendor-native implementations, and explicit CPU fallback, so the same PyTorch program runs across accelerators without workload changes. + +![Torch-FL architecture](assets/images/torch-fl.png) + +::::{grid} 1 2 2 3 +:gutter: 1 1 1 2 + +:::{grid-item-card} {octicon}`browser;1.5em;sd-mr-1` Overview +:link: overview/overview +:link-type: doc + +What Torch-FL is, its design principles, capabilities, and component architecture. + ++++ +[Learn more »](overview/overview.md) +::: + +:::{grid-item-card} {octicon}`book;1.5em;sd-mr-1` Getting Started +:link: getting_started/installation +:link-type: doc + +Platform selection, build requirements, environment setup, verification, and test markers. + ++++ +[Learn more »](getting_started/installation.md) +::: + +:::{grid-item-card} {octicon}`broadcast;1.5em;sd-mr-1` Distributed +:link: architecture/distributed +:link-type: doc + +`ProcessGroupFlagOS`, FlagCX, vendor fallbacks, DDP and DataParallel support. + ++++ +[Learn more »](architecture/distributed.md) +::: + +:::{grid-item-card} {octicon}`pulse;1.5em;sd-mr-1` Profiler +:link: architecture/profiler +:link-type: doc + +`torch.profiler` integration, vendor tracers, correlation ids, and parity with `torch.cuda`. + ++++ +[Learn more »](architecture/profiler.md) +::: + +:::{grid-item-card} {octicon}`zap;1.5em;sd-mr-1` torch.compile +:link: architecture/torch-compile +:link-type: doc + +Inductor integration for the `flagos` device, FlagTree compilation, and platform notes. + ++++ +[Learn more »](architecture/torch-compile.md) +::: + +:::{grid-item-card} {octicon}`gear;1.5em;sd-mr-1` Reference +:link: reference/compatibility +:link-type: doc + +Per-platform capability validation, environment variables, and operation-routing configuration. + ++++ +[Learn more »](reference/compatibility.md) +::: + +:::: + +--- + +```{toctree} +:caption: 📑 Release Notes +:maxdepth: 5 +:hidden: + +release_notes/release-notes.md +``` + +```{toctree} +:caption: 📚 Guides +:maxdepth: 5 +:hidden: + +overview/overview.md +overview/features.md +overview/architecture.md +getting_started/installation.md +getting_started/quickstart.md +architecture/distributed.md +architecture/profiler.md +architecture/torch-compile.md +reference/compatibility.md +reference/environment-variables.md +reference/platform-capability.md +reference/dtype-support.md +reference/troubleshooting.md +``` diff --git a/docs/torch_fl_en/overview/architecture.md b/docs/torch_fl_en/overview/architecture.md index 98f85e52d6..8a20f594aa 100644 --- a/docs/torch_fl_en/overview/architecture.md +++ b/docs/torch_fl_en/overview/architecture.md @@ -1,76 +1,76 @@ -# Architecture - -Torch-FL registers one PyTorch device and routes every operator that reaches it. - -![Torch-FL architecture](../assets/images/torch-fl.png) - -```text -PyTorch API - | -flagos device (PrivateUse1) - | -device runtime + per-operator routing - | -FlagGems/compiler kernels | compatibility boxing | vendor-native kernels | CPU fallback - | -accelerator runtime -``` - -## Device registration - -At import, Torch-FL claims the `PrivateUse1` dispatch key and publishes it under the name `flagos`: - -1. `torch.utils.rename_privateuse1_backend("flagos")` names the key. -2. `torch._register_device_module("flagos", flagos)` installs the device module, so `torch.flagos.*` works. -3. `torch.utils.generate_methods_for_privateuse1_backend(for_storage=True)` generates tensor and storage methods such as `.to("flagos")`. -4. The device module is also published as `torch_flagos` in `sys.modules`, satisfying `torch::utils::device_lazy_init`'s lookup by module name on platforms that trigger lazy init. - -The native extension (`torch_fl._C`) registers the `AutogradPrivateUse1` fallback and the operator implementations when it is loaded. Because that registration happens at `dlopen` time, the plugin checks that `PrivateUse1` is still unclaimed *before* loading it, and fails with an actionable message instead of an uncatchable abort when another vendor plugin has already claimed the key. - -## Import-time phases - -`import torch_fl` runs a fixed sequence of side effects. The order is load-bearing — a wrong order produces a `dlopen` abort or a wrong-vendor build rather than a Python exception — so it lives in one place: - -| Phase | What it does | -|---|---| -| 1. conf | Selects the operator-routing table for this build; stages the MetaX `libcudart` shim when enabled | -| 2. preload | Selects and preloads the vendor libtorch and CUDA assets before `import torch` | -| 3. claim | Imports `torch`, verifies the vendor runtime, claims `PrivateUse1`, loads `_C`, installs the device module | -| 4. vendor_compat | Installs vendor runtime shims and resolves the FlagGems vendor (`GEMS_VENDOR`) | -| 5. ecosystem | FlagGems registration prep, CUDA alias, distributed/DDP/DataParallel, compile backends | - -Two constraints are worth knowing when debugging an import: the vendor libtorch must be in place before `import torch`, and the `PrivateUse1` ownership check must precede the `_C` load. - -## Component layout - -| Path | Responsibility | -|---|---| -| `torch_fl/__init__.py` | Import-time phase pipeline, device module installation, ecosystem patches | -| `torch_fl/_env.py` | Single registry and reader for every `FLAGOS_*` variable | -| `torch_fl/flagos/` | Device module: streams, events, RNG, AMP, memory, meta kernels for tracing | -| `torch_fl/configs/` | Per-platform routing tables, `backends_.conf` | -| `torch_fl/accelerator/` | Per-vendor compatibility shims and runtime glue | -| `torch_fl/compile/` | Inductor backend, FlagTree shims, platform profiles, Triton guards | -| `torch_fl/comm/` | `ProcessGroupFlagOS` | -| `torch_fl/distributed.py` | `init_process_group`, `DistributedDataParallel`, buffer movement helpers | -| `torch_fl/quantization/` | Low-precision formats, conversion and modules | -| `torch_fl/tileops/` | Python side of the TileOps operator library (lazy import) | -| `torch_fl/compat/` | Apex and flex-attention compatibility layers | -| `csrc/aten/` | ATen layer: dispatcher, boxing, generated bindings, vendor backends | -| `csrc/runtime/` | Device runtime: allocators, guard, generator, per-accelerator sources | -| `csrc/profiler/` | Vendor-agnostic device tracer interface and per-vendor tracers | -| `csrc/include/flagos.h` | Unified runtime ABI (memory, stream, device, current-stream registry) | -| `scripts/codegen/` | Operator-binding generators (CUDA boxing, ACLNN, topsaten, mudnn, TileOps) | -| `scripts/tools/` | Preflight/validation tooling, including `torch-fl-preflight` | -| `tests/unit`, `tests/integration`, `tests/manual`, `tests/perf` | Unit, hardware, manual and benchmark suites | - -## Operator dispatch - -Operator implementations are reached through one dispatch key and one routing table: - -- The generated bindings under `csrc/aten/generated/` provide the CUDA-boxing kernels, the FlagGems Python and C++ callers, and the TileOps stubs, all registered against `PrivateUse1`. -- `csrc/aten/common.cc` reads the routing table at first dispatch: the table Torch-FL selected for its build, or the file named by `FLAGOS_BACKEND_CONFIG`. -- Vendor-native kernels live under `csrc/aten/backends//` and are generated for the operator surface that vendor library actually exports; operations the generator cannot match are skipped with a warning and fall through to routing or CPU fallback. -- Unrouted operators reach `cpu_fallback`, which runs the CPU reference implementation and copies the result back to the device. - -Routing decisions are reported per operator with `FLAGOS_LOG=dispatch`, and the table actually in use is reported by `torch_fl.backend_config_path()`. +# Architecture + +Torch-FL registers one PyTorch device and routes every operator that reaches it. + +![Torch-FL architecture](../assets/images/torch-fl.png) + +```text +PyTorch API + | +flagos device (PrivateUse1) + | +device runtime + per-operator routing + | +FlagGems/compiler kernels | compatibility boxing | vendor-native kernels | CPU fallback + | +accelerator runtime +``` + +## Device registration + +At import, Torch-FL claims the `PrivateUse1` dispatch key and publishes it under the name `flagos`: + +1. `torch.utils.rename_privateuse1_backend("flagos")` names the key. +2. `torch._register_device_module("flagos", flagos)` installs the device module, so `torch.flagos.*` works. +3. `torch.utils.generate_methods_for_privateuse1_backend(for_storage=True)` generates tensor and storage methods such as `.to("flagos")`. +4. The device module is also published as `torch_flagos` in `sys.modules`, satisfying `torch::utils::device_lazy_init`'s lookup by module name on platforms that trigger lazy init. + +The native extension (`torch_fl._C`) registers the `AutogradPrivateUse1` fallback and the operator implementations when it is loaded. Because that registration happens at `dlopen` time, the plugin checks that `PrivateUse1` is still unclaimed *before* loading it, and fails with an actionable message instead of an uncatchable abort when another vendor plugin has already claimed the key. + +## Import-time phases + +`import torch_fl` runs a fixed sequence of side effects. The order is load-bearing — a wrong order produces a `dlopen` abort or a wrong-vendor build rather than a Python exception — so it lives in one place: + +| Phase | What it does | +|---|---| +| 1. conf | Selects the operator-routing table for this build; stages the MetaX `libcudart` shim when enabled | +| 2. preload | Selects and preloads the vendor libtorch and CUDA assets before `import torch` | +| 3. claim | Imports `torch`, verifies the vendor runtime, claims `PrivateUse1`, loads `_C`, installs the device module | +| 4. vendor_compat | Installs vendor runtime shims and resolves the FlagGems vendor (`GEMS_VENDOR`) | +| 5. ecosystem | FlagGems registration prep, CUDA alias, distributed/DDP/DataParallel, compile backends | + +Two constraints are worth knowing when debugging an import: the vendor libtorch must be in place before `import torch`, and the `PrivateUse1` ownership check must precede the `_C` load. + +## Component layout + +| Path | Responsibility | +|---|---| +| `torch_fl/__init__.py` | Import-time phase pipeline, device module installation, ecosystem patches | +| `torch_fl/_env.py` | Single registry and reader for every `FLAGOS_*` variable | +| `torch_fl/flagos/` | Device module: streams, events, RNG, AMP, memory, meta kernels for tracing | +| `torch_fl/configs/` | Per-platform routing tables, `backends_.conf` | +| `torch_fl/accelerator/` | Per-vendor compatibility shims and runtime glue | +| `torch_fl/compile/` | Inductor backend, FlagTree shims, platform profiles, Triton guards | +| `torch_fl/comm/` | `ProcessGroupFlagOS` | +| `torch_fl/distributed.py` | `init_process_group`, `DistributedDataParallel`, buffer movement helpers | +| `torch_fl/quantization/` | Low-precision formats, conversion and modules | +| `torch_fl/tileops/` | Python side of the TileOps operator library (lazy import) | +| `torch_fl/compat/` | Apex and flex-attention compatibility layers | +| `csrc/aten/` | ATen layer: dispatcher, boxing, generated bindings, vendor backends | +| `csrc/runtime/` | Device runtime: allocators, guard, generator, per-accelerator sources | +| `csrc/profiler/` | Vendor-agnostic device tracer interface and per-vendor tracers | +| `csrc/include/flagos.h` | Unified runtime ABI (memory, stream, device, current-stream registry) | +| `scripts/codegen/` | Operator-binding generators (CUDA boxing, ACLNN, topsaten, mudnn, TileOps) | +| `scripts/tools/` | Preflight/validation tooling, including `torch-fl-preflight` | +| `tests/unit`, `tests/integration`, `tests/manual`, `tests/perf` | Unit, hardware, manual and benchmark suites | + +## Operator dispatch + +Operator implementations are reached through one dispatch key and one routing table: + +- The generated bindings under `csrc/aten/generated/` provide the CUDA-boxing kernels, the FlagGems Python and C++ callers, and the TileOps stubs, all registered against `PrivateUse1`. +- `csrc/aten/common.cc` reads the routing table at first dispatch: the table Torch-FL selected for its build, or the file named by `FLAGOS_BACKEND_CONFIG`. +- Vendor-native kernels live under `csrc/aten/backends//` and are generated for the operator surface that vendor library actually exports; operations the generator cannot match are skipped with a warning and fall through to routing or CPU fallback. +- Unrouted operators reach `cpu_fallback`, which runs the CPU reference implementation and copies the result back to the device. + +Routing decisions are reported per operator with `FLAGOS_LOG=dispatch`, and the table actually in use is reported by `torch_fl.backend_config_path()`. diff --git a/docs/torch_fl_en/overview/features.md b/docs/torch_fl_en/overview/features.md index 4df8db3669..d5ea71fa2d 100644 --- a/docs/torch_fl_en/overview/features.md +++ b/docs/torch_fl_en/overview/features.md @@ -1,98 +1,98 @@ -# Features - -## One device, standard PyTorch APIs - -Every supported accelerator is programmed through the `flagos` device. Model code, optimizers, and third-party libraries keep using standard PyTorch APIs; moving between accelerators does not require changing tensor device strings, kernel launches, or the model. - -The device module is installed at import time with `torch.utils.rename_privateuse1_backend("flagos")` and `torch._register_device_module()`, so `device="flagos"`, `torch.flagos.*` methods, tensor methods, and storage work the same way as a first-class PyTorch device. - -## Per-operator backend routing - -Each built wheel ships one routing table, `torch_fl/configs/backends_.conf`, holding `op = backend` entries for the platform it was built for. Routing is decided per operator — not per device or per model — and can mix backends inside one model: - -| Route family | Serves the operator with | -|---|---| -| `flagos_python` | FlagGems Triton kernels through the Python dispatcher | -| FlagGems C++ (`kFlagOs`) | FlagGems kernels through the C++ runtime, `liboperators.so` | -| `cuda` | CUDA compatibility-boxing kernels over an external or vendor `libtorch_cuda.so` | -| Vendor native (`ascend`, `gcu`, `musa`, …) | The vendor operator library (ACLNN, topsaten, mudnn, …) | -| `tileops` | TileOps/TileLang kernels (SM90 NVIDIA parts) | -| `cpu_fallback` | A correctness-first CPU implementation, copied back to the device | - -Two runtime variables change routing without rebuilding: - -- `FLAGOS_BACKEND_CONFIG` — point the process at a different routing table. -- `FLAGOS_OP_` — override one operator, e.g. `FLAGOS_OP_add__Tensor=cuda`. -- `FLAGOS_FORCE_BACKEND` — repin every operator onto one backend family (`flaggems`, `vendor`, `tileops`) for A/B measurement. - -Current routing is always queryable: `torch_fl.backend_config_path()` reports the table in use. - -## Execution paths - -Torch-FL implements four operator execution strategies, and a platform may combine several of them. These are implementation strategies, not user-selectable product tiers. - -- **Native vendor kernels** — direct calls into the vendor runtime and operator libraries (ACLNN for Ascend, topsaten for Enflame GCU, mudnn for Moore Threads MUSA). The plugin generates bindings to each vendor's C/C++ API. -- **Compatibility boxing** — zero-copy metadata conversion into an independent PyTorch dispatch key when the vendor stack exposes one that can coexist with `PrivateUse1`. CUDA boxing reuses NVIDIA kernels through an external `libtorch_cuda.so`; MetaX, PPU and Hygon DCU box their vendor torch builds the same way. -- **Portable compiler kernels** — FlagGems kernels generated through Triton or a compatible compiler backend (FlagTree), targeting multiple accelerator families without per-platform rewrites. -- **Explicit CPU fallback** — operators without a device kernel run on the CPU where PyTorch semantics permit. Coverage is documented per platform rather than presented as complete native support. - -## Device runtime and management API - -The `torch.flagos` module exposes a complete device interface: - -- Streams (`Stream`, `current_stream`, `stream`) and events (`Event`) -- Device queries: `device_count`, `current_device`, `set_device`, `get_device_properties` -- Synchronization: `synchronize` -- RNG: `manual_seed`, `manual_seed_all`, `get_rng_state`, `set_rng_state`, `initial_seed` -- Autocast: `get_amp_supported_dtype` -- Memory: `memory_allocated`, `memory_reserved`, `memory_stats`, `reset_peak_memory_stats`, `empty_cache` - -Device memory uses the rental/caching allocator by default (`FLAGOS_USE_CACHING_ALLOCATOR`, on by default), with `FLAGOS_USE_CACHING_ALLOCATOR=0` handing every allocation straight to the vendor runtime. - -## Training stack - -- **Eager execution and autograd** — forward and backward paths run on the device; `AutogradPrivateUse1` fallbacks are registered by the native extension. -- **`torch.autocast("flagos")` and `torch.amp.GradScaler("flagos")`** — FP16 and BF16 targets, with the standard PyTorch autocast policy groups. -- **`torch.compile`** — Inductor integration registering `flagos` as a first-class GPU device; see {doc}`torch.compile integration <../architecture/torch-compile>`. -- **Distributed** — `ProcessGroupFlagOS`, DDP, DataParallel and FSDP2 support; see {doc}`Distributed collectives <../architecture/distributed>`. - -## Low-precision and quantization - -`torch_fl.quantization` provides low-precision conversion helpers and modules: - -- `convert`, `FormatSpec`, `LowPrecisionFormat`, `get_format_spec`, `normalize_format`, `supported_formats` -- `SoftLowpLinear` — a linear layer that decodes low-precision weights before the matrix multiply - -On CUDA-boxing vendors shared by DCU and MetaX, the software low-precision matrix path (`soft_lowp`) serves `mm`, `bmm` and `addmm` for scalar FP8 formats (`float8_e4m3fn`, `float8_e5m2`, `float8_e4m3fnuz`, `float8_e5m2fnuz`, `float8_e8m0fnu`) and packed FP4 (`float4_e2m1fn_x2`). Values are decoded and accumulated in BF16: an unspecified output dtype defaults to BF16, an explicit one is honored. Block-scaled metadata formats (MXFP4, NVFP4, block FP4) and `_scaled_mm` families are outside that path. - -## Framework and ecosystem compatibility - -Import-time shims adapt other projects' assumptions about the device, all of them optional and measurable through their own switch: - -- **Apex** — the common `MultiTensorApply` entry point is patched so Apex's `amp_C` kernels receive zero-copy CUDA views of flagos tensors. -- **`torch.nn.attention.flex_attention`** — the hard-coded `{cuda, cpu, xpu, hpu}` device gate is relaxed for the `flagos` device; the fused template is reachable through the `flagos` compile backend. -- **`diffusers` Qwen-Image rotary embedding** — the `flagos` device is registered on both halves of the per-device RoPE table, keeping the rotation off the complex-exponential path `diffusers` would otherwise take on GCU and Ascend. -- **CUDA alias** — with `FLAGOS_ALIAS_CUDA` on by default, a `cuda` device string is accepted as `flagos`, so unmodified CUDA scripts run against the `flagos` device. - -## Observability - -- **Dispatch and fallback logging** — `FLAGOS_LOG=dispatch,fallback` prints the backend chosen per operator and every CPU-fallback dispatch. -- **Profiler** — `torch.profiler` collects a device timeline with flow arrows, per-operator device time, kernel metadata and runtime event names; see {doc}`Profiler integration <../architecture/profiler>`. -- **Wheel compatibility manifest** — every wheel carries `torch_fl/compatibility.json` recording the platform, kernel sets, bundled libtorch, build-time PyTorch and ABI, and observed FlagTree/FlagGems/FlagCX versions. The `torch-fl-preflight` CLI inspects it before a native extension is imported. -- **Unknown-variable warning** — `import torch_fl` scans the environment once and warns about misspelled `FLAGOS_*` names instead of silently ignoring them. - -## Hardware support - -| Platform | Execution path | Validated capabilities | Status | -|---|---|---|---| -| NVIDIA CUDA | CUDA boxing over an external `libtorch_cuda.so` | Eager, autograd, distributed (FlagCX/NCCL), profiler (CUPTI), FlagGems (Python + C++) | Stable | -| MetaX | CUDA boxing via cu-bridge against the vendor libtorch | Eager, autograd, AMP, low-precision matrix ops | Stable | -| Ascend | Native ACLNN backend, FlagGems via FlagTree (Triton 3.5) | Eager, autograd, RNG suite, profiler (MSPTI) | Beta | -| PPU | CUDA boxing against the PPU CUDA-13-compatible SDK | Eager, autograd, AMP | Experimental | -| Hygon DCU | CUDA boxing over the hipified DTK torch | Eager, autograd, FP16/BF16 AMP, profiler | Beta | -| Enflame GCU | Native topsaten backend, CPU fallback for unrouted/int64/float64 ops | Eager, AMP | Beta | -| Moore Threads MUSA | FlagGems-first Triton kernels, native mudnn fallback, CPU fallback | Eager, FP16/BF16 AMP | Experimental | -| D-Robotics BPU | No eager kernels; `torch.compile` graph path via hbdk4 | Graph compilation only | Runtime only | -| TsingMicro | Runtime build selector exists | No per-operator kernel set documented | Runtime only | - -Per-capability detail — including `torch.compile`, distributed and profiler status per platform — is in the {doc}`compatibility matrix <../reference/compatibility>`. +# Features + +## One device, standard PyTorch APIs + +Every supported accelerator is programmed through the `flagos` device. Model code, optimizers, and third-party libraries keep using standard PyTorch APIs; moving between accelerators does not require changing tensor device strings, kernel launches, or the model. + +The device module is installed at import time with `torch.utils.rename_privateuse1_backend("flagos")` and `torch._register_device_module()`, so `device="flagos"`, `torch.flagos.*` methods, tensor methods, and storage work the same way as a first-class PyTorch device. + +## Per-operator backend routing + +Each built wheel ships one routing table, `torch_fl/configs/backends_.conf`, holding `op = backend` entries for the platform it was built for. Routing is decided per operator — not per device or per model — and can mix backends inside one model: + +| Route family | Serves the operator with | +|---|---| +| `flagos_python` | FlagGems Triton kernels through the Python dispatcher | +| FlagGems C++ (`kFlagOs`) | FlagGems kernels through the C++ runtime, `liboperators.so` | +| `cuda` | CUDA compatibility-boxing kernels over an external or vendor `libtorch_cuda.so` | +| Vendor native (`ascend`, `gcu`, `musa`, …) | The vendor operator library (ACLNN, topsaten, mudnn, …) | +| `tileops` | TileOps/TileLang kernels (SM90 NVIDIA parts) | +| `cpu_fallback` | A correctness-first CPU implementation, copied back to the device | + +Two runtime variables change routing without rebuilding: + +- `FLAGOS_BACKEND_CONFIG` — point the process at a different routing table. +- `FLAGOS_OP_` — override one operator, e.g. `FLAGOS_OP_add__Tensor=cuda`. +- `FLAGOS_FORCE_BACKEND` — repin every operator onto one backend family (`flaggems`, `vendor`, `tileops`) for A/B measurement. + +Current routing is always queryable: `torch_fl.backend_config_path()` reports the table in use. + +## Execution paths + +Torch-FL implements four operator execution strategies, and a platform may combine several of them. These are implementation strategies, not user-selectable product tiers. + +- **Native vendor kernels** — direct calls into the vendor runtime and operator libraries (ACLNN for Ascend, topsaten for Enflame GCU, mudnn for Moore Threads MUSA). The plugin generates bindings to each vendor's C/C++ API. +- **Compatibility boxing** — zero-copy metadata conversion into an independent PyTorch dispatch key when the vendor stack exposes one that can coexist with `PrivateUse1`. CUDA boxing reuses NVIDIA kernels through an external `libtorch_cuda.so`; MetaX, PPU and Hygon DCU box their vendor torch builds the same way. +- **Portable compiler kernels** — FlagGems kernels generated through Triton or a compatible compiler backend (FlagTree), targeting multiple accelerator families without per-platform rewrites. +- **Explicit CPU fallback** — operators without a device kernel run on the CPU where PyTorch semantics permit. Coverage is documented per platform rather than presented as complete native support. + +## Device runtime and management API + +The `torch.flagos` module exposes a complete device interface: + +- Streams (`Stream`, `current_stream`, `stream`) and events (`Event`) +- Device queries: `device_count`, `current_device`, `set_device`, `get_device_properties` +- Synchronization: `synchronize` +- RNG: `manual_seed`, `manual_seed_all`, `get_rng_state`, `set_rng_state`, `initial_seed` +- Autocast: `get_amp_supported_dtype` +- Memory: `memory_allocated`, `memory_reserved`, `memory_stats`, `reset_peak_memory_stats`, `empty_cache` + +Device memory uses the rental/caching allocator by default (`FLAGOS_USE_CACHING_ALLOCATOR`, on by default), with `FLAGOS_USE_CACHING_ALLOCATOR=0` handing every allocation straight to the vendor runtime. + +## Training stack + +- **Eager execution and autograd** — forward and backward paths run on the device; `AutogradPrivateUse1` fallbacks are registered by the native extension. +- **`torch.autocast("flagos")` and `torch.amp.GradScaler("flagos")`** — FP16 and BF16 targets, with the standard PyTorch autocast policy groups. +- **`torch.compile`** — Inductor integration registering `flagos` as a first-class GPU device; see {doc}`torch.compile integration <../architecture/torch-compile>`. +- **Distributed** — `ProcessGroupFlagOS`, DDP, DataParallel and FSDP2 support; see {doc}`Distributed collectives <../architecture/distributed>`. + +## Low-precision and quantization + +`torch_fl.quantization` provides low-precision conversion helpers and modules: + +- `convert`, `FormatSpec`, `LowPrecisionFormat`, `get_format_spec`, `normalize_format`, `supported_formats` +- `SoftLowpLinear` — a linear layer that decodes low-precision weights before the matrix multiply + +On CUDA-boxing vendors shared by DCU and MetaX, the software low-precision matrix path (`soft_lowp`) serves `mm`, `bmm` and `addmm` for scalar FP8 formats (`float8_e4m3fn`, `float8_e5m2`, `float8_e4m3fnuz`, `float8_e5m2fnuz`, `float8_e8m0fnu`) and packed FP4 (`float4_e2m1fn_x2`). Values are decoded and accumulated in BF16: an unspecified output dtype defaults to BF16, an explicit one is honored. Block-scaled metadata formats (MXFP4, NVFP4, block FP4) and `_scaled_mm` families are outside that path. + +## Framework and ecosystem compatibility + +Import-time shims adapt other projects' assumptions about the device, all of them optional and measurable through their own switch: + +- **Apex** — the common `MultiTensorApply` entry point is patched so Apex's `amp_C` kernels receive zero-copy CUDA views of flagos tensors. +- **`torch.nn.attention.flex_attention`** — the hard-coded `{cuda, cpu, xpu, hpu}` device gate is relaxed for the `flagos` device; the fused template is reachable through the `flagos` compile backend. +- **`diffusers` Qwen-Image rotary embedding** — the `flagos` device is registered on both halves of the per-device RoPE table, keeping the rotation off the complex-exponential path `diffusers` would otherwise take on GCU and Ascend. +- **CUDA alias** — with `FLAGOS_ALIAS_CUDA` on by default, a `cuda` device string is accepted as `flagos`, so unmodified CUDA scripts run against the `flagos` device. + +## Observability + +- **Dispatch and fallback logging** — `FLAGOS_LOG=dispatch,fallback` prints the backend chosen per operator and every CPU-fallback dispatch. +- **Profiler** — `torch.profiler` collects a device timeline with flow arrows, per-operator device time, kernel metadata and runtime event names; see {doc}`Profiler integration <../architecture/profiler>`. +- **Wheel compatibility manifest** — every wheel carries `torch_fl/compatibility.json` recording the platform, kernel sets, bundled libtorch, build-time PyTorch and ABI, and observed FlagTree/FlagGems/FlagCX versions. The `torch-fl-preflight` CLI inspects it before a native extension is imported. +- **Unknown-variable warning** — `import torch_fl` scans the environment once and warns about misspelled `FLAGOS_*` names instead of silently ignoring them. + +## Hardware support + +| Platform | Execution path | Validated capabilities | Status | +|---|---|---|---| +| NVIDIA CUDA | CUDA boxing over an external `libtorch_cuda.so` | Eager, autograd, distributed (FlagCX/NCCL), profiler (CUPTI), FlagGems (Python + C++) | Stable | +| MetaX | CUDA boxing via cu-bridge against the vendor libtorch | Eager, autograd, AMP, low-precision matrix ops | Stable | +| Ascend | Native ACLNN backend, FlagGems via FlagTree (Triton 3.5) | Eager, autograd, RNG suite, profiler (MSPTI) | Beta | +| PPU | CUDA boxing against the PPU CUDA-13-compatible SDK | Eager, autograd, AMP | Experimental | +| Hygon DCU | CUDA boxing over the hipified DTK torch | Eager, autograd, FP16/BF16 AMP, profiler | Beta | +| Enflame GCU | Native topsaten backend, CPU fallback for unrouted/int64/float64 ops | Eager, AMP | Beta | +| Moore Threads MUSA | FlagGems-first Triton kernels, native mudnn fallback, CPU fallback | Eager, FP16/BF16 AMP | Experimental | +| D-Robotics BPU | No eager kernels; `torch.compile` graph path via hbdk4 | Graph compilation only | Runtime only | +| TsingMicro | Runtime build selector exists | No per-operator kernel set documented | Runtime only | + +Per-capability detail — including `torch.compile`, distributed and profiler status per platform — is in the {doc}`compatibility matrix <../reference/compatibility>`. diff --git a/docs/torch_fl_en/overview/overview.md b/docs/torch_fl_en/overview/overview.md index 188d4313ab..a1a385d453 100644 --- a/docs/torch_fl_en/overview/overview.md +++ b/docs/torch_fl_en/overview/overview.md @@ -1,51 +1,51 @@ -# Torch-FL Overview - -`torch_fl` is a custom PyTorch device plugin built on the `PrivateUse1` extension mechanism. It registers [FlagGems](https://github.com/flagos-ai/FlagGems) high-performance Triton operators, vendor-native operator libraries, and CUDA compatibility kernels behind one device name: `flagos`. - -Accelerator vendors ship different runtimes, compiler stacks, and PyTorch integration strategies. Torch-FL hides those differences behind a unified runtime and operator-routing layer, so users program against standard PyTorch APIs and a single device name, and the plugin selects a kernel implementation per operator based on platform capability and configuration. - -## Design principles - -Torch-FL is built on five principles: - -1. **PyTorch-native interface** — Standard PyTorch APIs work unchanged; users target the `flagos` device instead of vendor-specific extensions. -2. **One logical device** — A single device name (`flagos`) abstracts vendor differences. Platform-specific routing happens transparently at the operator level. -3. **Layered operator backends** — Each operation may dispatch to a different implementation. Routing decisions are per operator, not per device or per model. -4. **Reuse before reimplementation** — Established kernels and compiler stacks are integrated where their dispatch and ABI boundaries permit, rather than rewriting functionality that already exists. -5. **Explicit capability boundaries** — Unsupported operations and CPU fallback paths are documented rather than presented as complete native coverage. Status levels distinguish validated support from experimental integrations. - -## Quick start - -```python -import torch -import torch_fl - -# Create a tensor on the flagos device -x = torch.randn(4, 4, device="flagos:0") - -# Operations route to platform-appropriate kernels -y = torch.relu(x @ x) - -# Move the result back to CPU -print(y.cpu()) -``` - -Operator routing (FlagGems compiler kernels, vendor-native kernels, compatibility boxing, or CPU fallback) is determined by platform detection and runtime configuration. The code above works unchanged across all supported accelerators. - -## Status levels - -| Status | Meaning | -|---|---| -| Stable | Critical paths are continuously tested and the supported version combination is documented. | -| Beta | The primary path is validated, but coverage, packaging, or release procedures are not yet stable. | -| Experimental | Validation exists for a specific setup, model, or hardware environment; interfaces or build procedures may change. | -| Runtime only | Device runtime support exists, but the platform is not a general eager operator backend. | - -A capability existing in the Torch-FL codebase does not imply that every platform implements or validates it. See the {doc}`compatibility matrix <../reference/compatibility>` for per-platform detail. - -```{toctree} -:maxdepth: 2 - -features.md -architecture.md -``` +# Torch-FL Overview + +`torch_fl` is a custom PyTorch device plugin built on the `PrivateUse1` extension mechanism. It registers [FlagGems](https://github.com/flagos-ai/FlagGems) high-performance Triton operators, vendor-native operator libraries, and CUDA compatibility kernels behind one device name: `flagos`. + +Accelerator vendors ship different runtimes, compiler stacks, and PyTorch integration strategies. Torch-FL hides those differences behind a unified runtime and operator-routing layer, so users program against standard PyTorch APIs and a single device name, and the plugin selects a kernel implementation per operator based on platform capability and configuration. + +## Design principles + +Torch-FL is built on five principles: + +1. **PyTorch-native interface** — Standard PyTorch APIs work unchanged; users target the `flagos` device instead of vendor-specific extensions. +2. **One logical device** — A single device name (`flagos`) abstracts vendor differences. Platform-specific routing happens transparently at the operator level. +3. **Layered operator backends** — Each operation may dispatch to a different implementation. Routing decisions are per operator, not per device or per model. +4. **Reuse before reimplementation** — Established kernels and compiler stacks are integrated where their dispatch and ABI boundaries permit, rather than rewriting functionality that already exists. +5. **Explicit capability boundaries** — Unsupported operations and CPU fallback paths are documented rather than presented as complete native coverage. Status levels distinguish validated support from experimental integrations. + +## Quick start + +```python +import torch +import torch_fl + +# Create a tensor on the flagos device +x = torch.randn(4, 4, device="flagos:0") + +# Operations route to platform-appropriate kernels +y = torch.relu(x @ x) + +# Move the result back to CPU +print(y.cpu()) +``` + +Operator routing (FlagGems compiler kernels, vendor-native kernels, compatibility boxing, or CPU fallback) is determined by platform detection and runtime configuration. The code above works unchanged across all supported accelerators. + +## Status levels + +| Status | Meaning | +|---|---| +| Stable | Critical paths are continuously tested and the supported version combination is documented. | +| Beta | The primary path is validated, but coverage, packaging, or release procedures are not yet stable. | +| Experimental | Validation exists for a specific setup, model, or hardware environment; interfaces or build procedures may change. | +| Runtime only | Device runtime support exists, but the platform is not a general eager operator backend. | + +A capability existing in the Torch-FL codebase does not imply that every platform implements or validates it. See the {doc}`compatibility matrix <../reference/compatibility>` for per-platform detail. + +```{toctree} +:maxdepth: 2 + +features.md +architecture.md +``` diff --git a/docs/torch_fl_en/reference/compatibility.md b/docs/torch_fl_en/reference/compatibility.md index 990bcc6181..f2534ca665 100644 --- a/docs/torch_fl_en/reference/compatibility.md +++ b/docs/torch_fl_en/reference/compatibility.md @@ -1,58 +1,59 @@ -# Compatibility and Platform Support - -## Status definitions - -| Status | Meaning | -|---|---| -| Stable | Critical paths are continuously tested and the supported version combination is documented. | -| Beta | The primary path is validated, but coverage, packaging, or release procedures are not yet stable. | -| Experimental | Validation exists for a specific setup, model, or hardware environment; interfaces or build procedures may change. | -| Runtime only | Device runtime support exists, but the platform is not a general eager operator backend. | - -## Project compatibility - -| Component | Supported range | Notes | -|---|---|---| -| Python | 3.8 or later | Platform SDKs and available wheels may impose a narrower range | -| PyTorch | 2.10.x (`>=2.10,<2.11`) | Generated ATen bindings are tied to this minor line | -| FlagGems | Platform dependent | Installed from PyPI or a vendor-compatible build only where the platform route uses it | -| Triton / compiler | Platform dependent | Use the compiler distribution required by the selected accelerator; FlagTree where the platform is built on it | - -### ATen minor-line pinning - -Torch-FL generates native bindings to PyTorch's internal ATen operator registry. Those bindings are sensitive to C++ ABI and operator schema changes, so the project pins to a PyTorch minor line — currently **2.10.x**. A different minor version (for example 2.11.x) produces build or runtime failures; patch releases inside the same line (2.10.0 → 2.10.1) are compatible. - -### Wheel compatibility record - -Every built wheel carries `torch_fl/compatibility.json`, recording the selected platform and kernel sets, the bundled libtorch location, the build-time PyTorch version and C++ ABI flag, the FlagTree/FlagGems/FlagCX versions observed at build time, the declared vendor PyTorch version where one supplies device libraries, and the wheel's Python requirements. An SDK version is recorded only when the builder sets `FLAGOS_SDK_VERSION` to a verified value; an absent value means *unknown*, not "compatible with all". - -```bash -python scripts/tools/torch-fl-preflight --wheel dist/torch_fl-*.whl --platform cuda \ - --sdk-version 13.3 --check-installed -``` - -`torch-fl-preflight` runs without importing `torch_fl`, so its checks do not trigger backend import side effects. `--check-installed` compares declared dependency ranges with the installed environment, `--check-build-env` requires exact build versions, and `--require-sdk` rejects wheels without a declared SDK. A release table can be generated from final wheel files with `--release --markdown-table`. - -## Platform matrix - -| Platform | Build selector | Execution path | Eager and autograd | torch.compile | Distributed | Profiler | FlagGems | Status | -|---|---|---|---|---|---|---|---|---| -| NVIDIA CUDA | `FLAGOS_ACCELERATOR=cuda` (default) | CUDA boxing over an external `libtorch_cuda.so` | Stable | Experimental (Inductor GPU device registered; no CI step) | Beta (FlagCX + NCCL fallback, DDP live-verified) | Stable (CUPTI parity) | Beta (Python + C++ dispatch paths) | Stable | -| MetaX | `FLAGOS_ACCELERATOR=metax` | CUDA boxing via `cu-bridge` against the vendor libtorch | Stable (FP16/BF16 autocast and GradScaler measured in boxing mode) | Experimental (vendor Triton and FlagTree measured on C550) | Experimental (NCCL-shaped `mccl` fallback; not CI-covered) | Experimental (MCPTI parity measured on C550; not CI-covered) | Experimental (Python dispatch; not CI-tested on MetaX) | Stable | -| Ascend | `FLAGOS_ACCELERATOR=ascend` | Native ACLNN backend, FlagGems via FlagTree (Triton 3.5) | Stable (CI-covered ops, RNG suite) | Experimental (Inductor measured on 910 with triton-ascend only, not revalidated on FlagTree; no CI step) | Experimental (HCCL fallback; architectural routing only) | Beta (MSPTI events plus device-time linkage, CI-covered by the shared contract; parity suite excluded) | Beta (Python dispatch; float64 and bool `neg` routes fall back to ACLNN) | Beta | -| PPU | `FLAGOS_ACCELERATOR=ppu` | CUDA boxing against the PPU CUDA-13-compatible SDK, bundling its own libtorch | Experimental (FP16/BF16 autocast and GradScaler measured on PPU hardware, not in CI) | Not validated | Experimental (NCCL fallback via a vendor-adapted `libnccl.so.2`; not CI-covered) | Not validated on this vendor's tracer | Experimental (vendor-index Triton required) | Experimental | -| Hygon DCU | `FLAGOS_ACCELERATOR=dcu` | CUDA boxing over the hipified DTK torch build | Beta (including FP16/BF16 autocast and GradScaler) | Experimental (FlagTree HCU validated on `gfx936`; not in CI) | Experimental (RCCL via DTK; `all_reduce`/DDP measured on 2 cards) | Beta (parity suite runs in CI) | Beta (Python dispatch only) | Beta | -| Enflame GCU | `FLAGOS_ACCELERATOR=gcu` | Native `libtopsaten.so` backend, CPU fallback for unrouted/int64/float64 ops | Beta (operator, RNG, factory and AMP suites CI-guarded on S60) | Not validated | Not validated | Runtime only (TOPSPTI activities, no device events on a CPU-only Kineto build) | Experimental (Python dispatch, requires vendor Triton) | Beta | -| Moore Threads MUSA | `FLAGOS_ACCELERATOR=musa` | Native `mudnn` backend, CPU fallback for unrouted ops | Experimental (FP16/BF16 autocast and GradScaler measured on MTT S5000) | Experimental (FlagTree forward/backward measured on MTT S5000; vendor runtime required) | Not validated | Experimental (MUPTI device timeline measured on MTT S5000) | Experimental (Python dispatch, requires vendor Triton) | Experimental | -| D-Robotics BPU | `FLAGOS_ACCELERATOR=bpu` | No eager kernel sets; eager ops run on CPU | Runtime only (CPU fallback for eager) | Experimental (`torch.compile(backend="bpu")` graph path via hbdk4) | Not applicable | Not validated | Not applicable (no per-operator kernel build) | Runtime only | -| TsingMicro | `FLAGOS_ACCELERATOR=tsingmicro` | Runtime/build selector present; no per-operator kernel set documented | Runtime only | Not validated | Not validated | Not validated | Not applicable | Runtime only | - -## Reading the matrix - -- **Eager and autograd** is the primary operator path; a Stable rating means the platform's critical paths are continuously exercised. -- **torch.compile** is experimental on most platforms: it is validated on specific hardware, and several platforms have no CI step for it. See {doc}`torch.compile integration <../architecture/torch-compile>`. -- **Distributed** ratings reflect measured collective and DDP coverage rather than the presence of code; see {doc}`Distributed collectives <../architecture/distributed>`. -- **Profiler** ratings describe which parts of the `torch.profiler` contract a platform satisfies; see {doc}`Profiler integration <../architecture/profiler>`. -- **FlagGems** on a platform means the portable Triton kernel route is available there; it is measured separately for availability and correctness. - -Capability ratings here describe the platform, not one build. A wheel built with a subset of kernel sets (`FLAGOS_BUILD_*`) supports a subset of what the platform can do — the wheel's own record is authoritative for that wheel. +# Compatibility and Platform Support + +## Status definitions + +| Status | Meaning | +|---|---| +| Stable | Critical paths are continuously tested and the supported version combination is documented. | +| Beta | The primary path is validated, but coverage, packaging, or release procedures are not yet stable. | +| Experimental | Validation exists for a specific setup, model, or hardware environment; interfaces or build procedures may change. | +| Runtime only | Device runtime support exists, but the platform is not a general eager operator backend. | + +## Project compatibility + +| Component | Supported range | Notes | +|---|---|---| +| Python | One version per platform | Fixed in `setup.py`: a FlagTree build exists for exactly one cp tag and the wheel links it, so the interpreter is single-valued — 3.12 on CUDA/GCU/MetaX/PPU, 3.10 on DCU/MUSA, 3.11 on Ascend | +| PyTorch | 2.10.x (`>=2.10,<2.11`) | Generated ATen bindings are tied to this minor line | +| FlagGems | Exact pin (5.4.0) | Declared as an exact requirement, not a range: the per-op routing tables were generated against one cohort | +| FlagTree | Exact pin, per platform | Declared as an exact requirement and it *is* Triton — the package name carries the vendor's backend (e.g. `0.7.0+hcu3.6`, the trailing number being the Triton line) | +| FlagCX | Exact pin, per platform | Declared only where the vendor runtime has a build; PPU has none yet, and elsewhere the distributed path falls back to the NCCL-shaped route | + +### ATen minor-line pinning + +Torch-FL generates native bindings to PyTorch's internal ATen operator registry. Those bindings are sensitive to C++ ABI and operator schema changes, so the project pins to a PyTorch minor line — currently **2.10.x**. A different minor version (for example 2.11.x) produces build or runtime failures; patch releases inside the same line (2.10.0 → 2.10.1) are compatible. + +### Wheel compatibility record + +Every built wheel carries `torch_fl/compatibility.json`, recording the selected platform and kernel sets, the bundled libtorch location, the build-time PyTorch version and C++ ABI flag, the FlagTree/FlagGems/FlagCX versions observed at build time, the declared vendor PyTorch version where one supplies device libraries, and the wheel's Python requirements. An SDK version is recorded only when the builder sets `FLAGOS_SDK_VERSION` to a verified value; an absent value means *unknown*, not "compatible with all". + +```bash +python scripts/tools/torch-fl-preflight --wheel dist/torch_fl-*.whl --platform cuda \ + --sdk-version 13.3 --check-installed +``` + +`torch-fl-preflight` runs without importing `torch_fl`, so its checks do not trigger backend import side effects. `--check-installed` compares declared dependency ranges with the installed environment, `--check-build-env` requires exact build versions, and `--require-sdk` rejects wheels without a declared SDK. A release table can be generated from final wheel files with `--release --markdown-table`. + +## Platform matrix + +| Platform | Build selector | Execution path | Eager and autograd | torch.compile | Distributed | Profiler | FlagGems | Status | +|---|---|---|---|---|---|---|---|---| +| NVIDIA CUDA | `FLAGOS_ACCELERATOR=cuda` (default) | CUDA boxing over an external `libtorch_cuda.so` | Stable | Experimental (Inductor GPU device registered; validated by the integration test only) | Beta (FlagCX + NCCL fallback, DDP live-verified) | Stable (CUPTI parity) | Beta (Python + C++ dispatch paths) | Stable | +| MetaX | `FLAGOS_ACCELERATOR=metax` | CUDA boxing via `cu-bridge` against the vendor libtorch | Stable (FP16/BF16 autocast and GradScaler measured in boxing mode) | Experimental (vendor Triton and FlagTree measured on C550) | Experimental (NCCL-shaped `mccl` fallback; not continuously validated) | Experimental (MCPTI parity measured on C550; not continuously validated) | Experimental (Python dispatch; not validated on MetaX) | Stable | +| Ascend | `FLAGOS_ACCELERATOR=ascend` | Native ACLNN backend, FlagGems via FlagTree (Triton 3.5) | Stable (operator and RNG suites validated) | Experimental (Inductor measured on 910 with triton-ascend only, not revalidated on FlagTree) | Experimental (HCCL fallback; architectural routing only) | Beta (MSPTI events plus device-time linkage, covered by the shared contract; parity suite not included) | Beta (Python dispatch; float64 and bool `neg` routes fall back to ACLNN) | Beta | +| PPU | `FLAGOS_ACCELERATOR=ppu` | CUDA boxing against the PPU CUDA-13-compatible SDK, bundling its own libtorch | Experimental (FP16/BF16 autocast and GradScaler measured on PPU hardware) | Not validated | Experimental (NCCL fallback via a vendor-adapted `libnccl.so.2`; not continuously validated) | Not validated on this vendor's tracer | Experimental (vendor-index Triton required) | Experimental | +| Hygon DCU | `FLAGOS_ACCELERATOR=dcu` | CUDA boxing over the hipified DTK torch build | Beta (including FP16/BF16 autocast and GradScaler) | Experimental (FlagTree HCU validated on `gfx936`) | Experimental (RCCL via DTK; architectural routing, not continuously validated) | Beta (parity suite included) | Beta (Python dispatch only) | Beta | +| Enflame GCU | `FLAGOS_ACCELERATOR=gcu` | Native `libtopsaten.so` backend, CPU fallback for unrouted/int64/float64 ops | Beta (operator, RNG, factory and AMP suites exercised) | Not validated | Not validated | Runtime only (TOPSPTI activities, no device events on a CPU-only Kineto build) | Experimental (Python dispatch, requires vendor Triton) | Beta | +| Moore Threads MUSA | `FLAGOS_ACCELERATOR=musa` | Native `mudnn` backend, CPU fallback for unrouted ops | Experimental (FP16/BF16 autocast and GradScaler measured on MTT S5000) | Experimental (FlagTree forward/backward measured on MTT S5000; vendor runtime required) | Not validated | Experimental (MUPTI device timeline measured on MTT S5000) | Experimental (Python dispatch, requires vendor Triton) | Experimental | +| D-Robotics BPU | `FLAGOS_ACCELERATOR=bpu` | No eager kernel sets; eager ops run on CPU | Runtime only (CPU fallback for eager) | Experimental (`torch.compile(backend="bpu")` graph path via hbdk4) | Not applicable | Not validated | Not applicable (no per-operator kernel build) | Runtime only | +| TsingMicro | `FLAGOS_ACCELERATOR=tsingmicro` | Runtime/build selector present; no per-operator kernel set documented | Runtime only | Not validated | Not validated | Not validated | Not applicable | Runtime only | + +## Reading the matrix + +- **Eager and autograd** is the primary operator path; a Stable rating means the platform's critical paths are continuously exercised. +- **torch.compile** is experimental on most platforms: it is validated on specific hardware, and on several platforms it is exercised only by the integration test. See {doc}`torch.compile integration <../architecture/torch-compile>`. +- **Distributed** ratings reflect measured collective and DDP coverage rather than the presence of code; see {doc}`Distributed collectives <../architecture/distributed>`. +- **Profiler** ratings describe which parts of the `torch.profiler` contract a platform satisfies; see {doc}`Profiler integration <../architecture/profiler>`. +- **FlagGems** on a platform means the portable Triton kernel route is available there; it is measured separately for availability and correctness. + +The {doc}`platform capability matrix ` records what each accelerator builds and routes by default, and the {doc}`dtype support ` page covers per-platform dtype and AMP boundaries. Capability ratings here describe the platform, not one build. A wheel built with a subset of kernel sets (`FLAGOS_BUILD_*`) supports a subset of what the platform can do — the wheel's own record is authoritative for that wheel. diff --git a/docs/torch_fl_en/reference/dtype-support.md b/docs/torch_fl_en/reference/dtype-support.md new file mode 100644 index 0000000000..bfd34108f3 --- /dev/null +++ b/docs/torch_fl_en/reference/dtype-support.md @@ -0,0 +1,40 @@ +# Dtype support + +Torch-FL preserves the requested dtype for tensor **storage** and follows PyTorch's promotion rules for tensor-tensor operations. Compute coverage is bounded by the vendor library each backend uses, and AMP target support is a separate question from eager dtype support: a dtype the storage layer accepts is not automatically accepted by every operator. + +## Storage and compute + +| Backend | Storage and copies | Eager elementwise | Matmul family | +|---|---|---|---| +| NVIDIA CUDA | Native PyTorch CUDA dtype support | Native CUDA coverage | Native CUDA coverage | +| MetaX (boxing) | MACA libtorch CUDA-compatible coverage; FP8 and packed FP4 storage | MACA libtorch CUDA-compatible coverage | MACA coverage, plus software-emulated FP8 / packed FP4 `mm`/`bmm`/`addmm` | +| Hygon DCU | Vendor library coverage | Vendor library coverage | Vendor library coverage | +| Ascend | float16, bfloat16, float32, float64, integer, uint8, bool | Vendor coverage; unsupported ACLNN combinations use the CPU fallback | float16, bfloat16, float32 natively; float64 and unsupported types use the CPU fallback | +| Enflame GCU | float16, bfloat16, float32, float64, integer, bool | topsaten coverage; int64, float64 and unrouted operations use the CPU fallback | float16, bfloat16, float32 | +| Moore Threads MUSA | float16, bfloat16, float32, float64, integer, bool | mudnn coverage; unrouted operations use the CPU fallback | float16, bfloat16, float32 | + +## AMP targets + +`torch.autocast("flagos")` supports `torch.float16` and `torch.bfloat16` as lower-precision targets, using the standard PyTorch autocast policy groups: matmul and convolution prefer the selected lower-precision dtype, numerically sensitive operations (logarithm, normalization) use float32, and mixed inputs follow the promote policy. **Float32 and float64 are not valid autocast targets.** + +| Backend | AMP targets | Notes | +|---|---|---| +| NVIDIA CUDA | float16, bfloat16 | | +| MetaX (boxing) | float16, bfloat16 | Measured in CUDA-boxing mode; the legacy handwritten kernel mode is not covered | +| Hygon DCU | Backend-dependent | | +| Ascend | float16, bfloat16 | float64 is storable and works elementwise, but is neither an AMP target nor accepted by the native matmul API | +| Enflame GCU | float16, bfloat16 | | +| Moore Threads MUSA | float16, bfloat16 | | + +## Boundaries + +Where a vendor operator does not accept a dtype, the operator is served by the correctness-first CPU fallback, which computes on the host and copies the correctly typed result back to the device. The fallback is correctness-oriented and may be slower than a native kernel — it is a documented coverage boundary, not a failure. + +- **Ascend** — `aclnnMatmul`, `aclnnMm` and `aclnnBatchMatMul` reject float64 and integer inputs; those go through the fallback. `aclnnNeg` rejects int16, uint8 and bool. float64 storage, device copies, casts and elementwise operations stay float64 end to end. Complex and quantized dtypes are outside the supported contract. +- **Enflame GCU** — topsaten has no float64 and no int64 kernels, so both are stored natively but computed on the CPU. `topsatenNeg` additionally rejects uint8 and bool. Convolution follows the native topsaten path forward and the CPU fallback backward. +- **Moore Threads MUSA** — `GradScaler`'s unscale operation uses the correctness-oriented fallback (list and scalar operands are moved to CPU, the reference kernel runs, and the mutated values plus `found_inf` are copied back) rather than a native foreach kernel. +- **MetaX** — the low-precision matrix path is software emulation of *scalar* FP8 (`float8_e4m3fn`, `float8_e5m2`, `float8_e4m3fnuz`, `float8_e5m2fnuz`, `float8_e8m0fnu`) and packed FP4 (`float4_e2m1fn_x2`) for `mm`, `bmm` and `addmm`, including the `dtype`/`out` variants. Values decode and accumulate in BF16, so an unspecified output dtype defaults to BF16 and an explicit one is honored. Block-scaled metadata formats (MXFP4, NVFP4, block FP4) and the `_scaled_mm` families are **not** included. + +## What a passing test means + +The integration suites compare results against CPU references where a native kernel is unavailable, so a passing test means the operation follows the documented PyTorch contract — not necessarily that it used a native vendor kernel. diff --git a/docs/torch_fl_en/reference/environment-variables.md b/docs/torch_fl_en/reference/environment-variables.md index deff42c925..d9e34ae70b 100644 --- a/docs/torch_fl_en/reference/environment-variables.md +++ b/docs/torch_fl_en/reference/environment-variables.md @@ -1,97 +1,107 @@ -# Environment Variables - -Torch-FL has one namespace of its own, `FLAGOS_*`, plus a second set belonging to other projects (torch, FlagGems, FlagCX, TileLang, vendor SDKs) that it reads but does not own. This page documents the variables a user configures; the authoritative, complete list is the `VARIABLES` registry in `torch_fl/_env.py`, checked against the upstream `docs/reference/environment-variables.md` by a unit test. - -Nothing here is required to run a wheel: a wheel routes, compiles and runs with an empty environment. These variables select a different build, override a setting for measurement, or turn on a diagnostic. - -## How a value is read - -- **Booleans**: `1`/`true`/`on`/`yes` (any case) are on; `0`/`false`/`off`/`no` are off. Anything else is not a boolean — a warning is printed once and the variable's default is used instead of treating the value as truthy. -- **Empty means unset**: `FLAGOS_LOG=${EXTRA_LOG}` with `EXTRA_LOG` unset behaves as if `FLAGOS_LOG` were never exported, so a switch that defaults on stays on. -- **Enums**: a switch naming a mode rather than a boolean reports the alternatives and uses the default when given a value outside them. -- **Unknown names**: `import torch_fl` scans the environment once and warns about any `FLAGOS_*` name that is neither declared nor part of the dynamic `FLAGOS_OP_` family — a misspelled switch would otherwise be read by nobody and silently do nothing. - -## Build selection - -Inputs to `setup.py` and the CMake build; nothing at runtime reads them. The wheel records what it was built with, and that record is what runtime readers consult. - -| Variable | Default | Purpose | -|---|---|---| -| `FLAGOS_ACCELERATOR` | `cuda` | Platform the wheel is built for: `cuda`, `ppu`, `metax`, `ascend`, `tsingmicro`, `dcu`, `gcu`, `musa`, `bpu` | -| `FLAGOS_BUILD_VENDOR` | `ON`, `OFF` on `metax` | Compile the vendor's native kernels (no-op where the vendor ships none) | -| `FLAGOS_BUILD_FLAGGEMS` | `ON`, `OFF` on `bpu` | Compile the FlagGems Python kernel wrappers | -| `FLAGOS_BUILD_FLAGGEMS_CPP` | `ON` on `cuda`, `tsingmicro` | Compile the FlagGems C++ wrapper (`liboperators.so`) | -| `FLAGOS_BUILD_BOXING` | `ON`, `OFF` on `ascend`, `gcu`, `musa` | Compile the generated CUDA-boxing kernels | -| `FLAGOS_BUILD_TILEOPS` | `ON` on `cuda` | Compile the TileOps kernel wrappers (TileLang, SM90 NVIDIA) | -| `FLAGOS_BUILD_JOBS` | CPU count | Parallel jobs for the CMake build | -| `FLAGOS_WHEEL_LOCAL` | SDK-derived | Local version label, e.g. `metax3.8.1` | -| `FLAGOS_SKIP_CUDA_ASSETS` | `0` | Do not bundle an external `libtorch_cuda.so` | -| `FLAGOS_CUDA_ASSETS_DIR` | `.libtorch_cuda_assets` | Directory the external `libtorch_cuda.so` is copied from | -| `FLAGOS_DCU_VENDOR_CORE` | `0` | Use DTK's forked core libraries instead of the official PyTorch core (must match at build and import time) | - -An explicit value that contradicts a per-platform forced value is rejected with an error naming both, rather than letting whichever flag CMake saw last win. - -## Operator routing - -Which backend implementation each operator dispatches to. Routing is stated per operator in a generated table, `torch_fl/configs/backends_.conf`; the variables below override or widen it. - -| Variable | Default | Purpose | -|---|---|---| -| `FLAGOS_BACKEND_CONFIG` | none | Absolute path to a `backends_*.conf` file; overrides the table selected from the build record. For testing and debugging | -| `FLAGOS_OP_` | none | Per-operator override, e.g. `FLAGOS_OP_add__Tensor=cuda` (replace `.` with `__`) | -| `FLAGOS_FORCE_BACKEND` | none | Repin every operator onto one backend family (`flaggems`, `vendor`, `tileops`) for A/B measurement | -| `FLAGOS_DISABLE_FLAGGEMS_PY` | `0` | Leave the FlagGems Python layer unregistered (C++ stub-only mode) | - -`torch_fl.backend_config_path()` reports the table in use; `FLAGOS_BACKEND_CONFIG` holds only what the user exported, so reading it answers "did I override the table?". - -## Runtime diagnostics - -| Variable | Default | Purpose | -|---|---|---| -| `FLAGOS_LOG` | none | Comma-separated stderr diagnostics: `dispatch` (backend chosen per operator), `fallback` (each CPU-fallback dispatch), `op_cache` (Ascend operator-cache statistics) | -| `FLAGOS_TRACE` | `0` | Verbose logging in the device profiler shim | -| `FLAGOS_TRACER_LIBRARY` | auto-discovered | Override the tracer library the profiler shim loads | - -## Distributed - -| Variable | Default | Purpose | -|---|---|---| -| `FLAGOS_DIST_REDIRECT_GLOO` | `1` | Answer a plain `init_process_group(backend="gloo")` or `new_group` request with the flagos backend when the process accelerator is the flagos device | -| `FLAGOS_DIST_STAGED_GLOO` | `1` | Allow the host-staged gloo inner backend, the last fallback tier when no vendor communicator is available. Set `0` to fail loudly instead of staging | -| `FLAGOS_DIST_FORCE_NCCL` | `0` | In the manual MetaX distributed tests, skip FlagCX and use NCCL | - -## Vendor and framework compatibility - -Import-time shims that adapt a vendor runtime or another framework to the `flagos` device. None is needed on a stock CUDA box. - -| Variable | Default | Purpose | -|---|---|---| -| `FLAGOS_ALIAS_CUDA` | `1` | Alias the `cuda` device string to `flagos` for drop-in compatibility. Set `0` to opt out | -| `FLAGOS_DISABLE_CUDA_SHIM` | `0` | Skip registering the `torch.cuda` compatibility shim for generic GPU operations | -| `FLAGOS_METAX_CUDART_SHIM` | `0` | Preload the `libcudart` version-tag shim before `import torch`; needed on MetaX with generic PyTorch wheels | -| `FLAGOS_METAX_COMPAT` | `0` | Patch FlagGems `torch.cuda` device queries for MetaX compatibility | -| `FLAGOS_DCU_HIP_VERSION` | none | Override HIP version detection for the DCU runtime | -| `FLAGOS_DCU_SKIP_RUNTIME_CHECK` | `0` | Skip the DCU post-import checks, for deliberately testing a non-matching wheel pair | -| `FLAGOS_DCU_SDPA_FLASH` | `1` | On DCU, point DTK's SDPA selector at its CUTLASS flash adapter instead of the math decomposition | -| `FLAGOS_DISABLE_APEX_COMPAT` | `0` | Disable the optional Apex multi-tensor compatibility layer | -| `FLAGOS_DISABLE_QWENIMAGE_ROPE` | `0` | Leave `diffusers`' Qwen-Image rotary-embedding table alone, to measure the difference | -| `FLAGOS_DISABLE_FLEX_ATTENTION_COMPAT` | `0` | Leave flex-attention's hard-coded `{cuda, cpu, xpu, hpu}` device gate in place | - -## Assets, libraries and compilation - -| Variable | Default | Purpose | -|---|---|---| -| `FLAGOS_DISABLE_CUDA_ASSETS` | `0` | Skip preloading the bundled `libtorch_cuda.so` and CUDA libraries (for in-tree builds and for PPU) | -| `FLAGOS_VENDOR_TORCH_LIB` | auto-discovered | Path to the vendor torch's `lib` directory when no bundled vendor libtorch is present | -| `FLAGOS_USE_CACHING_ALLOCATOR` | `1` | Caching device allocator; set `0` to hand every allocation to the vendor runtime | -| `FLAGOS_USE_FLAGTREE` | `0` | Assert that a FlagTree build is the active Triton (required on Ascend when the compiler is FlagTree) | -| `FLAGOS_COMPILE_FALLBACK_EAGER` | `0` | Fall back to eager mode when `torch.compile` meets unsupported operations | -| `FLAGOS_TILEOPS_USE_L2` | `0` | Use the TileOps L2-cache tier | -| `FLAGOS_TILEOPS_CACHE_MAX` | `512` | TileOps instance-cache capacity | -| `FLAGOS_TILEOPS_DISABLE_ALL_CACHE` | `0` | Neutralize every TileLang cache (correct but slow; set before `tileops` is imported) | - -## BPU compiler - -The BPU graph path compiles through `hbdk4` on an x86 host. `FLAGOS_BPU_MARCH` selects the micro-architecture (`nash-p`, `nash-e`, `nash-m`), `FLAGOS_BPU_QUANTIZE` (on by default) keeps convolution on the device by inserting int8 quantization, `FLAGOS_BPU_CACHE` sets the compiler cache directory, and `FLAGOS_BPU_X86_PYTHON` / `FLAGOS_BPU_X86_EMULATOR` / `FLAGOS_BPU_X86_STUBS` describe the x86 host and emulator used for on-board compilation. - -The complete variable list — including the interoperability names owned by other projects, the codegen-only inputs, and retired names kept inert — is maintained in the upstream repository at `docs/reference/environment-variables.md`, alongside the registry in `torch_fl/_env.py` that the documentation is checked against. +# Environment Variables + +Torch-FL has one namespace of its own, `FLAGOS_*`, plus a second set belonging to other projects (torch, FlagGems, FlagCX, TileLang, vendor SDKs) that it reads but does not own. This page documents the variables a user configures; the authoritative, complete list is the `VARIABLES` registry in `torch_fl/_env.py`, checked against the upstream `docs/reference/environment-variables.md` by a unit test. + +Nothing here is required to run a wheel: a wheel routes, compiles and runs with an empty environment. These variables select a different build, override a setting for measurement, or turn on a diagnostic. + +## How a value is read + +- **Booleans**: `1`/`true`/`on`/`yes` (any case) are on; `0`/`false`/`off`/`no` are off. Anything else is not a boolean — a warning is printed once and the variable's default is used instead of treating the value as truthy. +- **Empty means unset**: `FLAGOS_LOG=${EXTRA_LOG}` with `EXTRA_LOG` unset behaves as if `FLAGOS_LOG` were never exported, so a switch that defaults on stays on. +- **Enums**: a switch naming a mode rather than a boolean reports the alternatives and uses the default when given a value outside them. +- **Unknown names**: `import torch_fl` scans the environment once and warns about any `FLAGOS_*` name that is neither declared nor part of the dynamic `FLAGOS_OP_` family — a misspelled switch would otherwise be read by nobody and silently do nothing. Names retired by an earlier release are deliberately excluded from that warning: an old export is inert (no alias and no deprecation window) rather than silently honoured. + +## Build selection + +Inputs to `setup.py` and the CMake build; nothing at runtime reads them. The wheel records what it was built with, and that record is what runtime readers consult. + +| Variable | Default | Purpose | +|---|---|---| +| `FLAGOS_ACCELERATOR` | `cuda` | Platform the wheel is built for: `cuda`, `ppu`, `metax`, `ascend`, `tsingmicro`, `dcu`, `gcu`, `musa`, `bpu` | +| `FLAGOS_BUILD_VENDOR` | `ON`, `OFF` on `metax` | Compile the vendor's native kernels (no-op where the vendor ships none) | +| `FLAGOS_BUILD_FLAGGEMS` | `ON`, `OFF` on `bpu` | Compile the FlagGems Python kernel wrappers | +| `FLAGOS_BUILD_FLAGGEMS_CPP` | `ON` on `cuda`, `tsingmicro` | Compile the FlagGems C++ wrapper (`liboperators.so`) | +| `FLAGOS_BUILD_BOXING` | `ON`, `OFF` on `ascend`, `gcu`, `musa` | Compile the generated CUDA-boxing kernels | +| `FLAGOS_BUILD_TILEOPS` | `ON` on `cuda` | Compile the TileOps kernel wrappers (TileLang, SM90 NVIDIA) | +| `FLAGOS_BUILD_JOBS` | CPU count | Parallel jobs for the CMake build; `MAX_JOBS` and `CMAKE_BUILD_PARALLEL_LEVEL` are honoured as lower-priority fallbacks | +| `FLAGOS_WHEEL_LOCAL` | SDK-derived | Local version label, e.g. `maca3.8.1.3` | +| `FLAGOS_SKIP_CUDA_ASSETS` | `0` | Do not bundle an external `libtorch_cuda.so` | +| `FLAGOS_CUDA_ASSETS_DIR` | `.libtorch_cuda_assets` | Directory the external `libtorch_cuda.so` is copied from | +| `FLAGOS_DCU_VENDOR_CORE` | `0` | Use DTK's forked core libraries instead of the official PyTorch core (must match at build and import time) | + +An explicit value that contradicts a per-platform forced value is rejected with an error naming both, rather than letting whichever flag CMake saw last win. + +## Operator routing + +Which backend implementation each operator dispatches to. Routing is stated per operator in a generated table, `torch_fl/configs/backends_.conf`; the variables below override or widen it. + +| Variable | Default | Purpose | +|---|---|---| +| `FLAGOS_BACKEND_CONFIG` | none | Absolute path to a `backends_*.conf` file; overrides the table selected from the build record. For testing and debugging | +| `FLAGOS_OP_` | none | Per-operator override, e.g. `FLAGOS_OP_add__Tensor=cuda` (replace `.` with `__`) | +| `FLAGOS_FORCE_BACKEND` | none | Repin every operator onto one backend family (`flaggems`, `vendor`, `tileops`) for A/B measurement | +| `FLAGOS_DISABLE_FLAGGEMS_PY` | `0` | Leave the FlagGems Python layer unregistered (C++ stub-only mode) | +| `FLAGOS_STARTUP_PROFILE` | `full` | `full` runs framework compatibility hooks during import; `minimal` leaves them for explicit activation. FlagTree, FlagGems and FlagCX stay required in both | + +`torch_fl.backend_config_path()` reports the table in use; `FLAGOS_BACKEND_CONFIG` holds only what the user exported, so reading it answers "did I override the table?". + +## Runtime diagnostics + +| Variable | Default | Purpose | +|---|---|---| +| `FLAGOS_LOG` | none | Comma-separated stderr diagnostics: `dispatch` (backend chosen per operator), `fallback` (each CPU-fallback dispatch), `op_cache` (Ascend operator-cache statistics) | +| `FLAGOS_TRACE` | `0` | Verbose logging in the device profiler shim | +| `FLAGOS_TRACER_LIBRARY` | auto-discovered | Override the tracer library the profiler shim loads | + +## Distributed + +| Variable | Default | Purpose | +|---|---|---| +| `FLAGOS_DIST_REDIRECT_GLOO` | `1` | Answer a plain `init_process_group(backend="gloo")` or `new_group` request with the flagos backend when the process accelerator is the flagos device | +| `FLAGOS_DIST_STAGED_GLOO` | `1` | Allow the host-staged gloo inner backend, the last fallback tier when no vendor communicator is available. Set `0` to fail loudly instead of staging | +| `FLAGOS_DIST_FORCE_NCCL` | `0` | Test-only: in the manual MetaX distributed tests, skip FlagCX and use NCCL | + +## Vendor and framework compatibility + +Import-time shims that adapt a vendor runtime or another framework to the `flagos` device. None is needed on a stock CUDA box. + +| Variable | Default | Purpose | +|---|---|---| +| `FLAGOS_ALIAS_CUDA` | `1` | Alias the `cuda` device string to `flagos` for drop-in compatibility. Set `0` to opt out | +| `FLAGOS_DISABLE_CUDA_SHIM` | `0` | Skip registering the `torch.cuda` compatibility shim for generic GPU operations | +| `FLAGOS_METAX_CUDART_SHIM` | `0` | Preload the `libcudart` version-tag shim before `import torch`; needed on MetaX with generic PyTorch wheels | +| `FLAGOS_METAX_COMPAT` | `0` | Patch FlagGems `torch.cuda` device queries for MetaX compatibility | +| `FLAGOS_DCU_HIP_VERSION` | none | Override HIP version detection for the DCU runtime | +| `FLAGOS_DCU_SKIP_RUNTIME_CHECK` | `0` | Skip the DCU post-import checks, for deliberately testing a non-matching wheel pair | +| `FLAGOS_DCU_SDPA_FLASH` | `1` | On DCU, point DTK's SDPA selector at its CUTLASS flash adapter instead of the math decomposition | +| `FLAGOS_DISABLE_APEX_COMPAT` | `0` | Disable the optional Apex multi-tensor compatibility layer | +| `FLAGOS_DISABLE_QWENIMAGE_ROPE` | `0` | Leave `diffusers`' Qwen-Image rotary-embedding table alone, to measure the difference | +| `FLAGOS_DISABLE_FLEX_ATTENTION_COMPAT` | `0` | Leave flex-attention's hard-coded `{cuda, cpu, xpu, hpu}` device gate in place | + +## Assets, libraries and compilation + +| Variable | Default | Purpose | +|---|---|---| +| `FLAGOS_DISABLE_CUDA_ASSETS` | `0` | Skip preloading the bundled `libtorch_cuda.so` and CUDA libraries (for in-tree builds and for PPU) | +| `FLAGOS_VENDOR_TORCH_LIB` | auto-discovered | Path to the vendor torch's `lib` directory when no bundled vendor libtorch is present | +| `FLAGOS_USE_CACHING_ALLOCATOR` | `1` | Caching device allocator; set `0` to hand every allocation to the vendor runtime | +| `FLAGOS_USE_FLAGTREE` | `0` | Assert that a FlagTree build is the active Triton (required on Ascend when the compiler is FlagTree) | +| `FLAGOS_COMPILE_FALLBACK_EAGER` | `0` | Fall back to eager mode when `torch.compile` meets unsupported operations | +| `FLAGOS_TILEOPS_USE_L2` | `0` | Use the TileOps L2-cache tier | +| `FLAGOS_TILEOPS_CACHE_MAX` | `512` | TileOps instance-cache capacity | +| `FLAGOS_TILEOPS_DISABLE_ALL_CACHE` | `0` | Neutralize every TileLang cache (correct but slow; set before `tileops` is imported) | + +## Code generation + +Inputs to the operator-binding generators; they matter when rebuilding a platform's generated routes, not at runtime. + +| Variable | Default | Purpose | +|---|---|---| +| `FLAGOS_EXEC_CACHE` | `1` | Cache Ascend operator-codegen execution results; `0` forces regeneration | +| `FLAGOS_CODEGEN_ALL` | `0` | Generate routes for the full leaf-CUDA operator set rather than the supported subset | + +## BPU compiler + +The BPU graph path compiles through `hbdk4` on an x86 host. `FLAGOS_BPU_MARCH` selects the micro-architecture (`nash-p`, `nash-e`, `nash-m`), `FLAGOS_BPU_QUANTIZE` (on by default) keeps convolution on the device by inserting int8 quantization, `FLAGOS_BPU_CACHE` sets the compiler cache directory, and `FLAGOS_BPU_X86_PYTHON` / `FLAGOS_BPU_X86_EMULATOR` / `FLAGOS_BPU_X86_STUBS` describe the x86 host and emulator used for on-board compilation. + +The complete variable list — including the interoperability names owned by other projects, the codegen-only inputs, and retired names kept inert — is maintained in the upstream repository at `docs/reference/environment-variables.md`, alongside the registry in `torch_fl/_env.py` that the documentation is checked against. diff --git a/docs/torch_fl_en/reference/platform-capability.md b/docs/torch_fl_en/reference/platform-capability.md new file mode 100644 index 0000000000..ab8567b4b5 --- /dev/null +++ b/docs/torch_fl_en/reference/platform-capability.md @@ -0,0 +1,51 @@ +# Platform capability matrix + +What each `FLAGOS_ACCELERATOR` value actually provides: which operator path its wheels build, where its device runtime comes from, and which routing decisions it makes by default. This is the per-platform answer to "does my accelerator support X at all", one level below the {doc}`compatibility matrix ` (which reports validation status per capability). + +## Operator path + +| Platform | `FLAGOS_ACCELERATOR` | Operator path | Device runtime sources | Native vendor kernel tree | Status | +|---|---|---|---|---|---| +| NVIDIA CUDA | `cuda` (default) | Native CUDA, FlagGems, generated CUDA boxing | `cuda` | — (uses CUDA/FlagGems paths) | Stable | +| PPU | `ppu` | CUDA-ABI boxing against the bundled PPU libtorch | `cuda` + bundled `lib_ppu/` | — (boxing only) | Experimental | +| MetaX | `metax` | CUDA-ABI boxing against the bundled MACA libtorch | `metax` | — (handwritten MetaX kernels retired) | Stable | +| Hygon DCU | `dcu` | CUDA-ABI boxing over the hipified DTK torch | `cuda` + DCU DTK-core ABI shim | — (boxing only) | Beta | +| Ascend | `ascend` | Vendor-native (ACLNN) with FlagGems via FlagTree | `ascend` | `ascend` | Beta | +| Enflame GCU | `gcu` | Vendor-native (topsaten) | `gcu` | `gcu` | Beta | +| Moore Threads MUSA | `musa` | Vendor-native (mudnn) | `musa` | `musa` | Experimental | +| TsingMicro | `tsingmicro` | CUDA-ABI boxing (Kuiper SDK) | `tsingmicro` | — | Runtime only | +| D-Robotics BPU | `bpu` | No per-operator kernels; whole-graph compilation | `bpu` | — | Runtime only | + +## Kernel sets + +`FLAGOS_BUILD_*` switches decide which kernel sets a wheel compiles. Which are on by default is a property of the platform, not a per-build choice: + +| Kernel set | What it provides | +|---|---| +| `vendor` | The platform's native operator tree, where one exists | +| `flaggems` | The FlagGems Python (Triton) dispatch path | +| `flaggems_cpp` | The FlagGems C++ path (`liboperators.so`) | +| `boxing` | Generated CUDA-boxing kernels for CUDA-ABI platforms | +| `tileops` | TileOps/TileLang kernels for SM90 NVIDIA parts | + +A wheel records the sets it was built with; runtime readers consult that record rather than the environment, so an exported `FLAGOS_BUILD_*` cannot make a wheel describe itself wrongly. See the {doc}`environment variable reference ` for the per-platform defaults. + +## Intentional asymmetries + +These are design decisions, not gaps waiting to be filled: + +- **BPU** has no per-operator kernels at all: eager operators reach the CPU fallback, and acceleration comes from `torch.compile(backend="bpu")`, which compiles a whole graph. +- **TsingMicro** is a build target with a runtime selector but no documented per-operator kernel set and no install guide. +- **MUSA** exposes its own accelerator directory for stream/event wrappers only — it has no `torch.cuda`-style compatibility module, so code that reaches for `torch.cuda.*` device queries on MUSA does not get a translation layer. +- **PPU** profiles through CUPTI but emits no `gpu_memset` activity, so its memset-related profiler parity cases are out of scope rather than failing. +- **CUDA boxing is shared** by CUDA, MetaX, PPU and DCU: one generated kernel set, four vendor runtime sources. + +## Device runtime and Python layers + +| Layer | Where it lives | Note | +|---|---|---| +| Generated ATen bindings | `csrc/aten/generated/` | Registered against `PrivateUse1`; the CUDA/boxing and FlagGems callers | +| Vendor-native kernels | `csrc/aten/backends//` | Generated per the operator surface the vendor library actually exports | +| Device runtime | `csrc/runtime/accelerator//` | DCU and PPU reuse the CUDA runtime tree and add their own pieces | +| Python device module | `torch_fl/flagos/` | Streams, events, RNG, AMP, memory, meta kernels | +| Routing tables | `torch_fl/configs/backends_.conf` | One `op = backend` entry per operator | diff --git a/docs/torch_fl_en/reference/troubleshooting.md b/docs/torch_fl_en/reference/troubleshooting.md new file mode 100644 index 0000000000..c94e47ac5c --- /dev/null +++ b/docs/torch_fl_en/reference/troubleshooting.md @@ -0,0 +1,63 @@ +# Troubleshooting + +Symptom-first fixes for the problems that recur across platforms. A device count of 0 with a working driver, a symbol error at import, or a crash in one operator usually traces back to import order, a missing vendor library, or the wrong compiler — the sections below cover each in turn. + +## Import order + +`import torch_fl` must come **before** `import torch` in a fresh process on every CUDA-ABI platform (CUDA, MetaX, PPU, DCU). At import, Torch-FL preloads the vendor `libtorch` and the CUDA assets; if `torch` is imported first, PyTorch caches its stub CUDA hooks and the preload has no effect. + +```python +import torch_fl # first +import torch +``` + +Symptoms of getting this wrong: + +| Symptom | Where | +|---|---| +| `Cannot initialize CUDA without ATen_cuda library` | CUDA | +| `undefined symbol` from `libtorch` | MetaX | +| Undefined `c10` symbol, or a crash instead of an exception | MUSA, when the wheel was built without `--no-build-isolation` | +| A wrong-vendor build resolved, or a `dlopen` abort | any CUDA-ABI platform | + +## `torch.flagos.device_count()` returns 0 + +Work through these in order: + +1. **Driver visible?** `nvidia-smi` (CUDA), `mx-smi` (MetaX), `hy-smi` (DCU), vendor tools on the others. If the driver does not see the device, no software fix helps. +2. **Runtime initialized?** A driver/runtime version skew is the most common cause. On CUDA, reinstall the matching `nvidia-*-cu12` runtime packages. +3. **Import order** — see above. +4. **Device nodes / visibility.** Ascend needs accessible `/dev/davinci*` nodes; CUDA honours `CUDA_VISIBLE_DEVICES`, and an out-of-range `device="flagos:N"` raises `invalid device ordinal`. +5. **Vendor library reachable.** A MetaX wheel that cannot find `/opt/maca`, or one whose `LD_LIBRARY_PATH` overrides the MACA runtime paths, reports a device-count mismatch between `torch.cuda` and `torch.flagos`. + +## Missing vendor libraries + +| Message | Platform | Cause and fix | +|---|---|---| +| `cannot open shared object file: libhydmi.so` | DCU | `hyhal` is not on `LD_LIBRARY_PATH`; add `/usr/local/hyhal/lib` (or `/opt/hyhal/lib`) | +| `MIOpen: librt.so not found` | DCU | DTK's MIOpen config references a removed `/usr/lib/.../librt.so`; the `FLAGOS_ACCELERATOR=dcu` build branch rewrites it — verify that selector is set and you are on current code | +| `libtorch_cuda.so not found` | PPU | PPU builds ship no bundled CUDA assets: export `FLAGOS_DISABLE_CUDA_ASSETS=1` before running Python | +| `CUDA_HOME not set` | PPU | export `CUDA_HOME=/usr/local/PPU_SDK/CUDA_SDK` before building | + +## Compiler and Triton + +| Message | Cause and fix | +|---|---| +| `No backend registered for 'hcu'` | The active `triton` is not the FlagTree build: uninstall every `triton` until clean, then install the FlagTree wheel for your platform from the FlagOS index | +| Two `triton` distributions in one environment | pip cannot remove a copy that ships no `dist-info`; delete `site-packages/triton` by path first, then reinstall | +| `Invalid cross-device link` (PPU) | pip cache and build dir are on different filesystems; download the wheel and install the file directly | +| A FlagGems-import error after setting `FLAGGEMS_DIR` to an incompatible build | install FlagGems and FlagTree from the same index for the same platform | + +FlagGems routes resolve kernels by name at dispatch time, so a FlagGems-backed operator raises rather than silently falling back when the package is missing. To confirm what actually served a call, run with `FLAGOS_LOG=dispatch`. + +## Distributed + +- A `ProcessGroupGloo` **rejects flagos tensors outright**. With `FLAGOS_DIST_REDIRECT_GLOO` (on by default) a plain `init_process_group(backend="gloo")` is answered with the flagos backend instead. +- If no vendor communicator (FlagCX / NCCL / HCCL / MCCL) is available, the host-staged gloo tier is used, which copies each operand device → host → device. Set `FLAGOS_DIST_STAGED_GLOO=0` to fail loudly instead of silently staging. +- `extended_api creator not found` when building FlagCX for PPU: rebuild with `FLAGCX_ADAPTOR=nvidia`. + +## Profiler and torch.compile + +- `torch.profiler` on a platform whose tracer library is not at its default path: point `FLAGOS_TRACER_LIBRARY` at the installed library. +- PPU and MUSA may show no device events when the installed PyTorch build is a CPU-only wheel: it supplies no `PrivateUse1` Kineto resolver, so collected activities never surface as device events. This is an environment limitation, not a tracer defect. +- `torch.compile` on a platform with no vendor Triton stack installed: the Inductor route needs the platform's FlagTree (or vendor Triton) build; `FLAGOS_COMPILE_FALLBACK_EAGER=1` falls back to eager for unsupported operations. diff --git a/docs/torch_fl_en/release_notes/release-notes.md b/docs/torch_fl_en/release_notes/release-notes.md index 1a79c5e5b9..2cd254b81a 100644 --- a/docs/torch_fl_en/release_notes/release-notes.md +++ b/docs/torch_fl_en/release_notes/release-notes.md @@ -1,14 +1,14 @@ -# Release Notes - -This section includes the release information for Torch-FL. - -## v0.1.0 - -Initial release of Torch-FL as part of FlagOS. - -- **One `flagos` device** — a `PrivateUse1`-based PyTorch device plugin; standard PyTorch APIs, tensor methods and storage work unchanged. -- **Per-operator backend routing** — FlagGems Triton kernels, vendor-native operator libraries, CUDA compatibility boxing and explicit CPU fallback behind one device name, with a per-platform routing table and per-operator overrides. -- **Multi-platform support** — NVIDIA CUDA, MetaX, Huawei Ascend, PPU, Hygon DCU, Enflame GCU, Moore Threads MUSA and D-Robotics BPU, each with its own build selector. -- **Training stack** — eager execution and autograd, `torch.autocast("flagos")` and `torch.amp.GradScaler("flagos")`, `torch.compile` integration, and DDP/FSDP support through `ProcessGroupFlagOS`. -- **Observability** — `torch.profiler` integration with a device timeline and flow arrows, per-operator dispatch and fallback logging, and a wheel compatibility manifest with the `torch-fl-preflight` inspector. -- **Platform compatibility** — CUDA boxing, native ACLNN and topsaten backends, and the FlagGems Python and C++ dispatch paths, with status levels recorded per platform. +# Release Notes + +This section includes the release information for Torch-FL. + +## v2.10.0 + +Initial release of Torch-FL as part of FlagOS. + +- **One `flagos` device** — a `PrivateUse1`-based PyTorch device plugin; standard PyTorch APIs, tensor methods and storage work unchanged. +- **Per-operator backend routing** — FlagGems Triton kernels, vendor-native operator libraries, CUDA compatibility boxing and explicit CPU fallback behind one device name, with a per-platform routing table and per-operator overrides. +- **Multi-platform support** — NVIDIA CUDA, MetaX, Huawei Ascend, PPU, Hygon DCU, Enflame GCU, Moore Threads MUSA and D-Robotics BPU, each with its own build selector. +- **Training stack** — eager execution and autograd, `torch.autocast("flagos")` and `torch.amp.GradScaler("flagos")`, `torch.compile` integration, and distributed support through `ProcessGroupFlagOS` (collectives and DDP; FSDP2 is validated on selected vendors). +- **Observability** — `torch.profiler` integration with a device timeline and flow arrows, per-operator dispatch and fallback logging, and a wheel compatibility manifest with the `torch-fl-preflight` inspector. +- **Platform compatibility** — CUDA boxing, native ACLNN and topsaten backends, and the FlagGems Python and C++ dispatch paths, with status levels recorded per platform. diff --git a/docs/torch_fl_zh/architecture/distributed.md b/docs/torch_fl_zh/architecture/distributed.md index e2a728302f..1af02e4339 100644 --- a/docs/torch_fl_zh/architecture/distributed.md +++ b/docs/torch_fl_zh/architecture/distributed.md @@ -1,70 +1,70 @@ -# 分布式集合通信 - -Torch-FL 通过 `ProcessGroupFlagOS` 为 `flagos` 设备提供分布式支持。它是原生的 `torch.distributed.ProcessGroup` 子类,在导入时完成注册,因此 `torch.distributed.init_process_group("flagos")` 可直接使用,无需对 `torch.distributed.*` 做任何 monkeypatch。 - -## 工作原理 - -`flagos` 张量与厂商张量共享同一块物理设备内存,因此集合通信只需要元数据转换,而不需要数据拷贝: - -1. 某个集合虚函数被调用,传入 `privateuseone` 张量。 -2. 在内部后端需要时,把张量转换为其期望的设备视图(基于同一 `data_ptr` 的零拷贝视图)。 -3. 调用被委派给包装的内部后端。 -4. 原样返回内部后端的 `Work` 对象,调用方(包括 DDP 的 reducer)因此拿到类型正确的 future。 - -`ProcessGroupFlagOS` 覆盖了所有集合虚函数 —— allreduce、allgather(列表形式与 into-tensor 形式)、reduce-scatter、all-to-all(普通与 single)、broadcast、reduce、gather、scatter、send/recv 及其立即版本,以及 barrier —— 而不是只覆盖少数 API,这样集合调用不会悄悄绕过转换。 - -## 后端选择 - -内部通信后端在创建通信组时按以下优先级解析: - -1. **FlagCX** — 异构集合通信库,可导入时优先使用。FlagCX 会为其自身设备注册后端(`flagcx`);`ProcessGroupFlagOS` 通过 `extended_api=True` 的创建接口构造其 `ProcessGroupFlagCX`。 -2. **厂商原生后端** — NVIDIA 与 MetaX 使用 `NCCL`,Ascend 使用 `HCCL`,摩尔线程 MUSA 使用 `MCCL`。 -3. **Host-staged gloo** — 最后一级回退,也是唯一不依赖任何厂商库的一级;它在每次集合通信时把操作数按 设备 → 主机 → 设备 拷贝。设置 `FLAGOS_DIST_STAGED_GLOO=0` 可拒绝该级别并直接失败。首次在该级别建立通信组时会输出一次告警。 - -没有点名 `flagos` 后端的请求也会被处理:`torch.distributed` 会把任何它不认识的设备类型路由到 gloo,而 `ProcessGroupGloo` 会直接拒绝 flagos 张量。在 `FLAGOS_DIST_REDIRECT_GLOO`(默认开启)下,当进程加速器是 flagos 设备时,普通的 `init_process_group(backend="gloo")` 或 `new_group` 请求会由 `flagos` 后端响应。 - -## 使用方式 - -```python -import torch -import torch_fl -import torch_fl.distributed as flagos_dist - -# "auto"(默认):优先 FlagCX,回退到厂商原生后端 -# "flagcx":强制使用 flagos / FlagCX,不可用时回退厂商后端 -# "nccl":强制 NCCL(NVIDIA、MetaX) -# "hccl":强制 HCCL(Ascend) -flagos_dist.init_process_group(backend="auto") - -model = MyModel().to("flagos:0") -model = flagos_dist.DistributedDataParallel(model) -flagos_dist.move_buffers_to_device(model, "flagos:0") -``` - -`torch_fl.distributed` 对外提供 `init_process_group`、`DistributedDataParallel` 与 `move_buffers_to_device`。在后端已注册的前提下,也可以直接调用 `torch.distributed.init_process_group("flagos")`,或由 `device_id=torch.device("privateuseone:0")` 自动选择。 - -### DDP - -导入时,Torch-FL 会补丁 `torch.nn.parallel.DistributedDataParallel.__init__`。当模型位于 `flagos` 设备上时,该补丁会: - -- 强制使用 Python reducer,绕过 C++ reducer 的 CUDA 断言; -- 替换默认的梯度累积钩子(其使用在 `privateuseone` 上没有分发的函数式集合通信),改为通过 `dist.all_reduce` 走 `ProcessGroupFlagOS`。 - -`torch.nn.DataParallel` 与函数式 `torch.nn.parallel.data_parallel` 也做了同样的补丁,使其副本放置到 flagos 设备上,而不是在设备类型探测处失败。 - -## 各厂商状态 - -| 厂商 | FlagCX 路径 | 原生回退 | 视图转换 | 说明 | -|---|---|---|---|---| -| NVIDIA | 可用 | NCCL | flagos → cuda 视图 | 集合通信与 DDP 梯度同步已在 2×/8× A100 上实测验证 | -| MetaX | 可复用 | 经 MACA libtorch 的 NCCL 形态 MCCL | flagos → cuda 视图 | 未纳入 CI | -| Ascend | 推荐的优先路径 | HCCL(自定义后端类型) | flagos → npu 视图 | Ascend 上没有 CUDA 兼容层。仅为架构层面路由,无集合级 CI 覆盖 | -| 海光 DCU | 可复用 | 经 DTK 的 RCCL | flagos → cuda 视图 | `all_reduce`/DDP 已在 2 卡上实测,未纳入 CI | -| 摩尔线程 MUSA | 可复用 | MCCL | flagos → cuda 视图 | host-staged gloo 已在 MTT S5000 上实测 | -| 燧原 GCU | 优先路径 | 无(仅 FlagCX) | 无需转换 | 已在两块 S60 上实测:集合通信、barrier、DDP 前反向与梯度同步、FSDP2 `fully_shard` 训练与分片 state-dict 存取 | - -## 限制 - -- 集合通信覆盖范围按厂商验证,缺口如实记录而非默认成立。在燧原 GCU 上,点对点通信、`gather`/`scatter` 的 root 参数、all-to-all、多机建联、进程故障恢复以及超过两台设备的部署尚未验证。 -- Ascend 的分布式支持属于架构层面:路由与视图逻辑已经存在,但没有集合级 CI 覆盖。 -- host-staged gloo 级别以正确性优先,每次集合通信都要付出一次设备 → 主机 → 设备的拷贝。 +# 分布式集合通信 + +Torch-FL 通过 `ProcessGroupFlagOS` 为 `flagos` 设备提供分布式支持。它是原生的 `torch.distributed.ProcessGroup` 子类,在导入时完成注册,因此 `torch.distributed.init_process_group("flagos")` 可直接使用,无需对 `torch.distributed.*` 做任何 monkeypatch。 + +## 工作原理 + +`flagos` 张量与厂商张量共享同一块物理设备内存,因此集合通信只需要元数据转换,而不需要数据拷贝: + +1. 某个集合虚函数被调用,传入 `privateuseone` 张量。 +2. 在内部后端需要时,把张量转换为其期望的设备视图(基于同一 `data_ptr` 的零拷贝视图)。 +3. 调用被委派给包装的内部后端。 +4. 原样返回内部后端的 `Work` 对象,调用方(包括 DDP 的 reducer)因此拿到类型正确的 future。 + +`ProcessGroupFlagOS` 覆盖了所有集合虚函数 —— allreduce、allgather(列表形式与 into-tensor 形式)、reduce-scatter、all-to-all(普通与 single)、broadcast、reduce、gather、scatter、send/recv 及其立即版本,以及 barrier —— 而不是只覆盖少数 API,这样集合调用不会悄悄绕过转换。 + +## 后端选择 + +内部通信后端在创建通信组时按以下优先级解析: + +1. **FlagCX** — 异构集合通信库,可导入时优先使用。FlagCX 会为其自身设备注册后端(`flagcx`);`ProcessGroupFlagOS` 通过 `extended_api=True` 的创建接口构造其 `ProcessGroupFlagCX`。 +2. **厂商原生后端** — NVIDIA 与 MetaX 使用 `NCCL`,Ascend 使用 `HCCL`,摩尔线程 MUSA 使用 `MCCL`。 +3. **Host-staged gloo** — 最后一级回退,也是唯一不依赖任何厂商库的一级;它在每次集合通信时把操作数按 设备 → 主机 → 设备 拷贝。设置 `FLAGOS_DIST_STAGED_GLOO=0` 可拒绝该级别并直接失败。首次在该级别建立通信组时会输出一次告警。 + +没有点名 `flagos` 后端的请求也会被处理:`torch.distributed` 会把任何它不认识的设备类型路由到 gloo,而 `ProcessGroupGloo` 会直接拒绝 flagos 张量。在 `FLAGOS_DIST_REDIRECT_GLOO`(默认开启)下,当进程加速器是 flagos 设备时,普通的 `init_process_group(backend="gloo")` 或 `new_group` 请求会由 `flagos` 后端响应。 + +## 使用方式 + +```python +import torch +import torch_fl +import torch_fl.distributed as flagos_dist + +# "auto"(默认):优先 FlagCX,回退到厂商原生后端 +# "flagcx":强制使用 flagos / FlagCX,不可用时回退厂商后端 +# "nccl":强制 NCCL(NVIDIA、MetaX) +# "hccl":强制 HCCL(Ascend) +flagos_dist.init_process_group(backend="auto") + +model = MyModel().to("flagos:0") +model = flagos_dist.DistributedDataParallel(model) +flagos_dist.move_buffers_to_device(model, "flagos:0") +``` + +`torch_fl.distributed` 对外提供 `init_process_group`、`DistributedDataParallel` 与 `move_buffers_to_device`。在后端已注册的前提下,也可以直接调用 `torch.distributed.init_process_group("flagos")`,或由 `device_id=torch.device("privateuseone:0")` 自动选择。 + +### DDP + +导入时,Torch-FL 会补丁 `torch.nn.parallel.DistributedDataParallel.__init__`。当模型位于 `flagos` 设备上时,该补丁会: + +- 强制使用 Python reducer,绕过 C++ reducer 的 CUDA 断言; +- 替换默认的梯度累积钩子(其使用在 `privateuseone` 上没有分发的函数式集合通信),改为通过 `dist.all_reduce` 走 `ProcessGroupFlagOS`。 + +`torch.nn.DataParallel` 与函数式 `torch.nn.parallel.data_parallel` 也做了同样的补丁,使其副本放置到 flagos 设备上,而不是在设备类型探测处失败。 + +## 各厂商状态 + +| 厂商 | FlagCX 路径 | 原生回退 | 视图转换 | 说明 | +|---|---|---|---|---| +| NVIDIA | 可用 | NCCL | flagos → cuda 视图 | 集合通信与 DDP 梯度同步已在 NVIDIA 多机多卡环境实测验证 | +| MetaX | 可复用 | 经 MACA libtorch 的 NCCL 形态 MCCL | flagos → cuda 视图 | 未持续验证 | +| Ascend | 推荐的优先路径 | HCCL(自定义后端类型) | flagos → npu 视图 | Ascend 上没有 CUDA 兼容层。仅为架构层面路由,无集合级验证 | +| 海光 DCU | 可复用 | 经 DTK 的 RCCL | flagos → cuda 视图 | `all_reduce`/DDP 已在 2 卡上实测,未持续验证 | +| 摩尔线程 MUSA | 可复用 | MCCL | flagos → cuda 视图 | FlagCX 优先路由与 MCCL 回退已实现;该主机上的端到端多进程集合通信尚未验证 | +| 燧原 GCU | 优先路径 | 无(仅 FlagCX) | 无需转换 | 已在两块 S60 上实测:集合通信、barrier、DDP 前反向与梯度同步、FSDP2 `fully_shard` 训练与分片 state-dict 存取 | + +## 限制 + +- 集合通信覆盖范围按厂商验证,缺口如实记录而非默认成立。在燧原 GCU 上,点对点通信、`gather`/`scatter` 的 root 参数、all-to-all、多机建联、进程故障恢复以及超过两台设备的部署尚未验证。 +- Ascend 的分布式支持属于架构层面:路由与视图逻辑已经存在,但没有集合级验证。 +- host-staged gloo 级别以正确性优先,每次集合通信都要付出一次设备 → 主机 → 设备的拷贝。 diff --git a/docs/torch_fl_zh/architecture/profiler.md b/docs/torch_fl_zh/architecture/profiler.md index eb88dcd077..70672c602d 100644 --- a/docs/torch_fl_zh/architecture/profiler.md +++ b/docs/torch_fl_zh/architecture/profiler.md @@ -1,73 +1,73 @@ -# Profiler 集成 - -`torch.profiler` 通过编译进 wheel 的设备追踪器支持 `flagos` 设备。`torch.profiler.profile(activities=[CPU, PrivateUse1])` 产出的 trace 与同一工作负载在 `torch.cuda` 上的 trace 结构等价: - -- **流向箭头**把每个 CPU 算子连接到它启动的设备内核。 -- **设备时间归因**:`prof.key_averages()` 报告逐算子的 `self_device_time_total`。 -- **完整的内核元数据**(grid/block、occupancy、共享内存、寄存器数量)以及已反修饰的内核名。 -- **运行期事件**携带由 callback id 解码出的真实 API 名称,而非占位符。 -- **memcpy 与 memset** 活动与内核一起采集。 - -## 三层架构 - -新增一个厂商意味着只写一个文件:满足与厂商无关接口的追踪器。 - -| 层次 | 文件 | 职责 | -|---|---|---| -| 与厂商无关的接口 | `csrc/profiler/device_tracer.h` | `DeviceTracer`、`DeviceEvent`、`EventKind` —— 厂商需要实现的完整契约 | -| 厂商追踪器 | `cupti_device_tracer.cc`、`cann_device_tracer.cc`、`musa_mupti_device_tracer.cc`、`roctracer_device_tracer.cc`、`gcu_topspti_device_tracer.cc`、`unavailable_device_tracer.cc` | 每个加速器一份实现,另有显式的「无设备活动」回退 | -| 通用适配层 | `flagos_kineto_profiler.{h,cc}` | 与厂商完全解耦的 Kineto/PyTorch profiler 适配器;由 `cupti_shim.h`、`mupti_shim.h`、`topspti_shim.h` 等 `dlopen` shim 绑定厂商活动库 | - -| 加速器 | 活动 API | 状态 | -|---|---|---| -| NVIDIA CUDA | CUPTI | 稳定,对等性套件纳入 CI | -| MetaX | MCPTI(MACA 中 CUDA 兼容的活动 API) | 实验性:C550 + MACA 3.8.0 上七项对等性断言全部通过(本地硬件验证,CI 无厂商 runner) | -| Ascend | MSPTI | Beta:内核/运行期/flow/memcpy 事件及设备时间关联由共享契约覆盖 CI;对等性套件本身不在 CI 中 | -| 海光 DCU | ROCtracer | Beta:对等性套件在 CI 中运行 | -| 摩尔线程 MUSA | MUPTI | 实验性:设备时间线已在 MTT S5000 上实测;CPU-Kineto 关联依赖具体环境 | -| 燧原 GCU | TOPSPTI | 仅运行时:TOPSPTI 能采集活动,但仅有 CPU 的 Kineto 构建不提供 PrivateUse1 resolver,活动无法呈现为设备事件 | -| 其他 | `unavailable_device_tracer.cc` | 显式的「无设备活动」回退 | - -## Correlation id - -一条 trace 中存在两套彼此独立的编号体系,都叫 “correlation”,外观相似但含义不同: - -| | `correlation_id` | `external_correlation_id` | -|---|---|---| -| 归属 | 活动 API 的编号 | PyTorch 的编号 | -| 关联对象 | 一次运行期调用与它产生的设备内核 | 一次设备/运行期活动与发起它的 CPU 算子 | -| 用途 | 绘制流向箭头 | 设备时间归因 | -| trace 字段 | `args["correlation"]` | `args["External id"]` | - -传错会静默失败:trace 仍能正常生成、箭头消失、`self_device_time_total` 读数为 0。设备时间必须通过 `getLinkedActivity` 回调(external id)解析,而流向箭头基于活动 correlation id。 - -## 调试 - -| 变量 | 默认值 | 作用 | -|---|---|---| -| `FLAGOS_TRACE` | `0` | 设备 profiler shim 的详细日志:事件排空、采集时间窗、追踪器绑定与注册 | -| `FLAGOS_TRACER_LIBRARY` | 自动探测 | 当默认路径与已安装驱动不匹配时,覆盖 profiler shim `dlopen` 的追踪器库 | - -有两类告警刻意不受 `FLAGOS_TRACE` 控制 —— 空的 linked-activity 回调与活动记录布局不匹配 —— 因为它们都会静默地把设备时间归零。 - -## 对等性测试与基线 - -`tests/integration/test_profiler_parity.py` 将 `flagos` trace 与在原生 `torch+cuda` 上采集的基线对比。七项断言全部只检查结构,不检查计数或耗时: - -| # | 断言 | 检查内容 | -|---|---|---| -| 1 | `test_category_coverage` | `flagos` trace 覆盖 `torch.cuda` 基线中的全部类别 | -| 2 | `test_flow_arrows_are_paired` | 每个流向箭头的起始半边都有对应的结束半边 | -| 3 | `test_arg_key_supersets` | 各类别的参数键是基线的超集 | -| 4 | `test_device_time_attribution` | 算子的 `self_device_time_total` 与其拥有的设备事件时长之和相吻合 | -| 5 | `test_kernel_names_are_demangled` | 不存在裸露的 C++ 修饰符号 | -| 6 | `test_runtime_names_come_from_cbid` | 运行期事件名由 callback id 解码得到 | -| 7 | `test_capture_window_containment` | 没有设备或运行期事件逸出采集时间窗 | - -基线位于 `tests/data/profiler_cuda_baseline.json`。第 2 项断言刻意比上游 `torch.cuda` 更严格:它是 Torch-FL 自身的约束,保留它是因为流向箭头出现悬空半边正是它要防的回归。 - -## 已知缺口 - -- 不采集 `overhead` 活动类别:它衡量的是 profiling 自身的开销而非用户工作负载。该缺口以已知项形式记录在基线中,而不是被静默省略。 -- MetaX 的 MCPTI 运行期 callback id 并非 NVIDIA CUPTI id,追踪器使用 MetaX 的 callback 命名空间,并把 API 名称解析推迟到活动刷出之后 —— 在 buffer 回调中调用解析器可能使 profiler 死锁。该扫描逻辑尚未在多个 MetaX SDK 版本、多种设备或非默认流上验证。 -- 燧原 GCU 的 profiler 支持仅为运行时级别,见上表。 +# Profiler 集成 + +`torch.profiler` 通过编译进 wheel 的设备追踪器支持 `flagos` 设备。`torch.profiler.profile(activities=[CPU, PrivateUse1])` 产出的 trace 与同一工作负载在 `torch.cuda` 上的 trace 结构等价: + +- **流向箭头**把每个 CPU 算子连接到它启动的设备内核。 +- **设备时间归因**:`prof.key_averages()` 报告逐算子的 `self_device_time_total`。 +- **完整的内核元数据**(grid/block、occupancy、共享内存、寄存器数量)以及已反修饰的内核名。 +- **运行期事件**携带由 callback id 解码出的真实 API 名称,而非占位符。 +- **memcpy 与 memset** 活动与内核一起采集。 + +## 三层架构 + +新增一个厂商意味着只写一个文件:满足与厂商无关接口的追踪器。 + +| 层次 | 文件 | 职责 | +|---|---|---| +| 与厂商无关的接口 | `csrc/profiler/device_tracer.h` | `DeviceTracer`、`DeviceEvent`、`EventKind` —— 厂商需要实现的完整契约 | +| 厂商追踪器 | `cupti_device_tracer.cc`、`cann_device_tracer.cc`、`musa_mupti_device_tracer.cc`、`roctracer_device_tracer.cc`、`gcu_topspti_device_tracer.cc`、`unavailable_device_tracer.cc` | 每个加速器一份实现,另有显式的「无设备活动」回退 | +| 通用适配层 | `flagos_kineto_profiler.{h,cc}` | 与厂商完全解耦的 Kineto/PyTorch profiler 适配器;由 `cupti_shim.h`、`mupti_shim.h`、`topspti_shim.h` 等 `dlopen` shim 绑定厂商活动库 | + +| 加速器 | 活动 API | 状态 | +|---|---|---| +| NVIDIA CUDA | CUPTI | 稳定,对等性套件已纳入 | +| MetaX | MCPTI(MACA 中 CUDA 兼容的活动 API) | 实验性:已在 C550 + MACA 3.8.0 上完成对等性验证(本地硬件验证,无厂商 runner) | +| Ascend | MSPTI | Beta:内核/运行期/flow/memcpy 事件及设备时间关联由共享契约覆盖;对等性套件本身未纳入 | +| 海光 DCU | ROCtracer | Beta:对等性套件已纳入 | +| 摩尔线程 MUSA | MUPTI | 实验性:设备时间线已在 MTT S5000 上实测;CPU-Kineto 关联依赖具体环境 | +| 燧原 GCU | TOPSPTI | 仅运行时:TOPSPTI 能采集活动,但仅有 CPU 的 Kineto 构建不提供 PrivateUse1 resolver,活动无法呈现为设备事件 | +| 其他 | `unavailable_device_tracer.cc` | 显式的「无设备活动」回退 | + +## Correlation id + +一条 trace 中存在两套彼此独立的编号体系,都叫 “correlation”,外观相似但含义不同: + +| | `correlation_id` | `external_correlation_id` | +|---|---|---| +| 归属 | 活动 API 的编号 | PyTorch 的编号 | +| 关联对象 | 一次运行期调用与它产生的设备内核 | 一次设备/运行期活动与发起它的 CPU 算子 | +| 用途 | 绘制流向箭头 | 设备时间归因 | +| trace 字段 | `args["correlation"]` | `args["External id"]` | + +传错会静默失败:trace 仍能正常生成、箭头消失、`self_device_time_total` 读数为 0。设备时间必须通过 `getLinkedActivity` 回调(external id)解析,而流向箭头基于活动 correlation id。 + +## 调试 + +| 变量 | 默认值 | 作用 | +|---|---|---| +| `FLAGOS_TRACE` | `0` | 设备 profiler shim 的详细日志:事件排空、采集时间窗、追踪器绑定与注册 | +| `FLAGOS_TRACER_LIBRARY` | 自动探测 | 当默认路径与已安装驱动不匹配时,覆盖 profiler shim `dlopen` 的追踪器库 | + +有两类告警刻意不受 `FLAGOS_TRACE` 控制 —— 空的 linked-activity 回调与活动记录布局不匹配 —— 因为它们都会静默地把设备时间归零。 + +## 对等性测试与基线 + +对等性测试将 `flagos` trace 与在原生 `torch+cuda` 上采集的基线对比。全部断言只检查结构,不检查计数或耗时: + +| # | 断言 | 检查内容 | +|---|---|---| +| 1 | `test_category_coverage` | `flagos` trace 覆盖 `torch.cuda` 基线中的全部类别 | +| 2 | `test_flow_arrows_are_paired` | 每个流向箭头的起始半边都有对应的结束半边 | +| 3 | `test_arg_key_supersets` | 各类别的参数键是基线的超集 | +| 4 | `test_device_time_attribution` | 算子的 `self_device_time_total` 与其拥有的设备事件时长之和相吻合 | +| 5 | `test_kernel_names_are_demangled` | 不存在裸露的 C++ 修饰符号 | +| 6 | `test_runtime_names_come_from_cbid` | 运行期事件名由 callback id 解码得到 | +| 7 | `test_capture_window_containment` | 没有设备或运行期事件逸出采集时间窗 | + +基线为在原生 `torch+cuda` 上采集的 trace。第 2 项断言刻意比上游 `torch.cuda` 更严格:它是 Torch-FL 自身的约束,保留它是因为流向箭头出现悬空半边正是它要防的回归。 + +## 已知缺口 + +- 不采集 `overhead` 活动类别:它衡量的是 profiling 自身的开销而非用户工作负载。该缺口以已知项形式记录在基线中,而不是被静默省略。 +- MetaX 的 MCPTI 运行期 callback id 并非 NVIDIA CUPTI id,追踪器使用 MetaX 的 callback 命名空间,并把 API 名称解析推迟到活动刷出之后 —— 在 buffer 回调中调用解析器可能使 profiler 死锁。该扫描逻辑尚未在多个 MetaX SDK 版本、多种设备或非默认流上验证。 +- 燧原 GCU 的 profiler 支持仅为运行时级别,见上表。 diff --git a/docs/torch_fl_zh/architecture/torch-compile.md b/docs/torch_fl_zh/architecture/torch-compile.md index 65366e9894..32919ca489 100644 --- a/docs/torch_fl_zh/architecture/torch-compile.md +++ b/docs/torch_fl_zh/architecture/torch-compile.md @@ -1,98 +1,98 @@ -# torch.compile 集成 - -`flagos` 设备支持 `torch.compile`,可实现自动内核融合并降低调度开销。计算图始终留在 `flagos` 设备上:不存在设备往返,也没有图边界处的拷贝。 - -## 快速开始 - -```python -import torch_fl # MetaX 与 Ascend 上必须先导入 -import torch - -model = torch.nn.Sequential( - torch.nn.Linear(512, 512), - torch.nn.ReLU(), - torch.nn.Linear(512, 512), -).to("flagos:0") - -model = torch.compile(model, backend="flagos") - -x = torch.randn(64, 512, device="flagos:0") -y = model(x) # 自动使用融合后的内核 -``` - -编译模式: - -```python -model = torch.compile(model, backend="flagos") # 默认 -model = torch.compile(model, backend="flagos", mode="max-autotune") # 编译更久,运行更快 -model = torch.compile(model, backend="flagos", options={"max_autotune": True}) -``` - -`mode` 与 `options` 会展开为仅作用于该次编译的 Inductor 配置补丁。该后端始终关闭 CUDA graphs,因此以 CUDA graphs 为主要手段的 `mode="reduce-overhead"` 在这里效果有限。 - -## FlagTree 编译 - -[FlagTree](https://github.com/flagos-ai/FlagTree) 是 Triton 的分支,其编译器面向多种厂商后端。它的集成方式是在**安装期替换**,这是理解它的关键: - -- 它的 wheel 名为 `flagtree`,但安装的模块是 `triton`。 -- 安装它会卸载官方 `triton` 并取而代之。 -- 因此 Inductor 自身的 `import triton` 在安装后就已经指向 FlagTree,Torch-FL 不需要对 `sys.modules` 做任何补丁。 - -后端编译器在 FlagTree 构建期通过 `FLAGTREE_BACKEND` 选择(NVIDIA 与 AMD 不设置),而不是运行期:同一份 Triton 内核代码为不同厂商后端编译。`is_flagtree_active()` 用于检测 FlagTree 构建,`FLAGOS_USE_FLAGTREE=1` 断言当前 Triton 必须是 FlagTree —— 不满足时直接报错,而不是静默使用官方 Triton 编译。 - -由于安装 FlagTree 会移除既有 `triton`,在 `triton` 已被 FlagGems 使用的机器上应将其构建到独立的虚拟环境中。FlagTree 0.6.2 及之后的 wheel 还会安装真实的 `flagtree` 包(FlagPrism 调试器/profiler 的宿主);访问 FlagTree 仍然通过 `triton`。 - -## 后端内部组成 - -| 组件 | 职责 | -|---|---| -| `torch_fl/compile/inductor_backend.py` | 向 `torch._dynamo` 注册 `flagos` 后端,并接入 Inductor 的设备接口 | -| `torch_fl/compile/device_interface.py` | `flagos` 设备的 Inductor GPU 设备注册 | -| `torch_fl/flagos/meta.py` | meta 内核,使追踪阶段能为缺少默认 meta 实现的算子推断输出形状 | -| `torch_fl/compile/flagtree_shim.py` | FlagTree 检测(`is_flagtree_active`、`require_flagtree`),不做导入补丁 | -| `torch_fl/compile/platform_profile.py` | 各平台代码生成 profile 与厂商绕过措施 | -| `torch_fl/compile/triton_*.py` | Triton 集成守卫:64 位索引、字节 load、libdevice、资源上限 | -| `torch_fl/compile/flagtree_ascend_policy.py` | FlagTree 策略注册表上的 Ascend 后端策略,从 `torch.flagos` 而非 `torch_npu` 取答案 | - -## 平台差异 - -- **Ascend** — 通过 FlagTree 的 Ascend 后端编译;插件自有的 `flagos` 策略从自身运行时回答 FlagTree 的策略名称,并强制 `TRITON_ENABLE_TASKQUEUE=false`(任务队列仅 `torch_npu` 支持)。未纳入 CI。 -- **PPU** — FlagTree 在异步 Inductor worker 中选取编译提示时会初始化 CUDA,可能在父进程已初始化 PPU context 后失败,因此 PPU 的 FlagTree 默认串行编译。仅在验证上游修复或刻意选择其他 worker 配置时才显式设置 `TORCHINDUCTOR_COMPILE_THREADS`。 -- **MetaX** — `torch.compile` 已在 CUDA boxing 模式下使用厂商 Triton 与 FlagTree main 验证。 -- **燧原 GCU** — 64 位代码生成守卫把「不支持 64 位」的失败转换为指名算子的可操作错误,见故障排查。 -- **地平线 BPU** — 编译是唯一的加速路径:`torch.compile(backend="bpu")` 追踪计算图,经 hbdk4 编译为 `.hbm` 产物并在 BPU 上执行,默认插入 int8 量化。 -- **CUDA** — `flagos` 已注册为一等 Inductor GPU 设备。CUDA 的 CI 任务中没有 `torch.compile` 步骤,该路径由集成测试而非 CI 覆盖。 - -## 性能 - -融合收益已做正确性验证(`tests/integration/test_compile.py`);与 CUDA 上官方 Inductor 的收益对比基准测试仍待开展。从结构上看两者应当接近 —— 相同的融合 pass、相同的 Triton 代码生成、因为计算图留在 flagos 而不存在每次调用的拷贝 —— 但这是预期而非实测结论。 - -```bash -python tests/perf/bench_compile.py --model=mlp --batch-size=64 -python tests/perf/bench_compile.py --model=transformer --compare-cuda -FLAGOS_USE_FLAGTREE=1 python tests/perf/bench_compile.py -``` - -## 环境变量 - -| 变量 | 默认值 | 作用 | -|---|---|---| -| `FLAGOS_USE_FLAGTREE` | `0` | 要求当前 Triton 为 FlagTree(断言,不切换) | -| `FLAGOS_COMPILE_FALLBACK_EAGER` | `0` | 编译出错时回退到 eager | -| `FLAGOS_TILEOPS_*` | 见参考 | TileOps/TileLang 的 L2 层级、实例缓存容量与缓存开关 | - -路由相关变量(`FLAGOS_BACKEND_CONFIG`、`FLAGOS_OP_`、`FLAGOS_FORCE_BACKEND`)同样作用于编译后的内核,因为编译出的计算图走同一张路由表。 - -## 故障排查 - -| 现象 | 原因与处理 | -|---|---| -| 图捕获或代码生成阶段报错 | 设置 `FLAGOS_COMPILE_FALLBACK_EAGER=1` 回退到 eager;检查是否存在不支持的算子(动态形状、自定义算子)以及缺失的 meta 实现 | -| GCU 上出现 `InductorError: ... has no 64-bit support` | 64 位代码生成守卫拒绝了在不支持 64 位索引的后端上需要 64 位索引的内核形状;缩小张量/索引规模,或让该算子留在原有路由上 | -| 相比 eager 没有加速 | 确认计算图确实编译成功(`TORCH_LOGS=inductor`),检查 Triton 自动调优是否仍在进行,并确认工作负载不是启动开销受限 | -| FlagTree 未生效 | `FLAGOS_USE_FLAGTREE=1` 在官方 Triton 下会报错;检查 `is_flagtree_active()` 并重新安装 FlagTree,它必须替换 `triton` 模块 | -| `torch.compile` 不可用 | 后端在 `torch._dynamo` 可导入时注册;检查 PyTorch 版本,并确认编译前已执行 `import torch_fl` | - -```bash -pytest tests/integration/test_compile.py -v --tb=short -``` +# torch.compile 集成 + +`flagos` 设备支持 `torch.compile`,可实现自动内核融合并降低调度开销。计算图始终留在 `flagos` 设备上:不存在设备往返,也没有图边界处的拷贝。 + +## 快速开始 + +```python +import torch_fl # MetaX 与 Ascend 上必须先导入 +import torch + +model = torch.nn.Sequential( + torch.nn.Linear(512, 512), + torch.nn.ReLU(), + torch.nn.Linear(512, 512), +).to("flagos:0") + +model = torch.compile(model, backend="flagos") + +x = torch.randn(64, 512, device="flagos:0") +y = model(x) # 自动使用融合后的内核 +``` + +编译模式: + +```python +model = torch.compile(model, backend="flagos") # 默认 +model = torch.compile(model, backend="flagos", mode="max-autotune") # 编译更久,运行更快 +model = torch.compile(model, backend="flagos", options={"max_autotune": True}) +``` + +`mode` 与 `options` 会展开为仅作用于该次编译的 Inductor 配置补丁。该后端始终关闭 CUDA graphs,因此以 CUDA graphs 为主要手段的 `mode="reduce-overhead"` 在这里效果有限。 + +## FlagTree 编译 + +[FlagTree](https://github.com/flagos-ai/FlagTree) 是 Triton 的分支,其编译器面向多种厂商后端。它的集成方式是在**安装期替换**,这是理解它的关键: + +- 它的 wheel 名为 `flagtree`,但安装的模块是 `triton`。 +- 安装它会卸载官方 `triton` 并取而代之。 +- 因此 Inductor 自身的 `import triton` 在安装后就已经指向 FlagTree,Torch-FL 不需要对 `sys.modules` 做任何补丁。 + +后端编译器在 FlagTree 构建期通过 `FLAGTREE_BACKEND` 选择(NVIDIA 与 AMD 不设置),而不是运行期:同一份 Triton 内核代码为不同厂商后端编译。`is_flagtree_active()` 用于检测 FlagTree 构建,`FLAGOS_USE_FLAGTREE=1` 断言当前 Triton 必须是 FlagTree —— 不满足时直接报错,而不是静默使用官方 Triton 编译。 + +由于安装 FlagTree 会移除既有 `triton`,在 `triton` 已被 FlagGems 使用的机器上应将其构建到独立的虚拟环境中。较新的 FlagTree wheel 还会安装真实的 `flagtree` 包(FlagPrism 调试器/profiler 的宿主);访问 FlagTree 仍然通过 `triton`。 + +## 后端内部组成 + +| 组件 | 职责 | +|---|---| +| `torch_fl/compile/inductor_backend.py` | 向 `torch._dynamo` 注册 `flagos` 后端,并接入 Inductor 的设备接口 | +| `torch_fl/compile/device_interface.py` | `flagos` 设备的 Inductor GPU 设备注册 | +| `torch_fl/flagos/meta.py` | meta 内核,使追踪阶段能为缺少默认 meta 实现的算子推断输出形状 | +| `torch_fl/compile/flagtree_shim.py` | FlagTree 检测(`is_flagtree_active`、`require_flagtree`),不做导入补丁 | +| `torch_fl/compile/platform_profile.py` | 各平台代码生成 profile 与厂商绕过措施 | +| `torch_fl/compile/triton_*.py` | Triton 集成守卫:64 位索引、字节 load、libdevice、资源上限 | +| `torch_fl/compile/flagtree_ascend_policy.py` | FlagTree 策略注册表上的 Ascend 后端策略,从 `torch.flagos` 而非 `torch_npu` 取答案 | + +## 平台差异 + +- **Ascend** — 通过 FlagTree 的 Ascend 后端编译;插件自有的 `flagos` 策略从自身运行时回答 FlagTree 的策略名称,并强制 `TRITON_ENABLE_TASKQUEUE=false`(任务队列仅 `torch_npu` 支持)。仅由集成测试覆盖。 +- **PPU** — FlagTree 在异步 Inductor worker 中选取编译提示时会初始化 CUDA,可能在父进程已初始化 PPU context 后失败,因此 PPU 的 FlagTree 默认串行编译。仅在验证上游修复或刻意选择其他 worker 配置时才显式设置 `TORCHINDUCTOR_COMPILE_THREADS`。 +- **MetaX** — `torch.compile` 已在 CUDA boxing 模式下使用厂商 Triton 与 FlagTree main 验证。 +- **燧原 GCU** — 64 位代码生成守卫把「不支持 64 位」的失败转换为指名算子的可操作错误,见故障排查。 +- **地平线 BPU** — 编译是唯一的加速路径:`torch.compile(backend="bpu")` 追踪计算图,经 hbdk4 编译为 `.hbm` 产物并在 BPU 上执行,默认插入 int8 量化。 +- **CUDA** — `flagos` 已注册为一等 Inductor GPU 设备。该路径由集成测试覆盖,而非专门的构建任务。 + +## 性能 + +融合收益已做正确性验证(`tests/integration/test_compile.py`);与 CUDA 上官方 Inductor 的收益对比基准测试仍待开展。从结构上看两者应当接近 —— 相同的融合 pass、相同的 Triton 代码生成、因为计算图留在 flagos 而不存在每次调用的拷贝 —— 但这是预期而非实测结论。 + +```bash +python tests/perf/bench_compile.py --model=mlp --batch-size=64 +python tests/perf/bench_compile.py --model=transformer --compare-cuda +FLAGOS_USE_FLAGTREE=1 python tests/perf/bench_compile.py +``` + +## 环境变量 + +| 变量 | 默认值 | 作用 | +|---|---|---| +| `FLAGOS_USE_FLAGTREE` | `0` | 要求当前 Triton 为 FlagTree(断言,不切换) | +| `FLAGOS_COMPILE_FALLBACK_EAGER` | `0` | 编译出错时回退到 eager | +| `FLAGOS_TILEOPS_*` | 见参考 | TileOps/TileLang 的 L2 层级、实例缓存容量与缓存开关 | + +路由相关变量(`FLAGOS_BACKEND_CONFIG`、`FLAGOS_OP_`、`FLAGOS_FORCE_BACKEND`)同样作用于编译后的内核,因为编译出的计算图走同一张路由表。 + +## 故障排查 + +| 现象 | 原因与处理 | +|---|---| +| 图捕获或代码生成阶段报错 | 设置 `FLAGOS_COMPILE_FALLBACK_EAGER=1` 回退到 eager;检查是否存在不支持的算子(动态形状、自定义算子)以及缺失的 meta 实现 | +| GCU 上出现 `InductorError: ... has no 64-bit support` | 64 位代码生成守卫拒绝了在不支持 64 位索引的后端上需要 64 位索引的内核形状;缩小张量/索引规模,或让该算子留在原有路由上 | +| 相比 eager 没有加速 | 确认计算图确实编译成功(`TORCH_LOGS=inductor`),检查 Triton 自动调优是否仍在进行,并确认工作负载不是启动开销受限 | +| FlagTree 未生效 | `FLAGOS_USE_FLAGTREE=1` 在官方 Triton 下会报错;检查 `is_flagtree_active()` 并重新安装 FlagTree,它必须替换 `triton` 模块 | +| `torch.compile` 不可用 | 后端在 `torch._dynamo` 可导入时注册;检查 PyTorch 版本,并确认编译前已执行 `import torch_fl` | + +```bash +pytest tests/integration/test_compile.py -v --tb=short +``` diff --git a/docs/torch_fl_zh/getting_started/installation.md b/docs/torch_fl_zh/getting_started/installation.md index 0141c7e4ef..132e22ad6d 100644 --- a/docs/torch_fl_zh/getting_started/installation.md +++ b/docs/torch_fl_zh/getting_started/installation.md @@ -1,223 +1,255 @@ -# 安装 - -## 选择平台 - -每个平台由一个构建变量 `FLAGOS_ACCELERATOR` 选择,并且各自有独立的执行路径: - -| 平台 | `FLAGOS_ACCELERATOR` | 执行路径 | 状态 | -|---|---|---|---| -| NVIDIA CUDA | `cuda`(默认) | 基于外部 `libtorch_cuda.so` 的 CUDA boxing | 稳定 | -| MetaX | `metax` | 通过 `cu-bridge` 对厂商 libtorch 进行 CUDA boxing | 稳定 | -| Ascend | `ascend` | 原生 ACLNN 算子后端,通过 FlagTree(Triton 3.5)使用 FlagGems | Beta | -| PPU | `ppu` | 与 NVIDIA CUDA 相同的 CUDA boxing 路径,针对 PPU CUDA 13 兼容 SDK,自带 libtorch | 实验性 | -| 海光 DCU | `dcu` | 基于 hipify 的 DTK torch 构建的 CUDA boxing | Beta | -| 燧原 GCU | `gcu` | 原生 `libtopsaten.so` 算子后端,未路由及 int64/float64 算子使用 CPU 回退 | Beta | -| 摩尔线程 MUSA | `musa` | FlagGems 优先的 Triton 内核,原生 `mudnn` 回退,未路由算子使用 CPU 回退 | 实验性 | -| 地平线 BPU | `bpu` | 不构建 eager 内核集合;eager 算子在 CPU 上执行,加速来自图编译路径 | 仅运行时 | -| 清微智能 | `tsingmicro` | 已提供运行时/构建选择器,尚无逐算子内核集合 | 仅运行时 | - -## 通用要求 - -所有平台都需要: - -- **Python**:3.8 或更高版本(平台 SDK 与可用 wheel 可能要求更窄的范围) -- **PyTorch**:2.10.x(`>=2.10,<2.11`)— 生成的 ATen 绑定与该次版本线绑定 -- **CMake**:3.18 或更高版本 -- **C++ 工具链**:可用的 C++17 编译器(GCC 7+、Clang 5+ 或 MSVC 2017+) -- **平台 SDK/运行时**:对应加速器的厂商 SDK、编译器与运行时库 - -同一 PyTorch 次版本线内的补丁版本(例如 2.10.0 → 2.10.1)互相兼容;跨次版本(例如 2.11.x)会在构建或运行时失败,因为生成的绑定对 C++ ABI 与算子 schema 变化敏感。 - -## 源码安装约定 - -各平台的安装遵循同一模式: - -```bash -FLAGOS_ACCELERATOR= pip install --no-build-isolation -e . -``` - -`--no-build-isolation` 是必需的:否则 pip 会创建隔离的构建环境,看不到当前环境中的 PyTorch 与平台 SDK,生成的本地绑定就会链接到错误的 torch,或找不到厂商 SDK。 - -上游仓库中每个平台目录(`docs/vendors//installation.md`)给出该平台的 SDK 环境变量与额外构建开关。各平台要点如下: - -### NVIDIA CUDA - -```bash -git clone https://github.com/flagos-ai/Torch-FL.git && cd Torch-FL - -pip install torch==2.10.0+cpu --index-url https://download.pytorch.org/whl/cpu - -FLAGOS_ACCELERATOR=cuda pip install --no-build-isolation -vvv -e . -``` - -该构建根据 PyTorch 的 ATen schema 生成 CUDA boxing 内核,把 `libtorch_cuda.so` 及相关 CUDA 调度库打包到 `torch_fl/lib/`,并固定与之匹配的 `nvidia-*-cu12` 运行期依赖。要求:计算能力 7.0 及以上的 NVIDIA GPU、470 及以上的驱动、用于构建的 CUDA 12.x 工具包,以及 `cmake`、`ninja`、`patchelf`。 - -可选的 FlagGems C++ 调度路径(开销最低的 FlagGems 路由): - -```bash -FLAGOS_ACCELERATOR=cuda \ - FLAGOS_BUILD_FLAGGEMS_CPP=1 \ - FLAGGEMS_DIR=/lib/cmake/FlagGems \ - pip install --no-build-isolation -vvv -e . -``` - -### MetaX - -MetaX 使用自包含的 boxing wheel:CUDA boxing 内核用宿主 `g++` 编译,MetaX 分支版 libtorch C++ 运行时直接打包进 wheel。目标机器只需要官方 `torch==2.10.0+cpu` wheel、`torch_fl` wheel 以及 `/opt/maca` 驱动运行时。 - -wheel 在具备完整 MACA SDK 与 `torch+metax` wheel(两者均来自 MetaX 开发者门户)的机器上构建:先生成 boxing 产物,再用 `scripts/vendor/bundle_maca_libtorch.sh` 打包分支版 libtorch,最后重新打包 wheel。产物体积较大(打包的 libtorch 使其超过 PyPI 的体积限制),通过私有源或直接传输分发。 - -目标主机上: - -```bash -pip install torch==2.10.0+cpu --index-url https://download.pytorch.org/whl/cpu -pip install torch_fl-+.whl -``` - -`FLAGOS_WHEEL_LOCAL` 会把目标 SDK 写入 wheel 的本地版本号(例如 `0.1.0+metax3.8.1`),避免两个 SDK 不兼容的 wheel 仅凭文件名无法区分。 - -**MetaX 上导入顺序很重要**:必须先导入 `torch_fl` 再 `import torch`。PyTorch 自带的 CUDA 12.x 运行时与 MACA 的 `cu-bridge` ABI 不兼容,`torch_fl` 会预加载一个提供所需符号版本的 shim。 - -### Ascend - -```bash -pip install torch==2.10.0 --index-url https://download.pytorch.org/whl/cpu -source /usr/local/Ascend/ascend-toolkit/set_env.sh - -FLAGOS_ACCELERATOR=ascend pip install --no-build-isolation -v -e . -``` - -要求:Ascend 910 与 CANN 9.0.0 或兼容版本、可访问的 `/dev/davinci*` 设备节点;若使用 FlagTree Ascend 3.5 wheel(仅提供 cp311),需要 Python 3.11,纯 ACLNN 构建则 Python 3.8+ 即可。 - -默认配置为已验证算子启用 FlagGems Python 路由,并以原生 ACLNN 内核作为回退。`scripts/codegen/codegen_ascend.py` 生成 ACLNN 内核;没有 ACLNN 映射的算子回退到 CPU。 - -Ascend 上的 FlagGems 运行在 **FlagTree**(FlagOS 的 Triton 分支)的 Ascend 3.5 线上。FlagTree 安装的模块名是 `triton`,因此需先移除官方或厂商 Triton,再从 FlagOS 源安装 FlagTree 与 FlagGems。该路径不会导入或链接 `torch_npu`:后端会安装轻量的 `torch_npu` stub,并在 FlagTree 的策略注册表上注册自己的 `flagos` 策略。 - -**Ascend 上导入顺序同样重要**:在可能注册设备后端的其他包之前导入 `torch_fl`。 - -### PPU - -PPU 表现为 CUDA 兼容设备:其 torch wheel 是完整的 CUDA 13 构建,并在 `CUDA` dispatch key 下注册算子,因此不需要官方 `+cpu` wheel,也不需要外部 `libtorch_cuda.so`。 - -```bash -FLAGOS_ACCELERATOR=ppu \ - CUDA_HOME=/usr/local/PPU_SDK/CUDA_SDK \ - FLAGOS_BUILD_FLAGGEMS_CPP=OFF \ - FLAGOS_BUILD_FLAGGEMS=OFF \ - FLAGOS_SKIP_CUDA_ASSETS=1 \ - pip install --no-build-isolation -vvv -e . -``` - -运行期需导出 `FLAGOS_DISABLE_CUDA_ASSETS=1`,使导入期对打包 `libtorch_cuda.so` 的预加载成为空操作(PPU torch 自带该运行时)。 - -### 海光 DCU - -```bash -source /opt/dtk/env.sh - -FLAGOS_ACCELERATOR=dcu pip install --no-build-isolation -vvv -e . -``` - -该构建为纯 boxing:DTK torch wheel 把 HIP 内核注册在 `CUDA` dispatch key 下,生成的 boxing 内核原样分发到 `libtorch_hip.so`,运行时源码用宿主 `g++` 直接编译 —— 不需要 `nvcc`、`hipcc` 或 hipify。FlagGems Python 路径默认开启;FlagGems C++ 路径保持关闭,因为 DTK 未提供 `liboperators.so`。 - -### 燧原 GCU - -```bash -pip install torch==2.10.0 --index-url https://download.pytorch.org/whl/cpu - -FLAGOS_ACCELERATOR=gcu pip install --no-build-isolation -v -e . -``` - -需要 TopsRider SDK(`libtopsrt.so` 运行时与 `libtopsaten.so` 算子库)。构建会运行 `scripts/codegen/codegen_gcu.py`,按已安装 `libtopsaten.so` 中实际存在的 `topsaten` 符号校验每个算子;SDK 中缺失的算子会给出告警并跳过。`torch-gcu` wheel 不能与 Torch-FL 同时使用,因为它会自行占用 `PrivateUse1`。 - -### 摩尔线程 MUSA - -```bash -pip install torch==2.10.0 --index-url https://download.pytorch.org/whl/cpu - -FLAGOS_ACCELERATOR=musa pip install --no-build-isolation -v -e . -``` - -需要 `/usr/local/musa` 下的 MUSA 工具包:`musart`(运行时)、`mudnn`(算子库)与 `murand`(设备随机数)。构建会运行 `scripts/codegen/codegen_mudnn.py`;覆盖范围为生成的算子集合加手写卷积内核,另有原生 RNG 内核,其余算子走 CPU 回退。此处 `--no-build-isolation` 是强制要求:否则 pip 的构建叠加层会解析出自己的 torch,导致 `import torch_fl` 因未定义的 `c10` 符号而失败。 - -### 地平线 BPU - -```bash -pip install torch==2.10.0+cpu --index-url https://download.pytorch.org/whl/cpu -FLAGOS_ACCELERATOR=bpu pip install --no-build-isolation -e . -``` - -BPU 平台只提供运行时加速:不存在逐算子 BPU 内核,eager 算子通过回退在 CPU 上执行,加速来自整图编译(`torch.compile(backend="bpu")`)或预构建 HBM 的 LLM 运行时。图编译需要 `hbdk4`,其 wheel 仅提供 x86_64 版本。 - -## 构建期开关 - -`setup.py` 会为内核集合开关强制指定各平台取值,并拒绝与之矛盾的环境变量显式取值: - -| 变量 | 默认值 | 作用 | -|---|---|---| -| `FLAGOS_ACCELERATOR` | `cuda` | wheel 面向的硬件平台 | -| `FLAGOS_BUILD_VENDOR` | `ON`,`metax` 上为 `OFF` | 编译厂商原生内核(厂商未提供时为空操作) | -| `FLAGOS_BUILD_FLAGGEMS` | `ON`,`bpu` 上为 `OFF` | 编译 FlagGems Python 内核包装 | -| `FLAGOS_BUILD_FLAGGEMS_CPP` | `cuda`/`tsingmicro` 上为 `ON` | 编译 FlagGems C++ 包装(`liboperators.so`) | -| `FLAGOS_BUILD_BOXING` | `ON`,`ascend`/`gcu`/`musa` 上为 `OFF` | 编译生成的 CUDA boxing 内核 | -| `FLAGOS_BUILD_TILEOPS` | `cuda` 上为 `ON` | 编译 TileOps 内核包装(NVIDIA SM90 芯片) | -| `FLAGOS_BUILD_JOBS` | CPU 核数 | CMake 构建并行任务数 | -| `FLAGOS_SKIP_CUDA_ASSETS` | `0` | 不打包外部 `libtorch_cuda.so` | -| `FLAGOS_WHEEL_LOCAL` | 由 SDK 推导 | wheel 的本地版本标记 | - -wheel 会把构建时使用的加速器与内核集合记录在 `torch_fl/_build_config.py` 中,运行期所有读取方都以该记录为准 —— 环境中残留的旧变量不会让 wheel 错误地描述自身。完整变量说明见 {doc}`环境变量参考 <../reference/environment-variables>`。 - -## 安装校验 - -```bash -python -c " -import torch_fl -import torch - -print(f'PyTorch version: {torch.__version__}') -print(f'flagos devices: {torch.flagos.device_count()}') -print(f'flagos available: {torch.flagos.is_available()}') - -x = torch.randn(4, 4, device='flagos:0') -y = (x @ x).sum() -print(f'Sample result: {y.cpu().item():.4f}') -" -``` - -安装正常时会输出当前机器的设备数量与一个浮点结果。若设备数量为 0,请检查厂商 SDK 安装、驱动、设备节点以及上文的导入顺序规则。 - -## 测试 - -```bash -# 单元测试:不依赖硬件 -pytest tests/unit -q - -# 当前后端上的算子正确性 -pytest tests/integration/ops/ -m main_ops -v --tb=short - -# 工厂算子是否遵循设备放置 -pytest tests/integration/test_factory_ops.py -v --tb=short -``` - -算子测试通过 `tests/integration/ops/conftest.py` 中注册的标记筛选: - -| 标记 | 含义 | -|---|---| -| `main_ops` | CI 冒烟子集中的代表性算子 | -| `anyplatform` | 可在任意加速器后端运行 | -| `cuda`、`metax`、`ascend`、`musa` | 需要对应后端的内核或硬件 | -| `flaggems` | 校验 `backends_.conf` 中的 FlagGems 路由 | -| `flaggems_python` | 需要 FlagGems Python 包装后端 | -| `flaggems_cpp` | 需要以 `FLAGOS_BUILD_FLAGGEMS_CPP=ON` 构建的 wheel | - -跨后端契约测试(profiler、AMP)通过 `tests/integration/conftest.py` 中的 `profiler*` 与 `amp*` 标记筛选;当当前平台不提供某项能力时,测试会带平台名的原因跳过,而不是伪造通过。 - -测试筛选是自动的:conftest 从 wheel 的构建记录、已安装平台标记或路由表名称探测平台,并跳过为不可用后端标记的测试。单元测试、模型测试与手动套件详见上游 `docs/development/testing.md`。 - -## 下一步 - -- {doc}`兼容性矩阵 <../reference/compatibility>` — 分平台能力验证 -- {doc}`环境变量 <../reference/environment-variables>` — 构建与运行期配置 -- {doc}`分布式集合通信 <../architecture/distributed>` — `ProcessGroupFlagOS` 与 FlagCX -- {doc}`Profiler <../architecture/profiler>` — `torch.profiler` 集成 -- {doc}`torch.compile <../architecture/torch-compile>` — Inductor 与 FlagTree 集成 +# 安装 + +## 选择平台 + +每个平台由一个构建变量 `FLAGOS_ACCELERATOR` 选择,并且各自有独立的执行路径: + +| 平台 | `FLAGOS_ACCELERATOR` | 执行路径 | 状态 | +|---|---|---|---| +| NVIDIA CUDA | `cuda`(默认) | 基于外部 `libtorch_cuda.so` 的 CUDA boxing | 稳定 | +| MetaX | `metax` | 通过 `cu-bridge` 对厂商 libtorch 进行 CUDA boxing | 稳定 | +| Ascend | `ascend` | 原生 ACLNN 算子后端,通过 FlagTree(Triton 3.5)使用 FlagGems | Beta | +| PPU | `ppu` | 与 NVIDIA CUDA 相同的 CUDA boxing 路径,针对 PPU CUDA 13 兼容 SDK,自带 libtorch | 实验性 | +| 海光 DCU | `dcu` | 基于 hipify 的 DTK torch 构建的 CUDA boxing | Beta | +| 燧原 GCU | `gcu` | 原生 `libtopsaten.so` 算子后端,未路由及 int64/float64 算子使用 CPU 回退 | Beta | +| 摩尔线程 MUSA | `musa` | FlagGems 优先的 Triton 内核,原生 `mudnn` 回退,未路由算子使用 CPU 回退 | 实验性 | +| 地平线 BPU | `bpu` | 不构建 eager 内核集合;eager 算子在 CPU 上执行,加速来自图编译路径 | 仅运行时 | +| 清微智能 | `tsingmicro` | 已提供运行时/构建选择器,尚无逐算子内核集合 | 仅运行时 | + +## 通用要求 + +所有平台都需要: + +- **Python**:每个平台只有一个解释器版本,不是区间 — CUDA/GCU/MetaX/PPU 为 3.12,DCU/MUSA 为 3.10,Ascend 为 3.11。FlagTree 只对单个 cp tag 发布,wheel 直接链接该版本,因此 `Requires-Python` 是单一版本 +- **PyTorch**:2.10.x(`>=2.10,<2.11`)— 生成的 ATen 绑定与该次版本线绑定 +- **CMake**:3.18 或更高版本 +- **C++ 工具链**:可用的 C++17 编译器(GCC 7+、Clang 5+ 或 MSVC 2017+) +- **平台 SDK/运行时**:对应加速器的厂商 SDK、编译器与运行时库 + +同一 PyTorch 次版本线内的补丁版本(例如 2.10.0 → 2.10.1)互相兼容;跨次版本(例如 2.11.x)会在构建或运行时失败,因为生成的绑定对 C++ ABI 与算子 schema 变化敏感。 + +## 源码安装约定 + +各平台的安装遵循同一模式: + +```bash +FLAGOS_ACCELERATOR= pip install --no-build-isolation -e . +``` + +`--no-build-isolation` 是必需的:否则 pip 会创建隔离的构建环境,看不到当前环境中的 PyTorch 与平台 SDK,生成的本地绑定就会链接到错误的 torch,或找不到厂商 SDK。 + +上游仓库中每个平台目录(`docs/vendors//installation.md`)给出该平台的 SDK 环境变量与额外构建开关。各平台要点如下: + +### NVIDIA CUDA + +```bash +git clone https://github.com/flagos-ai/Torch-FL.git && cd Torch-FL + +pip install torch==2.10.0+cpu --index-url https://download.pytorch.org/whl/cpu + +FLAGOS_ACCELERATOR=cuda pip install --no-build-isolation -vvv -e . +``` + +该构建根据 PyTorch 的 ATen schema 生成 CUDA boxing 内核,把 `libtorch_cuda.so` 及相关 CUDA 调度库打包到 `torch_fl/lib/`,并固定与之匹配的 `nvidia-*-cu12` 运行期依赖。要求:计算能力 7.0 及以上的 NVIDIA GPU;驱动需足够新以支持 wheel 打包的运行时(cu12.8 的 `libtorch_cuda.so` 及匹配的 `nvidia-*-cu12` wheels);构建需要提供 `nvcc` 与头文件的 CUDA 工具包;以及 `cmake`、`ninja`、`patchelf`。 + +可选的 FlagGems C++ 调度路径(开销最低的 FlagGems 路由): + +```bash +FLAGOS_ACCELERATOR=cuda \ + FLAGOS_BUILD_FLAGGEMS_CPP=1 \ + FLAGGEMS_DIR=/lib/cmake/FlagGems \ + pip install --no-build-isolation -vvv -e . +``` + +### MetaX + +MetaX 使用自包含的 boxing wheel:CUDA boxing 内核用宿主 `g++` 编译,MetaX 分支版 libtorch C++ 运行时直接打包进 wheel。目标机器只需要官方 `torch==2.10.0+cpu` wheel、`torch_fl` wheel 以及 `/opt/maca` 驱动运行时。 + +wheel 在具备完整 MACA SDK 与 `torch+metax` wheel(两者均来自 MetaX 开发者门户)的机器上构建:先生成 boxing 产物,再用 `scripts/vendor/bundle_maca_libtorch.sh` 打包分支版 libtorch,最后重新打包 wheel。产物体积较大(打包的 libtorch 使其超过 PyPI 的体积限制),通过私有源或直接传输分发。 + +目标主机上: + +```bash +pip install torch==2.10.0+cpu --index-url https://download.pytorch.org/whl/cpu +pip install torch_fl-+.whl +``` + +`FLAGOS_WHEEL_LOCAL` 会把目标 SDK 写入 wheel 的本地版本号(例如 `2.10.0+maca3.8.1.3`),避免两个 SDK 不兼容的 wheel 仅凭文件名无法区分。 + +**MetaX 上导入顺序很重要**:必须先导入 `torch_fl` 再 `import torch`。PyTorch 自带的 CUDA 12.x 运行时与 MACA 的 `cu-bridge` ABI 不兼容,`torch_fl` 会预加载一个提供所需符号版本的 shim。 + +### Ascend + +```bash +pip install torch==2.10.0 --index-url https://download.pytorch.org/whl/cpu +source /usr/local/Ascend/ascend-toolkit/set_env.sh + +FLAGOS_ACCELERATOR=ascend pip install --no-build-isolation -v -e . +``` + +要求:Ascend 910 与 CANN 9.0.0 或兼容版本、可访问的 `/dev/davinci*` 设备节点;且需要 Python 3.11 —— FlagTree Ascend wheel 仅提供 cp311,wheel 的 `Requires-Python` 即该单一版本。 + +默认配置为已验证算子启用 FlagGems Python 路由,并以原生 ACLNN 内核作为回退。`scripts/codegen/codegen_ascend.py` 生成 ACLNN 内核;没有 ACLNN 映射的算子回退到 CPU。 + +Ascend 上的 FlagGems 运行在 **FlagTree**(FlagOS 的 Triton 分支)的 Ascend 3.5 线上。FlagTree 安装的模块名是 `triton`,因此需先移除官方或厂商 Triton,再从 FlagOS 源安装 FlagTree 与 FlagGems。该路径不会导入或链接 `torch_npu`:后端会安装轻量的 `torch_npu` stub,并在 FlagTree 的策略注册表上注册自己的 `flagos` 策略。 + +**Ascend 上导入顺序同样重要**:在可能注册设备后端的其他包之前导入 `torch_fl`。 + +### PPU + +PPU 表现为 CUDA 兼容设备:其 torch wheel 是完整的 CUDA 13 构建,并在 `CUDA` dispatch key 下注册算子,因此不需要官方 `+cpu` wheel,也不需要外部 `libtorch_cuda.so`。 + +```bash +FLAGOS_ACCELERATOR=ppu \ + CUDA_HOME=/usr/local/PPU_SDK/CUDA_SDK \ + FLAGOS_BUILD_FLAGGEMS_CPP=OFF \ + FLAGOS_BUILD_FLAGGEMS=OFF \ + FLAGOS_SKIP_CUDA_ASSETS=1 \ + pip install --no-build-isolation -vvv -e . +``` + +运行期需导出 `FLAGOS_DISABLE_CUDA_ASSETS=1`,使导入期对打包 `libtorch_cuda.so` 的预加载成为空操作(PPU torch 自带该运行时)。 + +### 海光 DCU + +```bash +source /opt/dtk/env.sh + +FLAGOS_ACCELERATOR=dcu pip install --no-build-isolation -vvv -e . +``` + +该构建为纯 boxing:DTK torch wheel 把 HIP 内核注册在 `CUDA` dispatch key 下,生成的 boxing 内核原样分发到 `libtorch_hip.so`,运行时源码用宿主 `g++` 直接编译 —— 不需要 `nvcc`、`hipcc` 或 hipify。FlagGems Python 路径默认开启;FlagGems C++ 路径保持关闭,因为 DTK 未提供 `liboperators.so`。 + +### 燧原 GCU + +```bash +pip install torch==2.10.0 --index-url https://download.pytorch.org/whl/cpu + +FLAGOS_ACCELERATOR=gcu pip install --no-build-isolation -v -e . +``` + +需要 TopsRider SDK(`libtopsrt.so` 运行时与 `libtopsaten.so` 算子库)。构建会运行 `scripts/codegen/codegen_gcu.py`,按已安装 `libtopsaten.so` 中实际存在的 `topsaten` 符号校验每个算子;SDK 中缺失的算子会给出告警并跳过。`torch-gcu` wheel 不能与 Torch-FL 同时使用,因为它会自行占用 `PrivateUse1`。 + +### 摩尔线程 MUSA + +```bash +pip install torch==2.10.0 --index-url https://download.pytorch.org/whl/cpu + +FLAGOS_ACCELERATOR=musa pip install --no-build-isolation -v -e . +``` + +需要 `/usr/local/musa` 下的 MUSA 工具包:`musart`(运行时)、`mudnn`(算子库)与 `murand`(设备随机数)。构建会运行 `scripts/codegen/codegen_mudnn.py`;覆盖范围为生成的算子集合加手写卷积内核,另有原生 RNG 内核,其余算子走 CPU 回退。此处 `--no-build-isolation` 是强制要求:否则 pip 的构建叠加层会解析出自己的 torch,导致 `import torch_fl` 因未定义的 `c10` 符号而失败。 + +### 地平线 BPU + +```bash +pip install torch==2.10.0+cpu --index-url https://download.pytorch.org/whl/cpu +FLAGOS_ACCELERATOR=bpu pip install --no-build-isolation -e . +``` + +BPU 平台只提供运行时加速:不存在逐算子 BPU 内核,eager 算子通过回退在 CPU 上执行,加速来自整图编译(`torch.compile(backend="bpu")`)或预构建 HBM 的 LLM 运行时。图编译需要 `hbdk4`,其 wheel 仅提供 x86_64 版本。 + +## 运行期依赖与包索引 + +wheel 会声明三个本仓库之外构建、且缺一不可的包 — **FlagTree**(携带厂商后端的 Triton 构建)、**FlagGems**(算子来源)与 **FlagCX**(分布式后端)— 并按构建时所用版本精确固定,来源为 `.github/version-pins.env`。这三者都不是版本区间:`flagtree` 与 `flagcx` 根本不在 PyPI 上,FlagTree 构建是分平台的(包名携带厂商的 Triton 后端),而 PyPI 上的 `flag_gems` 版本比 `torch_fl/configs/backends_*.conf` 中逐算子路由表所依据的那批版本更旧。 + +因此索引必须包含多个位置: + +| 依赖 | 发布位置 | +|---|---| +| `torch_fl` | `flagos-pypi-` —— 由 wheel 本地版本号指定(`2.10.0+hygon` → `flagos-pypi-hygon`) | +| `flag_gems`、`flagcx` | 同一个厂商 lane | +| `flagtree` | `flagos-pypi-hosted`,所有平台通用 | +| `torch==2.10.0+cpu` | `https://download.pytorch.org/whl/cpu` | +| 其余(`packaging`、`PyYAML`、`numpy` 等) | PyPI(或镜像) | + +因此单个 `--index-url` 必须指向一个**包含上述全部内容的 group 仓库**。若没有配置这样的仓库,就逐条列出 —— DCU 就是这种情况: + +```bash +BASE=https://resource.flagos.net/repository +pip install \ + --index-url "$BASE/flagos-pypi-hygon/simple/" \ + --extra-index-url "$BASE/flagos-pypi-hosted/simple/" \ + --extra-index-url "$BASE/pypi-proxy/simple/" \ + --extra-index-url "https://download.pytorch.org/whl/cpu" \ + torch_fl==2.10.0+hygon +``` + +这里有两个容易出错的点:即使厂商 lane 已经包含全部三个 FlagOS 包,只用该 lane 也不够(`flag_gems` 自身声明了 `packaging>=26.0` 与 `PyYAML==6.0.1`,lane 不提供);而且 `flagtree` 不在大多数 lane 中 —— 所有平台都从 `flagos-pypi-hosted` 获取。 + +## 构建期开关 + +`setup.py` 会为内核集合开关强制指定各平台取值,并拒绝与之矛盾的环境变量显式取值: + +| 变量 | 默认值 | 作用 | +|---|---|---| +| `FLAGOS_ACCELERATOR` | `cuda` | wheel 面向的硬件平台 | +| `FLAGOS_BUILD_VENDOR` | `ON`,`metax` 上为 `OFF` | 编译厂商原生内核(厂商未提供时为空操作) | +| `FLAGOS_BUILD_FLAGGEMS` | `ON`,`bpu` 上为 `OFF` | 编译 FlagGems Python 内核包装 | +| `FLAGOS_BUILD_FLAGGEMS_CPP` | `cuda`/`tsingmicro` 上为 `ON` | 编译 FlagGems C++ 包装(`liboperators.so`) | +| `FLAGOS_BUILD_BOXING` | `ON`,`ascend`/`gcu`/`musa` 上为 `OFF` | 编译生成的 CUDA boxing 内核 | +| `FLAGOS_BUILD_TILEOPS` | `cuda` 上为 `ON` | 编译 TileOps 内核包装(NVIDIA SM90 芯片) | +| `FLAGOS_BUILD_JOBS` | CPU 核数 | CMake 构建并行任务数 | +| `FLAGOS_SKIP_CUDA_ASSETS` | `0` | 不打包外部 `libtorch_cuda.so` | +| `FLAGOS_WHEEL_LOCAL` | 由 SDK 推导 | wheel 的本地版本标记 | + +wheel 会把构建时使用的加速器与内核集合记录在 `torch_fl/_build_config.py` 中,运行期所有读取方都以该记录为准 —— 环境中残留的旧变量不会让 wheel 错误地描述自身。完整变量说明见 {doc}`环境变量参考 <../reference/environment-variables>`。 + +## 安装校验 + +```bash +python -c " +import torch_fl +import torch + +print(f'PyTorch version: {torch.__version__}') +print(f'flagos devices: {torch.flagos.device_count()}') +print(f'flagos available: {torch.flagos.is_available()}') + +x = torch.randn(4, 4, device='flagos:0') +y = (x @ x).sum() +print(f'Sample result: {y.cpu().item():.4f}') +" +``` + +安装正常时会输出当前机器的设备数量与一个浮点结果。若设备数量为 0,请检查厂商 SDK 安装、驱动、设备节点以及上文的导入顺序规则。 + +## 测试 + +```bash +# 单元测试:不依赖硬件 +pytest tests/unit -q + +# 当前后端上的算子正确性 +pytest tests/integration/ops/ -m main_ops -v --tb=short + +# 工厂算子是否遵循设备放置 +pytest tests/integration/test_factory_ops.py -v --tb=short +``` + +算子测试通过 `tests/integration/ops/conftest.py` 中注册的标记筛选: + +| 标记 | 含义 | +|---|---| +| `main_ops` | 冒烟子集中的代表性算子 | +| `anyplatform` | 可在任意加速器后端运行 | +| `cuda`、`metax`、`ascend`、`musa` | 需要对应后端的内核或硬件 | +| `flaggems` | 校验 `backends_.conf` 中的 FlagGems 路由 | +| `flaggems_python` | 需要 FlagGems Python 包装后端 | +| `flaggems_cpp` | 需要以 `FLAGOS_BUILD_FLAGGEMS_CPP=ON` 构建的 wheel | + +跨后端契约测试(profiler、AMP)通过 `tests/integration/conftest.py` 中的 `profiler*` 与 `amp*` 标记筛选;当当前平台不提供某项能力时,测试会带平台名的原因跳过,而不是伪造通过。 + +测试筛选是自动的:conftest 从 wheel 的构建记录、已安装平台标记或路由表名称探测平台,并跳过为不可用后端标记的测试。单元测试、模型测试与手动套件详见上游 `docs/development/testing.md`。 + +## 下一步 + +- {doc}`快速开始 ` — 与平台无关的用法 +- {doc}`兼容性矩阵 <../reference/compatibility>` — 分平台能力验证 +- {doc}`平台能力矩阵 <../reference/platform-capability>` — 各加速器构建与路由了什么 +- {doc}`数据类型支持 <../reference/dtype-support>` — 存储、AMP 目标与回退边界 +- {doc}`常见故障排查 <../reference/troubleshooting>` — 设备数为 0、库缺失、编译器报错 +- {doc}`环境变量 <../reference/environment-variables>` — 构建与运行期配置 +- {doc}`分布式集合通信 <../architecture/distributed>` — `ProcessGroupFlagOS` 与 FlagCX +- {doc}`Profiler <../architecture/profiler>` — `torch.profiler` 集成 +- {doc}`torch.compile <../architecture/torch-compile>` — Inductor 与 FlagTree 集成 diff --git a/docs/torch_fl_zh/getting_started/quickstart.md b/docs/torch_fl_zh/getting_started/quickstart.md new file mode 100644 index 0000000000..4667156b37 --- /dev/null +++ b/docs/torch_fl_zh/getting_started/quickstart.md @@ -0,0 +1,96 @@ +# 快速开始 + +本页给出与平台无关的用法。按你的平台完成安装后(见 {doc}`安装 `),同样的代码可在所有受支持的加速器上运行。 + +## 基本用法 + +```python +import torch +import torch_fl + +x = torch.randn(4, 4, device="flagos:0") +y = torch.relu(x @ x) +print(y.cpu()) +``` + +张量创建在第 0 个 `flagos` 设备上,矩阵乘与激活按平台路由到相应内核,结果再拷贝回 CPU 打印。 + +## 在设备之间搬运张量 + +```python +import torch +import torch_fl + +x = torch.randn(4, 4) # CPU 张量 +x_flagos = x.to("flagos") # 第 0 个设备 +x_flagos_1 = x.to("flagos:1") # 第 1 个设备 + +y = torch.randn(4, 4, device="flagos") +y_cpu = y.cpu() # 拷贝回 CPU +``` + +## 选择设备 + +`torch.flagos.device()` 设置当前设备上下文,之后 `device="flagos"` 的张量会落在该设备上: + +```python +import torch +import torch_fl + +with torch.flagos.device(0): + x = torch.randn(4, 4, device="flagos") + +with torch.flagos.device(1): + y = torch.randn(4, 4, device="flagos") +``` + +## 同步 + +与其他 PyTorch 设备一样,内核启动是异步的: + +```python +import torch +import torch_fl + +x = torch.randn(1000, 1000, device="flagos") +y = x @ x # 入队,未必已执行完 +torch.flagos.synchronize() # 等待设备排空 +``` + +## 设备查询 + +```python +import torch +import torch_fl + +if torch.flagos.is_available(): + print(f"Found {torch.flagos.device_count()} device(s)") + print(f"Current device: {torch.flagos.current_device()}") + + props = torch.flagos.get_device_properties(0) + print(f"Device name: {props.name}") + print(f"Total memory: {props.total_memory / 1024**3:.2f} GB") +else: + print("No flagos devices available") +``` + +若驱动正常但 `device_count()` 返回 0,见 {doc}`常见故障排查 <../reference/troubleshooting>`。 + +## 算子路由 + +`flagos` 张量上的运算按**算子**分发,而不是按设备或模型: + +- **可移植编译内核** — 平台路由启用时的 FlagGems Triton 内核 +- **厂商原生内核** — 厂商算子库(ACLNN、topsaten、mudnn 等) +- **兼容性 boxing** — 生成的内核,转发到外部厂商 `libtorch`(CUDA、MetaX、PPU、DCU) +- **CPU 回退** — 对没有设备内核的算子使用以实现正确性为先的 CPU 实现,再拷回设备 + +路由对代码透明:切换平台或内核来源时无需改动代码。要查看某次调用实际由哪个后端服务,设置 `FLAGOS_LOG=dispatch`(见{doc}`环境变量参考 <../reference/environment-variables>`)。 + +## 下一步 + +- {doc}`安装 ` — 分平台构建与验证 +- {doc}`平台能力矩阵 <../reference/platform-capability>` — 各加速器支持范围 +- {doc}`数据类型支持 <../reference/dtype-support>` — 存储、AMP 目标与回退边界 +- {doc}`分布式集合通信 <../architecture/distributed>` — `ProcessGroupFlagOS`、DDP 与 FSDP2 +- {doc}`性能分析 <../architecture/profiler>` — `torch.profiler` 集成 diff --git a/docs/torch_fl_zh/index.md b/docs/torch_fl_zh/index.md index 342a7dde18..157e4b13e6 100644 --- a/docs/torch_fl_zh/index.md +++ b/docs/torch_fl_zh/index.md @@ -1,102 +1,100 @@ -# Torch-FL - -Torch-FL(`torch_fl`)是面向 FlagOS 软件栈的 PyTorch 设备插件。它对外提供统一的 `flagos` 设备,并在可复用原生内核、可移植编译器内核、厂商原生实现和显式 CPU 回退之间路由算子,使同一份 PyTorch 程序无需修改即可运行在不同加速器上。 - -![Torch-FL 架构](assets/images/torch-fl.png) - -::::{grid} 1 2 2 3 -:gutter: 1 1 1 2 - -:::{grid-item-card} {octicon}`browser;1.5em;sd-mr-1` 概览 -:link: overview/overview -:link-type: doc - -Torch-FL 是什么、设计理念、能力范围与组件架构。 - -+++ -[了解更多 »](overview/overview.md) -::: - -:::{grid-item-card} {octicon}`book;1.5em;sd-mr-1` 快速入门 -:link: getting_started/installation -:link-type: doc - -平台选择、构建要求、环境配置、安装校验与测试标记。 - -+++ -[了解更多 »](getting_started/installation.md) -::: - -:::{grid-item-card} {octicon}`broadcast;1.5em;sd-mr-1` 分布式 -:link: architecture/distributed -:link-type: doc - -`ProcessGroupFlagOS`、FlagCX、厂商回退后端,以及 DDP 与 DataParallel 支持。 - -+++ -[了解更多 »](architecture/distributed.md) -::: - -:::{grid-item-card} {octicon}`pulse;1.5em;sd-mr-1` Profiler -:link: architecture/profiler -:link-type: doc - -`torch.profiler` 集成、厂商追踪器、correlation id,以及与 `torch.cuda` 的对等性。 - -+++ -[了解更多 »](architecture/profiler.md) -::: - -:::{grid-item-card} {octicon}`zap;1.5em;sd-mr-1` torch.compile -:link: architecture/torch-compile -:link-type: doc - -`flagos` 设备的 Inductor 集成、FlagTree 编译与各平台差异。 - -+++ -[了解更多 »](architecture/torch-compile.md) -::: - -:::{grid-item-card} {octicon}`gear;1.5em;sd-mr-1` 参考 -:link: reference/compatibility -:link-type: doc - -分平台能力验证、环境变量与算子路由配置。 - -+++ -[了解更多 »](reference/compatibility.md) -::: - -:::: - -- **代码仓库**:[flagos-ai/Torch-FL](https://github.com/flagos-ai/Torch-FL) -- **FlagGems**:[flagos-ai/FlagGems](https://github.com/flagos-ai/FlagGems) -- **FlagTree**:[flagos-ai/FlagTree](https://github.com/flagos-ai/FlagTree) -- **FlagCX**:[flagos-ai/FlagCX](https://github.com/flagos-ai/FlagCX) -- **许可证**:Apache License 2.0 - ---- - -```{toctree} -:caption: 📑 发布说明 -:maxdepth: 5 -:hidden: - -release_notes/release-notes.md -``` - -```{toctree} -:caption: 📚 指南 -:maxdepth: 5 -:hidden: - -overview/overview.md -overview/features.md -overview/architecture.md -getting_started/installation.md -architecture/distributed.md -architecture/profiler.md -architecture/torch-compile.md -reference/compatibility.md -reference/environment-variables.md -``` +# Torch-FL + +Torch-FL(`torch_fl`)是面向 FlagOS 软件栈的 PyTorch 设备插件。它对外提供统一的 `flagos` 设备,并在可复用原生内核、可移植编译器内核、厂商原生实现和显式 CPU 回退之间路由算子,使同一份 PyTorch 程序无需修改即可运行在不同加速器上。 + +![Torch-FL 架构](assets/images/torch-fl.png) + +::::{grid} 1 2 2 3 +:gutter: 1 1 1 2 + +:::{grid-item-card} {octicon}`browser;1.5em;sd-mr-1` 概览 +:link: overview/overview +:link-type: doc + +Torch-FL 是什么、设计理念、能力范围与组件架构。 + ++++ +[了解更多 »](overview/overview.md) +::: + +:::{grid-item-card} {octicon}`book;1.5em;sd-mr-1` 快速入门 +:link: getting_started/installation +:link-type: doc + +平台选择、构建要求、环境配置、安装校验与测试标记。 + ++++ +[了解更多 »](getting_started/installation.md) +::: + +:::{grid-item-card} {octicon}`broadcast;1.5em;sd-mr-1` 分布式 +:link: architecture/distributed +:link-type: doc + +`ProcessGroupFlagOS`、FlagCX、厂商回退后端,以及 DDP 与 DataParallel 支持。 + ++++ +[了解更多 »](architecture/distributed.md) +::: + +:::{grid-item-card} {octicon}`pulse;1.5em;sd-mr-1` Profiler +:link: architecture/profiler +:link-type: doc + +`torch.profiler` 集成、厂商追踪器、correlation id,以及与 `torch.cuda` 的对等性。 + ++++ +[了解更多 »](architecture/profiler.md) +::: + +:::{grid-item-card} {octicon}`zap;1.5em;sd-mr-1` torch.compile +:link: architecture/torch-compile +:link-type: doc + +`flagos` 设备的 Inductor 集成、FlagTree 编译与各平台差异。 + ++++ +[了解更多 »](architecture/torch-compile.md) +::: + +:::{grid-item-card} {octicon}`gear;1.5em;sd-mr-1` 参考 +:link: reference/compatibility +:link-type: doc + +分平台能力验证、环境变量与算子路由配置。 + ++++ +[了解更多 »](reference/compatibility.md) +::: + +:::: + +--- + +```{toctree} +:caption: 📑 发布说明 +:maxdepth: 5 +:hidden: + +release_notes/release-notes.md +``` + +```{toctree} +:caption: 📚 指南 +:maxdepth: 5 +:hidden: + +overview/overview.md +overview/features.md +overview/architecture.md +getting_started/installation.md +getting_started/quickstart.md +architecture/distributed.md +architecture/profiler.md +architecture/torch-compile.md +reference/compatibility.md +reference/environment-variables.md +reference/platform-capability.md +reference/dtype-support.md +reference/troubleshooting.md +``` diff --git a/docs/torch_fl_zh/overview/architecture.md b/docs/torch_fl_zh/overview/architecture.md index eb8e5ff291..85cd9eb470 100644 --- a/docs/torch_fl_zh/overview/architecture.md +++ b/docs/torch_fl_zh/overview/architecture.md @@ -1,76 +1,76 @@ -# 架构 - -Torch-FL 注册一个 PyTorch 设备,并对到达该设备的每个算子进行路由。 - -![Torch-FL 架构](../assets/images/torch-fl.png) - -```text -PyTorch API - | -flagos 设备(PrivateUse1) - | -设备运行时 + 按算子路由 - | -FlagGems/编译器内核 | 兼容性 boxing | 厂商原生内核 | CPU 回退 - | -加速器运行时 -``` - -## 设备注册 - -导入时,Torch-FL 接管 `PrivateUse1` dispatch key,并以 `flagos` 名称发布: - -1. `torch.utils.rename_privateuse1_backend("flagos")` 为 dispatch key 命名。 -2. `torch._register_device_module("flagos", flagos)` 安装设备模块,使 `torch.flagos.*` 可用。 -3. `torch.utils.generate_methods_for_privateuse1_backend(for_storage=True)` 生成 `.to("flagos")` 等张量与存储方法。 -4. 设备模块同时以 `torch_flagos` 名称发布到 `sys.modules`,以满足会触发惰性初始化的平台上 `torch::utils::device_lazy_init` 按模块名的查找。 - -原生扩展(`torch_fl._C`)在加载时注册 `AutogradPrivateUse1` 回退和算子实现。由于该注册发生在 `dlopen` 时,插件会在加载扩展**之前**检查 `PrivateUse1` 是否仍未被占用;若已被其他厂商插件接管,会给出可操作的错误信息,而不是无法捕获的中止。 - -## 导入期阶段 - -`import torch_fl` 会执行一串固定的副作用序列。这些步骤的顺序是关键的 —— 顺序错误会导致 `dlopen` 中止或选错厂商构建,而不是抛出 Python 异常 —— 因此集中在一处管理: - -| 阶段 | 作用 | -|---|---| -| 1. conf | 为本构建选择算子路由表;在启用时准备 MetaX `libcudart` shim | -| 2. preload | 在 `import torch` 之前选择并预加载厂商 libtorch 与 CUDA 资源 | -| 3. claim | 导入 `torch`、校验厂商运行时、接管 `PrivateUse1`、加载 `_C`、安装设备模块 | -| 4. vendor_compat | 安装厂商运行时 shim,并解析 FlagGems 厂商标识(`GEMS_VENDOR`) | -| 5. ecosystem | FlagGems 注册准备、CUDA 别名、分布式/DDP/DataParallel、编译后端 | - -调试导入问题时有两个约束值得注意:厂商 libtorch 必须在 `import torch` 之前就位;`PrivateUse1` 归属检查必须早于 `_C` 加载。 - -## 组件布局 - -| 路径 | 职责 | -|---|---| -| `torch_fl/__init__.py` | 导入期阶段流水线、设备模块安装、生态补丁 | -| `torch_fl/_env.py` | 所有 `FLAGOS_*` 变量的唯一注册表与读取入口 | -| `torch_fl/flagos/` | 设备模块:流、事件、随机数、AMP、内存、用于追踪的 meta 内核 | -| `torch_fl/configs/` | 分平台路由表 `backends_.conf` | -| `torch_fl/accelerator/` | 各厂商兼容 shim 与运行时适配 | -| `torch_fl/compile/` | Inductor 后端、FlagTree shim、平台 profile、Triton 守卫 | -| `torch_fl/comm/` | `ProcessGroupFlagOS` | -| `torch_fl/distributed.py` | `init_process_group`、`DistributedDataParallel`、缓冲区搬迁辅助函数 | -| `torch_fl/quantization/` | 低精度格式、转换与模块 | -| `torch_fl/tileops/` | TileOps 算子库的 Python 侧(惰性导入) | -| `torch_fl/compat/` | Apex 与 flex-attention 兼容层 | -| `csrc/aten/` | ATen 层:调度器、boxing、生成的绑定、厂商后端 | -| `csrc/runtime/` | 设备运行时:分配器、guard、generator、各加速器实现 | -| `csrc/profiler/` | 与厂商无关的设备追踪器接口及各厂商追踪器 | -| `csrc/include/flagos.h` | 统一运行时 ABI(内存、流、设备、当前流注册表) | -| `scripts/codegen/` | 算子绑定生成器(CUDA boxing、ACLNN、topsaten、mudnn、TileOps) | -| `scripts/tools/` | 预检与校验工具,包括 `torch-fl-preflight` | -| `tests/unit`、`tests/integration`、`tests/manual`、`tests/perf` | 单元、硬件集成、手动与基准测试套件 | - -## 算子调度 - -算子实现通过同一个 dispatch key 与同一张路由表到达: - -- `csrc/aten/generated/` 下的生成绑定提供 CUDA boxing 内核、FlagGems Python 与 C++ 调用方以及 TileOps stub,它们都注册到 `PrivateUse1`。 -- `csrc/aten/common.cc` 在首次调度时读取路由表:优先使用本构建选择的路由表,其次使用 `FLAGOS_BACKEND_CONFIG` 指定的文件。 -- 厂商原生内核位于 `csrc/aten/backends//` 下,按该厂商库实际导出的算子面生成;生成器无法匹配的算子在给出告警后被跳过,继续走路由或 CPU 回退。 -- 未路由的算子进入 `cpu_fallback`:执行 CPU 参考实现并把结果复制回设备。 - -`FLAGOS_LOG=dispatch` 会逐算子打印路由结果,`torch_fl.backend_config_path()` 返回实际使用的路由表。 +# 架构 + +Torch-FL 注册一个 PyTorch 设备,并对到达该设备的每个算子进行路由。 + +![Torch-FL 架构](../assets/images/torch-fl.png) + +```text +PyTorch API + | +flagos 设备(PrivateUse1) + | +设备运行时 + 按算子路由 + | +FlagGems/编译器内核 | 兼容性 boxing | 厂商原生内核 | CPU 回退 + | +加速器运行时 +``` + +## 设备注册 + +导入时,Torch-FL 接管 `PrivateUse1` dispatch key,并以 `flagos` 名称发布: + +1. `torch.utils.rename_privateuse1_backend("flagos")` 为 dispatch key 命名。 +2. `torch._register_device_module("flagos", flagos)` 安装设备模块,使 `torch.flagos.*` 可用。 +3. `torch.utils.generate_methods_for_privateuse1_backend(for_storage=True)` 生成 `.to("flagos")` 等张量与存储方法。 +4. 设备模块同时以 `torch_flagos` 名称发布到 `sys.modules`,以满足会触发惰性初始化的平台上 `torch::utils::device_lazy_init` 按模块名的查找。 + +原生扩展(`torch_fl._C`)在加载时注册 `AutogradPrivateUse1` 回退和算子实现。由于该注册发生在 `dlopen` 时,插件会在加载扩展**之前**检查 `PrivateUse1` 是否仍未被占用;若已被其他厂商插件接管,会给出可操作的错误信息,而不是无法捕获的中止。 + +## 导入期阶段 + +`import torch_fl` 会执行一串固定的副作用序列。这些步骤的顺序是关键的 —— 顺序错误会导致 `dlopen` 中止或选错厂商构建,而不是抛出 Python 异常 —— 因此集中在一处管理: + +| 阶段 | 作用 | +|---|---| +| 1. conf | 为本构建选择算子路由表;在启用时准备 MetaX `libcudart` shim | +| 2. preload | 在 `import torch` 之前选择并预加载厂商 libtorch 与 CUDA 资源 | +| 3. claim | 导入 `torch`、校验厂商运行时、接管 `PrivateUse1`、加载 `_C`、安装设备模块 | +| 4. vendor_compat | 安装厂商运行时 shim,并解析 FlagGems 厂商标识(`GEMS_VENDOR`) | +| 5. ecosystem | FlagGems 注册准备、CUDA 别名、分布式/DDP/DataParallel、编译后端 | + +调试导入问题时有两个约束值得注意:厂商 libtorch 必须在 `import torch` 之前就位;`PrivateUse1` 归属检查必须早于 `_C` 加载。 + +## 组件布局 + +| 路径 | 职责 | +|---|---| +| `torch_fl/__init__.py` | 导入期阶段流水线、设备模块安装、生态补丁 | +| `torch_fl/_env.py` | 所有 `FLAGOS_*` 变量的唯一注册表与读取入口 | +| `torch_fl/flagos/` | 设备模块:流、事件、随机数、AMP、内存、用于追踪的 meta 内核 | +| `torch_fl/configs/` | 分平台路由表 `backends_.conf` | +| `torch_fl/accelerator/` | 各厂商兼容 shim 与运行时适配 | +| `torch_fl/compile/` | Inductor 后端、FlagTree shim、平台 profile、Triton 守卫 | +| `torch_fl/comm/` | `ProcessGroupFlagOS` | +| `torch_fl/distributed.py` | `init_process_group`、`DistributedDataParallel`、缓冲区搬迁辅助函数 | +| `torch_fl/quantization/` | 低精度格式、转换与模块 | +| `torch_fl/tileops/` | TileOps 算子库的 Python 侧(惰性导入) | +| `torch_fl/compat/` | Apex 与 flex-attention 兼容层 | +| `csrc/aten/` | ATen 层:调度器、boxing、生成的绑定、厂商后端 | +| `csrc/runtime/` | 设备运行时:分配器、guard、generator、各加速器实现 | +| `csrc/profiler/` | 与厂商无关的设备追踪器接口及各厂商追踪器 | +| `csrc/include/flagos.h` | 统一运行时 ABI(内存、流、设备、当前流注册表) | +| `scripts/codegen/` | 算子绑定生成器(CUDA boxing、ACLNN、topsaten、mudnn、TileOps) | +| `scripts/tools/` | 预检与校验工具,包括 `torch-fl-preflight` | +| `tests/unit`、`tests/integration`、`tests/manual`、`tests/perf` | 单元、硬件集成、手动与基准测试套件 | + +## 算子调度 + +算子实现通过同一个 dispatch key 与同一张路由表到达: + +- `csrc/aten/generated/` 下的生成绑定提供 CUDA boxing 内核、FlagGems Python 与 C++ 调用方以及 TileOps stub,它们都注册到 `PrivateUse1`。 +- `csrc/aten/common.cc` 在首次调度时读取路由表:优先使用本构建选择的路由表,其次使用 `FLAGOS_BACKEND_CONFIG` 指定的文件。 +- 厂商原生内核位于 `csrc/aten/backends//` 下,按该厂商库实际导出的算子面生成;生成器无法匹配的算子在给出告警后被跳过,继续走路由或 CPU 回退。 +- 未路由的算子进入 `cpu_fallback`:执行 CPU 参考实现并把结果复制回设备。 + +`FLAGOS_LOG=dispatch` 会逐算子打印路由结果,`torch_fl.backend_config_path()` 返回实际使用的路由表。 diff --git a/docs/torch_fl_zh/overview/features.md b/docs/torch_fl_zh/overview/features.md index 21ba4c719f..c549b70074 100644 --- a/docs/torch_fl_zh/overview/features.md +++ b/docs/torch_fl_zh/overview/features.md @@ -1,98 +1,98 @@ -# 功能特性 - -## 统一设备与标准 PyTorch API - -所有已支持的加速器都通过 `flagos` 设备编程。模型代码、优化器和第三方库继续使用标准 PyTorch API;在不同加速器之间迁移不需要修改张量设备字符串、内核启动方式或模型代码。 - -设备模块在导入时通过 `torch.utils.rename_privateuse1_backend("flagos")` 与 `torch._register_device_module()` 安装,因此 `device="flagos"`、`torch.flagos.*` 方法、张量方法和存储的行为与 PyTorch 原生设备一致。 - -## 按算子路由后端 - -每个构建出的 wheel 只带一张路由表 `torch_fl/configs/backends_.conf`,其中是针对该平台编译的 `op = backend` 条目。路由以算子为粒度(而非设备或模型粒度),同一个模型内可以混用不同后端: - -| 路由族 | 为算子提供服务的是 | -|---|---| -| `flagos_python` | 经 Python 调度器的 FlagGems Triton 内核 | -| FlagGems C++(`kFlagOs`) | 经 C++ 运行时 `liboperators.so` 的 FlagGems 内核 | -| `cuda` | 基于外部或厂商 `libtorch_cuda.so` 的 CUDA 兼容性 boxing 内核 | -| 厂商原生(`ascend`、`gcu`、`musa` 等) | 厂商算子库(ACLNN、topsaten、mudnn 等) | -| `tileops` | TileOps/TileLang 内核(NVIDIA SM90 芯片) | -| `cpu_fallback` | 以正确性优先的 CPU 实现,结果复制回设备 | - -两个运行期变量可以在不重新构建的情况下改变路由: - -- `FLAGOS_BACKEND_CONFIG` — 让进程改用另一张路由表。 -- `FLAGOS_OP_` — 覆盖单个算子,例如 `FLAGOS_OP_add__Tensor=cuda`。 -- `FLAGOS_FORCE_BACKEND` — 将所有算子重新固定到某个后端族(`flaggems`、`vendor`、`tileops`),用于 A/B 测量。 - -当前路由始终可查询:`torch_fl.backend_config_path()` 返回正在使用的路由表。 - -## 执行路径 - -Torch-FL 实现四类算子执行策略,同一平台可以组合多种。这些属于内部实现策略,而不是用户选择的产品等级。 - -- **厂商原生内核** — 直接调用厂商运行时和算子库(Ascend 的 ACLNN、燧原 GCU 的 topsaten、摩尔线程 MUSA 的 mudnn)。插件为各厂商 C/C++ API 生成绑定代码。 -- **兼容性 boxing** — 当厂商栈提供可与 `PrivateUse1` 共存的独立 PyTorch dispatch key 时,以零拷贝方式转换张量元数据。CUDA boxing 通过外部 `libtorch_cuda.so` 复用 NVIDIA 内核;MetaX、PPU 和海光 DCU 以同样方式对各自的厂商 torch 构建进行 boxing。 -- **可移植编译器内核** — 使用 Triton 或兼容编译器后端(FlagTree)生成的 FlagGems 内核,在多个加速器系列之间复用,无需逐平台重写。 -- **显式 CPU 回退** — 在 PyTorch 语义允许时,没有设备内核的算子在 CPU 上执行。覆盖范围按平台记录,而不是被描述为完整原生支持。 - -## 设备运行时与管理 API - -`torch.flagos` 模块提供完整的设备接口: - -- 流(`Stream`、`current_stream`、`stream`)与事件(`Event`) -- 设备查询:`device_count`、`current_device`、`set_device`、`get_device_properties` -- 同步:`synchronize` -- 随机数:`manual_seed`、`manual_seed_all`、`get_rng_state`、`set_rng_state`、`initial_seed` -- 自动混合精度:`get_amp_supported_dtype` -- 内存:`memory_allocated`、`memory_reserved`、`memory_stats`、`reset_peak_memory_stats`、`empty_cache` - -设备内存默认使用缓存分配器(`FLAGOS_USE_CACHING_ALLOCATOR`,默认开启);设为 `0` 时每次分配都直接交给厂商运行时。 - -## 训练栈 - -- **Eager 执行与 autograd** — 前向与反向路径在设备上执行;原生扩展在加载时注册 `AutogradPrivateUse1` 回退。 -- **`torch.autocast("flagos")` 与 `torch.amp.GradScaler("flagos")`** — 支持 FP16 与 BF16 目标精度,并使用 PyTorch 标准的 autocast 策略分组。 -- **`torch.compile`** — Inductor 集成将 `flagos` 注册为一等 GPU 设备,详见 {doc}`torch.compile 集成 <../architecture/torch-compile>`。 -- **分布式** — 支持 `ProcessGroupFlagOS`、DDP、DataParallel 与 FSDP2,详见 {doc}`分布式集合通信 <../architecture/distributed>`。 - -## 低精度与量化 - -`torch_fl.quantization` 提供低精度转换工具与模块: - -- `convert`、`FormatSpec`、`LowPrecisionFormat`、`get_format_spec`、`normalize_format`、`supported_formats` -- `SoftLowpLinear` — 在矩阵乘之前解码低精度权重的线性层 - -在 DCU 与 MetaX 共用的 CUDA boxing 路径上,软件低精度矩阵路径(`soft_lowp`)为 `mm`、`bmm`、`addmm` 提供标量 FP8 格式(`float8_e4m3fn`、`float8_e5m2`、`float8_e4m3fnuz`、`float8_e5m2fnuz`、`float8_e8m0fnu`)与 packed FP4(`float4_e2m1fn_x2`)支持。数值以 BF16 解码并累加:未指定输出 dtype 时默认 BF16,显式指定时遵循指定值。块缩放元数据格式(MXFP4、NVFP4、block FP4)与 `_scaled_mm` 系列不在该路径范围内。 - -## 框架与生态兼容 - -导入时会安装若干兼容层,用于适配其他项目对设备的既有假设;它们都是可选的,并可通过各自开关进行测量: - -- **Apex** — 补丁作用于常见的 `MultiTensorApply` 入口,使 Apex 的 `amp_C` 内核接收到 flagos 张量的零拷贝 CUDA 视图。 -- **`torch.nn.attention.flex_attention`** — 放宽其硬编码的 `{cuda, cpu, xpu, hpu}` 设备白名单以接受 `flagos` 设备;融合模板可通过 `flagos` 编译后端使用。 -- **`diffusers` 的 Qwen-Image 旋转位置编码** — 在逐设备 RoPE 表的两处都注册 `flagos` 设备,使旋转不再走 GCU 与 Ascend 上原本会退化的复数指数路径。 -- **CUDA 别名** — `FLAGOS_ALIAS_CUDA` 默认开启,`cuda` 设备字符串会被接受为 `flagos`,未修改的 CUDA 脚本可直接运行在 `flagos` 设备上。 - -## 可观测性 - -- **调度与回退日志** — `FLAGOS_LOG=dispatch,fallback` 打印每个算子选中的后端,以及每次 CPU 回退调度。 -- **Profiler** — `torch.profiler` 可采集设备时间线,包含流向箭头、逐算子设备时间、内核元数据与运行期事件名称,详见 {doc}`Profiler 集成 <../architecture/profiler>`。 -- **wheel 兼容性清单** — 每个 wheel 都带有 `torch_fl/compatibility.json`,记录平台、内核集合、打包的 libtorch、构建期 PyTorch 与 ABI、以及构建时观测到的 FlagTree/FlagGems/FlagCX 版本;`torch-fl-preflight` 命令可在导入原生扩展前进行检查。 -- **未知变量告警** — `import torch_fl` 会扫描一次环境,对拼写错误的 `FLAGOS_*` 名称给出告警,而不是静默忽略。 - -## 硬件支持 - -| 平台 | 执行路径 | 已验证能力 | 状态 | -|---|---|---|---| -| NVIDIA CUDA | 基于外部 `libtorch_cuda.so` 的 CUDA boxing | Eager、autograd、分布式(FlagCX/NCCL)、profiler(CUPTI)、FlagGems(Python + C++) | 稳定 | -| MetaX | 通过 cu-bridge 对厂商 libtorch 进行 CUDA boxing | Eager、autograd、AMP、低精度矩阵运算 | 稳定 | -| Ascend | 原生 ACLNN 后端,通过 FlagTree(Triton 3.5)使用 FlagGems | Eager、autograd、RNG 套件、profiler(MSPTI) | Beta | -| PPU | 针对 PPU CUDA 13 兼容 SDK 的 CUDA boxing | Eager、autograd、AMP | 实验性 | -| 海光 DCU | 基于 hipify 的 DTK torch 的 CUDA boxing | Eager、autograd、FP16/BF16 AMP、profiler | Beta | -| 燧原 GCU | 原生 topsaten 后端,未路由及 int64/float64 算子使用 CPU 回退 | Eager、AMP | Beta | -| 摩尔线程 MUSA | FlagGems 优先的 Triton 内核,原生 mudnn 回退,未路由算子使用 CPU 回退 | Eager、FP16/BF16 AMP | 实验性 | -| 地平线 BPU | 无 eager 内核;通过 hbdk4 使用 `torch.compile` 图执行路径 | 仅图编译 | 仅运行时 | -| 清微智能 | 已提供运行时构建选择器 | 尚无逐算子内核集合 | 仅运行时 | - -包括 `torch.compile`、分布式与 profiler 在内的逐项能力状态,请参阅 {doc}`兼容性矩阵 <../reference/compatibility>`。 +# 功能特性 + +## 统一设备与标准 PyTorch API + +所有已支持的加速器都通过 `flagos` 设备编程。模型代码、优化器和第三方库继续使用标准 PyTorch API;在不同加速器之间迁移不需要修改张量设备字符串、内核启动方式或模型代码。 + +设备模块在导入时通过 `torch.utils.rename_privateuse1_backend("flagos")` 与 `torch._register_device_module()` 安装,因此 `device="flagos"`、`torch.flagos.*` 方法、张量方法和存储的行为与 PyTorch 原生设备一致。 + +## 按算子路由后端 + +每个构建出的 wheel 只带一张路由表 `torch_fl/configs/backends_.conf`,其中是针对该平台编译的 `op = backend` 条目。路由以算子为粒度(而非设备或模型粒度),同一个模型内可以混用不同后端: + +| 路由族 | 为算子提供服务的是 | +|---|---| +| `flagos_python` | 经 Python 调度器的 FlagGems Triton 内核 | +| FlagGems C++(`kFlagOs`) | 经 C++ 运行时 `liboperators.so` 的 FlagGems 内核 | +| `cuda` | 基于外部或厂商 `libtorch_cuda.so` 的 CUDA 兼容性 boxing 内核 | +| 厂商原生(`ascend`、`gcu`、`musa` 等) | 厂商算子库(ACLNN、topsaten、mudnn 等) | +| `tileops` | TileOps/TileLang 内核(NVIDIA SM90 芯片) | +| `cpu_fallback` | 以正确性优先的 CPU 实现,结果复制回设备 | + +两个运行期变量可以在不重新构建的情况下改变路由: + +- `FLAGOS_BACKEND_CONFIG` — 让进程改用另一张路由表。 +- `FLAGOS_OP_` — 覆盖单个算子,例如 `FLAGOS_OP_add__Tensor=cuda`。 +- `FLAGOS_FORCE_BACKEND` — 将所有算子重新固定到某个后端族(`flaggems`、`vendor`、`tileops`),用于 A/B 测量。 + +当前路由始终可查询:`torch_fl.backend_config_path()` 返回正在使用的路由表。 + +## 执行路径 + +Torch-FL 实现四类算子执行策略,同一平台可以组合多种。这些属于内部实现策略,而不是用户选择的产品等级。 + +- **厂商原生内核** — 直接调用厂商运行时和算子库(Ascend 的 ACLNN、燧原 GCU 的 topsaten、摩尔线程 MUSA 的 mudnn)。插件为各厂商 C/C++ API 生成绑定代码。 +- **兼容性 boxing** — 当厂商栈提供可与 `PrivateUse1` 共存的独立 PyTorch dispatch key 时,以零拷贝方式转换张量元数据。CUDA boxing 通过外部 `libtorch_cuda.so` 复用 NVIDIA 内核;MetaX、PPU 和海光 DCU 以同样方式对各自的厂商 torch 构建进行 boxing。 +- **可移植编译器内核** — 使用 Triton 或兼容编译器后端(FlagTree)生成的 FlagGems 内核,在多个加速器系列之间复用,无需逐平台重写。 +- **显式 CPU 回退** — 在 PyTorch 语义允许时,没有设备内核的算子在 CPU 上执行。覆盖范围按平台记录,而不是被描述为完整原生支持。 + +## 设备运行时与管理 API + +`torch.flagos` 模块提供完整的设备接口: + +- 流(`Stream`、`current_stream`、`stream`)与事件(`Event`) +- 设备查询:`device_count`、`current_device`、`set_device`、`get_device_properties` +- 同步:`synchronize` +- 随机数:`manual_seed`、`manual_seed_all`、`get_rng_state`、`set_rng_state`、`initial_seed` +- 自动混合精度:`get_amp_supported_dtype` +- 内存:`memory_allocated`、`memory_reserved`、`memory_stats`、`reset_peak_memory_stats`、`empty_cache` + +设备内存默认使用缓存分配器(`FLAGOS_USE_CACHING_ALLOCATOR`,默认开启);设为 `0` 时每次分配都直接交给厂商运行时。 + +## 训练栈 + +- **Eager 执行与 autograd** — 前向与反向路径在设备上执行;原生扩展在加载时注册 `AutogradPrivateUse1` 回退。 +- **`torch.autocast("flagos")` 与 `torch.amp.GradScaler("flagos")`** — 支持 FP16 与 BF16 目标精度,并使用 PyTorch 标准的 autocast 策略分组。 +- **`torch.compile`** — Inductor 集成将 `flagos` 注册为一等 GPU 设备,详见 {doc}`torch.compile 集成 <../architecture/torch-compile>`。 +- **分布式** — 支持 `ProcessGroupFlagOS`、DDP、DataParallel 与 FSDP2,详见 {doc}`分布式集合通信 <../architecture/distributed>`。 + +## 低精度与量化 + +`torch_fl.quantization` 提供低精度转换工具与模块: + +- `convert`、`FormatSpec`、`LowPrecisionFormat`、`get_format_spec`、`normalize_format`、`supported_formats` +- `SoftLowpLinear` — 在矩阵乘之前解码低精度权重的线性层 + +在 DCU 与 MetaX 共用的 CUDA boxing 路径上,软件低精度矩阵路径(`soft_lowp`)为 `mm`、`bmm`、`addmm` 提供标量 FP8 格式(`float8_e4m3fn`、`float8_e5m2`、`float8_e4m3fnuz`、`float8_e5m2fnuz`、`float8_e8m0fnu`)与 packed FP4(`float4_e2m1fn_x2`)支持。数值以 BF16 解码并累加:未指定输出 dtype 时默认 BF16,显式指定时遵循指定值。块缩放元数据格式(MXFP4、NVFP4、block FP4)与 `_scaled_mm` 系列不在该路径范围内。 + +## 框架与生态兼容 + +导入时会安装若干兼容层,用于适配其他项目对设备的既有假设;它们都是可选的,并可通过各自开关进行测量: + +- **Apex** — 补丁作用于常见的 `MultiTensorApply` 入口,使 Apex 的 `amp_C` 内核接收到 flagos 张量的零拷贝 CUDA 视图。 +- **`torch.nn.attention.flex_attention`** — 放宽其硬编码的 `{cuda, cpu, xpu, hpu}` 设备白名单以接受 `flagos` 设备;融合模板可通过 `flagos` 编译后端使用。 +- **`diffusers` 的 Qwen-Image 旋转位置编码** — 在逐设备 RoPE 表的两处都注册 `flagos` 设备,使旋转不再走 GCU 与 Ascend 上原本会退化的复数指数路径。 +- **CUDA 别名** — `FLAGOS_ALIAS_CUDA` 默认开启,`cuda` 设备字符串会被接受为 `flagos`,未修改的 CUDA 脚本可直接运行在 `flagos` 设备上。 + +## 可观测性 + +- **调度与回退日志** — `FLAGOS_LOG=dispatch,fallback` 打印每个算子选中的后端,以及每次 CPU 回退调度。 +- **Profiler** — `torch.profiler` 可采集设备时间线,包含流向箭头、逐算子设备时间、内核元数据与运行期事件名称,详见 {doc}`Profiler 集成 <../architecture/profiler>`。 +- **wheel 兼容性清单** — 每个 wheel 都带有 `torch_fl/compatibility.json`,记录平台、内核集合、打包的 libtorch、构建期 PyTorch 与 ABI、以及构建时观测到的 FlagTree/FlagGems/FlagCX 版本;`torch-fl-preflight` 命令可在导入原生扩展前进行检查。 +- **未知变量告警** — `import torch_fl` 会扫描一次环境,对拼写错误的 `FLAGOS_*` 名称给出告警,而不是静默忽略。 + +## 硬件支持 + +| 平台 | 执行路径 | 已验证能力 | 状态 | +|---|---|---|---| +| NVIDIA CUDA | 基于外部 `libtorch_cuda.so` 的 CUDA boxing | Eager、autograd、分布式(FlagCX/NCCL)、profiler(CUPTI)、FlagGems(Python + C++) | 稳定 | +| MetaX | 通过 cu-bridge 对厂商 libtorch 进行 CUDA boxing | Eager、autograd、AMP、低精度矩阵运算 | 稳定 | +| Ascend | 原生 ACLNN 后端,通过 FlagTree(Triton 3.5)使用 FlagGems | Eager、autograd、RNG 套件、profiler(MSPTI) | Beta | +| PPU | 针对 PPU CUDA 13 兼容 SDK 的 CUDA boxing | Eager、autograd、AMP | 实验性 | +| 海光 DCU | 基于 hipify 的 DTK torch 的 CUDA boxing | Eager、autograd、FP16/BF16 AMP、profiler | Beta | +| 燧原 GCU | 原生 topsaten 后端,未路由及 int64/float64 算子使用 CPU 回退 | Eager、AMP | Beta | +| 摩尔线程 MUSA | FlagGems 优先的 Triton 内核,原生 mudnn 回退,未路由算子使用 CPU 回退 | Eager、FP16/BF16 AMP | 实验性 | +| 地平线 BPU | 无 eager 内核;通过 hbdk4 使用 `torch.compile` 图执行路径 | 仅图编译 | 仅运行时 | +| 清微智能 | 已提供运行时构建选择器 | 尚无逐算子内核集合 | 仅运行时 | + +包括 `torch.compile`、分布式与 profiler 在内的逐项能力状态,请参阅 {doc}`兼容性矩阵 <../reference/compatibility>`。 diff --git a/docs/torch_fl_zh/overview/overview.md b/docs/torch_fl_zh/overview/overview.md index edf068f459..87f6c4d456 100644 --- a/docs/torch_fl_zh/overview/overview.md +++ b/docs/torch_fl_zh/overview/overview.md @@ -1,51 +1,51 @@ -# Torch-FL 概览 - -`torch_fl` 是基于 `PrivateUse1` 扩展机制的自定义 PyTorch 设备插件。它将 [FlagGems](https://github.com/flagos-ai/FlagGems) 高性能 Triton 算子、厂商原生算子库和 CUDA 兼容内核统一注册到同一个设备名称下:`flagos`。 - -不同加速器厂商提供的运行时、编译器栈和 PyTorch 集成方式各不相同。Torch-FL 通过统一的运行时和算子路由层屏蔽这些差异:用户只使用标准 PyTorch API 和单一设备名称,插件根据平台能力与配置为每个算子选择内核实现。 - -## 设计理念 - -Torch-FL 遵循五项原则: - -1. **PyTorch 原生接口** — 标准 PyTorch API 无需修改;用户面向 `flagos` 设备编程,而不是使用厂商专用扩展。 -2. **统一逻辑设备** — 单一设备名称(`flagos`)抽象厂商差异,平台相关路由在算子层透明完成。 -3. **分层算子后端** — 每个操作可以分发到不同实现。路由决策以算子为粒度,而不是以设备或模型为粒度。 -4. **优先复用而非重写** — 在 dispatch 和 ABI 边界允许的情况下集成成熟内核与编译器栈,避免重复实现已有能力。 -5. **明确能力边界** — 对不支持的操作和 CPU 回退路径进行明确说明,不将其描述为完整原生覆盖;通过状态等级区分已验证支持与实验性集成。 - -## 快速开始 - -```python -import torch -import torch_fl - -# 在 flagos 设备上创建张量 -x = torch.randn(4, 4, device="flagos:0") - -# 算子路由到与平台匹配的内核 -y = torch.relu(x @ x) - -# 结果传回 CPU -print(y.cpu()) -``` - -算子路由(FlagGems 编译器内核、厂商原生内核、兼容性 boxing 或 CPU 回退)由平台探测和运行期配置决定。上述代码在所有已支持的加速器上保持不变。 - -## 状态定义 - -| 状态 | 含义 | -|---|---| -| 稳定 | 关键路径持续接受测试,并已记录受支持的版本组合。 | -| Beta | 主要路径已经验证,但覆盖范围、打包或发布流程尚未稳定。 | -| 实验性 | 已在特定配置、模型或硬件环境中完成验证;接口或构建流程仍可能变化。 | -| 仅运行时 | 已提供设备运行时支持,但该平台不是通用 eager 算子后端。 | - -某项功能存在于 Torch-FL 代码库中,并不代表每个平台都已实现或验证该功能。分平台详情请参阅 {doc}`兼容性矩阵 <../reference/compatibility>`。 - -```{toctree} -:maxdepth: 2 - -features.md -architecture.md -``` +# Torch-FL 概览 + +`torch_fl` 是基于 `PrivateUse1` 扩展机制的自定义 PyTorch 设备插件。它将 [FlagGems](https://github.com/flagos-ai/FlagGems) 高性能 Triton 算子、厂商原生算子库和 CUDA 兼容内核统一注册到同一个设备名称下:`flagos`。 + +不同加速器厂商提供的运行时、编译器栈和 PyTorch 集成方式各不相同。Torch-FL 通过统一的运行时和算子路由层屏蔽这些差异:用户只使用标准 PyTorch API 和单一设备名称,插件根据平台能力与配置为每个算子选择内核实现。 + +## 设计理念 + +Torch-FL 遵循五项原则: + +1. **PyTorch 原生接口** — 标准 PyTorch API 无需修改;用户面向 `flagos` 设备编程,而不是使用厂商专用扩展。 +2. **统一逻辑设备** — 单一设备名称(`flagos`)抽象厂商差异,平台相关路由在算子层透明完成。 +3. **分层算子后端** — 每个操作可以分发到不同实现。路由决策以算子为粒度,而不是以设备或模型为粒度。 +4. **优先复用而非重写** — 在 dispatch 和 ABI 边界允许的情况下集成成熟内核与编译器栈,避免重复实现已有能力。 +5. **明确能力边界** — 对不支持的操作和 CPU 回退路径进行明确说明,不将其描述为完整原生覆盖;通过状态等级区分已验证支持与实验性集成。 + +## 快速开始 + +```python +import torch +import torch_fl + +# 在 flagos 设备上创建张量 +x = torch.randn(4, 4, device="flagos:0") + +# 算子路由到与平台匹配的内核 +y = torch.relu(x @ x) + +# 结果传回 CPU +print(y.cpu()) +``` + +算子路由(FlagGems 编译器内核、厂商原生内核、兼容性 boxing 或 CPU 回退)由平台探测和运行期配置决定。上述代码在所有已支持的加速器上保持不变。 + +## 状态定义 + +| 状态 | 含义 | +|---|---| +| 稳定 | 关键路径持续接受测试,并已记录受支持的版本组合。 | +| Beta | 主要路径已经验证,但覆盖范围、打包或发布流程尚未稳定。 | +| 实验性 | 已在特定配置、模型或硬件环境中完成验证;接口或构建流程仍可能变化。 | +| 仅运行时 | 已提供设备运行时支持,但该平台不是通用 eager 算子后端。 | + +某项功能存在于 Torch-FL 代码库中,并不代表每个平台都已实现或验证该功能。分平台详情请参阅 {doc}`兼容性矩阵 <../reference/compatibility>`。 + +```{toctree} +:maxdepth: 2 + +features.md +architecture.md +``` diff --git a/docs/torch_fl_zh/reference/compatibility.md b/docs/torch_fl_zh/reference/compatibility.md index 6954b6c41c..95b6f34796 100644 --- a/docs/torch_fl_zh/reference/compatibility.md +++ b/docs/torch_fl_zh/reference/compatibility.md @@ -1,58 +1,59 @@ -# 兼容性与平台支持 - -## 状态定义 - -| 状态 | 含义 | -|---|---| -| 稳定 | 关键路径持续接受测试,并已记录受支持的版本组合。 | -| Beta | 主要路径已经验证,但覆盖范围、打包或发布流程尚未稳定。 | -| 实验性 | 已在特定配置、模型或硬件环境中完成验证;接口或构建流程仍可能变化。 | -| 仅运行时 | 已提供设备运行时支持,但该平台不是通用 eager 算子后端。 | - -## 项目兼容性 - -| 组件 | 支持范围 | 说明 | -|---|---|---| -| Python | 3.8 或更高版本 | 平台 SDK 与可用 wheel 可能要求更窄的范围 | -| PyTorch | 2.10.x(`>=2.10,<2.11`) | 生成的 ATen 绑定与该次版本线绑定 | -| FlagGems | 取决于平台 | 仅在平台路由使用 FlagGems 时,从 PyPI 或厂商兼容构建安装 | -| Triton/编译器 | 取决于平台 | 使用所选加速器要求的编译器发行版;平台基于 FlagTree 时使用 FlagTree | - -### ATen 次版本线固定 - -Torch-FL 会为 PyTorch 内部的 ATen 算子注册表生成原生绑定。这些绑定对 C++ ABI 与算子 schema 变化敏感,因此项目固定在某个 PyTorch 次版本线上 —— 当前为 **2.10.x**。使用不同的次版本(例如 2.11.x)会导致构建或运行失败;同一次版本线内的补丁版本(2.10.0 → 2.10.1)互相兼容。 - -### wheel 兼容性记录 - -每个构建出的 wheel 都带有 `torch_fl/compatibility.json`,其中记录所选的平台与内核集合、打包的 libtorch 位置、构建期 PyTorch 版本与 C++ ABI 标记、构建时观测到的 FlagTree/FlagGems/FlagCX 版本、在由厂商提供设备库时显式声明的厂商 PyTorch 版本,以及 wheel 的 Python 依赖要求。只有当构建者把 `FLAGOS_SDK_VERSION` 设为已验证值时才记录 SDK 版本;缺失表示**未知**,而不是「兼容所有 SDK」。 - -```bash -python scripts/tools/torch-fl-preflight --wheel dist/torch_fl-*.whl --platform cuda \ - --sdk-version 13.3 --check-installed -``` - -`torch-fl-preflight` 运行时不会导入 `torch_fl`,因此检查不会触发后端导入副作用。`--check-installed` 对比声明的依赖范围与已安装环境,`--check-build-env` 要求精确的构建期版本,`--require-sdk` 拒绝未声明 SDK 的 wheel;使用 `--release --markdown-table` 可由最终 wheel 文件生成发布对照表。 - -## 平台矩阵 - -| 平台 | 构建选择器 | 执行路径 | Eager 与 autograd | torch.compile | 分布式 | Profiler | FlagGems | 状态 | -|---|---|---|---|---|---|---|---|---| -| NVIDIA CUDA | `FLAGOS_ACCELERATOR=cuda`(默认) | 基于外部 `libtorch_cuda.so` 的 CUDA boxing | 稳定 | 实验性(已注册 Inductor GPU 设备,CI 无该步骤) | Beta(FlagCX + NCCL 回退,DDP 已实测) | 稳定(CUPTI 对等性) | Beta(Python + C++ 调度路径) | 稳定 | -| MetaX | `FLAGOS_ACCELERATOR=metax` | 通过 `cu-bridge` 对厂商 libtorch 进行 CUDA boxing | 稳定(boxing 模式下实测 FP16/BF16 autocast 与 GradScaler) | 实验性(厂商 Triton 与 FlagTree 已在 C550 上实测) | 实验性(NCCL 形态的 `mccl` 回退,未纳入 CI) | 实验性(MCPTI 对等性已在 C550 上实测,未纳入 CI) | 实验性(Python 调度,MetaX 上未做 CI 测试) | 稳定 | -| Ascend | `FLAGOS_ACCELERATOR=ascend` | 原生 ACLNN 后端,通过 FlagTree(Triton 3.5)使用 FlagGems | 稳定(CI 覆盖的算子与 RNG 套件) | 实验性(仅在 910 + triton-ascend 上测过 Inductor,未在 FlagTree 上复验;CI 无该步骤) | 实验性(HCCL 回退;仅架构层面路由) | Beta(MSPTI 事件与设备时间关联由共享契约覆盖 CI,对等性套件未纳入) | Beta(Python 调度;float64 与 bool 的 `neg` 路由回退 ACLNN) | Beta | -| PPU | `FLAGOS_ACCELERATOR=ppu` | 针对 PPU CUDA 13 兼容 SDK 的 CUDA boxing,自带 libtorch | 实验性(FP16/BF16 autocast 与 GradScaler 已在 PPU 硬件上实测,未纳入 CI) | 未验证 | 实验性(经厂商适配的 `libnccl.so.2` 实现 NCCL 回退,未纳入 CI) | 未在该厂商追踪器上验证 | 实验性(需要厂商源 Triton) | 实验性 | -| 海光 DCU | `FLAGOS_ACCELERATOR=dcu` | 基于 hipify 的 DTK torch 构建的 CUDA boxing | Beta(含 FP16/BF16 autocast 与 GradScaler) | 实验性(FlagTree HCU 已在 `gfx936` 上验证,未纳入 CI) | 实验性(经 DTK 的 RCCL;`all_reduce`/DDP 已在 2 卡上实测) | Beta(对等性套件在 CI 中运行) | Beta(仅 Python 调度) | Beta | -| 燧原 GCU | `FLAGOS_ACCELERATOR=gcu` | 原生 `libtopsaten.so` 后端,未路由及 int64/float64 算子使用 CPU 回退 | Beta(算子、RNG、工厂与 AMP 套件在 S60 上由 CI 守卫) | 未验证 | 未验证 | 仅运行时(TOPSPTI 采集活动,但仅有 CPU 的 Kineto 构建不产生设备事件) | 实验性(Python 调度,需要厂商 Triton) | Beta | -| 摩尔线程 MUSA | `FLAGOS_ACCELERATOR=musa` | 原生 `mudnn` 后端,未路由算子使用 CPU 回退 | 实验性(FP16/BF16 autocast 与 GradScaler 已在 MTT S5000 上实测) | 实验性(FlagTree 前反向已在 MTT S5000 上实测,需要厂商运行时) | 未验证 | 实验性(MUPTI 设备时间线已在 MTT S5000 上实测) | 实验性(Python 调度,需要厂商 Triton) | 实验性 | -| 地平线 BPU | `FLAGOS_ACCELERATOR=bpu` | 不构建 eager 内核集合,eager 算子在 CPU 上执行 | 仅运行时(eager 走 CPU 回退) | 实验性(经 hbdk4 的 `torch.compile(backend="bpu")` 图路径) | 不适用 | 未验证 | 不适用(无逐算子内核构建) | 仅运行时 | -| 清微智能 | `FLAGOS_ACCELERATOR=tsingmicro` | 已提供运行时/构建选择器,无逐算子内核集合 | 仅运行时 | 未验证 | 未验证 | 未验证 | 不适用 | 仅运行时 | - -## 如何阅读本矩阵 - -- **Eager 与 autograd** 是主要算子路径;“稳定”表示该平台的关键路径持续接受测试。 -- **torch.compile** 在多数平台上仍属实验性:仅在特定硬件上验证,多个平台尚无对应 CI 步骤。详见 {doc}`torch.compile 集成 <../architecture/torch-compile>`。 -- **分布式**评级反映的是已实测的集合通信与 DDP 覆盖范围,而不是代码是否存在。详见 {doc}`分布式集合通信 <../architecture/distributed>`。 -- **Profiler** 评级描述平台满足 `torch.profiler` 契约的哪些部分。详见 {doc}`Profiler 集成 <../architecture/profiler>`。 -- **FlagGems** 表示该平台可用可移植的 Triton 内核路由;其可用性与正确性分别测量。 - -这里的评级描述的是平台,而不是某次构建。仅编译了部分内核集合(`FLAGOS_BUILD_*`)的 wheel 只支持该平台能力的一个子集 —— 对该 wheel 而言,其自身的构建记录才是权威依据。 +# 兼容性与平台支持 + +## 状态定义 + +| 状态 | 含义 | +|---|---| +| 稳定 | 关键路径持续接受测试,并已记录受支持的版本组合。 | +| Beta | 主要路径已经验证,但覆盖范围、打包或发布流程尚未稳定。 | +| 实验性 | 已在特定配置、模型或硬件环境中完成验证;接口或构建流程仍可能变化。 | +| 仅运行时 | 已提供设备运行时支持,但该平台不是通用 eager 算子后端。 | + +## 项目兼容性 + +| 组件 | 支持范围 | 说明 | +|---|---|---| +| Python | 每个平台唯一版本 | 在 `setup.py` 中固定:FlagTree 只对单个 cp tag 发布且 wheel 链接它,因此解释器版本唯一 — CUDA/GCU/MetaX/PPU 为 3.12,DCU/MUSA 为 3.10,Ascend 为 3.11 | +| PyTorch | 2.10.x(`>=2.10,<2.11`) | 生成的 ATen 绑定与该次版本线绑定 | +| FlagGems | 精确固定版本(5.4.0) | 声明为精确要求而非区间:逐算子路由表是针对同一批版本生成的 | +| FlagTree | 精确固定版本,分平台 | 声明为精确要求,且它**就是** Triton —— 包名携带厂商后端(如 `0.7.0+hcu3.6`,尾部数字即 Triton 版本线)。没有 FlagTree 构建的平台改为声明 `triton>=3.5.1` | +| FlagCX | 精确固定版本,分平台 | 仅在厂商运行时存在构建时声明;PPU 目前没有,其余平台分布式路径回退到 NCCL 形态的路由 | + +### ATen 次版本线固定 + +Torch-FL 会为 PyTorch 内部的 ATen 算子注册表生成原生绑定。这些绑定对 C++ ABI 与算子 schema 变化敏感,因此项目固定在某个 PyTorch 次版本线上 —— 当前为 **2.10.x**。使用不同的次版本(例如 2.11.x)会导致构建或运行失败;同一次版本线内的补丁版本(2.10.0 → 2.10.1)互相兼容。 + +### wheel 兼容性记录 + +每个构建出的 wheel 都带有 `torch_fl/compatibility.json`,其中记录所选的平台与内核集合、打包的 libtorch 位置、构建期 PyTorch 版本与 C++ ABI 标记、构建时观测到的 FlagTree/FlagGems/FlagCX 版本、在由厂商提供设备库时显式声明的厂商 PyTorch 版本,以及 wheel 的 Python 依赖要求。只有当构建者把 `FLAGOS_SDK_VERSION` 设为已验证值时才记录 SDK 版本;缺失表示**未知**,而不是「兼容所有 SDK」。 + +```bash +python scripts/tools/torch-fl-preflight --wheel dist/torch_fl-*.whl --platform cuda \ + --sdk-version 13.3 --check-installed +``` + +`torch-fl-preflight` 运行时不会导入 `torch_fl`,因此检查不会触发后端导入副作用。`--check-installed` 对比声明的依赖范围与已安装环境,`--check-build-env` 要求精确的构建期版本,`--require-sdk` 拒绝未声明 SDK 的 wheel;使用 `--release --markdown-table` 可由最终 wheel 文件生成发布对照表。 + +## 平台矩阵 + +| 平台 | 构建选择器 | 执行路径 | Eager 与 autograd | torch.compile | 分布式 | Profiler | FlagGems | 状态 | +|---|---|---|---|---|---|---|---|---| +| NVIDIA CUDA | `FLAGOS_ACCELERATOR=cuda`(默认) | 基于外部 `libtorch_cuda.so` 的 CUDA boxing | 稳定 | 实验性(已注册 Inductor GPU 设备,仅由集成测试覆盖) | Beta(FlagCX + NCCL 回退,DDP 已实测) | 稳定(CUPTI 对等性) | Beta(Python + C++ 调度路径) | 稳定 | +| MetaX | `FLAGOS_ACCELERATOR=metax` | 通过 `cu-bridge` 对厂商 libtorch 进行 CUDA boxing | 稳定(boxing 模式下实测 FP16/BF16 autocast 与 GradScaler) | 实验性(厂商 Triton 与 FlagTree 已在 C550 上实测) | 实验性(NCCL 形态的 `mccl` 回退,未持续验证) | 实验性(MCPTI 对等性已在 C550 上实测,未持续验证) | 实验性(Python 调度,MetaX 上未验证) | 稳定 | +| Ascend | `FLAGOS_ACCELERATOR=ascend` | 原生 ACLNN 后端,通过 FlagTree(Triton 3.5)使用 FlagGems | 稳定(算子与 RNG 套件已实测) | 实验性(仅在 910 + triton-ascend 上测过 Inductor,未在 FlagTree 上复验) | 实验性(HCCL 回退;仅架构层面路由) | Beta(MSPTI 事件与设备时间关联由共享契约覆盖,对等性套件未纳入) | Beta(Python 调度;float64 与 bool 的 `neg` 路由回退 ACLNN) | Beta | +| PPU | `FLAGOS_ACCELERATOR=ppu` | 针对 PPU CUDA 13 兼容 SDK 的 CUDA boxing,自带 libtorch | 实验性(FP16/BF16 autocast 与 GradScaler 已在 PPU 硬件上实测) | 未验证 | 实验性(经厂商适配的 `libnccl.so.2` 实现 NCCL 回退,未持续验证) | 未在该厂商追踪器上验证 | 实验性(需要厂商源 Triton) | 实验性 | +| 海光 DCU | `FLAGOS_ACCELERATOR=dcu` | 基于 hipify 的 DTK torch 构建的 CUDA boxing | Beta(含 FP16/BF16 autocast 与 GradScaler) | 实验性(FlagTree HCU 已在 `gfx936` 上验证) | 实验性(经 DTK 的 RCCL;`all_reduce`/DDP 已在 2 卡上实测) | Beta(对等性套件已纳入) | Beta(仅 Python 调度) | Beta | +| 燧原 GCU | `FLAGOS_ACCELERATOR=gcu` | 原生 `libtopsaten.so` 后端,未路由及 int64/float64 算子使用 CPU 回退 | Beta(算子、RNG、工厂与 AMP 套件已在设备上实测) | 未验证 | 未验证 | 仅运行时(TOPSPTI 采集活动,但仅有 CPU 的 Kineto 构建不产生设备事件) | 实验性(Python 调度,需要厂商 Triton) | Beta | +| 摩尔线程 MUSA | `FLAGOS_ACCELERATOR=musa` | 原生 `mudnn` 后端,未路由算子使用 CPU 回退 | 实验性(FP16/BF16 autocast 与 GradScaler 已在 MTT S5000 上实测) | 实验性(FlagTree 前反向已在 MTT S5000 上实测,需要厂商运行时) | 未验证 | 实验性(MUPTI 设备时间线已在 MTT S5000 上实测) | 实验性(Python 调度,需要厂商 Triton) | 实验性 | +| 地平线 BPU | `FLAGOS_ACCELERATOR=bpu` | 不构建 eager 内核集合,eager 算子在 CPU 上执行 | 仅运行时(eager 走 CPU 回退) | 实验性(经 hbdk4 的 `torch.compile(backend="bpu")` 图路径) | 不适用 | 未验证 | 不适用(无逐算子内核构建) | 仅运行时 | +| 清微智能 | `FLAGOS_ACCELERATOR=tsingmicro` | 已提供运行时/构建选择器,无逐算子内核集合 | 仅运行时 | 未验证 | 未验证 | 未验证 | 不适用 | 仅运行时 | + +## 如何阅读本矩阵 + +- **Eager 与 autograd** 是主要算子路径;“稳定”表示该平台的关键路径持续接受测试。 +- **torch.compile** 在多数平台上仍属实验性:仅在特定硬件上验证,多个平台仅由集成测试覆盖。详见 {doc}`torch.compile 集成 <../architecture/torch-compile>`。 +- **分布式**评级反映的是已实测的集合通信与 DDP 覆盖范围,而不是代码是否存在。详见 {doc}`分布式集合通信 <../architecture/distributed>`。 +- **Profiler** 评级描述平台满足 `torch.profiler` 契约的哪些部分。详见 {doc}`Profiler 集成 <../architecture/profiler>`。 +- **FlagGems** 表示该平台可用可移植的 Triton 内核路由;其可用性与正确性分别测量。 + +{doc}`平台能力矩阵 ` 记录各加速器默认构建与路由了什么,{doc}`数据类型支持 ` 覆盖各平台的 dtype 与 AMP 边界。这里的评级描述的是平台,而不是某次构建。仅编译了部分内核集合(`FLAGOS_BUILD_*`)的 wheel 只支持该平台能力的一个子集 —— 对该 wheel 而言,其自身的构建记录才是权威依据。 diff --git a/docs/torch_fl_zh/reference/dtype-support.md b/docs/torch_fl_zh/reference/dtype-support.md new file mode 100644 index 0000000000..664fd58401 --- /dev/null +++ b/docs/torch_fl_zh/reference/dtype-support.md @@ -0,0 +1,40 @@ +# 数据类型支持 + +Torch-FL 会为张量**存储**保留请求的数据类型,并对张量间运算遵循 PyTorch 的提升(promotion)规则。计算覆盖范围受各后端所用厂商库限制;此外,AMP 目标支持与 eager 类型支持是两个独立问题:存储层接受的类型,并不意味着每个算子都接受。 + +## 存储与计算 + +| 后端 | 存储与拷贝 | eager 逐元素 | 矩阵乘族 | +|---|---|---|---| +| NVIDIA CUDA | PyTorch CUDA 原生类型支持 | CUDA 原生覆盖 | CUDA 原生覆盖 | +| MetaX(boxing) | MACA libtorch 的 CUDA 兼容覆盖;支持 FP8 与 packed FP4 存储 | MACA libtorch 的 CUDA 兼容覆盖 | MACA 覆盖,外加软件模拟的 FP8 / packed FP4 `mm`/`bmm`/`addmm` | +| 海光 DCU | 厂商库覆盖 | 厂商库覆盖 | 厂商库覆盖 | +| Ascend | float16、bfloat16、float32、float64、整型、uint8、bool | 厂商覆盖;不支持的 ACLNN 组合走 CPU 回退 | 原生支持 float16、bfloat16、float32;float64 与不支持类型走 CPU 回退 | +| 燧原 GCU | float16、bfloat16、float32、float64、整型、bool | topsaten 覆盖;int64、float64 与未路由算子走 CPU 回退 | float16、bfloat16、float32 | +| 摩尔线程 MUSA | float16、bfloat16、float32、float64、整型、bool | mudnn 覆盖;未路由算子走 CPU 回退 | float16、bfloat16、float32 | + +## AMP 目标类型 + +`torch.autocast("flagos")` 支持 `torch.float16` 与 `torch.bfloat16` 作为低精度目标,采用 PyTorch 标准 autocast 策略组:矩阵乘与卷积优先使用所选低精度类型,数值敏感算子(对数、归一化)使用 float32,混合输入遵循 promote 策略。**float32 与 float64 不是合法的 autocast 目标。** + +| 后端 | AMP 目标 | 说明 | +|---|---|---| +| NVIDIA CUDA | float16、bfloat16 | | +| MetaX(boxing) | float16、bfloat16 | 在 CUDA-boxing 模式下测量;不覆盖旧版手写内核模式 | +| 海光 DCU | 因后端而异 | | +| Ascend | float16、bfloat16 | float64 可存储且逐元素可用,但既不是 AMP 目标,也不被原生矩阵乘 API 接受 | +| 燧原 GCU | float16、bfloat16 | | +| 摩尔线程 MUSA | float16、bfloat16 | | + +## 边界 + +当厂商算子不接受某类型时,该算子由以实现正确性为先的 CPU 回退承接:在主机上计算,再把类型正确的结果拷回设备。该回退以正确性为目标,可能慢于原生内核 —— 这是成文的覆盖边界,而非故障。 + +- **Ascend** — `aclnnMatmul`、`aclnnMm`、`aclnnBatchMatMul` 拒绝 float64 与整型输入(走回退);`aclnnNeg` 拒绝 int16、uint8 与 bool。float64 的存储、设备拷贝、类型转换与逐元素运算全程保持 float64。复数与量化类型不在支持范围内。 +- **燧原 GCU** — topsaten 没有 float64 与 int64 内核,两者可原生存储但在 CPU 上计算。`topsatenNeg` 另外拒绝 uint8 与 bool。卷积前向走原生 topsaten 路径,反向走 CPU 回退。 +- **摩尔线程 MUSA** — `GradScaler` 的 unscale 使用以实现正确性为先的回退(列表与标量操作数搬到 CPU,执行参考实现,再把变更后的数值与 `found_inf` 拷回),而非原生 foreach 内核。 +- **MetaX** — 低精度矩阵路径是**标量** FP8(`float8_e4m3fn`、`float8_e5m2`、`float8_e4m3fnuz`、`float8_e5m2fnuz`、`float8_e8m0fnu`)与 packed FP4(`float4_e2m1fn_x2`)在 `mm`、`bmm`、`addmm`(含 `dtype`/`out` 变体)上的软件模拟。数值以 BF16 解码并累加,因此未指定输出类型时默认 BF16,显式指定则遵循。块缩放元数据格式(MXFP4、NVFP4、block FP4)与 `_scaled_mm` 系列**不在**其中。 + +## 通过测试意味着什么 + +在缺少原生内核时,集成测试套件会将结果与 CPU 参考实现对比;因此测试通过意味着该算子遵循成文的 PyTorch 契约 —— 而**不一定**意味着它使用了厂商原生内核。 diff --git a/docs/torch_fl_zh/reference/environment-variables.md b/docs/torch_fl_zh/reference/environment-variables.md index fb70d9b013..06dafec323 100644 --- a/docs/torch_fl_zh/reference/environment-variables.md +++ b/docs/torch_fl_zh/reference/environment-variables.md @@ -1,97 +1,107 @@ -# 环境变量 - -Torch-FL 有自己的一套命名空间 `FLAGOS_*`,另外还会读取属于其他项目(torch、FlagGems、FlagCX、TileLang、厂商 SDK)的一组变量,但不拥有它们。本页记录用户需要配置的变量;完整且权威的清单是 `torch_fl/_env.py` 中的 `VARIABLES` 注册表,上游仓库的 `docs/reference/environment-variables.md` 与它逐项对齐,并由单元测试校验。 - -这里没有任何一项是运行 wheel 所必需的:wheel 在空环境下即可完成路由、编译与运行。这些变量用于选择不同的构建、为测量而覆盖某项设置,或打开诊断。 - -## 取值规则 - -- **布尔值**:`1`/`true`/`on`/`yes`(不区分大小写)为开;`0`/`false`/`off`/`no` 为关。其他取值不是布尔值 —— 会输出一次告警并使用该变量的默认值,而不会把该值当作真值。 -- **空值等于未设置**:`FLAGOS_LOG=${EXTRA_LOG}` 在 `EXTRA_LOG` 未设置时等同于从未导出 `FLAGOS_LOG`,因此默认开启的开关会保持开启。 -- **枚举**:命名模式的开关(而非布尔开关)在收到范围外的取值时会列出可选值并使用默认值。 -- **未知名称**:`import torch_fl` 会扫描一次环境,对既未声明、也不属于动态 `FLAGOS_OP_` 族的任何 `FLAGOS_*` 名称给出告警 —— 否则拼写错误的开关会无人读取、静默失效。 - -## 构建选择 - -这些变量是 `setup.py` 与 CMake 构建的输入,运行期没有任何代码读取它们。wheel 会记录构建时的取值,运行期读取方以该记录为准。 - -| 变量 | 默认值 | 作用 | -|---|---|---| -| `FLAGOS_ACCELERATOR` | `cuda` | wheel 面向的平台:`cuda`、`ppu`、`metax`、`ascend`、`tsingmicro`、`dcu`、`gcu`、`musa`、`bpu` | -| `FLAGOS_BUILD_VENDOR` | `ON`,`metax` 上为 `OFF` | 编译厂商原生内核(厂商未提供时为空操作) | -| `FLAGOS_BUILD_FLAGGEMS` | `ON`,`bpu` 上为 `OFF` | 编译 FlagGems Python 内核包装 | -| `FLAGOS_BUILD_FLAGGEMS_CPP` | `cuda`、`tsingmicro` 上为 `ON` | 编译 FlagGems C++ 包装(`liboperators.so`) | -| `FLAGOS_BUILD_BOXING` | `ON`,`ascend`、`gcu`、`musa` 上为 `OFF` | 编译生成的 CUDA boxing 内核 | -| `FLAGOS_BUILD_TILEOPS` | `cuda` 上为 `ON` | 编译 TileOps 内核包装(TileLang,NVIDIA SM90) | -| `FLAGOS_BUILD_JOBS` | CPU 核数 | CMake 构建并行任务数 | -| `FLAGOS_WHEEL_LOCAL` | 由 SDK 推导 | 本地版本标记,例如 `metax3.8.1` | -| `FLAGOS_SKIP_CUDA_ASSETS` | `0` | 不打包外部 `libtorch_cuda.so` | -| `FLAGOS_CUDA_ASSETS_DIR` | `.libtorch_cuda_assets` | 外部 `libtorch_cuda.so` 的拷贝来源目录 | -| `FLAGOS_DCU_VENDOR_CORE` | `0` | 使用 DTK 分支版核心库替代官方 PyTorch 核心库(构建与导入期必须一致) | - -与平台强制取值相矛盾的环境变量显式取值会被拒绝,并在错误信息中同时点名两者,而不是让 CMake 最后看到的那个生效。 - -## 算子路由 - -决定每个算子分发到哪个后端实现。路由在生成的 `torch_fl/configs/backends_.conf` 中逐算子声明,下列变量用于覆盖或扩展该表。 - -| 变量 | 默认值 | 作用 | -|---|---|---| -| `FLAGOS_BACKEND_CONFIG` | 无 | 指向某个 `backends_*.conf` 的绝对路径,覆盖由构建记录选择的路由表。用于测试与调试 | -| `FLAGOS_OP_` | 无 | 逐算子覆盖,例如 `FLAGOS_OP_add__Tensor=cuda`(算子名中的 `.` 替换为 `__`) | -| `FLAGOS_FORCE_BACKEND` | 无 | 把所有算子重新固定到某个后端族(`flaggems`、`vendor`、`tileops`),用于 A/B 测量 | -| `FLAGOS_DISABLE_FLAGGEMS_PY` | `0` | 不注册 FlagGems Python 层(仅 C++ 桩模式) | - -`torch_fl.backend_config_path()` 返回实际使用的路由表;`FLAGOS_BACKEND_CONFIG` 中只保留用户导出的值,因此读取它即可回答「我是否覆盖了路由表」。 - -## 运行期诊断 - -| 变量 | 默认值 | 作用 | -|---|---|---| -| `FLAGOS_LOG` | 无 | 逗号分隔的 stderr 诊断:`dispatch`(每个算子选中的后端)、`fallback`(每次 CPU 回退调度)、`op_cache`(Ascend 算子缓存统计) | -| `FLAGOS_TRACE` | `0` | 设备 profiler shim 的详细日志 | -| `FLAGOS_TRACER_LIBRARY` | 自动探测 | 覆盖 profiler shim 加载的追踪器库 | - -## 分布式 - -| 变量 | 默认值 | 作用 | -|---|---|---| -| `FLAGOS_DIST_REDIRECT_GLOO` | `1` | 当进程加速器是 flagos 设备时,用 flagos 后端响应普通的 `init_process_group(backend="gloo")` 或 `new_group` 请求 | -| `FLAGOS_DIST_STAGED_GLOO` | `1` | 允许 host-staged gloo 内部后端,即无厂商通信库时的最后一级回退。设为 `0` 时直接失败而不做 staged 拷贝 | -| `FLAGOS_DIST_FORCE_NCCL` | `0` | 在 MetaX 手动分布式测试中跳过 FlagCX 而使用 NCCL | - -## 厂商与框架兼容 - -导入期安装的兼容层,用于适配厂商运行时或其他框架对设备的假设。标准 CUDA 机器上无需任何一项。 - -| 变量 | 默认值 | 作用 | -|---|---|---| -| `FLAGOS_ALIAS_CUDA` | `1` | 为兼容性把 `cuda` 设备字符串别名到 `flagos`;设为 `0` 可关闭 | -| `FLAGOS_DISABLE_CUDA_SHIM` | `0` | 不注册面向通用 GPU 操作的 `torch.cuda` 兼容 shim | -| `FLAGOS_METAX_CUDART_SHIM` | `0` | 在 `import torch` 之前预加载 `libcudart` 版本标记 shim;MetaX 搭配通用 PyTorch wheel 时必需 | -| `FLAGOS_METAX_COMPAT` | `0` | 为 MetaX 兼容性补丁 FlagGems 的 `torch.cuda` 设备查询 | -| `FLAGOS_DCU_HIP_VERSION` | 无 | 覆盖 DCU 运行时的 HIP 版本探测 | -| `FLAGOS_DCU_SKIP_RUNTIME_CHECK` | `0` | 跳过 DCU 导入后检查,用于刻意测试不匹配的 wheel 组合 | -| `FLAGOS_DCU_SDPA_FLASH` | `1` | 在 DCU 上让 DTK 的 SDPA 选择器使用其 CUTLASS flash 适配器,而非强制数学分解 | -| `FLAGOS_DISABLE_APEX_COMPAT` | `0` | 关闭可选的 Apex 多张量兼容层 | -| `FLAGOS_DISABLE_QWENIMAGE_ROPE` | `0` | 保留 `diffusers` 的 Qwen-Image 旋转位置编码表原样,以便测量差异 | -| `FLAGOS_DISABLE_FLEX_ATTENTION_COMPAT` | `0` | 保留 flex-attention 硬编码的 `{cuda, cpu, xpu, hpu}` 设备白名单 | - -## 资源、库与编译 - -| 变量 | 默认值 | 作用 | -|---|---|---| -| `FLAGOS_DISABLE_CUDA_ASSETS` | `0` | 跳过对打包 `libtorch_cuda.so` 与 CUDA 库的预加载(用于树内构建与 PPU) | -| `FLAGOS_VENDOR_TORCH_LIB` | 自动探测 | 未打包厂商 libtorch 时,指向厂商 torch 的 `lib` 目录 | -| `FLAGOS_USE_CACHING_ALLOCATOR` | `1` | 缓存设备分配器;设为 `0` 时每次分配都交给厂商运行时 | -| `FLAGOS_USE_FLAGTREE` | `0` | 断言当前 Triton 是 FlagTree 构建(Ascend 上使用 FlagTree 编译器时必需) | -| `FLAGOS_COMPILE_FALLBACK_EAGER` | `0` | `torch.compile` 遇到不支持的算子时回退 eager | -| `FLAGOS_TILEOPS_USE_L2` | `0` | 使用 TileOps 的 L2 缓存层级 | -| `FLAGOS_TILEOPS_CACHE_MAX` | `512` | TileOps 实例缓存容量 | -| `FLAGOS_TILEOPS_DISABLE_ALL_CACHE` | `0` | 关闭所有 TileLang 缓存(正确但慢;须在导入 `tileops` 之前设置) | - -## BPU 编译器 - -BPU 的图路径在 x86 主机上经 `hbdk4` 编译。`FLAGOS_BPU_MARCH` 选择微架构(`nash-p`、`nash-e`、`nash-m`);`FLAGOS_BPU_QUANTIZE`(默认开启)通过插入 int8 量化让卷积留在设备上执行;`FLAGOS_BPU_CACHE` 设置编译器缓存目录;`FLAGOS_BPU_X86_PYTHON`、`FLAGOS_BPU_X86_EMULATOR`、`FLAGOS_BPU_X86_STUBS` 描述板上编译所用的 x86 主机与模拟器。 - -完整的变量清单 —— 包括由其他项目拥有、此处仅作互操作读取的名称,仅供代码生成使用的输入,以及保留但不再生效的历史名称 —— 维护在上游仓库的 `docs/reference/environment-variables.md`,并与 `torch_fl/_env.py` 中的注册表一一对应。 +# 环境变量 + +Torch-FL 有自己的一套命名空间 `FLAGOS_*`,另外还会读取属于其他项目(torch、FlagGems、FlagCX、TileLang、厂商 SDK)的一组变量,但不拥有它们。本页记录用户需要配置的变量;完整且权威的清单是 `torch_fl/_env.py` 中的 `VARIABLES` 注册表,上游仓库的 `docs/reference/environment-variables.md` 与它逐项对齐,并由单元测试校验。 + +这里没有任何一项是运行 wheel 所必需的:wheel 在空环境下即可完成路由、编译与运行。这些变量用于选择不同的构建、为测量而覆盖某项设置,或打开诊断。 + +## 取值规则 + +- **布尔值**:`1`/`true`/`on`/`yes`(不区分大小写)为开;`0`/`false`/`off`/`no` 为关。其他取值不是布尔值 —— 会输出一次告警并使用该变量的默认值,而不会把该值当作真值。 +- **空值等于未设置**:`FLAGOS_LOG=${EXTRA_LOG}` 在 `EXTRA_LOG` 未设置时等同于从未导出 `FLAGOS_LOG`,因此默认开启的开关会保持开启。 +- **枚举**:命名模式的开关(而非布尔开关)在收到范围外的取值时会列出可选值并使用默认值。 +- **未知名称**:`import torch_fl` 会扫描一次环境,对既未声明、也不属于动态 `FLAGOS_OP_` 族的任何 `FLAGOS_*` 名称给出告警 —— 否则拼写错误的开关会无人读取、静默失效。早期版本已废弃的名称被有意排除在该告警之外:旧导出是惰性的(不做别名、无弃用过渡期),不会被静默采纳。 + +## 构建选择 + +这些变量是 `setup.py` 与 CMake 构建的输入,运行期没有任何代码读取它们。wheel 会记录构建时的取值,运行期读取方以该记录为准。 + +| 变量 | 默认值 | 作用 | +|---|---|---| +| `FLAGOS_ACCELERATOR` | `cuda` | wheel 面向的平台:`cuda`、`ppu`、`metax`、`ascend`、`tsingmicro`、`dcu`、`gcu`、`musa`、`bpu` | +| `FLAGOS_BUILD_VENDOR` | `ON`,`metax` 上为 `OFF` | 编译厂商原生内核(厂商未提供时为空操作) | +| `FLAGOS_BUILD_FLAGGEMS` | `ON`,`bpu` 上为 `OFF` | 编译 FlagGems Python 内核包装 | +| `FLAGOS_BUILD_FLAGGEMS_CPP` | `cuda`、`tsingmicro` 上为 `ON` | 编译 FlagGems C++ 包装(`liboperators.so`) | +| `FLAGOS_BUILD_BOXING` | `ON`,`ascend`、`gcu`、`musa` 上为 `OFF` | 编译生成的 CUDA boxing 内核 | +| `FLAGOS_BUILD_TILEOPS` | `cuda` 上为 `ON` | 编译 TileOps 内核包装(TileLang,NVIDIA SM90) | +| `FLAGOS_BUILD_JOBS` | CPU 核数 | CMake 构建并行任务数;`MAX_JOBS` 与 `CMAKE_BUILD_PARALLEL_LEVEL` 作为低优先级回退 | +| `FLAGOS_WHEEL_LOCAL` | 由 SDK 推导 | 本地版本标记,例如 `maca3.8.1.3` | +| `FLAGOS_SKIP_CUDA_ASSETS` | `0` | 不打包外部 `libtorch_cuda.so` | +| `FLAGOS_CUDA_ASSETS_DIR` | `.libtorch_cuda_assets` | 外部 `libtorch_cuda.so` 的拷贝来源目录 | +| `FLAGOS_DCU_VENDOR_CORE` | `0` | 使用 DTK 分支版核心库替代官方 PyTorch 核心库(构建与导入期必须一致) | + +与平台强制取值相矛盾的环境变量显式取值会被拒绝,并在错误信息中同时点名两者,而不是让 CMake 最后看到的那个生效。 + +## 算子路由 + +决定每个算子分发到哪个后端实现。路由在生成的 `torch_fl/configs/backends_.conf` 中逐算子声明,下列变量用于覆盖或扩展该表。 + +| 变量 | 默认值 | 作用 | +|---|---|---| +| `FLAGOS_BACKEND_CONFIG` | 无 | 指向某个 `backends_*.conf` 的绝对路径,覆盖由构建记录选择的路由表。用于测试与调试 | +| `FLAGOS_OP_` | 无 | 逐算子覆盖,例如 `FLAGOS_OP_add__Tensor=cuda`(算子名中的 `.` 替换为 `__`) | +| `FLAGOS_FORCE_BACKEND` | 无 | 把所有算子重新固定到某个后端族(`flaggems`、`vendor`、`tileops`),用于 A/B 测量 | +| `FLAGOS_DISABLE_FLAGGEMS_PY` | `0` | 不注册 FlagGems Python 层(仅 C++ 桩模式) | +| `FLAGOS_STARTUP_PROFILE` | `full` | `full` 在导入期执行框架兼容钩子;`minimal` 将这些钩子留给显式激活。两种模式下 FlagTree、FlagGems、FlagCX 仍为必需依赖 | + +`torch_fl.backend_config_path()` 返回实际使用的路由表;`FLAGOS_BACKEND_CONFIG` 中只保留用户导出的值,因此读取它即可回答「我是否覆盖了路由表」。 + +## 运行期诊断 + +| 变量 | 默认值 | 作用 | +|---|---|---| +| `FLAGOS_LOG` | 无 | 逗号分隔的 stderr 诊断:`dispatch`(每个算子选中的后端)、`fallback`(每次 CPU 回退调度)、`op_cache`(Ascend 算子缓存统计) | +| `FLAGOS_TRACE` | `0` | 设备 profiler shim 的详细日志 | +| `FLAGOS_TRACER_LIBRARY` | 自动探测 | 覆盖 profiler shim 加载的追踪器库 | + +## 分布式 + +| 变量 | 默认值 | 作用 | +|---|---|---| +| `FLAGOS_DIST_REDIRECT_GLOO` | `1` | 当进程加速器是 flagos 设备时,用 flagos 后端响应普通的 `init_process_group(backend="gloo")` 或 `new_group` 请求 | +| `FLAGOS_DIST_STAGED_GLOO` | `1` | 允许 host-staged gloo 内部后端,即无厂商通信库时的最后一级回退。设为 `0` 时直接失败而不做 staged 拷贝 | +| `FLAGOS_DIST_FORCE_NCCL` | `0` | 仅测试用:在 MetaX 手动分布式测试中跳过 FlagCX 而使用 NCCL | + +## 厂商与框架兼容 + +导入期安装的兼容层,用于适配厂商运行时或其他框架对设备的假设。标准 CUDA 机器上无需任何一项。 + +| 变量 | 默认值 | 作用 | +|---|---|---| +| `FLAGOS_ALIAS_CUDA` | `1` | 为兼容性把 `cuda` 设备字符串别名到 `flagos`;设为 `0` 可关闭 | +| `FLAGOS_DISABLE_CUDA_SHIM` | `0` | 不注册面向通用 GPU 操作的 `torch.cuda` 兼容 shim | +| `FLAGOS_METAX_CUDART_SHIM` | `0` | 在 `import torch` 之前预加载 `libcudart` 版本标记 shim;MetaX 搭配通用 PyTorch wheel 时必需 | +| `FLAGOS_METAX_COMPAT` | `0` | 为 MetaX 兼容性补丁 FlagGems 的 `torch.cuda` 设备查询 | +| `FLAGOS_DCU_HIP_VERSION` | 无 | 覆盖 DCU 运行时的 HIP 版本探测 | +| `FLAGOS_DCU_SKIP_RUNTIME_CHECK` | `0` | 跳过 DCU 导入后检查,用于刻意测试不匹配的 wheel 组合 | +| `FLAGOS_DCU_SDPA_FLASH` | `1` | 在 DCU 上让 DTK 的 SDPA 选择器使用其 CUTLASS flash 适配器,而非强制数学分解 | +| `FLAGOS_DISABLE_APEX_COMPAT` | `0` | 关闭可选的 Apex 多张量兼容层 | +| `FLAGOS_DISABLE_QWENIMAGE_ROPE` | `0` | 保留 `diffusers` 的 Qwen-Image 旋转位置编码表原样,以便测量差异 | +| `FLAGOS_DISABLE_FLEX_ATTENTION_COMPAT` | `0` | 保留 flex-attention 硬编码的 `{cuda, cpu, xpu, hpu}` 设备白名单 | + +## 资源、库与编译 + +| 变量 | 默认值 | 作用 | +|---|---|---| +| `FLAGOS_DISABLE_CUDA_ASSETS` | `0` | 跳过对打包 `libtorch_cuda.so` 与 CUDA 库的预加载(用于树内构建与 PPU) | +| `FLAGOS_VENDOR_TORCH_LIB` | 自动探测 | 未打包厂商 libtorch 时,指向厂商 torch 的 `lib` 目录 | +| `FLAGOS_USE_CACHING_ALLOCATOR` | `1` | 缓存设备分配器;设为 `0` 时每次分配都交给厂商运行时 | +| `FLAGOS_USE_FLAGTREE` | `0` | 断言当前 Triton 是 FlagTree 构建(Ascend 上使用 FlagTree 编译器时必需) | +| `FLAGOS_COMPILE_FALLBACK_EAGER` | `0` | `torch.compile` 遇到不支持的算子时回退 eager | +| `FLAGOS_TILEOPS_USE_L2` | `0` | 使用 TileOps 的 L2 缓存层级 | +| `FLAGOS_TILEOPS_CACHE_MAX` | `512` | TileOps 实例缓存容量 | +| `FLAGOS_TILEOPS_DISABLE_ALL_CACHE` | `0` | 关闭所有 TileLang 缓存(正确但慢;须在导入 `tileops` 之前设置) | + +## 代码生成 + +算子绑定生成器的输入;它们在重新生成某平台路由时起作用,运行期无关。 + +| 变量 | 默认值 | 用途 | +|---|---|---| +| `FLAGOS_EXEC_CACHE` | `1` | 缓存 Ascend 算子 codegen 的执行结果;设为 `0` 强制重新生成 | +| `FLAGOS_CODEGEN_ALL` | `0` | 为完整的 leaf-CUDA 算子集合生成路由,而非仅受支持子集 | + +## BPU 编译器 + +BPU 的图路径在 x86 主机上经 `hbdk4` 编译。`FLAGOS_BPU_MARCH` 选择微架构(`nash-p`、`nash-e`、`nash-m`);`FLAGOS_BPU_QUANTIZE`(默认开启)通过插入 int8 量化让卷积留在设备上执行;`FLAGOS_BPU_CACHE` 设置编译器缓存目录;`FLAGOS_BPU_X86_PYTHON`、`FLAGOS_BPU_X86_EMULATOR`、`FLAGOS_BPU_X86_STUBS` 描述板上编译所用的 x86 主机与模拟器。 + +完整的变量清单 —— 包括由其他项目拥有、此处仅作互操作读取的名称,仅供代码生成使用的输入,以及保留但不再生效的历史名称 —— 维护在上游仓库的 `docs/reference/environment-variables.md`,并与 `torch_fl/_env.py` 中的注册表一一对应。 diff --git a/docs/torch_fl_zh/reference/platform-capability.md b/docs/torch_fl_zh/reference/platform-capability.md new file mode 100644 index 0000000000..badd05086c --- /dev/null +++ b/docs/torch_fl_zh/reference/platform-capability.md @@ -0,0 +1,51 @@ +# 平台能力矩阵 + +每个 `FLAGOS_ACCELERATOR` 取值实际提供什么:其 wheel 构建哪条算子路径、设备运行时来自哪里、默认做哪些路由决策。这是在{ref}`兼容性矩阵 `(按能力给出验证状态)之下一层的、按平台的「我的加速器到底支不支持某项能力」的答案。 + +## 算子路径 + +| 平台 | `FLAGOS_ACCELERATOR` | 算子路径 | 设备运行时来源 | 厂商原生内核树 | 状态 | +|---|---|---|---|---|---| +| NVIDIA CUDA | `cuda`(默认) | 原生 CUDA、FlagGems、生成的 CUDA boxing | `cuda` | —(走 CUDA/FlagGems 路径) | Stable | +| PPU | `ppu` | 针对打包的 PPU libtorch 做 CUDA-ABI boxing | `cuda` + 打包的 `lib_ppu/` | —(仅 boxing) | Experimental | +| MetaX | `metax` | 针对打包的 MACA libtorch 做 CUDA-ABI boxing | `metax` | —(手写 MetaX 内核已退役) | Stable | +| 海光 DCU | `dcu` | 基于 hipify 后的 DTK torch 做 CUDA-ABI boxing | `cuda` + DCU DTK-core ABI shim | —(仅 boxing) | Beta | +| Ascend | `ascend` | 厂商原生(ACLNN),FlagGems 经 FlagTree | `ascend` | `ascend` | Beta | +| 燧原 GCU | `gcu` | 厂商原生(topsaten) | `gcu` | `gcu` | Beta | +| 摩尔线程 MUSA | `musa` | 厂商原生(mudnn) | `musa` | `musa` | Experimental | +| TsingMicro | `tsingmicro` | CUDA-ABI boxing(Kuiper SDK) | `tsingmicro` | — | Runtime only | +| 地平线 BPU | `bpu` | 无逐算子内核;整图编译 | `bpu` | — | Runtime only | + +## 内核集合 + +`FLAGOS_BUILD_*` 开关决定 wheel 编译哪些内核集合。默认开启哪些是平台属性,而非每次构建的选择: + +| 内核集合 | 提供内容 | +|---|---| +| `vendor` | 平台的原生算子树(若存在) | +| `flaggems` | FlagGems Python(Triton)调度路径 | +| `flaggems_cpp` | FlagGems C++ 路径(`liboperators.so`) | +| `boxing` | 面向 CUDA-ABI 平台生成的 CUDA boxing 内核 | +| `tileops` | 面向 SM90 NVIDIA 型号的 TileOps/TileLang 内核 | + +wheel 会记录自身构建时启用的集合;运行期读取的是该记录而非环境变量,因此显式导出的 `FLAGOS_BUILD_*` 不会让 wheel 错误地描述自己。各平台默认值见{doc}`环境变量参考 `。 + +## 有意为之的不对称 + +以下是设计决策,不是待补齐的缺口: + +- **BPU** 完全没有逐算子内核:eager 算子走 CPU 回退,加速来自 `torch.compile(backend="bpu")` 的整图编译。 +- **TsingMicro** 是构建目标,有运行时选择器,但没有成文的逐算子内核集合,也没有安装指南。 +- **MUSA** 的加速器目录只提供 stream/event 包装 —— 没有 `torch.cuda` 风格的兼容模块,因此在 MUSA 上访问 `torch.cuda.*` 设备查询不会得到转换层。 +- **PPU** 通过 CUPTI 做性能分析,但不产生 `gpu_memset` 活动,因此与 memset 相关的对等性用例属于范围外,而非失败。 +- **CUDA boxing 是共享的**:一套生成内核,四个厂商运行时来源(CUDA、MetaX、PPU、DCU)。 + +## 设备运行时与 Python 层 + +| 层 | 位置 | 说明 | +|---|---|---| +| 生成的 ATen 绑定 | `csrc/aten/generated/` | 注册在 `PrivateUse1` 下;CUDA/boxing 与 FlagGems 调用方 | +| 厂商原生内核 | `csrc/aten/backends//` | 按厂商库实际导出的算子面生成 | +| 设备运行时 | `csrc/runtime/accelerator//` | DCU 与 PPU 复用 CUDA 运行时树并追加自身部分 | +| Python 设备模块 | `torch_fl/flagos/` | stream、event、RNG、AMP、显存、meta 内核 | +| 路由表 | `torch_fl/configs/backends_.conf` | 每个算子一条 `op = backend` | diff --git a/docs/torch_fl_zh/reference/troubleshooting.md b/docs/torch_fl_zh/reference/troubleshooting.md new file mode 100644 index 0000000000..df1bb38054 --- /dev/null +++ b/docs/torch_fl_zh/reference/troubleshooting.md @@ -0,0 +1,63 @@ +# 常见故障排查 + +按症状给出跨平台反复出现的问题的解决办法。驱动正常但设备数为 0、导入时符号错误、或单个算子崩溃,通常都指向导入顺序、缺失的厂商库、或编译器不匹配这三类原因,下列各节逐一覆盖。 + +## 导入顺序 + +在所有 CUDA-ABI 平台(CUDA、MetaX、PPU、DCU)上,新进程里 `import torch_fl` 必须**先于** `import torch`。导入时 Torch-FL 会预加载厂商 `libtorch` 与 CUDA 资源;若先导入 `torch`,PyTorch 会缓存其 stub CUDA hooks,预加载就失去作用。 + +```python +import torch_fl # 先 +import torch +``` + +顺序写错的典型症状: + +| 症状 | 平台 | +|---|---| +| `Cannot initialize CUDA without ATen_cuda library` | CUDA | +| `undefined symbol`(来自 `libtorch`) | MetaX | +| `c10` 符号未定义,或直接崩溃而非抛异常 | MUSA(wheel 未用 `--no-build-isolation` 构建时) | +| 解析到错误厂商的构建,或 `dlopen` 中止 | 任意 CUDA-ABI 平台 | + +## `torch.flagos.device_count()` 返回 0 + +按顺序排查: + +1. **驱动看得到设备吗?** CUDA 用 `nvidia-smi`,MetaX 用 `mx-smi`,DCU 用 `hy-smi`,其余用厂商工具。驱动都看不到设备时,软件层面无法补救。 +2. **运行时是否初始化?** 驱动/运行时版本错配是最常见原因。CUDA 上重装匹配的 `nvidia-*-cu12` 运行期包。 +3. **导入顺序** —— 见上节。 +4. **设备节点/可见性。** Ascend 需要可访问的 `/dev/davinci*` 节点;CUDA 受 `CUDA_VISIBLE_DEVICES` 影响,越界的 `device="flagos:N"` 会报 `invalid device ordinal`。 +5. **厂商库可达。** MetaX wheel 找不到 `/opt/maca`,或其 `LD_LIBRARY_PATH` 覆盖了 MACA 运行时路径时,`torch.cuda` 与 `torch.flagos` 的设备数会不一致。 + +## 缺失的厂商库 + +| 报错 | 平台 | 原因与处理 | +|---|---|---| +| `cannot open shared object file: libhydmi.so` | DCU | `hyhal` 不在 `LD_LIBRARY_PATH` 上;加入 `/usr/local/hyhal/lib`(或 `/opt/hyhal/lib`) | +| `MIOpen: librt.so not found` | DCU | DTK 的 MIOpen 配置引用了已移除的 `/usr/lib/.../librt.so`;`FLAGOS_ACCELERATOR=dcu` 构建分支会自动改写——确认选择器已设置且代码为最新 | +| `libtorch_cuda.so not found` | PPU | PPU 构建不打包 CUDA 资源:运行 Python 前导出 `FLAGOS_DISABLE_CUDA_ASSETS=1` | +| `CUDA_HOME not set` | PPU | 构建前导出 `CUDA_HOME=/usr/local/PPU_SDK/CUDA_SDK` | + +## 编译器与 Triton + +| 报错 | 原因与处理 | +|---|---| +| `No backend registered for 'hcu'` | 当前 `triton` 不是 FlagTree 构建:反复卸载直到干净,再从 FlagOS 索引安装对应平台的 FlagTree wheel | +| 同一环境里有两个 `triton` | 不带 `dist-info` 的副本 pip 删不掉;先按路径删除 `site-packages/triton` 再重装 | +| `Invalid cross-device link`(PPU) | pip 缓存与构建目录在不同文件系统;直接下载 wheel 后安装该文件 | +| `FLAGGEMS_DIR` 指向不兼容构建后 FlagGems 导入报错 | 从同一索引、同一平台安装 FlagGems 与 FlagTree | + +FlagGems 路由在分发时按名称解析内核,因此缺少该包时 FlagGems 支撑的算子会**报错**而不会静默回退。要确认某次调用实际由谁服务,用 `FLAGOS_LOG=dispatch` 运行。 + +## 分布式 + +- `ProcessGroupGloo` **会直接拒绝 flagos 张量**。在 `FLAGOS_DIST_REDIRECT_GLOO`(默认开启)下,`init_process_group(backend="gloo")` 会被换成 flagos 后端来应答。 +- 若没有可用的厂商通信库(FlagCX / NCCL / HCCL / MCCL),会落到宿主暂存 gloo 层,它把每个操作数做 device → host → device 拷贝。设 `FLAGOS_DIST_STAGED_GLOO=0` 可改为直接失败而非静默暂存。 +- 为 PPU 构建 FlagCX 时出现 `extended_api creator not found`:用 `FLAGCX_ADAPTOR=nvidia` 重新构建。 + +## 性能分析与 torch.compile + +- 某平台上 `torch.profiler` 找不到默认路径下的 tracer 库时,用 `FLAGOS_TRACER_LIBRARY` 指向实际安装的库。 +- 安装的是纯 CPU 版 PyTorch wheel 时,PPU 与 MUSA 可能看不到设备事件:该构建不提供 `PrivateUse1` 的 Kineto 解析器,采集到的活动不会呈现为设备事件。这是环境限制,不是 tracer 缺陷。 +- 平台上未安装厂商 Triton 栈时 `torch.compile` 不可用:Inductor 路径需要该平台的 FlagTree(或厂商 Triton)构建;`FLAGOS_COMPILE_FALLBACK_EAGER=1` 可让不支持的算子回退到 eager。 diff --git a/docs/torch_fl_zh/release_notes/release-notes.md b/docs/torch_fl_zh/release_notes/release-notes.md index 35688e045a..f32e2df58c 100644 --- a/docs/torch_fl_zh/release_notes/release-notes.md +++ b/docs/torch_fl_zh/release_notes/release-notes.md @@ -1,14 +1,14 @@ -# 发布说明 - -本节包含 Torch-FL 的发布信息。 - -## v0.1.0 - -Torch-FL 作为 FlagOS 一部分的初始版本。 - -- **统一的 `flagos` 设备** — 基于 `PrivateUse1` 的 PyTorch 设备插件;标准 PyTorch API、张量方法与存储无需修改即可使用。 -- **按算子路由后端** — 在同一设备名称下整合 FlagGems Triton 内核、厂商原生算子库、CUDA 兼容性 boxing 与显式 CPU 回退,并提供分平台路由表与逐算子覆盖。 -- **多平台支持** — NVIDIA CUDA、MetaX、华为 Ascend、PPU、海光 DCU、燧原 GCU、摩尔线程 MUSA 与地平线 BPU,各自拥有独立的构建选择器。 -- **训练栈** — eager 执行与 autograd、`torch.autocast("flagos")` 与 `torch.amp.GradScaler("flagos")`、`torch.compile` 集成,以及通过 `ProcessGroupFlagOS` 的 DDP/FSDP 支持。 -- **可观测性** — 带设备时间线与流向箭头的 `torch.profiler` 集成、逐算子调度与回退日志,以及带 `torch-fl-preflight` 检查工具的 wheel 兼容性清单。 -- **平台兼容性** — CUDA boxing、原生 ACLNN 与 topsaten 后端,以及 FlagGems Python 与 C++ 调度路径,并按平台记录状态等级。 +# 发布说明 + +本节包含 Torch-FL 的发布信息。 + +## v2.10.0 + +Torch-FL 作为 FlagOS 一部分的初始版本。 + +- **统一的 `flagos` 设备** — 基于 `PrivateUse1` 的 PyTorch 设备插件;标准 PyTorch API、张量方法与存储无需修改即可使用。 +- **按算子路由后端** — 在同一设备名称下整合 FlagGems Triton 内核、厂商原生算子库、CUDA 兼容性 boxing 与显式 CPU 回退,并提供分平台路由表与逐算子覆盖。 +- **多平台支持** — NVIDIA CUDA、MetaX、华为 Ascend、PPU、海光 DCU、燧原 GCU、摩尔线程 MUSA 与地平线 BPU,各自拥有独立的构建选择器。 +- **训练栈** — eager 执行与 autograd、`torch.autocast("flagos")` 与 `torch.amp.GradScaler("flagos")`、`torch.compile` 集成,以及通过 `ProcessGroupFlagOS` 的分布式支持(集合通信与 DDP;FSDP2 已在部分厂商平台验证)。 +- **可观测性** — 带设备时间线与流向箭头的 `torch.profiler` 集成、逐算子调度与回退日志,以及带 `torch-fl-preflight` 检查工具的 wheel 兼容性清单。 +- **平台兼容性** — CUDA boxing、原生 ACLNN 与 topsaten 后端,以及 FlagGems Python 与 C++ 调度路径,并按平台记录状态等级。