Skip to content

[TLE][Ascend] Add minimal INT8 GM-to-L1 ND2NZ primitive - #1323

Open
ggbondbest wants to merge 2 commits into
flagos-ai:triton_v3.5.xfrom
ggbondbest:ascend/int8-nd2nz-primitive
Open

ggbondbest wants to merge 2 commits into
flagos-ai:triton_v3.5.xfrom
ggbondbest:ascend/int8-nd2nz-primitive

Conversation

@ggbondbest

@ggbondbest ggbondbest commented Oct 2, 2026 •

Copy link
Copy Markdown

Add data_copy_gm_to_l1_nd2nz_int8, a TLE custom primitive for one signed INT8 GM-to-L1 ND→NZ transfer. It follows CANN 9.1's DataCopyGM2L1ND2NZImplBase and preserves all eight Nd2NzParams fields and their uint16_t types. C++ contains one copy_gm_to_cbuf_multi_nd2nz_b8 call; allocation, tiling, matrix multiplication, scaling and scheduling stay in the caller.

  • Register/export the primitive, include normal/mixed-core bitcode builds, and document the tl.dot consumer requirement.
  • Validate static parameter bounds and output capacity, including a static destination bound with dynamic source strides; reject Python booleans to avoid an i1/uint16_t ABI mismatch.
  • Add exact identity-matmul tests with random INT8 data, padded row strides, offsets and output guards, plus an optional equivalent-TLE A/B benchmark.

Validation on Ascend910B4-1 / CANN 9.1: all changed-file pre-commit checks, 3 valid and 33 rejected registration signatures, 6 device cases, and both manual and Ninja CMake custom-op bitcode builds passed. This builds the custom-op target, not the full compiler. Device coverage is nd_num=1; dynamic parameters have registration-only coverage. CANN 9.1 requires enable_legacy_insert_load_store_for_mix_cv=True for the existing custom-MTE2 output-memory-scope inference issue.

In the FlagGems-vLLM INT8 einsum consumer, switching only the input-copy path under identical compiler options gave median TLE/custom speedups of 1.307× (Flash B=128) and 1.095× (Pro B=4096), with per-call activation packing included and four paired blocks per shape. BF16 outputs matched exactly; FP32 checks passed. The compiler uses 4× temporary workspace for custom (256/512 KiB versus 64/128 KiB). Consumer PR: flagos-ai/FlagGems-vllm#894.

@CLAassistant

CLAassistant commented Oct 2, 2026 •

Copy link
Copy Markdown

CLA assistant check
All committers have signed the CLA.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CORE DOC Improvements or additions to documentation tle triton_v3.5.x

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants