Skip to content

Edge-frame kernels: packed-D tile IO (measurement-driven), sweepable warps - #15

Merged
alacour merged 6 commits into
perf/equivariant-linear-blockgemmfrom
perf/edge-frame-bwd-kernels
Sep 4, 2026
Merged

alacour merged 6 commits into
perf/equivariant-linear-blockgemmfrom
perf/edge-frame-bwd-kernels

Conversation

@alacour

@alacour alacour commented Sep 4, 2026

Copy link
Copy Markdown
Collaborator

Summary

Stacked on #14 (merge that first). Optimizes the edge-frame Triton kernel family — the largest remaining kernel cost (~25 ms/step at 512 atoms) — with each step driven by a measurement:

  1. ECENET_EF_WARPS: warps-per-program env knob on all launches. An A100 sweep (1/2/4/8) was flat → occupancy ruled out.
  2. ncu diagnosis: merged backward kernels at 75–86% L1/TEX with DRAM idle at 9–18% — L1-bound on the scalar per-element D gathers through the (l,m) column tables and matching dD scatter-stores.
  3. Packed-D kernels: D pre-packed once per step into dense (E,S,P)/(E,P,S) tensors (cached on the shared per-step D_block object; per-call packing was measured to erase the win) and dD unpacked with cached integer indices (boolean-mask indexing forced a host sync per backward — 10.6 ms CPU/call, caught by torch.profiler's CPU column). Kernels do only vectorized coalesced tile IO.
  4. Same treatment for the forward kernels (_ef_fwd, _ef_bwd_dx-as-PU-forward), which carry the identical gather pattern.

Measured (A100, 512-atom diamond, 44k edges, float32+TF32, all fusion flags):

  • _ef_bwd_merged 2.33 → 1.35 ms/call; _pu_bwd_merged 1.98 → 1.65
  • edge-frame fused forward 1.26 → 0.74 ms; pack+unrotate 1.09 → 0.63
  • joint E+F step 111.4 → 104.1 ms (4.60 → 4.92 atoms/ms; cumulative from the session's 218.6 baseline: 2.10×)

Packed-D is default on; ECENET_EF_PACKD=0 restores the table-gather kernels.

Test plan

  • tests/test_edge_frame_kernel.py passes on CPU; test_triton_packed_d runs the full fp64 comparison suite (EdgeFrameFused, PackUnrotateFused, e2n, single-source) under both packed settings on CUDA.
  • ⚠️ Run python tests/test_edge_frame_kernel.py once on a GPU box before merging.

🤖 Generated with Claude Code

https://claude.ai/code/session_019WyADHSVvKbQnYo1PAjGPK

alacour and others added 6 commits September 3, 2026 19:20
The edge-frame family is the largest remaining kernel cost (~25 ms/step at
512 atoms: _ef_bwd_merged 2.33 ms/call, _pu_bwd_merged 1.98, fwd ~1.0), all
running ~3.5x above their traffic bound. Unlike the RealSpace case the loads
are already coalesced tiles, so the gap is some mix of padded ieee tl.dot
arithmetic (9 valid of 16 in both dot dims), 36-byte row misalignment, and
per-edge program overhead — not separable without measurement. First step:
make warps-per-program env-tunable (default 4, the previous implicit value)
on all ten launch sites so the cheap axis can be swept before any redesign.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019WyADHSVvKbQnYo1PAjGPK
ncu on the merged backward kernels (A100, 512-atom box, 44k edges) settles
the question the flat warps sweep left open: L1/TEX throughput 75-86% with
DRAM at 9-18% and compute ~30% — the kernels are L1-bound, stalling on MIO
short-scoreboard, while DRAM idles. The amplification is the per-element
gathered D loads through the (l,m) column tables and the matching dD
scatter-stores: hundreds of non-vectorizable L1 transactions per program.

The packed variants spend the idle DRAM instead: the wrapper pre-packs D
into dense (E, S, P)/(E, P, S) tensors once per call (plain torch indexing,
~14 MB at 44k edges) and unpacks the packed dD afterwards, so the kernels
do only vectorized coalesced tile IO. Math and masks identical. Gated off
by default behind ECENET_EF_PACKD=1 (module flag, monkeypatchable) pending
an A/B on the A100; ncu's stall-fix estimate is ~33-37% on these kernels.

test_triton_packed_d reruns the full test_triton_paths comparison set (both
merged backwards, vs fp64 eager truth) with the flag on.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019WyADHSVvKbQnYo1PAjGPK
…block

The packed kernels won at the kernel level (ncu: _ef_bwd_merged 2.94 -> 1.68
ms, _pu_bwd_merged 2.49 -> 1.95 ms) but the joint step stayed flat: per-call
packing/transposing in the backward wrappers (~8 small torch ops x 7
backward calls/step) ate the entire ~5 ms win.

The model builds ONE D_block per step and hands the same Python object to
every fused edge-frame op, so _get_packed_D now caches the packed tensors as
an attribute on that object — packing runs once per step. Function forwards
fetch it and carry the tensors through save_for_backward (attributes do not
survive the re-wrapping); the backward wrappers receive them instead of
packing. A fresh step's fresh D_block starts clean, so no cross-step
staleness. Still gated behind ECENET_EF_PACKD=1 pending the A/B.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019WyADHSVvKbQnYo1PAjGPK
torch.profiler with the packed path on showed the GPU win arriving (packed
kernels 1.35/1.65 ms per call, CUDA total down) while wall time stayed flat,
and named the thief: EdgeFrameFusedBackward at 10.6 ms CPU per call, with
aten::index at 44% of CPU total. _unpack_dD's boolean-mask indexing
(cos_col[vc], dD[:, :, col[vc]]) forces a host-device sync on every backward
call — seven pipeline stalls per step. The masks are static, so the unpack
now uses integer index tensors precomputed once per (S, n_ang, device);
every op in the unpack is async.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019WyADHSVvKbQnYo1PAjGPK
…ault on

With the sync-free unpack the packed backward finally reached the wall
clock: joint force step 111.4 -> 107.8 ms on the A100 512-atom box
(EdgeFrameFusedBackward CPU 10.6 ms -> 0.29 ms per call, aten::index off
the CPU hot list, packed kernels at 1.35/1.65 ms in-app).

This extends the same treatment to the forward side, which carries the
identical table-gather pattern: _ef_fwd_packed_kernel (EdgeFrameFused
forward, ~1.04 ms x 4/step) and _ef_bwd_dx_packed_kernel (PackUnrotate's
forward contraction, ~1.17 ms x 3/step) read the pre-packed Dc/Ds / DcT/DsT
tiles instead of gathering through the column tables; amortization is free
since the packed tensors already live on the shared per-step D_block.

Default flipped ON (ECENET_EF_PACKD=0 restores the table-gather kernels);
test_triton_packed_d now runs the full fp64 comparison suite under both
settings.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019WyADHSVvKbQnYo1PAjGPK
@alacour
alacour merged commit fdf3e49 into perf/equivariant-linear-blockgemm Sep 4, 2026
1 check passed
@alacour
alacour deleted the perf/edge-frame-bwd-kernels branch September 4, 2026 03:34
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant