Repository navigation
Edge-frame kernels: packed-D tile IO (measurement-driven), sweepable warps - #15
Merged
alacour merged 6 commits intoSep 4, 2026
Merged
Conversation
The edge-frame family is the largest remaining kernel cost (~25 ms/step at 512 atoms: _ef_bwd_merged 2.33 ms/call, _pu_bwd_merged 1.98, fwd ~1.0), all running ~3.5x above their traffic bound. Unlike the RealSpace case the loads are already coalesced tiles, so the gap is some mix of padded ieee tl.dot arithmetic (9 valid of 16 in both dot dims), 36-byte row misalignment, and per-edge program overhead — not separable without measurement. First step: make warps-per-program env-tunable (default 4, the previous implicit value) on all ten launch sites so the cheap axis can be swept before any redesign. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019WyADHSVvKbQnYo1PAjGPK
ncu on the merged backward kernels (A100, 512-atom box, 44k edges) settles the question the flat warps sweep left open: L1/TEX throughput 75-86% with DRAM at 9-18% and compute ~30% — the kernels are L1-bound, stalling on MIO short-scoreboard, while DRAM idles. The amplification is the per-element gathered D loads through the (l,m) column tables and the matching dD scatter-stores: hundreds of non-vectorizable L1 transactions per program. The packed variants spend the idle DRAM instead: the wrapper pre-packs D into dense (E, S, P)/(E, P, S) tensors once per call (plain torch indexing, ~14 MB at 44k edges) and unpacks the packed dD afterwards, so the kernels do only vectorized coalesced tile IO. Math and masks identical. Gated off by default behind ECENET_EF_PACKD=1 (module flag, monkeypatchable) pending an A/B on the A100; ncu's stall-fix estimate is ~33-37% on these kernels. test_triton_packed_d reruns the full test_triton_paths comparison set (both merged backwards, vs fp64 eager truth) with the flag on. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019WyADHSVvKbQnYo1PAjGPK
…block The packed kernels won at the kernel level (ncu: _ef_bwd_merged 2.94 -> 1.68 ms, _pu_bwd_merged 2.49 -> 1.95 ms) but the joint step stayed flat: per-call packing/transposing in the backward wrappers (~8 small torch ops x 7 backward calls/step) ate the entire ~5 ms win. The model builds ONE D_block per step and hands the same Python object to every fused edge-frame op, so _get_packed_D now caches the packed tensors as an attribute on that object — packing runs once per step. Function forwards fetch it and carry the tensors through save_for_backward (attributes do not survive the re-wrapping); the backward wrappers receive them instead of packing. A fresh step's fresh D_block starts clean, so no cross-step staleness. Still gated behind ECENET_EF_PACKD=1 pending the A/B. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019WyADHSVvKbQnYo1PAjGPK
torch.profiler with the packed path on showed the GPU win arriving (packed kernels 1.35/1.65 ms per call, CUDA total down) while wall time stayed flat, and named the thief: EdgeFrameFusedBackward at 10.6 ms CPU per call, with aten::index at 44% of CPU total. _unpack_dD's boolean-mask indexing (cos_col[vc], dD[:, :, col[vc]]) forces a host-device sync on every backward call — seven pipeline stalls per step. The masks are static, so the unpack now uses integer index tensors precomputed once per (S, n_ang, device); every op in the unpack is async. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019WyADHSVvKbQnYo1PAjGPK
…ault on With the sync-free unpack the packed backward finally reached the wall clock: joint force step 111.4 -> 107.8 ms on the A100 512-atom box (EdgeFrameFusedBackward CPU 10.6 ms -> 0.29 ms per call, aten::index off the CPU hot list, packed kernels at 1.35/1.65 ms in-app). This extends the same treatment to the forward side, which carries the identical table-gather pattern: _ef_fwd_packed_kernel (EdgeFrameFused forward, ~1.04 ms x 4/step) and _ef_bwd_dx_packed_kernel (PackUnrotate's forward contraction, ~1.17 ms x 3/step) read the pre-packed Dc/Ds / DcT/DsT tiles instead of gathering through the column tables; amortization is free since the packed tensors already live on the shared per-step D_block. Default flipped ON (ECENET_EF_PACKD=0 restores the table-gather kernels); test_triton_packed_d now runs the full fp64 comparison suite under both settings. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019WyADHSVvKbQnYo1PAjGPK
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019WyADHSVvKbQnYo1PAjGPK
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Stacked on #14 (merge that first). Optimizes the edge-frame Triton kernel family — the largest remaining kernel cost (~25 ms/step at 512 atoms) — with each step driven by a measurement:
ECENET_EF_WARPS: warps-per-program env knob on all launches. An A100 sweep (1/2/4/8) was flat → occupancy ruled out.(E,S,P)/(E,P,S)tensors (cached on the shared per-stepD_blockobject; per-call packing was measured to erase the win) and dD unpacked with cached integer indices (boolean-mask indexing forced a host sync per backward — 10.6 ms CPU/call, caught by torch.profiler's CPU column). Kernels do only vectorized coalesced tile IO._ef_fwd,_ef_bwd_dx-as-PU-forward), which carry the identical gather pattern.Measured (A100, 512-atom diamond, 44k edges, float32+TF32, all fusion flags):
_ef_bwd_merged2.33 → 1.35 ms/call;_pu_bwd_merged1.98 → 1.65edge-frame fusedforward 1.26 → 0.74 ms;pack+unrotate1.09 → 0.63Packed-D is default on;
ECENET_EF_PACKD=0restores the table-gather kernels.Test plan
tests/test_edge_frame_kernel.pypasses on CPU;test_triton_packed_druns the full fp64 comparison suite (EdgeFrameFused, PackUnrotateFused, e2n, single-source) under both packed settings on CUDA.python tests/test_edge_frame_kernel.pyonce on a GPU box before merging.🤖 Generated with Claude Code
https://claude.ai/code/session_019WyADHSVvKbQnYo1PAjGPK