Conversation
Adds Strided.batched_permutedims!/plan_batched_permutedims: permute a batch of differently-shaped tensors with one shared perm in a single CPU pass or a single GPU kernel launch, with selectable GPU execution strategies (elementwise, thread-tile, cooperative shared-memory tile) and optional per-tensor alpha/beta scaled accumulation. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Member
|
Could we get some AMD benchmarks as well? |
Member
Author
|
I don't have AMD access unfortunately |
Member
|
I do if you can send me the script you used for the benchmarks above! |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
WIP — not ready for review, needs cleanup before merging.
Adds
batched_permutedims!/plan_batched_permutedims: permute a batch of tensors sharing onepermin a single CPU pass or GPU kernel launch, with selectable GPU strategies (elementwise, thread-tile, cooperative shared-memory tile) and optional per-tensoralpha/betascaled accumulation.Still to do before this is mergeable: general read-through for remaining rough edges, decide on
BP_AUTO's default strategy per family, and reconcile with upstreammain.Benchmarks (RTX A6000, batched transposes; achievable ceiling via plain
copyto!≈ 640/660 GB/s F32/F64, not the ~768 GB/s spec figure)GB/s (min-of-20,
plan_batched_permutedimsbuilt once and reused):BP_ELEMENTWISE(default)BP_THREADTILEBP_GROUPTILEBP_ELEMENTWISE(default)BP_THREADTILEBP_GROUPTILEBP_ELEMENTWISE(default)BP_THREADTILEBP_GROUPTILEBP_GROUPTILEwins on every one of these transpose batches (1.6–4x the default), but only applies to the transpose family;BP_THREADTILEis the better fit for non-transpose (payload) permutations.BP_AUTOconservatively defaults toBP_ELEMENTWISEfor now — picking a smarter per-family default is one of the open items above. The "many tiny tensors" case (last row) is dominated by per-call host-side bookkeeping rather than kernel time; that overhead was a specific target of later work in this branch's history and is no longer the main bottleneck it once was.Scaled accumulation (
alpha/beta): omitting them costs nothing (identical to the plain copy). Withbeta == 0, overhead is ~1–4%. Withbeta != 0(destination read+accumulate) on a transpose,BP_ELEMENTWISEcosts ~3.8–4x the unscaled call, whileBP_GROUPTILEcosts only ~1.5x and is faster in absolute terms.🤖 Generated with Claude Code