Skip to content

KA kernel for add_transform for Abelian symmetries - #533

Merged
kshyatt merged 16 commits into
mainfrom
ksh/add_transform_abelian
Sep 29, 2026
Merged

kshyatt merged 16 commits into
mainfrom
ksh/add_transform_abelian

Conversation

@kshyatt

@kshyatt kshyatt commented Sep 15, 2026

Copy link
Copy Markdown
Member

Split off the simpler Abelian component from #509. I've kept the device cache logic for now as it should still apply in this case (if I remember 3 weeks ago correctly...)

@kshyatt
kshyatt requested a review from lkdvos September 15, 2026 15:08
@kshyatt
kshyatt force-pushed the ksh/add_transform_abelian branch from 41a979f to 0782850 Compare September 15, 2026 15:14
@codecov

codecov Bot commented Sep 16, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

Files with missing lines Coverage Δ
ext/TensorKitGPUArraysExt.jl 93.68% <100.00%> (+4.59%) ⬆️
🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@lkdvos

lkdvos commented Sep 24, 2026

Copy link
Copy Markdown
Member

Update here: I'm trying to work out a way to make the whole TreeTransformer that ends up being cached also depend on the storagetype. This has the added benefit that we can store complex unitaries for complex tensors, therefore enabling this to go through BLAS, while also having the ability to return custom types for host vs device implementations. I'm hoping this will make this PR consist only of the actual kernel implementation

Comment thread ext/TensorKitGPUArraysExt.jl Outdated

@lkdvos lkdvos left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

question because I'm just trying to understand:

I got somewhat confused by the dense_strides, since I don't think we can guarantee the data is actually laid out in dense-stride-form, but after looking at this a bit more, am I correct that the idea is that you are creating an additional set of linear indices that are related to the GPU workers/threads, and then mapping that to both source and destination data linear indices?
In other words, there are 3 sets of 1:nwork, the first is the data index of the src, the second is the data index of the dst, and then the index of the worker?

It seems like it would be nice to be able to map the worker index to the dst index, so we can guarantee contiguous writes for the workers, but I do agree that it is not obvious how that mapping would work (and if it is even possible to do that efficiently).

Anyways, this is less of a comment about having to change implementation and more just me double-checking I understood this correctly

Comment thread ext/TensorKitGPUArraysExt.jl Outdated
Comment thread ext/TensorKitGPUArraysExt.jl
Comment thread ext/TensorKitGPUArraysExt.jl Outdated
Comment thread Project.toml Outdated
@kshyatt

kshyatt commented Sep 28, 2026

Copy link
Copy Markdown
Member Author

I got somewhat confused by the dense_strides, since I don't think we can guarantee the data is actually laid out in dense-stride-form, but after looking at this a bit more, am I correct that the idea is that you are creating an additional set of linear indices that are related to the GPU workers/threads, and then mapping that to both source and destination data linear indices?
In other words, there are 3 sets of 1:nwork, the first is the data index of the src, the second is the data index of the dst, and then the index of the worker?

Yes, that's the idea.

Comment thread ext/TensorKitGPUArraysExt.jl Outdated
Comment thread ext/TensorKitGPUArraysExt.jl Outdated
Comment thread ext/TensorKitGPUArraysExt.jl
Comment thread ext/TensorKitGPUArraysExt.jl Outdated
Comment thread ext/TensorKitGPUArraysExt.jl Outdated
Comment thread ext/TensorKitGPUArraysExt.jl Outdated
Comment thread ext/TensorKitGPUArraysExt.jl Outdated
@kshyatt
kshyatt force-pushed the ksh/add_transform_abelian branch from 36239fd to fde8b48 Compare September 29, 2026 06:37
@kshyatt

kshyatt commented Sep 29, 2026

Copy link
Copy Markdown
Member Author

Going to bypass the heckin slow ROCm queue here

@kshyatt
kshyatt merged commit 440722d into main Sep 29, 2026
65 of 66 checks passed
@kshyatt
kshyatt deleted the ksh/add_transform_abelian branch September 29, 2026 11:54
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants