KA kernel for add_transform for Abelian symmetries - #533
Conversation
41a979f to
0782850
Compare
Codecov Report✅ All modified and coverable lines are covered by tests.
🚀 New features to boost your workflow:
|
|
Update here: I'm trying to work out a way to make the whole |
d47e80f to
39153db
Compare
lkdvos
left a comment
There was a problem hiding this comment.
question because I'm just trying to understand:
I got somewhat confused by the dense_strides, since I don't think we can guarantee the data is actually laid out in dense-stride-form, but after looking at this a bit more, am I correct that the idea is that you are creating an additional set of linear indices that are related to the GPU workers/threads, and then mapping that to both source and destination data linear indices?
In other words, there are 3 sets of 1:nwork, the first is the data index of the src, the second is the data index of the dst, and then the index of the worker?
It seems like it would be nice to be able to map the worker index to the dst index, so we can guarantee contiguous writes for the workers, but I do agree that it is not obvious how that mapping would work (and if it is even possible to do that efficiently).
Anyways, this is less of a comment about having to change implementation and more just me double-checking I understood this correctly
Yes, that's the idea. |
Co-authored-by: Jutho <Jutho@users.noreply.github.com>
Co-authored-by: Jutho <Jutho@users.noreply.github.com>
36239fd to
fde8b48
Compare
|
Going to bypass the heckin slow ROCm queue here |
Split off the simpler Abelian component from #509. I've kept the device cache logic for now as it should still apply in this case (if I remember 3 weeks ago correctly...)