Repository navigation
Replies: 1 comment
|
Hi Matt, Sounds like a good proposal. It is on my TODO list now. However, this may not rank very high in priority compared with other features that I am planning on. My reasoning is as follows: In my understanding, the primary optimization for a batch tensordot would come from the reduction of the number of call overheads of torch functions. In the actual large scale computations, primary complexity comes from the actual BLAS operations. Therefore, reduction of call overheads will become marginal for sufficiently large scale scenarios. And the optimization works primarily for situations where tensors have many and small blocks. Please correct me if I have overlooked something. Also, it would help me better evaluate the priorities if you happen to have some real-application level benchmarks to show how much impact this could bring about. And thank you for initiating this discussion. |
Uh oh!
There was an error while loading. Please reload this page.
Hi, thanks for Nicole, it is a great library.
I noticed that the current contract() implementation dispatches one torch.matmul per surviving block pair. For Abelian (U(1)) tensors, multiple block pairs often share the same contracted charge sector and could be stacked into a single larger matrix and multiplied with one BLAS call. This is the approach used by TeNPy and ITensor internally, and it makes a significant difference once the number of block pairs grows.
As I understand this may be intentional given that Nicole targets SU(2) as well, where the intertwiner recoupling is block-pair-specific and cannot be naively batched. But for purely Abelian tensors (intw is None) the stacking is straightforward and would bring Nicole much closer to the performance of U(1)-specialized libraries.
Would you consider a fast path in contract() for the Abelian case? Happy to discuss.
Matt
All reactions