Skip to content

Add dedicated cached treebraider for AdjointTensorMap - #519

Closed
leburgel wants to merge 2 commits into
mainfrom
lb/adjoint_treebraider_treetransposer
Closed

leburgel wants to merge 2 commits into
mainfrom
lb/adjoint_treebraider_treetransposer

Conversation

@leburgel

@leburgel leburgel commented Aug 27, 2026 •

Copy link
Copy Markdown
Member

One option for a proper fix for #516, adding a dedicated cached conj_treebraider to handle permutations and transposes of AdjointTensorMaps with a plain TensorMap parent using the same infrastructure as the pure TensorMap path.

Some timings based on the reproducer of #516:

permute! — adjoint source vs plain TensorMap:

sector type before after plain ratio before ratio after speedup
fℤ₂ 0.0730 ms 0.0790 ms 0.0620 ms 1.32× 1.27× 0.92×
fℤ₂ ⊠ U(1) 0.1295 ms 0.0367 ms 0.0345 ms 3.78× 1.06× 3.53×
U(1) 0.1012 ms 0.0302 ms 0.0278 ms 3.64× 1.09× 3.35×
SU(2) 2.5470 ms 0.3867 ms 0.3621 ms 7.26× 1.07× 6.59×
fℤ₂ ⊠ U(1) ⊠ SU(3) 3.9624 ms 0.0499 ms 0.0500 ms 78.78× 1.00× 79.4×

Allocations for the adjoint path drop to exactly the plain path's: 794→310, 662→310, 26 883→2840, 42 653→645.

@leburgel leburgel changed the title Add dedicated cached treebraider and treetransposer for AdjointTensorMap Add dedicated cached treebraider for AdjointTensorMap Aug 27, 2026
@codecov

codecov Bot commented Aug 27, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

Files with missing lines Coverage Δ
src/spaces/homspace.jl 91.83% <100.00%> (+0.46%) ⬆️
src/tensors/indexmanipulations.jl 88.44% <100.00%> (+0.09%) ⬆️
src/tensors/treetransformers.jl 82.40% <100.00%> (-0.12%) ⬇️
🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

lkdvos added a commit that referenced this pull request Sep 9, 2026
Index manipulations now run through a single kernel that operates on
subblocks addressed by position: `StridedSubblocks` (sector-independent
views into the flat data of a `TensorMap`) or `TreeSubblocks` (any
`AbstractTensorMap`, through `subblock`), both carrying an optional lazy
conjugation. `TreeTransformer`s store only the mapping between subblock
positions and recoupling coefficients, alongside the subblock structures,
and are cached for every tensor type.

Adjoint sources and destinations, as well as `conj` in `tensoradd!`, are
folded into a conjugation flag, relabeled permutation and levels, and
conjugated scalars, so that `AdjointTensorMap` wrappers no longer force
the uncached generic path (fixes #516, supersedes #519 and #520).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@lkdvos

lkdvos commented Sep 9, 2026

Copy link
Copy Markdown
Member

Closing this in favor of #526

@lkdvos lkdvos closed this Sep 9, 2026
lkdvos added a commit that referenced this pull request Sep 9, 2026
Index manipulations now run through a single kernel that operates on
subblocks addressed by position: `StridedSubblocks` (sector-independent
views into the flat data of a `TensorMap`) or `TreeSubblocks` (any
`AbstractTensorMap`, through `subblock`), both carrying an optional lazy
conjugation. `TreeTransformer`s store only the mapping between subblock
positions and recoupling coefficients, alongside the subblock structures,
and are cached for every tensor type.

Adjoint sources and destinations, as well as `conj` in `tensoradd!`, are
folded into a conjugation flag, relabeled permutation and levels, and
conjugated scalars, so that `AdjointTensorMap` wrappers no longer force
the uncached generic path (fixes #516, supersedes #519 and #520).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
lkdvos added a commit that referenced this pull request Sep 15, 2026
Index manipulations now run through a single kernel that operates on
subblocks addressed by position: `StridedSubblocks` (sector-independent
views into the flat data of a `TensorMap`) or `TreeSubblocks` (any
`AbstractTensorMap`, through `subblock`), both carrying an optional lazy
conjugation. `TreeTransformer`s store only the mapping between subblock
positions and recoupling coefficients, alongside the subblock structures,
and are cached for every tensor type.

Adjoint sources and destinations, as well as `conj` in `tensoradd!`, are
folded into a conjugation flag, relabeled permutation and levels, and
conjugated scalars, so that `AdjointTensorMap` wrappers no longer force
the uncached generic path (fixes #516, supersedes #519 and #520).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
lkdvos added a commit that referenced this pull request Sep 16, 2026
Index manipulations now run through a single kernel that operates on
subblocks addressed by position: `StridedSubblocks` (sector-independent
views into the flat data of a `TensorMap`) or `TreeSubblocks` (any
`AbstractTensorMap`, through `subblock`), both carrying an optional lazy
conjugation. `TreeTransformer`s store only the mapping between subblock
positions and recoupling coefficients, alongside the subblock structures,
and are cached for every tensor type.

Adjoint sources and destinations, as well as `conj` in `tensoradd!`, are
folded into a conjugation flag, relabeled permutation and levels, and
conjugated scalars, so that `AdjointTensorMap` wrappers no longer force
the uncached generic path (fixes #516, supersedes #519 and #520).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
lkdvos added a commit that referenced this pull request Sep 17, 2026
Index manipulations now run through a single kernel that operates on
subblocks addressed by position: `StridedSubblocks` (sector-independent
views into the flat data of a `TensorMap`) or `TreeSubblocks` (any
`AbstractTensorMap`, through `subblock`), both carrying an optional lazy
conjugation. `TreeTransformer`s store only the mapping between subblock
positions and recoupling coefficients, alongside the subblock structures,
and are cached for every tensor type.

Adjoint sources and destinations, as well as `conj` in `tensoradd!`, are
folded into a conjugation flag, relabeled permutation and levels, and
conjugated scalars, so that `AdjointTensorMap` wrappers no longer force
the uncached generic path (fixes #516, supersedes #519 and #520).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
lkdvos added a commit that referenced this pull request Sep 20, 2026
…#526)

* Refactor index manipulation kernels around position-indexed subblocks

Index manipulations now run through a single kernel that operates on
subblocks addressed by position: `StridedSubblocks` (sector-independent
views into the flat data of a `TensorMap`) or `TreeSubblocks` (any
`AbstractTensorMap`, through `subblock`), both carrying an optional lazy
conjugation. `TreeTransformer`s store only the mapping between subblock
positions and recoupling coefficients, alongside the subblock structures,
and are cached for every tensor type.

Adjoint sources and destinations, as well as `conj` in `tensoradd!`, are
folded into a conjugation flag, relabeled permutation and levels, and
conjugated scalars, so that `AdjointTensorMap` wrappers no longer force
the uncached generic path (fixes #516, supersedes #519 and #520).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* blockiterator->blockiterators

* clean up trivial symmetry bypassing overhead

* Address review: normalize subblock parent, prime conjsrc

`StridedSubblocks` now stores its data the way `StridedView` parents it
(an `Array` becomes its underlying `Memory` on Julia >= 1.11), asking
`StridedView` itself rather than reproducing that rule. This makes the
view type a direct function of the type parameters, so `eltype` can be
written out instead of going through `Core.Compiler.return_type`.

As a consequence `storagetype` reports the normalized type, which is not
an `Array`, so the CPU branch of `_adapt_recoupling` is keyed on a
`CPUStorage` alias to keep the recoupling matrices off the `Adapt` path.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* Test fix: pick a multifusion-compatible space for the isometry test

`V1 ⊗ V2 ← V3 ⊗ V4` does not close the unit cycle that `GenericUnit`
sectors require, so constructing it threw a `SpaceMismatch` for the
`IsingBimodule` space lists. The sibling `Permutations: adjoint operands`
testset never hit this because it sits behind `symmetricbraiding`, which
is false for multifusion; this one builds its space unconditionally.

`V1 ⊗ V5 ← V2 ⊗ V4` does close the cycle, and is equally general for the
other sectors, so use that rather than skipping multifusion: the
transpose half of the test then covers them too, while the braid half
stays behind `hasbraiding` (multifusion is `NoBraiding`).

Only Windows and macOS saw this, since `default_spacelist` hands out
different space lists per OS on CI and only those two include the
multifusion entries.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* reorganize to fix docstring

* simplify getting number of transformer_threads

* mark `add_transform` as non-public

* fix (unrelated) docstring sentence

* all has_array_view

* clamp -> min

* docstring improvements

* simplify treetransformer implementations

* Rename `AbelianTreeTransformer` to `UniqueTreeTransformer`

"Abelian" is ambiguous for sectors: it can refer either to the fusion of two
sectors having a unique result, or to the commutativity of the fusion rules.
The transformer is selected on `FusionStyle(I) == UniqueFusion()`, so name it
after that.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* Pass conjugation as a flag instead of a subblock view op

`StridedSubblocks` and `TreeSubblocks` no longer apply `identity`/`conj` to
every view. Instead `conjsrc` is threaded through `add_transform_kernel!` into
`_add_transform_block!`, where it is handed to `TO.tensoradd!` as its `conjA`
argument, at the single-tree call and when packing a multi-tree block.

This drops a type parameter from both collections, so the kernel compiles to
one instance per (storage, numind) rather than one per conjugation. The runtime
flag is free: `flag2op` is union-split, and `conj` of a real-eltype
`StridedView` is a type-level no-op, which also makes the previous
`scalartype(t) <: Real` guard redundant.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* remove TreeSubblocks and route through StridedSubblocks

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants