Skip to content

Implicit gradient performance improvements - #430

Merged
leburgel merged 5 commits into
mainfrom
lb/implicit_gradient_speedups
Sep 25, 2026
Merged

leburgel merged 5 commits into
mainfrom
lb/implicit_gradient_speedups

Conversation

@leburgel

@leburgel leburgel commented Sep 20, 2026 •

Copy link
Copy Markdown
Member

Performance improvements for the computation of implicit contraction gradients. Two main improvements:

  • Restrict the pullback that we use in the hot loop of the implicit gradient linear problem to a pure environment pullback that doesn't track adjoints related to the state variables. The final step after solving the linear problem still requires the state pullback, which we now have to generate separately.
  • Switch to an "implicit" null space projection instead of actually constructing the null spaces of the left and right isometries and multiplying with their adjoints, i.e. apply $(\mathbb{1} - U U^\dagger) x$ instead of constructing $U_L$ and applying $U_L U_L^\dagger x$. This isn't beneficial for the C4v case, but makes a significant difference for the asymmetric characteristic equations, even though this means using larger tensors as variables in the actual linear problem.

Benchmark matrix: main vs lb/implicit_gradient_speedups

Per cell: 3 seeds × 5 reps × 2 arms = 15 pairs. 1 Julia thread, 8 BLAS threads. Linear solver krylovdim=20, maxiter=3, so the operator-application budget is nops=62. main and branch run adjacent in time for each key, so the relative gain is unaffected even on a contended machine.

Half-infinite — SimultaneousCTMRG + HalfInfiniteProjector

D χ n solve main solve branch solve gain e2e main e2e branch e2e gain bytes
3 20 15 6.45 4.82 -26.9% 7.12 5.96 -18.5% -21.7%
3 30 15 14.17 10.54 -26.0% 16.06 13.01 -19.8% -23.4%
3 40 15 21.60 16.01 -26.6% 25.41 20.21 -22.0% -23.8%
4 30 15 26.33 16.51 -37.6% 32.69 22.91 -29.5% -29.8%
4 40 15 41.68 26.90 -35.3% 52.21 38.10 -27.7% -30.3%
4 50 15 75.65 47.56 -38.0% 91.67 64.05 -31.0% -30.5%
5 40 15 84.67 52.33 -39.0% 110.26 78.99 -30.3% -33.4%
5 60 15 248.51 129.24 -47.4% 314.15 190.32 -40.2% -32.8%
5 80 15 584.44 316.51 -44.9% 704.70 435.47 -37.3% -34.3%

C4v — C4vCTMRG + C4vEighProjector, RotateReflect

D χ n solve main solve branch solve gain e2e main e2e branch e2e gain bytes
3 20 15 0.110 0.090 -16.8% 0.260 0.243 -4.1% -5.4%
3 30 15 0.823 0.785 -6.5% 1.12 1.10 -1.6% -4.7%
3 40 15 0.887 0.840 -0.0% 1.99 2.02 3.0% -5.3%
4 30 15 2.74 2.38 -12.4% 4.79 3.96 -5.3% -9.0%
4 40 15 3.99 3.01 -24.1% 7.25 6.47 -10.9% -9.8%
4 50 15 3.85 3.17 -20.3% 9.23 8.27 -12.7% -8.1%
5 40 15 9.89 7.20 -28.9% 18.13 15.80 -17.5% -11.9%
5 60 15 31.69 20.18 -35.0% 49.02 37.87 -22.4% -12.4%
5 80 15 47.05 33.29 -31.7% 77.00 65.92 -17.2% -12.0%
  • Identical nops between arms in every pair — the change never alters the iteration count.
  • All gradient checksums agree to ~1e-14.

@leburgel
leburgel marked this pull request as draft September 20, 2026 09:39
@codecov

codecov Bot commented Sep 20, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 97.72727% with 1 line in your changes missing coverage. Please review.

Files with missing lines Patch % Lines
...lgorithms/optimization/implicit_differentiation.jl 95.45% 1 Missing ⚠️
Files with missing lines Coverage Δ
...hms/contractions/ctmrg/characteristic_equations.jl 93.14% <100.00%> (-0.19%) ⬇️
...lgorithms/optimization/implicit_differentiation.jl 89.12% <95.45%> (-0.11%) ⬇️
🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@leburgel leburgel changed the title [WIP] Implicit gradient performance improvements Implicit gradient performance improvements Sep 21, 2026
@leburgel
leburgel requested a review from pbrehmer September 21, 2026 12:02
@leburgel
leburgel marked this pull request as ready for review September 21, 2026 12:03

@pbrehmer pbrehmer left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is really neat and the benchmarks look great! I only have very minor comments. Given that there is a substantial speedup, I was wondering whether we should default to implicit differentiation now? I think for C4vCTMRG that would certainly make sense, for SimultaneousCTMRG it's maybe not as obvious? I guess it would be good to have a comparative benchmark of the sped-up implicit versus fixed-point differentiation; so in case that's just a quick Claude prompt away, maybe we could find out? :-)

Comment thread src/algorithms/contractions/ctmrg/characteristic_equations.jl Outdated
Comment thread src/algorithms/contractions/ctmrg/characteristic_equations.jl Outdated
@leburgel

Copy link
Copy Markdown
Member Author

Given that there is a substantial speedup, I was wondering whether we should default to implicit differentiation now?

I'd be surprised if the implicit approach could beat fixed-point differentiation with the full pullbacks using the un-truncated decompositions for the bond dimensions we're currently usually dealing with, but I'll try to set up some comparison to see how big the gap still is at this point.

@pbrehmer

Copy link
Copy Markdown
Collaborator

I'd be surprised if the implicit approach could beat fixed-point differentiation with the full pullbacks using the un-truncated decompositions for the bond dimensions we're currently usually dealing with, but I'll try to set up some comparison to see how big the gap still is at this point.

Of course you're right. I forgot that the default case is using untruncated decompositions that are hard to beat. Still it'd be interesting to see, thanks!

@leburgel

Copy link
Copy Markdown
Member Author

I ran into some unexpected issues while benchmarking so I don't have a full result, but so far the implicit approach consistently beats even the full pullback for the C4v case for all bond dimensions I've tried. So there we could switch the default already.

For the asymmetric case, the implicit approach is still about a factor 2 slower than using full pullbacks. For small bond dimensions, it's even slower than using the MatrixAlgebraKit.jl truncated pullbacks, and this only changes at larger bond dimensions. I'll keep working on getting a full set of benchmarks once I resolve my remaining issues, but I think we can keep a default-switch for a follow up. Also to make it more visible.

I'm quite happy with this by itself though, should be good to go as is as far as I'm concerned.

@pbrehmer

Copy link
Copy Markdown
Collaborator

I guess the CTMRG failures on v1.13 are unrelated to this PR, right? Also good to go for me in that case.

@leburgel
leburgel merged commit ccc2d4a into main Sep 25, 2026
73 of 78 checks passed
@leburgel
leburgel deleted the lb/implicit_gradient_speedups branch September 25, 2026 06:11
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants