Implicit gradient performance improvements - #430
Conversation
Codecov Report❌ Patch coverage is
🚀 New features to boost your workflow:
|
pbrehmer
left a comment
There was a problem hiding this comment.
This is really neat and the benchmarks look great! I only have very minor comments. Given that there is a substantial speedup, I was wondering whether we should default to implicit differentiation now? I think for C4vCTMRG that would certainly make sense, for SimultaneousCTMRG it's maybe not as obvious? I guess it would be good to have a comparative benchmark of the sped-up implicit versus fixed-point differentiation; so in case that's just a quick Claude prompt away, maybe we could find out? :-)
Co-authored-by: Paul Brehmer <paul.brehmer@univie.ac.at>
I'd be surprised if the implicit approach could beat fixed-point differentiation with the full pullbacks using the un-truncated decompositions for the bond dimensions we're currently usually dealing with, but I'll try to set up some comparison to see how big the gap still is at this point. |
Of course you're right. I forgot that the default case is using untruncated decompositions that are hard to beat. Still it'd be interesting to see, thanks! |
|
I ran into some unexpected issues while benchmarking so I don't have a full result, but so far the implicit approach consistently beats even the full pullback for the C4v case for all bond dimensions I've tried. So there we could switch the default already. For the asymmetric case, the implicit approach is still about a factor 2 slower than using full pullbacks. For small bond dimensions, it's even slower than using the MatrixAlgebraKit.jl truncated pullbacks, and this only changes at larger bond dimensions. I'll keep working on getting a full set of benchmarks once I resolve my remaining issues, but I think we can keep a default-switch for a follow up. Also to make it more visible. I'm quite happy with this by itself though, should be good to go as is as far as I'm concerned. |
|
I guess the CTMRG failures on |
Performance improvements for the computation of implicit contraction gradients. Two main improvements:
Benchmark matrix:
mainvslb/implicit_gradient_speedupsPer cell: 3 seeds × 5 reps × 2 arms = 15 pairs. 1 Julia thread, 8 BLAS threads. Linear solver
krylovdim=20,maxiter=3, so the operator-application budget isnops=62.mainandbranchrun adjacent in time for each key, so the relative gain is unaffected even on a contended machine.Half-infinite —
SimultaneousCTMRG+HalfInfiniteProjectorC4v —
C4vCTMRG+C4vEighProjector,RotateReflectnopsbetween arms in every pair — the change never alters the iteration count.