Skip to content

Guard optimizers and linesearch against non-descent directions and zero steps - #48

Merged
Jutho merged 3 commits into
Jutho:masterfrom
lkdvos:harden-zero-step
Oct 2, 2026
Merged

Jutho merged 3 commits into
Jutho:masterfrom
lkdvos:harden-zero-step

Conversation

@lkdvos

@lkdvos lkdvos commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor

A non-descent search direction (e.g. from a preconditioner that is indefinite in floating point, see QuantumKitHub/MPSKit.jl#532) makes the linesearch return a zero step, which no optimizer handled:

  • ConjugateGradient hangs: β becomes 0/0 = NaN, the NaN direction passes the descent check, and the bracket expansion loops forever. A zero initial guess (2 * 0) loops there too.
  • LBFGS and GradientDescent retry the zero step until maxiter.
fg(x) = (dot(x, x) / 2, copy(x))
precondition(x, g) = [g[1], -g[2] / 2]
optimize(fg, [1.0, 0.1], ConjugateGradient(; gradtol = 1e-12); precondition)  # hangs

Changes

  • Optimizers check for descent before the linesearch. CG restarts from the preconditioned gradient on a bad or non-finite β; LBFGS resets its inverse Hessian. They stop with a warning if the preconditioned gradient is not a descent direction or the linesearch makes no progress. CG no longer doubles a zero step.
  • HagerZhangLineSearch rejects invalid initial guesses, treats a NaN slope as non-descent, and bounds the bracket expansion by maxfg.

Cost: one extra inner per iteration.

Tests

New tests for non-descent directions (all optimizers), CG with β = NaN or −1e3, and the linesearch guards. They fail or hang on master and pass here.

🤖 Generated with Claude Code

…ro steps

A non-descent search direction (for example from a preconditioner that is
indefinite in finite precision) made the linesearch return a zero step, which
none of the optimizers handled:

- ConjugateGradient: the next Hager-Zhang β is 0/0 = NaN; the resulting NaN
  direction passed the descent check of the linesearch, whose bracketing phase
  then loops forever. A zero initial guess (2 * 0) also loops forever there.
- LBFGS and GradientDescent: the same zero step is retried until `maxiter`.

The optimizers now check for descent before the linesearch. ConjugateGradient
restarts from the preconditioned gradient when β is not finite or the
conjugate direction is not a descent direction, and LBFGS resets its inverse
Hessian approximation. If the preconditioned gradient itself is not a descent
direction, or the linesearch makes no progress along it, they stop with a
warning. The linesearch rejects invalid initial guesses, treats a NaN slope as
non-descent, and bounds its bracket expansion by `maxfg`.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@codecov-commenter

codecov-commenter commented Oct 1, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 77.08333% with 11 lines in your changes missing coverage. Please review.
✅ Project coverage is 83.29%. Comparing base (4fc7181) to head (fd8870d).

Files with missing lines Patch % Lines
src/lbfgs.jl 64.28% 5 Missing ⚠️
src/cg.jl 87.50% 2 Missing ⚠️
src/gd.jl 71.42% 2 Missing ⚠️
src/linesearches.jl 77.77% 2 Missing ⚠️
Additional details and impacted files
@@            Coverage Diff             @@
##           master      #48      +/-   ##
==========================================
+ Coverage   82.31%   83.29%   +0.97%     
==========================================
  Files           8        8              
  Lines         752      802      +50     
==========================================
+ Hits          619      668      +49     
- Misses        133      134       +1     

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@lkdvos

lkdvos commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor Author

Closes #43.

Partially addresses #44: adds warnings, a CG restart and an LBFGS memory reset. GD and CG no longer get stuck: with a wrong-sign gradient (fg(x) = (x[1]^2, [-1.0])), master hangs and this PR stops after one iteration. LBFGS still runs to maxiter there, because the failing linesearch returns a tiny nonzero step within the ϵ tolerance. Fixing that needs the linesearch to report failure (see #15). The forced-step suggestion is not implemented.

Related, not closed: #15 (only the zero-step case is covered), #11 (bisect is already bounded on master; this PR bounds bracket), #16 (already works on master).

Does not supersede #19. safe_sqrt guards negative inner(x, v, v); this PR guards the sign of inner(x, g, η). The negative inner product in the MPSKit case is fixed at the source in QuantumKitHub/MPSKit.jl#532.

Comment thread src/cg.jl Outdated
Co-authored-by: Jutho <Jutho@users.noreply.github.com>
Comment thread src/linesearches.jl Outdated

@Jutho Jutho left a comment

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nice; great to see some of these infinite loops and other instabilities gone.

@lkdvos

lkdvos commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor Author

I don't have any write access here so you'll have to merge/tag this one I fear :)

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants