Skip to content

docs(fdfd): parallel ILU(0) measured on our matrices, both halves, and not used - #150

Merged
tachsin merged 1 commit into
mainfrom
perf/iterative-triangular
Oct 5, 2026
Merged

tachsin merged 1 commit into
mainfrom
perf/iterative-triangular

Conversation

@tachsin

@tachsin tachsin commented Oct 5, 2026

Copy link
Copy Markdown
Owner

Closes #141, by measurement: neither half of a deterministic parallel ILU(0) pays on photonoxide's matrices, so neither is used.

The factorization (Chow & Patel 2015, doi:10.1137/140968896)

Not built, measured first (comment on #141): the multigrid's smoothers already factorize slab by slab in parallel (about 1.3 s of Diel's hierarchy), and QMR + ILU(0) on the guide spends 4.0 of 5.6 s in its iterations. Chow and Patel's own results (their Section 4.2, Table 3) need 3 to 5 synchronous sweeps on matrices far from diagonally dominant.

The triangular solves (Anzt, Chow & Dongarra 2015, doi:10.1007/978-3-662-48096-0_50)

Built and measured: each triangular solve as k synchronous Jacobi sweeps, all rows in parallel, the same bits on any number of threads; on the transposed factors exactly the transpose of the same series (D⁻¹(I − TᵀD⁻¹)ʲ = (I − D⁻¹Tᵀ)ʲD⁻¹), so QMR stays consistent. The 40³ guide (Shin and Fan's operator, stretched PMLs), QMR + ILU(0) to 1e-8, 20 threads:

triangular solves exact 1 sweep 2 3 4 6 8
QMR iterations 195 1 163 747 730 929 1 782 2 322
time 3.3 s 12.9 s 10.4 s 11.5 s 15.5 s 32.7 s 48.1 s

3.5 times slower at best, and worse with more sweeps: the factors of an indefinite matrix with PMLs are far from diagonally dominant, and the sweeps' non-normal iteration grows before it converges, as the paper warns for such matrices.

What's in this PR

  • docs/methods/fdfd-3d.md, "What didn't pay": both halves, the table, both papers in its front matter; ROADMAP resolves the item as measured and not used.
  • The sweeps stay in the code, off unless a test turns them on (Ilu0::with_sweeps and IterativeSolver3d::with_ilu_sweeps are #[cfg(test)]), so the measurement can be repeated: SWEEP_CASE=guide cargo test --release fdfd::three::tests::ilu_sweeps -- --ignored --nocapture.
  • A test: the swept preconditioner is exactly its own transpose (uᵀ(M⁻¹v) = (M⁻ᵀu)ᵀv to 1e-12 for 1, 2 and 5 sweeps), and n sweeps on a factor of n rows give the exact solve.
  • cargo test --release (360 passed), clippy on both crates.

@tachsin tachsin added this to the 0.4.2 milestone Oct 5, 2026
@tachsin
tachsin merged commit 349be99 into main Oct 5, 2026
10 checks passed
@github-actions github-actions Bot mentioned this pull request Oct 5, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

perf: deterministic parallel ILU(0): factorization and triangular solves by fixed sweeps (A5)

1 participant