Skip to content

docs: photonoxide's direct solvers against PARDISO and MUMPS, and what would close the gap in pure Rust - #155

Merged
tachsin merged 1 commit into
mainfrom
docs/baselines
Oct 5, 2026
Merged

tachsin merged 1 commit into
mainfrom
docs/baselines

Conversation

@tachsin

@tachsin tachsin commented Oct 5, 2026

Copy link
Copy Markdown
Owner

Closes #143. The comparison itself; #153 added the exporter.

docs/baselines.md: photonoxide's direct solvers (faer's supernodal LU, COLAMD in 2D, nested dissection in 3D) against MKL PARDISO and MUMPS. All three ran on the same exported matrices on the same machine (Core Ultra 7 265K, WSL 2), each in a process of its own, on 1 and 20 threads. PARDISO and MUMPS were run as external programs: nothing linked, and their C drivers are kept outside the repository.

What it found

slab-2d strip-24 strip-32 strip-40
photonoxide's factor entries / PARDISO's 3.3 4.5 4.8 4.9
time / PARDISO's, 1 thread 2.1 4.7 5.0 4.6
time / PARDISO's, 20 threads 11 9.6 6.5 4.1
  • One thread: the fill, not the arithmetic.
    • faer's dense LU is within 10% of MKL's zgetrf at large-supernode sizes.
    • But faer's sparse LU pivots by rows, so it works in AᵀA's Cholesky structure (George & Ng 1987).
    • PARDISO and MUMPS work in A + Aᵀ's structure, with pivoting kept inside it.
  • Complex symmetric: all four matrices are diagonally similar to complex symmetric ones (feat(fdfd): QMR for complex symmetric matrices on the curl-curl operator's symmetric similarity, twice as fast #149). As such, PARDISO and MUMPS store half the entries and factor 1.1 to 1.9 times faster.
  • 20 threads: faer's threaded dense LU reaches 13 to 40% of MKL's speed. The penalty falls on the small systems: on the 40³ strip photonoxide scales as well as PARDISO, but on the 2D slab it is slower on 20 threads than on one.
  • Memory follows the fill: 16.8 GB against PARDISO's 3.8 GB on the 40³ strip.

What would close it, in pure Rust

  1. An LU on A + Aᵀ's structure (our nested dissection with one-step separators), with pivoting kept in the supernodes: static pivoting, which the fix: a 3D direct solve reaches round-off on any machine, finishing by QMR on an inaccurate factorization #82 refinement and QMR safety net already covers, or delayed pivots. Expected: 4 to 5 times fewer entries, and about PARDISO's one-thread time.
  2. A complex symmetric (unconjugated) LDLᵀ, Bunch–Kaufman, on B = S A S⁻¹. It halves the storage again. faer has only the Hermitian version, so this belongs upstream in faer.
  3. Parallelism over nested dissection's tree.
  4. faer's threaded dense LU, reported upstream.

Also

  • Linked from the roadmap (the benchmark-harness item is now ticked), the performance plan's "faer against MKL" section, and fdfd-3d's cost section.
  • factor_entries' docs say "about what faer stores": on strip-40 the structure's 1 065 M entries are 17.0 GB, against a 16.8 GB peak.
  • Practical notes in the doc:
    • Debian's MKL 2020 fails PARDISO's analysis on Intel OpenMP with two threads or more; it works with MKL_THREADING_LAYER=GNU.
    • The sequential MUMPS threads only through BLAS.

…he difference comes from, and what would close it in pure Rust (#143)
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

feat(bench): matrices exported for PARDISO and MUMPS as external baselines (needs the owner's licence decision)

1 participant