Skip to content

feat(bench): the bandwidth each iterative solve reaches, its bytes counted by the kernels, against the triad - #152

Merged
tachsin merged 1 commit into
mainfrom
feat/bench-bandwidth
Oct 5, 2026
Merged

tachsin merged 1 commit into
mainfrom
feat/bench-bandwidth

Conversation

@tachsin

@tachsin tachsin commented Oct 5, 2026

Copy link
Copy Markdown
Owner

Closes #142.

What

  • src/traffic.rs (crate-private): one counter of the bytes the iterative kernels move, by a stated model: each stored nonzero's value (16 B) and column index (8 B), each row pointer (8 B), each vector read or written (16 B a value), a gathered x counted once. Each kernel adds its bytes once a call, from all threads; the formulas are small functions, tested by hand.
  • Counted where the kernels are, so the count stays true when they change: the sparse product, QMR's fused passes (symmetric or not), GMRES's sums and combinations, the triangular solves (exact and swept), the multigrid's transfers and corrections.
  • bench::Phase::bytes (serde default, so older JSON still reads) for each iterating phase, Measurement::gigabytes_per_second, bench::bytes_moved.
  • photonoxide bench's report: a "Bandwidth (of the triad)" column, the GB/s reached and its share of the triad measured on the same threads (the roof, Williams et al. 2009, doi:10.1145/1498765.1498785), and the model in the report's header.

What it shows (docs/benchmarks.md, Core Ultra 7 265K)

problem 1 thread 20 threads
guide-qmr (symmetric QMR, 40³) 19.4 GB/s (59%) 84.4 GB/s (150%)
guide-ilu 15.3 GB/s (46%) 21.6 GB/s (38%)
guide-multigrid 12.7 GB/s (39%) 21.6 GB/s (38%)
diel-multigrid (864 k) 13.9 GB/s (42%) 37.1 GB/s (66%)
strip-ports-multigrid (1.8 M) 13.3 GB/s (40%) 34.3 GB/s (61%)
  • Above 100% on the 40³ guide: the count is the kernels' traffic from memory or cache, and the guide's vectors (3 MB each) fit in the 30 MB L3, where only its matrix (about 60 MB) streams from memory; the triad also counts 24 B an element where write-allocate moves 32. The header says so. The larger problems settle at 61–66%.
  • QMR + ILU(0) at 38% on 20 threads: its sequential triangular solves (docs(fdfd): parallel ILU(0) measured on our matrices, both halves, and not used #150 measured that sweeps don't help).
  • The plan's estimate of "about a fifth of the bandwidth" for the old QMR is replaced by measurement.

Checked

@tachsin tachsin added this to the 0.4.2 milestone Oct 5, 2026
@tachsin
tachsin merged commit c7c7637 into main Oct 5, 2026
10 checks passed
@github-actions github-actions Bot mentioned this pull request Oct 5, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

feat(bench): the bandwidth each iterative solve reaches, against the triad

1 participant