Skip to content

perf: what raising the minimum Rust version would gain, and what 1.88 can already do #402

Description

@tachsin

Summary

I looked at what Rust 1.89 to 1.99 would give genoxide (1.99 was released 2026-10-01) and at what the current minimum, 1.88 (since #315), already allows. My recommendation is to stay at 1.88 for now:

  • The two most useful newer APIs have 1.88 equivalents. core::hint::cold_path (1.95) and <[T]>::array_windows (1.94) can be written today. std's own source implements the first as a #[cold] function and the second with first_chunk (1.77).
  • I found nothing in 1.89 to 1.99 that brings a gain without a 1.88 equivalent. What can't be emulated is either clarity only (cfg_select!, if let guards, assert_matches!) or excluded by genoxide's rules (unsafe, FMA, non-deterministic float operations).
  • Raising the minimum doesn't speed anything up. Compiler improvements, LLVM 23 included, come with the toolchain users build with, whatever the declared minimum.
  • Several improvements need no bump. They're listed first below, mostly in the new linear algebra, L-BFGS-B and the test functions.

Version numbers are from the official release notes (RELEASES.md in rust-lang/rust) and the release posts on blog.rust-lang.org. Claims about how std implements something are from std's source on rust-lang/rust master.

Usable now, on 1.88

Performance first. None of these reorders a float sum or fuses operations, so seeded results should stay bit-identical; the bit-identical tests have to confirm it.

Change Where Gain
A cold-path hint without 1.95. An empty #[cold] fn cold_path() {} called at the start of a rare branch. That's how std builds core::hint::cold_path itself (library/core/src/hint.rs): the function is #[cold], with the comment "Even if for some reason the cold_path intrinsic is not visible to codegen, the coldness will ensure that branches this is in are still known to be cold". Its intrinsic's fallback body is #[cold] pub const fn cold_path() {}. Inlined everywhere:
• src/rng.rs:96: Lemire's rejection branch in below_u64 (:93).
• src/math.rs:55: ln's special cases.
• src/math.rs:133: exp's |x| ≥ 709.78 cases.
Linear algebra and local methods:
• src/linalg/cholesky.rs:222: a pivot that isn't positive.
• src/algorithm/lbfgsb.rs:790: NaN in projected_gradient.
• src/algorithm/lbfgsb/model.rs:137: a skipped correction pair.
• src/algorithm/nelder_mead.rs:514: NaN from an overflowing step.
Error paths:
• src/fitness.rs:110, :150: NaN errors.
genoxide already moves cold work into #[cold] functions (cmaes.rs:378, pso.rs:223, nelder_mead.rs:343); this marks a branch cold without moving code.
Performance, possibly: block layout of hot code. The std docs say the "exact effect to codegen is not guaranteed" and that it "can actually decrease performance", and advise benchmarks. Whether the 1.88 version gives the same machine code as 1.95's intrinsic isn't documented, so it has to be measured.
Fixed-size pairs without 1.94. In std's source, ArrayWindows::next is let ret = self.v.first_chunk(); if ret.is_some() { self.v = &self.v[1..]; } ret, using first_chunk (1.77). The same loop on 1.88: let mut rest = &x[..]; while let Some(&[a, b]) = rest.first_chunk::<2>() { …; rest = &rest[1..]; }, or as a small iter::from_fn helper. Test functions: src/problems/classic.rs:592 (Rosenbrock), :1078 (Dixon-Price), :1127 (Trid).
Analytic gradients: src/problems/gradients.rs:50, :168.
Multi-objective: src/multi/pareto.rs:355 (crowding distance, with first_chunk::<3>), src/multi/problems/classic.rs:387 (Kursawe).
CLI: src/bin/genoxide/builtin.rs:49.
Tests: cmaes.rs:1120, de.rs:2768, es.rs:1142, ga.rs:1724.
Performance, possibly: the items are &[f64; N], as with array_windows, so pair[1] needs no check. The same operations in the same order. Measure: as_chunks was slower in GA breeding (ga.rs:618).
Not equivalent: destructuring windows(2). let &[a, b] = pair else { unreachable!() } or <&[f64; 2]>::try_from(pair) the same sites windows yields &[T] of a length known only at run time, so both forms keep a length comparison per item in the source. Whether LLVM removes it can't be known without compiling, so I'm not claiming it does. The first_chunk loop above is the form std itself uses.
as_chunks_mut (1.88) for packing Strided writes with a bounds check per element:
• src/linalg/blas.rs:178 (gemm's A panel)
• src/linalg/cholesky.rs:176 (the update's panel)
• src/linalg/cholesky.rs:155-156 (the packed rows below a diagonal block)
Sketch: tile[kk * MR + r % MR] = x → for (slot, &x) in tile.as_chunks_mut::<MR>().0.iter_mut().zip(row) { slot[r % MR] = x }
Performance, possibly: the index is provably in bounds. Copies only.
Reslice or zip index loops gemv: src/linalg/blas.rs:61-67, for j in 0..n reading x[j] and four rows; x.len() == n is only a debug_assert.
L-BFGS-B: src/algorithm/lbfgsb.rs:424, :572, :587, :788.
L-BFGS-B model: src/algorithm/lbfgsb/model.rs:132, :447, :545, :622.
Sketch: let x = &x[..n]; before the loop, or zip. The first-order methods (src/algorithm/first_order.rs:644, :662, :703) already use zip.
Performance, possibly: fewer bounds checks in O(n) loops. Same order of operations.
f64::midpoint (1.85) src/algorithm/de.rs:1128: genoxide's own midpoint, used for DE's bounce-back (:1052, :1054) Clarity. Only if std's result equals (a + b) / 2.0 whenever that's finite, and a / 2.0 + b / 2.0 otherwise, bit for bit; to check before switching.
usize::isqrt (1.84) src/multi/spea2.rs:233: ((n as f64).sqrt() as usize) → n.isqrt() Clarity and exactness. The same k for any size genoxide allows (≤ 2^24).
<[T]>::split_off (1.87) src/checkpoint.rs:223-230: the take helper is bytes.split_off(..len) Clarity: removes a helper.
<[T]>::split_first_chunk (1.77) src/checkpoint.rs:151-152: take 8 bytes, then try_into → rest.split_first_chunk::<8>(). That's 1.93's as_array without the bump. Clarity.
is_multiple_of (1.87) Library: src/algorithm/open_es.rs:649, src/gp/boolean.rs:361, src/multi/problems/wfg.rs:723.
Python: python/src/problems.rs:417, :427. wfg.rs:233 and :239 already use it.
Consistency with #315.
#[expect(...)] (1.81) The 30 #[allow(...)]s (git grep -n '#\[allow('). Among them, the linear algebra's dead_code allows:
• src/linalg/blas.rs:10, cholesky.rs:10, triangular.rs:8, eigen.rs:13, each commented "L-BFGS-B (batch A2) and the Gaussian processes (batch B) are the first users".
• src/algorithm/line_search.rs:146, :344, :353.
L-BFGS-B now uses dot, gemv, cholesky and the triangular solves (src/algorithm/lbfgsb/model.rs:9-11), so some may be stale.
Tooling: an allow that's no longer needed becomes a warning.

Already in use, or no candidate:

  • In use: as_chunks and let chains (1.88), select_unpredictable (1.88), chunks_exact, repeat_n (1.82), div_ceil (1.73), OnceLock (1.70), and black_box in the benches.
  • get_disjoint_mut (1.86): the split_at_mut calls are prefix/suffix splits or equal parts, which suit them better (src/linalg/eigen.rs:132, :205, src/linalg/triangular.rs:16, :30, :46, :68, src/algorithm/lbfgsb/model.rs:179-181, :403, :753, src/algorithm/lbfgsb.rs:432).
  • Vec::extract_if (1.87): the retain calls drop items rather than collecting them.
  • No case found: inline const (1.79), use<..> (1.82). hint::assert_unchecked (1.81) is unsafe.

What the minimum version does and doesn't change

rust-version is the oldest compiler genoxide promises to build with. It gates which language features and std APIs genoxide's own code may use, and nothing else:

  • The compiler improvements come with the toolchain, not with the declared minimum:

    • LLVM 21 (1.91), LLVM 22 (1.95) and LLVM 23 (1.99). The notes for 1.96, 1.97 and 1.98 list no LLVM update; 1.97.1 "backports an LLVM submodule bump" to fix a miscompilation.
    • lld by default on x86_64-unknown-linux-gnu (1.90).
    • "Iteration on RangeInclusive … optimized better in some circumstances" (1.99).
    • The derive(PartialOrd) fast path (1.98).

    Anyone building genoxide with rustup's stable, including CI and the benchmarks, has all of them today, with rust-version = "1.88". Raising the minimum only stops people with older compilers from building.

  • The same goes for toolchain behavior changes. Rust uses "the v0 symbol mangling scheme by default" (1.97); the notes say this "may cause some tools (such as debuggers or profilers, especially with old versions) to fail to demangle symbols", which matters for the Callgrind benches. Cargo disables incremental compilation "by default when running in CI" (1.99). And the 1.93 notes say std stopped specializing on Copy, which "may result in some performance regressions". These reach CI's stable jobs whatever the minimum is; only the MSRV job stays on 1.88.

  • Lints follow the toolchain too. CI's stable clippy with -D warnings already runs every new lint. Cargo's build.warnings (1.97) could replace -D warnings in CI, but that depends on CI's toolchain, not on rust-version.

What a future bump would add: 1.89 to 1.99

For the record, if the minimum is raised later. The "1.88 equivalent" column is why none of this is needed now.

Would help genoxide's code:

Version Feature Where 1.88 equivalent
1.95 core::hint::cold_path the sites above An empty #[cold] function, std's own fallback; see above.
1.94 <[T]>::array_windows the sites above A first_chunk loop, std's own implementation; see above.
1.91 u64::carrying_mul src/rng.rs:95, :99 (u128::from(a) * u128::from(n), then the low and high halves) The current u128 code. Clarity only.
1.93 <[T]>::as_array src/checkpoint.rs:152 split_first_chunk (1.77).
1.95 cfg_select! Function pairs under #[cfg(feature = "parallel")] / #[cfg(not(...))]:
• src/algorithm.rs:81/:104
• src/algorithm/de.rs:930/:969
• src/algorithm/ga.rs:652/:688
• src/engine.rs:1100/:1128 and :1252/:1278
for_each_chunk: src/linalg.rs:54-61
None, but it's cosmetic: the #[cfg] pairs work.
1.96 assert_matches!, debug_assert_matches! 69 assert!(matches!(…)) in src/, tests/ and python/src, e.g. src/algorithm/cmaes.rs:1270, src/algorithm/ga.rs:1461. The 1.96 post: they panic "with a Debug representation of the value otherwise". assert!(matches!(v, …), "{v:?}"). Test messages only. Tests are built by the MSRV job (--all-targets), so this needs the bump.
1.97 <{integer}>::lowest_one, isolate_lowest_one and the rest src/rng.rs:169, :174: the set-bit walk with trailing_zeros and bits &= bits - 1 The current code. Clarity only; I haven't checked the new methods' exact signatures.
1.98 <{integer}>::format_into, NumBuffer src/bin/genoxide/process.rs:53: write!(line, "{number}") for each gene of integer, bit and permutation genomes sent to a fitness program. The 1.98 post: it "bypasses much of the dynamic dispatch that you would get with buffered write! formatting, which can be a boon to performance". A hand-written digit loop. CLI only, where each line goes to another process anyway.
1.95, 1.96 core::range::RangeInclusive (1.95), Range, RangeFrom, RangeToInclusive (1.96) Gene bounds are stored and passed as std::ops::RangeInclusive<f64>, e.g. first_order.rs:644. The 1.96 post explains the new types exist because the old ones implement Iterator and so can't be Copy. A two-field bounds type of genoxide's own. The public API takes std::ops::RangeInclusive, so either way it needs its own design.

Checked, and not applicable or excluded:

  • Excluded by genoxide's rules:
    • {fN}::algebraic_add, _sub, _mul, _div and _rem (1.98). The 1.98 post says "These methods are non-deterministic, since the compiler is free to choose different optimizations", which breaks bit-identical results.
    • Const-stable f64::mul_add (1.94): fused multiply-add; ruled out at src/linalg.rs:13-14.
    • AVX-512 target features (1.89) and avx512fp16 intrinsics (1.94): runtime dispatch needs unsafe under #![forbid(unsafe_code)]; intrinsics are safe only where the caller already has the features (1.86, 1.87). Compile-time -C target-feature is the user's choice (src/linalg.rs:18-21).
    • APIs that need unsafe: MaybeUninit, into_raw_parts, unchecked_*, as_ref_unchecked, Box::new_zeroed (1.92, 1.93, 1.95), Vec::from_parts (1.99), and C-variadic definitions (1.99).
  • No site in genoxide:
    • if let guards (1.95): no match arm re-tests a pattern after its guard.
    • Vec::push_mut (1.95): only a test pushes then calls last_mut() (wfg.rs:1371).
    • VecDeque::pop_front_if (1.93) and VecDeque::retain_back (1.99): the deques (cmaes.rs:696-703, local_search.rs:326) pop by length.
    • bool::ok_or (1.98): python/src/problems.rs:697, :714 want an Option<()>, not a Result.
    • strip_circumfix, subslice_range, substr_range (1.98): no matching parsing code.
    • String::from_utf8_lossy_owned (1.99): the two lossy conversions (checkpoint.rs:138, :147) are borrowed, on error paths.
    • IntoIterator for Box<[T; N]> (1.99), FusedIterator for StepBy (1.99), Default for RepeatN (1.97), From<T> for LazyLock (1.96), the new LazyLock methods (1.94; the statics at ctp.rs:495, mw.rs:431 need Self), strict_* (1.91), Duration::from_mins (1.91), bool: TryFrom<{integer}> (1.95), Peekable::next_if_map (1.94), BTreeMap::extract_if (1.91), RwLockWriteGuard::downgrade (1.92), fs::set_times (1.99), size_of_val_raw (1.99).
    • f64::consts::GOLDEN_RATIO and EULER_GAMMA (1.94): the line search is Moré-Thuente (src/algorithm/line_search.rs:248), not golden-section.
    • Const-stable float rounding (1.90): no const fn needs it.
  • Would break the public API: fmt::from_fn (1.93), because gp::tree::Display (src/gp/tree.rs:215) is a public type.
  • Toolchain, not minimum:
    • The lints added in 1.89 to 1.99, for example mismatched_lifetime_syntaxes (1.89), dead_code_pub_in_binary (1.97, allow-by-default; it could check the CLI binary), and unconditional_panic on a zero chunk or window size (1.99).
    • Cargo's include (1.94), build.warnings and resolver.lockfile-path (1.97), and the debug profile (1.99).
    • TOML 1.1 in Cargo.toml (1.94) would itself raise the development minimum, according to the notes.
  • Deprecations that don't affect genoxide: the legacy integral modules (std::i32::MAX and the like, fully deprecated in 1.99) and the std::char constants (deprecated in 1.97). Neither is used.

Costs of a bump, and who is affected

  • Distribution default compilers are already below 1.88:

    • Debian 13 "trixie", the current stable, ships rustc 1.85.1.
    • Ubuntu 24.04 and 22.04 LTS ship rustc 1.75.
    • Ubuntu 26.04 LTS ships rustc 1.93.1.
  • Ubuntu's versioned compiler packages are what 1.88 still serves (packages.ubuntu.com): rustc-1.89 and rustc-1.91 on 22.04 and 24.04, and rustc-1.91 to rustc-1.93 on 26.04.

    • A minimum of 1.95 would leave all of them out.
    • A minimum of 1.99 would also leave out the development release's rustc-1.95 to rustc-1.97. No Ubuntu package ships 1.99.
  • rustup users:

    • 1.95 (2026-04-16) is four releases behind stable.
    • 1.99 is the current stable: anyone who hasn't updated since 2026-10-01 couldn't build the next release.
  • Python users: the wheels are prebuilt (abi3, CPython ≥ 3.10), so only builds from source need the minimum.

  • Dependencies (rust-version of the locked versions) neither require nor block a bump:

    • rand 0.10.3, rand_chacha 0.10.0 and rand_core 0.10.1: 1.85.
    • rayon 1.12.0: 1.80.
    • serde 1.0.229: 1.56; serde_json 1.0.151: 1.71.
    • toml 0.9.12: 1.76; tracing 0.1.44: 1.65; libm 0.2.16: 1.63.
    • postcard 1.1.3: none set.
    • In python/: pyo3 0.29.2 and numpy 0.29.0: 1.83.
    • Development: criterion 0.8.2: 1.86; proptest 1.11.0: 1.85; gungraun 0.19.4: 1.85.1.

    So genoxide's own minimum, 1.88, is already above everything it depends on.

  • Releases: v0.11.0 has shipped, so a bump would land in 0.12 at the earliest. perf!: move to Rust 1.88, with small performance gains across the library #315 treated the last bump as breaking (perf!). The changelog parsers in release-plz.toml decide where such a PR title appears:

    cargo-semver-checks checks the public API; I wouldn't rely on it to notice a rust-version change.

1.95 or 1.99, if a bump happens:

  • 1.95 keeps four releases of slack. It gets every item in the table except assert_matches! (1.96) and format_into (1.98).
  • 1.99 adds those two, both test- or CLI-only, at the cost of requiring the newest stable.

Plan

  1. Stay at 1.88. Make the "usable now" changes in small PRs, one topic each:

    • the cold-path helper;
    • the first_chunk pairs;
    • the packing and index loops in linalg and L-BFGS-B;
    • the clarity items.
  2. Measure each change with the existing benchmarks:

    • Callgrind instruction counts (benches/instructions.rs, the gungraun CI gate);
    • criterion (benches/hot_paths.rs, benches/linalg.rs, benches/lbfgsb.rs, benches/first_order.rs);
    • the matched suite, via benchmarks/run.py versions.

    Each change must also pass the bit-identical tests, including the linear algebra's bit hashes (src/linalg/tests.rs). A change that doesn't help is dropped, as as_chunks was in GA breeding (ga.rs:618).

  3. If a later feature makes a bump worth it, it's one PR:

    docs/optimization-plan.md:971 records the 0.10 decision and can stay. README, CONTRIBUTING, AGENTS.md, examples/gpu/Cargo.toml and benchmarks/ set no version.

Not tested yet

This is research only: I haven't built anything, run any benchmark or changed any code. Every gain above is a candidate until it's measured. In particular, I don't know yet whether:

  • the #[cold] helper on 1.88 produces the same code as 1.95's intrinsic;
  • the first_chunk loop compiles to what array_windows does on a given toolchain;
  • destructuring windows(2) loses its length checks.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions