A cold-path hint without 1.95. An empty #[cold] fn cold_path() {} called at the start of a rare branch. That's how std builds core::hint::cold_path itself (library/core/src/hint.rs): the function is #[cold], with the comment "Even if for some reason the cold_path intrinsic is not visible to codegen, the coldness will ensure that branches this is in are still known to be cold". Its intrinsic's fallback body is #[cold] pub const fn cold_path() {}. |
Inlined everywhere: • src/rng.rs:96: Lemire's rejection branch in below_u64 (:93). • src/math.rs:55: ln's special cases. • src/math.rs:133: exp's |x| ≥ 709.78 cases. Linear algebra and local methods: • src/linalg/cholesky.rs:222: a pivot that isn't positive. • src/algorithm/lbfgsb.rs:790: NaN in projected_gradient. • src/algorithm/lbfgsb/model.rs:137: a skipped correction pair. • src/algorithm/nelder_mead.rs:514: NaN from an overflowing step. Error paths: • src/fitness.rs:110, :150: NaN errors. genoxide already moves cold work into #[cold] functions (cmaes.rs:378, pso.rs:223, nelder_mead.rs:343); this marks a branch cold without moving code. |
Performance, possibly: block layout of hot code. The std docs say the "exact effect to codegen is not guaranteed" and that it "can actually decrease performance", and advise benchmarks. Whether the 1.88 version gives the same machine code as 1.95's intrinsic isn't documented, so it has to be measured. |
Fixed-size pairs without 1.94. In std's source, ArrayWindows::next is let ret = self.v.first_chunk(); if ret.is_some() { self.v = &self.v[1..]; } ret, using first_chunk (1.77). The same loop on 1.88: let mut rest = &x[..]; while let Some(&[a, b]) = rest.first_chunk::<2>() { …; rest = &rest[1..]; }, or as a small iter::from_fn helper. |
Test functions: src/problems/classic.rs:592 (Rosenbrock), :1078 (Dixon-Price), :1127 (Trid). Analytic gradients: src/problems/gradients.rs:50, :168. Multi-objective: src/multi/pareto.rs:355 (crowding distance, with first_chunk::<3>), src/multi/problems/classic.rs:387 (Kursawe). CLI: src/bin/genoxide/builtin.rs:49. Tests: cmaes.rs:1120, de.rs:2768, es.rs:1142, ga.rs:1724. |
Performance, possibly: the items are &[f64; N], as with array_windows, so pair[1] needs no check. The same operations in the same order. Measure: as_chunks was slower in GA breeding (ga.rs:618). |
Not equivalent: destructuring windows(2). let &[a, b] = pair else { unreachable!() } or <&[f64; 2]>::try_from(pair) |
the same sites |
windows yields &[T] of a length known only at run time, so both forms keep a length comparison per item in the source. Whether LLVM removes it can't be known without compiling, so I'm not claiming it does. The first_chunk loop above is the form std itself uses. |
as_chunks_mut (1.88) for packing |
Strided writes with a bounds check per element: • src/linalg/blas.rs:178 (gemm's A panel) • src/linalg/cholesky.rs:176 (the update's panel) • src/linalg/cholesky.rs:155-156 (the packed rows below a diagonal block) Sketch: tile[kk * MR + r % MR] = x → for (slot, &x) in tile.as_chunks_mut::<MR>().0.iter_mut().zip(row) { slot[r % MR] = x } |
Performance, possibly: the index is provably in bounds. Copies only. |
| Reslice or zip index loops |
gemv: src/linalg/blas.rs:61-67, for j in 0..n reading x[j] and four rows; x.len() == n is only a debug_assert. L-BFGS-B: src/algorithm/lbfgsb.rs:424, :572, :587, :788. L-BFGS-B model: src/algorithm/lbfgsb/model.rs:132, :447, :545, :622. Sketch: let x = &x[..n]; before the loop, or zip. The first-order methods (src/algorithm/first_order.rs:644, :662, :703) already use zip. |
Performance, possibly: fewer bounds checks in O(n) loops. Same order of operations. |
f64::midpoint (1.85) |
src/algorithm/de.rs:1128: genoxide's own midpoint, used for DE's bounce-back (:1052, :1054) |
Clarity. Only if std's result equals (a + b) / 2.0 whenever that's finite, and a / 2.0 + b / 2.0 otherwise, bit for bit; to check before switching. |
usize::isqrt (1.84) |
src/multi/spea2.rs:233: ((n as f64).sqrt() as usize) → n.isqrt() |
Clarity and exactness. The same k for any size genoxide allows (≤ 2^24). |
<[T]>::split_off (1.87) |
src/checkpoint.rs:223-230: the take helper is bytes.split_off(..len) |
Clarity: removes a helper. |
<[T]>::split_first_chunk (1.77) |
src/checkpoint.rs:151-152: take 8 bytes, then try_into → rest.split_first_chunk::<8>(). That's 1.93's as_array without the bump. |
Clarity. |
is_multiple_of (1.87) |
Library: src/algorithm/open_es.rs:649, src/gp/boolean.rs:361, src/multi/problems/wfg.rs:723. Python: python/src/problems.rs:417, :427. wfg.rs:233 and :239 already use it. |
Consistency with #315. |
#[expect(...)] (1.81) |
The 30 #[allow(...)]s (git grep -n '#\[allow('). Among them, the linear algebra's dead_code allows: • src/linalg/blas.rs:10, cholesky.rs:10, triangular.rs:8, eigen.rs:13, each commented "L-BFGS-B (batch A2) and the Gaussian processes (batch B) are the first users". • src/algorithm/line_search.rs:146, :344, :353. L-BFGS-B now uses dot, gemv, cholesky and the triangular solves (src/algorithm/lbfgsb/model.rs:9-11), so some may be stale. |
Tooling: an allow that's no longer needed becomes a warning. |
Summary
I looked at what Rust 1.89 to 1.99 would give genoxide (1.99 was released 2026-10-01) and at what the current minimum, 1.88 (since #315), already allows. My recommendation is to stay at 1.88 for now:
core::hint::cold_path(1.95) and<[T]>::array_windows(1.94) can be written today. std's own source implements the first as a#[cold]function and the second withfirst_chunk(1.77).cfg_select!,if letguards,assert_matches!) or excluded by genoxide's rules (unsafe, FMA, non-deterministic float operations).Version numbers are from the official release notes (
RELEASES.mdin rust-lang/rust) and the release posts on blog.rust-lang.org. Claims about how std implements something are from std's source on rust-lang/rust master.Usable now, on 1.88
Performance first. None of these reorders a float sum or fuses operations, so seeded results should stay bit-identical; the bit-identical tests have to confirm it.
#[cold] fn cold_path() {}called at the start of a rare branch. That's how std buildscore::hint::cold_pathitself (library/core/src/hint.rs): the function is#[cold], with the comment "Even if for some reason the cold_path intrinsic is not visible to codegen, the coldness will ensure that branches this is in are still known to be cold". Its intrinsic's fallback body is#[cold] pub const fn cold_path() {}.•
src/rng.rs:96: Lemire's rejection branch inbelow_u64(:93).•
src/math.rs:55:ln's special cases.•
src/math.rs:133:exp's |x| ≥ 709.78 cases.Linear algebra and local methods:
•
src/linalg/cholesky.rs:222: a pivot that isn't positive.•
src/algorithm/lbfgsb.rs:790: NaN inprojected_gradient.•
src/algorithm/lbfgsb/model.rs:137: a skipped correction pair.•
src/algorithm/nelder_mead.rs:514: NaN from an overflowing step.Error paths:
•
src/fitness.rs:110,:150: NaN errors.genoxide already moves cold work into
#[cold]functions (cmaes.rs:378,pso.rs:223,nelder_mead.rs:343); this marks a branch cold without moving code.ArrayWindows::nextislet ret = self.v.first_chunk(); if ret.is_some() { self.v = &self.v[1..]; } ret, usingfirst_chunk(1.77). The same loop on 1.88:let mut rest = &x[..]; while let Some(&[a, b]) = rest.first_chunk::<2>() { …; rest = &rest[1..]; }, or as a smalliter::from_fnhelper.src/problems/classic.rs:592(Rosenbrock),:1078(Dixon-Price),:1127(Trid).Analytic gradients:
src/problems/gradients.rs:50,:168.Multi-objective:
src/multi/pareto.rs:355(crowding distance, withfirst_chunk::<3>),src/multi/problems/classic.rs:387(Kursawe).CLI:
src/bin/genoxide/builtin.rs:49.Tests:
cmaes.rs:1120,de.rs:2768,es.rs:1142,ga.rs:1724.&[f64; N], as witharray_windows, sopair[1]needs no check. The same operations in the same order. Measure:as_chunkswas slower in GA breeding (ga.rs:618).windows(2).let &[a, b] = pair else { unreachable!() }or<&[f64; 2]>::try_from(pair)windowsyields&[T]of a length known only at run time, so both forms keep a length comparison per item in the source. Whether LLVM removes it can't be known without compiling, so I'm not claiming it does. Thefirst_chunkloop above is the form std itself uses.as_chunks_mut(1.88) for packing•
src/linalg/blas.rs:178(gemm's A panel)•
src/linalg/cholesky.rs:176(the update's panel)•
src/linalg/cholesky.rs:155-156(the packed rows below a diagonal block)Sketch:
tile[kk * MR + r % MR] = x→for (slot, &x) in tile.as_chunks_mut::<MR>().0.iter_mut().zip(row) { slot[r % MR] = x }gemv:src/linalg/blas.rs:61-67,for j in 0..nreadingx[j]and four rows;x.len() == nis only adebug_assert.L-BFGS-B:
src/algorithm/lbfgsb.rs:424,:572,:587,:788.L-BFGS-B model:
src/algorithm/lbfgsb/model.rs:132,:447,:545,:622.Sketch:
let x = &x[..n];before the loop, orzip. The first-order methods (src/algorithm/first_order.rs:644,:662,:703) already usezip.f64::midpoint(1.85)src/algorithm/de.rs:1128: genoxide's ownmidpoint, used for DE's bounce-back (:1052,:1054)(a + b) / 2.0whenever that's finite, anda / 2.0 + b / 2.0otherwise, bit for bit; to check before switching.usize::isqrt(1.84)src/multi/spea2.rs:233:((n as f64).sqrt() as usize)→n.isqrt()kfor any size genoxide allows (≤ 2^24).<[T]>::split_off(1.87)src/checkpoint.rs:223-230: thetakehelper isbytes.split_off(..len)<[T]>::split_first_chunk(1.77)src/checkpoint.rs:151-152: take 8 bytes, thentry_into→rest.split_first_chunk::<8>(). That's 1.93'sas_arraywithout the bump.is_multiple_of(1.87)src/algorithm/open_es.rs:649,src/gp/boolean.rs:361,src/multi/problems/wfg.rs:723.Python:
python/src/problems.rs:417,:427.wfg.rs:233and:239already use it.#[expect(...)](1.81)#[allow(...)]s (git grep -n '#\[allow('). Among them, the linear algebra'sdead_codeallows:•
src/linalg/blas.rs:10,cholesky.rs:10,triangular.rs:8,eigen.rs:13, each commented "L-BFGS-B (batch A2) and the Gaussian processes (batch B) are the first users".•
src/algorithm/line_search.rs:146,:344,:353.L-BFGS-B now uses
dot,gemv,choleskyand the triangular solves (src/algorithm/lbfgsb/model.rs:9-11), so some may be stale.allowthat's no longer needed becomes a warning.Already in use, or no candidate:
as_chunksand let chains (1.88),select_unpredictable(1.88),chunks_exact,repeat_n(1.82),div_ceil(1.73),OnceLock(1.70), andblack_boxin the benches.get_disjoint_mut(1.86): thesplit_at_mutcalls are prefix/suffix splits or equal parts, which suit them better (src/linalg/eigen.rs:132,:205,src/linalg/triangular.rs:16,:30,:46,:68,src/algorithm/lbfgsb/model.rs:179-181,:403,:753,src/algorithm/lbfgsb.rs:432).Vec::extract_if(1.87): theretaincalls drop items rather than collecting them.const(1.79),use<..>(1.82).hint::assert_unchecked(1.81) isunsafe.What the minimum version does and doesn't change
rust-versionis the oldest compiler genoxide promises to build with. It gates which language features and std APIs genoxide's own code may use, and nothing else:The compiler improvements come with the toolchain, not with the declared minimum:
lldby default onx86_64-unknown-linux-gnu(1.90).RangeInclusive… optimized better in some circumstances" (1.99).derive(PartialOrd)fast path (1.98).Anyone building genoxide with rustup's stable, including CI and the benchmarks, has all of them today, with
rust-version = "1.88". Raising the minimum only stops people with older compilers from building.The same goes for toolchain behavior changes. Rust uses "the v0 symbol mangling scheme by default" (1.97); the notes say this "may cause some tools (such as debuggers or profilers, especially with old versions) to fail to demangle symbols", which matters for the Callgrind benches. Cargo disables incremental compilation "by default when running in CI" (1.99). And the 1.93 notes say std stopped specializing on
Copy, which "may result in some performance regressions". These reach CI's stable jobs whatever the minimum is; only the MSRV job stays on 1.88.Lints follow the toolchain too. CI's stable clippy with
-D warningsalready runs every new lint. Cargo'sbuild.warnings(1.97) could replace-D warningsin CI, but that depends on CI's toolchain, not onrust-version.What a future bump would add: 1.89 to 1.99
For the record, if the minimum is raised later. The "1.88 equivalent" column is why none of this is needed now.
Would help genoxide's code:
core::hint::cold_path#[cold]function, std's own fallback; see above.<[T]>::array_windowsfirst_chunkloop, std's own implementation; see above.u64::carrying_mulsrc/rng.rs:95,:99(u128::from(a) * u128::from(n), then the low and high halves)u128code. Clarity only.<[T]>::as_arraysrc/checkpoint.rs:152split_first_chunk(1.77).cfg_select!#[cfg(feature = "parallel")]/#[cfg(not(...))]:•
src/algorithm.rs:81/:104•
src/algorithm/de.rs:930/:969•
src/algorithm/ga.rs:652/:688•
src/engine.rs:1100/:1128and:1252/:1278for_each_chunk:src/linalg.rs:54-61#[cfg]pairs work.assert_matches!,debug_assert_matches!assert!(matches!(…))insrc/,tests/andpython/src, e.g.src/algorithm/cmaes.rs:1270,src/algorithm/ga.rs:1461. The 1.96 post: they panic "with aDebugrepresentation of the value otherwise".assert!(matches!(v, …), "{v:?}"). Test messages only. Tests are built by the MSRV job (--all-targets), so this needs the bump.<{integer}>::lowest_one,isolate_lowest_oneand the restsrc/rng.rs:169,:174: the set-bit walk withtrailing_zerosandbits &= bits - 1<{integer}>::format_into,NumBuffersrc/bin/genoxide/process.rs:53:write!(line, "{number}")for each gene of integer, bit and permutation genomes sent to a fitness program. The 1.98 post: it "bypasses much of the dynamic dispatch that you would get with bufferedwrite!formatting, which can be a boon to performance".core::range::RangeInclusive(1.95),Range,RangeFrom,RangeToInclusive(1.96)std::ops::RangeInclusive<f64>, e.g.first_order.rs:644. The 1.96 post explains the new types exist because the old ones implementIteratorand so can't beCopy.std::ops::RangeInclusive, so either way it needs its own design.Checked, and not applicable or excluded:
{fN}::algebraic_add,_sub,_mul,_divand_rem(1.98). The 1.98 post says "These methods are non-deterministic, since the compiler is free to choose different optimizations", which breaks bit-identical results.f64::mul_add(1.94): fused multiply-add; ruled out atsrc/linalg.rs:13-14.avx512fp16intrinsics (1.94): runtime dispatch needsunsafeunder#![forbid(unsafe_code)]; intrinsics are safe only where the caller already has the features (1.86, 1.87). Compile-time-C target-featureis the user's choice (src/linalg.rs:18-21).unsafe:MaybeUninit,into_raw_parts,unchecked_*,as_ref_unchecked,Box::new_zeroed(1.92, 1.93, 1.95),Vec::from_parts(1.99), and C-variadic definitions (1.99).if letguards (1.95): no match arm re-tests a pattern after its guard.Vec::push_mut(1.95): only a test pushes then callslast_mut()(wfg.rs:1371).VecDeque::pop_front_if(1.93) andVecDeque::retain_back(1.99): the deques (cmaes.rs:696-703,local_search.rs:326) pop by length.bool::ok_or(1.98):python/src/problems.rs:697,:714want anOption<()>, not aResult.strip_circumfix,subslice_range,substr_range(1.98): no matching parsing code.String::from_utf8_lossy_owned(1.99): the two lossy conversions (checkpoint.rs:138,:147) are borrowed, on error paths.IntoIterator for Box<[T; N]>(1.99),FusedIterator for StepBy(1.99),Default for RepeatN(1.97),From<T> for LazyLock(1.96), the newLazyLockmethods (1.94; the statics atctp.rs:495,mw.rs:431needSelf),strict_*(1.91),Duration::from_mins(1.91),bool: TryFrom<{integer}>(1.95),Peekable::next_if_map(1.94),BTreeMap::extract_if(1.91),RwLockWriteGuard::downgrade(1.92),fs::set_times(1.99),size_of_val_raw(1.99).f64::consts::GOLDEN_RATIOandEULER_GAMMA(1.94): the line search is Moré-Thuente (src/algorithm/line_search.rs:248), not golden-section.fmt::from_fn(1.93), becausegp::tree::Display(src/gp/tree.rs:215) is a public type.mismatched_lifetime_syntaxes(1.89),dead_code_pub_in_binary(1.97, allow-by-default; it could check the CLI binary), andunconditional_panicon a zero chunk or window size (1.99).include(1.94),build.warningsandresolver.lockfile-path(1.97), and thedebugprofile (1.99).Cargo.toml(1.94) would itself raise the development minimum, according to the notes.std::i32::MAXand the like, fully deprecated in 1.99) and thestd::charconstants (deprecated in 1.97). Neither is used.Costs of a bump, and who is affected
Distribution default compilers are already below 1.88:
Ubuntu's versioned compiler packages are what 1.88 still serves (packages.ubuntu.com):
rustc-1.89andrustc-1.91on 22.04 and 24.04, andrustc-1.91torustc-1.93on 26.04.rustc-1.95torustc-1.97. No Ubuntu package ships 1.99.rustup users:
Python users: the wheels are prebuilt (abi3, CPython ≥ 3.10), so only builds from source need the minimum.
Dependencies (
rust-versionof the locked versions) neither require nor block a bump:So genoxide's own minimum, 1.88, is already above everything it depends on.
Releases: v0.11.0 has shipped, so a bump would land in 0.12 at the earliest. perf!: move to Rust 1.88, with small performance gains across the library #315 treated the last bump as breaking (
perf!). The changelog parsers inrelease-plz.tomldecide where such a PR title appears:perf!:matches^perf(release-plz.toml:54) and goes under Performance, as perf!: move to Rust 1.88, with small performance gains across the library #315 did.chore!:orbuild!:match the skip rule (:57) and would be kept only byprotect_breaking_commits = true(:40).cargo-semver-checks checks the public API; I wouldn't rely on it to notice a
rust-versionchange.1.95 or 1.99, if a bump happens:
assert_matches!(1.96) andformat_into(1.98).Plan
Stay at 1.88. Make the "usable now" changes in small PRs, one topic each:
first_chunkpairs;linalgand L-BFGS-B;Measure each change with the existing benchmarks:
benches/instructions.rs, the gungraun CI gate);benches/hot_paths.rs,benches/linalg.rs,benches/lbfgsb.rs,benches/first_order.rs);benchmarks/run.py versions.Each change must also pass the bit-identical tests, including the linear algebra's bit hashes (
src/linalg/tests.rs). A change that doesn't help is dropped, asas_chunkswas in GA breeding (ga.rs:618).If a later feature makes a bump worth it, it's one PR:
Cargo.toml:18,python/Cargo.toml:6;.github/workflows/ci.yml:227,:231;ROADMAP.md:60, which cites issue Move to Rust 1.88, and look for small performance gains everywhere #314; the bump itself was PR perf!: move to Rust 1.88, with small performance gains across the library #315, andROADMAP.md:207cites both;ROADMAP.md:206-207.docs/optimization-plan.md:971records the 0.10 decision and can stay. README, CONTRIBUTING, AGENTS.md,examples/gpu/Cargo.tomlandbenchmarks/set no version.Not tested yet
This is research only: I haven't built anything, run any benchmark or changed any code. Every gain above is a candidate until it's measured. In particular, I don't know yet whether:
#[cold]helper on 1.88 produces the same code as 1.95's intrinsic;first_chunkloop compiles to whatarray_windowsdoes on a given toolchain;windows(2)loses its length checks.