Skip to content

A seeded solve gives the same bits on any number of threads or ranks; seed 0 is a seed; BERTINI_NUM_THREADS - #491

Merged
ofloveandhate merged 7 commits into
developfrom
fix/threaded-solves-are-reproducible
Oct 6, 2026
Merged

ofloveandhate merged 7 commits into
developfrom
fix/threaded-solves-are-reproducible

Conversation

@ofloveandhate

@ofloveandhate ofloveandhate commented Oct 6, 2026 •

Copy link
Copy Markdown
Contributor

A seeded solve now gives the same bits whatever runs it: one thread, many threads, or MPI ranks with any number of threads each. It also renames the thread-count variable to BERTINI_NUM_THREADS.

Fixes #378.
Fixes #475.

A path does not depend on the paths before it (#378, ADR-0071)

Each path already drew its random numbers from its own deterministic stream. But a tracker and an endgame are reused path after path, and five pieces of state flowed from one path into the next. With a thread pool, the scheduler decides which paths a thread runs before a given path, so a seeded threaded solve differed from run to run in step counts, precision and the last digits of its endpoints. A serial solve depended on path order: running cyclic-5's 120 paths in reverse changed 45 of them.

The five sources:

  1. An adaptive track built its first step at the precision the previous track ended in, and kept that precision.
  2. A drop from multiple precision to double re-precisioned only the end time. The step size, the time and its increment stayed at the old precision.
  3. The condition-number probe was rounded in place at every precision change, so a drop destroyed digits a later rise could not restore.
  4. The endgame's c/k probe was drawn the first time an endgame object needed it and kept for the object's life, or redrawn mid-path after an escalation.
  5. An endgame run in double left an earlier run's multiprecision times in place, so final_time_used reported another path's last time.

Each is fixed where it lives, not by giving every path a fresh tracker, which would leave the trap for anyone reusing a tracker or an endgame by hand:

  • A track builds its first step from its own start precision.
  • Every conversion re-precisions all internals.
  • Random directions are stored as drawn, and every working copy is derived from that copy.
  • The c/k probe is drawn at the start of every run, after the per-path reseed.
  • Both lanes of an adaptive endgame are cleared when a run begins.

Precision observers and the per-path precision record (ADR-0071)

The same investigation turned up a second kind of stale state: a record kept by an observer outliving the span it describes. The solver fills three fields of each path's metadata from precision observers attached to the tracker: precision_changed, time_of_first_prec_increase and max_precision_used. Those observers had three defects.

  1. A per-track record read as a per-path one. The solver read MinMaxPrecisionRecorder, which started afresh at every TrackingStarted. The endgame tracks a path in many calls to TrackPath, so a precision raised in an earlier endgame sub-track and lowered before the last went unreported. On seeded cyclic-5, power-series path 59 reported 30 digits after tracking at 40, and six Cauchy paths were under-reported.
  2. Recording from the wrong starting point. Both recorders started afresh at TrackingStarted, which a track sends only after its start point has been refined. So they missed any precision increase the refinement needed. A track whose refinement failed, such as one from a singular start point, never sends TrackingStarted at all, and was reported with the previous track's values.
  3. Unsubscribing. FirstPrecisionRecorder unsubscribed at its first increase. Attached once to a tracker, it reported that track forever after. Its members were also left uninitialized until its first event.

The fixes, each where the record lives:

  • FirstPrecisionRecorder and MinMaxPrecisionRecorder record one track each. They start afresh at the track's Initializing event, which comes before the start-point refinement, and they never unsubscribe. Attached once, they report every track, including one that fails to initialize. MinMaxPrecisionRecorder now counts both ends of every precision change, so a track's starting precision is included.
  • A new PathPrecisionRecorder records a whole path across every call to TrackPath. It starts afresh only when its owner calls Reset. The solver resets and attaches it when a path begins, reads it after each stage, and detaches it when the path ends. All three metadata fields now come from it.

The other recorded field this PR corrects is final_time_used (item 5 above). It came from the endgame's stale times, not from an observer.

Seed 0 is a seed (#475, ADR-0072)

set_random_seed(0) meant "draw a seed from entropy", so a script that set seed 0 to make a run reproducible got a different run each time. 0 is now an ordinary seed, as in numpy and Python's random. set_random_seed() or set_random_seed(None) draws a seed from entropy and returns it. A classic input file keeps Bertini 1's meaning: randomseed: 0, its default, still draws from entropy in the command-line program. Seed 0 is therefore the one value that does not carry between a Python script and a classic input file.

BERTINI_NUM_THREADS (ADR-0070)

b2 took its thread count from OMP_NUM_THREADS, although it uses no OpenMP. numpy's OpenBLAS reads that variable too, so pinning b2 to one thread pinned numpy, and tuning numpy changed b2. b2 now reads BERTINI_NUM_THREADS and ignores OMP_NUM_THREADS. The order is the variable if set, then num_threads, then every CPU the process may run on, which is OpenMP's default. MPI ranks follow the same rule and no longer default to one thread. Several ranks on one machine should set the variable, and the CLI help and the scaling tutorial say so.

Tests

Each regression test was run against the code without its fix and fails there.

  • Bit identity. A seeded cyclic-5 solve matches serial path by path at 1, 2, 3 and 8 threads, for both endgames: every endpoint, and every metadata field but the wall-clock time.
  • Hand-driven reuse. A reused tracker matches a fresh one bit for bit, for every pairing of previous and next start precision, ending in double and in multiple precision. A reused endgame matches a fresh one, for every flavor and tracker type.
  • Precision observers. A recorder attached once reports a failed singular-start track and then a regular one, each as itself. A reused tracker's recorders report a failed track exactly as a fresh tracker's do. Each path's max_precision_used equals the highest precision its tracker reported in any event, for both endgames. A Python test checks max_precision_used against the precision at every step, as recorded by SolutionPathCollector.
  • Seeds. Seed 0 reproduces a run, entropy seeds are reported and reproduce it, and classic randomseed: 0 still draws from entropy.
  • Python. Counterparts of the above, plus tests that BERTINI_NUM_THREADS decides the thread count and that OMP_NUM_THREADS does not reach b2. Under mpirun at 3 and 5 ranks, a distributed solve matches a serial one path by path.
  • CI. The MPI smoke test used to compare solution counts. With a fixed seed, it now requires main_data to be byte-identical across a serial run, a threaded run, and mpirun -n 2 with one and two threads per worker.

For users

  • Small numeric changes. Results differ from 4.0.0.dev1 in the last bits and, occasionally, in step counts. The old results depended on path order, so none of them is a reference.
  • The thread variable. Scripts and job files that set OMP_NUM_THREADS for b2 should set BERTINI_NUM_THREADS.
  • Precision metadata. max_precision_used, precision_changed and time_of_first_prec_increase can differ from 4.0.0.dev1, where they were wrong. Records written by earlier builds keep the values those builds computed.
  • Hand-attached observers. C++ code that attaches FirstPrecisionRecorder or MinMaxPrecisionRecorder to a tracker now gets a record of the most recent track, every time. To record a whole path, use PathPrecisionRecorder.

🤖 Generated with Claude Code

ofloveandhate and others added 7 commits October 5, 2026 22:20
b2 took its thread count from OMP_NUM_THREADS, although it uses no OpenMP. numpy's OpenBLAS reads that variable too, so pinning b2 to one thread pinned numpy, and tuning numpy changed how many paths b2 tracked at once. b2 now reads its own variable and ignores OMP_NUM_THREADS.

One rule serves a standalone solve and each MPI rank: BERTINI_NUM_THREADS if set, then num_threads, then the CPUs the process may run on (the affinity mask on Linux), which is OpenMP's default. MPI ranks no longer default to one thread; the CLI help, the scaling tutorial and the CHANGELOG tell users to set the variable when several ranks share a machine. Examples, tutorials, benchmark scripts, tools and the CI legs use the new name. ADR-0070 records the decision.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…solve is bit-identical to serial (#378)

Trackers and endgames are reused path after path, and five pieces of state flowed from one path into the next. The first step of an adaptive track was built at the precision the previous track ended in and kept that precision. A drop from multiple precision to double left the step size, time and increment at the old precision. The condition-number probe was rounded in place at every precision change. The endgame c/k probe was drawn the first time an endgame needed it and kept for its life, or redrawn mid-path after an escalation emptied the double lane. And an endgame run in double left an earlier run's multiprecision times in place, so final_time_used reported another path's time. A thread pool hands each thread a scheduler-chosen sequence of predecessors, so threaded solves were irreproducible, and a serial solve depended on path order.

Each piece now starts from the path's own inputs. The first step is built at the track's start precision; the downward conversion re-precisions every internal; both probes are stored as drawn and every working copy is derived from that; the c/k probe is drawn at the start of every run, after the per-path reseed; both lanes of an adaptive endgame are cleared when a run begins.

Tests: a used tracker matches a fresh one bit for bit for every pairing of previous and next start precision, ending in double and in multiple precision; the step size sits at the working precision after a track; the probe survives a drop in precision; a used endgame matches a fresh one, for every flavor and tracker type, including the c/k estimate; and a seeded cyclic-5 solve is bit-identical path by path, endpoint and every metadata field, at 1, 2, 3 and 8 threads, for both endgames. Each test was run against the code without its fix and fails there.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…depend on the paths before it (#378)

The ADR records the rule and its guardrails: a track or run sets what it uses from its own inputs, random directions are stored as drawn and every working copy derived, per-path directions are drawn at run start after the reseed, per-lane containers are cleared in every lane, and the fix is state hygiene rather than a fresh object per path, which would leave the trap for anyone reusing a tracker or endgame by hand.

Python tests: a threaded cyclic-5 solve matches serial bit for bit, endpoint and metadata, for both endgames; a reused tracker matches a fresh one; and under mpirun a distributed solve matches a serial one path by path (passes at 3 and 5 ranks).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…c input format keeps Bertini 1's meaning (#475)

set_random_seed(0) and SetGlobalSeed(0) meant "draw a seed from entropy", so a script that set seed 0 to make a run reproducible got a different run every time -- the reproducer in #475. 0 is now an ordinary seed, as in numpy and Python's random. set_random_seed() or set_random_seed(None) draws one from entropy and returns it; set_random_seed returns the seed in effect either way; C++ gains SetGlobalSeedFromEntropy. "Not set yet" is its own flag in random.cpp, no longer the value 0.

A classic input file keeps Bertini 1's meaning: randomseed: 0, its default, draws from entropy. The command-line program applies the classic seed through ApplyClassicRandomSeed. Seed 0 is therefore the one value that does not carry between a Python script and a classic input file; the docstrings, the changelog and ADR-0072 say so. Derived seeds still avoid 0, so a recorded seed replays through a classic input too.

Tests: seed 0 reproduces draws and the homotopy digest; an entropy seed is reported and reproduces the run; classic randomseed 0 draws from entropy and any other value is the seed; in Python, set_random_seed(0) is reproducible, set_random_seed() and None draw and return a seed, and the #475 reproducer is identical across trials at 1 and 4 threads.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… threads and ranks change no bit of the output

The MPI smoke test compared only solution counts, serial against mpirun -n 2. With a fixed randomseed it now demands byte-identical main_data across a serial run, a threaded run, and mpirun -n 2 with one and with two threads per worker, all set through BERTINI_NUM_THREADS, for both start-system families.

Python tests, which run in every wheel job on all three platforms, check that the variable decides the thread count: =1 makes a solve configured for 4 threads run every path on the calling thread, =3 makes one configured for 1 thread run on at most three pool threads, and OMP_NUM_THREADS=1 does not reach b2. They count the operating-system threads that deliver PathStarted, because comparing tracker wrappers with == does not compare the underlying trackers.

CLAUDE.md states the thread variable, the seed-0 convention, the CI legs, and the rule that a path does not depend on the paths before it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…erences and history move to comments

EffectiveThreadCount, RefreshCOverKProbe, SetGlobalSeed, ApplyClassicRandomSeed and RandomConfig carried issue numbers, ADR references and accounts of earlier behavior in their Doxygen blocks, which render as user documentation. The blocks now state the contract; the reasons sit in comments beside the code.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…rder reports every track it is attached for

The endgame tracks a path in many calls to TrackPath. The solver read MinMaxPrecisionRecorder, which started afresh at every TrackingStarted, so a precision raised in an earlier endgame sub-track and lowered before the last went unreported: on seeded cyclic-5, power-series path 59 reported 30 digits after tracking at 40, and six Cauchy paths were under-reported. The solver now keeps a PathPrecisionRecorder, reset only when a path begins and spanning all of its tracks, for precision_changed, time_of_first_prec_increase and max_precision_used.

The per-track recorders start afresh at the track's Initializing event, before the start point is refined, and never unsubscribe. FirstPrecisionRecorder unsubscribed at its first increase, so attached once it reported that track ever after; and starting afresh at TrackingStarted missed the refinement's increases and left a track that failed to initialize, which never sends TrackingStarted, reported with the previous track's values. Their members are now initialized.

Tests: a recorder attached once reports a failed singular-start track and then a regular one, each as itself; a used tracker's recorders report a failed track exactly as fresh ones do; and each path's max_precision_used equals the highest precision its tracker reported in any event, for both endgames. Each fails against the old behavior. A Python test checks max_precision_used against the per-path step collector. ADR-0071 gains the rule, and the changelog the fix.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@ofloveandhate

Copy link
Copy Markdown
Contributor Author

i was all like "time to do the work, just before 4.0, to get these threaded cellular decompositions to decompose the same as the serial ones". (i'm doing work in a yet-private repo on a re-implementation of bertini_real). and then it uncovered a series of bugs that demanded solving. so this fixes two issues in the core. and brings some interface consistency with numpy (0 just a seed, it doesn't mean "no seed" or "use entropy"). and, along the way, further uncovered some incorrect metadata was being computed -- because i had not correctly reset the observers at the beginning of tracking a path. like, seriously, i am loving using AI coding tools. i steer, i guide, and i use my deep insight. claude starts trying to go off and do something, and i'm like, no buddy, you stay with me. here. this is the place, this is the bug. and i fix it.

this result permits me to do a 4.0 release (i'm doing semantic versioning, and this interface will break interface -- so this is the right place to do the rename of the env var that controls threading, separating it from the openmp variable i had used early on because it was convenient).

yes, i used an ai tool. i have faith this code is correct. i steered the whole time, and i personally intervened and made decisions the whole time. i am proud of this work.

@ofloveandhate
ofloveandhate merged commit 271059f into develop Oct 6, 2026
51 checks passed
@ofloveandhate ofloveandhate mentioned this pull request Oct 6, 2026
ofloveandhate added a commit that referenced this pull request Oct 6, 2026
Bertini 2 4.0.0.

- `VERSION` goes from `4.0.0.dev1` to `4.0.0`.
- The CHANGELOG heading is dated: `## [4.0.0] - 2026-10-06`. The publish
workflow refuses a final release whose top block still says
`unreleased`, and `tools/release_notes.py --tag v4.0.0` accepts this
one.
- The changelog preamble now says that a seeded solve gives the same
bits however it runs (#491). It also names the two changes from that PR
that alter existing scripts without a compile error:
`BERTINI_NUM_THREADS` replaces `OMP_NUM_THREADS`, and MPI ranks no
longer default to one thread; `set_random_seed(0)` is an ordinary seed.

Both files are in the build workflow's `paths-ignore`, so this PR runs
no CI. After it merges, tagging `v4.0.0` on `develop` publishes to PyPI
and creates the GitHub release from the changelog block.

🤖 Generated with [Claude Code](https://claude.com/claude-code)
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

1 participant