Skip to content

Add a ml user friendly API for scenario sweep - #14

Merged
BDonnot merged 9 commits into
mainfrom
pf_ml
Sep 30, 2026
Merged

BDonnot merged 9 commits into
mainfrom
pf_ml

Conversation

@BDonnot

@BDonnot BDonnot commented Sep 11, 2026

Copy link
Copy Markdown
Contributor
  • a diffirentiable powerflow there

Assisted-by: Claude Code (Sonnet 5)

Signed-off-by: DONNOT Benjamin <benjamin.donnot@rte-france.com>
Signed-off-by: DONNOT Benjamin <benjamin.donnot@rte-france.com>
Assisted-by: Claude Code (Fable 5.1)

Signed-off-by: DONNOT Benjamin <benjamin.donnot@rte-france.com>
…ernel

Assisted-by: Claude Code (Opus 5)

Signed-off-by: DONNOT Benjamin <benjamin.donnot@rte-france.com>
Three bugs in the persistent-driver path of PR #14, each with a regression
test:

- ScenarioSweepSession::run(): a hot run (only injections / gen_v changed)
  keeps the live batch source but used to clear every row's disconnected /
  masked_buses flags, which only a new source recomputes. Islanded rows
  still came back NaN while get_disconnected() and n_disconnected reported
  nothing. The path is now chosen first and the flags are reset only when
  a new source is built.
- ScenarioSweepSession::set_branch_data(): the live driver uploads the
  branch admittances / flow buffers once, so data set after the first run
  never reached the GPU (wrong flows, and an out-of-bounds write in
  zero_branch_flows_kernel if the branch count grew). It now invalidates the
  driver's cached branch data so the next run / compute_flows re-uploads.
- BatchPowerFlow: the topology and generator masks were cached before the
  session received them; a call raising in between (e.g. a bad gen_status
  shape) left the cache ahead of the session, and the next call with the
  same line_status was treated as unchanged and solved on the old trips.
  The masks are now committed only once the session accepted them.

Assisted-by: Claude Code
Claude-Session: https://claude.ai/code/session_01FZwHuL9k42bdoU93GqaPQW
Signed-off-by: Benjamin Donnot <benjamin.donnot@rte-france.com>
Found by a second review of PR #14 (including cc28dd5); each fix comes
with a regression test.

- BatchPowerFlow: the "session still holds a trip list / generator mask
  for another row count" check compared against _last_n_scen, which is
  updated before the op runs. After a call that failed inside the op, the
  next call with the same row count sent nothing and every run() was then
  refused ("row count no longer matches"). The row count of what the
  session holds is now recorded when the session accepts it.
- ScenarioSweepSession::run(): a warm run kept the chunk capacity the cold
  run sized for its own active rows (1 when all but one row was islanded),
  so a later warm run with every row active was solved in one-row chunks.
  The driver is now rebuilt when the live capacity would need more than
  twice a fresh driver's chunk count; the old driver is freed before the
  new one is allocated.
- compute_limit_violations = False had no effect on a reused driver: the
  fused check stayed armed with the old limits. run() now disarms it.
- set_branch_data(): refuses data that drops a branch the current topology
  trips (before replacing anything), rebuilds the per-row Ybus patches from
  the new admittances, and drops per-branch limits of a different length so
  compute_limit_violations asks for set_limits() again instead of reading
  past their end.
- differentiable compute_flows(): a half-open branch's open end (bus -1)
  read V[-1] and vn_kv[-1], i.e. the last bus. It now follows
  compute_branch_flows_kernel: V = 0 and no terminal current on that side,
  base current from the live endpoint.

Assisted-by: Claude Code
Claude-Session: https://claude.ai/code/session_01FZwHuL9k42bdoU93GqaPQW
Signed-off-by: Benjamin Donnot <benjamin.donnot@rte-france.com>
Each fix comes with a regression test.

- BatchPfDriver::solve_JT_batch: refuse a forward that ran in several
  chunks. Every adjoint buffer holds one chunk; with a caller-supplied J
  snapshot (which passes the shape check) the rhs gather wrote n_active
  rows into it, out of bounds, and the gen_v kernel read past d_V_batch /
  d_Ybus_values_batch.
- ScenarioSweepGPU: send gen_v to the session with the generators that
  set_contingency_gens() takes out of a row NaN'd there (re-sent when the
  mask changes). The conflict check already ignored them, but the session
  still applied their set-point, so an off generator could impose its
  voltage on a bus another connected generator keeps PV.
- ScenarioSweepSession::run(): with keep_final_jacobian, also rebuild the
  driver when the live capacity needs more than one chunk and a fresh
  driver would not; the forward otherwise refused a batch_size that was
  already large enough, on every topology change.
- ScenarioSweepSession::run(): count run_counter on entry and clear
  last_run_kept_jacobian_. A run() that threw after replacing the source or
  the driver left the counter unchanged, so a pending alias-mode
  BatchPowerFlow backward read the new buffers as its own.
- acpf_nr.cu: the single-system ||F||_inf reductions (NR loop and the
  presolved check) used thrust::maximum from 0, which drops NaN: an all-NaN
  F read as 0 and was reported converged. They now propagate NaN like the
  batched compute_residuals_kernel. test_matching_alg's NaN case now also
  asserts converged is False; docstrings updated.

Assisted-by: Claude Code
Claude-Session: https://claude.ai/code/session_01FZwHuL9k42bdoU93GqaPQW
Signed-off-by: Benjamin Donnot <benjamin.donnot@rte-france.com>
Signed-off-by: Benjamin Donnot <benjamin.donnot@rte-france.com>
Signed-off-by: Benjamin Donnot <benjamin.donnot@rte-france.com>
@BDonnot
BDonnot merged commit 2614562 into main Sep 30, 2026
19 checks passed
BDonnot added a commit that referenced this pull request Sep 30, 2026
Brings in the build CI (#16), the Modal GPU test job (#17) and the
scenario-sweep / adjoint fixes of #14.

Conflict resolutions:
- scenario_sweep_session.cu: keep main's warm path (rebuild the driver
  when the live chunk capacity would split the active rows into too many
  chunks, or break keep_final_jacobian's single chunk) and its reset_rows
  lambda; drop main's per-row slack weights from the ScenarioSweepBatch
  constructor, since this branch hands them over after the build through
  set_slack_redistribution.
- _batch_pf.py: take main's rule that _topology_mask only changes once
  the session accepted the topology; _row_masks (redistribute_slack's
  pre-pass, which runs before the op hands the topology over) now reads
  the queued topology and keys its cache on that tensor.
- _flows.py: per-terminal ampere bases (this branch) with main's
  half-open handling; a -1 side falls back to the live end's vn_kv, as
  set_branch_data does.
- scenario_sweep/gpu_facade.py: main's _push_gen_v helper, with this
  branch's gen_v_bus mapping.

Assisted-by: Claude Code
Claude-Session: https://claude.ai/code/session_01Afi372v35Q3tbaDtypd9Fa
Signed-off-by: Benjamin Donnot <benjamin.donnot@rte-france.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant