Skip to content

Light environment, topological actions in ScenarioSweep, and batch slack fixes - #226

Open
BDonnot wants to merge 52 commits into
dev_1.1.1from
fast-env
Open

BDonnot wants to merge 52 commits into
dev_1.1.1from
fast-env

Conversation

@BDonnot

@BDonnot BDonnot commented Oct 5, 2026

Copy link
Copy Markdown
Collaborator
added removed
C++ (src/core, src/bindings) +3338 -177
Python (package) +191 -8
Tests +2370 -2
Docs + changelog +405 -0
Build / other +294 -0
total +6598 -187

fast-env already has dev_1.1.1 merged in (7e30c26), so this merges without conflicts.

What it brings

Light environment (LightEnv, C++ with Python bindings): a grid2op-like environment running on an LSGrid.

  • init_actions registers checked topological actions; step(act_id) plays them with grid2op-like cooldowns.
  • Protections against overflow.
  • LightEnvObservation: read-only views on the env's state, with no copy.
  • Cheap copies; noexcept moves in C++.
  • A docs page (docs/light_env.rst) and a benchmark against grid2op (benchmarks/light_env.py).

ScenarioSweep.set_topo_actions: one topological action per row (disconnections, reconnections, a bus split or merged), still on a single symbolic analysis.

  • The solver labelling is built for the union of the buses the rows use.
  • The admittance entries a row writes are reserved up front.
  • Each row is value edits only.
  • What is still refused is listed under [TODO] in the changelog.

Core: parameters that took bare references to Eigen objects now take Eigen::Ref.

Fixes made after merging dev_1.1.1

Each code fix came with a test that failed before it.

  • The slack pre-pass now counts each element where the row's action puts it (redistribute_slack + set_topo_actions). A moved generator still injects; a load or storage unit the action disconnects, or leaves alone on a busbar, is lost; a reactivated generator is a gain. A unit that can take part in the slack takes its share on its new bus.
  • redistribute_slack is refused with a single-slack algorithm (NRSing_*, Gauss-Seidel). Those ignore the weights, so before, the voltages came out unchanged while compute_physical_violations read targets the solve never used. New capability flag: BaseAlgo::distributes_slack().
  • Refactor fallback when a row has slack weights of its own. A unit saturated by the pre-pass has its weight set to 0. With KLU that zero could land on a pivot kept from the base factorization, and the row then failed to converge without handle_disconnected_grid. Reproduced on l2rpn_case14_sandbox with NR_KLU.
  • Topological actions on an algorithm that cannot mask a bus now get an error that names set_topo_actions (it used to name handle_disconnected_grid). Actions that change nothing are plain rows on any algorithm.
  • C++ tests: TempFile names carry a per-process random token. Before, ctest -j processes could write the same file: a parallel stress run failed 13 of 24 file-writing tests, and now passes.
  • Housekeeping: shorter changelog entries; removed a root-level scratch script.

Testing

  • C++ (Catch2): 398/398, serially and with ctest -j 16 --repeat until-fail:3.
  • Python (unittest, grid2op installed from source): test_ScenarioSweep_topology, test_batch_redistribute_slack, test_ScenarioSweep, test_ScenarioSweep_gen_contingency, test_ContingencyAnalysis_split, test_ContingencyAnalysis_limit_violations, test_physical_violations, test_lsgrid_violations, test_main_component_slack, test_can_participate_slack, test_cap_slack_at_active_limits, test_binary_serialization, test_storage_pypowsybl, test_voltage_control_batch, test_hvdc_batch, test_fdpf, test_solver_control, test_SecurityAnlysis, test_LightEnv.
  • The dist_slack_algorithm plugin example builds and its test_plugin.py passes.

Note for review: DCO

Six older commits on this branch (February 2025: 525803d, 6a286c5, 5f080d0, 2a34c6b, a0ca58a, 1d6c68e) have no Signed-off-by: trailer, so the DCO check is expected to fail until they are signed off.

🤖 Generated with Claude Code

https://claude.ai/code/session_01ApvmWkD2kwsiB7HscvzfuV


Generated by Claude Code

BDonnot and others added 30 commits December 18, 2024 15:54
Signed-off-by: DONNOT Benjamin <benjamin.donnot@rte-france.com>
Signed-off-by: DONNOT Benjamin <benjamin.donnot@rte-france.com>
Signed-off-by: DONNOT Benjamin <benjamin.donnot@rte-france.com>
Signed-off-by: DONNOT Benjamin <benjamin.donnot@rte-france.com>
Signed-off-by: DONNOT Benjamin <benjamin.donnot@rte-france.com>
Signed-off-by: DONNOT Benjamin <benjamin.donnot@rte-france.com>
Signed-off-by: DONNOT Benjamin <benjamin.donnot@rte-france.com>
Signed-off-by: DONNOT Benjamin <benjamin.donnot@rte-france.com>
Signed-off-by: DONNOT Benjamin <benjamin.donnot@rte-france.com>
Signed-off-by: DONNOT Benjamin <benjamin.donnot@rte-france.com>
add a mode to keep only the main cc when performing steps
Signed-off-by: DONNOT Benjamin <benjamin.donnot@rte-france.com>
Signed-off-by: DONNOT Benjamin <benjamin.donnot@rte-france.com>
Signed-off-by: DONNOT Benjamin <benjamin.donnot@rte-france.com>
Signed-off-by: DONNOT Benjamin <benjamin.donnot@rte-france.com>
Signed-off-by: DONNOT Benjamin <benjamin.donnot@rte-france.com>
Signed-off-by: DONNOT Benjamin <benjamin.donnot@rte-france.com>
Sync the fast-env branch with the head of the fix_derivatives branch.

The PR renamed GridModel to LSGrid, moved the core under src/core and
split the pybind11 module into src/bindings/python/binding_*.cpp, so the
light-env work carried on this branch is ported onto that layout:

- src/light_env/*.hpp move to src/core/light_env/, are wrapped in the
  ls2g namespace, and target LSGrid (change_bus_* now takes a
  GridModelBusId).
- the LightEnv / Protections / TopoAction / ElementType bindings that
  lived at the end of src/main.cpp become src/bindings/python/
  binding_light_env.cpp, registered through bind_light_env().
- test_light_env.py imports from lightsim2grid.lightsim2grid_cpp.

The three modify/delete conflicts (src/GridModel.hpp,
src/element_container/SGenContainer.cpp, src/main.cpp) are resolved by
taking the deletion: the fast-env edits there were a template parameter
rename, a blank line, and the bindings ported above.

Assisted-by: Claude Code (claude-fable-5-1)
Signed-off-by: DONNOT Benjamin <benjamin.donnot@rte-france.com>
…cooldowns

LightEnv.init_actions (lightsim2grid/lightEnv.py) takes a list of grid2op
actions (set_bus / set_line_status only) or TopoAction objects. Every
action is checked in C++ against the initial grid: a non-existing element,
a busbar outside -1 / 0 / 1..n_busbar_per_sub (eg -2), a bad status or a
contradiction between set_line_status and set_bus on a line end raises a
ValueError naming the action, and nothing is registered.

TopoAction now speaks grid2op: local busbar ids resolved through the
substation of each element (new get_subid getters on the containers),
grid2op line numbering (powerlines then transformers), set_line_status.

step(act_id) applies the action when legal. Cooldowns follow grid2op's
rules and update order: an action touching a substation / line still in
cooldown is illegal (info["is_illegal"]), elements the action touched get
nb_timestep_cooldown_sub / nb_timestep_cooldown_line, lines tripped by the
protections get nb_timestep_reconnection. reset restores the initial
topology and the protections' counters.

Two pre-existing bugs fixed on the way: the protections looped for ever
once a line had tripped (its overflow counter never went down), and the
observation returned by reset was a stale rho.

Tests (lightsim2grid/tests/test_LightEnv.py) compare every step with a
grid2op env on the same grid and chronics: illegal flag, cooldown vectors,
line statuses, load / gen buses and flows.

Assisted-by: Claude Code (claude-fable-5-1)
Signed-off-by: DONNOT Benjamin <benjamin.donnot@rte-france.com>
ScenarioSweep.set_topo_actions takes one TopoAction per row (grid2op
semantics: set_bus / set_line_status; the grid2op wrapper converts
grid2op actions), checked against the grid when registered.

This first stage plays disconnections only, through the same
fixed-sparsity paths as the masks, so the whole sweep still runs on one
symbolic analysis:
- a branch (set_line_status -1, set_bus -1 on an end) joins the row's
  disconnected-branch list (YbusPolicy::Contingency::topo_branches_off,
  merged by branch_ids_for_row: same Ybus edit, connectivity, current
  checks and flow cleaning as a mask entry);
- a generator (set_bus -1) joins the generator-contingency path
  (SbusPolicy::Vary::topo_gen_off, read with gen_off through gen_off_in:
  injection, slack re-weighting, PV -> PQ pinning);
- a load or a storage unit (set_bus -1) is taken back out of the row's
  injection (topo_loads_off / topo_storages_off).

A row naming an element both in a mask and in its action, a move
(set_bus > 0), a reconnection and the DC algorithm are refused with a
message; the changelog TODO section lists them. TopoPlan.hpp resolves an
action into per-element placements against the base grid (the
representation the next stages extend). TopoAction exposes its resolved
entries and is bound with apply_to_gridmodel.

Tests compare every row with a one-off ac_pf on a copy of the grid with
the action really applied (voltages and flows), and pin nb_analyze == 1,
threads agreement, the masked / NOT_SIMULATED split-row behaviour and the
refusals: lightsim2grid/tests/test_ScenarioSweep_topology.py and
src/tests/test_scenario_sweep_topology.cpp.

Assisted-by: Claude Code (claude-fable-5-1)
Signed-off-by: DONNOT Benjamin <benjamin.donnot@rte-france.com>
…PQ -> PV)

A row may put back a generator disconnected in the base grid, on the bus
it was last on (set_bus to that busbar). Its bus turns PV for the row at
constant sparsity: a base-PQ bus already owns a Vm unknown and a Q
equation, so the row only pins that equation (set_pv_pinned_buses) and
seeds |V| there at the generator's set-point -- the row's own when
modify_gen_v gives one. A generator that does not regulate voltage is an
injection and its bus keeps its label.

- _maybe_prepare_gen_contingency now builds a per-row pinned set
  (_row_pv_pinned_: the switchable buses still PV plus the buses a
  reactivated generator pins) and per-row |V| seeds (_row_vm_seed_,
  applied by _apply_step_topo_seed after _apply_step_gen_v). A base-PV
  bus whose own controllers the row takes out stays PV through a
  reactivated generator, provided both ask the same magnitude.
- SbusPolicy::Vary::topo_gens_on adds the generator's P (and Q if it does
  not regulate) to the row; VoltageSourceContainer gains
  would_be_local_voltage_controller (the status-free predicate).
- refused with a message: a slack participant, a slack bus, a remote or
  group-held controller, another busbar, keep_jacobian /
  compute_physical_violations with a reactivation.

Tests: the stage 1 file runs again on a base grid with a generator off
(TestScenarioSweepGenReactivation) plus the reactivation rows against a
one-off ac_pf, with and without a per-row set-point, the refusals, and
nb_analyze == 1; the Catch2 test adds a non-regulating base-off
generator and the same cases.

Assisted-by: Claude Code (claude-fable-5-1)
Signed-off-by: DONNOT Benjamin <benjamin.donnot@rte-france.com>
…, reconnections)

A row may now move elements between busbars and reconnect branches --
what creates a bus (a busbar nothing stood on, now used) or merges two --
with the whole sweep still on one symbolic analysis. The union layout:

- LSGrid::init_converter_bus_id / build_solver_input take an optional
  list of buses to keep in the solved system although empty in the grid;
  the sweep hands it every bus any row uses (_maybe_topo_union, L1).
- the admittance entries a row writes (the diagonal of every extra bus,
  the four entries of every branch placement) are reserved as stored
  zeros right after the cache is built (_maybe_reserve_union_pattern),
  so the base solve builds the superset Jacobian sparsity; extra buses
  are seeded with a finite |V| and masked (identity rows) in the base
  case and in every row that leaves them empty (_base_masked_, the
  masked loop's resting state -- the masked loop runs whenever actions
  are registered).
- a row is value edits: YbusPolicy::Contingency::topo_branches_moved
  (the base contribution out, the raw block in at the row's buses,
  negated coefficients through the existing -= path);
  SbusPolicy::Vary::topo_loads_on / topo_storages_on / topo_gens_on
  (with their gridmodel bus); the generator pinning of stage 2 (a bus a
  generator lands on is held, a bus its last generator leaves is
  solved for); per-row connectivity by the search whenever a row places
  a branch, with connected meaning nothing stranded beyond the base
  mask, and NOT_SIMULATED kept without handle_disconnected_grid.
- flows and current-limit checks read a row's own branch placements
  (BranchBusOverride, BaseBatchSolverSynch::_row_branch_overrides_).

Refused: moving or reactivating a slack participant, a generator on a
slack bus or one regulating remotely / in a control group, a regulating
storage unit, keep_jacobian / compute_physical_violations with a
generator move, the DC algorithm (changelog TODO).

Tests: substation splits and merges, a dangling bus, a generator moved
to a new busbar, an island of one bus (masked), reconnections where the
branch was and on a new busbar, two substations rewired in one row,
flows and current violations of moved lines, four threads vs one, and
nb_analyze == 1 throughout -- every row against a one-off ac_pf on the
rewired grid (lightsim2grid/tests/test_ScenarioSweep_topology.py,
src/tests/test_scenario_sweep_topology.cpp).

Assisted-by: Claude Code (claude-fable-5-1)
Signed-off-by: DONNOT Benjamin <benjamin.donnot@rte-france.com>
BDonnot and others added 22 commits September 18, 2026 18:08
Assisted-by: Claude Code (Fable 5.1)
Signed-off-by: DONNOT Benjamin <benjamin.donnot@rte-france.com>
Follows 57e73e1 (compute_physical_violations following a row's generator
placements), whose Python tests did not run:

- the pandapower loader expects every busbar to exist in the network
  (n_sub x n_busbar buses): the test adds busbar 2 of each substation
  before loading case14 with n_busbar_per_sub=2;
- the reference rebuilt from the base grid instead of the grid under
  test, so the reactivation case compared with the generator on;
- the "every overloaded bus is reported" check now skips the slack bus,
  which the reactive check does not cover on a plain row either.

Assisted-by: Claude Code (claude-opus-5)
Signed-off-by: DONNOT Benjamin <benjamin.donnot@rte-france.com>
step() returned current_step_ / max_iter_ as reward: an integer division,
and by the powerflow iteration cap rather than the episode length, so the
reward was 0 for most of an episode. info["survival_time"] was
current_step_ / max_step_, also an integer division, so always "0" or "1".
Both are now survival_ratio(), the fraction of the episode survived as a
double (1 on success, current_step / max_step otherwise).

The bindings of LightEnv.reset / assign_time_series and of the thermal
limit and max overflow getters / setters of Protections had "TODO" as
docstring; they are documented (units, numbering, which injections are
actually applied), and the arguments are named.

Tests: test_LightEnv (reward over a short episode, survival_time on
success and on a divergence).

Assisted-by: Claude Code (claude-opus-5-5)
Claude-Session: https://claude.ai/code/session_019LN3o2bJBuNuNf9sN2tcNb
Signed-off-by: Benjamin Donnot <benjamin.donnot@rte-france.com>
reset() and step() returned a copy of rho. They now return the env's
LightEnvObservation, which holds only a pointer to its env; every getter
returns an Eigen::Ref<const ...> and the bindings expose them as read-only
numpy views (def_property_readonly, so reference_internal: the array keeps
the observation alive, which keeps the env alive).

What it exposes: rho, p / q / a on both sides of every line (grid2op
numbering, powerlines then trafos, currents in kA like the thermal
limits), topo_vect (local busbar ids at their grid2op position), the line
and substation cooldowns, load_p, gen_p and current_step.

Where the data lives:
- rho, load_p, gen_p, the cooldowns: read where they already are (the
  protections, the grid results, the env), nothing is copied.
- the flows and topo_vect are not stored in grid2op order anywhere (lines
  and trafos are separate containers, the grid knows buses, not local
  busbars), so extract_observation writes them once per step into
  env-owned buffers, which the observation then views.

Lifetime: the env buffers (flows, topo_vect, cooldowns) are allocated
once in the constructor and only written in place (reset_cooldowns now
setZero()s instead of reallocating), so a view on them stays valid for
the life of the env. rho and the grid results are reallocated at reset
(new LSGrid, Protections::check_validity) and on a divergence
(reset_results), so a view is documented as valid until the next reset
or a step ending the episode by a divergence. LightEnv is made
non-copyable and non-movable since its observation points to it.

topo_vect needs the grid2op positions (set_*_pos_topo_vect, done by
LightSimBackend): new const getters get_pos_topo_vect[_side_1/2] on the
containers, and SubstationContainer::gridmodel_to_local, the inverse of
local_to_gridmodel. Without the positions, topo_vect raises.

Tests: test_LightEnv. assert_same_state now also compares the light
observation with the grid2op one at every step of the existing tests
(topo_vect, flows, load_p, gen_p, cooldowns); new tests check the arrays
are read-only views that follow the env across a step, and that a view
keeps its env alive.

Assisted-by: Claude Code (claude-opus-5-5)
Claude-Session: https://claude.ai/code/session_019LN3o2bJBuNuNf9sN2tcNb
Signed-off-by: Benjamin Donnot <benjamin.donnot@rte-france.com>
A LightEnv could not be copied: its live grid is a unique_ptr (LSGrid has
a copy constructor but no copy assignment) and its observation points back
to it. A copy is what a tree search (or any look-ahead) on the env needs.

The state now lives in a LightEnvState base, copied member by member, and
LightEnv only adds the observation: its copy / move constructors and
assignments copy the base and keep the observation pointing to the env it
belongs to. Having the state apart means a member added later cannot be
forgotten by a hand-written copy. The live grid is held by a small
DeepCopyPtr (a unique_ptr whose copy clones).

What never changes during an episode is shared between copies through a
shared_ptr<const ...>: the initial grid, the time series (now one
TimeSeries struct) and the registered actions. assign_time_series /
init_actions replace that env's pointer only, so the copies do not see
each other's changes. Everything else (live grid, protections and their
counters, cooldowns, step, observation buffers) is copied: a copy is an
independent env at the same point of the same episode. On case14 a copy
takes ~11 us, a step ~17 us.

reset() without assign_time_series now raises instead of reading row 0 of
an empty matrix.

Python: LightEnv(other), __copy__ / __deepcopy__ in the bindings, and
LightEnv.copy() / __copy__ / __deepcopy__ in lightEnv.py so a copy keeps
the python subclass (its init_actions accepting grid2op actions).

Tests: test_LightEnv (a copy has the same state in its own memory, steps
without touching the original and gives the same result on the same
step, outlives the original, and does not share a later
assign_time_series).

Assisted-by: Claude Code (claude-opus-5-5)
Claude-Session: https://claude.ai/code/session_019LN3o2bJBuNuNf9sN2tcNb
Signed-off-by: Benjamin Donnot <benjamin.donnot@rte-france.com>
The move constructor / assignment added with the copies were not noexcept,
so a std::vector<LightEnv> copied every env when it grew instead of moving
them. Every member of LightEnvState is nothrow movable (LSGrid itself is
not, it is behind DeepCopyPtr), so the moves are now noexcept; two
static_asserts next to LightEnvState keep it that way when a member is
added.

A move copies nothing: the grid, the protections and the observation
buffers change owner, and the observation of the env moved to views them
at the same address.

A moved-from env had a null grid: get_grid(), reset(), step() or
obs.load_p / gen_p dereferenced it. It can now only be destroyed or
assigned to (which makes it usable again); anything else throws
std::logic_error naming the call. step() also checks, so it does not ask
for a reset() that would throw in turn.

Tests: new src/tests/test_light_env_copy_move.cpp (Catch2), since the
moves only exist on the C++ side: copies are independent and step like
the original; a move keeps the buffers and the grid at their address and
steps like the original would have; a moved-from env throws and is usable
again once assigned to; envs in a growing std::vector keep their own
observation. Full C++ suite green (353), and the [light_env] cases under
the ASan + UBSan build (__SANITIZE=1).

Assisted-by: Claude Code (claude-opus-5-5)
Claude-Session: https://claude.ai/code/session_019LN3o2bJBuNuNf9sN2tcNb
Signed-off-by: Benjamin Donnot <benjamin.donnot@rte-france.com>
The Windows wheels failed on the static_assert guarding the noexcept moves:

  light_env.hpp(195,66): error C2338: static_assert failed:
  'LightEnvState should be nothrow move constructible'

MSVC's std::unordered_map allocates a sentinel node, so its move
constructor is not noexcept (libstdc++'s is, hence green on Linux and
macOS), and LightEnvState held one: the step info.

That map was never state: reset() and step() cleared it first thing,
filled it and returned a copy of it. It is now a local of step() (reset()
returns an empty one), which removes the member and a map copy per step.
Every remaining member is a scalar, an Eigen vector, a std::vector, a
shared_ptr or a unique_ptr, all nothrow movable on MSVC too. The comment
on the static_asserts now names the containers to keep out of the state.

Checked: the three failed Windows jobs all stop on that one assert and
nothing else; C++ [light_env] cases and test_LightEnv green on Linux.

Assisted-by: Claude Code (claude-opus-5-5)
Claude-Session: https://claude.ai/code/session_019LN3o2bJBuNuNf9sN2tcNb
Signed-off-by: Benjamin Donnot <benjamin.donnot@rte-france.com>
light env: zero-copy observation, copyable / movable LightEnv, reward fix
Recent commits (light env, ScenarioSweep topology, slack participation,
voltage-control override) introduced parameters taking bare `const RealVect &`
/ `RealMat &` / `CplxVect &` where the rest of the code base uses Eigen::Ref.
Aligned them with the convention:

- read-only inputs -> `const Eigen::Ref<const T> &`: InjAction ctor,
  LightEnv::assign_time_series (bound to Python), LightEnv::aux_fill_topo_vect,
  TopoAction::aux_check_subid / aux_resolve_one_side, and the
  set_voltage_control_v_set chain (BaseAlgo, NRAlgo, NRSystem,
  AlgorithmSelector, VoltageControl::set_v_set_override);
- in-place outputs that are never resized -> `Eigen::Ref<T>`:
  SlackParticipation::accumulate_raw / split, the generator and storage
  accumulate_slack_weights_solver, BaseBatchSweep::_maybe_seed_extra_buses
  and _apply_step_topo_seed.

Left as they are on purpose: getters returning a const reference to a member
(get_subid, get_pos_topo_vect...), local aliases of members, and output
buffers that are resized or returned by reference.

No behaviour change.

Assisted-by: Claude Code
Claude-Session: https://claude.ai/code/session_01Ne9bEVwN1gxzvfKmccCj3g
Signed-off-by: Benjamin Donnot <benjamin.donnot@rte-france.com>
docs/light_env.rst: what LightEnv is for, a quick start building it from a
grid2op env, how to set up the time series / protections / cooldowns /
actions, what reset and step return, the observation (a zero-copy view),
copying for lookahead, the list of what it does not model compared with
grid2op, and the benchmark results.

benchmarks/light_env.py: the same l2rpn_case14_sandbox chronics played on a
grid2op env with LightSimBackend and on a LightEnv built with the same
thermal limits, protections and cooldowns (grid2op's hard overflow threshold
raised out of reach, the light env has none). Three workloads: do nothing,
a random unitary set_bus action every 10 steps, and a one-step lookahead
over 10 candidates (obs.simulate vs light_env.copy().step). Both sides play
the same number of steps and end the same episodes.

On the test chronics a step is 25 to 50 times faster than grid2op's, a
lookahead evaluation about 20 times faster than obs.simulate. The first step
of a copy is ~3x slower than a warm step: a copied LSGrid starts with a cold
solver.

Assisted-by: Claude Code
Claude-Session: https://claude.ai/code/session_01Ne9bEVwN1gxzvfKmccCj3g
Signed-off-by: Benjamin Donnot <benjamin.donnot@rte-france.com>
set_thermal_limit_or / _ex and set_max_line_time_step_overflow took a
mutable Eigen::Ref, which pybind11 can only bind to a writeable array of the
exact dtype: grid2op's thermal limits (float32) were refused, and so was any
read-only array. They only read their input, so they now take
`const Eigen::Ref<const T> &`, the convention elsewhere in the core.

Found while checking the snippets of the new documentation page.

Assisted-by: Claude Code
Claude-Session: https://claude.ai/code/session_01Ne9bEVwN1gxzvfKmccCj3g
Signed-off-by: Benjamin Donnot <benjamin.donnot@rte-france.com>
Eigen::Ref convention in recent core code; light env docs page and benchmark
Conflicts resolved:

- BaseBatchSweep.hpp: dev_1.1.1 moved the operational limit checks to
  OperationalCheck.hpp (and added violation_rel_tol); the per-row branch
  placement fast-env gave check_current_violations (BranchBusOverride, a
  branch a topological action moved read with its raw admittance block) is
  ported there. BranchBusOverride moves with it, so that LSGrid.cpp can keep
  including OperationalCheck.hpp without the batch headers.
  _select_ref_slack_and_masks keeps dev_1.1.1's solver-numbered reference
  slack and fast-env's "no handle_disconnected_grid: a row stranding a base
  bus stays NOT_SIMULATED" line; the call sites pass both rel_tol and the
  row's overrides; the topology and slack-redistribution members and helpers
  are both kept.
- StorageContainer.hpp: dev_1.1.1's storage_off mask on
  accumulate_slack_weights_solver, with fast-env's Eigen::Ref output.
- CHANGELOG.rst: dev_1.1.1's rewritten file, with fast-env's unreleased
  entries moved under [1.1.1] and its TODO entry kept.

Assisted-by: Claude Code
Claude-Session: https://claude.ai/code/session_01ApvmWkD2kwsiB7HscvzfuV
Signed-off-by: Benjamin Donnot <benjamin.donnot@rte-france.com>
With redistribute_slack, _prepare_slack_redistribution works out what a row
loses before the solve and shares it on the slack units within their limits.
It read every element off the base grid, so a row carrying a topological
action (set_topo_actions) was mis-counted:

- a generator the action moves is in topo_gen_off, so it was counted as lost
  power although it still injects on its new bus;
- a load or storage unit the action disconnects, or leaves alone on a busbar
  (masked), was not counted at all, nor a generator it reactivates;
- the elements moved off a busbar the row leaves empty were counted as
  stranded on that (masked) bus;
- a unit only flagged "can participate in the slack" was given its share on
  its base bus (storage) or left out (generator) when the action moved it.

The pre-pass now adds a term (d): for every unit the action places,
disconnects or reactivates, what the base balance counted for it minus what
the row's main component gets from it. Terms (a) and (b) leave those units
out, and the participants the action places stand on the bus it gives them.
A row without an action is computed exactly as before.

Tested against a one-off grid with the action applied and the true lost
power redistributed (test_ScenarioSweep_topology, the RedistributeSlack
classes): every row but the plain one and the generator disconnection failed
before.

Assisted-by: Claude Code
Claude-Session: https://claude.ai/code/session_01ApvmWkD2kwsiB7HscvzfuV
Signed-off-by: Benjamin Donnot <benjamin.donnot@rte-france.com>
…ithm

A single-slack algorithm (NRSing_*, Gauss-Seidel) ignores the slack weights
and leaves the whole imbalance on the reference bus. With redistribute_slack
the pre-pass still moved the set-points by those weights and took the units
it saturated out of them: the solve did not use any of it, the voltages came
out bit for bit as without the option, but compute_physical_violations read
the pre-pass targets -- a generator pushed far below its min_p was no longer
reported. This is the default algorithm of a grid2op environment's grid.

compute() now refuses the combination. The capability is a new BaseAlgo
flag, distributes_slack(), following the existing ones: true for the NR with
the MultiSlack extension (a compile-time trait on NRSystem), the DC and the
fast-decoupled algorithms; false for NRSing_*, Gauss-Seidel and, by default,
a plugin -- the dist_slack_algorithm example declares it.

Assisted-by: Claude Code
Claude-Session: https://claude.ai/code/session_01ApvmWkD2kwsiB7HscvzfuV
Signed-off-by: Benjamin Donnot <benjamin.donnot@rte-france.com>
Registered topological actions turn the masked row loop on (_mask_mode)
without handle_disconnected_grid. On an algorithm that cannot mask a bus
(Gauss-Seidel, the fast-decoupled one), compute() then failed with "the
`handle_disconnected_grid` mode requires a Newton-Raphson algorithm", an
option the caller never set. The error now says that set_topo_actions needs
a Newton-Raphson algorithm, and names the active one.

Assisted-by: Claude Code
Claude-Session: https://claude.ai/code/session_01ApvmWkD2kwsiB7HscvzfuV
Signed-off-by: Benjamin Donnot <benjamin.donnot@rte-france.com>
… algorithm

_mask_mode() turned the masked row loop on as soon as topological actions
were registered, even when every one of them changes nothing. On an
algorithm that cannot mask a bus (Gauss-Seidel, the fast-decoupled one) such
a batch was then refused, although nothing is ever masked -- while the DC
test already pins "a do-nothing action is a plain row". It now keys on
_topology_active_ (some row's action does something), which
_maybe_resolve_topology settles before any caller reads it.

Assisted-by: Claude Code
Claude-Session: https://claude.ai/code/session_01ApvmWkD2kwsiB7HscvzfuV
Signed-off-by: Benjamin Donnot <benjamin.donnot@rte-france.com>
A development script (it loads l2rpn_idf_2023 and times a LightEnv), not a
unittest and referenced nowhere. benchmarks/light_env.py and
docs/light_env.rst now cover what it was for.

Assisted-by: Claude Code
Claude-Session: https://claude.ai/code/session_01ApvmWkD2kwsiB7HscvzfuV
Signed-off-by: Benjamin Donnot <benjamin.donnot@rte-france.com>
The fast-env entries under [1.1.1] and the set_topo_actions [TODO] item ran
to six to thirteen lines each; brought to the one-to-four-line register
CLAUDE.md asks for. The detail lives in the docs and the commit messages.

Assisted-by: Claude Code
Claude-Session: https://claude.ai/code/session_01ApvmWkD2kwsiB7HscvzfuV
Signed-off-by: Benjamin Donnot <benjamin.donnot@rte-france.com>
…ts own

A row whose slack pre-pass saturates a slack unit (redistribute_slack) solves
with that unit's weight at 0. The weight is a value of the Jacobian's
slack_absorbed column, which the base ("n") factorization had non-zero, and
KLU's refactorization keeps the base pivot sequence: when the pivot sits on
that entry, klu_refactor meets a zero and the row fails. With
handle_disconnected_grid the rows took the masked path, which already turns
the fallback on, so the failure only showed without it.

Observed on l2rpn_case14_sandbox with NR_KLU, generator 1 joining the slack
of generator 5 with a tight range: every row saturating it failed, any load
change. SparseLU (it re-pivots) and NRRefactorRetry_KLU (one extra factorize)
solved the same rows. Whether a unit hits the pivot depends on KLU's ordering
of the grid: generator 0 in its place does not. dev_1.1.1's tests clamp a
unit that does not.

compute() now turns set_refactor_fallback on, for the member algorithm and
the threads' ones, whenever some row solves with slack weights of its own (a
participant its pre-pass saturates, or one it disconnects -- the same zero,
not reproduced there), as it already does for masked buses and PV / PQ
switches. The fallback only runs on a failure.

Assisted-by: Claude Code
Claude-Session: https://claude.ai/code/session_01ApvmWkD2kwsiB7HscvzfuV
Signed-off-by: Benjamin Donnot <benjamin.donnot@rte-france.com>
ctest runs every TEST_CASE in a process of its own, and with `ctest -j` they
run side by side in the build directory. TempFile named its files from a
per-process counter only, so the first file of two concurrent processes was
ls2g_unit_test_0.lsb in both, and one test read or removed the other's file.

Reproduced by running the 24 tests that write files with `ctest -j 16
--repeat until-fail:30`: 13 of them failed. With a 64-bit token drawn once per
process in the name (std::random_device mixed with the clock and the address
of a static, plain C++14, no platform #ifdef) all 24 pass the same run, and
the whole suite passes `ctest -j 16 --repeat until-fail:3`.

Assisted-by: Claude Code
Claude-Session: https://claude.ai/code/session_01ApvmWkD2kwsiB7HscvzfuV
Signed-off-by: Benjamin Donnot <benjamin.donnot@rte-france.com>
…e.com>

I, DONNOT Benjamin <benjamin.donnot@rte-france.com>, hereby add my Signed-off-by to this commit: 525803d
I, DONNOT Benjamin <benjamin.donnot@rte-france.com>, hereby add my Signed-off-by to this commit: 6a286c5
I, DONNOT Benjamin <benjamin.donnot@rte-france.com>, hereby add my Signed-off-by to this commit: 5f080d0
I, DONNOT Benjamin <benjamin.donnot@rte-france.com>, hereby add my Signed-off-by to this commit: 2a34c6b
I, DONNOT Benjamin <benjamin.donnot@rte-france.com>, hereby add my Signed-off-by to this commit: a0ca58a
I, DONNOT Benjamin <benjamin.donnot@rte-france.com>, hereby add my Signed-off-by to this commit: 1d6c68e

Assisted-by: Claude Code
Claude-Session: https://claude.ai/code/session_01ApvmWkD2kwsiB7HscvzfuV
Signed-off-by: DONNOT Benjamin <benjamin.donnot@rte-france.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant