From 8cbcf0cfb45a620b03f20625c7d6142e8ada65f9 Mon Sep 17 00:00:00 2001 From: Bobby Powers Date: Mon, 7 Sep 2026 16:24:33 -0700 Subject: [PATCH 01/10] engine: relative loop scores per partition, one owner A relative loop score is a loop's share of its cycle partition: its raw score over the sum of |score| across every loop of the partition (reference section 4.4). The exhaustive path grouped by (partition, slot) instead -- an arrayed loop's slot k competed only with other loops' slot k, and a scalar loop was broadcast into every slot bucket -- so on any model whose elements are coupled the denominators missed most of the partition's members. On test/cross_element_ltm the births loop read 1/2 at NYC and 2/3 at Boston where the partition holds four active members and the share is 1/3; on the arrays investigation's sliced_sum corpus model the dominant-period verdict was wrong for the whole run. A slot index is a position in one loop's own dimension space, so two arrayed loops over different dimension lists, or a scalar loop, attach no shared meaning to "slot k"; the partition is the only sound group. ltm_post::compute_rel_loop_scores is now the single owner: every (loop, slot) with a loop_score column is one member of its slot's partition (an unresolved slot is a Solo member of its own), the denominator is the partition-wide sum, a scalar loop has exactly one series and an arrayed loop one per slot, laid out step-major with the loop's own slot count as stride. The stride/broadcast machinery, the slot-0 convenience view, and the FFI's streaming per-element pair are deleted; libsimlin caches the owner's whole output per SimState and reads the requested slot (or the argmax-abs collapse the layout also uses) out of it, and discovery's rank_and_filter accumulates its partition totals and divides through the same group_totals / relative_series pair, so the two surfaces cannot disagree on the rule. The link-level normalization shares the same accumulator (NaN excluded, Inf kept, finite overflow saturating). Numbers pinned by tests, through the real pipeline on cross_element_ltm at every checked step: births 1/3 at both elements, the NYC migration_out loop -1/6, the cross-element migration_in loop +1/6, partition magnitudes summing to 1 -- from the engine owner, the libsimlin FFI (subscripted and bare ids), and pysimlin (Sim and Run.loops). A property test over generated per-slot partition vectors checks the owner bit-for-bit against a naive per-member reference and the partition identity at every step. The arrays investigation's de-subscripting oracle (report B), rerun on this tree with its dumper example rebuilt against the new owner: the relative-score rows that were wrong become exact on cross_element (4 -> 0 bad), sliced_sum (12 -> 0), per_element_eqns (4 -> 0), disjoint_dim (5 -> 0), mapped_dims (2 -> 0), share_sum (3 -> 0), bare_reducer, scalar_cofactor, variable_backed and dt_quarter (3 -> 0 each); raw scores and edges are unchanged on every model. The rows still wrong (cross_agg, cross_agg4, and the reducer-split models mean/min/max/stddev_loop, scalar_cofactor_inline, two_d, migration_matrix) are the stitching and split-partial defects, not the normalization. The oracle's summary lines: cross-agg-ltm: values 6 checked/0 bad | edges ok=12 bad=0 missing=0 split=0 | loops matched=7 raw_bad=0 rel_bad=7 only_arrayed=0 only_scalar=1 cross-element-ltm: values 11 checked/0 bad | edges ok=16 bad=0 missing=0 split=0 | loops matched=8 raw_bad=0 rel_bad=0 only_arrayed=0 only_scalar=0 arrayed-population-ltm: values 15 checked/0 bad | edges ok=12 bad=0 missing=0 split=0 | loops matched=6 raw_bad=0 rel_bad=0 only_arrayed=0 only_scalar=0 a2a_diagonal: values 15 checked/0 bad | edges ok=12 bad=0 missing=0 split=0 | loops matched=6 raw_bad=0 rel_bad=0 only_arrayed=0 only_scalar=0 bare_reducer: values 7 checked/0 bad | edges ok=12 bad=0 missing=0 split=0 | loops matched=3 raw_bad=0 rel_bad=0 only_arrayed=3 only_scalar=3 cross_agg4: values 8 checked/0 bad | edges ok=20 bad=0 missing=0 split=0 | loops matched=15 raw_bad=0 rel_bad=15 only_arrayed=0 only_scalar=9 disjoint_dim: values 10 checked/0 bad | edges ok=14 bad=0 missing=0 split=0 | loops matched=6 raw_bad=0 rel_bad=0 only_arrayed=0 only_scalar=0 dt_quarter: values 9 checked/0 bad | edges ok=18 bad=0 missing=0 split=3 | loops matched=13 raw_bad=0 rel_bad=0 only_arrayed=0 only_scalar=1 fixed_broadcast: values 7 checked/0 bad | edges ok=5 bad=1 missing=0 split=0 | loops matched=3 raw_bad=1 rel_bad=0 only_arrayed=0 only_scalar=0 mapped_dims: values 9 checked/0 bad | edges ok=12 bad=0 missing=0 split=0 | loops matched=6 raw_bad=0 rel_bad=0 only_arrayed=0 only_scalar=0 max_loop: values 6 checked/0 bad | edges ok=11 bad=1 missing=0 split=3 | loops matched=10 raw_bad=4 rel_bad=6 only_arrayed=0 only_scalar=1 mean_loop: values 6 checked/0 bad | edges ok=9 bad=3 missing=0 split=3 | loops matched=10 raw_bad=6 rel_bad=10 only_arrayed=0 only_scalar=1 migration_matrix: values 21 checked/0 bad | edges ok=27 bad=36 missing=0 split=9 | loops matched=283 raw_bad=274 rel_bad=283 only_arrayed=0 only_scalar=503 min_loop: values 6 checked/0 bad | edges ok=11 bad=1 missing=0 split=3 | loops matched=10 raw_bad=4 rel_bad=6 only_arrayed=0 only_scalar=1 per_element_eqns: values 6 checked/0 bad | edges ok=8 bad=0 missing=0 split=0 | loops matched=4 raw_bad=0 rel_bad=0 only_arrayed=0 only_scalar=0 scalar_cofactor_inline: values 7 checked/0 bad | edges ok=12 bad=6 missing=0 split=3 | loops matched=13 raw_bad=9 rel_bad=13 only_arrayed=0 only_scalar=13 scalar_cofactor: values 8 checked/0 bad | edges ok=14 bad=0 missing=0 split=0 | loops matched=3 raw_bad=0 rel_bad=0 only_arrayed=4 only_scalar=4 share_sum: values 9 checked/0 bad | edges ok=18 bad=0 missing=0 split=3 | loops matched=13 raw_bad=0 rel_bad=0 only_arrayed=0 only_scalar=1 sliced_sum: values 17 checked/0 bad | edges ok=24 bad=0 missing=0 split=0 | loops matched=12 raw_bad=0 rel_bad=0 only_arrayed=0 only_scalar=0 stddev_loop: values 6 checked/0 bad | edges ok=9 bad=3 missing=0 split=3 | loops matched=10 raw_bad=6 rel_bad=10 only_arrayed=0 only_scalar=1 two_d: values 12 checked/0 bad | edges ok=30 bad=12 missing=0 split=6 | loops matched=84 raw_bad=64 rel_bad=84 only_arrayed=0 only_scalar=352 variable_backed: values 7 checked/0 bad | edges ok=12 bad=0 missing=0 split=0 | loops matched=3 raw_bad=0 rel_bad=0 only_arrayed=3 only_scalar=3 Discovery's ranking statistic (mean_relative_contribution) divides through relative_series too, and the enumerated universe's totals (retain_circuits) go through add_to_total, so the summand rule has one owner on every path. --- docs/design/engine-performance.md | 4 +- docs/design/ltm--loops-that-matter.md | 55 +- docs/reference/ltm--loops-that-matter.md | 14 +- src/libsimlin/simlin.h | 8 +- src/libsimlin/src/analysis.rs | 219 +- src/libsimlin/src/lib.rs | 34 +- src/libsimlin/src/simulation.rs | 12 +- src/libsimlin/tests/integration/analysis.rs | 112 +- src/pysimlin/simlin/sim.py | 4 + .../tests/test_relative_loop_scores.py | 107 + src/simlin-engine/src/db.rs | 6 +- src/simlin-engine/src/db/analysis.rs | 4 +- src/simlin-engine/src/db/ltm/mod.rs | 10 +- src/simlin-engine/src/db/ltm_unified_tests.rs | 4 +- .../src/db/ltm_value_gate_tests.rs | 2 +- .../src/layout/detect_ltm_loops.rs | 64 +- src/simlin-engine/src/ltm/types.rs | 6 +- .../src/ltm_augment_zero_slot.rs | 2 +- src/simlin-engine/src/ltm_finding.rs | 108 +- src/simlin-engine/src/ltm_finding_enum.rs | 18 +- src/simlin-engine/src/ltm_post.rs | 2712 +++++------------ src/simlin-engine/tests/integration/layout.rs | 26 +- .../tests/integration/ltm_relative_scores.rs | 230 ++ src/simlin-engine/tests/integration/main.rs | 1 + .../tests/integration/simulate_ltm.rs | 37 +- 25 files changed, 1412 insertions(+), 2387 deletions(-) create mode 100644 src/pysimlin/tests/test_relative_loop_scores.py create mode 100644 src/simlin-engine/tests/integration/ltm_relative_scores.rs diff --git a/docs/design/engine-performance.md b/docs/design/engine-performance.md index c13e5e119..94ac22777 100644 --- a/docs/design/engine-performance.md +++ b/docs/design/engine-performance.md @@ -1009,8 +1009,8 @@ short of it rather than taking a ~2-3%. `db::ltm_value_gate_tests::a_nonfinite_target_arm_is_omitted_to_zero_not_nan`. Whether `0` is the better answer is **open** and tracked as #1022: `src/float.rs` argues an - engine-manufactured NaN is noise, while GH #542 built the `denom_summand` - exclusion specifically to preserve a `NaN` score as a per-loop "undefined + engine-manufactured NaN is noise, while GH #542 built the `ltm_post::group_totals` + NaN exclusion specifically to preserve a `NaN` score as a per-loop "undefined here" signal. The signal survives on the target's own series and on every live arm, so what changes is confined to arms with no causal dependence on their source. diff --git a/docs/design/ltm--loops-that-matter.md b/docs/design/ltm--loops-that-matter.md index 17cfcc63d..c3e540604 100644 --- a/docs/design/ltm--loops-that-matter.md +++ b/docs/design/ltm--loops-that-matter.md @@ -19,7 +19,7 @@ The implementation is split across these modules in `src/simlin-engine/src/`: | `ltm_finding.rs` | Post-simulation loop discovery for models too large for exhaustive enumeration: scoring, retention, ranking, and the cap | | `ltm_finding_enum.rs` | Discovery's exact candidate generator: union-graph elementary-circuit enumeration and its retention pass | | `ltm_finding_fallback.rs` | Discovery's shortest-path candidate generator, used when the enumeration cannot finish within its budgets or the caller's deadline | -| `ltm_post.rs` | Post-simulation computation: normalizes loop scores into relative loop scores using the cycle-partition mapping produced during LTM compilation | +| `ltm_post.rs` | Post-simulation computation: the one owner of relative loop-score normalization (`compute_rel_loop_scores`, every `(loop, slot)` a member of its slot's cycle partition) plus the `group_totals` / `relative_series` pair discovery normalizes through | The production entry point is the `model_ltm_variables` tracked function in `db/ltm/mod.rs`, invoked as part of `compile_project_incremental`. LTM compilation @@ -177,16 +177,26 @@ disconnected stock groups, each subcomponent has a separate loop dominance profi ### How Partitions Are Used -- **Exhaustive mode**: `generate_loop_score_variables()` records each loop's - partition on the emitted `loop_score` `LtmSyntheticVar`. Post-simulation, - `compute_rel_loop_scores()` (`ltm_post.rs`) groups loops by partition and - normalizes each loop score against the sum of absolute scores within its own - partition, ensuring structurally independent stock groups don't dilute each - other's scores. +- **Exhaustive mode**: `model_ltm_variables` records each loop's per-slot + partition vector (`LtmVariablesResult::loop_partitions`). Post-simulation, + `compute_rel_loop_scores()` (`ltm_post.rs`) makes every `(loop, slot)` a + member of its slot's partition -- a scalar loop one member, an arrayed loop + one per element -- and normalizes each member against the sum of absolute + scores over all members of that partition, so structurally independent + stock groups don't dilute each other's scores and an arrayed loop's element + competes with its siblings and with scalar loops exactly as the + de-subscripted model's N scalar loops would. The group is the partition and + nothing finer: a slot index is a position in one loop's own dimension space, + so keying on it too splits a partition into denominators that miss most of + its members. - **Discovery mode**: `rank_and_filter()` computes per-partition, per-timestep - score totals. A loop is retained if at any single timestep its absolute score - is >= `MIN_CONTRIBUTION` of its partition's total. This prevents globally tiny - but partition-dominant loops from being filtered out. + score totals with the same accumulator (`ltm_post::add_to_total`, through + `group_totals` for the discovered set and `retain_circuits` for the + enumerated universe) and divides through the same + `ltm_post::relative_series`. A loop is retained if at any single + timestep its absolute score is >= `MIN_CONTRIBUTION` of its partition's + total. This prevents globally tiny but partition-dominant loops from being + filtered out. ### Module-Internal Stocks and Partitions @@ -1131,8 +1141,8 @@ Over the materialized loops: non-survivor's mass is still in the denominator, matching exhaustive mode, where the enumerated set IS the universe. On the fallback path there is no universe to measure against, so the discovered set supplies its own totals. - `NaN` summands are excluded and `Inf` kept, mirroring - `ltm_post::denom_summand`. + `NaN` summands are excluded and `Inf` kept, the one accumulator + `ltm_post::add_to_total` applies on every path. 3. **Retention filter**, peak semantics: keep a loop if at ANY single step its |score| is >= `MIN_CONTRIBUTION` (0.1%) of its group's total there. This runs BEFORE any cap (GH #310), so a loop dominant in a small partition but @@ -1706,11 +1716,14 @@ dominance profiles. The loop-id → cycle-partition mapping is cached as `LtmVariablesResult::loop_partitions: HashMap>>` -- *per slot* of an A2A loop, since two elements of the same A2A loop can land in different cycle partitions (the slot's stocks differ). Relative loop scores are -derived post-simulation by `compute_rel_loop_scores` consumers (e.g. -`libsimlin::analysis`), normalizing each `(partition, slot)` loop score against -the sum of absolute scores in that partition at that slot -- so an independent -A2A loop's normalization does not cross-pollute a sibling A2A loop that -happens to share a loop ID but lives in a different partition. +derived post-simulation by `ltm_post::compute_rel_loop_scores` -- the one +owner every reader (`libsimlin::analysis`, the layout's importance series) +goes through -- which normalizes each slot against the sum of absolute scores +over every member of the slot's partition: the loop's sibling slots, other +arrayed loops' slots and scalar loops alike. Two slots of one A2A loop that +live in different partitions therefore never normalize against each other, +while two coupled slots do, and a scalar loop in the partition is one member +with one series. **Cross-element / mixed loops**: Circuits containing scalar nodes or with inconsistent variable-level structures. Each circuit becomes its own scalar @@ -2098,9 +2111,11 @@ cases remain deliberate carve-outs: IF-THEN-ELSE, loop score equations, generated variable structure - **`ltm_post.rs`**: Post-simulation relative loop score computation -- - partition grouping, SAFEDIV-0 semantics on empty-denominator timesteps, - property-based equivalence with the reference compile-time formula on - synthetic loop-score matrices + per-partition grouping of every `(loop, slot)`, Solo groups for unresolved + slots, NaN exclusion / Inf retention / saturating totals, SAFEDIV-0 on + empty-denominator timesteps, and a property test against a naive + per-member reference on generated per-slot partition vectors that also + checks the partition identity (each partition's magnitudes sum to 1) - **`ltm_finding_tests.rs`** (the `#[cfg(test)]` sibling of `ltm_finding.rs`), by family: diff --git a/docs/reference/ltm--loops-that-matter.md b/docs/reference/ltm--loops-that-matter.md index 7236c7cc6..f3178851c 100644 --- a/docs/reference/ltm--loops-that-matter.md +++ b/docs/reference/ltm--loops-that-matter.md @@ -612,7 +612,19 @@ The **relative loop score** normalizes by the sum of absolute loop scores: RelativeLoopScore(L) = LoopScore(L) / sum_Y(|LoopScore(Y)|) ``` -where the sum runs over all loops Y in the same cycle partition. +where the sum runs over all loops Y in the same cycle partition. In an arrayed model +the members of a partition are the loops of the de-subscripted model: each element of +an apply-to-all loop is one member, as is each scalar or cross-element loop, and every +member divides by the same partition sum (Section 15.4). + +> **Simlin implementation note.** `ltm_post::compute_rel_loop_scores` is the single +> owner of this normalization: every `(loop, slot)` with a `loop_score` column is one +> member of its slot's partition, and the libsimlin accessors and the layout's importance +> series read it rather than dividing on their own; discovery's ranking accumulates its +> partition totals with the same `add_to_total` (through `group_totals` for the +> discovered set, `retain_circuits` for the enumerated universe) and divides with the +> same `relative_series`. The group is the partition and nothing finer -- +> a slot index names an element only within one loop's own dimension space. Properties: - Normalized to range [-1, 1] diff --git a/src/libsimlin/simlin.h b/src/libsimlin/simlin.h index 071d1d327..956309f18 100644 --- a/src/libsimlin/simlin.h +++ b/src/libsimlin/simlin.h @@ -676,10 +676,10 @@ void simlin_free_links(SimlinLinks *links); // slab. // // The wasm-backend twin of `simlin_analyze_get_relative_loop_score`. Both -// FFIs funnel through `rel_loop_score_series` (extracted in Subcomponent A) -// over an `engine::Results` and the `(loop_partitions, loop_element_index)` -// snapshots, so the per-loop time series they produce cannot diverge by -// construction. +// FFIs resolve the loop id against the `loop_element_index` snapshot and +// funnel through `rel_loop_score_series` over an `engine::Results` and the +// `loop_partitions` snapshot, so the per-loop time series they produce +// cannot diverge by construction. // // Unlike the links twin, the rel-loop-score path needs the snapshots // `model_ltm_variables` derives (the per-loop partition map and slot diff --git a/src/libsimlin/src/analysis.rs b/src/libsimlin/src/analysis.rs index 2cc6d2a19..01657233e 100644 --- a/src/libsimlin/src/analysis.rs +++ b/src/libsimlin/src/analysis.rs @@ -1375,173 +1375,52 @@ pub unsafe extern "C" fn simlin_free_links(links: *mut SimlinLinks) { } } -/// The partition of loop `pv` at slot `k`. For an arrayed loop this is -/// `pv[k]` (out-of-range slots and genuinely-`None` partitions both yield -/// `None`); for a scalar loop (`len <= 1`) the single partition `pv[0]` is -/// broadcast across every slot it is compared in -- a scalar loop has no -/// elements, so it carries its one partition into every slot. This is the -/// same `slot_partition` convention `ltm_post::compute_rel_loop_scores_per_element` -/// uses to bucket loops into the `(partition, slot)` grid. -fn slot_partition_at(pv: &[Option], k: usize) -> Option { - if pv.len() <= 1 { - pv.first().copied().flatten() - } else { - pv.get(k).copied().flatten() - } -} - -/// Compute the per-(partition, slot) denominator series, lazily populating -/// the cache. A loop `other` is a member of bucket `(partition_key, -/// element_k)` iff its slot-`element_k` partition equals `partition_key`; -/// each member then contributes `|loop_score[other, k']|` where `k'` is -/// `effective_slot(n_slots[other], element_k)` -- 0 for a broadcast scalar -/// member, `element_k` for an arrayed member with `element_k < n_slots`, and -/// skipped entirely for an arrayed member past its own slots. This -/// reproduces `ltm_post::compute_rel_loop_scores_per_element`'s bucket sums -/// exactly via the streaming `compute_partition_denominator_for_element` -/// helper, just amortized across repeated FFI queries on the same bucket. -/// -/// GH #750: an unresolved (`None`) partition is a PER-LOOP singleton group -/// (the engine's `NormGroup::Solo` -- unrelated module-internal-stock / -/// lagged-stockless loops must not cross-normalize), so when `partition_key` -/// is `None` the only member is `querying_loop_id` itself. That bucket -/// bypasses the shared cache: its `(None, element_k)` key would collide -/// across distinct solo loops, and the single-member denominator is cheap to -/// recompute anyway. -fn ensure_denom_for_element( - cache: &mut HashMap<(Option, usize), Vec>, - results: &engine::Results, - loop_partitions: &engine::indexmap::IndexMap>>, - element_index_map: &HashMap, - partition_key: Option, - element_k: usize, - querying_loop_id: &str, -) -> Vec { - let loop_n_slots = |id: &str| -> usize { - element_index_map - .get(id) - .map(|m| m.n_slots) - .unwrap_or(1) - .max(1) - }; - if partition_key.is_none() { - return engine::ltm_post::compute_partition_denominator_for_element( - results, - std::iter::once((querying_loop_id, loop_n_slots(querying_loop_id))), - element_k, - ); - } - if let Some(cached) = cache.get(&(partition_key, element_k)) { - return cached.clone(); - } - let members: Vec<(&str, usize)> = loop_partitions - .iter() - .filter_map(|(id, pv)| { - if slot_partition_at(pv, element_k) == partition_key { - Some((id.as_str(), loop_n_slots(id))) - } else { - None - } - }) - .collect(); - let denom = engine::ltm_post::compute_partition_denominator_for_element( - results, - members.iter().copied(), - element_k, - ); - cache.insert((partition_key, element_k), denom.clone()); - denom -} - /// Backend-agnostic relative-loop-score time series for a *resolved* -/// (base, element_index, n_slots) query. -/// -/// `element_index` follows the FFI dispatch convention: `Some(k)` requests -/// a single slot (a scalar loop's only slot, or an arrayed loop's specific -/// element resolved from a `r1[Boston]`-style subscript); `None` requests -/// the argmax-abs aggregator across all `n_slots` (a bare ID on an arrayed -/// loop). This split is the *only* analytic dispatch in libsimlin's -/// rel-loop-score path, so concentrating it here ensures the VM FFI -/// (`simlin_analyze_get_relative_loop_score`) and the from-wasm FFI added -/// in a later task drive identical math. -/// -/// `cache` is a `&mut HashMap<(Option, usize), Vec>` shaped -/// exactly like `SimState::cached_partition_denominators`. The VM FFI -/// passes the persistent on-state cache (so repeated queries amortize); -/// the from-wasm FFI passes a stack-local empty cache (no persistent -/// sim). The split-borrow against `&mut state` is preserved -- callers -/// borrow `results`, `loop_partitions`, and `loop_element_index` from -/// `state` immutably while passing `&mut state.cached_partition_denominators` -/// as the cache argument. -/// -/// Returns `None` only for the engine's own missing-data signal -/// (`compute_rel_loop_score_for_element` / `compute_rel_loop_score_argmax_abs` -/// both `None` when `loop_score_{loop_id}` is absent from `results.offsets`); -/// the FFI shell maps that to a `DoesNotExist` error. Lookup of `loop_id` -/// in `loop_partitions` is the FFI shell's responsibility (its absence is -/// also `DoesNotExist`, but with a different message), so the core -/// indexes `loop_partitions[loop_id]` directly. -/// -/// Divergence from the design doc's single-core proposal: see the note -/// on `OwnedLink` and the phase plan -- splitting into two focused cores -/// (this one + the links core) more honestly satisfies AC5.1 than a -/// single signature carrying snapshots that the links analysis never -/// reads. +/// `(base, element_index)` query. +/// +/// The series come from the engine's one owner of the partition +/// normalization, `ltm_post::compute_rel_loop_scores`, computed over every +/// loop of `results` at once and kept in `cache` for the next query (the VM +/// FFI passes `SimState::cached_rel_loop_scores`, which `run_to_end`, `reset` +/// and `set_value_by_offset` reset alongside `results`; the from-wasm FFI +/// passes a fresh `None` per call). This function only READS a slot out of +/// the owner's per-`(loop, slot)` output; it never divides on its own, so the +/// FFI cannot drift from the layout's importance series or the engine tests. +/// +/// `element_index` follows the FFI dispatch convention: `Some(k)` requests a +/// single slot (a scalar loop's only slot, or an arrayed loop's specific +/// element resolved from a `r1[Boston]`-style subscript); `None` requests the +/// argmax-abs aggregate across all slots (a bare id on an arrayed loop), +/// `ltm_post::argmax_abs_by_step` -- the same collapse the layout applies. +/// +/// Returns `None` when `loop_id` has no `loop_score` column in `results` +/// (the owner omits it), or when `k` is past the slot count the owner laid +/// the loop out with (only reachable through a snapshot whose element index +/// and partition vector disagree, which valid compilation never produces); +/// the FFI shell maps that to a `DoesNotExist` error. pub(crate) fn rel_loop_score_series( results: &engine::Results, loop_partitions: &engine::indexmap::IndexMap>>, - loop_element_index: &HashMap, - cache: &mut HashMap<(Option, usize), Vec>, + cache: &mut Option>>, loop_id: &str, element_index: Option, - n_slots: usize, ) -> Option> { - let partitions = loop_partitions.get(loop_id)?; + let rel = cache + .get_or_insert_with(|| engine::ltm_post::compute_rel_loop_scores(results, loop_partitions)); + let series = rel.get(loop_id)?; + let step_count = results.step_count; match element_index { Some(k) => { - // Group the denominator by the *queried slot's* partition (slot 0 - // for a scalar loop), so an uncoupled A2A loop normalizes per - // element rather than against a pooled slot-0 bucket. - let partition_key = slot_partition_at(partitions, k); - let denom = ensure_denom_for_element( - cache, - results, - loop_partitions, - loop_element_index, - partition_key, - k, - loop_id, - ); - engine::ltm_post::compute_rel_loop_score_for_element( - results, loop_id, n_slots, k, &denom, - ) - } - None => { - // Argmax-abs aggregator over all slots; each slot's denominator is - // keyed on *that slot's* partition (matching the per-element - // helper), not slot 0's. - let mut denoms: Vec> = Vec::with_capacity(n_slots); - for k in 0..n_slots { - let partition_key = slot_partition_at(partitions, k); - let denom = ensure_denom_for_element( - cache, - results, - loop_partitions, - loop_element_index, - partition_key, - k, - loop_id, - ); - denoms.push(denom); + let n_slots = series.len().checked_div(step_count).unwrap_or(1).max(1); + // A resolved slot past the owner's layout can only come from a + // snapshot whose element index and partition vector disagree; + // report the loop as unknown rather than read another loop's data. + if k >= n_slots { + return None; } - let denom_refs: Vec<&[f64]> = denoms.iter().map(|d| d.as_slice()).collect(); - engine::ltm_post::compute_rel_loop_score_argmax_abs( - results, - loop_id, - n_slots, - &denom_refs, - ) + Some((0..step_count).map(|t| series[t * n_slots + k]).collect()) } + None => Some(engine::ltm_post::argmax_abs_by_step(series, step_count)), } } @@ -1674,8 +1553,8 @@ pub(crate) struct LtmSnapshots { /// project db's discovery flag). A discovery-mode blob carries loop-score /// columns for pinned loops only while this snapshot names every enumerated /// loop: a query for one of those resolves here, then fails the -/// `results.offsets` lookup in the core (`DoesNotExist`), and a partition -/// denominator omits the members without a column. +/// `results.offsets` lookup in the core (`DoesNotExist`), and the +/// partition totals omit the members without a column. /// /// The maps are empty when the derivation scored no loop; the caller's /// downstream `rel_loop_score_series` then naturally fails the @@ -1856,10 +1735,10 @@ pub(crate) fn resolve_loop_query<'a>( /// slab. /// /// The wasm-backend twin of `simlin_analyze_get_relative_loop_score`. Both -/// FFIs funnel through `rel_loop_score_series` (extracted in Subcomponent A) -/// over an `engine::Results` and the `(loop_partitions, loop_element_index)` -/// snapshots, so the per-loop time series they produce cannot diverge by -/// construction. +/// FFIs resolve the loop id against the `loop_element_index` snapshot and +/// funnel through `rel_loop_score_series` over an `engine::Results` and the +/// `loop_partitions` snapshot, so the per-loop time series they produce +/// cannot diverge by construction. /// /// Unlike the links twin, the rel-loop-score path needs the snapshots /// `model_ltm_variables` derives (the per-loop partition map and slot @@ -2032,21 +1911,19 @@ pub unsafe extern "C" fn simlin_analyze_rel_loop_score_from_wasm_results( } }; - // No persistent simulation backs this FFI, so the partition-denominator - // cache lives on the stack for this call. Repeated FFI queries against - // the same model will recompute denominators each time -- a trade-off the + // No persistent simulation backs this FFI, so the relative-score cache + // lives on the stack for this call. Repeated FFI queries against the + // same model recompute the normalization each time -- a trade-off the // from-wasm path accepts in exchange for not maintaining a per-call cache // on the wasm side (the wasm interactive-scrubbing flow re-runs the blob // and reaches this FFI fresh). - let mut cache: HashMap<(Option, usize), Vec> = HashMap::new(); + let mut cache: Option>> = None; let series = match rel_loop_score_series( &results, &loop_partitions, - &loop_element_index, &mut cache, resolved.base, resolved.element_index, - resolved.n_slots, ) { Some(s) => s, None => { @@ -2182,15 +2059,13 @@ pub unsafe extern "C" fn simlin_analyze_get_relative_loop_score( }; // Split-borrow `&mut state`: the cache (mutated) is borrowed disjointly - // from `results`/`loop_partitions`/`loop_element_index` (read-only). + // from `results`/`loop_partitions` (read-only). let series = match rel_loop_score_series( results, &state.loop_partitions, - &state.loop_element_index, - &mut state.cached_partition_denominators, + &mut state.cached_rel_loop_scores, resolved.base, resolved.element_index, - resolved.n_slots, ) { Some(s) => s, None => { diff --git a/src/libsimlin/src/lib.rs b/src/libsimlin/src/lib.rs index 2bb9d68ed..068b41a7e 100644 --- a/src/libsimlin/src/lib.rs +++ b/src/libsimlin/src/lib.rs @@ -550,13 +550,12 @@ pub(crate) struct SimState { /// `loop_partitions` intentionally), or when compilation itself /// failed. /// - /// `simlin_analyze_get_relative_loop_score` keys the rel-loop-score - /// denominator on the *queried slot's* partition (`loop_partitions[id][k]`) - /// -- so an element-wise-uncoupled A2A loop normalizes per element, exactly - /// as `ltm_post::compute_rel_loop_scores_per_element` does for the engine's - /// own consumers. When populated alongside `loop_element_index`, the two - /// agree on slot count for every loop (see the `debug_assert!` in - /// `simlin_sim_new`). + /// `simlin_analyze_get_relative_loop_score` normalizes every `(loop, slot)` + /// within *that slot's* partition (`loop_partitions[id][k]`) through the + /// engine's one owner, `ltm_post::compute_rel_loop_scores`, so the FFI + /// reads exactly the numbers the engine's own consumers see. When + /// populated alongside `loop_element_index`, the two agree on slot count + /// for every loop (see the `debug_assert!` in `simlin_sim_new`). pub(crate) loop_partitions: engine::indexmap::IndexMap>>, /// Snapshot of per-loop dimension metadata taken at /// `simlin_sim_new` time. Used by the FFI subscript resolver to @@ -575,22 +574,15 @@ pub(crate) struct SimState { /// tell exhaustive Johnson enumeration apart from the auto-flipped /// post-simulation discovery. pub(crate) ltm_mode: Option, - /// Per-(partition, slot) denominator series cached across FFI calls to - /// `simlin_analyze_get_relative_loop_score`. The rel-loop-score - /// definition is `loop_score / Σ|loop_score|` *within a cycle - /// partition*, evaluated at a specific element slot for arrayed - /// loops. The key is `(partition-of-slot-k, k)` -- the partition - /// component is the cycle partition of *that slot* (so an uncoupled - /// A2A loop's slots key into different partitions), matching - /// `ltm_post::compute_rel_loop_scores_per_element`'s bucket grid. - /// Keying this way lets repeated FFI queries against the same - /// bucket reuse the expensive sum. pysimlin's `_populate_loop_behavior` - /// walks every loop in a project; with this cache the per-partition - /// sum is computed once per slot and reused across all member loops. - /// Invalidated in lockstep with `results`: cleared on + /// The per-`(loop, slot)` relative loop scores of `results` + /// (`ltm_post::compute_rel_loop_scores` over `loop_partitions`), computed + /// on the first `simlin_analyze_get_relative_loop_score` query and reused + /// by every later one. pysimlin's `_populate_loop_behavior` walks every + /// loop in a project; with this cache the partition sums are accumulated + /// once. Invalidated in lockstep with `results`: reset to `None` on /// `simlin_sim_run_to_end`, `simlin_sim_reset`, and /// `simlin_sim_set_value_by_offset`. - pub(crate) cached_partition_denominators: HashMap<(Option, usize), Vec>, + pub(crate) cached_rel_loop_scores: Option>>, /// Resolved conveyor plans for a conveyor model (`None`/empty otherwise). /// `run_to_end` consumes the VM (`into_results`) and `reset` recreates it /// from `compiled`; a plain `Vm::new(compiled)` would drop the conveyor diff --git a/src/libsimlin/src/simulation.rs b/src/libsimlin/src/simulation.rs index 2fab86c46..73af8d62d 100644 --- a/src/libsimlin/src/simulation.rs +++ b/src/libsimlin/src/simulation.rs @@ -227,7 +227,7 @@ pub unsafe extern "C" fn simlin_sim_new( loop_partitions: captured_loop_partitions, loop_element_index: captured_loop_element_index, ltm_mode: captured_ltm_mode, - cached_partition_denominators: HashMap::new(), + cached_rel_loop_scores: None, conveyor_plans, queue_plans, }), @@ -300,7 +300,7 @@ pub unsafe extern "C" fn simlin_sim_run_to_end( match vm.run_to_end() { Ok(_) => { state.results = Some(vm.into_results()); - state.cached_partition_denominators.clear(); + state.cached_rel_loop_scores = None; } Err(err) => { state.vm = Some(vm); @@ -365,7 +365,7 @@ pub unsafe extern "C" fn simlin_sim_reset(sim: *mut SimlinSim, out_error: *mut * let mut state = sim_ref.state.lock().unwrap(); state.results = None; - state.cached_partition_denominators.clear(); + state.cached_rel_loop_scores = None; if let Some(ref mut vm) = state.vm { // Fast path: reuse existing VM allocation @@ -684,10 +684,10 @@ pub unsafe extern "C" fn simlin_sim_set_value_by_offset( *slot = val; // Defensive invalidation: only constant slots are writable now, // and a constant is never a `loop_score` input to the cached - // partition denominators -- but clearing the cache is cheap (the - // next FFI call repopulates lazily) and keeps this write path + // relative loop scores -- but dropping the cache is cheap (the + // next FFI call recomputes lazily) and keeps this write path // trivially safe against future loosening of the gate. - state.cached_partition_denominators.clear(); + state.cached_rel_loop_scores = None; return; } } diff --git a/src/libsimlin/tests/integration/analysis.rs b/src/libsimlin/tests/integration/analysis.rs index 6fb7f6e39..7f11ddc60 100644 --- a/src/libsimlin/tests/integration/analysis.rs +++ b/src/libsimlin/tests/integration/analysis.rs @@ -2207,7 +2207,7 @@ fn test_get_loop_element_count_arrayed_vs_scalar() { // `None` cohort and cross-normalized to a pooled value (each ~0.5); the per-slot // `loop_partitions` keep them separate. This test confirms the FFI exposes the // per-slot data correctly (the subscripted-loop-id accessor) and that what it -// returns matches the engine's `compute_rel_loop_scores_per_element` exactly -- +// returns matches the engine's `compute_rel_loop_scores` exactly -- // i.e. a round trip through the C API preserves the per-slot partitions. /// Build the two-A2A-subsystem datamodel project. @@ -2255,8 +2255,7 @@ fn engine_reference_rel_per_element( vm.run_to_end().unwrap(); let results = vm.into_results(); - let rel = - simlin_engine::ltm_post::compute_rel_loop_scores_per_element(&results, &loop_partitions); + let rel = simlin_engine::ltm_post::compute_rel_loop_scores(&results, &loop_partitions); (rel, n_slots_by_loop) } @@ -2352,7 +2351,7 @@ fn test_two_a2a_subsystems_per_slot_rel_score_round_trips() { assert_eq!( ffi_v, engine_v, "FFI rel-loop-score for {loop_id}[{elem_name}] at step {s} must match \ - compute_rel_loop_scores_per_element ({ffi_v} vs {engine_v})" + compute_rel_loop_scores ({ffi_v} vs {engine_v})" ); // AC2.1: each loop is alone in its (per-element) partition, // so the score is +1 once dynamics are nonzero -- NOT the @@ -3830,3 +3829,108 @@ fn polarity_label(p: SimlinLoopPolarity) -> &'static str { SimlinLoopPolarity::Undetermined => "U", } } + +// The FFI relative-loop-score accessor reads the engine's one normalization +// owner (`ltm_post::compute_rel_loop_scores`), so on a COUPLED arrayed model +// every `(loop, slot)` of a partition divides by the same partition sum. +// `test/cross_element_ltm` is the hand-computed case (the engine-side twin is +// `ltm_relative_scores::cross_element_loops_normalize_over_the_whole_partition`): +// one partition holds both regions; the births loop scores +1 at both slots, +// the NYC migration_out loop -0.5, the cross-element migration_in loop +0.5, +// every other loop 0; so the shares are 1/3, 1/3, -1/6 and +1/6 and their +// magnitudes sum to 1. + +/// The FFI loop whose variable set is exactly `vars`; returns its id. +unsafe fn ffi_loop_id_with_variables(loops: *mut SimlinLoops, vars: &[&str]) -> String { + let want: std::collections::HashSet<&str> = vars.iter().copied().collect(); + let loop_slice = std::slice::from_raw_parts((*loops).loops, (*loops).count); + let mut seen = Vec::new(); + for l in loop_slice { + let names: Vec = std::slice::from_raw_parts(l.variables, l.var_count) + .iter() + .map(|v| CStr::from_ptr(*v).to_str().unwrap().to_string()) + .collect(); + let id = CStr::from_ptr(l.id).to_str().unwrap().to_string(); + if names + .iter() + .map(String::as_str) + .collect::>() + == want + { + return id; + } + seen.push((id, names)); + } + panic!("no loop over {vars:?}; loops: {seen:?}"); +} + +#[test] +fn test_cross_element_rel_scores_share_one_partition_denominator_via_ffi() { + let xml = std::fs::read_to_string("../../test/cross_element_ltm/cross_element.stmx").unwrap(); + let project = simlin_engine::open_xmile(&mut std::io::BufReader::new(xml.as_bytes())) + .expect("the cross_element fixture parses"); + let pb = engine_serde::serialize(&project).unwrap(); + let mut buf = Vec::new(); + pb.encode(&mut buf).unwrap(); + + unsafe { + let (proj, model, sim) = open_arrayed_sim_with_ltm(&buf); + + let mut err: *mut SimlinError = ptr::null_mut(); + let loops = simlin_analyze_get_loops(model, &mut err); + assert!(err.is_null()); + let births = ffi_loop_id_with_variables(loops, &["population", "births"]); + let out_loop = ffi_loop_id_with_variables( + loops, + &["population", "migration_pressure", "migration_out"], + ); + let cross = ffi_loop_id_with_variables( + loops, + &[ + "population[nyc]", + "migration_pressure[boston]", + "migration_in[nyc]", + ], + ); + simlin_free_loops(loops); + + let read = |id: &str| { + read_relative_loop_series(sim, id) + .unwrap_or_else(|(c, m)| panic!("reading {id} failed: {c:?} {m}")) + }; + let births_nyc = read(&format!("{births}[NYC]")); + let births_boston = read(&format!("{births}[Boston]")); + let births_bare = read(&births); + let out_nyc = read(&format!("{out_loop}[NYC]")); + let out_boston = read(&format!("{out_loop}[Boston]")); + let out_bare = read(&out_loop); + let cross_series = read(&cross); + + // The ratios are time-invariant on this fixture (both populations + // grow at exactly 2% per step), so every step from the first active + // one reads the same shares. + for step in [2usize, 10, 30] { + let close = |got: f64, want: f64, what: &str| { + assert!( + (got - want).abs() < 1e-9, + "step {step}: {what} = {got}, expected {want}" + ); + }; + close(births_nyc[step], 1.0 / 3.0, "births[NYC]"); + close(births_boston[step], 1.0 / 3.0, "births[Boston]"); + close(births_bare[step], 1.0 / 3.0, "births (bare id, argmax-abs)"); + close(out_nyc[step], -1.0 / 6.0, "migration_out loop[NYC]"); + close(out_boston[step], 0.0, "migration_out loop[Boston]"); + close(out_bare[step], -1.0 / 6.0, "migration_out loop (bare id)"); + close( + cross_series[step], + 1.0 / 6.0, + "cross-element migration_in loop", + ); + } + + simlin_sim_unref(sim); + simlin_model_unref(model); + simlin_project_unref(proj); + } +} diff --git a/src/pysimlin/simlin/sim.py b/src/pysimlin/simlin/sim.py index 2b12e999e..a84405e57 100644 --- a/src/pysimlin/simlin/sim.py +++ b/src/pysimlin/simlin/sim.py @@ -450,6 +450,10 @@ def get_relative_loop_score( The relative loop score normalizes a loop's raw ``loop_score`` by the magnitudes of all loops that share its cycle-partition, so it reads as the loop's fractional contribution to model behavior at each step. + Every element of an arrayed loop is one member of its element's + partition, alongside the loop's other elements, other arrayed loops' + elements and scalar loops, so the magnitudes over a partition's + members sum to 1 at every active step. This requires the simulation to have been run with enable_ltm=True. diff --git a/src/pysimlin/tests/test_relative_loop_scores.py b/src/pysimlin/tests/test_relative_loop_scores.py new file mode 100644 index 000000000..131413183 --- /dev/null +++ b/src/pysimlin/tests/test_relative_loop_scores.py @@ -0,0 +1,107 @@ +"""Relative loop scores on a coupled arrayed model. + +The engine normalizes every (loop, element) of a cycle partition against the +same partition sum -- sibling elements of an arrayed loop, other arrayed +loops' elements and scalar loops alike -- and both ``Sim.get_relative_loop_score`` +and ``Run.loops`` read that one normalization. ``test/cross_element_ltm`` is +the hand-computed case: its two regions are coupled through migration, so one +partition holds every loop of the model, and because both populations grow at +exactly 2% per step the raw loop scores are constant once active: +1 at both +elements of the births loop, -0.5 at NYC for the migration_out loop (0 at +Boston, whose migration flows are clamped to 0), +0.5 for the cross-element +migration_in loop, 0 for every other loop. The partition sum is 3, so the +shares are 1/3, 1/3, -1/6 and +1/6, and their magnitudes sum to 1. +""" + +from __future__ import annotations + +import numpy as np +import pytest + +import simlin + +STEPS = (2, 10, 30) + + +def _loop_with_variables(loops, variables: set[str]): + for loop in loops: + if set(loop.variables) == variables: + return loop + raise AssertionError( + f"no loop over {sorted(variables)}; loops: {[(lp.id, lp.variables) for lp in loops]}" + ) + + +@pytest.fixture +def cross_element_model(cross_element_ltm_path): + return simlin.load(cross_element_ltm_path) + + +class TestCrossElementRelativeScores: + def test_per_element_shares_normalize_over_the_whole_partition( + self, cross_element_model + ) -> None: + loops = cross_element_model.loops + births = _loop_with_variables(loops, {"population", "births"}) + out_loop = _loop_with_variables( + loops, {"population", "migration_pressure", "migration_out"} + ) + cross = _loop_with_variables( + loops, {"population[nyc]", "migration_pressure[boston]", "migration_in[nyc]"} + ) + + with cross_element_model.simulate(enable_ltm=True) as sim: + sim.run_to_end() + assert sim.get_loop_element_count(births.id) == 2 + assert sim.get_loop_element_count(out_loop.id) == 2 + assert sim.get_loop_element_count(cross.id) == 1 + + births_nyc = sim.get_relative_loop_score(births.id, element="NYC") + births_boston = sim.get_relative_loop_score(births.id, element="Boston") + out_nyc = sim.get_relative_loop_score(out_loop.id, element="NYC") + out_boston = sim.get_relative_loop_score(out_loop.id, element="Boston") + cross_series = sim.get_relative_loop_score(cross.id) + + for step in STEPS: + assert births_nyc[step] == pytest.approx(1 / 3, abs=1e-9) + assert births_boston[step] == pytest.approx(1 / 3, abs=1e-9) + assert out_nyc[step] == pytest.approx(-1 / 6, abs=1e-9) + assert out_boston[step] == pytest.approx(0.0, abs=1e-9) + assert cross_series[step] == pytest.approx(1 / 6, abs=1e-9) + + # The partition identity over EVERY member: each element of each + # loop, summed, is 1. + for step in STEPS: + total = 0.0 + for loop in loops: + n = sim.get_loop_element_count(loop.id) + if n == 1: + total += abs(sim.get_relative_loop_score(loop.id)[step]) + else: + for element in ("NYC", "Boston"): + total += abs( + sim.get_relative_loop_score(loop.id, element=element)[step] + ) + assert total == pytest.approx(1.0, abs=1e-9) + + def test_run_loops_behavior_is_the_dominant_elements_share(self, cross_element_model) -> None: + """``Run.loops`` carries the bare-id series: the argmax-abs across an + arrayed loop's elements, read from the same per-element normalization. + """ + run = cross_element_model.run(analyze_loops=True) + assert run.ltm_mode == "exhaustive" + loops = run.loops + births = _loop_with_variables(loops, {"population", "births"}) + out_loop = _loop_with_variables( + loops, {"population", "migration_pressure", "migration_out"} + ) + cross = _loop_with_variables( + loops, {"population[nyc]", "migration_pressure[boston]", "migration_in[nyc]"} + ) + assert births.partition == out_loop.partition == cross.partition + + for step in STEPS: + assert births.behavior_time_series[step] == pytest.approx(1 / 3, abs=1e-9) + assert out_loop.behavior_time_series[step] == pytest.approx(-1 / 6, abs=1e-9) + assert cross.behavior_time_series[step] == pytest.approx(1 / 6, abs=1e-9) + assert np.all(np.isfinite(births.behavior_time_series)) diff --git a/src/simlin-engine/src/db.rs b/src/simlin-engine/src/db.rs index 7088c944a..41d93c49f 100644 --- a/src/simlin-engine/src/db.rs +++ b/src/simlin-engine/src/db.rs @@ -595,8 +595,8 @@ impl From for LtmOverlay { /// its cycle-partition index **per slot**: length 1 for scalar/cross-element/ /// mixed loops, one entry per element (in the runtime's row-major slot order) /// for A2A loops, matching `ltm_post::build_loop_element_index`'s `n_slots`. -/// Slots sharing a `(partition, slot)` key form the denominator when -/// `ltm_post::compute_rel_loop_scores*` normalizes; an element-wise-uncoupled +/// Every slot is one member of its partition's denominator when +/// `ltm_post::compute_rel_loop_scores` normalizes; an element-wise-uncoupled /// A2A loop's entries are N distinct partitions (the per-slot fix, GH #487), /// a coupled one's coincide, a `None` entry is a slot below the parent graph /// (e.g. a pure module-internal loop). Populated only in exhaustive LTM @@ -605,7 +605,7 @@ impl From for LtmOverlay { /// It is an `IndexMap` (not a `HashMap`) so iteration order is the loops' /// **emission order** -- the content-derived order `assign_loop_ids` produces /// and `model_ltm_variables` inserts in (enumerated loops first, then pinned). -/// The post-sim rel-loop-score denominator (`ltm_post::compute_rel_loop_scores*`) +/// The post-sim rel-loop-score denominator (`ltm_post::compute_rel_loop_scores`) /// sums `|loop_score|` in this order, so preserving emission order keeps that /// IEEE-754 (non-associative) sum bit-for-bit identical to the pre-#461 /// compile-time emitter, which accumulated in the same `detected_loops` order diff --git a/src/simlin-engine/src/db/analysis.rs b/src/simlin-engine/src/db/analysis.rs index db86e6987..935b861e4 100644 --- a/src/simlin-engine/src/db/analysis.rs +++ b/src/simlin-engine/src/db/analysis.rs @@ -2876,8 +2876,8 @@ pub fn reclassify_loops_from_results( // An A2A loop's loop_score occupies `n_slots` consecutive offsets; // a scalar/cross-element/mixed loop has exactly one. The slot count - // comes from the loop's partition vector (1 when absent), matching - // `ltm_post`'s `loop_n_slots`. + // comes from the loop's partition vector (1 when absent), the slot + // count `ltm_post::compute_rel_loop_scores` lays the loop out with. let n_slots = loop_partitions .get(&loop_item.id) .map(|p| p.len().max(1)) diff --git a/src/simlin-engine/src/db/ltm/mod.rs b/src/simlin-engine/src/db/ltm/mod.rs index d6a218884..b3cff01dd 100644 --- a/src/simlin-engine/src/db/ltm/mod.rs +++ b/src/simlin-engine/src/db/ltm/mod.rs @@ -1604,15 +1604,15 @@ pub fn model_ltm_variables( }; // Capture each loop's per-slot partition vector before consuming - // `partitions` so post-sim `compute_rel_loop_scores*` can group slots - // into the same `(partition, slot)` denominator bins. The vector's + // `partitions` so post-sim `compute_rel_loop_scores` can put each slot + // in its own partition's normalization group. The vector's // length must match the loop_score series' slot count -- 1 for a // scalar/cross-element/mixed loop, the dimension-element-space size // for an A2A loop -- which is the same `n_slots` that // `ltm_post::build_loop_element_index` derives from // `LtmSyntheticVar.dimensions` + the project dims; both feed - // `compute_rel_loop_scores_per_element`, so a length mismatch would - // desync the per-element normalization. + // `compute_rel_loop_scores`, so a length mismatch would desync the + // per-slot normalization. for l in detected_loops.iter() { let parts = partitions.partition_for_loop(l, dm_dims); debug_assert!( @@ -1632,7 +1632,7 @@ pub fn model_ltm_variables( }, "loop {:?}: per-slot partition vector length {} disagrees with the loop's slot \ count; it must equal `build_loop_element_index`'s n_slots (both feed \ - `compute_rel_loop_scores_per_element`)", + `compute_rel_loop_scores`)", l.id, parts.len(), ); diff --git a/src/simlin-engine/src/db/ltm_unified_tests.rs b/src/simlin-engine/src/db/ltm_unified_tests.rs index a5fc0fd49..d26cfbeb9 100644 --- a/src/simlin-engine/src/db/ltm_unified_tests.rs +++ b/src/simlin-engine/src/db/ltm_unified_tests.rs @@ -2356,8 +2356,8 @@ fn a2a_loop_partitions_have_one_entry_per_element() { // indices. Pre-#487 the A2A loop carried variable-level stocks // (`"pop"`) so `partition_for_loop` returned a single `None`; now it // returns `[Some(p0), Some(p1), Some(p2)]` in the runtime's row-major - // slot order -- so the rel-loop-score normalizer can keep the three - // per-element subsystems in separate `(partition, slot)` buckets. + // slot order -- so the rel-loop-score normalizer keeps each of the three + // per-element subsystems in its own partition's group. let project = TestProject::new("a2a_partition") .named_dimension("Region", &["NYC", "Boston", "LA"]) .array_stock("pop[Region]", "100", &["births"], &[], None) diff --git a/src/simlin-engine/src/db/ltm_value_gate_tests.rs b/src/simlin-engine/src/db/ltm_value_gate_tests.rs index 115a4c4fc..478dcc8d8 100644 --- a/src/simlin-engine/src/db/ltm_value_gate_tests.rs +++ b/src/simlin-engine/src/db/ltm_value_gate_tests.rs @@ -427,7 +427,7 @@ fn a_nested_freeze_arm_is_not_a_structural_zero() { /// comes from the guard form's own `NaN - NaN`, not from the modeller's /// equation -- and the arm has no causal dependence on the source at all, so /// `0` is the structurally known answer rather than a guess. -/// * GH #542 points the other way. `ltm_post::denom_summand` excludes a `NaN` +/// * GH #542 points the other way. `ltm_post::group_totals` excludes a `NaN` /// summand from its partition denominator specifically so that one undefined /// score does not poison its siblings, while the bad loop's OWN numerator /// stays `NaN` -- described there as "the honest per-loop 'undefined here' diff --git a/src/simlin-engine/src/layout/detect_ltm_loops.rs b/src/simlin-engine/src/layout/detect_ltm_loops.rs index 8cfcdbe04..41c3862d6 100644 --- a/src/simlin-engine/src/layout/detect_ltm_loops.rs +++ b/src/simlin-engine/src/layout/detect_ltm_loops.rs @@ -10,8 +10,6 @@ //! importance series and cycle partitions -- from the incremental salsa LTM //! pipeline, falling back to persisted `loop_metadata` when any step fails. -use std::collections::HashMap; - use crate::ltm_dominance::{FeedbackLoop, LoopPolarity}; /// Try to detect feedback loops using LTM analysis via the incremental @@ -61,60 +59,32 @@ fn try_detect_ltm_loops_incremental( Some(vm) }); - // Capture the loop_partitions mapping AND per-loop slot counts off the - // same `model_ltm_variables` derivation the VM's program was assembled - // from. Per-element rel scores need both the partition map (which loops - // normalize together) and the per-loop slot count (how many elements - // each A2A loop occupies). - let (loop_partitions, n_slots_by_loop) = if vm_result.is_some() { - let ltm_vars = crate::db::model_ltm_variables(db, source_model, source_project); - let dm_dims = crate::db::project_datamodel_dims(db, source_project); - let dim_size: HashMap<&str, usize> = dm_dims.iter().map(|d| (d.name(), d.len())).collect(); - let prefix = "$\u{205A}ltm\u{205A}loop_score\u{205A}"; - let n_slots: HashMap = ltm_vars - .vars - .iter() - .filter_map(|v| { - let id = v.name.strip_prefix(prefix)?; - let n = if v.dimensions.is_empty() { - 1 - } else { - v.dimensions - .iter() - .map(|d| dim_size.get(d.as_str()).copied().unwrap_or(1)) - .product() - }; - Some((id.to_string(), n)) - }) - .collect(); - (ltm_vars.loop_partitions.clone(), n_slots) + // Capture the per-slot loop_partitions mapping off the same + // `model_ltm_variables` derivation the VM's program was assembled from: + // it says which `(loop, slot)`s normalize together, and its per-loop + // vector length is the loop's slot count. + let loop_partitions = if vm_result.is_some() { + crate::db::model_ltm_variables(db, source_model, source_project) + .loop_partitions + .clone() } else { - (indexmap::IndexMap::new(), HashMap::new()) + indexmap::IndexMap::new() }; let vm = vm_result?; let results = vm.into_results(); - // `rel_loop_score` is no longer a VM variable; derive it post-sim from - // the `loop_score` series the VM does emit, using the per-slot partition - // mapping cached on `model_ltm_variables`. See - // `docs/design-plans/2026-04-18-ltm-cap-lift-diagnosis.md`. - // - // For arrayed (A2A) loops we compute per-element rel scores then - // aggregate to a single signed series via argmax-abs across slots -- - // i.e. each step's importance is the dominant element's contribution, - // with sign preserved. For scalar loops this reduces to identity. - // The aggregation is delegated to `ltm_post::aggregate_per_element_argmax_abs` - // so the partition-stride handling (mixed partitions where stride > - // per-loop n_slots) is centralized and unit-testable. See issue #463. - // `compute_rel_loop_scores_per_element` derives each loop's slot count - // from `loop_partitions[id].len()`, so no separate slot-count map is - // threaded; `aggregate_per_element_argmax_abs` still takes one. + // Relative loop scores are derived post-sim from the `loop_score` series + // the VM emits, by the one owner of the partition normalization + // (`ltm_post::compute_rel_loop_scores`). An arrayed (A2A) loop's + // per-slot series is collapsed to one signed importance series by + // argmax-abs across its slots -- each step's importance is the dominant + // element's contribution, sign preserved; a scalar loop's series is + // already one per step (issue #463). let per_element_rel_scores = - crate::ltm_post::compute_rel_loop_scores_per_element(&results, &loop_partitions); + crate::ltm_post::compute_rel_loop_scores(&results, &loop_partitions); let importance_by_loop = crate::ltm_post::aggregate_per_element_argmax_abs( &per_element_rel_scores, - &n_slots_by_loop, results.step_count, ); diff --git a/src/simlin-engine/src/ltm/types.rs b/src/simlin-engine/src/ltm/types.rs index 393c5fe0f..33353c2e3 100644 --- a/src/simlin-engine/src/ltm/types.rs +++ b/src/simlin-engine/src/ltm/types.rs @@ -101,9 +101,9 @@ pub struct Link { /// `partition_for_loop` looks up stocks in `model_element_cycle_partitions`, /// whose `stock_partition` map is element-keyed. Variable-level names here /// would cause `partition_for_loop` to return `None`, silently corrupting -/// per-loop / per-slot normalization in `ltm_post::compute_rel_loop_scores*` -/// (the loop would bucket into the catch-all `None` group instead of its -/// actual SCC). +/// the per-slot normalization in `ltm_post::compute_rel_loop_scores` (the +/// loop would become a Solo member instead of joining its actual SCC's +/// group). /// /// `assign_loop_ids` derives loop IDs from `links` (sorted distinct /// variable names), not `stocks`, so the element-level `stocks` granularity diff --git a/src/simlin-engine/src/ltm_augment_zero_slot.rs b/src/simlin-engine/src/ltm_augment_zero_slot.rs index 199860ddc..d5131f3dd 100644 --- a/src/simlin-engine/src/ltm_augment_zero_slot.rs +++ b/src/simlin-engine/src/ltm_augment_zero_slot.rs @@ -61,7 +61,7 @@ pub(crate) enum ZeroSlotPolicy { /// noise in a channel practitioners debug by hand, and this NaN is /// engine-made -- the guard form's own subtraction -- on an arm with no /// causal dependence on its source, so `0` is the structurally known answer. - /// GH #542 points the other way: `ltm_post::denom_summand` excludes a `NaN` + /// GH #542 points the other way: `ltm_post::group_totals` excludes a `NaN` /// score from its partition denominator precisely so the bad entry's own /// numerator can stay `NaN` as "the honest per-loop 'undefined here' /// signal". Replacing some of those with `0` partially undoes that. diff --git a/src/simlin-engine/src/ltm_finding.rs b/src/simlin-engine/src/ltm_finding.rs index 481b772b1..5d0408422 100644 --- a/src/simlin-engine/src/ltm_finding.rs +++ b/src/simlin-engine/src/ltm_finding.rs @@ -37,7 +37,7 @@ use crate::common::{Canonical, Ident, Result}; use crate::datamodel; use crate::db::LtmSyntheticVar; use crate::ltm::{CausalGraph, CyclePartitions, Link, LinkPolarity, Loop, LoopPolarity}; -use crate::ltm_post::NormGroup; +use crate::ltm_post::{NormGroup, group_totals, relative_series}; use crate::results::Results; // Union-graph circuit enumeration: discovery's primary candidate generator @@ -3521,8 +3521,8 @@ fn loop_partition_slot(fl: &FoundLoop, partitions: &CyclePartitions) -> Option>, /// Per-partition count of circuits in the enumerated universe that @@ -3644,6 +3644,9 @@ fn subtract_reported_mass_from_totals( /// no active steps returns `NaN` (it sorts last). `Inf/Inf = NaN` at a /// dominance inflection is naturally excluded since `NaN` is not active. fn mean_relative_contribution(fl: &FoundLoop, totals: &[f64]) -> f64 { + // The same division every relative series gets (`ltm_post::relative_series`, + // SAFEDIV-0); the active-step mask below is what this statistic adds. + let relative = relative_series(fl.scores.iter().map(|&(_, score)| score), totals); let mut sum = 0.0; let mut active = 0usize; for (i, &(_, score)) in fl.scores.iter().enumerate() { @@ -3661,10 +3664,10 @@ fn mean_relative_contribution(fl: &FoundLoop, totals: &[f64]) -> f64 { if inactive { continue; } - let rel = score.abs() / total; - // rel is in [0, 1] for a finite score (the loop's own |score| is part of - // total). An Inf score makes total Inf, so rel == Inf/Inf == NaN and the - // step drops out here; guard against any residual NaN to be safe. + // |rel| is in [0, 1] for a finite score (the loop's own |score| is part + // of total). An Inf score makes total Inf, so rel == Inf/Inf == NaN and + // the step drops out here; guard against any residual NaN to be safe. + let rel = relative[i].abs(); if rel.is_nan() { continue; } @@ -3813,8 +3816,8 @@ fn cmp_relative_importance(a: &RelativeImportance, b: &RelativeImportance) -> st /// cross-partition equivalence is by design -- the relative key measures /// in-partition dominance, not how long the partition itself stays active. /// -/// **NaN/Inf handling** mirrors `ltm_post.rs::denom_summand` (GH #542) so the -/// two LTM paths agree: a `NaN` `score[t]` contributes nothing to a partition +/// **NaN/Inf handling** is the shared accumulator's (`ltm_post::group_totals`, +/// GH #542), so the two LTM paths agree: a `NaN` `score[t]` contributes nothing to a partition /// total and that step is skipped in the loop's own mean; an `Inf` `score[t]` /// stays in the partition total (a real dominance-inflection signal), so the /// loop's own `Inf/Inf = NaN` step is skipped and dominated siblings see a @@ -3885,7 +3888,7 @@ fn rank_and_filter( let loop_groups: Vec = found_loops .iter() .enumerate() - .map(|(i, fl)| NormGroup::for_loop(slot0(fl), i)) + .map(|(i, fl)| NormGroup::for_member(slot0(fl), i)) .collect(); // Group loops by normalization group over the FULL discovered set @@ -3928,44 +3931,34 @@ fn rank_and_filter( .collect(); // Per-group per-timestep totals: Σ|score_j[t]| over the group's loops, - // NaN excluded (an undefined score is not signal; matches GH #542's - // denom_summand). Inf is kept -- a real divergence at a dominance - // inflection. Computed over ALL discovered loops, before any cap, so the - // denominator reflects the whole partition (the truncate-before-filter - // order of GH #310 used to compute totals over only the top-200 - // survivors). A Solo-group total is just that loop's own |score| series. - // On the enumeration path, a Partition group's totals instead come from - // `external_totals` -- the full enumerated universe's mass (see the fn - // doc above). + // accumulated by the shared owner (`ltm_post::group_totals`: NaN + // excluded, Inf kept, finite overflow saturating -- the same rule the + // exhaustive series are normalized under, so the two surfaces agree). + // Computed over ALL discovered loops, before any cap, so the + // denominator reflects the whole partition (a truncate-before-filter + // order would total only the survivors, GH #310). A Solo-group total + // is just that loop's own |score| series. On the enumeration path, a + // Partition group's totals instead come from the universe -- the full + // enumerated population's mass (see the fn doc above). let mut partition_totals: HashMap> = HashMap::new(); if step_count > 0 { - for (&group, indices) in &partition_groups { + for &group in partition_groups.keys() { if let (NormGroup::Partition(p), Some(stats)) = (group, universe) && let Some(totals) = stats.totals.get(&p) { debug_assert_eq!(totals.len(), step_count); partition_totals.insert(group, totals.clone()); - continue; - } - let mut totals = vec![0.0; step_count]; - for &idx in indices { - for (i, &(_, score)) in found_loops[idx].scores.iter().enumerate() { - if !score.is_nan() { - // Saturating, as retention's own bank is: a finite - // sum overflowing to Inf would zero every finite share. - let mass = score.abs(); - let sum = totals[i] + mass; - totals[i] = - if sum.is_infinite() && mass.is_finite() && totals[i].is_finite() { - f64::MAX - } else { - sum - }; - } - } } - partition_totals.insert(group, totals); } + let discovered_totals = group_totals( + loop_groups + .iter() + .zip(found_loops.iter()) + .filter(|(group, _)| !partition_totals.contains_key(group)) + .map(|(&group, fl)| (group, fl.scores.iter().map(|&(_, score)| score))), + step_count, + ); + partition_totals.extend(discovered_totals); } // Partition-aware MIN_CONTRIBUTION retention filter (peak semantics, @@ -4085,26 +4078,6 @@ fn attach_partition_metadata( meta } -/// The SIGNED per-timestep partition-relative loop score series for one loop. -/// -/// `rel[t] = score[t] / totals[t]`, with `totals[t]` the loop's cycle-partition -/// denominator (`Σ_{j in partition} |score_j[t]|`, NaN summands already -/// excluded by `rank_and_filter`). SAFEDIV-0 (`totals[t] == 0` -> `0.0`) and a -/// `NaN` numerator propagating to `NaN` both match -/// `ltm_post::compute_rel_loop_scores` exactly, so the discovery and pinned-loop -/// relative-score surfaces agree. Sign is preserved (a balancing loop reads -/// negative), giving a value in `[-1, 1]` for a finite score. -fn signed_relative_scores(fl: &FoundLoop, totals: &[f64]) -> Vec { - fl.scores - .iter() - .enumerate() - .map(|(t, &(_, score))| { - let total = totals.get(t).copied().unwrap_or(0.0); - if total == 0.0 { 0.0 } else { score / total } - }) - .collect() -} - /// The deepest per-step rank a loop can hold and still be anchored (AC5.1). /// /// `k = 1` is the guarantee: every step's dominant loop in a competing group @@ -4136,7 +4109,7 @@ const ANCHOR_SHARE_OF_CAP: f64 = 0.5; /// One ranked loop, as the coverage-aware cap sees it. /// /// `rel` is the loop's signed per-step relative score series -/// ([`signed_relative_scores`]); the selection reads magnitudes and treats a +/// (`FoundLoop::rel_scores`); the selection reads magnitudes and treats a /// `NaN` as absent (a `NaN` score is an undefined contribution, exactly as in /// [`mean_relative_contribution`]). struct SelectionRow<'a> { @@ -4318,12 +4291,13 @@ fn rank_truncate_and_id( // permutation and the key survives ID assignment (no recomputation). // // While the per-partition denominators are in hand, also attach each loop's - // SIGNED per-timestep relative score series (`rel_scores`) -- the same - // `score[t] / partition_total[t]` normalization, SAFEDIV-0, that - // `ltm_post::compute_rel_loop_scores` applies on the pinned-loop path. This - // is the [-1, 1] importance series `analysis::to_loop_summary` / - // `to_feedback_loop` surface, so dominance/ranking is partition-relative - // (comparable across partitions) rather than raw-magnitude-biased. + // SIGNED per-timestep relative score series (`rel_scores`) through the + // shared owner (`ltm_post::relative_series`: `score[t] / + // partition_total[t]`, SAFEDIV-0), the same division the exhaustive + // series get in `ltm_post::compute_rel_loop_scores`. This is the [-1, 1] + // importance series `analysis::to_loop_summary` / `to_feedback_loop` + // surface, so dominance/ranking is partition-relative (comparable across + // partitions) rather than raw-magnitude-biased. let mut keyed: Vec<(RelativeImportance, FoundLoop)> = std::mem::take(found_loops) .into_iter() .enumerate() @@ -4331,7 +4305,7 @@ fn rank_truncate_and_id( let totals = &partition_totals[&loop_groups[idx]]; let mean_rel = mean_relative_contribution(&fl, totals); let key = loop_sort_key(&fl.loop_info); - fl.rel_scores = signed_relative_scores(&fl, totals); + fl.rel_scores = relative_series(fl.scores.iter().map(|&(_, score)| score), totals); ( RelativeImportance { mean_rel, diff --git a/src/simlin-engine/src/ltm_finding_enum.rs b/src/simlin-engine/src/ltm_finding_enum.rs index 74ae08bd7..6c0d9f979 100644 --- a/src/simlin-engine/src/ltm_finding_enum.rs +++ b/src/simlin-engine/src/ltm_finding_enum.rs @@ -1330,17 +1330,13 @@ pub(super) fn retain_circuits( if mass != 0.0 { banked_mass = true; } - // Saturating: a sum of FINITE masses that would overflow to Inf - // would make every finite share read as 0 and drop a real loop - // universe wholesale; capping the total at f64::MAX keeps shares - // finite (and merely compressed) there. A genuinely infinite mass - // still makes the total Inf, the dominance-inflection convention. - let sum = totals[t] + mass; - totals[t] = if sum.is_infinite() && mass.is_finite() && totals[t].is_finite() { - f64::MAX - } else { - sum - }; + // The one accumulator every partition total goes through + // (`ltm_post::add_to_total`): a sum of FINITE masses that would + // overflow saturates at f64::MAX so every finite share stays + // finite (merely compressed) instead of reading 0 and dropping a + // real loop universe wholesale; a genuinely infinite mass still + // makes the total Inf, the dominance-inflection convention. + crate::ltm_post::add_to_total(&mut totals[t], scratch[t]); if totals[t] > 0.0 { // `max` drops a NaN ratio (`Inf / Inf` at a dominance // inflection), which is right: the exact test rejects that step diff --git a/src/simlin-engine/src/ltm_post.rs b/src/simlin-engine/src/ltm_post.rs index ad0590406..263669d59 100644 --- a/src/simlin-engine/src/ltm_post.rs +++ b/src/simlin-engine/src/ltm_post.rs @@ -2,96 +2,172 @@ // Use of this source code is governed by the Apache License, // Version 2.0, that can be found in the LICENSE file. -//! Post-simulation computation of LTM relative loop scores. +//! Post-simulation relative loop scores: the one owner of the cycle-partition +//! normalization every surface reads. //! -//! Historical context: exhaustive LTM used to emit a synthetic -//! `$⁚ltm⁚rel_loop_score⁚{id}` variable for every loop whose equation -//! normalized that loop's `loop_score` against the partition sum of -//! `|loop_score_j|`. Emission was O(P²) text per partition (see -//! `docs/design-plans/2026-04-18-ltm-cap-lift-diagnosis.md`) and -//! dominated compile memory for dense models. Option B of the cap-lift -//! design plan moves the normalization here, executed post-simulation -//! against the O(P × save_steps) `loop_score` timeseries that the VM -//! already writes to `Results`. - -use std::collections::{BTreeMap, BTreeSet, HashMap}; +//! A loop's raw `loop_score` is the product of its link scores; its +//! *relative* score at a step is that score divided by the sum of `|score|` +//! over every loop of the same cycle partition (reference section 4.4). +//! Exhaustive mode records the raw series in [`Results`] (one column per +//! slot of the `$⁚ltm⁚loop_score⁚{id}` synthetic); discovery holds them on +//! its `FoundLoop`s. Both feed [`group_totals`] and [`relative_series`], so +//! the two surfaces cannot disagree on what "relative" means, and every +//! reader of the exhaustive series -- the libsimlin FFI, the layout's +//! importance series, the tests -- goes through [`compute_rel_loop_scores`] +//! rather than dividing on its own. +//! +//! Normalization is per partition and nothing finer. Every `(loop, slot)` +//! whose stocks resolve to partition `p` is a member of `p`'s group: a scalar +//! loop is one member, an arrayed loop one member per element slot. This is +//! the de-subscripted reading the reference's section 15.4 promises -- an +//! arrayed loop's element `k` competes with its sibling elements and with +//! every scalar loop of the partition exactly as the N scalar loops of the +//! hand-expanded model would. Never key the group on the slot index as +//! well: a slot index is a position in one loop's own dimension space, so +//! two arrayed loops over different dimension lists (or a scalar loop, which +//! has no elements) attach no shared meaning to "slot k", and grouping by it +//! splits one partition into denominators that each miss most of its +//! members: on `test/cross_element_ltm` a per-slot key reads the births +//! loop's share as 1/2 at one element and 2/3 at the other, where the +//! partition holds four active members and the share is 1/3 at both. +//! +//! The relative scores are computed here rather than emitted as synthetic +//! VM variables: a `rel_loop_score` equation names every sibling's +//! `loop_score`, O(P²) text per partition, which dominated compile memory on +//! dense models (`docs/design-plans/2026-04-18-ltm-cap-lift-diagnosis.md`). +//! Post-simulation the cost is O(P × save_steps) over series the VM writes +//! anyway. + +use std::collections::HashMap; +use std::hash::Hash; use indexmap::IndexMap; use crate::common::{Canonical, Ident}; use crate::results::Results; -/// A `(group, slot)` bucket key for the per-element normalization grid. -type BucketKey = (NormGroup, usize); -/// A `(loop_index, read_slot)` pair: which loop contributes to a bucket and -/// which of its own `loop_score` slots is read (0 for a broadcast scalar -/// loop, the bucket's slot for an arrayed loop). -type BucketMember = (usize, usize); - -/// The normalization group a loop's slot belongs to (GH #750). +/// The normalization group one member of a relative-score computation +/// belongs to (GH #750). /// -/// Loops whose cycle partition resolves share a [`NormGroup::Partition`] -/// group and normalize against each other. A slot whose partition is -/// unresolved (`None` -- a loop genuinely below the parent stock graph: -/// module-internal-stock loops, PREVIOUS-lagged stockless loops) gets a -/// [`NormGroup::Solo`] group keyed by the loop's emission-order index, so -/// two UNRELATED unpartitioned loops can never share a denominator -- the -/// GH #487-class cross-pollution the old shared default `None` bucket -/// reintroduced. A solo loop's relative score collapses to the documented -/// lone-pin degeneracy: sign-preserving `+/-1` when active, `0` via -/// SAFEDIV-0 when not. (Loops genuinely coupled through module-internal -/// state are under-merged by this rule -- each normalizes alone instead of -/// against its true siblings; resolving partitions for state-carrying -/// module instances is the tracked refinement that would move those loops -/// out of the `None` bucket entirely.) +/// A member whose cycle partition resolves is in that +/// [`NormGroup::Partition`] and normalizes against every other member of +/// it. A member whose partition is unresolved (`None` -- a loop genuinely +/// below the parent stock graph: module-internal-stock loops, +/// PREVIOUS-lagged stockless loops) gets a [`NormGroup::Solo`] group of its +/// own, so two UNRELATED unpartitioned members can never share a +/// denominator -- the GH #487-class cross-pollution a shared default `None` +/// bucket reintroduces. A solo member's relative score collapses to the +/// lone-pin degeneracy documented on [`compute_rel_loop_scores`]: +/// sign-preserving `+/-1` when active, `0` via SAFEDIV-0 when not. Loops +/// genuinely coupled through module-internal state are under-merged by this +/// rule -- each normalizes alone instead of against its true siblings; +/// resolving partitions for state-carrying module instances is the +/// refinement that would move those loops out of the `None` bucket. #[derive(Clone, Copy, PartialEq, Eq, PartialOrd, Ord, Hash, Debug)] pub(crate) enum NormGroup { /// A resolved cycle partition (the engine-internal partition index). Partition(usize), - /// An unresolved-partition loop, alone in its own group. The payload is - /// the loop's index in the caller's emission-order loop list, which is - /// unique per loop -- the value itself is never interpreted. + /// An unresolved-partition member, alone in its own group. The payload + /// is the member's index in the caller's member list -- a `(loop, slot)` + /// in exhaustive mode, a discovered loop in discovery -- which is unique + /// per member; the value itself is never interpreted. Solo(usize), } impl NormGroup { - /// Build the group for loop `loop_idx`'s partition value. - pub(crate) fn for_loop(partition: Option, loop_idx: usize) -> NormGroup { + /// The group of the member at `member_idx` whose cycle partition is + /// `partition`. + pub(crate) fn for_member(partition: Option, member_idx: usize) -> NormGroup { match partition { Some(p) => NormGroup::Partition(p), - None => NormGroup::Solo(loop_idx), + None => NormGroup::Solo(member_idx), } } } -/// One loop's `|loop_score|` contribution to a partition-sum denominator, -/// with `NaN` summands excluded. +/// Add one member's `|score|` at a step to its group's running total. /// -/// A `NaN` loop_score at a step means that one loop's score is *undefined* -/// there; it is not signal, so it must not flow into the partition sum -- -/// otherwise a single bad loop turns the whole partition's denominator into -/// `NaN` (the `denom == 0.0` SAFEDIV guard does not fire on `NaN`, since -/// `NaN == 0.0` is false), poisoning every sibling's relative score (GH -/// #542). Dropping the `NaN` summand lets healthy siblings normalize -/// against the healthy denominator; the bad loop's *own* numerator stays -/// `NaN`, so as long as a healthy sibling keeps the denominator non-zero -/// its own relative score stays `NaN` -- the honest per-loop "undefined -/// here" signal. (When the `NaN` loop is the partition's only contributor, -/// or every member is `NaN` at a step, excluding it collapses the -/// denominator to `0.0` and the SAFEDIV-0 guard yields `0.0` instead -- the -/// lone-pin degeneracy documented on `compute_rel_loop_scores`.) This -/// matches the discovery path, where a `NaN` link score marks its edge -/// inactive and contributes nothing. +/// A `NaN` score means that member's score is *undefined* at the step; it +/// is not signal, so it contributes nothing -- otherwise a single bad member +/// turns the whole group's denominator into `NaN` (the `total == 0.0` +/// SAFEDIV guard does not fire on `NaN`), poisoning every sibling's relative +/// score (GH #542). The bad member's own numerator stays `NaN`, so as long +/// as a healthy sibling keeps the total non-zero its own relative score +/// stays `NaN` -- the honest per-member "undefined here" signal. (When the +/// `NaN` member is the group's only contributor the total stays `0.0` and +/// SAFEDIV-0 yields `0.0` instead.) This is also the discovery rule: a +/// `NaN` link score marks its edge inactive and contributes nothing. /// -/// `+/-Inf` is deliberately NOT excluded: a raw loop score legitimately -/// diverges at a dominance inflection (the link-score denominators go to -/// zero there), so an `Inf` summand is real signal that the loop dominates. -/// Keeping it in the sum sends the dominated siblings to `0` (`finite/Inf`) -/// and the dominant loop to `NaN` (`Inf/Inf`) -- the same inflection-point -/// behaviour the removed SAFEDIV equation produced, preserved bug-for-bug. +/// `+/-Inf` is deliberately kept: a raw loop score legitimately diverges at +/// a dominance inflection (the link-score denominators go to zero there), +/// so an `Inf` summand is real signal that the member dominates. Keeping +/// it sends the dominated siblings to `0` (`finite/Inf`) and the dominant +/// member to `NaN` (`Inf/Inf`). A FINITE sum that overflows saturates to +/// `f64::MAX` instead of becoming `Inf`: every summand was finite, so an +/// infinite total would zero every finite share for no member's benefit. #[inline] -fn denom_summand(v: f64) -> f64 { - if v.is_nan() { 0.0 } else { v.abs() } +pub(crate) fn add_to_total(total: &mut f64, score: f64) { + if score.is_nan() { + return; + } + let mass = score.abs(); + let sum = *total + mass; + *total = if sum.is_infinite() && mass.is_finite() && total.is_finite() { + f64::MAX + } else { + sum + }; +} + +/// Per-group, per-step `Σ|score|` over `members`: the denominator of every +/// relative score. +/// +/// `members` yields `(group, scores)` pairs; each member's `|score|` is +/// added to its group's total at every step through `add_to_total` (`NaN` +/// excluded, `Inf` kept, finite overflow saturating). Members are +/// accumulated in the order given, and IEEE-754 addition is not +/// associative, so callers pass them in a deterministic order -- the +/// exhaustive path's loop emission order (GH #468), the discovery path's +/// candidate order. A series shorter than `step_count` contributes nothing +/// past its end. +/// +/// The key type is generic so the loop-level normalization (keyed by +/// [`NormGroup`]) and the link-level one (keyed by the link's target) share +/// this single accumulator rather than each restating the summand rule. +pub fn group_totals(members: I, step_count: usize) -> HashMap> +where + K: Hash + Eq, + I: IntoIterator, + S: IntoIterator, +{ + let mut totals: HashMap> = HashMap::new(); + for (group, scores) in members { + let total = totals + .entry(group) + .or_insert_with(|| vec![0.0_f64; step_count]); + for (t, score) in scores.into_iter().take(step_count).enumerate() { + add_to_total(&mut total[t], score); + } + } + totals +} + +/// `score[t] / totals[t]` at every step, with SAFEDIV-0 semantics: a `0` +/// total yields `0` rather than `NaN`. A `NaN` numerator propagates (the +/// member is undefined at that step); `Inf/Inf` is `NaN` at a dominance +/// inflection. A step past the end of `totals` reads a `0` total. +pub fn relative_series(scores: S, totals: &[f64]) -> Vec +where + S: IntoIterator, +{ + scores + .into_iter() + .enumerate() + .map(|(t, score)| { + let total = totals.get(t).copied().unwrap_or(0.0); + if total == 0.0 { 0.0 } else { score / total } + }) + .collect() } /// Build the canonical identifier of a loop's `loop_score` synthetic variable. @@ -105,368 +181,166 @@ pub(crate) fn loop_score_ident(loop_id: &str) -> Ident { Ident::new(&name) } -/// Slot count of a loop's `loop_score` series from its per-slot partition -/// vector: 1 for a scalar / cross-element / mixed loop, the dimension -/// element-space size for an A2A loop. Mirrors -/// `ltm_post::build_loop_element_index`'s `n_slots`, since both are derived -/// from the same `LtmSyntheticVar` metadata in `model_ltm_variables`. -fn loop_n_slots(loop_partitions: &IndexMap>>, id: &str) -> usize { - loop_partitions.get(id).map(|v| v.len()).unwrap_or(1).max(1) -} - -/// The partition (`Option`) of loop `id` at slot `k`. -/// -/// For an arrayed loop this is `loop_partitions[id][k]`; for a scalar loop -/// (`n_slots == 1`) it is `loop_partitions[id][0]` broadcast across every -/// `k` -- a scalar loop has no elements, so it carries its single partition -/// into every slot it is compared in (the same broadcast the pre-PR -/// compile-time emitter applied when a scalar `loop_score` was referenced -/// from an arrayed `rel_loop_score` equation). `None` (out-of-range `k` on -/// an arrayed loop, or a genuinely-`None` partition) means "no contribution -/// at this slot for the purpose of *that loop's own* series", though it -/// still buckets into the `None` cohort. -fn slot_partition( - loop_partitions: &IndexMap>>, - id: &str, - k: usize, -) -> Option { - let v = loop_partitions.get(id)?; - if v.len() <= 1 { - v.first().copied().flatten() - } else { - v.get(k).copied().flatten() - } -} - -/// Compute per-loop, per-timestep relative loop scores from simulated -/// `loop_score` data -- the **slot-0 convenience view**. -/// -/// For each loop whose `loop_score` series is present in `results`, the -/// returned value is: -/// -/// ```text -/// rel_loop_score[i, t] = loop_score[i, t, 0] / Σ_{j : slot-0 partition of j == slot-0 partition of i} |loop_score[j, t, 0]| -/// ``` +/// Per-loop, per-slot, per-step relative loop scores from an exhaustive-mode +/// [`Results`] -- the owner every reader of those scores goes through. /// -/// `loop_partitions` maps each loop ID to its **per-slot** cycle-partition -/// vector (as produced by `model_ltm_variables`; length 1 for a -/// scalar/cross-element/mixed loop, one entry per element for an A2A loop). -/// This function reports only slot 0 for every loop and groups loops by -/// their *slot-0* partition (`loop_partitions[id][0]`). A loop whose -/// partition is unresolved (`None` -- genuinely below the parent stock -/// graph: module-internal-stock loops, PREVIOUS-lagged stockless loops) -/// normalizes ALONE (GH #750, [`NormGroup::Solo`]): unrelated subsystems -/// must not cross-normalize, so its relative score collapses to the lone-pin -/// degeneracy below. This preserves the pre-Phase-2 scalar contract (one -/// series per loop), so existing libsimlin/pysimlin/TS callers see no shape -/// change; callers that want genuine per-element normalization use -/// [`compute_rel_loop_scores_per_element`]. +/// `loop_partitions` maps each loop id to its **per-slot** cycle-partition +/// vector as `model_ltm_variables` produces it: length 1 for a scalar, +/// cross-element or mixed loop, one entry per element for an A2A loop, in +/// the row-major slot order the `loop_score` columns use. Every `(loop, +/// slot)` with a `loop_score` column is one member of the group +/// [`NormGroup::for_member`] assigns to its slot's partition, and its +/// relative score is [`relative_series`] against that group's +/// [`group_totals`]. A scalar loop therefore has exactly one series and an +/// arrayed loop one per slot, each normalized against every member of the +/// slot's own partition -- sibling slots of the same loop, other arrayed +/// loops' slots, and scalar loops alike. /// -/// The denominator uses SAFEDIV-0 semantics: when the partition sum at -/// slot 0 is `0` the result is `0` rather than `NaN`. +/// The returned series for loop `id` is flat and step-major: the value at +/// step `s`, slot `k` is `series[s * n_slots + k]` with `n_slots == +/// loop_partitions[id].len()` (1 for a scalar loop, so the series is just +/// `step_count` long). Loops whose `loop_score` is absent from `results` +/// (LTM disabled for that loop, or a discovery-mode compilation) are +/// omitted. /// -/// Non-finite `loop_score` handling (GH #542): a `NaN` summand is -/// *excluded* from the partition sum via [`denom_summand`], so one loop -/// whose score is undefined at a step no longer poisons every sibling's -/// relative score in that partition. The bad loop's own numerator is -/// still `NaN`, so when a healthy sibling keeps the denominator non-zero -/// its own relative score stays `NaN` -- the honest per-loop "undefined -/// here" signal -- while the healthy siblings normalize against the -/// healthy denominator. (When the `NaN` loop is the partition's only -/// contributor, or every member is `NaN` at the step, excluding it leaves -/// a `0.0` denominator and the SAFEDIV-0 guard yields `0.0` instead -- see -/// the lone-pin degeneracy section below.) This matches the discovery -/// path's "a `NaN` link contributes nothing" rule (a `NaN` score marks its -/// edge inactive there). `+/-Inf` is deliberately -/// *kept* in the sum: a raw loop score legitimately diverges at a -/// dominance inflection, so `Inf` is real signal -- it sends the -/// dominated siblings to `0` and the dominant loop to `NaN` (`Inf/Inf`), -/// the same inflection-point behaviour the removed SAFEDIV equation -/// produced. Earlier this whole module propagated `NaN` through normal -/// IEEE-754 arithmetic only because the post-simulation refactor -/// preserved the removed synthetic equation's semantics bug-for-bug; the -/// removed SAFEDIV never promised `NaN`-resilience either. -/// -/// Loops whose `loop_score` is absent from `results` (e.g. LTM disabled -/// for that loop, or discovery-mode compilation) are omitted from the -/// returned map. +/// Members are accumulated in `loop_partitions`' insertion order -- loop +/// emission order, slots ascending within a loop -- NOT a re-sort. The +/// per-partition sum is bit-significant (IEEE-754 addition is +/// non-associative) and emission order is the content-derived order +/// `assign_loop_ids` produces, deterministic across salsa cache +/// invalidations and processes (GH #468); a bare lex sort would order a +/// partition as `b1, b10, b2, ...` and perturb the sum at the ULP. /// /// # Lone-pin degeneracy /// -/// A modeler-pinned loop (`pin{n}` id) is registered in its own single-slot -/// `loop_partitions` entry. When a pin is the *only* loop in its partition -- -/// always the case in discovery mode (no enumerated loop scores exist there) -/// and in exhaustive mode whenever the pin is the lone loop through its stock -/// -- the partition sum equals `|loop_score[pin]|`, so the relative score -/// collapses to exactly `+1` (or `-1`, carrying the raw score's sign) whenever -/// the loop is active and `0` (via SAFEDIV-0) when its raw score is `0`. This -/// is intentional: a fraction-of-all-known-loops normalization is undefined -/// for a partition of one. Callers that want a pinned loop's actual magnitude -/// should read its **raw** `loop_score` series directly. Two or more pins (or -/// a pin plus enumerated loops) on stocks in the *same* SCC partition normalize +/// A modeler-pinned loop (`pin{n}` id) that is the *only* member of its +/// group -- always the case in discovery mode (no enumerated loop scores +/// exist there) and in exhaustive mode whenever the pin is the lone loop +/// through its stock -- has a total equal to its own `|loop_score|`, so its +/// relative score collapses to exactly `+1` (or `-1`, carrying the raw +/// score's sign) whenever the loop is active and `0` (via SAFEDIV-0) when +/// its raw score is `0`. This is intentional: a fraction-of-all-known-loops +/// normalization is undefined for a group of one. Callers that want a +/// pinned loop's actual magnitude read its **raw** `loop_score` series +/// ([`compute_raw_loop_score_for_element`]). Two or more pins (or a pin +/// plus enumerated loops) on stocks in the *same* partition normalize /// against each other normally. pub fn compute_rel_loop_scores( results: &Results, loop_partitions: &IndexMap>>, ) -> HashMap> { - // Iterate in `loop_partitions`' *emission* order (its `IndexMap` insertion - // order), NOT a re-sort. The partition-sum denominator below accumulates - // `|loop_score|` over the loops in a partition in exactly this order, and - // IEEE-754 addition is non-associative, so the order is bit-significant. - // Emission order is the content-derived order `assign_loop_ids` produced - // (see `ltm::graph::loop_id_sort_key`), which the pre-#461 compile-time - // emitter also accumulated in -- so iterating it here restores bit-for-bit - // parity with that removed emitter (GH #468). A bare lex sort would order - // the same partition as `b1, b10, b2, ...`, perturbing the sum at the ULP. - // Emission order is itself deterministic across salsa cache invalidations - // and across processes (`assign_loop_ids` is a pure function of loop - // content, not of `HashMap` enumeration), so this is just as deterministic - // as the old lex sort while also being bit-parity-correct. - let loop_ids: Vec<&String> = loop_partitions.keys().collect(); - - let offsets: Vec> = loop_ids - .iter() - .map(|id| results.offsets.get(&loop_score_ident(id)).copied()) - .collect(); - - // Group loops by their slot-0 partition (the convenience-view key). An - // unresolved (`None`) partition gets a per-loop Solo group (GH #750): - // unpartitioned loops have no provable coupling, so they never share a - // denominator. - let mut partition_groups: HashMap> = HashMap::new(); - for (i, id) in loop_ids.iter().enumerate() { - partition_groups - .entry(NormGroup::for_loop( - slot_partition(loop_partitions, id, 0), - i, - )) - .or_default() - .push(i); - } - - // One output series per loop, parallel to `loop_ids`. Loops without - // a known offset get an empty Vec so we can skip them when - // assembling the final map. - let mut series: Vec> = offsets - .iter() - .map(|o| { - if o.is_some() { - Vec::with_capacity(results.step_count) - } else { - Vec::new() - } - }) - .collect(); - - for row in results.iter() { - for indices in partition_groups.values() { - let denom: f64 = indices - .iter() - .filter_map(|&i| offsets[i].map(|off| denom_summand(row[off]))) - .sum(); - - for &i in indices { - let Some(off) = offsets[i] else { continue }; - let num = row[off]; - let val = if denom == 0.0 { 0.0 } else { num / denom }; - series[i].push(val); - } + /// One `(loop, slot)` member: its group and its `Results` column. + struct Member { + group: NormGroup, + column: usize, + } + // Loops with a `loop_score` column, in emission order, as + // `(id, n_slots)`; their members are contiguous in `members`, in the + // same order. + let mut loops: Vec<(&String, usize)> = Vec::with_capacity(loop_partitions.len()); + let mut members: Vec = Vec::new(); + for (id, partitions) in loop_partitions { + let Some(&base) = results.offsets.get(&loop_score_ident(id)) else { + continue; + }; + let n_slots = partitions.len().max(1); + debug_assert!( + base + n_slots <= results.step_size, + "loop {id}: {n_slots} slots at column {base} exceed the {} columns per step", + results.step_size + ); + loops.push((id, n_slots)); + for k in 0..n_slots { + let partition = partitions.get(k).copied().flatten(); + members.push(Member { + group: NormGroup::for_member(partition, members.len()), + column: base + k, + }); } } - let mut out: HashMap> = HashMap::with_capacity(loop_ids.len()); - for (i, id) in loop_ids.iter().enumerate() { - if offsets[i].is_some() { - out.insert((*id).clone(), std::mem::take(&mut series[i])); + let column = |c: usize| results.iter().map(move |row| row[c]); + let totals = group_totals( + members.iter().map(|m| (m.group, column(m.column))), + results.step_count, + ); + + let mut out: HashMap> = HashMap::with_capacity(loops.len()); + let mut first_member = 0usize; + for (id, n_slots) in loops { + let mut series = vec![0.0_f64; results.step_count * n_slots]; + for (k, member) in members[first_member..first_member + n_slots] + .iter() + .enumerate() + { + let rel = relative_series(column(member.column), &totals[&member.group]); + for (t, v) in rel.into_iter().enumerate() { + series[t * n_slots + k] = v; + } } + first_member += n_slots; + out.insert(id.clone(), series); } out } -/// Per-timestep, per-slot relative loop scores, grouped by -/// `(partition, slot)`. -/// -/// [`compute_rel_loop_scores`] collapses every loop's `loop_score` to -/// slot 0. This function keeps every slot, and -- crucially -- groups -/// slots by `(NormGroup, k)` (the slot's partition, or the loop's own -/// [`NormGroup::Solo`] group when unresolved -- GH #750) rather than by a -/// single per-loop partition. So an A2A loop over an element-wise-coupled -/// dimension (every slot in partition `p`) lands in buckets `(p, 0)`, -/// `(p, 1)`, ...; an A2A loop over an element-wise-uncoupled dimension -/// spreads across `(p0, 0)`, `(p1, 1)`, ... -- which is precisely why two -/// disconnected per-element feedback subsystems over the same dimension -/// stop cross-normalizing (GH #487); and an unresolved (`None`) slot -/// normalizes against that loop alone, so unrelated -/// module-internal-stock / lagged-stockless loops do not cross-normalize -/// either (the same #487-class pollution, in the old shared `None` -/// bucket). -/// -/// `loop_partitions` is the per-slot partition map from -/// `model_ltm_variables`; the loop's slot count is `loop_partitions[id].len()` -/// (no separate slot-count map is threaded). Returns a flat `Vec` per -/// loop id; the value at step `s`, slot `k` is at index `s * stride + k`, -/// where `stride` is the loop's own slot count for an arrayed loop, and for -/// a scalar loop the largest slot index its (slot-0) partition covers + 1 -/// (1 if no arrayed loop shares that partition). A scalar loop broadcasts -/// its single value into every slot of its partition's buckets -- the same -/// broadcast the pre-PR compile-time emitter applied when a scalar -/// `loop_score` was referenced from an arrayed `rel_loop_score` equation. -/// -/// Denominator at bucket `(p, k)` at step `s` is `Σ |loop_score[j, s, rs_j]|` -/// over the members of that bucket, where `rs_j` is `0` for a scalar member -/// (broadcast) and `k` for an arrayed member with `k < n_slots[j]` (an -/// arrayed loop with `k >= n_slots[j]` is not a member of slot-`k` buckets). -/// SAFEDIV-0 semantics, per-bucket `NaN` exclusion, and `Inf` retention -/// all match [`compute_rel_loop_scores`] (via [`denom_summand`]): a `NaN` -/// at one `(loop, slot)` does not poison the rest of its `(partition, -/// slot)` bucket (GH #542), while an `Inf` at one slot stays in that -/// bucket's denominator. Within each bucket the members are pushed in -/// `loop_partitions`' emission order, so the per-bucket `Σ|loop_score|` -/// accumulates in that order (bit-parity with the pre-#461 emitter, GH #468); -/// the `BTreeMap` over the bucket *grid* keeps the order the buckets -/// themselves are visited deterministic across runs. -pub fn compute_rel_loop_scores_per_element( - results: &Results, - loop_partitions: &IndexMap>>, -) -> HashMap> { - // Emission order, not a re-sort: the per-bucket denominator (built over the - // `members` grid below) sums `|loop_score|` in this order, and IEEE-754 - // addition is non-associative. See `compute_rel_loop_scores` for the full - // bit-parity rationale (GH #468). - let loop_ids: Vec<&String> = loop_partitions.keys().collect(); - - let offsets: Vec> = loop_ids - .iter() - .map(|id| results.offsets.get(&loop_score_ident(id)).copied()) - .collect(); - let n_slots: Vec = loop_ids - .iter() - .map(|id| loop_n_slots(loop_partitions, id)) - .collect(); - - // For each normalization group, the set of slot indices where some loop - // is in it. A scalar loop contributes its single slot 0; an arrayed - // loop contributes its per-slot partitions. This drives the broadcast - // stride for scalar loops (a scalar loop in partition `p` is "compared - // in" every slot of `p`'s buckets) and lets us pre-build the bucket - // membership. An unresolved (`None`) slot maps to the loop's own Solo - // group (GH #750), so a scalar `None` loop never broadcasts into an - // unrelated arrayed `None` loop's slots -- and never stretches its own - // stride to that loop's slot count. - let mut partition_slots: BTreeMap> = BTreeMap::new(); - for (i, id) in loop_ids.iter().enumerate() { - for k in 0..n_slots[i] { - partition_slots - .entry(NormGroup::for_loop( - slot_partition(loop_partitions, id, k), - i, - )) - .or_default() - .insert(k); - } - } - - // Per-loop output stride: an arrayed loop's own slot count; a scalar - // loop's (slot-0) partition's largest covered slot index + 1. - let strides: Vec = loop_ids - .iter() - .enumerate() - .map(|(i, id)| { - if n_slots[i] > 1 { - n_slots[i] - } else { - let p = NormGroup::for_loop(slot_partition(loop_partitions, id, 0), i); - partition_slots - .get(&p) - .and_then(|ks| ks.iter().max().copied()) - .map(|m| m + 1) - .unwrap_or(1) - .max(1) - } - }) - .collect(); - - // Pre-build the `(partition, slot)` -> [(loop_idx, read_slot)] grid. - // `read_slot` is 0 for a scalar member (broadcast) and `k` for an arrayed - // member; arrayed members past their own `n_slots` are not in any - // slot-`k` bucket (no OOB read past their own `loop_score` slots). - let mut members: BTreeMap> = BTreeMap::new(); - for (i, id) in loop_ids.iter().enumerate() { - if offsets[i].is_none() { - continue; - } - if n_slots[i] <= 1 { - // Scalar loop: appears in every slot of its (slot-0) group, - // always reading slot 0. A Solo group covers only slot 0, so a - // `None`-partition scalar loop is its own single bucket. - let p = NormGroup::for_loop(slot_partition(loop_partitions, id, 0), i); - if let Some(ks) = partition_slots.get(&p) { - for &k in ks { - members.entry((p, k)).or_default().push((i, 0)); +/// Collapse a per-slot series laid out as [`compute_rel_loop_scores`] +/// returns it (`series[t * n_slots + k]`, `n_slots = series.len() / +/// step_count`) to one signed value per step: the slot with the largest +/// `|value|`, its sign preserved, the lowest slot index on ties. A step +/// whose winner is non-finite (every slot `NaN`, or an `Inf`) reports `0.0`, +/// the "inactive" reading the layout and FFI consumers of an importance +/// series want; a `NaN` slot never displaces a finite candidate (the +/// comparison is false). A scalar loop (`n_slots == 1`) is returned +/// unchanged apart from that non-finite mapping. An empty series yields an +/// empty result. +pub fn argmax_abs_by_step(series: &[f64], step_count: usize) -> Vec { + if series.is_empty() || step_count == 0 { + return Vec::new(); + } + let n_slots = (series.len() / step_count).max(1); + (0..step_count) + .map(|t| { + let mut best = 0.0_f64; + let mut best_abs = -1.0_f64; + for &v in &series[t * n_slots..(t + 1) * n_slots] { + // `>` (not `>=`) keeps the lowest-index slot on ties; a NaN + // compares false and never becomes the winner. + if v.abs() > best_abs { + best_abs = v.abs(); + best = v; } } - } else { - for k in 0..n_slots[i] { - let p = NormGroup::for_loop(slot_partition(loop_partitions, id, k), i); - members.entry((p, k)).or_default().push((i, k)); - } - } - } - - let mut series: Vec> = offsets - .iter() - .enumerate() - .map(|(i, o)| { - if o.is_some() { - vec![0.0_f64; results.step_count * strides[i]] - } else { - Vec::new() - } + if best.is_finite() { best } else { 0.0 } }) - .collect(); - - for (step, row) in results.iter().enumerate() { - for (&(_p, k), member_list) in &members { - let denom: f64 = member_list - .iter() - .filter_map(|&(i, rs)| offsets[i].map(|off| denom_summand(row[off + rs]))) - .sum(); - for &(i, rs) in member_list { - let Some(off) = offsets[i] else { continue }; - let num = row[off + rs]; - let val = if denom == 0.0 { 0.0 } else { num / denom }; - series[i][step * strides[i] + k] = val; - } - } - } + .collect() +} - let mut out: HashMap> = HashMap::with_capacity(loop_ids.len()); - for (i, id) in loop_ids.iter().enumerate() { - if offsets[i].is_some() { - out.insert((*id).clone(), std::mem::take(&mut series[i])); - } - } - out +/// [`argmax_abs_by_step`] over every loop of a [`compute_rel_loop_scores`] +/// map: one signed importance series per loop, `step_count` long. This is +/// what a bare (unsubscripted) arrayed loop id means on every surface -- the +/// layout's `FeedbackLoop::importance_series` and the FFI's +/// `simlin_analyze_get_relative_loop_score("r1")` -- so the two cannot pick +/// different slots. +pub fn aggregate_per_element_argmax_abs( + per_element_rel_scores: &HashMap>, + step_count: usize, +) -> HashMap> { + per_element_rel_scores + .iter() + .map(|(id, series)| (id.clone(), argmax_abs_by_step(series, step_count))) + .collect() } -/// Resolve the slot offset to read for a loop with `n_slots` slots when -/// the partition is being queried at `element_index`. -/// -/// - Scalar loops (`n_slots <= 1`) → `Some(0)` -- slot 0 broadcasts -/// across every partition element. -/// - Arrayed loops with `element_index < n_slots` → `Some(element_index)`. -/// - Arrayed loops with `element_index >= n_slots` → `None` -- the loop -/// has no own element at this partition index and must not contribute -/// (matches the gating that [`compute_rel_loop_scores_per_element`] -/// applies in the full-sweep path). +/// The slot a raw per-element read of a loop with `n_slots` slots resolves +/// `element_index` to. /// -/// Returning `None` rather than clamping to `n_slots - 1` matters for -/// mixed-stride partitions where two arrayed loops have different -/// dimensionalities: the loop that runs out of slots first does NOT -/// stand in for the larger loop's later elements. Callers (the -/// streaming partition denominator and per-loop helpers) skip -/// `None`-returning members so the FFI's amortised path stays -/// bit-for-bit consistent with the full-sweep helper. +/// - Scalar loops (`n_slots <= 1`) -> `Some(0)`: their single slot answers +/// every element index. +/// - Arrayed loops with `element_index < n_slots` -> `Some(element_index)`. +/// - Arrayed loops with `element_index >= n_slots` -> `None`: the loop has +/// no such element, and the caller must not read past its own columns +/// into a neighboring loop's data. fn effective_slot(n_slots: usize, element_index: usize) -> Option { if n_slots <= 1 { Some(0) @@ -477,105 +351,6 @@ fn effective_slot(n_slots: usize, element_index: usize) -> Option { } } -/// Streaming per-element partition denominator: the amortized path the -/// libsimlin FFI's per-partition cache uses. -/// -/// For each `(loop_id, n_slots)` in the iterator whose `loop_score` -/// variable is present in `results`, contributes -/// `|row[off + slot]|` to the partition sum at every step, where -/// `slot` is determined by [`effective_slot`]: -/// - Scalar loops (`n_slots <= 1`) contribute slot 0 (broadcast). -/// - Arrayed loops with `element_index < n_slots` contribute their -/// own slot at `element_index`. -/// - Arrayed loops with `element_index >= n_slots` do NOT contribute -/// -- the loop has no own element at this partition index. -/// -/// A contributing slot's value goes through [`denom_summand`], so a -/// `NaN` slot is excluded from the sum and an `Inf` slot retained (GH -/// #542) -- the same per-bucket `NaN`-isolation the full-sweep -/// [`compute_rel_loop_scores_per_element`] applies. -/// -/// This skip-vs-clamp distinction matters for mixed-stride partitions -/// (two arrayed loops with different dimensionalities sharing a -/// partition). Producing the same partition sums as the full-sweep -/// [`compute_rel_loop_scores_per_element`] is the contract the -/// libsimlin FFI per-partition cache relies on; the streaming pair -/// must be a strictly cheaper path to the same numbers, not an -/// approximation. -/// -/// It exists so the libsimlin FFI per-partition cache can amortize across -/// element-aware queries (cache key `(partition, element_index)`) without -/// falling back to the non-streaming -/// [`compute_rel_loop_scores_per_element`]. This element-aware form is the -/// ONLY streaming entry point: `libsimlin::analysis` reaches for it even in -/// the scalar case, passing `n_slots = 1`, so a scalar-specialized twin -/// would have no caller. -pub fn compute_partition_denominator_for_element<'a, I>( - results: &Results, - loop_id_slots: I, - element_index: usize, -) -> Vec -where - I: IntoIterator, -{ - let entries: Vec<(usize, usize)> = loop_id_slots - .into_iter() - .filter_map(|(id, n_slots)| { - let off = results.offsets.get(&loop_score_ident(id)).copied()?; - let slot = effective_slot(n_slots, element_index)?; - Some((off, slot)) - }) - .collect(); - - let mut denom = vec![0.0_f64; results.step_count]; - for (t, row) in results.iter().enumerate() { - denom[t] = entries - .iter() - .map(|&(off, slot)| denom_summand(row[off + slot])) - .sum(); - } - denom -} - -/// One loop's relative-score series at element `k`, given a pre-computed -/// per-element partition denominator. -/// -/// Reads `row[off + slot]` as the numerator at each step, where `slot` -/// is determined by [`effective_slot`]: -/// - Scalar loops (`n_slots <= 1`) read slot 0 (broadcast). -/// - Arrayed loops with `element_index < n_slots` read their own slot. -/// - Arrayed loops with `element_index >= n_slots` return all zeros -/// -- the loop has no own element at this partition index, matching -/// the zero-fill that [`compute_rel_loop_scores_per_element`] -/// applies in the full-sweep path. -/// -/// Paired with [`compute_partition_denominator_for_element`] for SAFEDIV -/// normalisation. Returns `None` only when the loop's `loop_score` -/// variable is entirely absent from `results` (matching the scalar -/// streaming helper's "absent loop" contract); a present-but-no-element -/// query yields all-zeros, not `None`. -pub fn compute_rel_loop_score_for_element( - results: &Results, - loop_id: &str, - n_slots: usize, - element_index: usize, - denominator: &[f64], -) -> Option> { - let off = results.offsets.get(&loop_score_ident(loop_id)).copied()?; - let Some(slot) = effective_slot(n_slots, element_index) else { - // This loop has no own element at the queried partition index. - // Return zero-fill rather than reading another loop's slot. - return Some(vec![0.0; results.step_count]); - }; - let mut out = Vec::with_capacity(results.step_count); - for (t, row) in results.iter().enumerate() { - let num = row[off + slot]; - let denom = denominator[t]; - out.push(if denom == 0.0 { 0.0 } else { num / denom }); - } - Some(out) -} - /// Per-loop dimension metadata used by the FFI subscript resolver to /// turn `r1[Boston]` (or `r1[Boston, 2]`) into a concrete slot offset. /// @@ -775,164 +550,16 @@ pub fn build_loop_element_index( out } -/// Aggregate the per-element rel-score map produced by -/// [`compute_rel_loop_scores_per_element`] into a single signed series -/// per loop, via signed argmax-abs across each loop's own slots. -/// -/// Used by the layout's `compute_metadata` to populate -/// `FeedbackLoop::importance_series` with one value per saved step. -/// -/// ## Stride handling -/// -/// [`compute_rel_loop_scores_per_element`] lays each loop's series out -/// row-major as `series[t * stride + k]`, where: -/// -/// - for an **arrayed** loop, `stride == n_slots` -- the loop's own -/// slot count. Every slot index in `0..n_slots` is a real element, -/// so there are no padding positions and `n == stride`. -/// - for a **scalar** loop, `stride` is the largest slot index its -/// (slot-0) partition covers + 1 (1 if no arrayed loop shares that -/// partition); the loop's own `n_slots` is 1, so `stride >= n` with -/// positions `1..stride` being broadcast padding the scalar loop's -/// own value never occupies. -/// -/// This helper recovers `stride` from `series.len() / step_count` so -/// consumers don't need to track partition stride independently. -/// -/// The inner argmax-abs iterates only the loop's *own* `n_slots` -/// (`n_slots_by_loop[loop_id]`, default 1), reading `series[t * stride + k]` -/// for `k` in `0..n_slots`. For a scalar loop that is slot 0 only -- -/// the canonical scalar view, matching the pre-PR `compute_rel_loop_scores` -/// behaviour; the partition's broadcast-padding positions `1..stride` -/// are skipped. For an arrayed loop `stride == n_slots`, so the loop -/// over `0..n_slots` covers exactly the loop's own elements with no -/// out-of-bounds read. (Pre-Phase-2, arrayed loops were padded to the -/// partition's max stride and had to skip `n_slots..stride`; that -/// padding no longer exists.) -/// -/// ## Output -/// -/// Per loop: a `Vec` of length `step_count` (or empty if the -/// input series is empty). Non-finite picks are mapped to `0.0`, -/// matching the existing layout filter. -/// -/// Loops present in `per_element_rel_scores` but absent from -/// `n_slots_by_loop` default to `n_slots = 1` (scalar) -- legacy -/// callers that haven't snapshotted dim metadata still get a -/// well-formed result. -pub fn aggregate_per_element_argmax_abs( - per_element_rel_scores: &HashMap>, - n_slots_by_loop: &HashMap, - step_count: usize, -) -> HashMap> { - let mut out = HashMap::with_capacity(per_element_rel_scores.len()); - for (loop_id, series) in per_element_rel_scores { - if series.is_empty() { - out.insert(loop_id.clone(), Vec::new()); - continue; - } - let n = n_slots_by_loop.get(loop_id).copied().unwrap_or(1).max(1); - // Recover the helper's actual stride from the input length. - // For a scalar loop in a mixed partition stride > n == 1 - // (broadcast padding); for an arrayed loop and for a scalar - // loop alone in its partition stride == n. - let stride = (series.len() / step_count.max(1)).max(1); - let mut agg = Vec::with_capacity(step_count); - for t in 0..step_count { - let mut best = 0.0_f64; - let mut best_abs = -1.0_f64; - // Iterate this loop's own slots only. `n <= stride` always - // by construction (an arrayed loop has stride == n_slots; a - // scalar loop has n == 1 and stride >= 1), so the index - // never exceeds the series bounds. - for k in 0..n { - let v = series[t * stride + k]; - // `>` (not `>=`) keeps lowest-index slot on ties; NaN - // comparisons are always false so a NaN never displaces - // a finite candidate. - if v.abs() > best_abs { - best_abs = v.abs(); - best = v; - } - } - agg.push(if best.is_finite() { best } else { 0.0 }); - } - out.insert(loop_id.clone(), agg); - } - out -} - -/// Aggregate a multi-slot loop's relative-score series down to a single -/// signed series by picking the element with the largest `|rel[k, t]|` -/// at each step and emitting that element's *signed* value. -/// -/// Sign is preserved across steps even when the dominant element flips -/// (e.g. argmax may be slot 0 at one step and slot 1 at the next). -/// Ties on `|rel|` are broken by lowest slot index, matching the -/// stable-first-wins convention `max_by_key` provides. -/// -/// `denominators_per_element` must have length `n_slots`, with each -/// entry's length equal to `results.step_count`. The element-`k` -/// denominator is the partition sum at element `k` (e.g. produced by -/// [`compute_partition_denominator_for_element`] called with the -/// member loops and `element_index = k`). -/// -/// Scalar (`n_slots == 1`) reduces to identity: the aggregator returns -/// the single slot's series unchanged. Used -/// by both the layout single-line importance metric and the FFI -/// dispatch when callers pass a bare arrayed loop ID without a -/// subscript. -pub fn compute_rel_loop_score_argmax_abs( - results: &Results, - loop_id: &str, - n_slots: usize, - denominators_per_element: &[&[f64]], -) -> Option> { - let off = results.offsets.get(&loop_score_ident(loop_id)).copied()?; - let slots = n_slots.max(1); - debug_assert_eq!( - denominators_per_element.len(), - slots, - "argmax-abs needs one denominator series per element slot" - ); - let mut out = Vec::with_capacity(results.step_count); - for (t, row) in results.iter().enumerate() { - let mut best: f64 = 0.0; - let mut best_abs: f64 = -1.0; - for (k, denom_series) in denominators_per_element.iter().take(slots).enumerate() { - // `k` is iterated 0..slots and slots == max(n_slots, 1), so - // `effective_slot` always returns `Some`. We pass it through - // anyway for consistency with the streaming helpers and to - // catch any future caller that constructs `denominators_per_element` - // longer than the loop's own `n_slots`. - let Some(slot) = effective_slot(n_slots, k) else { - continue; - }; - let num = row[off + slot]; - let denom = denom_series[t]; - let rel = if denom == 0.0 { 0.0 } else { num / denom }; - // `>` (not `>=`) keeps the lowest-index slot when ties occur. - if rel.abs() > best_abs { - best_abs = rel.abs(); - best = rel; - } - } - out.push(best); - } - Some(out) -} - -/// The RAW sibling of [`compute_rel_loop_score_for_element`] (GH #998): -/// reads element `element_index` of `loop_id`'s raw `loop_score` series with -/// NO normalization. This is the accessor the lone-pin workaround needs -- -/// a modeler-pinned loop alone in its cycle partition has a relative score -/// of exactly `±1` by construction, so its RAW series is the informative -/// one, and before this the raw read (`get_series` on the synthetic name) -/// silently resolved to element 0 only. +/// Element `element_index` of `loop_id`'s RAW `loop_score` series, with NO +/// normalization (GH #998). This is the accessor the lone-pin workaround +/// needs -- a modeler-pinned loop alone in its cycle partition has a +/// relative score of exactly `±1` by construction, so its RAW series is the +/// informative one, and a plain `get_series` on the synthetic name resolves +/// to element 0 only. /// -/// Conventions match the relative sibling: scalar loops (`n_slots <= 1`) -/// read slot 0; an `element_index` past `n_slots` yields zero-fill rather -/// than reading a neighboring column; `None` only when the loop's +/// Scalar loops (`n_slots <= 1`) read slot 0 whatever the element index; an +/// `element_index` past `n_slots` yields zero-fill rather than reading a +/// neighboring column ([`effective_slot`]); `None` only when the loop's /// `loop_score` variable is absent from `results`. pub fn compute_raw_loop_score_for_element( results: &Results, @@ -947,18 +574,17 @@ pub fn compute_raw_loop_score_for_element( Some(results.iter().map(|row| row[off + slot]).collect()) } -/// The RAW sibling of [`compute_rel_loop_score_argmax_abs`] (GH #998): the -/// bare-id aggregate over an arrayed loop's raw `loop_score` slots -- each -/// step emits the SIGNED value of the slot with the largest magnitude, ties -/// broken to the lowest slot index. Scalar (`n_slots <= 1`) reduces to -/// identity, so a bare id means the same thing on the raw and relative -/// accessors (the dominant element's contribution, sign preserved). Two -/// deliberate divergences from the relative sibling: the dominant slot is -/// picked by |raw| here and by |relative| there (the two aggregates can -/// select different slots at a step where slots sit in different partitions -/// with different denominators), and a step where EVERY slot is NaN stays -/// NaN here (honest raw data) while the relative aggregate reports 0.0 -/// (its SAFEDIV "inactive" convention). +/// The bare-id aggregate over an arrayed loop's RAW `loop_score` slots (GH +/// #998): each step emits the SIGNED value of the slot with the largest +/// magnitude, ties broken to the lowest slot index. Scalar (`n_slots <= 1`) +/// reduces to identity, so a bare id means the same thing on the raw and +/// relative accessors (the dominant element's contribution, sign preserved). +/// Two deliberate divergences from the relative aggregate +/// ([`argmax_abs_by_step`]): the dominant slot is picked by |raw| here and by +/// |relative| there (the two can select different slots at a step where +/// slots sit in different partitions with different denominators), and a +/// step where EVERY slot is NaN stays NaN here (honest raw data) while the +/// relative aggregate reports 0.0 (its SAFEDIV "inactive" convention). pub fn compute_raw_loop_score_argmax_abs( results: &Results, loop_id: &str, @@ -970,12 +596,9 @@ pub fn compute_raw_loop_score_argmax_abs( for row in results.iter() { // A step where EVERY slot is NaN stays NaN: the raw accessor's // contract is honest data, and a fabricated finite 0.0 would hide - // undefined values the per-element accessors report. This is a - // DELIBERATE divergence from `compute_rel_loop_score_argmax_abs`, - // whose 0.0-on-undefined matches the relative score's SAFEDIV - // "inactive" convention (and whose consumers coerce non-finite to 0 - // anyway). A step with any finite slot picks the finite argmax -- - // NaN slots are skipped by the comparison below. + // undefined values the per-element accessors report. A step with + // any finite slot picks the finite argmax -- NaN slots are skipped + // by the comparison below. let mut best: Option = None; let mut best_abs: f64 = -1.0; for k in 0..slots { @@ -1027,10 +650,11 @@ pub struct RelLinkInput<'a> { /// ``` /// /// This is the link-level analogue of [`compute_rel_loop_scores`]'s -/// per-partition normalization and yields a value in `[-1, 1]` that *is* -/// comparable across the whole model -- the fraction of the target's change -/// attributable to that input. (LTM ref 13.3 / 13.12, Schoenberg 2020 -/// section 4.) +/// per-partition normalization -- the same [`group_totals`] and +/// [`relative_series`], keyed by target instead of by partition -- and +/// yields a value in `[-1, 1]` that *is* comparable across the whole model: +/// the fraction of the target's change attributable to that input. (LTM +/// ref 13.3 / 13.12, Schoenberg 2020 section 4.) /// /// **Signed, not absolute.** The literature phrases the relative magnitude /// as a `[0, 1]` fraction, but we keep the link's sign (so the result is in @@ -1039,14 +663,14 @@ pub struct RelLinkInput<'a> { /// or down, and dropping it would discard information the raw series /// carries. Callers that want the pure magnitude take `.abs()`. /// -/// **NaN/Inf semantics** match the loop-level denominator -/// ([`denom_summand`], GH #542): a `NaN` summand is excluded from the -/// per-target sum (one input being undefined at a step must not poison its -/// siblings' relative scores), while `Inf` is retained (a genuinely -/// diverging input dominates its target, so the finite siblings normalize -/// to `0` and the diverging one to `NaN` via `Inf/Inf`). A `0.0` -/// denominator yields `0.0` (SAFEDIV-0), and a link's own `NaN` numerator -/// stays `NaN` -- the honest "undefined here" per-link signal. +/// **NaN/Inf semantics** are the shared accumulator's (GH #542): a `NaN` +/// summand is excluded from the per-target sum (one input being undefined +/// at a step must not poison its siblings' relative scores), while `Inf` is +/// retained (a genuinely diverging input dominates its target, so the +/// finite siblings normalize to `0` and the diverging one to `NaN` via +/// `Inf/Inf`). A `0.0` denominator yields `0.0` (SAFEDIV-0), and a link's +/// own `NaN` numerator stays `NaN` -- the honest "undefined here" per-link +/// signal. /// /// **Denominator scope.** The sum runs over the input links that *have* a /// score series; `None`-score links (constants/parameters; out-of-loop @@ -1064,58 +688,27 @@ pub struct RelLinkInput<'a> { /// /// The return value is parallel to `links`: `Some(series)` of length /// `step_count` for each link that had a score, `None` for each link that -/// did not (mirroring the input's `score: None`). +/// did not (mirroring the input's `score: None`). Within a target the +/// denominator accumulates in the caller's link order. pub fn compute_rel_link_scores( links: &[RelLinkInput<'_>], step_count: usize, ) -> Vec>> { - // Group the indices of scored links by their `to` target. A - // `BTreeMap` keeps the float-summation order of each denominator - // deterministic across runs (the link order within a target is the - // caller's stable input order, preserved by pushing in iteration order). - let mut groups: BTreeMap<&str, Vec> = BTreeMap::new(); - for (i, link) in links.iter().enumerate() { - if link.score.is_some() { - groups.entry(link.to).or_default().push(i); - } - } - - // Per-target denominator series: Σ |score| over the target's scored - // links, with NaN excluded / Inf kept (denom_summand), one entry per - // saved step. - let mut denom_by_target: HashMap<&str, Vec> = HashMap::with_capacity(groups.len()); - for (target, indices) in &groups { - let mut denom = vec![0.0_f64; step_count]; - for &i in indices { - let series = links[i].score.expect("grouped links always have a score"); - // A score series shorter than step_count contributes 0 past its - // end (zip stops at the shorter); this never happens in production - // (every column is results.step_count long) but keeps the helper - // total. - for (d, &v) in denom.iter_mut().zip(series.iter()) { - *d += denom_summand(v); - } - } - denom_by_target.insert(target, denom); - } - + let totals = group_totals( + links + .iter() + .filter_map(|link| link.score.map(|s| (link.to, s.iter().copied()))), + step_count, + ); links .iter() .map(|link| { let series = link.score?; - let denom = &denom_by_target[link.to]; - // `denom` has length `step_count`; a `series` shorter than that - // (never in production -- every column is results.step_count long) - // is read as 0 past its end via `series.get(t)`. - let out: Vec = denom - .iter() - .enumerate() - .map(|(t, &d)| { - let num = series.get(t).copied().unwrap_or(0.0); - if d == 0.0 { 0.0 } else { num / d } - }) - .collect(); - Some(out) + // Pad a short series with zeros so the output is always + // `step_count` long (never in production -- every column is + // `results.step_count` long -- but it keeps the helper total). + let padded = (0..step_count).map(|t| series.get(t).copied().unwrap_or(0.0)); + Some(relative_series(padded, &totals[link.to])) }) .collect() } @@ -1137,7 +730,7 @@ pub fn compute_rel_link_scores( pub fn rel_link_group_sizes(links: &[RelLinkInput<'_>]) -> Vec { // A link CONTRIBUTES when its series adds at least one summand to its // target's denominator -- i.e. it has at least one non-NaN entry - // (`denom_summand` excludes NaN and keeps Inf, and Inf is not NaN). An + // (`add_to_total` excludes NaN and keeps Inf, and Inf is not NaN). An // all-NaN series never contributes, so counting it would over-report // competition: its finite sibling normalizes against itself alone and // reads +/-1 by construction, which is exactly the degeneracy this @@ -1200,7 +793,6 @@ mod tests { row[i + 1] = ser[step]; } } - let sim_specs = SimSpecs { start: 0.0, stop: (step_count.saturating_sub(1)) as f64, @@ -1209,7 +801,6 @@ mod tests { sim_method: SimMethod::Euler, time_units: None, }; - Results { offsets, data: data.into_boxed_slice(), @@ -1241,88 +832,82 @@ mod tests { .collect() } - /// Inlined reference implementation of `compute_rel_loop_scores`'s - /// slot-0-convenience SAFEDIV formula: each loop normalizes its slot-0 - /// value against the sum of slot-0 values over loops sharing its slot-0 - /// partition. The proptest compares against this to catch any numeric - /// divergence. - /// - /// `NaN` summands are excluded from the partition sum here too (GH - /// #542), independently re-deriving the production `denom_summand` - /// rule, so the oracle stays correct if a future generator introduces - /// non-finite samples (the current generators are finite-only, so this - /// branch is latent but kept honest). `+/-Inf` is kept in the sum, - /// matching the production semantics. - fn reference_rel_loop_scores( - loop_ids: &[String], - loop_partitions: &IndexMap>>, - series: &[Vec], - ) -> Vec> { - let step_count = series.first().map(|s| s.len()).unwrap_or(0); - let mut groups: HashMap, Vec> = HashMap::new(); + /// Build a `Results` with each loop occupying a configurable number of + /// slots. Layout: `time | loop0 slot 0..n0 | loop1 slot 0..n1 | ...`. + /// `loop_data[i][step][slot]` is the value at (step, slot) for loop i. + fn make_arrayed_results( + loop_ids: &[&str], + slots_per_loop: &[usize], + loop_data: &[Vec>], + ) -> Results { + assert_eq!(loop_ids.len(), slots_per_loop.len()); + assert_eq!(loop_ids.len(), loop_data.len()); + let step_count = loop_data[0].len(); + for d in loop_data.iter() { + assert_eq!(d.len(), step_count); + } + let total_slots: usize = slots_per_loop.iter().sum(); + let step_size = 1 + total_slots; + let mut data = vec![0.0_f64; step_count * step_size]; + let mut offsets: HashMap, usize> = HashMap::new(); + offsets.insert(Ident::new("time"), 0); + let mut cursor = 1; + let mut loop_offsets = Vec::with_capacity(loop_ids.len()); for (i, id) in loop_ids.iter().enumerate() { - let key = loop_partitions - .get(id) - .and_then(|v| v.first().copied().flatten()); - groups.entry(key).or_default().push(i); + offsets.insert(loop_score_ident(id), cursor); + loop_offsets.push(cursor); + cursor += slots_per_loop[i]; } - let mut out: Vec> = (0..loop_ids.len()) - .map(|_| Vec::with_capacity(step_count)) - .collect(); - // `t` is an index into every per-loop series simultaneously, so - // the range-based form is clearer than an iterator over one series. - #[allow(clippy::needless_range_loop)] - for t in 0..step_count { - for indices in groups.values() { - let denom: f64 = indices - .iter() - .map(|&i| { - let v = series[i][t]; - if v.is_nan() { 0.0 } else { v.abs() } - }) - .sum(); - for &i in indices { - let num = series[i][t]; - let val = if denom == 0.0 { 0.0 } else { num / denom }; - out[i].push(val); + for step in 0..step_count { + let row = &mut data[step * step_size..(step + 1) * step_size]; + row[0] = step as f64; + for (i, &off) in loop_offsets.iter().enumerate() { + let slots = &loop_data[i][step]; + assert_eq!(slots.len(), slots_per_loop[i]); + for (slot, &v) in slots.iter().enumerate() { + row[off + slot] = v; } } } - out - } - - /// Naive, from-first-principles reference for - /// `compute_rel_loop_scores_per_element`. Where the engine pre-builds - /// a `(partition, slot)` BTreeMap grid, this loops directly over each - /// `(loop, output-slot, step)` and re-derives the SAFEDIV denominator - /// by scanning *every* loop and asking "is it a member of this bucket?" - /// -- a structurally different computation, so the proptest is a real - /// oracle, not a paraphrase of the implementation. - /// - /// Member rule (mirrors the engine's `slot_partition` broadcast and - /// `effective_slot` gating, spelled out inline rather than via the - /// engine helpers): for a bucket at partition `p`, output slot `k`, - /// loop `j` is a member iff - /// - `j` is **scalar** (`n_slots == 1`): `slots[j][0] == p` -- it - /// broadcasts its single value into every slot of partition `p`; - /// it reads slot 0; OR - /// - `j` is **arrayed** (`n_slots > 1`) and `k < n_slots[j]` and - /// `slots[j][k] == p` -- it reads its own slot `k`; an arrayed - /// loop past its own slot count is not a member of slot-`k` - /// buckets (no out-of-bounds read into another loop's data). + let sim_specs = SimSpecs { + start: 0.0, + stop: (step_count.saturating_sub(1)) as f64, + dt: Dt::Dt(1.0), + save_step: None, + sim_method: SimMethod::Euler, + time_units: None, + }; + Results { + offsets, + data: data.into_boxed_slice(), + step_size, + step_count, + specs: Specs::from(&sim_specs), + is_vensim: false, + } + } + + /// Naive, from-first-principles reference for + /// [`compute_rel_loop_scores`]. Where the engine accumulates group + /// totals in one pass, this loops directly over each `(loop, slot, + /// step)` and re-derives the SAFEDIV denominator by scanning *every* + /// other `(loop, slot)` and asking "is it in my partition?" -- a + /// structurally different computation, so the proptest is a real + /// oracle, not a paraphrase of the implementation. /// - /// Output stride per loop: an arrayed loop's own slot count; a scalar - /// loop's (slot-0) partition's largest covered slot index + 1 (so a - /// scalar loop alone in its partition has stride 1). Slots a scalar - /// loop's partition does not cover stay 0.0 (gaps in the covered set). + /// Member rule, spelled out inline rather than via the engine helpers: + /// `(j, k')` is in `(i, k)`'s group iff both resolve to the same + /// `Some(p)`; an unresolved (`None`) slot is in a group of its own + /// (GH #750), so it normalizes against itself only. `NaN` summands are + /// excluded (GH #542) and `Inf` kept, re-derived inline too. /// /// `slots[i]` is loop `i`'s per-slot partition vector (length 1 for a /// scalar loop); `series[i][step][slot]` its `loop_score`. Returns one - /// flat `Vec` per loop in `loop_ids` order; element `step * - /// stride_i + k`. Float summation walks loops in ascending index -- - /// the same order the engine's sorted-`loop_id` member lists produce - /// when the ids are `L{i}` with `i < 10` -- so the comparison is exact. - fn reference_rel_loop_scores_per_element( + /// flat `Vec` per loop in `loop_ids` order, `step * n_slots_i + k` + /// indexed. The denominator walks members in ascending `(loop, slot)` + /// order -- the production accumulation order -- so the comparison is + /// bit-exact, not merely close. + fn reference_rel_loop_scores( loop_ids: &[String], slots: &[Vec>], series: &[Vec>], @@ -1330,108 +915,33 @@ mod tests { ) -> Vec> { let n = loop_ids.len(); let n_slots: Vec = slots.iter().map(|v| v.len().max(1)).collect(); - // The partition loop `i` carries into slot `k`: an arrayed loop's - // own per-slot entry; a scalar loop broadcasts slot 0's partition. - let slot_part = |i: usize, k: usize| -> Option { - if n_slots[i] <= 1 { - slots[i].first().copied().flatten() - } else { - slots[i].get(k).copied().flatten() - } - }; - // The normalization group of loop `i`'s slot `k`: the resolved - // partition, or the loop's own Solo group when unresolved (GH #750 - // -- unrelated unpartitioned loops never share a bucket). - // Re-derived inline so the oracle stays structurally independent of - // the production grouping. - let slot_group = |i: usize, k: usize| -> NormGroup { - match slot_part(i, k) { - Some(p) => NormGroup::Partition(p), - None => NormGroup::Solo(i), - } - }; - // For each group, the set of slot indices some loop occupies in it - // (a scalar loop occupies only slot 0): used for the scalar - // broadcast stride and to know whether a scalar loop "appears" at a - // given output slot. - let mut partition_slots: BTreeMap> = BTreeMap::new(); - for (i, &ns) in n_slots.iter().enumerate() { - for k in 0..ns { - partition_slots - .entry(slot_group(i, k)) - .or_default() - .insert(k); - } - } - let strides: Vec = (0..n) - .map(|i| { - if n_slots[i] > 1 { - n_slots[i] - } else { - let p = slot_group(i, 0); - partition_slots - .get(&p) - .and_then(|ks| ks.iter().max().copied()) - .map(|m| m + 1) - .unwrap_or(1) - .max(1) - } - }) - .collect(); - // Is loop `j` a member of the bucket (group `p`, output slot - // `k`)? If so, which of its own slots does it read? - // - // A scalar loop is in `(p, k)` iff its (slot-0) group is `p` - // AND `k` is a slot index *some* loop occupies in `p` - // (`partition_slots[p]`) -- the engine only pushes a scalar member - // into the slots its partition actually spans, so a stride that - // overshoots a gap leaves that output position at 0.0. - let read_slot_in_bucket = |j: usize, p: NormGroup, k: usize| -> Option { - if n_slots[j] <= 1 { - let covers_k = partition_slots.get(&p).is_some_and(|ks| ks.contains(&k)); - if slot_group(j, 0) == p && covers_k { - Some(0) - } else { - None - } - } else if k < n_slots[j] && slot_group(j, k) == p { - Some(k) - } else { - None - } - }; + let partition = + |i: usize, k: usize| -> Option { slots[i].get(k).copied().flatten() }; let mut out: Vec> = (0..n) - .map(|i| vec![0.0_f64; step_count * strides[i]]) + .map(|i| vec![0.0_f64; step_count * n_slots[i]]) .collect(); for i in 0..n { - for k in 0..strides[i] { - // Which bucket is loop `i`'s output slot `k` in, and does - // `i` actually occupy it? An arrayed loop occupies every - // `k < n_slots` (and stride == n_slots, so no overshoot); - // a scalar loop occupies only the slots its partition - // covers (its stride may overshoot a gap). - let p = slot_group(i, k); - let Some(read_i) = read_slot_in_bucket(i, p, k) else { - continue; + for k in 0..n_slots[i] { + // The members of (i, k)'s group, in ascending (loop, slot) + // order: every slot sharing its resolved partition, or just + // itself when unresolved. + let members: Vec<(usize, usize)> = match partition(i, k) { + Some(p) => (0..n) + .flat_map(|j| (0..n_slots[j]).map(move |kk| (j, kk))) + .filter(|&(j, kk)| partition(j, kk) == Some(p)) + .collect(), + None => vec![(i, k)], }; - // Bucket members, scanned over every loop. - let bucket: Vec<(usize, usize)> = (0..n) - .filter_map(|j| read_slot_in_bucket(j, p, k).map(|rs| (j, rs))) - .collect(); for step in 0..step_count { - // `NaN` summands excluded (GH #542), re-derived inline - // so this oracle stays structurally independent of the - // production `denom_summand`; `Inf` kept in the sum. - let denom: f64 = bucket + let denom: f64 = members .iter() - .map(|&(j, rs)| { - let v = series[j][step][rs]; + .map(|&(j, kk)| { + let v = series[j][step][kk]; if v.is_nan() { 0.0 } else { v.abs() } }) .sum(); - let num = series[i][step][read_i]; - let val = if denom == 0.0 { 0.0 } else { num / denom }; - out[i][step * strides[i] + k] = val; + let num = series[i][step][k]; + out[i][step * n_slots[i] + k] = if denom == 0.0 { 0.0 } else { num / denom }; } } } @@ -1467,8 +977,7 @@ mod tests { /// its per-partition `Σ|loop_score|` denominator in the `loop_partitions` /// IndexMap's *emission* (insertion) order, NOT a re-sort of the loop ids. /// IEEE-754 addition is non-associative, so a lex re-sort would perturb the - /// denominator at the ULP and break bit-for-bit parity with the pre-#461 - /// compile-time emitter. + /// denominator at the ULP and break the bit-for-bit stability of the sum. /// /// The fixture puts 12 same-prefix loops (`r1..r12`) in one partition with /// scores chosen so the emission-order abs-sum and the lex-order abs-sum are @@ -1636,14 +1145,15 @@ mod tests { } #[test] - fn nan_loop_score_isolated_in_per_element_bucket() { - // Per-element twin of `nan_loop_score_isolated_to_its_own_loop`: - // a NaN at one (loop, slot) must not poison the sibling sharing - // that `(partition, slot)` bucket. Two coupled A2A loops, 2 - // slots each; plant a NaN at A's slot 0 at step 0. + fn nan_slot_isolated_to_its_own_member() { + // Arrayed twin of `nan_loop_score_isolated_to_its_own_loop`: a NaN + // at one (loop, slot) is dropped from the partition's total, so + // every other member of the partition -- the loop's own sibling + // slot included -- normalizes against the healthy total. Two + // coupled A2A loops, 2 slots each, all four slots in partition 0; + // plant a NaN at A's slot 0 at step 0. let n_slots: usize = 2; - // A slot0 = [NaN, 9], A slot1 = [4, 4] - // B slot0 = [3, 3], B slot1 = [6, 6] + // A = [NaN, 4] then [9, 4]; B = [3, 6] then [3, 6]. let loop_data = vec![ vec![vec![f64::NAN, 4.0], vec![9.0, 4.0]], vec![vec![3.0, 6.0], vec![3.0, 6.0]], @@ -1652,26 +1162,35 @@ mod tests { let partitions = mapping_per_slot(&[("A", vec![Some(0); n_slots]), ("B", vec![Some(0); n_slots])]); - let rel = compute_rel_loop_scores_per_element(&results, &partitions); + let rel = compute_rel_loop_scores(&results, &partitions); let a = rel.get("A").unwrap(); let b = rel.get("B").unwrap(); - - // step 0, slot 0: bucket {A, B}; A is NaN, so it is dropped from - // the denom (denom = |3| = 3). A's own rel score = NaN/3 = NaN; - // B's = 3/3 = 1 (healthy, NOT poisoned). let at = |step: usize, k: usize| step * n_slots + k; - assert!(a[at(0, 0)].is_nan(), "NaN loop keeps its own NaN rel score"); + + // step 0: total over the partition = |4| + |3| + |6| = 13 (the NaN + // slot dropped). A[0] = NaN/13 = NaN; the other three members + // share the healthy 13. + assert!(a[at(0, 0)].is_nan(), "the NaN slot keeps its own NaN"); assert!( - (b[at(0, 0)] - 1.0).abs() < 1e-12, - "healthy slot-0 sibling normalizes against the healthy denom: got {}", + (a[at(0, 1)] - 4.0 / 13.0).abs() < 1e-12, + "got {}", + a[at(0, 1)] + ); + assert!( + (b[at(0, 0)] - 3.0 / 13.0).abs() < 1e-12, + "got {}", b[at(0, 0)] ); - // step 0, slot 1: both finite; denom = |4| + |6| = 10. - assert!((a[at(0, 1)] - 0.4).abs() < 1e-12); - assert!((b[at(0, 1)] - 0.6).abs() < 1e-12); - // step 1: all finite; slot 0 denom = |9| + |3| = 12, slot 1 = 10. - assert!((a[at(1, 0)] - (9.0 / 12.0)).abs() < 1e-12); - assert!((b[at(1, 0)] - (3.0 / 12.0)).abs() < 1e-12); + assert!( + (b[at(0, 1)] - 6.0 / 13.0).abs() < 1e-12, + "got {}", + b[at(0, 1)] + ); + // step 1: all finite; total = 9 + 4 + 3 + 6 = 22. + assert!((a[at(1, 0)] - 9.0 / 22.0).abs() < 1e-12); + assert!((a[at(1, 1)] - 4.0 / 22.0).abs() < 1e-12); + assert!((b[at(1, 0)] - 3.0 / 22.0).abs() < 1e-12); + assert!((b[at(1, 1)] - 6.0 / 22.0).abs() < 1e-12); } #[test] @@ -1681,9 +1200,7 @@ mod tests { // denominators go to zero there). Unlike a NaN, an Inf is NOT // filtered from the partition sum -- it stays, so the dominated // siblings correctly go to 0 (finite/Inf) and the dominant loop - // momentarily reads NaN (Inf/Inf). This preserves the legitimate - // inflection-point behaviour of the removed SAFEDIV equation - // bug-for-bug; only the NaN-poisoning case (GH #542) changed. + // momentarily reads NaN (Inf/Inf). let inf = f64::INFINITY; let series_a = &[inf, 2.0][..]; let series_b = &[5.0, 3.0][..]; @@ -1708,23 +1225,22 @@ mod tests { assert!((rel_b[1] - 0.6).abs() < 1e-12); } + /// A finite partition total that overflows saturates to `f64::MAX` + /// rather than becoming `Inf`: every summand was finite, so the + /// members keep finite (tiny) shares instead of all reading `0`. A + /// genuine `Inf` summand still makes the total `Inf`. #[test] - fn inf_loop_score_kept_in_per_element_bucket() { - // Per-element twin of `inf_loop_score_kept_in_denominator`: an - // +Inf at one (loop, slot) stays in that bucket's denominator, - // so the dominated sibling -> 0 and the dominant loop -> NaN. - let n_slots: usize = 1; - // A slot0 = [Inf], B slot0 = [5]. - let loop_data = vec![vec![vec![f64::INFINITY]], vec![vec![5.0]]]; - let results = make_arrayed_results(&["A", "B"], &[n_slots, n_slots], &loop_data); - let partitions = - mapping_per_slot(&[("A", vec![Some(0); n_slots]), ("B", vec![Some(0); n_slots])]); - - let rel = compute_rel_loop_scores_per_element(&results, &partitions); - let a = rel.get("A").unwrap(); - let b = rel.get("B").unwrap(); - assert!(a[0].is_nan(), "dominant +Inf loop -> NaN in its bucket"); - assert_eq!(b[0], 0.0, "dominated sibling -> 0 in the same bucket"); + fn group_totals_saturate_on_finite_overflow_but_keep_inf() { + let totals = group_totals( + [ + (0usize, vec![f64::MAX, 1.0]), + (0usize, vec![f64::MAX, f64::INFINITY]), + ], + 2, + ); + let total = &totals[&0]; + assert_eq!(total[0], f64::MAX, "finite overflow saturates"); + assert_eq!(total[1], f64::INFINITY, "a real Inf summand is kept"); } #[test] @@ -1734,9 +1250,9 @@ mod tests { // loops), so two `None`-partition loops must NOT cross-normalize -- // they may be entirely unrelated subsystems. Each gets its own // singleton group, collapsing to the documented lone-pin degeneracy - // (sign-preserving +/-1, or 0 via SAFEDIV-0). The old behavior - // pooled them into one default bucket (rel = 0.75 / 0.25 here), the - // GH #487-class cross-pollution. + // (sign-preserving +/-1, or 0 via SAFEDIV-0). A shared default + // bucket would pool them (rel = 0.75 / 0.25 here), the GH #487-class + // cross-pollution. let series_a = &[3.0, 0.0][..]; let series_b = &[-1.0, 2.0][..]; let results = make_results_for_loops(&[("A", series_a), ("B", series_b)]); @@ -1770,171 +1286,119 @@ mod tests { } #[test] - fn unpartitioned_loops_do_not_pool_per_element() { - // Per-element twin of `unpartitioned_loops_normalize_independently` - // (GH #750): a scalar `None`-partition loop must not broadcast into - // an unrelated arrayed `None`-partition loop's slot buckets (which - // would both dilute the arrayed loop's slots and stretch the scalar - // loop's stride to the arrayed loop's slot count). + fn unpartitioned_arrayed_slots_are_each_solo() { + // Arrayed twin of `unpartitioned_loops_normalize_independently`: an + // unresolved slot is a Solo member on its own, so the two `None` + // slots of one arrayed loop do not normalize against each other, + // and a `None` scalar loop does not join either of them. let n_slots: usize = 2; - // One step: A is scalar (one slot, score 3); B is arrayed over 2 - // slots (scores 4 and 6). Data shape is `loop_data[i][step][slot]`. + // One step: A is scalar (score 3); B is arrayed over 2 slots (4, 6). let loop_data = vec![vec![vec![3.0]], vec![vec![4.0, 6.0]]]; let results = make_arrayed_results(&["A", "B"], &[1, n_slots], &loop_data); let partitions = mapping_per_slot(&[("A", vec![None]), ("B", vec![None, None])]); - let rel = compute_rel_loop_scores_per_element(&results, &partitions); + let rel = compute_rel_loop_scores(&results, &partitions); let a = rel.get("A").unwrap(); let b = rel.get("B").unwrap(); - // A stays a single-slot series normalized against itself only. - assert_eq!( - a.len(), - 1, - "a None-partition scalar loop keeps stride 1 (no broadcast into \ - unrelated None loops' slots); got {a:?}" - ); + assert_eq!(a.len(), 1, "a scalar loop has exactly one series"); assert!((a[0] - 1.0).abs() < 1e-12, "got {}", a[0]); - // B's slots normalize against B alone (each slot its own bucket). assert_eq!(b.len(), n_slots); assert!((b[0] - 1.0).abs() < 1e-12, "got {}", b[0]); assert!((b[1] - 1.0).abs() < 1e-12, "got {}", b[1]); } - /// Per-element variant: two A2A loops over an element-wise-coupled - /// dimension (every slot in partition 0), each with 3 element slots. - /// At every element k both loops' slot k lands in bucket `(0, k)`, so - /// the sum of absolute rel-scores at element k must equal 1.0 (non-zero - /// elements) or 0.0 (zero-denominator elements) independently -- the - /// scalar path collapses to slot 0 and would sum to 1.0 only for - /// element 0. + /// The headline rule: every `(loop, slot)` of a partition shares ONE + /// denominator. A coupled A2A loop (both slots in partition 0) and a + /// scalar loop in the same partition: at each step the total is + /// `|A[0]| + |A[1]| + |S|`, the scalar loop has exactly one series, and + /// the three shares sum to 1. Grouping by `(partition, slot)` instead + /// would give A[0] a denominator of `|A[0]| + |S|` and A[1] one of + /// `|A[1]| + |S|`, each missing a member, and would hand the scalar loop + /// two different series. #[test] - fn per_element_helper_normalizes_within_each_slot() { - let n_slots: usize = 3; - let step_count: usize = 4; - // Two A2A loops with distinct per-element magnitudes so each - // element has a meaningful partition split. - // A: [1, 3, 5, 2, ...] per element 0, 1, 2, ... - // B: [3, 1, 15, 6, ...] per element 0, 1, 2, ... - // Constructing by steps * elements and writing directly into - // a Results layout avoids coupling to the rest of the engine. - let mut data = vec![0.0_f64; step_count * (2 * n_slots + 1)]; - let step_size = 2 * n_slots + 1; - let a_off = 1; - let b_off = 1 + n_slots; - for step in 0..step_count { - let row = &mut data[step * step_size..(step + 1) * step_size]; - row[0] = step as f64; // time - for k in 0..n_slots { - row[a_off + k] = ((step + 1) * (k + 1)) as f64; - row[b_off + k] = ((step + 1) * (k + 2)) as f64; - } - } - let mut offsets: HashMap, usize> = HashMap::new(); - offsets.insert(Ident::new("time"), 0); - offsets.insert(loop_score_ident("A"), a_off); - offsets.insert(loop_score_ident("B"), b_off); - - let sim_specs = crate::datamodel::SimSpecs { - start: 0.0, - stop: (step_count - 1) as f64, - dt: crate::datamodel::Dt::Dt(1.0), - save_step: None, - sim_method: crate::datamodel::SimMethod::Euler, - time_units: None, - }; - let results = Results { - offsets, - data: data.into_boxed_slice(), - step_size, - step_count, - specs: crate::results::Specs::from(&sim_specs), - is_vensim: false, - }; + fn arrayed_slots_and_scalar_loops_share_one_partition_denominator() { + let n_slots: usize = 2; + // step 0: A = [3, 6], S = 1 -> total 10 + // step 1: A = [1, 4], S = -5 -> total 10 + let loop_data = vec![ + vec![vec![3.0, 6.0], vec![1.0, 4.0]], + vec![vec![1.0], vec![-5.0]], + ]; + let results = make_arrayed_results(&["A", "S"], &[n_slots, 1], &loop_data); + let partitions = mapping_per_slot(&[("A", vec![Some(0); n_slots]), ("S", vec![Some(0)])]); - // Both loops coupled: every slot in partition 0. - let partitions = - mapping_per_slot(&[("A", vec![Some(0); n_slots]), ("B", vec![Some(0); n_slots])]); + let rel = compute_rel_loop_scores(&results, &partitions); + let a = rel.get("A").unwrap(); + let s = rel.get("S").unwrap(); + assert_eq!(a.len(), 2 * n_slots); + assert_eq!(s.len(), 2, "a scalar loop has one series, step_count long"); - let rel = compute_rel_loop_scores_per_element(&results, &partitions); - let a = rel.get("A").expect("A must have a series"); - let b = rel.get("B").expect("B must have a series"); - assert_eq!(a.len(), step_count * n_slots); - assert_eq!(b.len(), step_count * n_slots); + let at = |step: usize, k: usize| step * n_slots + k; + assert!((a[at(0, 0)] - 0.3).abs() < 1e-12, "got {}", a[at(0, 0)]); + assert!((a[at(0, 1)] - 0.6).abs() < 1e-12, "got {}", a[at(0, 1)]); + assert!((s[0] - 0.1).abs() < 1e-12, "got {}", s[0]); + assert!((a[at(1, 0)] - 0.1).abs() < 1e-12, "got {}", a[at(1, 0)]); + assert!((a[at(1, 1)] - 0.4).abs() < 1e-12, "got {}", a[at(1, 1)]); + assert!((s[1] - (-0.5)).abs() < 1e-12, "got {}", s[1]); + } + + /// Two arrayed loops of different widths in one partition: all five + /// slots share the denominator. A slot index means nothing across + /// loops here -- A's slot 2 has no counterpart in B and still competes + /// with every member. + #[test] + fn arrayed_loops_of_different_widths_share_a_partition() { + // One step: A = [1, 2, 3] (3 slots), B = [4, 10] (2 slots); total 20. + let loop_data = vec![vec![vec![1.0, 2.0, 3.0]], vec![vec![4.0, 10.0]]]; + let results = make_arrayed_results(&["A", "B"], &[3, 2], &loop_data); + let partitions = mapping_per_slot(&[("A", vec![Some(0); 3]), ("B", vec![Some(0); 2])]); - for step in 0..step_count { - for k in 0..n_slots { - let idx = step * n_slots + k; - let sum = a[idx].abs() + b[idx].abs(); - // Magnitudes per element are finite and non-zero here, - // so the sum of absolute rel-scores must be 1.0 with - // full float precision. - assert!( - (sum - 1.0).abs() < 1e-12, - "step {step} elem {k}: |a|+|b| = {sum}, not 1.0" - ); - } + let rel = compute_rel_loop_scores(&results, &partitions); + let a = rel.get("A").unwrap(); + let b = rel.get("B").unwrap(); + assert_eq!(a.len(), 3); + assert_eq!(b.len(), 2); + for (got, expected) in a.iter().zip([0.05, 0.10, 0.15]) { + assert!( + (got - expected).abs() < 1e-12, + "got {got}, expected {expected}" + ); + } + for (got, expected) in b.iter().zip([0.20, 0.50]) { + assert!( + (got - expected).abs() < 1e-12, + "got {got}, expected {expected}" + ); } } - /// Per-element variant, the headline GH #487 case: two A2A loops over - /// element-wise-*uncoupled* dimensions -- each slot of each loop is in - /// its own partition, and no slot of A shares a partition with any slot - /// of B. Each loop's slot k therefore normalizes against itself only, - /// so every rel score is ±1.0 -- the two loops do NOT cross-normalize - /// even though `compute_rel_loop_scores`'s slot-0-pooled view used to - /// (pre-fix) lump them when both had `None` partitions. + /// The GH #487 case: two A2A loops over element-wise-*uncoupled* + /// dimensions -- each slot of each loop is in its own partition, and no + /// slot of A shares a partition with any slot of B. Each slot + /// therefore normalizes against itself only, so every rel score is + /// ±1.0 -- the two loops do NOT cross-normalize. #[test] - fn per_element_uncoupled_a2a_loops_do_not_cross_normalize() { + fn uncoupled_a2a_slots_self_normalize() { let step_count: usize = 3; - // A has 2 slots, B has 3 slots; A's slots are partitions 0,1 and - // B's slots are partitions 2,3,4 -- all distinct, none shared. - // Layout: time | A slot0 | A slot1 | B slot0..2 - let step_size = 1 + 2 + 3; - let a_off = 1; - let b_off = 3; - let mut data = vec![0.0_f64; step_count * step_size]; - for step in 0..step_count { - let row = &mut data[step * step_size..(step + 1) * step_size]; - row[0] = step as f64; - // Distinct, non-zero, per-step-varying magnitudes. - for k in 0..2 { - row[a_off + k] = ((step + 2) * (k + 1)) as f64; - } - for k in 0..3 { - row[b_off + k] = -(((step + 3) * (k + 1)) as f64); - } - } - let mut offsets: HashMap, usize> = HashMap::new(); - offsets.insert(Ident::new("time"), 0); - offsets.insert(loop_score_ident("A"), a_off); - offsets.insert(loop_score_ident("B"), b_off); - let sim_specs = crate::datamodel::SimSpecs { - start: 0.0, - stop: (step_count - 1) as f64, - dt: crate::datamodel::Dt::Dt(1.0), - save_step: None, - sim_method: crate::datamodel::SimMethod::Euler, - time_units: None, - }; - let results = Results { - offsets, - data: data.into_boxed_slice(), - step_size, - step_count, - specs: crate::results::Specs::from(&sim_specs), - is_vensim: false, - }; - + // A has 2 slots (partitions 0, 1), B has 3 slots (partitions 2, 3, + // 4); distinct, non-zero, per-step-varying magnitudes. + let a_data: Vec> = (0..step_count) + .map(|step| (0..2).map(|k| ((step + 2) * (k + 1)) as f64).collect()) + .collect(); + let b_data: Vec> = (0..step_count) + .map(|step| (0..3).map(|k| -(((step + 3) * (k + 1)) as f64)).collect()) + .collect(); + let results = make_arrayed_results(&["A", "B"], &[2, 3], &[a_data, b_data]); let partitions = mapping_per_slot(&[ ("A", vec![Some(0), Some(1)]), ("B", vec![Some(2), Some(3), Some(4)]), ]); - let rel = compute_rel_loop_scores_per_element(&results, &partitions); + + let rel = compute_rel_loop_scores(&results, &partitions); let a = rel.get("A").expect("A must have a series"); let b = rel.get("B").expect("B must have a series"); assert_eq!(a.len(), step_count * 2); assert_eq!(b.len(), step_count * 3); - // Every slot of every loop normalizes against itself only -> ±1.0. for &v in a.iter().chain(b.iter()) { assert!( (v.abs() - 1.0).abs() < 1e-12, @@ -1943,83 +1407,41 @@ mod tests { } } - /// Mixed partition: one scalar loop and one A2A loop. The scalar - /// loop's single slot broadcasts into every element of the - /// partition's max-slots denominator -- this matches the pre-PR - /// compile-time emitter, which expanded a scalar loop_score - /// reference across the arrayed rel_loop_score target. + /// The discovery path's inputs -- `(time, score)` pairs per loop with a + /// per-loop group -- normalize through the same two functions the + /// exhaustive owner uses, so feeding the same numbers to both yields the + /// same relative series. This is the parity the two surfaces rely on. #[test] - fn per_element_helper_broadcasts_scalar_across_elements() { - let n_slots: usize = 2; - let step_count: usize = 2; - // Layout: time | A (scalar, 1 slot) | B (A2A, 2 slots) - let step_size = 1 + 1 + n_slots; - let a_off = 1; - let b_off = 2; - let mut data = vec![0.0_f64; step_count * step_size]; - // A[t=0] = 2, B[t=0] = [3, 6]; denominators = [5, 8] - // A[t=1] = 1, B[t=1] = [1, 4]; denominators = [2, 5] - let a_vals = [2.0_f64, 1.0]; - let b_vals = [[3.0_f64, 6.0], [1.0, 4.0]]; - for step in 0..step_count { - let row = &mut data[step * step_size..(step + 1) * step_size]; - row[0] = step as f64; - row[a_off] = a_vals[step]; - row[b_off..b_off + n_slots].copy_from_slice(&b_vals[step][..n_slots]); + fn discovery_style_series_normalize_through_the_same_functions() { + let scores_a: Vec<(f64, f64)> = vec![(0.0, 1.0), (1.0, -4.0), (2.0, 0.0)]; + let scores_b: Vec<(f64, f64)> = vec![(0.0, 3.0), (1.0, 4.0), (2.0, 0.0)]; + let groups = [ + NormGroup::for_member(Some(0), 0), + NormGroup::for_member(Some(0), 1), + ]; + fn score(pair: &(f64, f64)) -> f64 { + pair.1 } - let mut offsets: HashMap, usize> = HashMap::new(); - offsets.insert(Ident::new("time"), 0); - offsets.insert(loop_score_ident("A"), a_off); - offsets.insert(loop_score_ident("B"), b_off); - - let sim_specs = crate::datamodel::SimSpecs { - start: 0.0, - stop: (step_count - 1) as f64, - dt: crate::datamodel::Dt::Dt(1.0), - save_step: None, - sim_method: crate::datamodel::SimMethod::Euler, - time_units: None, - }; - let results = Results { - offsets, - data: data.into_boxed_slice(), - step_size, - step_count, - specs: crate::results::Specs::from(&sim_specs), - is_vensim: false, - }; - - // A is scalar (one slot in partition 0); B is A2A coupled (both - // slots in partition 0). A broadcasts its single value into both - // of B's slots, so A's series is padded to B's stride. - let partitions = mapping_per_slot(&[("A", vec![Some(0)]), ("B", vec![Some(0), Some(0)])]); - - let rel = compute_rel_loop_scores_per_element(&results, &partitions); - let a = rel.get("A").unwrap(); - let b = rel.get("B").unwrap(); - assert_eq!(a.len(), step_count * n_slots); - assert_eq!(b.len(), step_count * n_slots); - - let at = |step: usize, k: usize| step * n_slots + k; - - // Element 0: denom t0 = |2| + |3| = 5; denom t1 = |1| + |1| = 2. - assert!((a[at(0, 0)] - (2.0 / 5.0)).abs() < 1e-12); - assert!((b[at(0, 0)] - (3.0 / 5.0)).abs() < 1e-12); - assert!((a[at(1, 0)] - (1.0 / 2.0)).abs() < 1e-12); - assert!((b[at(1, 0)] - (1.0 / 2.0)).abs() < 1e-12); + let totals = group_totals( + [ + (groups[0], scores_a.iter().map(score)), + (groups[1], scores_b.iter().map(score)), + ], + 3, + ); + let rel_a = relative_series(scores_a.iter().map(score), &totals[&groups[0]]); + let rel_b = relative_series(scores_b.iter().map(score), &totals[&groups[1]]); - // Element 1: scalar A broadcasts its slot-0 value. denom t0 = - // |2| + |6| = 8; denom t1 = |1| + |4| = 5. This is the - // property that the scalar-only helpers cannot express. - assert!((a[at(0, 1)] - (2.0 / 8.0)).abs() < 1e-12); - assert!((b[at(0, 1)] - (6.0 / 8.0)).abs() < 1e-12); - assert!((a[at(1, 1)] - (1.0 / 5.0)).abs() < 1e-12); - assert!((b[at(1, 1)] - (4.0 / 5.0)).abs() < 1e-12); + let results = make_results_for_loops(&[("A", &[1.0, -4.0, 0.0]), ("B", &[3.0, 4.0, 0.0])]); + let owner = compute_rel_loop_scores(&results, &mapping(&[("A", Some(0)), ("B", Some(0))])); + assert_eq!(&rel_a, owner.get("A").unwrap()); + assert_eq!(&rel_b, owner.get("B").unwrap()); + assert_eq!(rel_a, vec![0.25, -0.5, 0.0]); + assert_eq!(rel_b, vec![0.75, 0.5, 0.0]); } - /// `aggregate_per_element_argmax_abs`: pure-arrayed sanity check. - /// Single arrayed loop with 3 elements, stride==n; output should - /// be argmax-abs across the 3 elements at each step. + /// `argmax_abs_by_step` on a 3-slot series: the slot with the largest + /// magnitude wins each step, sign preserved. #[test] fn aggregate_pure_arrayed_argmax_abs() { let mut per_elem = HashMap::new(); @@ -2027,87 +1449,30 @@ mod tests { // step 0: [0.1, 0.5, -0.2] -> argmax-abs picks 0.5 // step 1: [0.3, -0.4, 0.0] -> argmax-abs picks -0.4 per_elem.insert("L".to_string(), vec![0.1, 0.5, -0.2, 0.3, -0.4, 0.0]); - let mut n_slots = HashMap::new(); - n_slots.insert("L".to_string(), 3); - let out = aggregate_per_element_argmax_abs(&per_elem, &n_slots, 2); + let out = aggregate_per_element_argmax_abs(&per_elem, 2); let agg = out.get("L").expect("L must have aggregate"); assert_eq!(agg, &vec![0.5, -0.4]); } - /// `aggregate_per_element_argmax_abs`: scalar in mixed partition. - /// The helper input has stride=3 (partition max) but n=1 for the - /// scalar loop. Output should have length step_count, with each - /// value taken from the loop's own slot 0 (canonical scalar view). - /// Pre-fix the layout used the wrong stride and produced a series - /// of length step_count*3 with misaligned values. - #[test] - fn aggregate_scalar_in_mixed_partition_returns_step_count_values() { - let mut per_elem = HashMap::new(); - // Scalar A in mixed partition (stride=3); 4 steps × 3 elements. - // Each step has the same value at index 0 (slot 0 broadcast), - // but indices 1,2 carry the per-element rel-scores from - // distinct partition denominators. - // step 0: [0.10, 0.20, 0.30] - // step 1: [0.15, 0.25, 0.35] - // step 2: [0.12, 0.22, 0.32] - // step 3: [0.18, 0.28, 0.38] - per_elem.insert( - "A".to_string(), - vec![ - 0.10, 0.20, 0.30, // step 0 - 0.15, 0.25, 0.35, // step 1 - 0.12, 0.22, 0.32, // step 2 - 0.18, 0.28, 0.38, // step 3 - ], - ); - let mut n_slots = HashMap::new(); - n_slots.insert("A".to_string(), 1); - - let out = aggregate_per_element_argmax_abs(&per_elem, &n_slots, 4); - let agg = out.get("A").expect("A must have aggregate"); - // Length must match step_count, NOT step_count*stride. Each - // value is the scalar's own slot 0 at that step. - assert_eq!(agg.len(), 4); - assert_eq!(agg, &vec![0.10, 0.15, 0.12, 0.18]); - } - - /// `aggregate_per_element_argmax_abs`: the recovered stride and the - /// mapped `n_slots` need not agree. Post-Phase-2 - /// `compute_rel_loop_scores_per_element` lays an arrayed loop out at - /// `stride == n_slots`, but the aggregator is defensive: if a caller - /// supplies a series whose recovered stride (5 here) exceeds the - /// loop's mapped `n_slots` (2) -- e.g. a partially-snapshotted - /// `n_slots_by_loop` -- the argmax-abs iterates only the mapped 2 - /// slots, never the trailing padding. + /// A scalar loop's series (stride 1) passes through unchanged. #[test] - fn aggregate_arrayed_in_mixed_partition_iterates_own_slots_only() { + fn aggregate_scalar_series_is_identity() { let mut per_elem = HashMap::new(); - // 2 steps × 5-stride. Loop is mapped to n=2 own slots. - // Positions 2..5 are stale/padding and must NOT be included in - // argmax-abs. To prove the iterator stops at n=2, place a trap - // value (1e6) at position 4 -- if the helper ever iterates - // 0..stride it would pick this and we'd notice. - per_elem.insert( - "B".to_string(), - vec![ - 0.1, -0.3, 0.0, 0.0, 1.0e6, // step 0; trap at index 4 - 0.4, 0.2, 0.0, 0.0, 1.0e6, // step 1; trap at index 4 - ], - ); - let mut n_slots = HashMap::new(); - n_slots.insert("B".to_string(), 2); + per_elem.insert("A".to_string(), vec![0.10, 0.15, -0.12, 0.18]); + let out = aggregate_per_element_argmax_abs(&per_elem, 4); + assert_eq!(out.get("A").unwrap(), &vec![0.10, 0.15, -0.12, 0.18]); + } - let out = aggregate_per_element_argmax_abs(&per_elem, &n_slots, 2); - let agg = out.get("B").expect("B must have aggregate"); - assert_eq!(agg.len(), 2); - // step 0: argmax-abs over [0.1, -0.3] = -0.3 (sign preserved). - // step 1: argmax-abs over [0.4, 0.2] = 0.4. - assert_eq!(agg, &vec![-0.3, 0.4]); + /// Ties (two elements with equal |rel|) are broken deterministically: + /// the lowest slot index wins. + #[test] + fn argmax_abs_ties_broken_by_lowest_index() { + assert_eq!(argmax_abs_by_step(&[0.4, -0.4], 1), vec![0.4]); } - /// Non-finite values (NaN, ±Inf) get mapped to 0.0 in the output, - /// matching the existing layout filter behavior. + /// Non-finite values (NaN, ±Inf) map to 0.0 in the output; a NaN never + /// displaces a finite candidate. #[test] fn aggregate_filters_non_finite() { let mut per_elem = HashMap::new(); @@ -2115,569 +1480,181 @@ mod tests { "L".to_string(), vec![f64::NAN, 0.5, f64::INFINITY, -f64::INFINITY], ); - let mut n_slots = HashMap::new(); - n_slots.insert("L".to_string(), 2); - let out = aggregate_per_element_argmax_abs(&per_elem, &n_slots, 2); + let out = aggregate_per_element_argmax_abs(&per_elem, 2); let agg = out.get("L").expect("L must have aggregate"); - // step 0: [NaN, 0.5]. argmax-abs comparison with NaN is false, - // so 0.5 wins. Output is finite (0.5). - // step 1: [Inf, -Inf]. Both non-finite; output 0.0. - assert_eq!(agg.len(), 2); - assert_eq!(agg[0], 0.5); - assert_eq!(agg[1], 0.0); + // step 0: [NaN, 0.5] -> 0.5 (NaN compares false). + // step 1: [Inf, -Inf] -> both non-finite -> 0.0. + assert_eq!(agg, &vec![0.5, 0.0]); } - /// An empty per-loop series yields an empty aggregate (matches - /// the existing layout's "no series" branch). + /// An empty per-loop series yields an empty aggregate (the layout's + /// "no series" branch). #[test] fn aggregate_empty_series_yields_empty_output() { let mut per_elem = HashMap::new(); per_elem.insert("L".to_string(), Vec::new()); - let mut n_slots = HashMap::new(); - n_slots.insert("L".to_string(), 1); - - let out = aggregate_per_element_argmax_abs(&per_elem, &n_slots, 5); - let agg = out.get("L").expect("L must have aggregate (even if empty)"); - assert!(agg.is_empty()); + let out = aggregate_per_element_argmax_abs(&per_elem, 5); + assert!(out.get("L").unwrap().is_empty()); } - /// Mismatched-dim arrayed partition: loop A has n=2 and loop B has - /// n=3 in the same partition. At partition element k=2, A has no - /// own element; the helper must NOT OOB-read past A's allocated - /// slots into B's data, and A's series stays its own length (n=2) - /// rather than being padded to B's. - /// - /// We pick B's slot 0 as a sentinel (999.0) so that an OOB read of - /// `row[off_A + 2]` (which equals `row[off_B + 0]` in our layout) - /// pulls this clearly-wrong value -- the test fails loudly rather - /// than passing by accident on whatever uninitialised data the - /// allocator happens to return. - #[test] - fn per_element_helper_handles_arrayed_with_smaller_n() { - // Layout: time | A slots 0..2 | B slots 0..3 - // step 0: | 1.0 2.0 | 999.0 20.0 30.0 - let loop_data = vec![vec![vec![1.0, 2.0]], vec![vec![999.0, 20.0, 30.0]]]; - let results = make_arrayed_results(&["A", "B"], &[2, 3], &loop_data); + proptest! { + #![proptest_config(ProptestConfig::with_cases(128))] - // Both coupled: A's two slots and B's three slots all in partition 0. - let partitions = mapping_per_slot(&[ - ("A", vec![Some(0), Some(0)]), - ("B", vec![Some(0), Some(0), Some(0)]), - ]); + /// `compute_rel_loop_scores` must match the naive per-member + /// reference for arbitrary per-slot partition vectors -- coupled + /// (all entries the same `Some(p)`), uncoupled (distinct `Some(p)` + /// per slot), `None`-laced, and scalar -- mixed across loops of + /// different widths so slots of different loops really do share + /// partitions. The one-pass accumulation and the scan-every-member + /// reference are computed independently, so any divergence in the + /// grouping, the Solo rule, or the SAFEDIV-0 handling shows up here. + /// + /// Per-loop `spec[i] = (kind, len, base, vals)` builds the + /// partition vector: + /// - kind 0: scalar `[Some(base)]`. + /// - kind 1: scalar `[None]`. + /// - kind 2: coupled arrayed `[Some(base); len]`. + /// - kind 3: uncoupled arrayed `[Some(base), Some(base+1), ...]` + /// (distinct consecutive partitions). + /// - kind 4: `None`-laced arrayed -- `vals[k]` chooses + /// `Some(vals[k])` or `None` per slot: the vector + /// `partition_for_loop` returns for an A2A loop some of whose + /// slots resolve to no parent-level stock partition (a slot + /// whose only state is module-internal), the rest to + /// arbitrary partitions. + /// Partition indices stay in a small pool (so coupling across + /// *different* loops actually happens); lengths stay tiny so 128 + /// cases run in well under a second on a debug build. + #[test] + fn rel_loop_scores_match_naive_reference( + specs in prop::collection::vec( + ( + 0usize..=4, // kind + 1usize..=3, // arrayed length + 0usize..=3, // base partition + prop::collection::vec(0usize..=4, 3), // per-slot None/Some chooser (>=4 => None) + ), + 1..=4, + ), + num_steps in 1usize..=4, + // Flat pool of loop_score samples; sliced per (loop, slot, step). + raw_vals in prop::collection::vec(-50.0_f64..=50.0_f64, 1..=300), + ) { + let n = specs.len(); + let loop_ids: Vec = (0..n).map(|i| format!("L{i}")).collect(); - let rel = compute_rel_loop_scores_per_element(&results, &partitions); - let a = rel.get("A").expect("A must have a series"); - let b = rel.get("B").expect("B must have a series"); - // Each loop's series has its own slot count -- A's is 2, not B's 3. - assert_eq!(a.len(), 2); - assert_eq!(b.len(), 3); - - // k=0,1: bucket (0, k) = {A.slotk, B.slotk}. - // denom_0 = |1| + |999| = 1000; a[0] = 1/1000, b[0] = 999/1000. - // denom_1 = |2| + |20| = 22; a[1] = 2/22, b[1] = 20/22. - assert!((a[0] - 1.0 / 1000.0).abs() < 1e-12); - assert!((b[0] - 999.0 / 1000.0).abs() < 1e-12); - assert!((a[1] - 2.0 / 22.0).abs() < 1e-12); - assert!((b[1] - 20.0 / 22.0).abs() < 1e-12); - - // k=2: bucket (0, 2) = {B.slot2} only -- A has no slot 2, so it's - // not in any slot-2 bucket and doesn't OOB-read into B's data. - // denom_2 = |B[2]| = 30 -> b[2] = 30/30 = 1.0. - assert!( - (b[2] - 1.0).abs() < 1e-12, - "B's slot 2 should normalise against itself only (A has no slot 2); got {}", - b[2] - ); - } + // Materialize each loop's per-slot partition vector. + let slots: Vec>> = specs + .iter() + .map(|(kind, len, base, vals)| match kind { + 0 => vec![Some(*base)], + 1 => vec![None], + 2 => vec![Some(*base); *len], + 3 => (0..*len).map(|k| Some(*base + k)).collect(), + _ => (0..*len) + .map(|k| { + let v = vals[k % vals.len()]; + if v >= 4 { None } else { Some(v) } + }) + .collect(), + }) + .collect(); + let n_slots: Vec = slots.iter().map(|v| v.len().max(1)).collect(); - /// Build a `Results` with each loop occupying a configurable number of - /// slots. Layout: `time | loop0 slot 0..n0 | loop1 slot 0..n1 | ...`. - /// `loop_data[i][step][slot]` is the value at (step, slot) for loop i. - fn make_arrayed_results( - loop_ids: &[&str], - slots_per_loop: &[usize], - loop_data: &[Vec>], - ) -> Results { - assert_eq!(loop_ids.len(), slots_per_loop.len()); - assert_eq!(loop_ids.len(), loop_data.len()); - let step_count = loop_data[0].len(); - for d in loop_data.iter() { - assert_eq!(d.len(), step_count); - } - let total_slots: usize = slots_per_loop.iter().sum(); - let step_size = 1 + total_slots; - let mut data = vec![0.0_f64; step_count * step_size]; - let mut offsets: HashMap, usize> = HashMap::new(); - offsets.insert(Ident::new("time"), 0); - let mut cursor = 1; - let mut loop_offsets = Vec::with_capacity(loop_ids.len()); - for (i, id) in loop_ids.iter().enumerate() { - offsets.insert(loop_score_ident(id), cursor); - loop_offsets.push(cursor); - cursor += slots_per_loop[i]; - } - for step in 0..step_count { - let row = &mut data[step * step_size..(step + 1) * step_size]; - row[0] = step as f64; - for (i, &off) in loop_offsets.iter().enumerate() { - let slots = &loop_data[i][step]; - assert_eq!(slots.len(), slots_per_loop[i]); - for (slot, &v) in slots.iter().enumerate() { - row[off + slot] = v; - } - } - } - let sim_specs = SimSpecs { - start: 0.0, - stop: (step_count.saturating_sub(1)) as f64, - dt: Dt::Dt(1.0), - save_step: None, - sim_method: SimMethod::Euler, - time_units: None, - }; - Results { - offsets, - data: data.into_boxed_slice(), - step_size, - step_count, - specs: Specs::from(&sim_specs), - is_vensim: false, - } - } - - /// Per-element streaming partition denominator must read the queried - /// element from each member loop, NOT slot 0. Two A2A loops with - /// distinct per-slot values: the denominator at element 1 must equal - /// `|loop0[t, 1]| + |loop1[t, 1]|`, which differs from element 0's. - #[test] - fn per_element_partition_denominator_reads_queried_slot() { - // 2 loops, 3 slots each, 2 timesteps. - // loop0[step][slot] and loop1[step][slot] chosen so element-1 values - // differ visibly from element-0 values. - let loop_data = vec![ - // loop0: [step][slot] - vec![ - vec![1.0, 7.0, 2.0], // step 0: slots 0,1,2 - vec![2.0, 9.0, 3.0], // step 1 - ], - // loop1 - vec![vec![4.0, 5.0, 6.0], vec![8.0, 11.0, 12.0]], - ]; - let results = make_arrayed_results(&["A", "B"], &[3, 3], &loop_data); - - let denom_e0 = compute_partition_denominator_for_element( - &results, - [("A", 3_usize), ("B", 3_usize)], - 0, - ); - let denom_e1 = compute_partition_denominator_for_element( - &results, - [("A", 3_usize), ("B", 3_usize)], - 1, - ); - - // Element 0: |1|+|4|=5, |2|+|8|=10. - assert_eq!(denom_e0, vec![5.0, 10.0]); - // Element 1: |7|+|5|=12, |9|+|11|=20. - assert_eq!(denom_e1, vec![12.0, 20.0]); - // Sanity: element 1 is genuinely different from element 0. - assert_ne!(denom_e0, denom_e1); - } - - /// Scalar loops in mixed partitions must broadcast slot 0 regardless - /// of the queried element_index. This matches the pre-PR compile-time - /// emitter that expanded a scalar loop_score reference across an - /// arrayed rel_loop_score target. - #[test] - fn per_element_partition_denominator_broadcasts_scalar() { - // Loop A is scalar (1 slot); loop B is A2A (3 slots). - let loop_data = vec![ - vec![vec![3.0], vec![5.0]], // A scalar - vec![vec![1.0, 7.0, 2.0], vec![2.0, 9.0, 3.0]], // B arrayed - ]; - let results = make_arrayed_results(&["A", "B"], &[1, 3], &loop_data); - - // Element 0 query: A contributes |3|=3 (its only slot), B contributes |1|=1. - let denom_e0 = compute_partition_denominator_for_element( - &results, - [("A", 1_usize), ("B", 3_usize)], - 0, - ); - assert_eq!(denom_e0, vec![3.0 + 1.0, 5.0 + 2.0]); - - // Element 1 query: A still contributes |3|=3 (broadcast), B contributes |7|=7. - let denom_e1 = compute_partition_denominator_for_element( - &results, - [("A", 1_usize), ("B", 3_usize)], - 1, - ); - assert_eq!(denom_e1, vec![3.0 + 7.0, 5.0 + 9.0]); - } + // Build per-(loop, step, slot) loop_score data from the flat + // pool, advancing a single cursor so successive slots get + // distinct samples. `series[i][step][slot]`. + let mut cursor = 0usize; + let mut series: Vec>> = Vec::with_capacity(n); + for &ns in &n_slots { + let mut per_step = Vec::with_capacity(num_steps); + for _ in 0..num_steps { + let mut per_slot = Vec::with_capacity(ns); + for _ in 0..ns { + per_slot.push(raw_vals[cursor % raw_vals.len()]); + cursor += 1; + } + per_step.push(per_slot); + } + series.push(per_step); + } - /// `compute_rel_loop_score_for_element` paired with - /// `compute_partition_denominator_for_element` must reproduce the - /// per-element view that the full-sweep - /// `compute_rel_loop_scores_per_element` produces. Bit-for-bit - /// agreement is the contract the libsimlin per-partition cache - /// relies on -- the streaming pair is meant to be a strictly cheaper - /// path to the same numbers, not an approximation. - #[test] - fn per_element_streaming_matches_full_sweep() { - let n_slots: usize = 3; - let step_count: usize = 4; - // Reuse the fixture from `per_element_helper_normalizes_within_each_slot`: - // A: row[a_off + k] = (step+1) * (k+1) - // B: row[b_off + k] = (step+1) * (k+2) - let mut a_data = Vec::with_capacity(step_count); - let mut b_data = Vec::with_capacity(step_count); - for step in 0..step_count { - let a_row: Vec = (0..n_slots) - .map(|k| ((step + 1) * (k + 1)) as f64) - .collect(); - let b_row: Vec = (0..n_slots) - .map(|k| ((step + 1) * (k + 2)) as f64) + let results = make_arrayed_results( + &loop_ids.iter().map(|s| s.as_str()).collect::>(), + &n_slots, + &series, + ); + let loop_partitions: IndexMap>> = loop_ids + .iter() + .zip(slots.iter()) + .map(|(id, v)| (id.clone(), v.clone())) .collect(); - a_data.push(a_row); - b_data.push(b_row); - } - let results = make_arrayed_results(&["A", "B"], &[n_slots, n_slots], &[a_data, b_data]); - - // Both A2A loops coupled (every slot in partition 0), so slot k of - // each lands in bucket (0, k) -- the streaming helper, called with - // both loops as members at element k, sums the same two slot-k - // values into the denominator. - let partitions = - mapping_per_slot(&[("A", vec![Some(0); n_slots]), ("B", vec![Some(0); n_slots])]); - let full = compute_rel_loop_scores_per_element(&results, &partitions); + let actual = compute_rel_loop_scores(&results, &loop_partitions); + let expected = reference_rel_loop_scores(&loop_ids, &slots, &series, num_steps); - for k in 0..n_slots { - let denom = compute_partition_denominator_for_element( - &results, - [("A", n_slots), ("B", n_slots)], - k, - ); - let rel_a = compute_rel_loop_score_for_element(&results, "A", n_slots, k, &denom) - .expect("A must have a series"); - let rel_b = compute_rel_loop_score_for_element(&results, "B", n_slots, k, &denom) - .expect("B must have a series"); - - for step in 0..step_count { - let full_idx = step * n_slots + k; - let full_a = full.get("A").unwrap()[full_idx]; - let full_b = full.get("B").unwrap()[full_idx]; - // Bit-for-bit: same arithmetic order, same rounding. - assert_eq!( - rel_a[step], full_a, - "loop A step {step} elem {k}: streaming {} vs full {}", - rel_a[step], full_a - ); - assert_eq!( - rel_b[step], full_b, - "loop B step {step} elem {k}: streaming {} vs full {}", - rel_b[step], full_b + for (i, id) in loop_ids.iter().enumerate() { + let a = actual.get(id).expect("every loop has a series"); + let e = &expected[i]; + prop_assert_eq!( + a.len(), + e.len(), + "loop {}: series length {} vs reference {}", + id, + a.len(), + e.len() ); + for (idx, (&av, &ev)) in a.iter().zip(e.iter()).enumerate() { + if av.is_nan() && ev.is_nan() { + continue; + } + // Both walk the members in ascending (loop, slot) order, + // so the result is bit-identical, not merely close. + prop_assert_eq!( + av, ev, + "loop {} flat-index {}: actual {} vs reference {}", id, idx, av, ev + ); + } } - } - } - - /// Mixed-stride parity: for two coupled A2A loops with different - /// `n_slots` sharing the same per-slot partition, the streaming pair - /// (`compute_partition_denominator_for_element` + - /// `compute_rel_loop_score_for_element`) must produce the same - /// per-element rel-scores as the full-sweep - /// `compute_rel_loop_scores_per_element`. This is the contract the - /// libsimlin FFI per-partition cache relies on -- the streaming - /// pair must be a strictly cheaper path to the same numbers. - /// - /// Each loop's full-sweep series has its own slot count (A: 3, B: 2); - /// at slot 2 only A is a member of bucket (0, 2), so the streaming - /// helper -- called with both loops as members at element 2 but B - /// gated out by `effective_slot(2, 2) == None` -- agrees. - #[test] - fn streaming_helpers_match_full_sweep_in_mixed_stride_partition() { - // A has n=3, B has n=2, both coupled (every slot in partition 0). - // Multi-step so we exercise more than one row; distinct-per-step - // values so any wrong-stride bug shows up loudly. - // step 0: A = [1.0, 2.0, 5.0], B = [10.0, 7.0] - // step 1: A = [1.5, 2.5, 6.0], B = [11.0, 8.0] - // step 2: A = [2.0, 3.0, 7.0], B = [12.0, 9.0] - let loop_data = vec![ - vec![ - vec![1.0, 2.0, 5.0], - vec![1.5, 2.5, 6.0], - vec![2.0, 3.0, 7.0], - ], - vec![vec![10.0, 7.0], vec![11.0, 8.0], vec![12.0, 9.0]], - ]; - let results = make_arrayed_results(&["A", "B"], &[3, 2], &loop_data); - let partitions = mapping_per_slot(&[ - ("A", vec![Some(0), Some(0), Some(0)]), - ("B", vec![Some(0), Some(0)]), - ]); - - let full = compute_rel_loop_scores_per_element(&results, &partitions); - let full_a = full.get("A").unwrap(); - let full_b = full.get("B").unwrap(); - // A's series is 3-strided, B's is 2-strided. - assert_eq!(full_a.len(), results.step_count * 3); - assert_eq!(full_b.len(), results.step_count * 2); - - // The members of bucket (0, k) for k in 0..3 are {A, B} for k<2, - // {A} for k==2; the streaming helper expresses this via - // `effective_slot(n_b, k)` returning None for B at k>=n_b. - for k in 0..3 { - let denom = compute_partition_denominator_for_element( - &results, - [("A", 3_usize), ("B", 2_usize)], - k, - ); - let rel_a = compute_rel_loop_score_for_element(&results, "A", 3, k, &denom) - .expect("A must have a series"); - for step in 0..results.step_count { - let full_v = full_a[step * 3 + k]; - assert_eq!( - rel_a[step], full_v, - "loop A step {step} elem {k}: streaming {} vs full {}", - rel_a[step], full_v - ); + // The partition identity: at every step where a group's total + // is finite and non-zero, the magnitudes of its members' shares + // sum to exactly 1. Groups are re-derived here from the + // partition vectors (a Solo member is its own group). + let mut group_of: Vec<(NormGroup, usize, usize)> = Vec::new(); + for (i, v) in slots.iter().enumerate() { + for k in 0..n_slots[i] { + let g = NormGroup::for_member(v.get(k).copied().flatten(), group_of.len()); + group_of.push((g, i, k)); + } } - if k < 2 { - let rel_b = compute_rel_loop_score_for_element(&results, "B", 2, k, &denom) - .expect("B must have a series"); - for step in 0..results.step_count { - let full_v = full_b[step * 2 + k]; - assert_eq!( - rel_b[step], full_v, - "loop B step {step} elem {k}: streaming {} vs full {}", - rel_b[step], full_v + let mut groups: HashMap> = HashMap::new(); + for &(g, i, k) in &group_of { + groups.entry(g).or_default().push((i, k)); + } + for members in groups.values() { + for step in 0..num_steps { + let total: f64 = members + .iter() + .map(|&(i, k)| series[i][step][k].abs()) + .sum(); + if total == 0.0 { + continue; + } + let share: f64 = members + .iter() + .map(|&(i, k)| actual[&loop_ids[i]][step * n_slots[i] + k].abs()) + .sum(); + prop_assert!( + (share - 1.0).abs() < 1e-9, + "step {}: partition shares sum to {} not 1", step, share ); } - } else { - // B has no slot k>=2; the streaming helper returns all-zeros - // for it and the full-sweep series doesn't have that index. - let rel_b = compute_rel_loop_score_for_element(&results, "B", 2, k, &denom) - .expect("B must have a series"); - assert!(rel_b.iter().all(|&v| v == 0.0)); } } } - /// Streaming `compute_partition_denominator_for_element` must skip - /// arrayed members at partition indices past their own n_slots, - /// matching the gating now applied by the full-sweep - /// `compute_rel_loop_scores_per_element`. Pre-fix the streaming - /// helper clamped to the loop's last slot via `effective_slot`, - /// which silently disagreed with the full-sweep helper for any - /// mixed-stride partition (one arrayed loop with `n_a` slots - /// sharing a partition with another arrayed loop with `n_b < n_a`). - /// We plant a sentinel at B's last slot so a clamp would pull - /// it loudly into the denom; the principled "skip" semantic - /// excludes B entirely at element 2. - #[test] - fn streaming_partition_denominator_skips_arrayed_loops_past_own_slots() { - // step_count = 1. A has n=3 with values [1, 2, 5]; B has n=2 - // with values [10, 999.0] (sentinel at slot 1). At partition - // element k=2, B has no slot -- the principled denom is - // |A[2]| = 5, NOT |A[2]| + |B[1]| = 5 + 999 = 1004. - let loop_data = vec![vec![vec![1.0, 2.0, 5.0]], vec![vec![10.0, 999.0]]]; - let results = make_arrayed_results(&["A", "B"], &[3, 2], &loop_data); - - let denom_at_2 = compute_partition_denominator_for_element( - &results, - [("A", 3_usize), ("B", 2_usize)], - 2, - ); - assert_eq!( - denom_at_2, - vec![5.0], - "B has no slot at partition index 2; its sentinel must NOT \ - pollute the denominator (skip, not clamp)" - ); - - // For sanity, k=0 and k=1 should include both members. - let denom_at_0 = compute_partition_denominator_for_element( - &results, - [("A", 3_usize), ("B", 2_usize)], - 0, - ); - assert_eq!(denom_at_0, vec![1.0 + 10.0]); - let denom_at_1 = compute_partition_denominator_for_element( - &results, - [("A", 3_usize), ("B", 2_usize)], - 1, - ); - assert_eq!(denom_at_1, vec![2.0 + 999.0]); - } - - /// `compute_rel_loop_score_for_element` queried at an element this - /// loop doesn't have (n=2, queried at k=2) must return all-zeros - /// rather than clamping to the loop's last slot. This matches the - /// full-sweep helper's "zero-fill at positions n..max_slots" rule - /// so any future caller that directly queries past a loop's range - /// gets the right answer. - #[test] - fn streaming_rel_score_returns_zeros_when_loop_has_no_own_element() { - // B is arrayed with n=2 and a sentinel value at slot 1. When - // queried at element_index=2 it has no own element; the - // result must be all-zeros, not the slot-1 sentinel rel-score. - let loop_data = vec![vec![vec![10.0, 999.0]]]; - let results = make_arrayed_results(&["B"], &[2], &loop_data); - - // Use a denom that would produce a clearly-wrong rel-score if - // the helper clamped: |sentinel|/|denom| would be ~999, but - // the principled answer is 0.0. - let denom = vec![1.0_f64]; - let rel = compute_rel_loop_score_for_element(&results, "B", 2, 2, &denom) - .expect("B has a series"); - assert_eq!( - rel, - vec![0.0], - "queried at index 2 (past B's n=2), result must be 0 not clamped" - ); - } - - /// SAFEDIV-0 semantics propagate per element: a partition where every - /// member's queried slot is identically zero must yield 0 (not NaN) - /// for the rel-score, matching the scalar-helper contract. - #[test] - fn per_element_streaming_safediv_zero() { - // Both loops have slot 0 = 0 across all steps but slot 1 != 0; - // querying element 0 must yield 0 (no panic, no NaN). - let loop_data = vec![ - vec![vec![0.0, 5.0], vec![0.0, 4.0]], - vec![vec![0.0, 3.0], vec![0.0, 2.0]], - ]; - let results = make_arrayed_results(&["A", "B"], &[2, 2], &loop_data); - - let denom = compute_partition_denominator_for_element( - &results, - [("A", 2_usize), ("B", 2_usize)], - 0, - ); - assert_eq!(denom, vec![0.0, 0.0]); - - let rel = compute_rel_loop_score_for_element(&results, "A", 2, 0, &denom).unwrap(); - for v in rel { - assert_eq!(v, 0.0, "SAFEDIV-0 must yield 0, got {v}"); - } - } - - /// The streaming FFI denominator (`compute_partition_denominator_for_element`) - /// must exclude a `NaN` summand and keep an `Inf` one, exactly like the - /// full-sweep helper -- this is the path libsimlin's - /// `simlin_analyze_get_relative_loop_score` cache drives, so the - /// GH #542 fix must hold there too. - #[test] - fn per_element_streaming_denominator_excludes_nan_keeps_inf() { - // step 0: A slot0 = NaN, B slot0 = 3 -> denom = 3 (NaN dropped). - // step 1: A slot0 = Inf, B slot0 = 3 -> denom = Inf (Inf kept). - let loop_data = vec![ - vec![vec![f64::NAN, 7.0], vec![f64::INFINITY, 7.0]], - vec![vec![3.0, 1.0], vec![3.0, 1.0]], - ]; - let results = make_arrayed_results(&["A", "B"], &[2, 2], &loop_data); - - let denom = compute_partition_denominator_for_element( - &results, - [("A", 2_usize), ("B", 2_usize)], - 0, - ); - assert_eq!(denom[0], 3.0, "NaN summand excluded from streaming denom"); - assert_eq!( - denom[1], - f64::INFINITY, - "Inf summand retained in streaming denom" - ); - - // The healthy sibling B normalizes against the NaN-free denom at - // step 0 (3/3 = 1) and goes to 0 against the +Inf denom at step 1. - let rel_b = compute_rel_loop_score_for_element(&results, "B", 2, 0, &denom).unwrap(); - assert!( - (rel_b[0] - 1.0).abs() < 1e-12, - "healthy B not poisoned: {}", - rel_b[0] - ); - assert_eq!(rel_b[1], 0.0, "B dominated by +Inf sibling -> 0"); - } - - /// An absent loop_score variable returns `None`, matching the - /// "omit absent loops" contract of the full-sweep API. - #[test] - fn per_element_streaming_absent_loop_returns_none() { - let results = make_arrayed_results(&["A"], &[2], &[vec![vec![1.0, 2.0], vec![3.0, 4.0]]]); - let denom = compute_partition_denominator_for_element(&results, [("A", 2_usize)], 0); - assert!(compute_rel_loop_score_for_element(&results, "missing", 2, 0, &denom).is_none()); - } - - /// Signed argmax-abs aggregator: at each step, return the signed - /// rel-score of the element with the largest `|rel[k, t]|`. The - /// dominant element can switch between steps; the sign is preserved - /// from whichever element won that step. - #[test] - fn argmax_abs_picks_dominant_element_with_sign() { - // 2 elements, 2 steps. loop_score chosen so element 0 dominates - // at step 0 and element 1 dominates at step 1, with opposite signs. - // step 0: slot 0 = 5, slot 1 = -1 -> rel = 0.5, -0.1 - // step 1: slot 0 = 1, slot 1 = -8 -> rel = 0.1, -0.8 - let loop_data = vec![vec![vec![5.0, -1.0], vec![1.0, -8.0]]]; - let results = make_arrayed_results(&["L"], &[2], &loop_data); - - // Constant denom of 10 per step per element. - let denoms = vec![vec![10.0_f64; 2]; 2]; - let denom_refs: Vec<&[f64]> = denoms.iter().map(|d| d.as_slice()).collect(); - - let agg = compute_rel_loop_score_argmax_abs(&results, "L", 2, &denom_refs) - .expect("L must have a series"); - - // step 0: argmax-abs is slot 0 (|0.5| > |-0.1|), signed value = +0.5. - // step 1: argmax-abs is slot 1 (|-0.8| > |0.1|), signed value = -0.8. - assert_eq!(agg, vec![0.5, -0.8]); - } - - /// Scalar (n_slots == 1) reduces to identity: the aggregator returns - /// the single slot's series unchanged. - #[test] - fn argmax_abs_scalar_reduces_to_identity() { - let loop_data = vec![vec![vec![3.0], vec![-7.0]]]; - let results = make_arrayed_results(&["L"], &[1], &loop_data); - let denoms = [vec![10.0_f64, 10.0]]; - let denom_refs: Vec<&[f64]> = denoms.iter().map(|d| d.as_slice()).collect(); - - let agg = compute_rel_loop_score_argmax_abs(&results, "L", 1, &denom_refs) - .expect("L must have a series"); - assert_eq!(agg, vec![0.3, -0.7]); - } - - /// Ties (two elements with equal |rel|) are broken deterministically: - /// the lowest slot index wins. This matches Rust's stable - /// `Ord`/`max_by_key` convention for "first hit on equal". - #[test] - fn argmax_abs_ties_broken_by_lowest_index() { - // Both slots have equal magnitude at step 0; slot 0 has positive - // sign and slot 1 has negative sign. The output must be slot 0's - // value (+0.4), not slot 1's. - let loop_data = vec![vec![vec![4.0, -4.0]]]; - let results = make_arrayed_results(&["L"], &[2], &loop_data); - let denoms = [vec![10.0_f64], vec![10.0_f64]]; - let denom_refs: Vec<&[f64]> = denoms.iter().map(|d| d.as_slice()).collect(); - - let agg = compute_rel_loop_score_argmax_abs(&results, "L", 2, &denom_refs) - .expect("L must have a series"); - assert_eq!(agg, vec![0.4]); - } - - /// Absent loop returns `None`, matching the streaming-helper contract. - #[test] - fn argmax_abs_absent_loop_returns_none() { - let results = make_arrayed_results(&["L"], &[2], &[vec![vec![1.0, 2.0]]]); - let denoms = [vec![10.0_f64], vec![10.0]]; - let denom_refs: Vec<&[f64]> = denoms.iter().map(|d| d.as_slice()).collect(); - assert!(compute_rel_loop_score_argmax_abs(&results, "missing", 2, &denom_refs).is_none()); - } - /// `LoopElementIndex` for a scalar loop reports empty dimensions /// and `n_slots = 1`. Used by the libsimlin FFI dispatch to detect /// "this loop is not arrayed, reject subscripted IDs." @@ -2965,221 +1942,6 @@ mod tests { assert!(index.contains_key("r1")); } - /// SAFEDIV-0 propagates per element: an element with zero denom at a - /// given step contributes a 0 rel-score for that element at that step. - /// If every element's denom is zero, the aggregator returns 0. - #[test] - fn argmax_abs_safediv_zero_per_element() { - // step 0: both elements have denom 0 -> both rel = 0 -> agg = 0. - // step 1: element 0 has denom 0 (rel=0); element 1 has denom 10 - // and slot value -7 -> rel = -0.7, dominant. - let loop_data = vec![vec![vec![5.0, 5.0], vec![5.0, -7.0]]]; - let results = make_arrayed_results(&["L"], &[2], &loop_data); - let denoms = [ - vec![0.0_f64, 0.0], // element 0 - vec![0.0_f64, 10.0], // element 1 - ]; - let denom_refs: Vec<&[f64]> = denoms.iter().map(|d| d.as_slice()).collect(); - - let agg = compute_rel_loop_score_argmax_abs(&results, "L", 2, &denom_refs) - .expect("L must have a series"); - assert_eq!(agg.len(), 2); - // step 0: all elements safe-div to 0 -> agg = 0. - assert_eq!(agg[0], 0.0); - // step 1: only element 1 is non-zero, signed value -0.7. - assert!((agg[1] - (-0.7)).abs() < 1e-12, "got {}", agg[1]); - } - - proptest! { - #![proptest_config(ProptestConfig::with_cases(128))] - - /// For any small random model, `compute_rel_loop_scores` must - /// agree with the reference SAFEDIV formula to within 1e-10. - /// Generators: - /// - 1..=6 loops, assigned to 1..=3 partitions. - /// - 1..=10 timesteps. - /// - loop_score samples in [-100, 100]. - #[test] - fn matches_reference_formula( - num_loops in 1usize..=6, - num_partitions in 1usize..=3, - num_steps in 1usize..=10, - raw_values in prop::collection::vec( - prop::collection::vec(-100.0_f64..=100.0_f64, 1..=10), - 1..=6, - ), - raw_partitions in prop::collection::vec(0usize..=2, 1..=6), - ) { - let num_loops = num_loops.min(raw_values.len()).min(raw_partitions.len()); - let num_steps = num_steps.min(raw_values[0].len()); - let num_partitions = num_partitions.max(1); - - // Build per-loop series with uniform step count. - let series: Vec> = (0..num_loops) - .map(|i| raw_values[i].iter().copied().take(num_steps).collect()) - .collect(); - for s in &series { - prop_assume!(s.len() == num_steps); - } - - let loop_ids: Vec = (0..num_loops).map(|i| format!("L{i}")).collect(); - // Scalar-loop shape: one slot per loop (the slot-0 convenience - // view ignores anything beyond slot 0 anyway). - let loop_partitions: IndexMap>> = loop_ids - .iter() - .enumerate() - .map(|(i, id)| (id.clone(), vec![Some(raw_partitions[i] % num_partitions)])) - .collect(); - - // Build Results matching the series. - let pair_refs: Vec<(&str, &[f64])> = loop_ids - .iter() - .zip(series.iter()) - .map(|(id, s)| (id.as_str(), s.as_slice())) - .collect(); - let results = make_results_for_loops(&pair_refs); - - let scored = compute_rel_loop_scores(&results, &loop_partitions); - let expected = reference_rel_loop_scores(&loop_ids, &loop_partitions, &series); - - for (i, id) in loop_ids.iter().enumerate() { - let actual_series = scored.get(id).expect("every loop has a series"); - prop_assert_eq!(actual_series.len(), num_steps); - for t in 0..num_steps { - let a = actual_series[t]; - let e = expected[i][t]; - // Both NaN counts as a match (shouldn't occur given - // the finite generator range, but safeguard anyway). - if a.is_nan() && e.is_nan() { - continue; - } - prop_assert!( - (a - e).abs() <= 1e-10, - "loop {} t={}: actual={} expected={}", id, t, a, e - ); - } - } - } - - /// `compute_rel_loop_scores_per_element` must match the naive - /// per-`(partition, slot)` reference for arbitrary multi-slot - /// partition vectors -- coupled (all entries the same `Some(p)`), - /// uncoupled (distinct `Some(p)` per slot), `None`-laced, and - /// scalar. This is the regression net for the GH #487 bucket - /// grouping: the optimized BTreeMap-grid implementation and the - /// scan-all-loops reference are computed independently, so any - /// divergence in the broadcast stride, the slot gating, or the - /// SAFEDIV-0 handling shows up here. - /// - /// Per-loop `spec[i] = (kind, len, base, vals)` builds the - /// partition vector: - /// - kind 0: scalar `[Some(base)]`. - /// - kind 1: scalar `[None]`. - /// - kind 2: coupled arrayed `[Some(base); len]`. - /// - kind 3: uncoupled arrayed `[Some(base), Some(base+1), ...]` - /// (distinct consecutive partitions). - /// - kind 4: `None`-laced arrayed -- `vals[k]` chooses - /// `Some(vals[k])` or `None` per slot. - /// Partition indices stay in a small pool (so coupling across - /// *different* loops actually happens); lengths stay tiny so 128 - /// cases run in well under a second on a debug build. - #[test] - fn per_element_matches_naive_reference( - specs in prop::collection::vec( - ( - 0usize..=4, // kind - 1usize..=3, // arrayed length - 0usize..=3, // base partition - prop::collection::vec(0usize..=4, 3), // per-slot None/Some chooser (>=4 => None) - ), - 1..=4, - ), - num_steps in 1usize..=4, - // Flat pool of loop_score samples; sliced per (loop, slot, step). - // Includes 0.0 so the SAFEDIV-0 path is exercised. - raw_vals in prop::collection::vec(-50.0_f64..=50.0_f64, 1..=300), - ) { - let n = specs.len(); - let loop_ids: Vec = (0..n).map(|i| format!("L{i}")).collect(); - - // Materialize each loop's per-slot partition vector. - let slots: Vec>> = specs - .iter() - .map(|(kind, len, base, vals)| match kind { - 0 => vec![Some(*base)], - 1 => vec![None], - 2 => vec![Some(*base); *len], - 3 => (0..*len).map(|k| Some(*base + k)).collect(), - _ => (0..*len) - .map(|k| { - let v = vals[k % vals.len()]; - if v >= 4 { None } else { Some(v) } - }) - .collect(), - }) - .collect(); - let n_slots: Vec = slots.iter().map(|v| v.len().max(1)).collect(); - - // Build per-(loop, step, slot) loop_score data from the flat - // pool, advancing a single cursor so successive slots get - // distinct samples. `series[i][step][slot]`. - let mut cursor = 0usize; - let mut series: Vec>> = Vec::with_capacity(n); - for &ns in &n_slots { - let mut per_step = Vec::with_capacity(num_steps); - for _ in 0..num_steps { - let mut per_slot = Vec::with_capacity(ns); - for _ in 0..ns { - per_slot.push(raw_vals[cursor % raw_vals.len()]); - cursor += 1; - } - per_step.push(per_slot); - } - series.push(per_step); - } - - let results = make_arrayed_results( - &loop_ids.iter().map(|s| s.as_str()).collect::>(), - &n_slots, - &series, - ); - let loop_partitions: IndexMap>> = loop_ids - .iter() - .zip(slots.iter()) - .map(|(id, v)| (id.clone(), v.clone())) - .collect(); - - let actual = compute_rel_loop_scores_per_element(&results, &loop_partitions); - let expected = - reference_rel_loop_scores_per_element(&loop_ids, &slots, &series, num_steps); - - for (i, id) in loop_ids.iter().enumerate() { - let a = actual.get(id).expect("every loop has a series"); - let e = &expected[i]; - prop_assert_eq!( - a.len(), - e.len(), - "loop {}: series length {} vs reference {}", - id, - a.len(), - e.len() - ); - for (idx, (&av, &ev)) in a.iter().zip(e.iter()).enumerate() { - if av.is_nan() && ev.is_nan() { - continue; - } - // The two paths sum the same floats in the same order - // (sorted-loop-id member lists; ids are `L{i}`, i < 4), - // so the result is bit-identical, not merely close. - prop_assert_eq!( - av, ev, - "loop {} flat-index {}: actual {} vs reference {}", id, idx, av, ev - ); - } - } - } - } - // --- Raw per-element loop scores (GH #998) --------------------------- #[test] @@ -3234,7 +1996,7 @@ mod tests { #[test] fn raw_loop_score_out_of_range_element_is_zero_fill() { - // Matches the relative helper's convention (`effective_slot`): an + // `effective_slot`'s convention: an // ARRAYED loop queried past its own slot count yields zeros, not a // read of a neighboring loop's column -- while a SCALAR loop // broadcasts any element index to its single slot. diff --git a/src/simlin-engine/tests/integration/layout.rs b/src/simlin-engine/tests/integration/layout.rs index 87672afde..53c375060 100644 --- a/src/simlin-engine/tests/integration/layout.rs +++ b/src/simlin-engine/tests/integration/layout.rs @@ -815,12 +815,10 @@ fn test_constant_variable_has_no_deps_with_salsa() { /// Codex review regression (PR #472): every detected loop's /// `importance_series` must have length exactly `results.step_count`, -/// regardless of the partition stride that -/// `compute_rel_loop_scores_per_element` happens to write its output -/// at. Pre-fix, the layout divided `series.len()` by the loop's own -/// `n_slots` to derive `n_steps`, which silently produced -/// `step_count * stride`-long importance_series whenever the -/// helper's stride exceeded the loop's own slot count. +/// whatever slot count `compute_rel_loop_scores` lays the loop out with: +/// an arrayed loop's per-slot series is `step_count * n_slots` long and +/// the argmax-abs collapse must reduce it to one value per step, never +/// pass the per-slot layout through as an importance series. /// /// Note on coverage: at the time of writing, the engine's partition /// logic uses *element-level* stock SCCs in `model_element_cycle_partitions` @@ -988,11 +986,11 @@ fn test_arrayed_loop_importance_matches_argmax_abs_aggregation() { vm.run_to_end().unwrap(); let results = vm.into_results(); - // `compute_rel_loop_scores_per_element` derives each loop's slot count - // from `loop_partitions[id].len()`; `n_slots_by_loop` is still used below - // (and by `aggregate_per_element_argmax_abs` inside `compute_metadata`). - let per_elem = ltm_post::compute_rel_loop_scores_per_element(&results, &loop_partitions); - let slot0_only = ltm_post::compute_rel_loop_scores(&results, &loop_partitions); + // `compute_rel_loop_scores` lays each loop out with its own slot count + // (`loop_partitions[id].len()`); `n_slots_by_loop` recomputes it from + // the dimensions independently so the stride the test indexes with is + // not the owner's own. + let per_elem = ltm_post::compute_rel_loop_scores(&results, &loop_partitions); // (1) Contract: for every detected loop, importance_series must equal // the argmax-abs aggregation of the per-element series. @@ -1058,10 +1056,12 @@ fn test_arrayed_loop_importance_matches_argmax_abs_aggregation() { if n_slots_by_loop.get(&fl.name).copied().unwrap_or(1) <= 1 { return false; } - let slot0 = slot0_only.get(&fl.name).cloned().unwrap_or_default(); + let n = n_slots_by_loop.get(&fl.name).copied().unwrap_or(1).max(1); + let series = per_elem.get(&fl.name).cloned().unwrap_or_default(); + let slot0 = series.iter().copied().step_by(n); fl.importance_series .iter() - .zip(&slot0) + .zip(slot0) .any(|(a, s)| (a - s).abs() > 1e-9) }); assert!( diff --git a/src/simlin-engine/tests/integration/ltm_relative_scores.rs b/src/simlin-engine/tests/integration/ltm_relative_scores.rs new file mode 100644 index 000000000..7b2ee664d --- /dev/null +++ b/src/simlin-engine/tests/integration/ltm_relative_scores.rs @@ -0,0 +1,230 @@ +// Copyright 2026 The Simlin Authors. All rights reserved. +// Use of this source code is governed by the Apache License, +// Version 2.0, that can be found in the LICENSE file. + +//! Relative loop scores through the real pipeline, checked against hand +//! calculations on a fixture whose per-element loops share one cycle +//! partition. +//! +//! The owner of the normalization is `ltm_post::compute_rel_loop_scores`: +//! every `(loop, slot)` of a partition divides by the same sum of `|score|` +//! over all of the partition's members -- scalar loops and every slot of +//! every arrayed loop alike. These tests pin that rule on +//! `test/cross_element_ltm`, where the two stocks are coupled through +//! migration and so every loop of the model, arrayed or not, lives in one +//! partition. + +use std::collections::HashSet; +use std::fs::File; +use std::io::BufReader; + +use simlin_engine::db::{ + DetectedLoop, SimlinDb, compile_project_incremental, model_detected_loops, model_ltm_variables, + sync_from_datamodel_incremental, +}; +use simlin_engine::indexmap::IndexMap; +use simlin_engine::{Results, Vm, ltm_post, xmile}; + +fn load_xmile_model(path: &str) -> simlin_engine::datamodel::Project { + let f = File::open(path).unwrap_or_else(|e| panic!("failed to open {path}: {e}")); + let mut f = BufReader::new(f); + xmile::project_from_reader(&mut f) + .unwrap_or_else(|e| panic!("failed to parse XMILE from {path}: {e}")) +} + +/// Compile `project` with the LTM overlay, run it, and return the results +/// with the loop list and the per-slot partition map the same derivation +/// produced -- exactly what production feeds the owner. +fn simulate_with_ltm( + project: &simlin_engine::datamodel::Project, +) -> ( + Results, + Vec, + IndexMap>>, +) { + let mut db = SimlinDb::default(); + let sync = sync_from_datamodel_incremental(&mut db, project, None); + let compiled = + compile_project_incremental(&db, sync.project, "main", simlin_engine::db::LtmOverlay::On) + .expect("the fixture compiles with LTM"); + let source_model = sync.models["main"].source_model; + let loop_partitions = model_ltm_variables(&db, source_model, sync.project) + .loop_partitions + .clone(); + let loops = model_detected_loops(&db, source_model, sync.project) + .loops + .clone(); + let mut vm = Vm::new(compiled).unwrap(); + vm.run_to_end().expect("the fixture simulates"); + (vm.into_results(), loops, loop_partitions) +} + +/// The loop whose variable set is exactly `vars` (subscripts included). +fn loop_with_variables<'a>(loops: &'a [DetectedLoop], vars: &[&str]) -> &'a DetectedLoop { + let want: HashSet<&str> = vars.iter().copied().collect(); + loops + .iter() + .find(|l| { + l.variables + .iter() + .map(String::as_str) + .collect::>() + == want + }) + .unwrap_or_else(|| { + panic!( + "no loop over {vars:?}; loops: {:?}", + loops + .iter() + .map(|l| (l.id.as_str(), l.variables.clone())) + .collect::>() + ) + }) +} + +/// `test/cross_element_ltm` at dt = 1. Region = {NYC, Boston}; +/// `population[NYC]` starts at 1000, `population[Boston]` at 500. +/// +/// Hand calculation. Because `migration_in[NYC]` and `migration_out[NYC]` +/// are both `(pop[NYC] - pop[Boston]) * 0.01` while Boston's two migration +/// flows are clamped to 0 by `MAX`, each stock's net flow is exactly +/// `0.02 * pop[r]`: both populations grow at 2% per step forever, and so +/// does their difference. Every link score is therefore a ratio of +/// quantities that all scale by 1.02 per step, and the raw loop scores are +/// constant once active: +/// +/// - births loop, per element: `pop[r] -> births[r]` is `1` (linear), and +/// `births[r] -> pop[r]` is `Δbirths / Δnet = 1` (births IS the net flow): +/// raw `+1` at both slots. +/// - migration_out loop at NYC (`pop[NYC] -> mp[NYC] -> out[NYC] -> +/// pop[NYC]`): `Δ_{pop[NYC]} mp[NYC] / Δmp[NYC] = 0.01*Δpop[NYC] / +/// (0.01*(Δpop[NYC] - Δpop[Boston]))` = `20 / (20 - 10) = 2`; `mp -> out` +/// is `1`; `out -> pop` is `-(Δout / Δnet) = -(0.1 / 0.4) = -0.25`: raw +/// `-0.5`. At Boston every migration flow is 0, so that slot scores 0. +/// - the cross-element loop `pop[NYC] -> mp[Boston] -> in[NYC] -> pop[NYC]`: +/// `pop[NYC] -> mp[Boston]` is `-2` (same magnitudes, opposite sign), +/// `mp[Boston] -> in[NYC]` is `-1`, `in[NYC] -> pop[NYC]` is `+0.25`: +/// raw `+0.5`. +/// - every other loop crosses a Boston migration flow, whose link scores +/// are 0, so it scores 0. +/// +/// One partition holds both stocks, so the partition sum is +/// `1 + 1 + 0.5 + 0.5 = 3` and the shares are `1/3, 1/3, -1/6, +1/6`. +/// Grouping by `(partition, slot)` instead gives the births loop `1/2` at +/// NYC and `2/3` at Boston, and hands the scalar cross-element loop two +/// different series (`1/4` and `1/3`): none of those numbers is the share of +/// anything. +#[test] +fn cross_element_loops_normalize_over_the_whole_partition() { + let project = load_xmile_model("../../test/cross_element_ltm/cross_element.stmx"); + let (results, loops, loop_partitions) = simulate_with_ltm(&project); + let rel = ltm_post::compute_rel_loop_scores(&results, &loop_partitions); + + let births = loop_with_variables(&loops, &["population", "births"]); + let out_loop = loop_with_variables( + &loops, + &["population", "migration_pressure", "migration_out"], + ); + let cross = loop_with_variables( + &loops, + &[ + "population[nyc]", + "migration_pressure[boston]", + "migration_in[nyc]", + ], + ); + assert_eq!( + loop_partitions[&births.id].len(), + 2, + "births is A2A over Region" + ); + assert_eq!( + loop_partitions[&out_loop.id].len(), + 2, + "migration_out loop is A2A" + ); + assert_eq!( + loop_partitions[&cross.id].len(), + 1, + "the cross-element loop is scalar" + ); + let partition: HashSet> = loop_partitions.values().flatten().copied().collect(); + assert_eq!( + partition.len(), + 1, + "the coupled stocks put every slot of every loop in one partition: {loop_partitions:?}" + ); + + // Both flow-to-stock scores are defined from the second saved step on; + // the ratios are time-invariant, so any later step reads the same + // shares. Steps 2, 10 and 30 (t = 2, 10, 30 at dt = 1). + for step in [2usize, 10, 30] { + let at = |id: &str, n_slots: usize, k: usize| rel[id][step * n_slots + k]; + let close = |got: f64, want: f64, what: &str| { + assert!( + (got - want).abs() < 1e-9, + "step {step}: {what} = {got}, expected {want}" + ); + }; + close(at(&births.id, 2, 0), 1.0 / 3.0, "births[nyc]"); + close(at(&births.id, 2, 1), 1.0 / 3.0, "births[boston]"); + close( + at(&out_loop.id, 2, 0), + -1.0 / 6.0, + "migration_out loop[nyc]", + ); + close(at(&out_loop.id, 2, 1), 0.0, "migration_out loop[boston]"); + assert_eq!( + rel[&cross.id].len(), + results.step_count, + "a scalar loop has one series" + ); + close( + at(&cross.id, 1, 0), + 1.0 / 6.0, + "cross-element migration_in loop", + ); + + // The partition identity: the magnitudes over EVERY member -- each + // slot of each loop -- sum to 1, and the four members above are the + // only active ones. + let mut total = 0.0; + for (id, series) in &rel { + let n_slots = loop_partitions[id].len().max(1); + for k in 0..n_slots { + let v = series[step * n_slots + k]; + assert!(v.is_finite(), "step {step}: {id} slot {k} is {v}"); + total += v.abs(); + } + } + assert!( + (total - 1.0).abs() < 1e-9, + "step {step}: partition shares sum to {total}" + ); + } +} + +/// The bare-id aggregate of an arrayed loop is the argmax-abs across its +/// slots, taken from the same per-slot series: for the births loop both +/// slots read `1/3`, for the migration_out loop the NYC slot's `-1/6` wins +/// over Boston's `0`. +#[test] +fn cross_element_bare_id_aggregate_picks_the_dominant_slot() { + let project = load_xmile_model("../../test/cross_element_ltm/cross_element.stmx"); + let (results, loops, loop_partitions) = simulate_with_ltm(&project); + let rel = ltm_post::compute_rel_loop_scores(&results, &loop_partitions); + let agg = ltm_post::aggregate_per_element_argmax_abs(&rel, results.step_count); + + let births = loop_with_variables(&loops, &["population", "births"]); + let out_loop = loop_with_variables( + &loops, + &["population", "migration_pressure", "migration_out"], + ); + for step in [2usize, 10, 30] { + assert!((agg[&births.id][step] - 1.0 / 3.0).abs() < 1e-9); + assert!((agg[&out_loop.id][step] - (-1.0 / 6.0)).abs() < 1e-9); + } + for series in agg.values() { + assert_eq!(series.len(), results.step_count); + } +} diff --git a/src/simlin-engine/tests/integration/main.rs b/src/simlin-engine/tests/integration/main.rs index 725447ed1..f9ef6d2de 100644 --- a/src/simlin-engine/tests/integration/main.rs +++ b/src/simlin-engine/tests/integration/main.rs @@ -33,6 +33,7 @@ mod layout; mod ltm_array_agg; mod ltm_discovery_large_models; mod ltm_dt_invariance; +mod ltm_relative_scores; // Compares xmutil-based MDL parsing against the native Rust parser, so it // needs the optional xmutil C++ converter compiled in. #[cfg(feature = "xmutil")] diff --git a/src/simlin-engine/tests/integration/simulate_ltm.rs b/src/simlin-engine/tests/integration/simulate_ltm.rs index bb23a5c04..78f6af663 100644 --- a/src/simlin-engine/tests/integration/simulate_ltm.rs +++ b/src/simlin-engine/tests/integration/simulate_ltm.rs @@ -40,11 +40,9 @@ fn compile_ltm_incremental( } /// Compile with LTM enabled and capture the per-slot loop_partitions -/// mapping `compute_rel_loop_scores*` need to derive relative scores -/// post-sim. Since rel_loop_score is no longer emitted as a VM variable -/// (see docs/design-plans/2026-04-18-ltm-cap-lift-diagnosis.md), tests -/// that used to filter `results.offsets` for `$⁚ltm⁚rel_loop_score⁚{id}` -/// must now invoke `ltm_post::compute_rel_loop_scores(results, loop_partitions)`. +/// mapping `ltm_post::compute_rel_loop_scores` needs to derive relative +/// scores post-sim (relative scores are not VM variables; see +/// docs/design-plans/2026-04-18-ltm-cap-lift-diagnosis.md). fn compile_ltm_incremental_with_partitions( project: &simlin_engine::datamodel::Project, ) -> ( @@ -3926,18 +3924,6 @@ fn find_loop_score_offsets(results: &Results) -> Vec<(String, usize)> { entries } -/// Test helper: thin forwarder to the production per-element helper. -/// Retained so the existing A2A integration tests keep calling the -/// same name; they now pin the production code rather than a parallel -/// implementation. The per-slot `loop_partitions` carries each loop's -/// slot count (its `len()`), so no separate slot-count map is threaded. -fn compute_rel_loop_scores_per_element( - results: &Results, - loop_partitions: &IndexMap>>, -) -> HashMap> { - ltm_post::compute_rel_loop_scores_per_element(results, loop_partitions) -} - /// AC6.1 + AC6.4 + AC6.5: Pure A2A loop scores for an arrayed feedback model. /// /// Model: population[Region] (3 regions) with a reinforcing birth loop: @@ -4109,14 +4095,11 @@ fn test_a2a_two_loop_relative_scores_sum_to_100() { "Number of loop partitions should equal number of loop score vars" ); - // For each element, the absolute values of the per-element relative - // loop scores across all loops should sum to approximately 1.0. We - // use the per-element helper because the A2A case requires per-element - // normalization, while the scalar view (`ltm_post::compute_rel_loop_scores`) - // collapses to element 0. Both A2A loops pass through `population[r]`, - // so at each element their slots land in the same `(partition, slot)` - // bucket and self-normalize together. - let rel_per_element = compute_rel_loop_scores_per_element(&results, &loop_partitions); + // Both A2A loops pass through `population[r]` and the elements are + // uncoupled, so each element is its own partition holding exactly the + // two loops' slots for that element: the magnitudes of the two + // per-element relative scores sum to 1.0 there. + let rel_per_element = ltm_post::compute_rel_loop_scores(&results, &loop_partitions); for elem in 0..n_elements { // Pick a timestep late enough to have meaningful values (skip @@ -4228,7 +4211,7 @@ fn test_disconnected_a2a_loops_normalize_per_partition() { // Both subsystems are purely reinforcing, so every nonzero relative score is // exactly +1.0 -- NOT the pre-fix pooled value the two loops would share if // they cross-normalized. - let rel_per_element = compute_rel_loop_scores_per_element(&results, &loop_partitions); + let rel_per_element = ltm_post::compute_rel_loop_scores(&results, &loop_partitions); assert_eq!( rel_per_element.len(), 2, @@ -5291,7 +5274,7 @@ fn test_arrayed_population_ltm_exhaustive() { ); // This is a pure-A2A model over `Region`, so every loop has // `n_elements` slots and its rel-score series strides by `n_elements`. - let rel_per_element = compute_rel_loop_scores_per_element(&results, &loop_partitions); + let rel_per_element = ltm_post::compute_rel_loop_scores(&results, &loop_partitions); // Check that relative loop scores per element sum to ~1.0 at some // timestep after initialization. From 05e4dae43a046f55d30dc5036503d1c840a2e908 Mon Sep 17 00:00:00 2001 From: Bobby Powers Date: Mon, 7 Sep 2026 16:53:32 -0700 Subject: [PATCH 02/10] engine: score flow-to-stock links off a net-flow aux MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The flow-to-stock link score was the 2023 paper's Eq. 3 written over stock differences, `time_step * (PREVIOUS(flow) - PREVIOUS(PREVIOUS(flow)))` over the stock's second difference. Under Euler that is the paper's value one dt late, on a different window from every other link in the loop (GH #544); it breaks the paper's aggregation invariance at a fixed t; and it cost two nested-lag capture helpers per link and a two-step startup guard. Every stock with a scored flow-to-stock edge now gets a synthetic net-flow aux `$⁚ltm⁚net⁚{stock} = (inflows) - (outflows)`, shaped like the stock and minted beside the score by `shaped_link_score` (deduplicated by name), and the score is the ordinary instantaneous link score of that aux with respect to the flow: the paper's implementation option (b), section 4.4. The aux is a linear sum, so the partial is closed-form, `sign * |Δflow / Δnet|`, read over the same [t - dt, t] window as every other link, with no dt factor, no stock history, and a first value at the step after the start. The aux is not a causal node (it is in no loop and no link) and sorts ahead of every score in the evaluation order. A structural flow-to-stock edge has one emission owner: the Bare shape of the per-shape emitter, routed before the shape-driven emitters and without consulting the reference-site IR. That closes two wrong-number paths: a scalar flow into an arrayed stock took the scalar-to-arrayed emitter and was scored as a partial of the stock's initial-value equation (a stock's ast() is its initial); and a stock whose per-element initial reads its own flow had that initial's sites taken as the edge's shapes, minting one identical arrayed score per element, or, for a pinned-element read in a two-dimensional initial, diverting the edge to the per-element arm and dropping every loop through the stock. C-LEARN has two such scalar-flow edges (CH4 and CO2 emissions into their arrayed atmosphere stocks). Pinned in tests/integration/ltm_flow_to_stock.rs (new): the paper's table (in 5 -> 10, out 4 -> 5) scores 1.25 / -0.25 at the step of the change; a stock with separate nonlinear flows and its twin with one net flow agree on LS(flow -> stock) and on the relative loop scores to 1e-12 at the same t; births 0.1a / deaths a/20 split 2/3 : -1/3 from t = 1; a scalar flow into an arrayed stock scores 5/5.45 and 5/7.4 per element, the loop through it +1 with its drain loop at -2; an outflow-only stock scores exactly -1; a per-element initial reading its own flow yields one arrayed score with no warning; the net aux is absent from every loop and causal edge. The isolated-loop invariant holds at dt in {1, 0.5, 0.25, 0.125} with a one-step startup guard. The logistic golden keeps its +dt relabel: the shift is a labeling convention for every link score (the reference tool labels the score computed over [t, t + dt] with t), not a flow-to-stock timing effect, and the design and reference docs now say so. Re-baselined, each deliberately: the two LTM character goldens (only the flow-to-stock arms move, to the new guard form); the value-gate golden (growth -> pop is exactly 1 from t = 1, the net aux's slots appear); the two fragment goldens (fixture 10 gained a dynamic-index read so it still reaches the implicit-helper emission site, since the retired nested lag was its only helper -- the same reason five other fixtures now mint their helper through `PREVIOUS(weight[PREVIOUS(idx, idx)])`); the DELAY3 bound-port pin, 0, 1, 2, 4, 8 (was 0, 0, 1, 2, 4: the one-step relabel); and C-LEARN's LTM guardrail, 6,193 -> 6,224 variables (+35 net auxes, +2 arrayed scores, -6 per-element scalars) and 29,447 -> 28,725 slots (-888 nested-lag helper columns, +166 net-aux slots), derived in its rustdoc from `ltm_var_dump` and a column diff of `simlin simulate --ltm`. The Euler-only refusal keeps firing on flow-to-stock emission; its docs no longer cite the retired numerator as the reason. --- .../2026-09-04-link-scores-from-fragments.md | 14 +- docs/design/ltm--loops-that-matter.md | 104 +-- docs/reference/ltm--loops-that-matter.md | 65 +- docs/tech-debt.md | 2 +- src/simlin-engine/CLAUDE.md | 2 +- src/simlin-engine/src/common.rs | 4 +- src/simlin-engine/src/db/assemble_tests.rs | 25 +- src/simlin-engine/src/db/diagnostic.rs | 19 +- .../ltm_loop_discovery.txt | 523 +++++++++++---- .../ltm_loop_exhaustive.txt | 279 ++++---- .../src/db/fragment_char_tests.rs | 97 +-- .../src/db/fragment_input_tests.rs | 14 +- src/simlin-engine/src/db/ltm/compile.rs | 91 ++- src/simlin-engine/src/db/ltm/link_scores.rs | 145 +++-- src/simlin-engine/src/db/ltm/mod.rs | 77 +-- .../db/ltm_char_golden/agg_nested_reducer.txt | 2 +- .../ltm_char_golden/arrayed_agg_to_target.txt | 2 +- src/simlin-engine/src/db/ltm_tests.rs | 75 +-- src/simlin-engine/src/db/ltm_unified_tests.rs | 27 +- .../src/db/ltm_value_golden/value_gate.txt | 11 +- src/simlin-engine/src/db/prev_init_tests.rs | 18 +- src/simlin-engine/src/ltm_augment.rs | 259 +++++--- src/simlin-engine/src/ltm_augment_tests.rs | 264 +++++--- .../src/ltm_augment_with_lookup.rs | 3 +- .../integration/ltm_discovery_large_models.rs | 49 +- .../tests/integration/ltm_dt_invariance.rs | 121 ++-- .../tests/integration/ltm_flow_to_stock.rs | 605 ++++++++++++++++++ src/simlin-engine/tests/integration/main.rs | 1 + .../tests/integration/simulate.rs | 29 +- .../tests/integration/simulate_ltm.rs | 50 +- .../tests/integration/simulate_ltm_pinned.rs | 4 +- .../tests/integration/test_helpers.rs | 96 +++ 32 files changed, 2225 insertions(+), 852 deletions(-) create mode 100644 src/simlin-engine/tests/integration/ltm_flow_to_stock.rs diff --git a/docs/design-plans/2026-09-04-link-scores-from-fragments.md b/docs/design-plans/2026-09-04-link-scores-from-fragments.md index 33177fd51..f4685948e 100644 --- a/docs/design-plans/2026-09-04-link-scores-from-fragments.md +++ b/docs/design-plans/2026-09-04-link-scores-from-fragments.md @@ -28,9 +28,11 @@ variables. Because the rewrite is keyed per link on the target and its live range, the whole-model list becomes metadata (which links exist, their names and dimensions) while each link's program lives in its own memo, so an edit invalidates only the links that touch the edited variable. The families that -are not partials (flow-to-stock, black-box, aggregate, module composite and -loop scores) become typed builders over slot reads on the same footing, which -is what lets the text generators be deleted rather than merely bypassed. Two +are not partials (black-box, aggregate, module composite and loop scores) +become typed builders over slot reads on the same footing, which is what lets +the text generators be deleted rather than merely bypassed; a flow-to-stock +score is the ordinary partial of the stock's net-flow aux and needs no builder +of its own. Two run-time consequences follow from having a single place that emits a score: the 37-opcode guard scaffold collapses to one `LinkScore` opcode reading per-variable deltas computed once per step, and, on native and separably, the @@ -311,7 +313,7 @@ typed builder with no text: | family | today | this plan | |---|---|---| -| flow -> stock (second-order structural formula, `generate_flow_to_stock_equation`) | text | builder over `LoadVar`/`SymLoadPrev` of the flow and stock | +| flow -> stock (the partial of the stock's net-flow aux `$⁚ltm⁚net⁚{stock}`, `generate_flow_to_stock_equation` / `generate_net_flow_equation`) | text (closed form) | an ordinary link-score program over the net aux's fragment, live range the flow; the aux is a compiled fragment like any variable | | black-box unit transfer (`black_box_unit_transfer_equation`) | text | builder | | aggregate nodes (`$⁚ltm⁚agg⁚n`, `AggNode::reducer_expr0`) | typed `Expr0`, then parsed tiers | the reducer's own compiled fragment, live range per feeder | | module composites (`m·$⁚ltm⁚composite⁚port`) | text over pathway products | builder over the pathway link scores' slots | @@ -444,8 +446,8 @@ fragment or one of its per-element segments. ### Phase 4: The formula families, and no text anywhere -**Goal:** flow-to-stock, black-box, aggregate, composite and loop scores are -typed builders; the text generators, `LtmArm`, `LtmEquation` and +**Goal:** the net-flow auxes, black-box, aggregate, composite and loop scores +are typed builders; the text generators, `LtmArm`, `LtmEquation` and `model_ltm_implicit_var_info` are gone. **Components:** diff --git a/docs/design/ltm--loops-that-matter.md b/docs/design/ltm--loops-that-matter.md index 17cfcc63d..e8a6156f3 100644 --- a/docs/design/ltm--loops-that-matter.md +++ b/docs/design/ltm--loops-that-matter.md @@ -317,24 +317,46 @@ For a link from `x` to `z` where `z = f(x, y, ...)`: ### Flow-to-Stock Links -`generate_flow_to_stock_equation()` in `ltm_augment.rs`. - -Implements the corrected 2023 formula (Schoenberg et al., Eq. 3). The numerator -uses `PREVIOUS()` to align timing: at time t, `PREVIOUS(flow)` is the flow value -at t-1 that drove the stock change from t-1 to t. +`generate_flow_to_stock_equation()` and `generate_net_flow_equation()` in +`ltm_augment.rs`. + +A stock's flows reach it through its wiring, not through an equation, so the +score is built around a synthetic net-flow auxiliary: every stock with a scored +flow-to-stock edge gets `$⁚ltm⁚net⁚{stock} = (inflows) - (outflows)`, shaped +like the stock (`Equation::ApplyToAll` over an arrayed stock's dimensions; `0` +for a side with no flows), minted beside the score by `shaped_link_score` and +deduplicated by name (every flow of the stock mints the same aux). The score +for `flow -> stock` is the ordinary instantaneous link score of that aux with +respect to the flow. Because the aux is a linear sum its ceteris-paribus +partial is closed-form -- `Δ_flow net = +Δflow` for an inflow, `-Δflow` for an +outflow -- so the emitted equation is the standard guard form with that +numerator: ``` -numerator = PREVIOUS(flow) - PREVIOUS(PREVIOUS(flow)) -denominator = (stock - PREVIOUS(stock)) - (PREVIOUS(stock) - PREVIOUS(PREVIOUS(stock))) -link_score = sign * ABS(SAFEDIV(numerator, denominator, 0)) +if (TIME = INITIAL_TIME) then 0 +else if ((net - PREVIOUS(net)) = 0) OR ((flow - PREVIOUS(flow)) = 0) then 0 +else SAFEDIV(+/-(flow - PREVIOUS(flow)), ABS(net - PREVIOUS(net)), 0) * SIGN(flow - PREVIOUS(flow)) ``` -The denominator is the second-order change in the stock (its "acceleration"). -The ratio is wrapped in `ABS()` because flow-to-stock polarity is structural: -inflows always contribute positively (+1), outflows negatively (-1). The sign -is applied outside the absolute value. This equation returns 0 for the first -two timesteps (insufficient history for second-order differences), guarded by -`TIME = INITIAL_TIME` and `PREVIOUS(TIME, INITIAL_TIME) = INITIAL_TIME`. +which evaluates to `sign * |Δflow / Δnet|`: the 2023 paper's Eq. 3 (its +denominator `Δ(S_t) - Δ(S_{t-dt})` is `Δnet`) in the paper's own +implementation option (b) (section 4.4: aggregate the flows into a net flow, +score each flow into it, and let the net flow's link into the stock be 1). +Polarity is structural: inflows +1, outflows -1. Both deltas are read over +`[t - dt, t]`, the window of every other link score, so a loop's link scores +all describe one interval, the score is the same whether a stock's flows are +written separately or as one net flow, and no `dt` appears (an isolated loop +scores exactly `+/-1` at every `dt`, `tests/integration/ltm_dt_invariance.rs`). +Like every other score it is 0 at `TIME = INITIAL_TIME` and defined from the +first step after the start. The net aux is LTM machinery, not a causal node: +the causal graph keeps its `flow -> stock` edges, and the aux appears in no +loop and no link. + +A scalar flow into an arrayed stock broadcasts into every element's net flow, +so its score is one arrayed variable over the stock's dimensions +(`link_score_dimensions`). The structural edge is routed to the per-shape +emitter ahead of the shape-driven emitters (`emit_link_scores_for_edge`), which +would otherwise score it as a partial of the stock's initial-value equation. ### Stock-to-Flow Links @@ -1915,9 +1937,11 @@ enumerated loops. ### Euler Integration Only -The corrected flow-to-stock formula uses discrete differences that assume Euler -integration. The papers note compatibility with Runge-Kutta "in principle" but -this has not been explored in the implementation. +`assemble_simulation` refuses the overlay under RK2/RK4 (GH #486) when any +instantiated model emits a flow-to-stock score. The scores are differences of +saved-step values, so the guard keeps them on Euler-stepped trajectories; the +2020 paper (section 6.1) says the method is compatible with Runge-Kutta "in +principle", and Simlin has not established that for RK-stepped runs. ### Performance on Very Large Models @@ -2051,26 +2075,32 @@ cases remain deliberate carve-outs: synthesized compile-time equations, avoiding O(P^2) equation-text growth on models with very large same-partition loop sets (e.g. WRLD3). -7. **Flow-to-stock numerator timing and `time_step` scaling**: The flow-to-stock - link score numerator uses `time_step * (PREVIOUS(flow) - PREVIOUS(PREVIOUS(flow)))` - rather than the published bare `flow - PREVIOUS(flow)`. Two deliberate changes: - - *1-DT shift*: in Euler integration, the flow at t-1 drove the stock change - from t-1 to t, so `PREVIOUS(flow)` aligns the numerator and denominator to - the same causal interval. This produces results shifted by one DT compared - to reference SD software (Stella/iThink). The integration test - (`tests/integration/simulate_ltm.rs`) compensates by shifting reference - timestamps forward by DT when loading golden data. This convention is also - documented for end users in the reference doc's Section 3.2 (the "numerator - timing convention" note under - [Flow-to-Stock Link Score](../reference/ltm--loops-that-matter.md)). - - *`time_step` factor*: the denominator (second-order stock change) is - `dt * (netflow(t-dt) - netflow(t-2dt))` in Euler, carrying one `dt` that the - raw flow delta in the numerator lacks; without the factor every - flow-to-stock link score is `1/dt` too large and the error compounds once - per stock in a loop. The published Eq. 3 omits the factor because every - worked example in the papers uses dt=1. Verified empirically: at dt=0.25 - an isolated loop scores +1.0 with the factor (4.0 without), and at dt=1 - the paper's Table 3 values (1.25 / -0.25) reproduce exactly. +7. **Time labeling, and the flow-to-stock score's window** -- two separate + facts: + - *Time labeling, for every link score*: Simlin labels the score computed + over `[t - dt, t]` (its reading of `PREVIOUS`) with `t`; the + Stella-derived golden data in `test/logistic_growth_ltm` labels the score + computed over `[t, t + dt]` with `t`. The integration test + (`tests/integration/simulate_ltm.rs`) therefore shifts the reference + timestamps forward by one DT when loading. That fixture has one flow per + stock, so its flow-to-stock scores are identically 1 in either convention + and the whole shift sits in the instantaneous links: with the relabel the + relative loop scores agree to the file's rounding, and inverting + `|b1|/|r1| = (pop/1000)/(1 - pop/1000)` on each golden column reproduces + `pop` at the column's own `t`. The convention is documented for end users + in the reference doc's Section 3.2 note under + [Flow-to-Stock Link Score](../reference/ltm--loops-that-matter.md). + - *Flow-to-stock window*: the flow-to-stock score is the paper's + implementation option (b) -- the instantaneous score of the stock's + net-flow aux with respect to the flow, `sign * |Δflow / Δnet|` -- read + over the same `[t - dt, t]` window as every other link, with no `dt` + factor, no stock history and a first value at `start + dt`. The paper's + Eq. 3 is written over stock differences, each of which carries one `dt` + under Euler; expressing `Δnet` as a flow difference removes the factor + rather than compensating for it. At dt=1 the paper's table values + (1.25 / -0.25) reproduce at the step of the change + (`tests/integration/ltm_flow_to_stock.rs`), and an isolated loop scores + `+/-1` at every dt (`tests/integration/ltm_dt_invariance.rs`). 8. **Ceteris-paribus via AST transformation**: The papers describe re-evaluating equations with current values of one input and previous values of all others. diff --git a/docs/reference/ltm--loops-that-matter.md b/docs/reference/ltm--loops-that-matter.md index 7236c7cc6..55260d5ec 100644 --- a/docs/reference/ltm--loops-that-matter.md +++ b/docs/reference/ltm--loops-that-matter.md @@ -345,7 +345,9 @@ flow, not the flow's value. This makes it invariant to how flows are structurall specified: combining two separate flows into one net flow (a purely cosmetic choice with no mathematical effect on the model) produces identical LTM results. This ensures that analysis depends only on the model's mathematical behavior, not on the modeler's structural -presentation choices. +presentation choices. Simlin holds this at every time step: a stock with separate flows and +its twin with one net flow give the same flow-to-stock link scores and the same relative +loop scores at the same `t` (`tests/integration/ltm_flow_to_stock.rs`). This aggregation invariance is the key 2023 correction. Any value-based flow-to-stock formulation is deprecated and should not be used. @@ -379,42 +381,29 @@ instantaneous link score. Implementors can either: (a) use this formula directly disaggregated flows, or (b) automatically aggregate all flows into net flows and use link score 1 for all net-flow-to-stock links. Both produce identical results. -> **Simlin implementation note: numerator timing convention.** The discrete numerator -> `Delta(i)` has two valid renderings, and Simlin's choice produces output shifted by one -> DT relative to reference SD software (Stella/iThink) for the same model: +> **Simlin implementation note.** Simlin takes option (b): every stock with a scored +> flow-to-stock link gets a synthetic net-flow auxiliary `$⁚ltm⁚net⁚{stock} = inflows - +> outflows`, and the flow-to-stock link score is the ordinary instantaneous score of that +> auxiliary with respect to the flow -- `sign * |Delta(flow) / Delta(net)|`, the polarity +> structural (+1 inflow, -1 outflow). `Delta(net)` is the `Delta(S_t) - Delta(S_{t-dt})` +> above expressed as a flow difference, so no `dt` appears and the score reads the same +> `[t - dt, t]` window as every other link score in the model: a loop's link scores all +> describe one interval, the score is the same whether the stock's flows are written +> separately or as one net flow, and it is defined from the first step after the start like +> every other link (0 at `TIME = INITIAL_TIME`). > -> - `i(t) - i(t-dt)` -- the change in the *current* flow rate (the Stella convention). -> - `i(t-dt) - i(t-2dt)` -- the change in the flow over the interval that *drove* the -> measured stock change (the Simlin convention). -> -> Simlin uses the second form. The rationale is causal-interval alignment: under Euler -> integration the flow value at `t-dt` is what drove the stock change measured at `t` -> (`S(t) - S(t-dt) = dt * i(t-dt)`), so taking the change in *that* flow puts the numerator -> and the second-order stock-change denominator on the same causal interval. Concretely, -> Simlin generates the numerator as `time_step * (PREVIOUS(i) - PREVIOUS(PREVIOUS(i)))`, -> i.e. `i(t-dt) - i(t-2dt)` (the `time_step` factor is the separate dimensional -> correction discussed in the design doc, not part of the timing choice). -> -> Both renderings are mathematically valid and have identical continuous-time limits: as -> `dt -> 0` both converge to the Section 3.3 continuous ratio `|di/dt / d^2S/dt^2|` -> (Eq. 6 of Schoenberg, Hayward, and Eberlein 2023), which is -> timing-convention-agnostic. The discrete divergence is purely a one-DT phase shift of an -> otherwise numerically-identical series. -> -> **User-visible consequence.** A user comparing Simlin LTM output against Stella output -> for the same model and inputs will see numerically identical link/loop scores shifted by -> one DT. This is not a bug in either tool. To reconcile the two when comparing across -> tools, shift one set of timestamps by `+/-DT`: Simlin's value at time `t` corresponds to -> Stella's value at time `t-DT`. (Simlin's own LTM integration tests apply exactly this -> compensation when validating against reference golden data -- they shift the reference -> timestamps forward by DT before comparing.) -> -> This convention is documented per flow-to-stock link. A separate, distinct issue -> concerns window consistency *within* a loop's chain of links; that is tracked -> independently and is not the same as this numerator timing choice. See divergence #6, -> "Flow-to-stock numerator timing and `time_step` scaling," in -> [the design doc](../design/ltm--loops-that-matter.md) for the implementation details and -> the empirical verification. +> **Time labeling, for every link score.** Simlin labels the score computed over +> `[t - dt, t]` with `t` (its reading of `PREVIOUS`). Stella labels the score computed over +> `[t, t + dt]` with `t`: on the logistic-growth fixture (`test/logistic_growth_ltm`, one +> flow per stock, so its flow-to-stock scores are 1 in any convention) Stella's relative +> loop scores match Simlin's relabelled by `+dt` to the file's rounding, and inverting +> `|b1|/|r1| = (pop/1000)/(1 - pop/1000)` on each Stella column reproduces `pop` at that +> column's own `t`. A user comparing Simlin output against Stella output for the same model +> therefore sees the same link and loop scores one DT apart: Simlin's value at `t` is +> Stella's at `t - dt`. Simlin's LTM integration test applies exactly that shift to the +> reference timestamps. Whether Stella computes a multi-flow stock's flow-to-stock score +> over the same window as its other links is not verifiable from the data in this +> repository; the paper's aggregation-invariance argument requires it. ### 3.3 Continuous-Time Form @@ -905,7 +894,9 @@ ones, as demonstrated by the three-party arms race model (Section 12.2). - In principle compatible with Runge-Kutta and other integration methods - The flow-to-stock formula requires values from two previous timesteps (Delta(S_t) and Delta(S_{t-dt})), so link scores for flow-to-stock links are undefined for the first - two timesteps + two timesteps (Simlin: written as the flow difference `Delta(net)` of the stock's + net-flow auxiliary, the score needs one previous timestep and is defined from the first + step after the start, like every other link) ### 9.2 Equation Re-evaluation Cost diff --git a/docs/tech-debt.md b/docs/tech-debt.md index 02cb53872..94c57aa17 100644 --- a/docs/tech-debt.md +++ b/docs/tech-debt.md @@ -280,7 +280,7 @@ Known debt items consolidated from CLAUDE.md files and codebase analysis. Each e - **Component**: simlin-engine (src/simlin-engine/src/ltm_augment.rs flow-to-stock path) - **Severity**: medium -- **Description**: The 2023 flow-to-stock link-score formula assumes Euler integration: `PREVIOUS(flow) - PREVIOUS(PREVIOUS(flow))` aligns the numerator to the causal interval that drove the stock change from t-1 to t. Under RK2/RK4 this alignment breaks and link scores become mathematically nonsensical. Nothing currently prevents a user from setting `integration_method = RK4` and `ltm_enabled = true`; they'd get numbers that look plausible but are wrong. Fix: emit a compile-time diagnostic (preferably an Error) when LTM is enabled on a model whose sim specs select a non-Euler integrator. +- **Description**: LTM link scores are differences of saved-step values (a flow-to-stock score is `|Δflow / Δnet|` over the stock's net-flow aux), and whether they describe an RK-stepped trajectory faithfully has not been established, so `assemble_simulation` refuses LTM under RK2/RK4 (GH #486) when any instantiated model emits a flow-to-stock score. Open question: lift the refusal once the RK case is validated (the 2020 paper's section 6.1 calls the method compatible with Runge-Kutta "in principle"). - **Tracked in**: #486 (LTM tracking epic: #488) - **Owner**: unassigned - **Last reviewed**: 2026-04-29 diff --git a/src/simlin-engine/CLAUDE.md b/src/simlin-engine/CLAUDE.md index ba4caeee8..e4b96e4a0 100644 --- a/src/simlin-engine/CLAUDE.md +++ b/src/simlin-engine/CLAUDE.md @@ -135,7 +135,7 @@ LTM instruments a model with synthetic variables (link scores, loop scores, aggr - `equation.rs`: `LtmEquation`, the typed per-arm `Expr0` a synthetic variable carries (a text generator's arm is parsed once, `LtmArm::new`; a typed generator's arm is its tree, `LtmArm::from_typed`). `link_scores.rs`: the emitters, each parsing its arm once and reading the completeness guard off that arm; `shaped_link_score` is the per-shape entry. - `loops.rs` (including cross-aggregate loop recovery), `pinned.rs` (`LOOPSCORE` pins), `compile.rs` (`lower_ltm_variable`, the one lowering of a generated equation or a generated capture under the shapes of its reads; fragment compilation; the compile-failure diagnostic pass), `parse.rs`. - `LtmImplicitVarMeta` is captures only: a generated equation is built from an expanded tree and mints no module instance, so the module universe under LTM is the source models' own (`db::assemble::enumerate_module_instances`). -- `ltm_augment.rs` and its `ltm_augment_*.rs` children -- the ceteris-paribus equation generators: wrap every non-source reference of the target's equation in `PREVIOUS`, pin arrayed dependencies to the element the target reads (`subscript_idents_in_expr0`, on the tree), and freeze subscript indices. Production never parses a target equation; it lowers the target's `Expr2` back to `Expr0` (`patch::expr2_to_expr0`) and consults the IR by structural path, and a target with no lowered body is declined (`PartialEquationErrorKind::MissingTypedTarget`), never scored around user text or a `0` body. An aggregate node's reducer is the classified `BuiltinFn` (`AggNode::reducer`, projected by `reducer_expr0`); `AggNode::reducer_key` is its printed identity for the dedup, the IR's routing and the wrap's matching, never an equation to parse. +- `ltm_augment.rs` and its `ltm_augment_*.rs` children -- the ceteris-paribus equation generators: wrap every non-source reference of the target's equation in `PREVIOUS`, pin arrayed dependencies to the element the target reads (`subscript_idents_in_expr0`, on the tree), and freeze subscript indices. A flow-to-stock score is the closed-form partial of the stock's synthetic net-flow aux `$⁚ltm⁚net⁚{stock} = inflows - outflows` (`generate_net_flow_equation`, minted beside the score by `shaped_link_score` and deduplicated by name), `sign * |Δflow / Δnet|` over the same `[t - dt, t]` window as every other link; the aux is not a causal node. Production never parses a target equation; it lowers the target's `Expr2` back to `Expr0` (`patch::expr2_to_expr0`) and consults the IR by structural path, and a target with no lowered body is declined (`PartialEquationErrorKind::MissingTypedTarget`), never scored around user text or a `0` body. An aggregate node's reducer is the classified `BuiltinFn` (`AggNode::reducer`, projected by `reducer_expr0`); `AggNode::reducer_key` is its printed identity for the dedup, the IR's routing and the wrap's matching, never an equation to parse. - `ltm/` -- the vocabulary (`Link`, `Loop`, polarity in `types.rs`), `CausalGraph` (`graph.rs`), Johnson's circuit enumerator and Tarjan SCC (`indexed.rs`), cycle partitions (`partitions.rs`), static polarity analysis (`polarity.rs`). - `ltm_finding.rs` with `ltm_finding_enum.rs` and `ltm_finding_fallback.rs` -- post-simulation discovery: exact union-graph enumeration, a shortest-path sampling fallback under a wall-clock and memory budget, retention against the full loop universe, and the coverage-aware cap. - `ltm_post.rs`, `ltm_dominance.rs` -- relative loop and link scores (normalized within a cycle partition, in emission order so the IEEE sum is stable) and dominant-period selection. diff --git a/src/simlin-engine/src/common.rs b/src/simlin-engine/src/common.rs index 579b8a1f2..33bb424bb 100644 --- a/src/simlin-engine/src/common.rs +++ b/src/simlin-engine/src/common.rs @@ -490,8 +490,8 @@ pub enum ErrorCode { QueueOverflowNotOnQueue, /// LTM (Loops That Matter) analysis was requested on a model containing a /// queue. A queue is a stock with non-INTEG dynamics (a FIFO of batches), - /// so the flow-to-stock link-score numerator assumes plain INTEG under - /// Euler and any score touching the queue may be wrong. Emitted as a + /// so the flow-to-stock link score, which treats the stock's net flow as + /// its rate of change, and any score touching the queue may be wrong. Emitted as a /// Warning naming the queue, mirroring `ConveyorLtmDegraded` /// (docs/design/queues.md §10.5). QueueLtmDegraded, diff --git a/src/simlin-engine/src/db/assemble_tests.rs b/src/simlin-engine/src/db/assemble_tests.rs index 2e0e061dd..8843e3ad9 100644 --- a/src/simlin-engine/src/db/assemble_tests.rs +++ b/src/simlin-engine/src/db/assemble_tests.rs @@ -379,7 +379,7 @@ fn generated_ltm_helpers_are_captures_only() { let ltm_implicit = model_ltm_implicit_var_info(&db, model, sync.project); assert!( !ltm_implicit.is_empty(), - "the flow-to-stock score synthesizes PREVIOUS capture helpers" + "the gap -> adjustment score's frozen dynamic-index read synthesizes a capture helper" ); assert!( ltm_implicit @@ -704,8 +704,10 @@ fn results_offsets_are_the_assembled_layouts_offsets_on_a_module_bearing_model() /// A goal-seeking loop through a stdlib SMTH1 instance plus an arrayed /// growth loop: under LTM the layout grows a synthetic-variable section -/// (scalar and arrayed link scores, a loop score) and an LTM implicit section -/// (the flow-to-stock score's nested `PREVIOUS` capture helpers). +/// (scalar and arrayed link scores, the stocks' net-flow auxes, a loop score) +/// and an LTM implicit section (the capture helper the `gap -> adjustment` +/// score's frozen dynamic-index read `PREVIOUS(pop[PREVIOUS(idx, idx)])` +/// synthesizes). fn ltm_project() -> datamodel::Project { let mut project = x_project( sim_specs(), @@ -716,7 +718,8 @@ fn ltm_project() -> datamodel::Project { x_stock("level", "50", &["adjustment"], &[], None), x_aux("smoothed_level", "SMTH1(level, 3)", None), x_aux("gap", "goal - smoothed_level", None), - x_flow("adjustment", "gap / 5", None), + x_aux("idx", "1", None), + x_flow("adjustment", "gap / 5 + pop[idx] / 1000", None), datamodel::Variable::Stock(datamodel::Stock { ident: "pop".to_string(), equation: datamodel::Equation::ApplyToAll( @@ -795,7 +798,7 @@ fn results_offsets_are_the_assembled_layouts_offsets_under_ltm() { ); assert!( any_with("arg0"), - "the flow-to-stock score's nested PREVIOUS capture helpers are saved series: {keys:?}" + "the gap -> adjustment score's capture helper is a saved series: {keys:?}" ); // The arrayed `grow -> pop` link score occupies two slots and is keyed once. @@ -879,8 +882,10 @@ fn dedup_consecutive(names: Vec) -> Vec { /// A resolved recurrence SCC (`ref.mdl`-shaped `ce`/`ecc`, whose element graph /// is acyclic), a stock initialized from it so the SCC's members are scheduled -/// in the initials too, a goal-seeking loop through a stdlib SMTH1 instance, -/// and LTM enabled. +/// in the initials too, a goal-seeking loop through a stdlib SMTH1 instance +/// whose flow reads an arrayed weight by a dynamic index (so the `gap -> +/// inflow` score mints a capture helper and the LTM implicit tail is +/// populated), and LTM enabled. fn scc_stdlib_ltm_project() -> datamodel::Project { let mut project = x_project( sim_specs(), @@ -901,8 +906,10 @@ fn scc_stdlib_ltm_project() -> datamodel::Project { ("t3", "ce[t3] + 1"), ], ), + x_arrayed("w", "t", &[("t1", "1"), ("t2", "2"), ("t3", "3")]), + x_aux("idx", "1", None), x_stock("acc", "ecc[t3] * 10", &["inflow"], &[], None), - x_flow("inflow", "gap / 5", None), + x_flow("inflow", "gap / 5 + w[idx] / 1000", None), x_aux("goal", "100", None), x_aux("smoothed", "SMTH1(acc, 3)", None), x_aux("gap", "goal - smoothed", None), @@ -972,7 +979,7 @@ fn each_program_emits_in_runlist_order_then_the_ltm_tail() { ); assert!( !ltm_implicit.is_empty(), - "the flow-to-stock score yields nested PREVIOUS capture helpers" + "the gap -> inflow score's frozen dynamic-index read yields a capture helper" ); let synthetic_tail: Vec = ltm_vars .vars diff --git a/src/simlin-engine/src/db/diagnostic.rs b/src/simlin-engine/src/db/diagnostic.rs index ee6a4b105..5523fa236 100644 --- a/src/simlin-engine/src/db/diagnostic.rs +++ b/src/simlin-engine/src/db/diagnostic.rs @@ -94,8 +94,9 @@ use crate::common::{Error, UnitError}; /// NaN spreads through arithmetic into whatever reads it. /// 5. When LTM is enabled, `emit_conveyor_ltm_degraded_warnings` and /// `emit_queue_ltm_degraded_warnings` -- one `Warning` per conveyor stock -/// and per queue stock in THIS model, because LTM's flow-to-stock link-score -/// formula assumes plain INTEG but both are non-INTEG stock types +/// and per queue stock in THIS model, because LTM's flow-to-stock link score +/// treats a stock's net flow as its rate of change (plain INTEG) but both +/// are non-INTEG stock types /// (docs/design/conveyors.md §9.6, docs/design/queues.md §10.5). Emitted here /// rather than inside `model_ltm_variables` so each fires exactly once even /// for a module-referenced sub-model (see those functions' rustdoc for the @@ -401,11 +402,11 @@ fn emit_duplicate_variable_diagnostics(db: &dyn Db, model: SourceModel) { /// /// Both stock types have non-INTEG dynamics -- a conveyor's material rides a /// fixed-length belt and exits after the transit time, a queue is a FIFO of -/// batches whose outflow is demand-driven -- so the change from t-1 to t is -/// NOT `dt * inflow(t-1)`. LTM's flow-to-stock link-score numerator -/// (`PREVIOUS(flow) - PREVIOUS(PREVIOUS(flow))`) assumes plain INTEG under -/// Euler, so any link or loop score touching such a stock would be silently -/// wrong. The salsa DIAGNOSTIC path never expands either stock type into its +/// batches whose outflow is demand-driven -- so the stock's rate of change is +/// NOT `inflows - outflows`. LTM's flow-to-stock link score is the partial of +/// exactly that net flow (`ltm_augment::generate_flow_to_stock_equation`, +/// which assumes plain INTEG), so any link or loop score touching such a +/// stock would be silently wrong. The salsa DIAGNOSTIC path never expands either stock type into its /// hidden variables + native pass (only the special-stock build path /// `queue_compile::build_vm` does, which CLEARS the marker), so the `Compat` /// marker is still present here and the stock would be scored as plain INTEG. @@ -450,8 +451,8 @@ fn emit_ltm_degraded_warnings( let msg = format!( "LTM (Loops That Matter) analysis over {noun} stock '{name}' is degraded: a {noun} \ is a stock with non-INTEG dynamics{dynamics_detail}, but the flow-to-stock \ - link-score numerator `PREVIOUS(flow) - PREVIOUS(PREVIOUS(flow))` assumes plain \ - INTEG under Euler, so any link or loop score touching '{name}' may be wrong. \ + link score treats the stock's net flow `inflows - outflows` as its rate of \ + change (plain INTEG), so any link or loop score touching '{name}' may be wrong. \ Treat scores involving this {noun} as advisory ({doc_ref})." ); Diagnostic { diff --git a/src/simlin-engine/src/db/fragment_char_golden/ltm_loop_discovery.txt b/src/simlin-engine/src/db/fragment_char_golden/ltm_loop_discovery.txt index 09e896f1f..4cff17a51 100644 --- a/src/simlin-engine/src/db/fragment_char_golden/ltm_loop_discovery.txt +++ b/src/simlin-engine/src/db/fragment_char_golden/ltm_loop_discovery.txt @@ -1,16 +1,23 @@ ########## model main (module inputs: []) ########## == layout == - n_slots: 9 - $⁚$⁚ltm⁚link_score⁚growth→level⁚0⁚arg0: [6, 7) - $⁚$⁚ltm⁚link_score⁚growth→level⁚1⁚arg0: [7, 8) - $⁚$⁚ltm⁚link_score⁚growth→level⁚2⁚arg0: [8, 9) - $⁚ltm⁚link_score⁚growth→level: [3, 4) - $⁚ltm⁚link_score⁚level→growth: [4, 5) - $⁚ltm⁚link_score⁚rate→growth: [5, 6) + n_slots: 17 + $⁚$⁚ltm⁚link_score⁚idx→growth⁚0⁚arg0: [12, 13) + $⁚$⁚ltm⁚link_score⁚level→growth⁚0⁚arg0: [13, 14) + $⁚$⁚ltm⁚link_score⁚rate→growth⁚0⁚arg0: [14, 15) + $⁚$⁚ltm⁚link_score⁚weight→growth⁚0⁚arg0: [15, 16) + $⁚$⁚ltm⁚link_score⁚weight→growth⁚1⁚arg0: [16, 17) + $⁚ltm⁚link_score⁚growth→level: [6, 7) + $⁚ltm⁚link_score⁚idx→growth: [7, 8) + $⁚ltm⁚link_score⁚level→growth: [8, 9) + $⁚ltm⁚link_score⁚rate→growth: [9, 10) + $⁚ltm⁚link_score⁚weight→growth: [10, 11) + $⁚ltm⁚net⁚level: [11, 12) growth: [0, 1) - level: [1, 2) - rate: [2, 3) -== main::$⁚$⁚ltm⁚link_score⁚growth→level⁚0⁚arg0 [ltm-implicit-helper] : flow == + idx: [1, 2) + level: [2, 3) + rate: [3, 4) + weight: [4, 6) +== main::$⁚$⁚ltm⁚link_score⁚idx→growth⁚0⁚arg0 [ltm-implicit-helper] : flow == initial: flow: literals: [] @@ -20,26 +27,85 @@ module_decls: [] static_views: [] code: - 0000 LoadGlobalVar off=0 (time) - 0001 AssignCurr $⁚$⁚ltm⁚link_score⁚growth→level⁚0⁚arg0@0 - 0002 Ret + 0000 LoadVar idx@0 + 0001 PushSubscriptIndex bounds=2 + 0002 LoadSubscript weight@0 + 0003 AssignCurr $⁚$⁚ltm⁚link_score⁚idx→growth⁚0⁚arg0@0 + 0004 Ret stock: -== main::$⁚$⁚ltm⁚link_score⁚growth→level⁚1⁚arg0 [ltm-implicit-helper] : flow == +== main::$⁚$⁚ltm⁚link_score⁚level→growth⁚0⁚arg0 [ltm-implicit-helper] : flow == initial: flow: - literals: [0.0] + literals: [] temp_sizes: [] dim_lists: [] graphical_functions: [] module_decls: [] static_views: [] code: - 0000 LoadConstant #0 (=0.0) - 0001 LoadPrev growth@0 - 0002 AssignCurr $⁚$⁚ltm⁚link_score⁚growth→level⁚1⁚arg0@0 - 0003 Ret + 0000 LoadVar idx@0 + 0001 LoadPrev idx@0 + 0002 PushSubscriptIndex bounds=2 + 0003 LoadSubscript weight@0 + 0004 AssignCurr $⁚$⁚ltm⁚link_score⁚level→growth⁚0⁚arg0@0 + 0005 Ret + stock: +== main::$⁚$⁚ltm⁚link_score⁚rate→growth⁚0⁚arg0 [ltm-implicit-helper] : flow == + initial: + flow: + literals: [] + temp_sizes: [] + dim_lists: [] + graphical_functions: [] + module_decls: [] + static_views: [] + code: + 0000 LoadVar idx@0 + 0001 LoadPrev idx@0 + 0002 PushSubscriptIndex bounds=2 + 0003 LoadSubscript weight@0 + 0004 AssignCurr $⁚$⁚ltm⁚link_score⁚rate→growth⁚0⁚arg0@0 + 0005 Ret stock: -== main::$⁚$⁚ltm⁚link_score⁚growth→level⁚2⁚arg0 [ltm-implicit-helper] : flow == +== main::$⁚$⁚ltm⁚link_score⁚weight→growth⁚0⁚arg0 [ltm-implicit-helper] : flow == + initial: + flow: + literals: [] + temp_sizes: [] + dim_lists: [[2]] + graphical_functions: [] + module_decls: [] + static_views: [] + code: + 0000 PushVarViewDirect weight@0 dim_list=0 + 0001 LoadVar idx@0 + 0002 LoadPrev idx@0 + 0003 ViewSubscriptDynamic dim=0 + 0004 ArraySum + 0005 PopView + 0006 AssignCurr $⁚$⁚ltm⁚link_score⁚weight→growth⁚0⁚arg0@0 + 0007 Ret + stock: +== main::$⁚$⁚ltm⁚link_score⁚weight→growth⁚1⁚arg0 [ltm-implicit-helper] : flow == + initial: + flow: + literals: [] + temp_sizes: [] + dim_lists: [[2]] + graphical_functions: [] + module_decls: [] + static_views: [] + code: + 0000 PushVarViewDirect weight@0 dim_list=0 + 0001 LoadVar idx@0 + 0002 LoadPrev idx@0 + 0003 ViewSubscriptDynamic dim=0 + 0004 ArraySum + 0005 PopView + 0006 AssignCurr $⁚$⁚ltm⁚link_score⁚weight→growth⁚1⁚arg0@0 + 0007 Ret + stock: +== main::$⁚ltm⁚link_score⁚growth→level [ltm-synthetic] : flow == initial: flow: literals: [0.0] @@ -50,11 +116,48 @@ static_views: [] code: 0000 LoadConstant #0 (=0.0) - 0001 LoadPrev level@0 - 0002 AssignCurr $⁚$⁚ltm⁚link_score⁚growth→level⁚2⁚arg0@0 - 0003 Ret + 0001 LoadConstant #0 (=0.0) + 0002 LoadVar growth@0 + 0003 LoadConstant #0 (=0.0) + 0004 LoadPrev growth@0 + 0005 Op2 Sub + 0006 LoadVar $⁚ltm⁚net⁚level@0 + 0007 LoadConstant #0 (=0.0) + 0008 LoadPrev $⁚ltm⁚net⁚level@0 + 0009 Op2 Sub + 0010 Apply Abs + 0011 LoadConstant #0 (=0.0) + 0012 Apply SafeDiv + 0013 LoadVar growth@0 + 0014 LoadConstant #0 (=0.0) + 0015 LoadPrev growth@0 + 0016 Op2 Sub + 0017 Apply Sign + 0018 Op2 Mul + 0019 LoadVar $⁚ltm⁚net⁚level@0 + 0020 LoadConstant #0 (=0.0) + 0021 LoadPrev $⁚ltm⁚net⁚level@0 + 0022 Op2 Sub + 0023 LoadConstant #0 (=0.0) + 0024 Op2 Eq + 0025 LoadVar growth@0 + 0026 LoadConstant #0 (=0.0) + 0027 LoadPrev growth@0 + 0028 Op2 Sub + 0029 LoadConstant #0 (=0.0) + 0030 Op2 Eq + 0031 Op2 Or + 0032 SetCond + 0033 If + 0034 LoadGlobalVar off=0 (time) + 0035 LoadGlobalVar off=2 (initial_time) + 0036 Op2 Eq + 0037 SetCond + 0038 If + 0039 AssignCurr $⁚ltm⁚link_score⁚growth→level@0 + 0040 Ret stock: -== main::$⁚ltm⁚link_score⁚growth→level [ltm-synthetic] : flow == +== main::$⁚ltm⁚link_score⁚idx→growth [ltm-synthetic] : flow == initial: flow: literals: [0.0] @@ -65,38 +168,53 @@ static_views: [] code: 0000 LoadConstant #0 (=0.0) - 0001 LoadGlobalVar off=1 (dt) + 0001 LoadConstant #0 (=0.0) 0002 LoadConstant #0 (=0.0) - 0003 LoadPrev growth@0 + 0003 LoadPrev level@0 0004 LoadConstant #0 (=0.0) - 0005 LoadPrev $⁚$⁚ltm⁚link_score⁚growth→level⁚1⁚arg0@0 - 0006 Op2 Sub - 0007 Op2 Mul - 0008 LoadVar level@0 - 0009 LoadConstant #0 (=0.0) - 0010 LoadPrev level@0 - 0011 Op2 Sub - 0012 LoadConstant #0 (=0.0) - 0013 LoadPrev level@0 + 0005 LoadPrev rate@0 + 0006 Op2 Mul + 0007 LoadConstant #0 (=0.0) + 0008 LoadPrev $⁚$⁚ltm⁚link_score⁚idx→growth⁚0⁚arg0@0 + 0009 Op2 Mul + 0010 LoadConstant #0 (=0.0) + 0011 LoadPrev growth@0 + 0012 Op2 Sub + 0013 LoadVar growth@0 0014 LoadConstant #0 (=0.0) - 0015 LoadPrev $⁚$⁚ltm⁚link_score⁚growth→level⁚2⁚arg0@0 + 0015 LoadPrev growth@0 0016 Op2 Sub - 0017 Op2 Sub + 0017 Apply Abs 0018 LoadConstant #0 (=0.0) 0019 Apply SafeDiv - 0020 Apply Abs - 0021 LoadGlobalVar off=0 (time) - 0022 LoadGlobalVar off=2 (initial_time) - 0023 Op2 Eq - 0024 LoadGlobalVar off=2 (initial_time) - 0025 LoadPrev $⁚$⁚ltm⁚link_score⁚growth→level⁚0⁚arg0@0 - 0026 LoadGlobalVar off=2 (initial_time) - 0027 Op2 Eq - 0028 Op2 Or - 0029 SetCond - 0030 If - 0031 AssignCurr $⁚ltm⁚link_score⁚growth→level@0 - 0032 Ret + 0020 LoadVar idx@0 + 0021 LoadConstant #0 (=0.0) + 0022 LoadPrev idx@0 + 0023 Op2 Sub + 0024 Apply Sign + 0025 Op2 Mul + 0026 LoadVar growth@0 + 0027 LoadConstant #0 (=0.0) + 0028 LoadPrev growth@0 + 0029 Op2 Sub + 0030 LoadConstant #0 (=0.0) + 0031 Op2 Eq + 0032 LoadVar idx@0 + 0033 LoadConstant #0 (=0.0) + 0034 LoadPrev idx@0 + 0035 Op2 Sub + 0036 LoadConstant #0 (=0.0) + 0037 Op2 Eq + 0038 Op2 Or + 0039 SetCond + 0040 If + 0041 LoadGlobalVar off=0 (time) + 0042 LoadGlobalVar off=2 (initial_time) + 0043 Op2 Eq + 0044 SetCond + 0045 If + 0046 AssignCurr $⁚ltm⁚link_score⁚idx→growth@0 + 0047 Ret stock: == main::$⁚ltm⁚link_score⁚level→growth [ltm-synthetic] : flow == initial: @@ -115,43 +233,46 @@ 0004 LoadPrev rate@0 0005 Op2 Mul 0006 LoadConstant #0 (=0.0) - 0007 LoadPrev growth@0 - 0008 Op2 Sub - 0009 LoadVar growth@0 - 0010 LoadConstant #0 (=0.0) - 0011 LoadPrev growth@0 - 0012 Op2 Sub - 0013 Apply Abs - 0014 LoadConstant #0 (=0.0) - 0015 Apply SafeDiv - 0016 LoadVar level@0 + 0007 LoadPrev $⁚$⁚ltm⁚link_score⁚level→growth⁚0⁚arg0@0 + 0008 Op2 Mul + 0009 LoadConstant #0 (=0.0) + 0010 LoadPrev growth@0 + 0011 Op2 Sub + 0012 LoadVar growth@0 + 0013 LoadConstant #0 (=0.0) + 0014 LoadPrev growth@0 + 0015 Op2 Sub + 0016 Apply Abs 0017 LoadConstant #0 (=0.0) - 0018 LoadPrev level@0 - 0019 Op2 Sub - 0020 Apply Sign - 0021 Op2 Mul - 0022 LoadVar growth@0 - 0023 LoadConstant #0 (=0.0) - 0024 LoadPrev growth@0 - 0025 Op2 Sub + 0018 Apply SafeDiv + 0019 LoadVar level@0 + 0020 LoadConstant #0 (=0.0) + 0021 LoadPrev level@0 + 0022 Op2 Sub + 0023 Apply Sign + 0024 Op2 Mul + 0025 LoadVar growth@0 0026 LoadConstant #0 (=0.0) - 0027 Op2 Eq - 0028 LoadVar level@0 + 0027 LoadPrev growth@0 + 0028 Op2 Sub 0029 LoadConstant #0 (=0.0) - 0030 LoadPrev level@0 - 0031 Op2 Sub + 0030 Op2 Eq + 0031 LoadVar level@0 0032 LoadConstant #0 (=0.0) - 0033 Op2 Eq - 0034 Op2 Or - 0035 SetCond - 0036 If - 0037 LoadGlobalVar off=0 (time) - 0038 LoadGlobalVar off=2 (initial_time) - 0039 Op2 Eq - 0040 SetCond - 0041 If - 0042 AssignCurr $⁚ltm⁚link_score⁚level→growth@0 - 0043 Ret + 0033 LoadPrev level@0 + 0034 Op2 Sub + 0035 LoadConstant #0 (=0.0) + 0036 Op2 Eq + 0037 Op2 Or + 0038 SetCond + 0039 If + 0040 LoadGlobalVar off=0 (time) + 0041 LoadGlobalVar off=2 (initial_time) + 0042 Op2 Eq + 0043 SetCond + 0044 If + 0045 AssignCurr $⁚ltm⁚link_score⁚level→growth@0 + 0046 Ret stock: == main::$⁚ltm⁚link_score⁚rate→growth [ltm-synthetic] : flow == initial: @@ -170,43 +291,132 @@ 0004 LoadVar rate@0 0005 Op2 Mul 0006 LoadConstant #0 (=0.0) - 0007 LoadPrev growth@0 - 0008 Op2 Sub - 0009 LoadVar growth@0 - 0010 LoadConstant #0 (=0.0) - 0011 LoadPrev growth@0 - 0012 Op2 Sub - 0013 Apply Abs - 0014 LoadConstant #0 (=0.0) - 0015 Apply SafeDiv - 0016 LoadVar rate@0 + 0007 LoadPrev $⁚$⁚ltm⁚link_score⁚rate→growth⁚0⁚arg0@0 + 0008 Op2 Mul + 0009 LoadConstant #0 (=0.0) + 0010 LoadPrev growth@0 + 0011 Op2 Sub + 0012 LoadVar growth@0 + 0013 LoadConstant #0 (=0.0) + 0014 LoadPrev growth@0 + 0015 Op2 Sub + 0016 Apply Abs 0017 LoadConstant #0 (=0.0) - 0018 LoadPrev rate@0 - 0019 Op2 Sub - 0020 Apply Sign - 0021 Op2 Mul - 0022 LoadVar growth@0 - 0023 LoadConstant #0 (=0.0) - 0024 LoadPrev growth@0 - 0025 Op2 Sub + 0018 Apply SafeDiv + 0019 LoadVar rate@0 + 0020 LoadConstant #0 (=0.0) + 0021 LoadPrev rate@0 + 0022 Op2 Sub + 0023 Apply Sign + 0024 Op2 Mul + 0025 LoadVar growth@0 0026 LoadConstant #0 (=0.0) - 0027 Op2 Eq - 0028 LoadVar rate@0 + 0027 LoadPrev growth@0 + 0028 Op2 Sub 0029 LoadConstant #0 (=0.0) - 0030 LoadPrev rate@0 - 0031 Op2 Sub + 0030 Op2 Eq + 0031 LoadVar rate@0 0032 LoadConstant #0 (=0.0) - 0033 Op2 Eq - 0034 Op2 Or - 0035 SetCond - 0036 If - 0037 LoadGlobalVar off=0 (time) - 0038 LoadGlobalVar off=2 (initial_time) - 0039 Op2 Eq - 0040 SetCond - 0041 If - 0042 AssignCurr $⁚ltm⁚link_score⁚rate→growth@0 - 0043 Ret + 0033 LoadPrev rate@0 + 0034 Op2 Sub + 0035 LoadConstant #0 (=0.0) + 0036 Op2 Eq + 0037 Op2 Or + 0038 SetCond + 0039 If + 0040 LoadGlobalVar off=0 (time) + 0041 LoadGlobalVar off=2 (initial_time) + 0042 Op2 Eq + 0043 SetCond + 0044 If + 0045 AssignCurr $⁚ltm⁚link_score⁚rate→growth@0 + 0046 Ret + stock: +== main::$⁚ltm⁚link_score⁚weight→growth [ltm-synthetic] : flow == + initial: + flow: + literals: [0.0] + temp_sizes: [] + dim_lists: [[2], [2]] + graphical_functions: [] + module_decls: [] + static_views: [] + code: + 0000 LoadConstant #0 (=0.0) + 0001 LoadConstant #0 (=0.0) + 0002 LoadConstant #0 (=0.0) + 0003 LoadPrev level@0 + 0004 LoadConstant #0 (=0.0) + 0005 LoadPrev rate@0 + 0006 Op2 Mul + 0007 LoadVar idx@0 + 0008 LoadPrev idx@0 + 0009 PushSubscriptIndex bounds=2 + 0010 LoadSubscript weight@0 + 0011 Op2 Mul + 0012 LoadConstant #0 (=0.0) + 0013 LoadPrev growth@0 + 0014 Op2 Sub + 0015 LoadVar growth@0 + 0016 LoadConstant #0 (=0.0) + 0017 LoadPrev growth@0 + 0018 Op2 Sub + 0019 Apply Abs + 0020 LoadConstant #0 (=0.0) + 0021 Apply SafeDiv + 0022 PushVarViewDirect weight@0 dim_list=0 + 0023 LoadVar idx@0 + 0024 LoadPrev idx@0 + 0025 ViewSubscriptDynamic dim=0 + 0026 ArraySum + 0027 PopView + 0028 LoadConstant #0 (=0.0) + 0029 LoadPrev $⁚$⁚ltm⁚link_score⁚weight→growth⁚1⁚arg0@0 + 0030 Op2 Sub + 0031 Apply Sign + 0032 Op2 Mul + 0033 LoadVar growth@0 + 0034 LoadConstant #0 (=0.0) + 0035 LoadPrev growth@0 + 0036 Op2 Sub + 0037 LoadConstant #0 (=0.0) + 0038 Op2 Eq + 0039 PushVarViewDirect weight@0 dim_list=1 + 0040 LoadVar idx@0 + 0041 LoadPrev idx@0 + 0042 ViewSubscriptDynamic dim=0 + 0043 ArraySum + 0044 PopView + 0045 LoadConstant #0 (=0.0) + 0046 LoadPrev $⁚$⁚ltm⁚link_score⁚weight→growth⁚0⁚arg0@0 + 0047 Op2 Sub + 0048 LoadConstant #0 (=0.0) + 0049 Op2 Eq + 0050 Op2 Or + 0051 SetCond + 0052 If + 0053 LoadGlobalVar off=0 (time) + 0054 LoadGlobalVar off=2 (initial_time) + 0055 Op2 Eq + 0056 SetCond + 0057 If + 0058 AssignCurr $⁚ltm⁚link_score⁚weight→growth@0 + 0059 Ret + stock: +== main::$⁚ltm⁚net⁚level [ltm-synthetic] : flow == + initial: + flow: + literals: [0.0] + temp_sizes: [] + dim_lists: [] + graphical_functions: [] + module_decls: [] + static_views: [] + code: + 0000 LoadVar growth@0 + 0001 LoadConstant #0 (=0.0) + 0002 BinOpAssignCurr Sub $⁚ltm⁚net⁚level@0 + 0003 Ret stock: == main::growth [explicit] : flow == initial: @@ -220,8 +430,25 @@ code: 0000 LoadVar level@0 0001 LoadVar rate@0 - 0002 BinOpAssignCurr Mul growth@0 - 0003 Ret + 0002 Op2 Mul + 0003 LoadVar idx@0 + 0004 PushSubscriptIndex bounds=2 + 0005 LoadSubscript weight@0 + 0006 BinOpAssignCurr Mul growth@0 + 0007 Ret + stock: +== main::idx [explicit] : flow == + initial: + flow: + literals: [1.0] + temp_sizes: [] + dim_lists: [] + graphical_functions: [] + module_decls: [] + static_views: [] + code: + 0000 AssignConstCurr idx@0 #0 (=1.0) + 0001 Ret stock: == main::level [explicit] : initial+stock == initial: @@ -264,46 +491,84 @@ 0000 AssignConstCurr rate@0 #0 (=0.1) 0001 Ret stock: +== main::weight [explicit] : flow == + initial: + flow: + literals: [1.0, 2.0] + temp_sizes: [] + dim_lists: [] + graphical_functions: [] + module_decls: [] + static_views: [] + code: + 0000 AssignConstCurr weight@0 #0 (=1.0) + 0001 AssignConstCurr weight@1 #1 (=2.0) + 0002 Ret + stock: ########## runtime ########## step 0: - $⁚$⁚ltm⁚link_score⁚growth→level⁚0⁚arg0 = 0.0 - $⁚$⁚ltm⁚link_score⁚growth→level⁚1⁚arg0 = 0.0 - $⁚$⁚ltm⁚link_score⁚growth→level⁚2⁚arg0 = 0.0 + $⁚$⁚ltm⁚link_score⁚idx→growth⁚0⁚arg0 = 1.0 + $⁚$⁚ltm⁚link_score⁚level→growth⁚0⁚arg0 = 1.0 + $⁚$⁚ltm⁚link_score⁚rate→growth⁚0⁚arg0 = 1.0 + $⁚$⁚ltm⁚link_score⁚weight→growth⁚0⁚arg0 = 1.0 + $⁚$⁚ltm⁚link_score⁚weight→growth⁚1⁚arg0 = 1.0 $⁚ltm⁚link_score⁚growth→level = 0.0 + $⁚ltm⁚link_score⁚idx→growth = 0.0 $⁚ltm⁚link_score⁚level→growth = 0.0 $⁚ltm⁚link_score⁚rate→growth = 0.0 + $⁚ltm⁚link_score⁚weight→growth = 0.0 + $⁚ltm⁚net⁚level = 1.0 dt = 1.0 final_time = 2.0 growth = 1.0 + idx = 1.0 initial_time = 0.0 level = 10.0 rate = 0.1 time = 0.0 + weight[d1] = 1.0 + weight[d2] = 2.0 step 1: - $⁚$⁚ltm⁚link_score⁚growth→level⁚0⁚arg0 = 1.0 - $⁚$⁚ltm⁚link_score⁚growth→level⁚1⁚arg0 = 1.0 - $⁚$⁚ltm⁚link_score⁚growth→level⁚2⁚arg0 = 10.0 - $⁚ltm⁚link_score⁚growth→level = 0.0 + $⁚$⁚ltm⁚link_score⁚idx→growth⁚0⁚arg0 = 1.0 + $⁚$⁚ltm⁚link_score⁚level→growth⁚0⁚arg0 = 1.0 + $⁚$⁚ltm⁚link_score⁚rate→growth⁚0⁚arg0 = 1.0 + $⁚$⁚ltm⁚link_score⁚weight→growth⁚0⁚arg0 = 1.0 + $⁚$⁚ltm⁚link_score⁚weight→growth⁚1⁚arg0 = 1.0 + $⁚ltm⁚link_score⁚growth→level = 1.0 + $⁚ltm⁚link_score⁚idx→growth = 0.0 $⁚ltm⁚link_score⁚level→growth = 1.0 $⁚ltm⁚link_score⁚rate→growth = 0.0 + $⁚ltm⁚link_score⁚weight→growth = 0.0 + $⁚ltm⁚net⁚level = 1.1 dt = 1.0 final_time = 2.0 growth = 1.1 + idx = 1.0 initial_time = 0.0 level = 11.0 rate = 0.1 time = 1.0 + weight[d1] = 1.0 + weight[d2] = 2.0 step 2: - $⁚$⁚ltm⁚link_score⁚growth→level⁚0⁚arg0 = 2.0 - $⁚$⁚ltm⁚link_score⁚growth→level⁚1⁚arg0 = 1.1 - $⁚$⁚ltm⁚link_score⁚growth→level⁚2⁚arg0 = 11.0 - $⁚ltm⁚link_score⁚growth→level = 1.0000000000000044 + $⁚$⁚ltm⁚link_score⁚idx→growth⁚0⁚arg0 = 1.0 + $⁚$⁚ltm⁚link_score⁚level→growth⁚0⁚arg0 = 1.0 + $⁚$⁚ltm⁚link_score⁚rate→growth⁚0⁚arg0 = 1.0 + $⁚$⁚ltm⁚link_score⁚weight→growth⁚0⁚arg0 = 1.0 + $⁚$⁚ltm⁚link_score⁚weight→growth⁚1⁚arg0 = 1.0 + $⁚ltm⁚link_score⁚growth→level = 1.0 + $⁚ltm⁚link_score⁚idx→growth = 0.0 $⁚ltm⁚link_score⁚level→growth = 1.0 $⁚ltm⁚link_score⁚rate→growth = 0.0 + $⁚ltm⁚link_score⁚weight→growth = 0.0 + $⁚ltm⁚net⁚level = 1.21 dt = 1.0 final_time = 2.0 growth = 1.21 + idx = 1.0 initial_time = 0.0 level = 12.1 rate = 0.1 time = 2.0 + weight[d1] = 1.0 + weight[d2] = 2.0 diff --git a/src/simlin-engine/src/db/fragment_char_golden/ltm_loop_exhaustive.txt b/src/simlin-engine/src/db/fragment_char_golden/ltm_loop_exhaustive.txt index 61a40221e..20f4e1850 100644 --- a/src/simlin-engine/src/db/fragment_char_golden/ltm_loop_exhaustive.txt +++ b/src/simlin-engine/src/db/fragment_char_golden/ltm_loop_exhaustive.txt @@ -1,16 +1,17 @@ ########## model main (module inputs: []) ########## == layout == - n_slots: 9 - $⁚$⁚ltm⁚link_score⁚growth→level⁚0⁚arg0: [6, 7) - $⁚$⁚ltm⁚link_score⁚growth→level⁚1⁚arg0: [7, 8) - $⁚$⁚ltm⁚link_score⁚growth→level⁚2⁚arg0: [8, 9) - $⁚ltm⁚link_score⁚growth→level: [3, 4) - $⁚ltm⁚link_score⁚level→growth: [4, 5) - $⁚ltm⁚loop_score⁚r1: [5, 6) + n_slots: 11 + $⁚$⁚ltm⁚link_score⁚level→growth⁚0⁚arg0: [10, 11) + $⁚ltm⁚link_score⁚growth→level: [6, 7) + $⁚ltm⁚link_score⁚level→growth: [7, 8) + $⁚ltm⁚loop_score⁚r1: [8, 9) + $⁚ltm⁚net⁚level: [9, 10) growth: [0, 1) - level: [1, 2) - rate: [2, 3) -== main::$⁚$⁚ltm⁚link_score⁚growth→level⁚0⁚arg0 [ltm-implicit-helper] : flow == + idx: [1, 2) + level: [2, 3) + rate: [3, 4) + weight: [4, 6) +== main::$⁚$⁚ltm⁚link_score⁚level→growth⁚0⁚arg0 [ltm-implicit-helper] : flow == initial: flow: literals: [] @@ -20,39 +21,12 @@ module_decls: [] static_views: [] code: - 0000 LoadGlobalVar off=0 (time) - 0001 AssignCurr $⁚$⁚ltm⁚link_score⁚growth→level⁚0⁚arg0@0 - 0002 Ret - stock: -== main::$⁚$⁚ltm⁚link_score⁚growth→level⁚1⁚arg0 [ltm-implicit-helper] : flow == - initial: - flow: - literals: [0.0] - temp_sizes: [] - dim_lists: [] - graphical_functions: [] - module_decls: [] - static_views: [] - code: - 0000 LoadConstant #0 (=0.0) - 0001 LoadPrev growth@0 - 0002 AssignCurr $⁚$⁚ltm⁚link_score⁚growth→level⁚1⁚arg0@0 - 0003 Ret - stock: -== main::$⁚$⁚ltm⁚link_score⁚growth→level⁚2⁚arg0 [ltm-implicit-helper] : flow == - initial: - flow: - literals: [0.0] - temp_sizes: [] - dim_lists: [] - graphical_functions: [] - module_decls: [] - static_views: [] - code: - 0000 LoadConstant #0 (=0.0) - 0001 LoadPrev level@0 - 0002 AssignCurr $⁚$⁚ltm⁚link_score⁚growth→level⁚2⁚arg0@0 - 0003 Ret + 0000 LoadVar idx@0 + 0001 LoadPrev idx@0 + 0002 PushSubscriptIndex bounds=2 + 0003 LoadSubscript weight@0 + 0004 AssignCurr $⁚$⁚ltm⁚link_score⁚level→growth⁚0⁚arg0@0 + 0005 Ret stock: == main::$⁚ltm⁚link_score⁚growth→level [ltm-synthetic] : flow == initial: @@ -65,38 +39,46 @@ static_views: [] code: 0000 LoadConstant #0 (=0.0) - 0001 LoadGlobalVar off=1 (dt) - 0002 LoadConstant #0 (=0.0) - 0003 LoadPrev growth@0 - 0004 LoadConstant #0 (=0.0) - 0005 LoadPrev $⁚$⁚ltm⁚link_score⁚growth→level⁚1⁚arg0@0 - 0006 Op2 Sub - 0007 Op2 Mul - 0008 LoadVar level@0 - 0009 LoadConstant #0 (=0.0) - 0010 LoadPrev level@0 - 0011 Op2 Sub - 0012 LoadConstant #0 (=0.0) - 0013 LoadPrev level@0 + 0001 LoadConstant #0 (=0.0) + 0002 LoadVar growth@0 + 0003 LoadConstant #0 (=0.0) + 0004 LoadPrev growth@0 + 0005 Op2 Sub + 0006 LoadVar $⁚ltm⁚net⁚level@0 + 0007 LoadConstant #0 (=0.0) + 0008 LoadPrev $⁚ltm⁚net⁚level@0 + 0009 Op2 Sub + 0010 Apply Abs + 0011 LoadConstant #0 (=0.0) + 0012 Apply SafeDiv + 0013 LoadVar growth@0 0014 LoadConstant #0 (=0.0) - 0015 LoadPrev $⁚$⁚ltm⁚link_score⁚growth→level⁚2⁚arg0@0 + 0015 LoadPrev growth@0 0016 Op2 Sub - 0017 Op2 Sub - 0018 LoadConstant #0 (=0.0) - 0019 Apply SafeDiv - 0020 Apply Abs - 0021 LoadGlobalVar off=0 (time) - 0022 LoadGlobalVar off=2 (initial_time) - 0023 Op2 Eq - 0024 LoadGlobalVar off=2 (initial_time) - 0025 LoadPrev $⁚$⁚ltm⁚link_score⁚growth→level⁚0⁚arg0@0 - 0026 LoadGlobalVar off=2 (initial_time) - 0027 Op2 Eq - 0028 Op2 Or - 0029 SetCond - 0030 If - 0031 AssignCurr $⁚ltm⁚link_score⁚growth→level@0 - 0032 Ret + 0017 Apply Sign + 0018 Op2 Mul + 0019 LoadVar $⁚ltm⁚net⁚level@0 + 0020 LoadConstant #0 (=0.0) + 0021 LoadPrev $⁚ltm⁚net⁚level@0 + 0022 Op2 Sub + 0023 LoadConstant #0 (=0.0) + 0024 Op2 Eq + 0025 LoadVar growth@0 + 0026 LoadConstant #0 (=0.0) + 0027 LoadPrev growth@0 + 0028 Op2 Sub + 0029 LoadConstant #0 (=0.0) + 0030 Op2 Eq + 0031 Op2 Or + 0032 SetCond + 0033 If + 0034 LoadGlobalVar off=0 (time) + 0035 LoadGlobalVar off=2 (initial_time) + 0036 Op2 Eq + 0037 SetCond + 0038 If + 0039 AssignCurr $⁚ltm⁚link_score⁚growth→level@0 + 0040 Ret stock: == main::$⁚ltm⁚link_score⁚level→growth [ltm-synthetic] : flow == initial: @@ -115,43 +97,46 @@ 0004 LoadPrev rate@0 0005 Op2 Mul 0006 LoadConstant #0 (=0.0) - 0007 LoadPrev growth@0 - 0008 Op2 Sub - 0009 LoadVar growth@0 - 0010 LoadConstant #0 (=0.0) - 0011 LoadPrev growth@0 - 0012 Op2 Sub - 0013 Apply Abs - 0014 LoadConstant #0 (=0.0) - 0015 Apply SafeDiv - 0016 LoadVar level@0 + 0007 LoadPrev $⁚$⁚ltm⁚link_score⁚level→growth⁚0⁚arg0@0 + 0008 Op2 Mul + 0009 LoadConstant #0 (=0.0) + 0010 LoadPrev growth@0 + 0011 Op2 Sub + 0012 LoadVar growth@0 + 0013 LoadConstant #0 (=0.0) + 0014 LoadPrev growth@0 + 0015 Op2 Sub + 0016 Apply Abs 0017 LoadConstant #0 (=0.0) - 0018 LoadPrev level@0 - 0019 Op2 Sub - 0020 Apply Sign - 0021 Op2 Mul - 0022 LoadVar growth@0 - 0023 LoadConstant #0 (=0.0) - 0024 LoadPrev growth@0 - 0025 Op2 Sub + 0018 Apply SafeDiv + 0019 LoadVar level@0 + 0020 LoadConstant #0 (=0.0) + 0021 LoadPrev level@0 + 0022 Op2 Sub + 0023 Apply Sign + 0024 Op2 Mul + 0025 LoadVar growth@0 0026 LoadConstant #0 (=0.0) - 0027 Op2 Eq - 0028 LoadVar level@0 + 0027 LoadPrev growth@0 + 0028 Op2 Sub 0029 LoadConstant #0 (=0.0) - 0030 LoadPrev level@0 - 0031 Op2 Sub + 0030 Op2 Eq + 0031 LoadVar level@0 0032 LoadConstant #0 (=0.0) - 0033 Op2 Eq - 0034 Op2 Or - 0035 SetCond - 0036 If - 0037 LoadGlobalVar off=0 (time) - 0038 LoadGlobalVar off=2 (initial_time) - 0039 Op2 Eq - 0040 SetCond - 0041 If - 0042 AssignCurr $⁚ltm⁚link_score⁚level→growth@0 - 0043 Ret + 0033 LoadPrev level@0 + 0034 Op2 Sub + 0035 LoadConstant #0 (=0.0) + 0036 Op2 Eq + 0037 Op2 Or + 0038 SetCond + 0039 If + 0040 LoadGlobalVar off=0 (time) + 0041 LoadGlobalVar off=2 (initial_time) + 0042 Op2 Eq + 0043 SetCond + 0044 If + 0045 AssignCurr $⁚ltm⁚link_score⁚level→growth@0 + 0046 Ret stock: == main::$⁚ltm⁚loop_score⁚r1 [ltm-synthetic] : flow == initial: @@ -168,6 +153,21 @@ 0002 BinOpAssignCurr Mul $⁚ltm⁚loop_score⁚r1@0 0003 Ret stock: +== main::$⁚ltm⁚net⁚level [ltm-synthetic] : flow == + initial: + flow: + literals: [0.0] + temp_sizes: [] + dim_lists: [] + graphical_functions: [] + module_decls: [] + static_views: [] + code: + 0000 LoadVar growth@0 + 0001 LoadConstant #0 (=0.0) + 0002 BinOpAssignCurr Sub $⁚ltm⁚net⁚level@0 + 0003 Ret + stock: == main::growth [explicit] : flow == initial: flow: @@ -180,8 +180,25 @@ code: 0000 LoadVar level@0 0001 LoadVar rate@0 - 0002 BinOpAssignCurr Mul growth@0 - 0003 Ret + 0002 Op2 Mul + 0003 LoadVar idx@0 + 0004 PushSubscriptIndex bounds=2 + 0005 LoadSubscript weight@0 + 0006 BinOpAssignCurr Mul growth@0 + 0007 Ret + stock: +== main::idx [explicit] : flow == + initial: + flow: + literals: [1.0] + temp_sizes: [] + dim_lists: [] + graphical_functions: [] + module_decls: [] + static_views: [] + code: + 0000 AssignConstCurr idx@0 #0 (=1.0) + 0001 Ret stock: == main::level [explicit] : initial+stock == initial: @@ -224,46 +241,66 @@ 0000 AssignConstCurr rate@0 #0 (=0.1) 0001 Ret stock: +== main::weight [explicit] : flow == + initial: + flow: + literals: [1.0, 2.0] + temp_sizes: [] + dim_lists: [] + graphical_functions: [] + module_decls: [] + static_views: [] + code: + 0000 AssignConstCurr weight@0 #0 (=1.0) + 0001 AssignConstCurr weight@1 #1 (=2.0) + 0002 Ret + stock: ########## runtime ########## step 0: - $⁚$⁚ltm⁚link_score⁚growth→level⁚0⁚arg0 = 0.0 - $⁚$⁚ltm⁚link_score⁚growth→level⁚1⁚arg0 = 0.0 - $⁚$⁚ltm⁚link_score⁚growth→level⁚2⁚arg0 = 0.0 + $⁚$⁚ltm⁚link_score⁚level→growth⁚0⁚arg0 = 1.0 $⁚ltm⁚link_score⁚growth→level = 0.0 $⁚ltm⁚link_score⁚level→growth = 0.0 $⁚ltm⁚loop_score⁚r1 = 0.0 + $⁚ltm⁚net⁚level = 1.0 dt = 1.0 final_time = 2.0 growth = 1.0 + idx = 1.0 initial_time = 0.0 level = 10.0 rate = 0.1 time = 0.0 + weight[d1] = 1.0 + weight[d2] = 2.0 step 1: - $⁚$⁚ltm⁚link_score⁚growth→level⁚0⁚arg0 = 1.0 - $⁚$⁚ltm⁚link_score⁚growth→level⁚1⁚arg0 = 1.0 - $⁚$⁚ltm⁚link_score⁚growth→level⁚2⁚arg0 = 10.0 - $⁚ltm⁚link_score⁚growth→level = 0.0 + $⁚$⁚ltm⁚link_score⁚level→growth⁚0⁚arg0 = 1.0 + $⁚ltm⁚link_score⁚growth→level = 1.0 $⁚ltm⁚link_score⁚level→growth = 1.0 - $⁚ltm⁚loop_score⁚r1 = 0.0 + $⁚ltm⁚loop_score⁚r1 = 1.0 + $⁚ltm⁚net⁚level = 1.1 dt = 1.0 final_time = 2.0 growth = 1.1 + idx = 1.0 initial_time = 0.0 level = 11.0 rate = 0.1 time = 1.0 + weight[d1] = 1.0 + weight[d2] = 2.0 step 2: - $⁚$⁚ltm⁚link_score⁚growth→level⁚0⁚arg0 = 2.0 - $⁚$⁚ltm⁚link_score⁚growth→level⁚1⁚arg0 = 1.1 - $⁚$⁚ltm⁚link_score⁚growth→level⁚2⁚arg0 = 11.0 - $⁚ltm⁚link_score⁚growth→level = 1.0000000000000044 + $⁚$⁚ltm⁚link_score⁚level→growth⁚0⁚arg0 = 1.0 + $⁚ltm⁚link_score⁚growth→level = 1.0 $⁚ltm⁚link_score⁚level→growth = 1.0 - $⁚ltm⁚loop_score⁚r1 = 1.0000000000000044 + $⁚ltm⁚loop_score⁚r1 = 1.0 + $⁚ltm⁚net⁚level = 1.21 dt = 1.0 final_time = 2.0 growth = 1.21 + idx = 1.0 initial_time = 0.0 level = 12.1 rate = 0.1 time = 2.0 + weight[d1] = 1.0 + weight[d2] = 2.0 diff --git a/src/simlin-engine/src/db/fragment_char_tests.rs b/src/simlin-engine/src/db/fragment_char_tests.rs index 3c027cef1..3de1c77e1 100644 --- a/src/simlin-engine/src/db/fragment_char_tests.rs +++ b/src/simlin-engine/src/db/fragment_char_tests.rs @@ -1651,8 +1651,12 @@ fn char_resolved_recurrence_scc() { // The two `db/ltm/compile.rs` emission sites -- each an inline COPY of the // shared compile+symbolize tail that `db::assemble` also carries -- are what // this fixture reaches: `compile_ltm_synthetic_fragment` for the link/loop -// score variables, and `compile_ltm_implicit_var_fragment` for the PREVIOUS -// capture helpers those score equations synthesize. +// score variables and the stock's net-flow aux, and +// `compile_ltm_implicit_var_fragment` for the PREVIOUS capture helper a +// score equation synthesizes -- here the `level -> growth` partial's frozen +// dynamic-index read `PREVIOUS(weight[PREVIOUS(idx, idx)])`, which is why +// the flow reads its rate through `weight[idx]` (`weight[d1] = 1`, so the +// simulated values are those of `growth = level * rate`). // // Both modes are rendered because they emit DIFFERENT synthetic variables from // different arms: exhaustive enumeration produces one link score per LOOP edge @@ -1665,10 +1669,21 @@ fn char_resolved_recurrence_scc() { // --------------------------------------------------------------------------- fn ltm_loop_model() -> datamodel::Project { + ltm_loop_model_with("0.1", "level * rate * weight[idx]") +} + +/// [`ltm_loop_model`] with `rate`'s constant and `growth`'s equation as +/// given -- the one builder behind the fixture and the edits the +/// incrementality half applies to it, so an edit changes exactly what it +/// names. +fn ltm_loop_model_with(rate: &str, growth: &str) -> datamodel::Project { TestProject::new("frag_ltm_loop") .with_sim_time(0.0, 2.0, 1.0) - .aux("rate", "0.1", None) - .flow("growth", "level * rate", None) + .named_dimension("d", &["d1", "d2"]) + .array_with_ranges("weight[d]", vec![("d1", "1"), ("d2", "2")]) + .aux("idx", "1", None) + .aux("rate", rate, None) + .flow("growth", growth, None) .stock("level", "10", &["growth"], &[], None) .build_datamodel() } @@ -1690,26 +1705,30 @@ fn char_ltm_fragments_exhaustive() { FixtureExpect { models: &[("main", &[])], phases: &[ - ("main::$⁚$⁚ltm⁚link_score⁚growth→level⁚0⁚arg0", "flow"), - ("main::$⁚$⁚ltm⁚link_score⁚growth→level⁚1⁚arg0", "flow"), - ("main::$⁚$⁚ltm⁚link_score⁚growth→level⁚2⁚arg0", "flow"), + ("main::$⁚$⁚ltm⁚link_score⁚level→growth⁚0⁚arg0", "flow"), ("main::$⁚ltm⁚link_score⁚growth→level", "flow"), ("main::$⁚ltm⁚link_score⁚level→growth", "flow"), ("main::$⁚ltm⁚loop_score⁚r1", "flow"), + ("main::$⁚ltm⁚net⁚level", "flow"), ("main::growth", "flow"), + ("main::idx", "flow"), ("main::level", "initial+stock"), ("main::rate", "flow"), + ("main::weight", "flow"), ], why: "Exhaustive enumeration finds the single circuit \ `level -> growth -> level` and emits a link score for each of \ its TWO edges plus one `loop_score⁚r1` for the circuit -- so \ - `rate→growth`, a causal edge no circuit traverses, gets no \ - score here (contrast the discovery fixture). Every synthetic \ - is a scalar aux, hence flow-only. Only the `growth→level` \ - score synthesizes PREVIOUS capture helpers, and a PREVIOUS \ - capture is flow-only too: its kind is its phase demand, and \ - the intrinsic's fallback covers every read before the first \ - step commits.", + `rate→growth`, `idx→growth` and `weight→growth`, causal edges \ + no circuit traverses, get no score here (contrast the \ + discovery fixture). The `growth→level` score reads the \ + stock's net-flow aux `$⁚ltm⁚net⁚level`, minted beside it. \ + Every synthetic is a scalar aux, hence flow-only. Only the \ + `level→growth` score synthesizes a PREVIOUS capture helper \ + (its frozen dynamic-index read), and a PREVIOUS capture is \ + flow-only too: its kind is its phase demand, and the \ + intrinsic's fallback covers every read before the first step \ + commits.", spot_checks: LTM_LOOP_SPOT_CHECKS, ltm: FixtureLtm::Exhaustive, expect_one_resolved_scc: false, @@ -1725,26 +1744,36 @@ fn char_ltm_fragments_discovery() { FixtureExpect { models: &[("main", &[])], phases: &[ - ("main::$⁚$⁚ltm⁚link_score⁚growth→level⁚0⁚arg0", "flow"), - ("main::$⁚$⁚ltm⁚link_score⁚growth→level⁚1⁚arg0", "flow"), - ("main::$⁚$⁚ltm⁚link_score⁚growth→level⁚2⁚arg0", "flow"), + ("main::$⁚$⁚ltm⁚link_score⁚idx→growth⁚0⁚arg0", "flow"), + ("main::$⁚$⁚ltm⁚link_score⁚level→growth⁚0⁚arg0", "flow"), + ("main::$⁚$⁚ltm⁚link_score⁚rate→growth⁚0⁚arg0", "flow"), + ("main::$⁚$⁚ltm⁚link_score⁚weight→growth⁚0⁚arg0", "flow"), + ("main::$⁚$⁚ltm⁚link_score⁚weight→growth⁚1⁚arg0", "flow"), ("main::$⁚ltm⁚link_score⁚growth→level", "flow"), + ("main::$⁚ltm⁚link_score⁚idx→growth", "flow"), ("main::$⁚ltm⁚link_score⁚level→growth", "flow"), ("main::$⁚ltm⁚link_score⁚rate→growth", "flow"), + ("main::$⁚ltm⁚link_score⁚weight→growth", "flow"), + ("main::$⁚ltm⁚net⁚level", "flow"), ("main::growth", "flow"), + ("main::idx", "flow"), ("main::level", "initial+stock"), ("main::rate", "flow"), + ("main::weight", "flow"), ], - why: "Discovery mode scores every CAUSAL edge, so `rate→growth` \ - appears here even though no circuit traverses it, and no \ - loop-score variable is emitted at all (discovery finds \ - and ranks loops after the run instead of enumerating \ - circuits at compile time). \ - Every score is a scalar aux, hence flow-only. Only the \ - `growth→level` score -- the stock-update edge, whose \ - ceteris-paribus numerator re-integrates the stock -- \ - synthesizes PREVIOUS capture helpers, and a PREVIOUS capture \ - is flow-only too: its kind is its phase demand.", + why: "Discovery mode scores every CAUSAL edge, so `rate→growth`, \ + `idx→growth` and `weight→growth` appear here even though no \ + circuit traverses them, and no loop-score variable is emitted \ + at all (discovery finds and ranks loops after the run instead \ + of enumerating circuits at compile time). The `growth→level` \ + score reads the stock's net-flow aux `$⁚ltm⁚net⁚level`, \ + minted beside it. Every score is a scalar aux, hence \ + flow-only. Every score into `growth` but the one from \ + `level`'s co-read `idx` freezes the dynamic-index read \ + `weight[idx]` -- the whole read for `level`/`rate`, the \ + index alone or the head alone for `weight` -- and each such \ + freeze is a PREVIOUS capture helper, flow-only too: its kind \ + is its phase demand.", spot_checks: LTM_LOOP_SPOT_CHECKS, ltm: FixtureLtm::Discovery, expect_one_resolved_scc: false, @@ -2646,12 +2675,7 @@ fn implicit_and_ltm_fragment_cache_granularity() { // Step A: change a CONSTANT. A link score's equation is derived from the // TARGET's equation structure, not from any source's value, so no link // score's text moves and no LTM fragment recompiles. - let rate_edited = TestProject::new("frag_ltm_loop") - .with_sim_time(0.0, 2.0, 1.0) - .aux("rate", "0.2", None) - .flow("growth", "level * rate", None) - .stock("level", "10", &["growth"], &[], None) - .build_datamodel(); + let rate_edited = ltm_loop_model_with("0.2", "level * rate * weight[idx]"); let (ltm_state3, rate_execs) = resync_and_assemble( &mut ltm_db, &rate_edited, @@ -2691,12 +2715,7 @@ fn implicit_and_ltm_fragment_cache_granularity() { // unchanged dependency set, which is a salsa dependency-verification // question rather than a fragment-compiler one. Asserting the full set // here would land a flaky test. - let growth_edited = TestProject::new("frag_ltm_loop") - .with_sim_time(0.0, 2.0, 1.0) - .aux("rate", "0.2", None) - .flow("growth", "level * rate * 1", None) - .stock("level", "10", &["growth"], &[], None) - .build_datamodel(); + let growth_edited = ltm_loop_model_with("0.2", "level * rate * weight[idx] * 1"); let (_ltm_state4, growth_execs) = resync_and_assemble( &mut ltm_db, &growth_edited, diff --git a/src/simlin-engine/src/db/fragment_input_tests.rs b/src/simlin-engine/src/db/fragment_input_tests.rs index fec3ebbb3..a55f2550a 100644 --- a/src/simlin-engine/src/db/fragment_input_tests.rs +++ b/src/simlin-engine/src/db/fragment_input_tests.rs @@ -151,11 +151,17 @@ fn smooth_of_module_output_project() -> datamodel::Project { project } +/// One feedback loop whose flow reads an arrayed weight by a dynamic index, +/// so the `level -> growth` score's partial freezes `weight[idx]` as +/// `PREVIOUS(weight[PREVIOUS(idx, idx)])` and synthesizes a capture helper. fn ltm_loop_project() -> datamodel::Project { TestProject::new("fragment_input_ltm") .with_sim_time(0.0, 2.0, 1.0) + .named_dimension("d", &["d1", "d2"]) + .array_with_ranges("weight[d]", vec![("d1", "1"), ("d2", "2")]) + .aux("idx", "1", None) .aux("rate", "0.1", None) - .flow("growth", "level * rate", None) + .flow("growth", "level * rate * weight[idx]", None) .stock("level", "10", &["growth"], &[], None) .build_datamodel() } @@ -355,8 +361,10 @@ fn ltm_implicit_constructor_is_compile_ltm_implicit_var_fragments_input() { let helpers = model_ltm_implicit_var_info(&db, model, sync.project); assert!( - helpers.contains_key("$⁚$⁚ltm⁚link_score⁚growth→level⁚0⁚arg0"), - "the growth->level score synthesizes PREVIOUS capture helpers" + helpers.contains_key("$⁚$⁚ltm⁚link_score⁚level→growth⁚0⁚arg0"), + "the level->growth score's frozen dynamic-index read synthesizes a PREVIOUS \ + capture helper; got {:?}", + helpers.keys().collect::>() ); for (name, meta) in helpers.iter() { let production = crate::db::ltm::compile_ltm_implicit_var_fragment( diff --git a/src/simlin-engine/src/db/ltm/compile.rs b/src/simlin-engine/src/db/ltm/compile.rs index 6e0b97489..9bae31d46 100644 --- a/src/simlin-engine/src/db/ltm/compile.rs +++ b/src/simlin-engine/src/db/ltm/compile.rs @@ -28,7 +28,7 @@ use crate::db::{ Db, Diagnostic, DiagnosticError, LtmLinkId, LtmSyntheticVar, RefShape, SourceModel, SourceProject, VarFragmentResult, canonical_module_input_set, compile_phase_to_per_var_bytecodes, lowered_variable_by_name, project_converted_dimensions, - project_dimensions_context, variable_tables, + project_datamodel_dims, project_dimensions_context, variable_tables, }; use super::parse::{parse_ltm_equation, scalarize_ltm_equation}; @@ -165,16 +165,19 @@ pub(crate) fn shaped_link_score_executions() -> usize { #[cfg_attr(feature = "debug-derive", derive(Debug))] #[derive(Clone, PartialEq)] pub enum ShapedLinkScore { - /// The link-score variable was generated. `freeze_helpers` carries the - /// GH #995 array-freeze helper variables the score's partial references - /// (usually empty); the emission loop pushes them alongside the score, - /// deduplicated by their content-derived names. + /// The link-score variable was generated. `helpers` carries the + /// companion synthetic variables the score reads at their current-step + /// value, which the emission loop registers beside it, deduplicated by + /// name: the GH #995 array-freeze helpers the partial references (usually + /// none) and, for a flow-to-stock score, the stock's net-flow aux + /// (`ltm_augment::generate_net_flow_equation`), which every flow of that + /// stock mints identically. Scored { /// Boxed to keep the enum small next to its dataless variants /// (clippy `large_enum_variant`); `LtmSyntheticVar` carries whole /// parsed equations. var: Box, - freeze_helpers: Vec, + helpers: Vec, }, /// A `PartialEquationError` made the edge unscoreable: the warning /// naming why, which the caller records with the edge. Boxed so the @@ -275,10 +278,11 @@ pub fn shaped_link_score<'db>( }) { // Module-link partials thread no dep-dims table (their GH #526 // check keeps the permissive legacy collapse), so no array - // freeze can be materialized on this arm. + // freeze can be materialized on this arm, and a module is never + // a stock's flow. Some(lsv) => ShapedLinkScore::Scored { var: Box::new(lsv), - freeze_helpers: vec![], + helpers: vec![], }, None => ShapedLinkScore::NoVariable, }; @@ -411,6 +415,13 @@ pub fn shaped_link_score<'db>( } }; + let mut helpers: Vec = raw_freeze_helpers + .into_iter() + .map(freeze_helper_var) + .collect(); + if let Some(net) = net_flow_aux(db, model, project, from_var.as_deref(), &to_var) { + helpers.push(net); + } ShapedLinkScore::Scored { var: Box::new(LtmSyntheticVar { name: var_name, @@ -418,11 +429,67 @@ pub fn shaped_link_score<'db>( dimensions: vec![], compile_directly: false, }), - freeze_helpers: raw_freeze_helpers - .into_iter() - .map(freeze_helper_var) - .collect(), + helpers, + } +} + +/// The net-flow aux a flow-to-stock score reads +/// (`ltm_augment::generate_net_flow_equation`), when `(from, to)` is a flow +/// into its stock; `None` for every other link. +/// +/// Built from the stock's declared flows alone, so every flow of one stock +/// mints the same variable and the emission loop's name dedup keeps one. Its +/// dimensions are the stock's, in the datamodel casing every arrayed LTM +/// variable is tagged with, so the score's apply-to-all body over those same +/// dimensions reads it element for element. +fn net_flow_aux( + db: &dyn Db, + model: SourceModel, + project: SourceProject, + from_var: Option<&crate::variable::Variable>, + to_var: &crate::variable::Variable, +) -> Option { + let (inflows, outflows) = crate::ltm_augment::flow_to_stock_wiring(from_var, to_var)?; + let stock = to_var.ident.as_str(); + let resolve = |flows: &[Ident]| { + flows + .iter() + .map(|flow| { + ( + flow.clone(), + lowered_variable_by_name(db, model, project, flow.as_str()), + ) + }) + .collect::>() + }; + let inflow_vars = resolve(inflows); + let outflow_vars = resolve(outflows); + fn refs( + flows: &[( + Ident, + Option>, + )], + ) -> Vec<(&str, Option<&crate::variable::Variable>)> { + flows + .iter() + .map(|(flow, var)| (flow.as_str(), var.as_deref())) + .collect() } + let equation = crate::ltm_augment::generate_net_flow_equation( + to_var, + &refs(&inflow_vars), + &refs(&outflow_vars), + ); + let dims = super::link_scores::datamodel_dim_names( + &endpoint_dimensions(db, model, project, stock).unwrap_or_default(), + project_datamodel_dims(db, project), + ); + Some(LtmSyntheticVar { + name: crate::ltm_augment::net_flow_var_name(stock), + equation: super::parse::retarget_ltm_equation_dims(equation, &dims), + dimensions: dims, + compile_directly: false, + }) } /// Convert a wrap-produced [`crate::ltm_augment::ArrayFreezeHelper`] into the diff --git a/src/simlin-engine/src/db/ltm/link_scores.rs b/src/simlin-engine/src/db/ltm/link_scores.rs index 0475642d9..d048c9637 100644 --- a/src/simlin-engine/src/db/ltm/link_scores.rs +++ b/src/simlin-engine/src/db/ltm/link_scores.rs @@ -219,6 +219,27 @@ impl ArrayedSlotMap { } } +/// The declared dimension names of `dims` in the project's datamodel casing +/// -- the names `parse_ltm_equation` feeds into `Equation::ApplyToAll`, which +/// `get_dimensions` resolves by exact string match against the project's +/// datamodel dimensions. A dimension the datamodel does not declare keeps its +/// canonical name. +pub(super) fn datamodel_dim_names( + dims: &[crate::dimensions::Dimension], + dm_dims: &[crate::datamodel::Dimension], +) -> Vec { + dims.iter() + .map(|d| { + let canonical = d.name(); + dm_dims + .iter() + .find(|dm| crate::common::canonicalize(dm.name()).as_ref() == canonical) + .map(|dm| dm.name().to_string()) + .unwrap_or_else(|| canonical.to_string()) + }) + .collect() +} + /// Whether `(from, to)` is a stock's structural inflow/outflow edge -- the one /// Bare edge the wiring's pairing (`db::BareSpelling::StockFlow`) applies to. fn is_structural_flow_to_stock(db: &dyn Db, model: SourceModel, from: &str, to: &str) -> bool { @@ -271,18 +292,25 @@ pub(super) fn link_score_dimensions( let from_dims = endpoint_dimensions(db, model, project, from).unwrap_or_default(); - // Scalar source -> arrayed target: NOT handled here. The main - // link-score loop routes these to `try_scalar_to_arrayed_link_scores` - // (one scalar link score per target element) before - // `emit_per_shape_link_scores` is reached. Returning empty here is - // the safe fallback if that routing is ever bypassed (e.g. the - // target failed to lower): a scalar Bare link score + // Scalar source -> arrayed target. A scalar FLOW feeds every element of + // its arrayed stock, and its flow-to-stock score is one arrayed variable + // over the stock's dimensions (the flow broadcast into each element's net + // flow): the element graph emits `flow -> stock[e]` for every element and + // discovery's `expand_a2a_link_offsets` attaches a scalar source to every + // target-element slot, so the arrayed score is the shape both read. Every + // other scalar-source edge is NOT handled here: the main link-score loop + // routes it to `try_scalar_to_arrayed_link_scores` (one scalar link score + // per target element) before `emit_per_shape_link_scores` is reached, and + // returning empty is the safe fallback if that routing is ever bypassed + // (e.g. the target failed to lower): a scalar Bare link score // (`{from}→{to}`, no dims) parses to the useless-but-harmless edge - // `(from, to)`, whereas a Bare-A2A var would make - // `expand_a2a_link_offsets` invent a phantom `from[elem]` node that - // breaks loops through `from` in the search graph. + // `(from, to)`. if from_dims.is_empty() { - return vec![]; + return if is_structural_flow_to_stock(db, model, from, to) { + datamodel_dim_names(to_dims, dm_dims) + } else { + vec![] + }; } // Same-dimension A2A: both have identical dimension(s). @@ -378,17 +406,7 @@ pub(super) fn link_score_dimensions( if dims_compatible { // Map canonical dimension names back to their original // datamodel names for correct equation parsing. - to_dims - .iter() - .map(|d| { - let canonical = d.name(); - dm_dims - .iter() - .find(|dm| crate::common::canonicalize(dm.name()).as_ref() == canonical) - .map(|dm| dm.name().to_string()) - .unwrap_or_else(|| canonical.to_string()) - }) - .collect() + datamodel_dim_names(to_dims, dm_dims) } else { // Cross-dimensional (arrayed-to-scalar, or mismatched // dimensions). These edges are handled by @@ -1591,17 +1609,7 @@ pub(super) fn try_scalar_to_arrayed_link_scores( // Map the owner's canonical dim names back to their // datamodel casing for correct equation parsing (the // same mapping `link_score_dimensions` applies). - to_dims - .iter() - .map(|d| { - let canonical = d.name(); - dm_dims - .iter() - .find(|dm| crate::common::canonicalize(dm.name()).as_ref() == canonical) - .map(|dm| dm.name().to_string()) - .unwrap_or_else(|| canonical.to_string()) - }) - .collect() + datamodel_dim_names(&to_dims, dm_dims) } else { agg.result_dims.clone() }; @@ -2481,10 +2489,7 @@ pub(super) fn try_disjoint_dim_arrayed_link_scores( return Some(vec![]); } ShapedLinkScore::NoVariable => continue, - ShapedLinkScore::Scored { - var: lsv, - freeze_helpers, - } => { + ShapedLinkScore::Scored { var: lsv, helpers } => { let mut lsv = *lsv; // `lsv.name` is already `link_score_var_name(from, to, &shape)` // from the shaped path -- no need to re-derive it here. @@ -2492,7 +2497,7 @@ pub(super) fn try_disjoint_dim_arrayed_link_scores( lsv.dimensions = dims.clone(); lsv.equation = retarget_ltm_equation_dims(lsv.equation, &dims); lsv.compile_directly = false; - vars.extend(freeze_helpers); + vars.extend(helpers); vars.push(lsv); } } @@ -2657,6 +2662,34 @@ pub(super) fn emit_per_shape_link_scores( return; } + emit_shaped_link_scores( + db, from, to, shapes, vars_start, model, project, dm_dims, vars, warnings, + ); +} + +/// Emit one link-score variable per shape of `shapes` for the `(from, to)` +/// edge, through the per-shape query (`shaped_link_score`), with the +/// companions each score reads registered beside it. +/// +/// The two callers decide the shape set: [`emit_per_shape_link_scores`] +/// reads it off the edge's classified reference sites, and the structural +/// flow-to-stock branch of [`emit_link_scores_for_edge`] passes `Bare` alone. +/// `vars_start` is `vars`' length before anything for this edge was pushed: +/// a GH #780 doom discards back to it, so no already-emitted sibling var +/// survives for an edge whose loops are dropped. +#[allow(clippy::too_many_arguments)] // threads the emission context +fn emit_shaped_link_scores( + db: &dyn Db, + from: &str, + to: &str, + shapes: Vec, + vars_start: usize, + model: SourceModel, + project: SourceProject, + dm_dims: &[crate::datamodel::Dimension], + vars: &mut Vec, + warnings: &mut LtmWarnings, +) { let target_dims = link_score_dimensions(db, from, to, model, project, dm_dims); // GH #758: when BOTH endpoints are arrayed non-module variables but @@ -2739,14 +2772,11 @@ pub(super) fn emit_per_shape_link_scores( // composite-less module link); not an unscoreable edge. continue; } - ShapedLinkScore::Scored { - var: lsv, - freeze_helpers, - } => { + ShapedLinkScore::Scored { var: lsv, helpers } => { let mut lsv = *lsv; // Set the canonical name and dimensions per Phase 3 Task 4/5. lsv.name = crate::ltm_augment::link_score_var_name(from, to, &shape); - vars.extend(freeze_helpers); + vars.extend(helpers); // Every shape takes the target's dimensions: for FixedIndex // each per-element link score is scalar when the target is // scalar and arrayed when the target is arrayed; Bare (and @@ -4284,6 +4314,37 @@ pub(super) fn emit_link_scores_for_edge( vars: &mut Vec, warnings: &mut LtmWarnings, ) { + // A flow into its stock. A stock's only causal in-edges are its wiring, + // and the flow reaches the stock through that wiring rather than through + // an equation, so the edge has exactly one shape, `Bare`, and one owner: + // `shaped_link_score` builds the flow-to-stock score and the stock's + // net-flow aux. The reference-site IR is deliberately NOT consulted: a + // stock's classified sites are those of its INITIAL-VALUE equation (a + // stock's `ast()` is its initial), so a per-element initial reading its + // own flow (`s[a] = f[a] * 5`) would mint one per-element score per + // slot, all byte-identical, and a pinned-element read in a + // two-dimensional initial would divert the edge to the per-element arm + // and drop every loop through it. Routed FIRST because the shape + // emitters below decide by endpoint dimensions alone: a scalar flow into + // an arrayed stock looks like a scalar-to-arrayed aux edge to + // `try_scalar_to_arrayed_link_scores`, which would score it as a partial + // of that same initial-value equation, a plausible-looking wrong number. + if is_structural_flow_to_stock(db, model, from, to) { + let vars_start = vars.len(); + emit_shaped_link_scores( + db, + from, + to, + vec![RefShape::Bare], + vars_start, + model, + project, + dm_dims, + vars, + warnings, + ); + return; + } // The set of synthetic aggs `(from, to)` routes through, read off // the reference-site IR (the unique `ThroughAgg` `AggRef`s of this // edge's classified sites, in first-occurrence order). The routing diff --git a/src/simlin-engine/src/db/ltm/mod.rs b/src/simlin-engine/src/db/ltm/mod.rs index d6a218884..552f662e0 100644 --- a/src/simlin-engine/src/db/ltm/mod.rs +++ b/src/simlin-engine/src/db/ltm/mod.rs @@ -188,15 +188,7 @@ pub(crate) fn endpoint_dimensions( } /// The single integration method the assembled simulation actually runs, when -/// it is NOT Euler. -/// -/// LTM's 2023 flow-to-stock link-score formula -/// (`PREVIOUS(flow) - PREVIOUS(PREVIOUS(flow))`) only aligns its numerator to -/// the causal interval that drove the stock change from t-1 to t under Euler -/// integration; under RK2/RK4 the sub-stepped stock update breaks that -/// alignment and the link scores become mathematically meaningless. The VM -/// and wasm backends both genuinely honor RK2/RK4 (distinct stepping loops), -/// so the bad scores would look plausible while being wrong. +/// it is NOT Euler -- the method the GH #486 guard keeps LTM off. /// /// A `CompiledSimulation` has exactly ONE `Specs.method`, resolved by /// `assemble_simulation` from the MAIN (root) model's `model_sim_specs` @@ -223,15 +215,13 @@ pub(super) fn effective_non_euler_method( /// Whether `model_ltm_variables` emits at least one flow-to-stock link score /// for this model -- the EXACT precondition the GH #486/#663 non-Euler guard -/// must gate on. +/// gates on. /// -/// A flow-to-stock link score is the only LTM synthetic var the Euler-only -/// `PREVIOUS(flow) - PREVIOUS(PREVIOUS(flow))` numerator drives; under RK2/RK4 -/// it is mathematically meaningless. The guard previously gated on "the model -/// has a stock", but that over-rejects a loop-free model (GH #663): in +/// The guard keys on a flow-to-stock score rather than on "the model has a +/// stock" because the latter over-rejects a loop-free model (GH #663): in /// exhaustive mode LTM scores only the edges of detected feedback loops, so an /// open-loop stock (a constant inflow that never reads the stock back) emits -/// NO flow-to-stock score and nothing is corrupted. +/// NO flow-to-stock score and there is nothing to guard. /// /// Crucially this is mode-aware where a loop-presence proxy is NOT: in /// DISCOVERY mode (user-forced or auto-flipped) and in any model with input @@ -289,12 +279,7 @@ fn sim_method_display_name(method: datamodel::SimMethod) -> &'static str { pub(super) fn ltm_non_euler_diagnostic_message(method: datamodel::SimMethod) -> String { format!( "LTM (Loops That Matter) analysis requires Euler integration, but this model uses \ - {}. The flow-to-stock link-score formula assumes the Euler update \ - `stock(t) = stock(t-1) + dt * flow(t-1)`, so its numerator \ - `PREVIOUS(flow) - PREVIOUS(PREVIOUS(flow))` only aligns to the causal interval that \ - drove the stock change under Euler. Higher-order integrators sub-step the stock \ - update, so the link scores would be mathematically meaningless. Switch the \ - integration method to Euler, or disable LTM analysis.", + {}. Switch the integration method to Euler, or disable LTM analysis.", sim_method_display_name(method), ) } @@ -2026,29 +2011,33 @@ pub fn model_ltm_variables( }); } - // Freeze helpers (GH #995) are minted per referencing partial with - // content-derived names, so the same frozen slice reached from several - // link scores emits byte-identical duplicates. Collapse them to one -- - // duplicate names would mint colliding layout slots -- keeping the first - // occurrence; a same-name pair with DIFFERENT content would be a - // naming-scheme bug, so it is a debug panic rather than silently kept. + // A score's companion variables are minted per score with content-derived + // names -- a freeze helper (GH #995) once per partial that references the + // frozen slice, a stock's net-flow aux once per flow of the stock -- so + // the same companion reached from several link scores emits byte-identical + // duplicates. Collapse them to one -- duplicate names would mint colliding + // layout slots -- keeping the first occurrence; a same-name pair with + // DIFFERENT content would be a naming-scheme bug, so it is a debug panic + // rather than silently kept. { - let mut seen_freeze: HashMap = HashMap::new(); + let mut seen_companion: HashMap = HashMap::new(); vars.retain(|v| { - if !v.name.starts_with(crate::ltm_augment::FREEZE_HELPER_PREFIX) { + if !v.name.starts_with(crate::ltm_augment::FREEZE_HELPER_PREFIX) + && !v.name.starts_with(crate::ltm_augment::NET_FLOW_PREFIX) + { return true; } - match seen_freeze.get(&v.name) { + match seen_companion.get(&v.name) { Some(first) => { debug_assert!( first == v, - "freeze helper name collision with differing content: {}", + "companion variable name collision with differing content: {}", v.name ); false } None => { - seen_freeze.insert(v.name.clone(), v.clone()); + seen_companion.insert(v.name.clone(), v.clone()); true } } @@ -2058,14 +2047,12 @@ pub fn model_ltm_variables( // Sort by evaluation-order category so the VM's sequential flow // evaluation respects the dependency chain: composites reference paths // which reference loop scores which reference link scores, and link - // scores referencing an aggregate node OR a freeze helper read its - // current-step value, so those fragments must run first. Within each - // category, sort lexically for determinism. (`compute_layout` section 3 - // re-sorts LTM vars purely by name -- `$⁚ltm⁚agg⁚{n}` < - // `$⁚ltm⁚freeze⁚...` < `$⁚ltm⁚link_score⁚...` lexically, so both get - // their layout slots before any consumer there too -- but the runlist - // order is what the same-timestep ordering hazard turns on, and that - // comes from this sort.) + // scores referencing an aggregate node, a freeze helper OR a stock's + // net-flow aux read its current-step value, so those fragments must run + // first. Within each category, sort lexically for determinism. + // (`compute_layout` section 3 re-sorts LTM vars purely by name; layout + // order assigns slots, but the runlist order is what the same-timestep + // ordering hazard turns on, and that comes from this sort.) vars.sort_by(|a, b| { fn category(name: &str) -> u8 { // The agg check uses the `$⁚ltm⁚agg⁚` *prefix*, not a substring @@ -2075,10 +2062,12 @@ pub fn model_ltm_variables( // agg aux it references. if crate::ltm_agg::is_synthetic_agg_name(name) { 0 // aggregate nodes: before everything that may reference them - } else if name.starts_with(crate::ltm_augment::FREEZE_HELPER_PREFIX) { - // Array-freeze helpers (GH #995): pure `PREVIOUS` reads of - // model variables, referenced by link scores at their - // current-step value -- run before every score. + } else if name.starts_with(crate::ltm_augment::FREEZE_HELPER_PREFIX) + || name.starts_with(crate::ltm_augment::NET_FLOW_PREFIX) + { + // Array-freeze helpers (GH #995) and net-flow auxes: pure + // reads of model variables, referenced by link scores at + // their current-step value -- run before every score. 1 } else if name.contains("\u{205A}composite\u{205A}") { 5 diff --git a/src/simlin-engine/src/db/ltm_char_golden/agg_nested_reducer.txt b/src/simlin-engine/src/db/ltm_char_golden/agg_nested_reducer.txt index 21d0d7086..1cc5b9765 100644 --- a/src/simlin-engine/src/db/ltm_char_golden/agg_nested_reducer.txt +++ b/src/simlin-engine/src/db/ltm_char_golden/agg_nested_reducer.txt @@ -3,7 +3,7 @@ scalar: if (TIME = INITIAL_TIME) then 0 else if ((growth[region·boston] - PREVI $⁚ltm⁚link_score⁚$⁚ltm⁚agg⁚0→growth[nyc] scalar: if (TIME = INITIAL_TIME) then 0 else if ((growth[region·nyc] - PREVIOUS(growth[region·nyc])) = 0) OR (("$⁚ltm⁚agg⁚0" - PREVIOUS("$⁚ltm⁚agg⁚0")) = 0) then 0 else SAFEDIV(((sum(PREVIOUS(matrix[region·nyc, *]) * PREVIOUS(other[region·nyc, *]) * "$⁚ltm⁚agg⁚0") * 0.001) - PREVIOUS(growth[region·nyc])), ABS((growth[region·nyc] - PREVIOUS(growth[region·nyc]))), 0) * SIGN(("$⁚ltm⁚agg⁚0" - PREVIOUS("$⁚ltm⁚agg⁚0"))) $⁚ltm⁚link_score⁚growth→stock dims=[Region] -a2a[Region]: if (TIME = INITIAL_TIME) OR (PREVIOUS(TIME, INITIAL_TIME) = INITIAL_TIME) then 0 else ABS(SAFEDIV((time_step * (PREVIOUS(growth[Region]) - PREVIOUS(PREVIOUS(growth[Region])))), ((stock[Region] - PREVIOUS(stock[Region])) - (PREVIOUS(stock[Region]) - PREVIOUS(PREVIOUS(stock[Region])))), 0)) +a2a[Region]: if (TIME = INITIAL_TIME) then 0 else if (("$⁚ltm⁚net⁚stock"[Region] - PREVIOUS("$⁚ltm⁚net⁚stock"[Region])) = 0) OR ((growth[Region] - PREVIOUS(growth[Region])) = 0) then 0 else SAFEDIV((growth[Region] - PREVIOUS(growth[Region])), ABS(("$⁚ltm⁚net⁚stock"[Region] - PREVIOUS("$⁚ltm⁚net⁚stock"[Region]))), 0) * SIGN((growth[Region] - PREVIOUS(growth[Region]))) $⁚ltm⁚link_score⁚pop[boston]→$⁚ltm⁚agg⁚0 scalar: if (TIME = INITIAL_TIME) then 0 else if (("$⁚ltm⁚agg⁚0" - PREVIOUS("$⁚ltm⁚agg⁚0")) = 0) OR ((pop[region·boston] - PREVIOUS(pop[region·boston])) = 0) then 0 else SAFEDIV((PREVIOUS("$⁚ltm⁚agg⁚0") + (pop[region·boston] - PREVIOUS(pop[region·boston])) - PREVIOUS("$⁚ltm⁚agg⁚0")), ABS(("$⁚ltm⁚agg⁚0" - PREVIOUS("$⁚ltm⁚agg⁚0"))), 0) * SIGN((pop[region·boston] - PREVIOUS(pop[region·boston]))) $⁚ltm⁚link_score⁚pop[nyc]→$⁚ltm⁚agg⁚0 diff --git a/src/simlin-engine/src/db/ltm_char_golden/arrayed_agg_to_target.txt b/src/simlin-engine/src/db/ltm_char_golden/arrayed_agg_to_target.txt index e21230e96..8c620da54 100644 --- a/src/simlin-engine/src/db/ltm_char_golden/arrayed_agg_to_target.txt +++ b/src/simlin-engine/src/db/ltm_char_golden/arrayed_agg_to_target.txt @@ -11,4 +11,4 @@ scalar: if (TIME = INITIAL_TIME) then 0 else if (("$⁚ltm⁚agg⁚0"[nyc] - PRE $⁚ltm⁚link_score⁚matrix[nyc,b]→$⁚ltm⁚agg⁚0[nyc] scalar: if (TIME = INITIAL_TIME) then 0 else if (("$⁚ltm⁚agg⁚0"[nyc] - PREVIOUS("$⁚ltm⁚agg⁚0"[nyc])) = 0) OR ((matrix[region·nyc,dim2·b] - PREVIOUS(matrix[region·nyc,dim2·b])) = 0) then 0 else SAFEDIV((PREVIOUS("$⁚ltm⁚agg⁚0"[nyc]) + (matrix[region·nyc,dim2·b] - PREVIOUS(matrix[region·nyc,dim2·b])) - PREVIOUS("$⁚ltm⁚agg⁚0"[nyc])), ABS(("$⁚ltm⁚agg⁚0"[nyc] - PREVIOUS("$⁚ltm⁚agg⁚0"[nyc]))), 0) * SIGN((matrix[region·nyc,dim2·b] - PREVIOUS(matrix[region·nyc,dim2·b]))) $⁚ltm⁚link_score⁚outflow→stock dims=[Region] -a2a[Region]: if (TIME = INITIAL_TIME) OR (PREVIOUS(TIME, INITIAL_TIME) = INITIAL_TIME) then 0 else ABS(SAFEDIV((time_step * (PREVIOUS(outflow[Region]) - PREVIOUS(PREVIOUS(outflow[Region])))), ((stock[Region] - PREVIOUS(stock[Region])) - (PREVIOUS(stock[Region]) - PREVIOUS(PREVIOUS(stock[Region])))), 0)) +a2a[Region]: if (TIME = INITIAL_TIME) then 0 else if (("$⁚ltm⁚net⁚stock"[Region] - PREVIOUS("$⁚ltm⁚net⁚stock"[Region])) = 0) OR ((outflow[Region] - PREVIOUS(outflow[Region])) = 0) then 0 else SAFEDIV((outflow[Region] - PREVIOUS(outflow[Region])), ABS(("$⁚ltm⁚net⁚stock"[Region] - PREVIOUS("$⁚ltm⁚net⁚stock"[Region]))), 0) * SIGN((outflow[Region] - PREVIOUS(outflow[Region]))) diff --git a/src/simlin-engine/src/db/ltm_tests.rs b/src/simlin-engine/src/db/ltm_tests.rs index 14a3e54b2..d05bc6048 100644 --- a/src/simlin-engine/src/db/ltm_tests.rs +++ b/src/simlin-engine/src/db/ltm_tests.rs @@ -152,10 +152,11 @@ fn ltm_capture_helpers_compile_exactly_the_phases_their_kind_demands() { } } -/// The helpers the score generator actually mints -- the flow-to-stock -/// score's nested `PREVIOUS` captures -- are flow-only: no initial fragment, -/// since `PREVIOUS`'s fallback covers every read before the first step -/// commits. +/// The helpers the score generator actually mints -- here the frozen +/// dynamic-index read `PREVIOUS(weight[PREVIOUS(idx, idx)])` of the +/// `level -> growth` partial, whose argument is no static slot -- are +/// flow-only: no initial fragment, since `PREVIOUS`'s fallback covers every +/// read before the first step commits. #[test] fn generated_ltm_capture_helpers_are_flow_only() { use super::compile::{compile_ltm_implicit_var_fragment, ltm_helper_phases_present}; @@ -163,8 +164,11 @@ fn generated_ltm_capture_helpers_are_flow_only() { use crate::capture::CaptureKind; let project = TestProject::new("generated_ltm_capture_phases") + .named_dimension("d", &["d1", "d2"]) + .array_with_ranges("weight[d]", vec![("d1", "1"), ("d2", "2")]) + .aux("idx", "1", None) .stock("level", "100", &["growth"], &[], None) - .flow("growth", "level * rate", None) + .flow("growth", "level * rate * weight[idx]", None) .aux("rate", "0.1", None) .build_datamodel(); let db = SimlinDb::default(); @@ -177,7 +181,7 @@ fn generated_ltm_capture_helpers_are_flow_only() { .collect(); assert!( !captures.is_empty(), - "the flow-to-stock score mints nested PREVIOUS captures" + "the level -> growth score's frozen dynamic-index read mints a capture" ); for meta in captures { assert_eq!( @@ -2698,33 +2702,28 @@ fn delay3_ramp_project(in_submodel: bool) -> datamodel::Project { } /// The `input -> stock` link score of a `DELAY3` instance reads the bound -/// port's two-step lag through a capture helper, at the main level and inside -/// a sub-model alike. +/// port through `PREVIOUS(input)`, at the main level and inside a sub-model +/// alike. /// -/// The flow-to-stock score (`ltm_augment::generate_flow_to_stock_equation`, -/// Schoenberg et al. 2023 Eq. 3) is -/// `|dt * (PREVIOUS(input) - PREVIOUS(PREVIOUS(input)))| / |second difference -/// of stock|`, zero for the first two steps. The nested lag is a capture -/// helper (`..⁚1⁚arg0 = PREVIOUS(input, 0)`) inside the `stdlib⁚delay3` -/// instance, whose body snapshots the instance's BOUND port `input`; lowering -/// resolves that port to its own slot (`Context::snapshot_storage`), so the -/// helper is `0, 3, 4, 5, 6` (the fallback, then the lagged ramp). +/// The score (`ltm_augment::generate_flow_to_stock_equation`) is +/// `|Δinput / Δnet|` over the instance's net-flow aux +/// `$⁚ltm⁚net⁚stock = input - flow_1`. `PREVIOUS(input)` snapshots the +/// instance's BOUND port `input`; lowering resolves that port to its own slot +/// (`Context::snapshot_storage`), so the lag reads the ramp one step back. /// /// Derived from the template (`stdlib/delay3.stmx`: `stock` inflow `input`, /// outflow `flow_1 = stock / (delay_time / 3)`, `delay_time = 2`, so -/// `stock(0) = 3 * 2/3 = 2`): `stock = 2, 2, 3, 3.5, 4.25`, the numerator is -/// `1` from t=2 on (the ramp rises by 1 per step), and the second differences -/// are `1, -0.5, 0.25`, so the score is `1, 2, 4` at t=2..4. +/// `stock(0) = 3 * 2/3 = 2`) with `input = 3, 4, 5, 6, 7`: +/// `stock = 2, 2, 3, 3.5, 4.25`, `flow_1 = 3, 3, 4.5, 5.25, 6.375`, so +/// `net = 0, 1, 0.5, 0.75, 0.625`; `Δinput` is `1` at every step and +/// `Δnet = 1, -0.5, 0.25, -0.125`, so the score is `1, 2, 4, 8` at t=1..4. /// -/// A parse that captured the port, or a lowering that could not address it, -/// gave the helper no fragment at all -- the LTM tail appends helpers by -/// bytecode presence, so the score silently read an unwritten 0 for the -/// two-step lag and printed `4, 10, 24` (the numerator degenerated to the -/// one-step-lagged ramp `4, 5, 6`) with no diagnostic. That is the -/// silent-wrong-number class the invariant forbids, which is why the values -/// are pinned here rather than only the helper's presence. +/// A lowering that could not address the bound port would read an unwritten +/// slot for the lag with no diagnostic -- the silent-wrong-number class the +/// invariant forbids -- which is why the values are pinned, with the net aux +/// beside them. #[test] -fn delay3_input_to_stock_link_score_uses_the_two_step_lag_of_the_bound_port() { +fn delay3_input_to_stock_link_score_reads_the_bound_port() { for (in_submodel, prefix) in [(false, ""), (true, "sub\u{00B7}")] { let project = delay3_ramp_project(in_submodel); let db = SimlinDb::default(); @@ -2743,9 +2742,9 @@ fn delay3_input_to_stock_link_score_uses_the_two_step_lag_of_the_bound_port() { results.get(&name).cloned().unwrap_or_else(|| { let candidates: Vec<&String> = results .keys() - .filter(|k| k.contains("input\u{2192}stock")) + .filter(|k| k.contains("delay3\u{00B7}$\u{205A}ltm")) .collect(); - panic!("{name} not in results; input->stock names: {candidates:?}") + panic!("{name} not in results; the instance's LTM names: {candidates:?}") }) }; let score = series(format!( @@ -2753,19 +2752,15 @@ fn delay3_input_to_stock_link_score_uses_the_two_step_lag_of_the_bound_port() { )); assert_eq!( score, - vec![0.0, 0.0, 1.0, 2.0, 4.0], - "in_submodel={in_submodel}: the input -> stock link score must use the \ - two-step lag of the bound port (it was 0, 0, 4, 10, 24 when the nested \ - lag's helper silently failed to compile)" + vec![0.0, 1.0, 2.0, 4.0, 8.0], + "in_submodel={in_submodel}: the input -> stock link score is |Δinput / Δnet| \ + with the bound port's one-step lag" ); - let helper = series(format!( - "{prefix}$⁚delayed⁚0⁚delay3\u{00B7}$⁚$⁚ltm⁚link_score⁚input\u{2192}stock⁚1⁚arg0" - )); + let net = series(format!("{prefix}$⁚delayed⁚0⁚delay3\u{00B7}$⁚ltm⁚net⁚stock")); assert_eq!( - helper, - vec![0.0, 3.0, 4.0, 5.0, 6.0], - "in_submodel={in_submodel}: the nested-lag helper holds PREVIOUS(input, 0) \ - of the bound port" + net, + vec![0.0, 1.0, 0.5, 0.75, 0.625], + "in_submodel={in_submodel}: the instance's net-flow aux is input - flow_1" ); } } diff --git a/src/simlin-engine/src/db/ltm_unified_tests.rs b/src/simlin-engine/src/db/ltm_unified_tests.rs index a5fc0fd49..805f4a18d 100644 --- a/src/simlin-engine/src/db/ltm_unified_tests.rs +++ b/src/simlin-engine/src/db/ltm_unified_tests.rs @@ -1149,14 +1149,21 @@ fn test_model_ltm_fragment_diagnostics_emits_warning() { /// The failure is injected via `LtmFragmentFailureGuard` (the same GH #547 /// mechanism the synthetic-var tests use, extended to the helper compile /// path) so this test survives every real helper-compile bug being fixed. -/// The fixture's flow-to-stock link score genuinely mints `arg0` helpers -/// (its `PREVIOUS(PREVIOUS(cap_flow))` nested capture), pinned below so the -/// guard pattern cannot silently match nothing. +/// The fixture's `cap_stock -> aux_0` link score genuinely mints an `arg0` +/// helper (the frozen dynamic-index read `PREVIOUS(weight[PREVIOUS(idx, +/// idx)])`), pinned below so the guard pattern cannot silently match nothing. #[test] fn test_model_ltm_fragment_diagnostics_covers_implicit_helpers() { use crate::db::{LtmFragmentFailureGuard, model_ltm_implicit_var_info}; - let project = build_chain_scc_project("implicit_frag_fail", 5); + let project = crate::test_common::TestProject::new("implicit_frag_fail") + .named_dimension("d", &["d1", "d2"]) + .array_with_ranges("weight[d]", vec![("d1", "1"), ("d2", "2")]) + .aux("idx", "1", None) + .aux("aux_0", "cap_stock * weight[idx]", None) + .flow("cap_flow", "aux_0", None) + .stock("cap_stock", "0", &["cap_flow"], &[], None) + .build_datamodel(); let db = SimlinDb::default(); let (source_project, model) = { let sync = sync_from_datamodel(&db, &project); @@ -3888,15 +3895,9 @@ fn test_ltm_var_name_index_matches_vars() { // ── GH #486: LTM requires Euler integration ───────────────────────────── // -// The 2023 flow-to-stock link-score formula -// (`PREVIOUS(flow) - PREVIOUS(PREVIOUS(flow))`) only aligns the numerator to -// the causal interval that drove the stock change from t-1 to t under Euler -// integration. RK2/RK4 sub-step the stock update, so that alignment breaks -// and the resulting link scores are mathematically meaningless. Non-Euler -// integration IS honored at runtime (the VM and wasm backends both have -// distinct RK2/RK4 stepping loops), so the wrong scores would look plausible -// but be wrong -- a silent-correctness hazard. The engine therefore rejects -// the combination at sim-compile time. +// The engine keeps LTM on Euler-stepped runs and rejects the combination of +// the overlay with RK2/RK4 at sim-compile time, gated on a flow-to-stock +// score actually being emitted (`model_emits_flow_to_stock_score`). // // CRITICAL granularity (the multi-model probes below): the VM has a SINGLE // global integration method, taken from the MAIN (root) model's diff --git a/src/simlin-engine/src/db/ltm_value_golden/value_gate.txt b/src/simlin-engine/src/db/ltm_value_golden/value_gate.txt index ba1376562..bf7a11b0a 100644 --- a/src/simlin-engine/src/db/ltm_value_golden/value_gate.txt +++ b/src/simlin-engine/src/db/ltm_value_golden/value_gate.txt @@ -1,8 +1,8 @@ $⁚ltm⁚link_score⁚alt[a1]→growth[0] 0.000000000000e0 6.666666666667e-1 6.600660066007e-1 6.535306996046e-1 6.470600986184e-1 6.406535629885e-1 $⁚ltm⁚link_score⁚alt[a1]→growth[1] 0.000000000000e0 1.000000000000e0 1.000000000000e0 1.000000000000e0 1.000000000000e0 1.000000000000e0 $⁚ltm⁚link_score⁚alt[a1]→growth[2] 0.000000000000e0 0.000000000000e0 0.000000000000e0 0.000000000000e0 0.000000000000e0 0.000000000000e0 -$⁚ltm⁚link_score⁚growth→pop[0] 0.000000000000e0 0.000000000000e0 1.000000000000e0 9.999999999998e-1 1.000000000000e0 9.999999999999e-1 -$⁚ltm⁚link_score⁚growth→pop[1] 0.000000000000e0 0.000000000000e0 9.999999999997e-1 9.999999999999e-1 1.000000000000e0 1.000000000000e0 +$⁚ltm⁚link_score⁚growth→pop[0] 0.000000000000e0 1.000000000000e0 1.000000000000e0 1.000000000000e0 1.000000000000e0 1.000000000000e0 +$⁚ltm⁚link_score⁚growth→pop[1] 0.000000000000e0 1.000000000000e0 1.000000000000e0 1.000000000000e0 1.000000000000e0 1.000000000000e0 $⁚ltm⁚link_score⁚growth→pop[2] 0.000000000000e0 0.000000000000e0 0.000000000000e0 0.000000000000e0 0.000000000000e0 0.000000000000e0 $⁚ltm⁚link_score⁚pop[boston]→pop_total[0] 0.000000000000e0 1.578947368421e-1 1.649162354629e-1 1.715697797158e-1 1.778798792994e-1 1.838689119179e-1 $⁚ltm⁚link_score⁚pop[la]→pop_total[0] 0.000000000000e0 3.157894736842e-1 3.073927967621e-1 2.993785051921e-1 2.917210936525e-1 2.843972770067e-1 @@ -12,5 +12,8 @@ $⁚ltm⁚link_score⁚pop[nyc]→growth[2] 0.000000000000e0 0.000000000000e0 0. $⁚ltm⁚link_score⁚pop[nyc]→pop_total[0] 0.000000000000e0 5.263157894737e-1 5.276909677750e-1 5.290517150920e-1 5.303990270480e-1 5.317338110755e-1 $⁚ltm⁚link_score⁚pop_total→alt[a1][0] 0.000000000000e0 1.000000000000e0 1.000000000000e0 1.000000000000e0 1.000000000000e0 1.000000000000e0 $⁚ltm⁚link_score⁚pop_total→alt[a2][0] 0.000000000000e0 1.000000000000e0 1.000000000000e0 1.000000000000e0 1.000000000000e0 1.000000000000e0 -$⁚ltm⁚loop_score⁚r1[0] 0.000000000000e0 0.000000000000e0 1.649162354628e-1 1.715697797158e-1 1.778798792994e-1 1.838689119179e-1 -$⁚ltm⁚loop_score⁚r2[0] 0.000000000000e0 0.000000000000e0 1.000000000000e0 9.999999999998e-1 1.000000000000e0 9.999999999999e-1 +$⁚ltm⁚loop_score⁚r1[0] 0.000000000000e0 1.578947368421e-1 1.649162354629e-1 1.715697797158e-1 1.778798792994e-1 1.838689119179e-1 +$⁚ltm⁚loop_score⁚r2[0] 0.000000000000e0 1.000000000000e0 1.000000000000e0 1.000000000000e0 1.000000000000e0 1.000000000000e0 +$⁚ltm⁚net⁚pop[0] 1.000000000000e-1 1.030000000000e-1 1.060300000000e-1 1.090903000000e-1 1.121812030000e-1 1.153030150300e-1 +$⁚ltm⁚net⁚pop[1] 3.000000000000e-2 3.219000000000e-2 3.438519000000e-2 3.658560519000e-2 3.879128109519e-2 4.100225357929e-2 +$⁚ltm⁚net⁚pop[2] 6.000000000000e-2 6.000000000000e-2 6.000000000000e-2 6.000000000000e-2 6.000000000000e-2 6.000000000000e-2 diff --git a/src/simlin-engine/src/db/prev_init_tests.rs b/src/simlin-engine/src/db/prev_init_tests.rs index d10721fc2..fd02a8643 100644 --- a/src/simlin-engine/src/db/prev_init_tests.rs +++ b/src/simlin-engine/src/db/prev_init_tests.rs @@ -728,18 +728,14 @@ fn test_ltm_bare_element_subscripts_no_helpers() { ); let info = model_ltm_implicit_var_info(&db, source_model, sync.project); - // The only helpers allowed are the flow-to-stock link score's nested - // PREVIOUS(PREVIOUS(...)) captures (semantically necessary: the VM keeps - // one step of history, so a two-step lag needs a helper that re-lags a - // lagged value). Bare-element subscripts (`rate[b2]`, `pop[b2]`) must not - // synthesize any. - let non_flow_to_stock: Vec<&String> = info - .keys() - .filter(|name| !name.contains("grow\u{2192}pop")) - .collect(); + // Bare-element subscripts (`rate[b2]`, `pop[b2]`) are static slots and + // must synthesize no helper; every other generated read here is a static + // slot too (a flow-to-stock score reads only the flow and the stock's + // net-flow aux), so the model mints no LTM helper at all. assert!( - non_flow_to_stock.is_empty(), - "only the flow-to-stock nested-PREVIOUS helpers may remain; unexpected: {non_flow_to_stock:?}" + info.is_empty(), + "a bare-element subscript must not synthesize a helper; unexpected: {:?}", + info.keys().collect::>() ); } diff --git a/src/simlin-engine/src/ltm_augment.rs b/src/simlin-engine/src/ltm_augment.rs index 7205078d3..c5f2f30f7 100644 --- a/src/simlin-engine/src/ltm_augment.rs +++ b/src/simlin-engine/src/ltm_augment.rs @@ -6,8 +6,9 @@ //! //! This module generates synthetic variables for Loops That Matter (LTM) analysis. //! The generated equations use the intrinsic two-argument `PREVIOUS(value, initial)` -//! function. First- and second-timestep guards are expressed explicitly with -//! `TIME = INITIAL_TIME` and `PREVIOUS(TIME, INITIAL_TIME) = INITIAL_TIME`. +//! function. Every score reads one step of history, so the first-timestep guard +//! is expressed explicitly with `TIME = INITIAL_TIME` and every score's first +//! value is at the step after the start. use crate::ast::{Expr0, IndexExpr0, print_eqn}; use crate::builtins::UntypedBuiltinFn; @@ -3100,7 +3101,8 @@ fn target_iterated_dim_names_canonical(to_var: &Variable) -> Vec { /// other-dep correspondence check; `None` keeps the historical permissive /// collapse for every dep (legacy / db-less callers). /// -/// Flow-to-stock links use a fixed structural formula and ignore `shape`, +/// A flow-to-stock link is the closed-form partial of the stock's net-flow +/// aux ([`generate_flow_to_stock_equation`]) and ignores `shape`, /// `source_dim_elements`, `dim_ctx`, and `dep_dims`. #[allow(clippy::too_many_arguments)] // threads the link-score generation context pub(crate) fn generate_link_score_equation_for_link( @@ -3134,8 +3136,8 @@ pub(crate) fn generate_link_score_equation_for_link( /// Returns `Err([`PartialEquationError`])` when the target's equation text /// cannot be parsed for the ceteris-paribus partial (GH #311); the /// db-bearing caller turns this into a `Warning` and skips the variable. -/// The flow-to-stock branch uses a fixed structural formula with no parse, -/// so it is infallible and always returns `Ok`. +/// The flow-to-stock branch is a closed-form partial with no parse, so it +/// is infallible and always returns `Ok`. #[allow(clippy::too_many_arguments)] // threads the link-score generation context fn generate_link_score_equation( from: &Ident, @@ -3157,14 +3159,14 @@ fn generate_link_score_equation( // Binding `flow_var` here -- rather than computing an `is_flow_to_stock` // bool and re-fetching the (proven-present) flow variable -- lets the // generator take a plain `&Variable`. - if to_var.is_stock() - && let Some(flow_var) = all_vars.get(from) - && matches!(flow_var.kind, VarKind::Aux { is_flow: true, .. }) + if let Some(flow_var) = all_vars + .get(from) + .filter(|fv| flow_to_stock_wiring(Some(fv), to_var).is_some()) { - // Flow-to-stock uses a fixed structural formula -- no AST parse, - // so neither `shape` nor `source_dim_elements` matter here. The - // flow variable is passed in only for its declared dimensions, so - // an arrayed flow can be referenced with an explicit subscript. + // The flow-to-stock score is closed-form -- no AST parse, so + // neither `shape` nor `source_dim_elements` matter here. The flow + // variable is passed in only for its declared dimensions, so an + // arrayed flow can be referenced with an explicit subscript. Ok(generate_flow_to_stock_equation( from.as_str(), to.as_str(), @@ -3223,7 +3225,9 @@ fn link_score_guard_form(partial_eq: &str, target_ref: &str, source_ref: &str) - /// the changed-first form (numerator `(partial - PREVIOUS(target))`), the /// changed-last form (numerator `(target - frozen)` -- the /// [`shaped_guard_form_text`] fallback and -/// [`generate_scalar_feeder_to_agg_equation`]), and any future attribution +/// [`generate_scalar_feeder_to_agg_equation`]), the flow-to-stock score +/// (numerator `+/-Δflow`, the folded partial of the stock's net-flow aux -- +/// [`generate_flow_to_stock_equation`]), and any future attribution /// convention with the same guard structure. fn link_score_guard_form_with_numerator( numerator: &str, @@ -3774,105 +3778,154 @@ fn dimension_subscript_suffix(var: &Variable) -> String { } } -/// Generate flow-to-stock link score equation. -/// -/// The structural inflow/outflow formula has no per-element equation -/// text -- the compiler applies it element-wise when the stock and flow -/// are arrayed -- so the result is `Equation::Scalar` for a scalar stock -/// and `Equation::ApplyToAll(stock_dims, _)` for an arrayed stock (the -/// shared formula evaluated per element). -/// -/// For an arrayed stock every stock/flow reference is emitted with an -/// explicit dimension subscript (`stock[Dim]`, `flow[Dim]`) rather than a -/// bare arrayed name. A bare arrayed name nested inside -/// `PREVIOUS(PREVIOUS(...))` does not survive the apply-to-all expansion: -/// the inner `PREVIOUS(name)` is an *expression* argument, so -/// `builtins_visitor` routes it through a synthesized *scalar* helper aux -/// whose equation is `PREVIOUS(name, 0)` -- and a bare arrayed name has no -/// scalar meaning, so that helper fragment fails to compile and the LTM -/// fragment compiler silently stubs it to 0 (the score then collapses to -/// a wrong constant -- `1/9` for the canonical pop/growth model instead -/// of the isolated-loop invariant `1`). An explicit subscript keeps every -/// occurrence a scalar per-element access the helper aux can hold. Each -/// variable is subscripted by its *own* declared dimensions; a valid -/// arrayed inflow/outflow shares the stock's dimensions, so those names -/// are all bound by the `ApplyToAll` iteration. A scalar stock/flow has -/// no dimensions, so its references stay bare -- the pre-fix behavior. -/// -/// NOTE: the engine captures a snapshot-only apply-to-all body structurally -/// (an apply-to-all capture over the target's dimensions), so the bare form -/// compiles too and this generator-side subscripting is not load-bearing. -/// It is intentionally retained: the engine fix is a strict superset -/// (an already-subscripted reference stays on the unchanged scalar-helper -/// path), this output is pinned by dedicated tests, and re-baselining the -/// LTM equation text for every arrayed flow-to-stock link score across the -/// corpus would be a broad change with no behavioral benefit. +/// The name prefix of a stock's synthetic net-flow auxiliary, +/// `$⁚ltm⁚net⁚{stock}` -- see [`generate_net_flow_equation`]. +pub(crate) const NET_FLOW_PREFIX: &str = "$\u{205A}ltm\u{205A}net\u{205A}"; + +/// The synthetic net-flow auxiliary of `stock` (canonical name). +pub(crate) fn net_flow_var_name(stock: &str) -> String { + format!("{NET_FLOW_PREFIX}{stock}") +} + +/// A stock's declared `(inflows, outflows)`, as [`flow_to_stock_wiring`] +/// hands them out. +pub(crate) type StockFlows<'a> = (&'a [Ident], &'a [Ident]); + +/// The declared `(inflows, outflows)` of the stock a `(from, to)` link feeds, +/// when `to` is a stock and `from` a flow; `None` for every other link. A +/// stock's only causal in-edges are its wiring (`model_causal_edges`: a +/// stock's equation is its initial value, so nothing else points at it), so +/// `Some` is also "the edge is structural". The one statement of the pair +/// test, read by the generator branch, the flow-to-stock generator (which +/// signs the flow by the side it sits on) and the net-flow aux's emitter. +pub(crate) fn flow_to_stock_wiring<'a>( + from_var: Option<&Variable>, + to_var: &'a Variable, +) -> Option> { + let from_is_flow = + from_var.is_some_and(|v| matches!(v.kind, VarKind::Aux { is_flow: true, .. })); + match &to_var.kind { + VarKind::Stock { + inflows, outflows, .. + } if from_is_flow => Some((inflows, outflows)), + _ => None, + } +} + +/// How a stock's flow is spelled inside the stock-shaped LTM equations (the +/// net-flow aux and the flow-to-stock score): with its own declared +/// dimensions when they are the stock's (`flow[Region]`, a scalar +/// per-element access under the equation's iteration over the stock's +/// axes), else bare. +/// +/// A flow declared over OTHER dimensions than its stock's (`inflow[dimb]` +/// into `level[suba]`, `dimb -> dima` with `suba` inside `dima`) is bare +/// because under the iteration over the stock's dimensions the compiler +/// resolves a bare arrayed name through its implicit subscripts +/// (`get_implicit_subscripts`, the pairing the wiring itself uses to fold +/// the flow into the stock), where `inflow[dimb]` names an axis that +/// iteration does not carry and does not lower. A scalar flow into an +/// arrayed stock is bare too: it broadcasts into every element's net flow. +/// A flow the model could not lower (`None`) is bare as well, so the +/// compiler's own refusal of the flow reaches the fragment diagnostics +/// rather than a guess about its shape. +fn stock_flow_ref(flow: &str, flow_var: Option<&Variable>, stock_var: &Variable) -> String { + let flow_q = quote_ident(flow); + match flow_var { + Some(fv) if target_equation_dims(fv) == target_equation_dims(stock_var) => { + format!("{flow_q}{}", dimension_subscript_suffix(fv)) + } + _ => flow_q, + } +} + +/// The stock's net-flow auxiliary `$⁚ltm⁚net⁚{stock}`: `(inflows) - (outflows)`, +/// shaped like the stock (`Equation::ApplyToAll` over an arrayed stock's +/// dimensions, scalar otherwise), with `0` for a side the stock has no flows +/// on. +/// +/// This is the 2023 paper's implementation option (b) (Schoenberg, Hayward +/// and Eberlein 2023, section 4.4): aggregate every stock's flows into one +/// net flow and score each flow's link into the net flow with the ordinary +/// instantaneous formula, the net flow's own link into the stock being +/// exactly 1. The aux is LTM machinery, not a causal node: the causal graph +/// keeps its `flow -> stock` edges, and this variable appears in no loop and +/// no link. It is evaluated each step ahead of the scores that read it (the +/// evaluation-order sort in `model_ltm_variables`), so a score reads both its +/// current value and `PREVIOUS(net)`. `inflows` and `outflows` are the +/// stock's declared flows, in declaration order, each with its lowered +/// variable when the model has one (see [`stock_flow_ref`]). +pub(crate) fn generate_net_flow_equation( + stock_var: &Variable, + inflows: &[(&str, Option<&Variable>)], + outflows: &[(&str, Option<&Variable>)], +) -> LtmEquation { + let side = |flows: &[(&str, Option<&Variable>)]| -> String { + if flows.is_empty() { + "0".to_string() + } else { + flows + .iter() + .map(|(flow, flow_var)| stock_flow_ref(flow, *flow_var, stock_var)) + .collect::>() + .join(" + ") + } + }; + let text = format!("({}) - ({})", side(inflows), side(outflows)); + link_score_equation_for_target(text, stock_var) +} + +/// Generate the flow-to-stock link score equation. +/// +/// The score is the ordinary instantaneous link score +/// ([`link_score_guard_form_with_numerator`]) of the stock's net-flow aux +/// ([`generate_net_flow_equation`]) with respect to `flow`. Because the aux +/// is a linear sum its ceteris-paribus partial is closed-form -- `Δ_flow net` +/// is `+Δflow` for an inflow and `-Δflow` for an outflow -- so the score is +/// `sign * |Δflow / Δnet|` with the structural polarity (+1 inflow, -1 +/// outflow): the 2023 paper's Eq. 3 (Schoenberg, Hayward and Eberlein 2023, +/// section 4.1), whose denominator `Δ(S_t) - Δ(S_{t-dt})` is `Δnet`. Both +/// deltas are read over `[t - dt, t]`, the window of every other link score, +/// so a loop's link scores all describe one interval and the score is +/// invariant to whether the stock's flows are written separately or as one +/// net flow (the paper's section 4.4). No `dt` appears: the ratio of two +/// flow deltas is dimensionless as written, and an isolated loop scores +/// exactly `+/-1` at every `dt` (`tests/integration/ltm_dt_invariance.rs`). +/// +/// The result is shaped like the stock: `Equation::Scalar` for a scalar +/// stock, `Equation::ApplyToAll(stock_dims, _)` for an arrayed one, the flow +/// spelled per [`stock_flow_ref`] and the net aux subscripted by the stock's +/// own dimensions. fn generate_flow_to_stock_equation( flow: &str, stock: &str, flow_var: &Variable, stock_var: &Variable, ) -> LtmEquation { - // Check if this flow is an inflow or outflow - let is_inflow = if let VarKind::Stock { inflows, .. } = &stock_var.kind { - inflows.iter().any(|f| f.as_str() == flow) - } else { - true // Default to inflow + // Polarity is structural: an inflow raises the stock, an outflow lowers + // it. A flow the stock does not list as an inflow is an outflow. + let Some((inflows, _)) = flow_to_stock_wiring(Some(flow_var), stock_var) else { + unreachable!( + "generate_flow_to_stock_equation is reached only through \ + `flow_to_stock_wiring`, which admits a flow into a stock alone" + ); }; - - let sign = if is_inflow { "" } else { "-" }; - - // Reference an arrayed stock/flow by its own declared dimensions so - // every occurrence is a scalar per-element access; see the function - // doc for why a bare arrayed name breaks the nested-PREVIOUS terms. - // For a scalar stock/flow the suffix is empty and the references stay - // bare, exactly as before. A flow declared over OTHER dimensions than - // its stock's (`inflow[dimb]` into `level[suba]`, `dimb -> dima` with - // `suba` inside `dima`) is spelled bare too: under the score's - // iteration over the stock's dimensions the compiler resolves a bare - // arrayed name through its implicit subscripts (`get_implicit_subscripts`, - // the pairing the wiring itself uses to fold the flow into the stock), - // where `inflow[dimb]` names an axis that iteration does not carry and - // does not lower. - let stock_ref = format!("{stock}{}", dimension_subscript_suffix(stock_var)); - let flow_ref = if target_equation_dims(flow_var) == target_equation_dims(stock_var) { - format!("{flow}{}", dimension_subscript_suffix(flow_var)) + let is_inflow = inflows.iter().any(|f| f.as_str() == flow); + let flow_ref = stock_flow_ref(flow, Some(flow_var), stock_var); + let net_ref = format!( + "{}{}", + quote_ident(&net_flow_var_name(stock)), + dimension_subscript_suffix(stock_var) + ); + // The changed-first numerator `net(flow_t, others_{t-1}) - net_{t-1}` of + // a linear sum, folded: the other flows cancel and only this flow's own + // delta remains, signed by the side of the sum it sits on. + let numerator = if is_inflow { + format!("({flow_ref} - PREVIOUS({flow_ref}))") } else { - flow.to_string() + format!("(PREVIOUS({flow_ref}) - {flow_ref})") }; - - // Per the corrected 2023 formula (Schoenberg et al., Eq. 3): - // LS(inflow -> S) = |Delta(i) / (Delta(S_t) - Delta(S_{t-dt}))| * (+1) - // LS(outflow -> S) = |Delta(o) / (Delta(S_t) - Delta(S_{t-dt}))| * (-1) - // - // The polarity is structural (fixed +1/-1), not dynamic. ABS ensures - // the magnitude is always positive; the sign is applied outside. - // - // The numerator uses PREVIOUS values to align timing with the denominator. - // At time t, the flow at t-1 (PREVIOUS(flow)) is what drove the stock change from t-1 to t. - // We measure the change in that causal flow: flow(t-1) - flow(t-2). - // - // The `time_step` factor makes the score the dimensionally-correct - // discretization of the continuous form `|di/dt / d^2S/dt^2|` - // (Schoenberg et al. 2023, Eq. 6): the denominator below is the - // second-order stock change `dt * (netflow(t-1) - netflow(t-2))`, which - // already carries one `dt`; the raw flow delta in the numerator carries - // none, so without this factor the score is `1/dt` too large and the - // error compounds once per flow-to-stock link in a loop. The published - // Eq. 3 omits `dt` because every worked example in the papers uses dt=1. - let numerator = - format!("(time_step * (PREVIOUS({flow_ref}) - PREVIOUS(PREVIOUS({flow_ref}))))"); - let denominator = format!( - "(({stock_ref} - PREVIOUS({stock_ref})) - (PREVIOUS({stock_ref}) - PREVIOUS(PREVIOUS({stock_ref}))))" - ); - - // Return 0 for the first two timesteps when we don't have enough history for second-order differences - let text = format!( - "if \ - (TIME = INITIAL_TIME) OR (PREVIOUS(TIME, INITIAL_TIME) = INITIAL_TIME) \ - then 0 \ - else {sign}ABS(SAFEDIV({numerator}, {denominator}, 0))" - ); + let text = link_score_guard_form_with_numerator(&numerator, &net_ref, &flow_ref); link_score_equation_for_target(text, stock_var) } diff --git a/src/simlin-engine/src/ltm_augment_tests.rs b/src/simlin-engine/src/ltm_augment_tests.rs index 9d15b18b6..f4f463960 100644 --- a/src/simlin-engine/src/ltm_augment_tests.rs +++ b/src/simlin-engine/src/ltm_augment_tests.rs @@ -3984,17 +3984,43 @@ fn flow_to_stock_test_flow(ident: &str, eqn: Equation) -> Variable { } } -/// LTM deep-review Finding 2: for an *arrayed* stock the flow-to-stock -/// link-score equation must reference the stock and flow with explicit -/// dimension subscripts. A *bare* arrayed name nested inside -/// `PREVIOUS(PREVIOUS(...))` is routed through a synthesized *scalar* -/// helper aux (see `builtins_visitor`) that cannot hold an arrayed -/// value -- the fragment then fails to compile and the LTM compiler -/// silently stubs it to 0, collapsing the score to a wrong constant -/// (`1/9` for the canonical pop/growth model instead of the -/// isolated-loop invariant `1`). -#[test] -fn test_flow_to_stock_arrayed_subscripts_references() { +/// The `"{net}"` reference a flow-to-stock score reads: the stock's net-flow +/// aux, quoted (its name needs quoting), with `suffix` (`""` or `[Dim]`). +fn net_ref(stock: &str, suffix: &str) -> String { + format!("\"{}\"{suffix}", net_flow_var_name(stock)) +} + +/// The text of a scalar stock's flow-to-stock score: the standard guard form +/// of the stock's net-flow aux with respect to the flow, the numerator the +/// flow's own delta (the folded partial of a linear sum; an inflow keeps its +/// sign), every reference bare. Nothing reads the stock, and no `dt` appears. +#[test] +fn flow_to_stock_scalar_inflow_is_the_net_flow_partial() { + let stock = + flow_to_stock_test_stock("s", Equation::Scalar("100".to_string()), &["births"], &[]); + let flow = flow_to_stock_test_flow("births", Equation::Scalar("s * 0.1".to_string())); + + let equation = generate_flow_to_stock_equation("births", "s", &flow, &stock); + let LtmEquation::Scalar(arm) = &equation else { + panic!("scalar stock must yield Equation::Scalar; got: {equation:?}"); + }; + let net = net_ref("s", ""); + assert_eq!( + &*arm.text, + format!( + "if (TIME = INITIAL_TIME) then 0 else if (({net} - PREVIOUS({net})) = 0) OR \ + ((births - PREVIOUS(births)) = 0) then 0 else SAFEDIV((births - \ + PREVIOUS(births)), ABS(({net} - PREVIOUS({net}))), 0) * SIGN((births - \ + PREVIOUS(births)))" + ) + ); +} + +/// An arrayed stock's score is `Equation::ApplyToAll` over the stock's +/// dimensions, the flow and the net aux both subscripted by them (a scalar +/// per-element access under the iteration). +#[test] +fn flow_to_stock_arrayed_inflow_subscripts_flow_and_net() { let stock = flow_to_stock_test_stock( "pop", Equation::ApplyToAll(vec!["region".to_string()], "100".to_string()), @@ -4007,95 +4033,185 @@ fn test_flow_to_stock_arrayed_subscripts_references() { ); let equation = generate_flow_to_stock_equation("growth", "pop", &flow, &stock); - let text = match &equation { - LtmEquation::ApplyToAll(dims, arm) => { - assert_eq!(dims, &vec!["region".to_string()]); - &arm.text - } - other => panic!("arrayed stock must yield LtmEquation::ApplyToAll; got: {other:?}"), + let LtmEquation::ApplyToAll(dims, arm) = &equation else { + panic!("arrayed stock must yield LtmEquation::ApplyToAll; got: {equation:?}"); }; + assert_eq!(dims, &vec!["region".to_string()]); + let net = net_ref("pop", "[region]"); + assert_eq!( + &*arm.text, + format!( + "if (TIME = INITIAL_TIME) then 0 else if (({net} - PREVIOUS({net})) = 0) OR \ + ((growth[region] - PREVIOUS(growth[region])) = 0) then 0 else \ + SAFEDIV((growth[region] - PREVIOUS(growth[region])), ABS(({net} - \ + PREVIOUS({net}))), 0) * SIGN((growth[region] - PREVIOUS(growth[region])))" + ) + ); +} - // Every stock/flow occurrence carries the dimension subscript -- - // including the nested-PREVIOUS terms, which are exactly the ones - // that break with a bare arrayed name. - assert!( - text.contains("PREVIOUS(PREVIOUS(growth[region]))"), - "nested-PREVIOUS flow term must be subscripted; got: {text}" +/// An outflow's polarity is structural: its numerator is the negated flow +/// delta, `PREVIOUS(flow) - flow`, and nothing else changes. +#[test] +fn flow_to_stock_outflow_negates_the_numerator() { + let stock = flow_to_stock_test_stock( + "pop", + Equation::ApplyToAll(vec!["region".to_string()], "100".to_string()), + &[], + &["deaths"], ); - assert!( - text.contains("PREVIOUS(PREVIOUS(pop[region]))"), - "nested-PREVIOUS stock term must be subscripted; got: {text}" + let flow = flow_to_stock_test_flow( + "deaths", + Equation::ApplyToAll(vec!["region".to_string()], "pop[region] * 0.05".to_string()), ); - // ...and no bare arrayed name survives as a PREVIOUS argument. + + let equation = generate_flow_to_stock_equation("deaths", "pop", &flow, &stock); + let LtmEquation::ApplyToAll(_, arm) = &equation else { + panic!("arrayed stock must yield LtmEquation::ApplyToAll; got: {equation:?}"); + }; assert!( - !text.contains("PREVIOUS(growth)") && !text.contains("PREVIOUS(growth,"), - "no bare arrayed flow reference may remain; got: {text}" + arm.text + .contains("SAFEDIV((PREVIOUS(deaths[region]) - deaths[region]), ABS(("), + "an outflow's numerator is the negated flow delta; got: {}", + arm.text ); assert!( - !text.contains("PREVIOUS(pop)") && !text.contains("PREVIOUS(pop,"), - "no bare arrayed stock reference may remain; got: {text}" + arm.text + .ends_with("* SIGN((deaths[region] - PREVIOUS(deaths[region])))"), + "the sign factor is the flow's own delta, unnegated; got: {}", + arm.text ); } -/// Guard: a *scalar* stock's flow-to-stock equation must NOT gain -/// subscripts -- it stays the bare-name `Equation::Scalar` form so the -/// scalar isolated-loop invariant (pinned by `ltm_dt_invariance.rs`) -/// is unaffected. +/// A flow declared over other dimensions than its stock's is spelled bare +/// (the compiler resolves it through the wiring's implicit subscripts), while +/// the net aux keeps the stock's subscript; so is a scalar flow into an +/// arrayed stock, which broadcasts into every element's net flow. #[test] -fn test_flow_to_stock_scalar_stays_bare() { - let stock = - flow_to_stock_test_stock("s", Equation::Scalar("100".to_string()), &["births"], &[]); - let flow = flow_to_stock_test_flow("births", Equation::Scalar("s * 0.1".to_string())); +fn flow_to_stock_flow_over_other_dims_or_scalar_is_spelled_bare() { + let stock = flow_to_stock_test_stock( + "level", + Equation::ApplyToAll(vec!["suba".to_string()], "100".to_string()), + &["inflow", "fill"], + &[], + ); + let mapped = flow_to_stock_test_flow( + "inflow", + Equation::ApplyToAll(vec!["dimb".to_string()], "1".to_string()), + ); + let scalar = flow_to_stock_test_flow("fill", Equation::Scalar("2".to_string())); + let net = net_ref("level", "[suba]"); - let equation = generate_flow_to_stock_equation("births", "s", &flow, &stock); - let text = match &equation { - LtmEquation::Scalar(arm) => &arm.text, - other => panic!("scalar stock must yield Equation::Scalar; got: {other:?}"), - }; + for (name, flow) in [("inflow", &mapped), ("fill", &scalar)] { + let equation = generate_flow_to_stock_equation(name, "level", flow, &stock); + let LtmEquation::ApplyToAll(dims, arm) = &equation else { + panic!("arrayed stock must yield LtmEquation::ApplyToAll; got: {equation:?}"); + }; + assert_eq!(dims, &vec!["suba".to_string()]); + assert_eq!( + &*arm.text, + format!( + "if (TIME = INITIAL_TIME) then 0 else if (({net} - PREVIOUS({net})) = 0) OR \ + (({name} - PREVIOUS({name})) = 0) then 0 else SAFEDIV(({name} - \ + PREVIOUS({name})), ABS(({net} - PREVIOUS({net}))), 0) * SIGN(({name} - \ + PREVIOUS({name})))" + ) + ); + } +} - assert!( - text.contains("PREVIOUS(PREVIOUS(births))"), - "scalar flow term must stay bare; got: {text}" - ); - assert!( - text.contains("PREVIOUS(PREVIOUS(s))"), - "scalar stock term must stay bare; got: {text}" +/// The net-flow aux of a scalar stock: `(inflows) - (outflows)`, each side the +/// declared flows in declaration order. +#[test] +fn net_flow_equation_sums_inflows_minus_outflows() { + let stock = flow_to_stock_test_stock( + "s", + Equation::Scalar("100".to_string()), + &["births", "immigration"], + &["deaths"], ); - assert!( - !text.contains('['), - "scalar flow-to-stock equation must have no subscripts; got: {text}" + let births = flow_to_stock_test_flow("births", Equation::Scalar("1".to_string())); + let immigration = flow_to_stock_test_flow("immigration", Equation::Scalar("1".to_string())); + let deaths = flow_to_stock_test_flow("deaths", Equation::Scalar("1".to_string())); + + let equation = generate_net_flow_equation( + &stock, + &[ + ("births", Some(&births)), + ("immigration", Some(&immigration)), + ], + &[("deaths", Some(&deaths))], ); + let LtmEquation::Scalar(arm) = &equation else { + panic!("scalar stock must yield Equation::Scalar; got: {equation:?}"); + }; + assert_eq!(&*arm.text, "(births + immigration) - (deaths)"); } -/// An arrayed *outflow* keeps the negative structural sign while still -/// being subscripted: the sign is applied outside `ABS()`, independent -/// of the subscripting. +/// A side the stock has no flows on is `0`, so a one-sided stock's net flow +/// is its flows' sum with the structural sign and the score's `Δnet` is +/// well-defined. #[test] -fn test_flow_to_stock_arrayed_outflow_sign() { +fn net_flow_equation_uses_zero_for_a_missing_side() { + let births = flow_to_stock_test_flow("births", Equation::Scalar("1".to_string())); + let deaths = flow_to_stock_test_flow("deaths", Equation::Scalar("1".to_string())); + + let inflow_only = + flow_to_stock_test_stock("s", Equation::Scalar("100".to_string()), &["births"], &[]); + let LtmEquation::Scalar(arm) = + generate_net_flow_equation(&inflow_only, &[("births", Some(&births))], &[]) + else { + panic!("scalar stock must yield Equation::Scalar"); + }; + assert_eq!(&*arm.text, "(births) - (0)"); + + let outflow_only = + flow_to_stock_test_stock("s", Equation::Scalar("100".to_string()), &[], &["deaths"]); + let LtmEquation::Scalar(arm) = + generate_net_flow_equation(&outflow_only, &[], &[("deaths", Some(&deaths))]) + else { + panic!("scalar stock must yield Equation::Scalar"); + }; + assert_eq!(&*arm.text, "(0) - (deaths)"); +} + +/// The net-flow aux is shaped like the stock, and its flows are spelled as +/// the score spells them: subscripted when their dimensions are the stock's, +/// bare for a flow over other dimensions, a scalar flow, or a flow the model +/// could not lower. +#[test] +fn net_flow_equation_is_shaped_like_the_stock() { let stock = flow_to_stock_test_stock( "pop", Equation::ApplyToAll(vec!["region".to_string()], "100".to_string()), - &[], + &["growth", "mapped", "fill", "broken"], &["deaths"], ); - let flow = flow_to_stock_test_flow( - "deaths", - Equation::ApplyToAll(vec!["region".to_string()], "pop[region] * 0.05".to_string()), + let region = |eqn: &str| Equation::ApplyToAll(vec!["region".to_string()], eqn.to_string()); + let growth = flow_to_stock_test_flow("growth", region("1")); + let deaths = flow_to_stock_test_flow("deaths", region("1")); + let mapped = flow_to_stock_test_flow( + "mapped", + Equation::ApplyToAll(vec!["dimb".to_string()], "1".to_string()), ); + let fill = flow_to_stock_test_flow("fill", Equation::Scalar("1".to_string())); - let equation = generate_flow_to_stock_equation("deaths", "pop", &flow, &stock); - let text = match &equation { - LtmEquation::ApplyToAll(_, arm) => &arm.text, - other => panic!("arrayed stock must yield Equation::ApplyToAll; got: {other:?}"), - }; - - assert!( - text.contains("-ABS(SAFEDIV("), - "outflow link score must carry the negative structural sign; got: {text}" + let equation = generate_net_flow_equation( + &stock, + &[ + ("growth", Some(&growth)), + ("mapped", Some(&mapped)), + ("fill", Some(&fill)), + ("broken", None), + ], + &[("deaths", Some(&deaths))], ); - assert!( - text.contains("PREVIOUS(PREVIOUS(deaths[region]))"), - "outflow must still be subscripted; got: {text}" + let LtmEquation::ApplyToAll(dims, arm) = &equation else { + panic!("arrayed stock must yield LtmEquation::ApplyToAll; got: {equation:?}"); + }; + assert_eq!(dims, &vec!["region".to_string()]); + assert_eq!( + &*arm.text, + "(growth[region] + mapped + fill + broken) - (deaths[region])" ); } diff --git a/src/simlin-engine/src/ltm_augment_with_lookup.rs b/src/simlin-engine/src/ltm_augment_with_lookup.rs index 33361c3ad..27145bad2 100644 --- a/src/simlin-engine/src/ltm_augment_with_lookup.rs +++ b/src/simlin-engine/src/ltm_augment_with_lookup.rs @@ -69,7 +69,8 @@ use super::quote_ident; /// | [`super::build_element_reducer_link_score`], `!is_bare` / RANK / un-pinnable-body arms | 2 | never | /// /// Not numerator producers, and so nothing to wrap: `generate_flow_to_stock_equation` -/// (a fixed structural formula whose target is a `Variable::Stock`, which +/// (the closed-form partial of the stock's net-flow aux, a linear sum of +/// flows with no graphical function; its target is a `Variable::Stock`, which /// `is_implicit_with_lookup` excludes) and the module composite / black-box /// scores (`Δoutput`-shaped transfer formulas with no target-equation partial). /// diff --git a/src/simlin-engine/tests/integration/ltm_discovery_large_models.rs b/src/simlin-engine/tests/integration/ltm_discovery_large_models.rs index 2ee0e5e86..5615d65a5 100644 --- a/src/simlin-engine/tests/integration/ltm_discovery_large_models.rs +++ b/src/simlin-engine/tests/integration/ltm_discovery_large_models.rs @@ -37,22 +37,22 @@ //! tractable on the variable graph could still be intractable on the //! element graph, and a variable-level test would not catch that. //! -//! ## The 2-step startup guard +//! ## The startup guard //! -//! LTM link scores are `PREVIOUS()`-based, so the first two saved -//! timesteps (indices 0 and 1) are startup-degenerate. Step 0's link -//! scores can be NaN -- both discovery generators skip step 0 for exactly -//! that reason -- and steps 0-1 carry no -//! positive loop contribution, so `rank_and_filter` drops every loop -//! whose only timesteps are those two. Step index 2 is the first -//! genuinely discoverable timestep. `FIRST_DISCOVERABLE_STEP` and -//! `TRUNCATED_STEP_COUNT` encode this; the discovery tests truncate -//! results to `TRUNCATED_STEP_COUNT` (3) so the window holds exactly one -//! discoverable timestep, and `assert_discovery_contract` only inspects -//! score values at step indices `>= FIRST_DISCOVERABLE_STEP`. (The same -//! "step 2 is the first real step" fact is relied on by the existing -//! discovery test in `simulate_ltm.rs`, which iterates `for step in -//! 2..`.) +//! LTM link scores are `PREVIOUS()`-based, so the first saved timestep +//! (index 0) is startup-degenerate: every score is pinned to 0 there by its +//! `TIME = INITIAL_TIME` guard, step 0's link scores can be NaN -- both +//! discovery generators skip step 0 for exactly that reason -- and it +//! carries no positive loop contribution, so `rank_and_filter` drops every +//! loop whose only timestep is that one. Step index 1 is the first +//! genuinely discoverable timestep: every score reads one step of history, +//! the flow-to-stock score included (it is `|Δflow / Δnet|` over the +//! stock's net-flow aux, with no stock history behind it). +//! `FIRST_DISCOVERABLE_STEP` and `TRUNCATED_STEP_COUNT` encode this; the +//! discovery tests truncate results to `TRUNCATED_STEP_COUNT` (2) so the +//! window holds exactly one discoverable timestep, and +//! `assert_discovery_contract` only inspects score values at step indices +//! `>= FIRST_DISCOVERABLE_STEP`. //! //! ## Discovery is tractable on World3 (was GH #540, now closed) //! @@ -133,11 +133,11 @@ const CLEARN_MDL: &str = "../../test/xmutil_test_models/C-LEARN v77 for Vensim.m /// Index of the first genuinely discoverable saved timestep. /// /// LTM link scores are `PREVIOUS()`-based: steps 0 and 1 are -/// startup-degenerate (step 0's link scores can be NaN; steps 0-1 carry -/// no positive loop contribution). Step index 2 is the first timestep -/// whose link scores -- and therefore discovered loop scores -- are -/// meaningful. See the module-level "2-step startup guard" section. -const FIRST_DISCOVERABLE_STEP: usize = 2; +/// startup-degenerate (step 0's link scores are guarded to 0 and can be +/// NaN; it carries no positive loop contribution). Step index 1 is the +/// first timestep whose link scores -- and therefore discovered loop scores +/// -- are meaningful. See the module-level "startup guard" section. +const FIRST_DISCOVERABLE_STEP: usize = 1; /// Number of saved timesteps to keep when truncating results for a /// single-discoverable-timestep discovery run: the two startup-guard @@ -242,11 +242,10 @@ fn truncate_results(results: &Results, n_steps: usize) -> Results { /// unrelated bug. /// /// Score finiteness is checked only at step indices `>= -/// FIRST_DISCOVERABLE_STEP`: steps 0-1 are startup-degenerate and step 0 -/// in particular is allowed to be NaN by the LTM algorithm (see the -/// module-level "2-step startup guard" section), so asserting finiteness -/// there would be asserting on behavior the algorithm treats as -/// undefined. +/// FIRST_DISCOVERABLE_STEP`: step 0 is startup-degenerate and is allowed +/// to be NaN by the LTM algorithm (see the module-level "startup guard" +/// section), so asserting finiteness there would be asserting on behavior +/// the algorithm treats as undefined. fn assert_discovery_contract(found: &[ltm_finding::FoundLoop]) { assert!( !found.is_empty(), diff --git a/src/simlin-engine/tests/integration/ltm_dt_invariance.rs b/src/simlin-engine/tests/integration/ltm_dt_invariance.rs index 101610358..3872e8275 100644 --- a/src/simlin-engine/tests/integration/ltm_dt_invariance.rs +++ b/src/simlin-engine/tests/integration/ltm_dt_invariance.rs @@ -3,7 +3,7 @@ // Version 2.0, that can be found in the LICENSE file. //! Regression tests for the `dt`-invariance of LTM (Loops That Matter) -//! loop scores -- "Finding 1" of the LTM deep review. +//! loop scores. //! //! ## The invariant being pinned //! @@ -13,30 +13,30 @@ //! loop acting on its stock(s) -- is exactly `+/-1` at every timestep, //! regardless of loop gain *and regardless of the integration step `dt`*. //! -//! ## The bug these tests guard against +//! ## Why it holds, and what a regression looks like //! -//! `generate_flow_to_stock_equation` in `ltm_augment.rs` builds the -//! flow-to-stock link score. The 2023 paper's Eq. 3 writes it as -//! `sign * |Delta(flow) / (Delta(S_t) - Delta(S_{t-dt}))|`. Taken -//! literally that is only correct for `dt = 1`: the denominator is the -//! second-order stock change, which under Euler integration is -//! `dt * (netflow(t-dt) - netflow(t-2dt))` and so already carries one -//! factor of `dt`, while the raw flow delta in the numerator carries -//! none. The dimensionally-correct discretization of the paper's -//! continuous form (Eq. 6, `|di/dt / d^2S/dt^2|`, which is manifestly -//! dimensionless) multiplies the numerator by `dt`. +//! A flow-to-stock link score is the partial of the stock's net-flow aux +//! `$⁚ltm⁚net⁚{stock}` with respect to the flow, `sign * |Δflow / Δnet|` +//! (`generate_flow_to_stock_equation` in `ltm_augment.rs`; the 2023 +//! paper's Eq. 3). For a stock with one flow the net flow IS that flow, so +//! the score is `|Δf / Δf| = 1` identically: no `dt` appears, because both +//! deltas are flow deltas over the same `[t - dt, t]` window. The +//! stock-to-flow link of an isolated loop is likewise `+/-1` (the stock is +//! the flow's only moving input), so the product is `+/-1` whatever the +//! step size. //! -//! Without that factor each flow-to-stock link score is `1/dt` too -//! large, and because a loop score is the product of its link scores -//! the error compounds once per flow-to-stock link: an isolated -//! one-stock loop reads `1/dt`, an isolated two-stock loop reads -//! `1/dt^2`, and so on. Normalization into relative loop scores does -//! *not* cancel it when a partition mixes loops of different stock -//! counts -- so at `dt != 1` the dominant loop of a partition can flip -//! purely as an artifact of the step size. +//! A regression is any score that reads the stock's own history instead: +//! a stock changes by `dt * net` per step, so a score written over stock +//! differences (`|Δflow / Δ(ΔS)|`) is `1/dt` too large, and because a loop +//! score is the product of its link scores the error compounds once per +//! flow-to-stock link -- an isolated one-stock loop reads `1/dt`, a +//! two-stock loop `1/dt^2`. Normalization into relative loop scores does +//! *not* cancel it when a partition mixes loops of different stock counts, +//! so at `dt != 1` the dominant loop of a partition can flip purely as an +//! artifact of the step size. //! //! The three tests below pin, respectively: the one-stock isolated-loop -//! value, the two-stock isolated-loop value (proving the error does not +//! value, the two-stock isolated-loop value (proving a regression does not //! merely rescale every loop equally), and the `dt`-invariance of the //! relative scores in a mixed-stock-count partition (the real-world //! impact). @@ -111,16 +111,15 @@ fn loop_score_series(results: &Results, name: &str) -> Vec { results.iter().map(|row| row[offset]).collect() } -/// Number of saved steps at the start of a flow-to-stock link score -/// series that are pinned to `0` by the equation's startup guard -/// (`TIME = INITIAL_TIME` and `PREVIOUS(TIME, INITIAL_TIME) = -/// INITIAL_TIME`): the second-order stock difference in the score's -/// denominator needs two steps of history before it is defined. -const STARTUP_STEPS: usize = 2; +/// Number of saved steps at the start of a link-score series that are +/// pinned to `0` by the equation's startup guard (`TIME = INITIAL_TIME`): +/// every score reads one step of history, so its first value is at the +/// step after the start. +const STARTUP_STEPS: usize = 1; /// Absolute tolerance for the isolated-loop `+/-1` check. Observed -/// floating-point error in the settled region is ~1e-13; the bug this -/// guards against shifts the score by a whole `(1/dt)^k - 1 >= 1`, so +/// floating-point error in the settled region is ~1e-13; the regression +/// this guards against shifts the score by a whole `(1/dt)^k - 1 >= 1`, so /// any tolerance well below 1 catches it. const SETTLED_TOL: f64 = 1e-6; @@ -147,22 +146,22 @@ fn assert_isolated_loop_series(series: &[f64], expected: f64, dt: f64) { (value - expected).abs() < SETTLED_TOL, "dt={dt}: settled step {step} has loop_score {value}, expected {expected}. \ An isolated loop's raw score is exactly +/-1 at every dt; a value scaled \ - by (1/dt)^k means the flow-to-stock link score lost its `dt` factor \ - (LTM review Finding 1)." + by (1/dt)^k means a flow-to-stock link score is reading the stock's own \ + history instead of the net flow's." ); } } } -/// Finding 1, one-stock case: an isolated single-stock reinforcing loop +/// One-stock case: an isolated single-stock reinforcing loop /// (`s -> births -> s`) has a raw loop score of exactly `+1` at every -/// `dt`. `dt = 1.0` is the control -- the missing `dt` factor is -/// invisible there, so this case also guards against an over-correction -/// -- while `dt = 0.5` and `dt = 0.25` are where a regressed formula -/// would read `1/dt` (`2.0` and `4.0` respectively). +/// `dt`. `dt = 1.0` is the control -- a stock-history score is invisible +/// there, so this case also guards against an over-correction -- while the +/// fractional steps are where a regressed formula would read `1/dt` +/// (`2.0`, `4.0` and `8.0` respectively). #[test] fn isolated_one_stock_loop_raw_score_is_one_at_every_dt() { - for dt in [1.0_f64, 0.5, 0.25] { + for dt in [1.0_f64, 0.5, 0.25, 0.125] { let project = TestProject::new("iso_one_stock") .with_sim_time(0.0, 8.0, dt) .stock("s", "100", &["births"], &[], None) @@ -185,17 +184,18 @@ fn isolated_one_stock_loop_raw_score_is_one_at_every_dt() { } } -/// Finding 1, two-stock case: an isolated two-stock reinforcing loop +/// Two-stock case: an isolated two-stock reinforcing loop /// (`a -> in_b -> b -> in_a -> a`) also has a raw loop score of exactly -/// `+1` at every `dt`. This is the key test that the missing-`dt` error +/// `+1` at every `dt`. This is the key test that a stock-history score /// does *not* simply rescale every loop equally: the loop has two /// flow-to-stock links, so a regressed formula compounds the error to -/// `1/dt^2` (`4.0` at `dt = 0.5`, `16.0` at `dt = 0.25`) -- a different -/// power of `dt` than the one-stock loop, which is precisely why -/// normalization cannot rescue a mixed-stock-count partition. +/// `1/dt^2` (`4.0` at `dt = 0.5`, `16.0` at `dt = 0.25`, `64.0` at +/// `dt = 0.125`) -- a different power of `dt` than the one-stock loop, +/// which is precisely why normalization cannot rescue a mixed-stock-count +/// partition. #[test] fn isolated_two_stock_loop_raw_score_is_one_at_every_dt() { - for dt in [1.0_f64, 0.5, 0.25] { + for dt in [1.0_f64, 0.5, 0.25, 0.125] { let project = TestProject::new("iso_two_stock") .with_sim_time(0.0, 8.0, dt) .stock("a", "100", &["in_a"], &[], None) @@ -272,17 +272,18 @@ fn dominant_loop(scores: &HashMap) -> String { .expect("at least one loop") } -/// Finding 1, real-world impact: in a single cycle partition that mixes -/// a one-stock and a two-stock loop, the *relative* loop scores must be -/// near-invariant to `dt` for a near-linear model. +/// Real-world impact: in a single cycle partition that mixes a one-stock +/// and a two-stock loop, the *relative* loop scores must be near-invariant +/// to `dt` for a near-linear model. /// /// Normalization divides each loop score by the partition sum, so a -/// per-loop factor that is the *same* for every loop cancels. The -/// missing-`dt` error is *not* the same for every loop: it is +/// per-loop factor that is the *same* for every loop cancels. A +/// stock-history score's error is *not* the same for every loop: it is /// `(1/dt)^(stock count)`, so it survives normalization whenever a -/// partition mixes stock counts. Pre-fix, the relative split of this -/// model swung ~0.33 between `dt = 1.0` and `dt = 0.25` and the dominant -/// loop flipped; post-fix the drift is ~0.01 and the ordering is stable. +/// partition mixes stock counts -- with such a score the relative split of +/// this model swings ~0.33 between `dt = 1.0` and `dt = 0.25` and the +/// dominant loop flips, where the net-flow score drifts ~0.01 and keeps +/// the ordering. #[test] fn relative_loop_scores_are_dt_invariant_for_mixed_stock_counts() { let at_full_dt = relative_scores_at_dt(1.0); @@ -304,10 +305,10 @@ fn relative_loop_scores_are_dt_invariant_for_mixed_stock_counts() { "the same loop ids must be present at both dt values" ); - // Per-loop drift bound. Observed post-fix drift is ~0.01; the - // pre-fix drift was ~0.33. 0.05 leaves comfortable margin above the - // floating-point / near-linearity noise floor while staying far - // below the regressed behavior. + // Per-loop drift bound. The observed drift is ~0.01 and a stock-history + // score's is ~0.33; 0.05 leaves comfortable margin above the + // floating-point / near-linearity noise floor while staying far below + // the regressed behavior. const MAX_DRIFT: f64 = 0.05; for (id, full_score) in &at_full_dt { let quarter_score = at_quarter_dt[id]; @@ -315,14 +316,14 @@ fn relative_loop_scores_are_dt_invariant_for_mixed_stock_counts() { assert!( drift < MAX_DRIFT, "loop {id}: relative score drifted {drift:.4} between dt=1.0 ({full_score:.4}) \ - and dt=0.25 ({quarter_score:.4}). A drift this large means the flow-to-stock \ - link score lost its `dt` factor, inflating each loop by (1/dt)^(stock count) \ - so the error survives partition normalization (LTM review Finding 1)." + and dt=0.25 ({quarter_score:.4}). A drift this large means a flow-to-stock \ + link score is reading the stock's own history, inflating each loop by \ + (1/dt)^(stock count) so the error survives partition normalization." ); } - // The headline symptom of Finding 1: the dominant loop of the - // partition flipping purely because the integration step changed. + // The headline symptom: the dominant loop of the partition flipping + // purely because the integration step changed. let dominant_full = dominant_loop(&at_full_dt); let dominant_quarter = dominant_loop(&at_quarter_dt); assert_eq!( diff --git a/src/simlin-engine/tests/integration/ltm_flow_to_stock.rs b/src/simlin-engine/tests/integration/ltm_flow_to_stock.rs new file mode 100644 index 000000000..788afff52 --- /dev/null +++ b/src/simlin-engine/tests/integration/ltm_flow_to_stock.rs @@ -0,0 +1,605 @@ +// Copyright 2026 The Simlin Authors. All rights reserved. +// Use of this source code is governed by the Apache License, +// Version 2.0, that can be found in the LICENSE file. + +//! The flow-to-stock link score, pinned against the 2023 correction paper's +//! own numbers and its aggregation-invariance argument. +//! +//! Schoenberg, Hayward and Eberlein (2023), "Improving Loops that Matter", +//! define the flow-to-stock link score (their Eq. 3) as the change in the +//! flow over the change in the stock's net flow, and observe (their section +//! 4.4) that an implementation may equivalently aggregate every stock's flows +//! into one net flow and score each flow's link into that net flow with the +//! ordinary instantaneous formula, the net flow's own link into the stock +//! being exactly 1. Simlin takes that second form: every stock with a scored +//! flow-to-stock edge gets a synthetic net-flow auxiliary +//! `$⁚ltm⁚net⁚{stock}` and the score is `sign * |Δflow / Δnet|`, read over +//! the same `[t - dt, t]` window as every other link score in the model. +//! +//! These tests pin the numbers that form implies -- the paper's worked table +//! at the step of the change (not one step later), exact agreement between a +//! model with separate flows and its twin with one net flow at the same +//! time, the 67/33 split of the births/deaths model, the per-element score +//! of a scalar flow feeding an arrayed stock and the loop through it, an +//! outflow-only stock, a stock whose initial value reads its own flow, and +//! the net aux's absence from every loop and link surface -- through the +//! real compile-and-simulate pipeline. +//! +//! The family's other arms are pinned elsewhere: a flow declared over +//! dimensions that MAP onto its stock's, per slot, in `ltm_array_agg.rs` +//! (`a_flow_feeding_its_stock_through_a_parent_mapping_scores_per_slot`); a +//! flow inside a module instance, whose net aux lives in the instance's +//! namespace, in `db/ltm_tests.rs` +//! (`delay3_input_to_stock_link_score_reads_the_bound_port`); the generated +//! text of every spelling (scalar, arrayed, outflow, other-dimensioned and +//! scalar flows, missing sides) in `ltm_augment_tests.rs` (`flow_to_stock_*` +//! and `net_flow_equation_*`). + +use simlin_engine::datamodel::{Equation, Variable}; +use simlin_engine::ltm_finding; +use simlin_engine::ltm_post; +use simlin_engine::test_common::TestProject; + +use crate::test_helpers::{ltm_discovery_inputs, ltm_run, ltm_series}; + +/// `$⁚ltm⁚link_score⁚{from}→{to}`. +fn link_score(from: &str, to: &str) -> String { + format!("$\u{205A}ltm\u{205A}link_score\u{205A}{from}\u{2192}{to}") +} + +/// Every step's absolute difference between two same-length series. +fn max_abs_diff(a: &[f64], b: &[f64]) -> f64 { + assert_eq!(a.len(), b.len(), "series lengths differ"); + a.iter() + .zip(b) + .map(|(x, y)| (x - y).abs()) + .fold(0.0, f64::max) +} + +/// The 2023 paper's table (section 4.3): a stock with inflow `in` stepping +/// 5 -> 10 and outflow `out` stepping 4 -> 5 at one step, dt = 1. +/// +/// Hand calculation at the step of the change (t = 2): +/// +/// ```text +/// Δin = 10 - 5 = 5 +/// Δout = 5 - 4 = 1 +/// net = in - out: 1 at t = 1, 5 at t = 2, so Δnet = 4 +/// LS(in -> S) = +|Δin / Δnet| = +5/4 = +1.25 +/// LS(out -> S) = -|Δout / Δnet| = -1/4 = -0.25 +/// ``` +/// +/// which are the paper's 1.25 and 0.25 (its magnitudes) with the structural +/// signs. Every other step has `Δin = Δout = 0`, so both scores are 0 there +/// by the zero-delta guard -- including t = 3, where a score computed over +/// the previous interval would put the paper's values one step late. +/// +/// The `0 * s` terms close a (causally inert) loop through each flow so +/// exhaustive mode scores the two flow-to-stock edges. +#[test] +fn paper_2023_table_scores_the_step_of_the_change() { + let project = TestProject::new("paper_2023_table") + .with_sim_time(0.0, 5.0, 1.0) + .stock("s", "100", &["inflow"], &["outflow"], None) + .flow("inflow", "5 + STEP(5, 2) + 0 * s", None) + .flow("outflow", "4 + STEP(1, 2) + 0 * s", None) + .build_datamodel(); + let results = ltm_run(&project, false).results; + + let inflow_score = ltm_series(&results, &link_score("inflow", "s"), 0); + let outflow_score = ltm_series(&results, &link_score("outflow", "s"), 0); + assert_eq!( + inflow_score, + vec![0.0, 0.0, 1.25, 0.0, 0.0, 0.0], + "LS(inflow -> s) at t = 0..5" + ); + assert_eq!( + outflow_score, + vec![0.0, 0.0, -0.25, 0.0, 0.0, 0.0], + "LS(outflow -> s) at t = 0..5" + ); +} + +/// The 2023 paper's aggregation invariance (its section 4.4), at the same +/// time step: a stock with separate nonlinear inflow and outflow, and its +/// twin where the same two expressions are auxiliaries feeding one net flow, +/// give the same flow-to-stock link score and the same relative loop scores +/// at every t. +/// +/// In the aggregated twin the net flow is the stock's only flow, so its +/// flow-to-stock score is `|Δnet / Δnet| = 1` wherever it is non-zero and +/// the chain `LS(out -> net) * LS(net -> s)` reduces to the ordinary +/// instantaneous score of `net = in - out` with respect to `out`, which is +/// `-|Δout / Δnet|` -- the disaggregated twin's `LS(out -> s)` by +/// construction. The remaining differences are floating-point rounding in +/// the two spellings of `Δnet`, far below the tolerance. +#[test] +fn flow_to_stock_scores_are_aggregation_invariant_at_the_same_step() { + let dt = 0.5; + let disaggregated = TestProject::new("agg_dis") + .with_sim_time(0.0, 20.0, dt) + .stock("s", "10", &["inflow"], &["outflow"], None) + .flow("inflow", "0.3 * s ^ 0.8", None) + .flow("outflow", "0.02 * s ^ 1.3", None) + .build_datamodel(); + let aggregated = TestProject::new("agg_agg") + .with_sim_time(0.0, 20.0, dt) + .stock("s", "10", &["net"], &[], None) + .aux("inflow", "0.3 * s ^ 0.8", None) + .aux("outflow", "0.02 * s ^ 1.3", None) + .flow("net", "inflow - outflow", None) + .build_datamodel(); + let dis = ltm_run(&disaggregated, false); + let agg = ltm_run(&aggregated, false); + + let net_to_s = ltm_series(&agg.results, &link_score("net", "s"), 0); + for flow in ["inflow", "outflow"] { + let direct = ltm_series(&dis.results, &link_score(flow, "s"), 0); + let through_net: Vec = ltm_series(&agg.results, &link_score(flow, "net"), 0) + .iter() + .zip(&net_to_s) + .map(|(a, b)| a * b) + .collect(); + // The stock moves at every step of this run, so the invariance is + // exercised on non-zero scores, not on a vacuous pair of zero series. + assert!( + direct.iter().skip(1).all(|v| v.abs() > 1e-3), + "LS({flow} -> s) is active at every step after the first: {direct:?}" + ); + let diff = max_abs_diff(&direct, &through_net); + assert!( + diff < 1e-12, + "LS({flow} -> s) differs from LS({flow} -> net) * LS(net -> s) by {diff:e}: \ + {direct:?} vs {through_net:?}" + ); + } + + let dis_rel = ltm_post::compute_rel_loop_scores(&dis.results, &dis.loop_partitions); + let agg_rel = ltm_post::compute_rel_loop_scores(&agg.results, &agg.loop_partitions); + // Each twin has exactly one loop through the inflow and one through the + // outflow (the static polarity of `s ^ p` is unknown, so the ids are + // `u{n}` and their numbering is not shared across the twins). + for flow in ["inflow", "outflow"] { + let dis_id = dis.loop_through(flow); + let agg_id = agg.loop_through(flow); + let diff = max_abs_diff(&dis_rel[dis_id], &agg_rel[agg_id]); + assert!( + diff < 1e-12, + "relative score of the loop through {flow} ({dis_id} / {agg_id}) differs \ + between the twins by {diff:e}" + ); + } +} + +/// Seamlessly Integrating (Schoenberg et al. 2020), figure 5: a stock with +/// births `0.1 * a` and deaths `a / 20` splits 67/33 between its reinforcing +/// and balancing loops. +/// +/// Hand calculation, valid at every step after the first (`Δa > 0`): +/// +/// ```text +/// LS(a -> births) = +1 (births = 0.1 a exactly, so Δ_a births = Δbirths) +/// LS(births -> a) = +|Δbirths / Δnet| = |0.1 Δa / 0.05 Δa| = +2 +/// LS(a -> deaths) = +1 +/// LS(deaths -> a) = -|Δdeaths / Δnet| = -|0.05 Δa / 0.05 Δa| = -1 +/// loop r1 = +2, loop b1 = -1, so the relative scores are +2/3 and -1/3. +/// ``` +/// +/// The split holds from the first step after the start: every score's +/// window is `[t - dt, t]`, so no score needs a second step of history. +#[test] +fn births_and_deaths_split_two_thirds_one_third_from_the_first_step() { + let project = TestProject::new("births_deaths") + .with_sim_time(0.0, 10.0, 1.0) + .stock("a", "100", &["births"], &["deaths"], None) + .flow("births", "0.1 * a", None) + .flow("deaths", "a / 20", None) + .build_datamodel(); + let run = ltm_run(&project, false); + let results = &run.results; + let rel = ltm_post::compute_rel_loop_scores(results, &run.loop_partitions); + + let expected = |value: f64| -> Vec { + let mut v = vec![value; results.step_count]; + v[0] = 0.0; + v + }; + for (id, value) in [("r1", 2.0 / 3.0), ("b1", -1.0 / 3.0)] { + let diff = max_abs_diff(&rel[id], &expected(value)); + assert!( + diff < 1e-9, + "relative score of {id} should be {value} from t = 1 on, got {:?}", + rel[id] + ); + } +} + +/// A scalar flow feeding an arrayed stock: the score is one arrayed +/// variable over the stock's dimensions (the scalar flow broadcast into +/// every element's net flow), never a partial of the stock's initial-value +/// equation. +/// +/// `tank[D]` has the scalar inflow `fill = 5 + STEP(5, 2)` and the arrayed +/// outflow `drain[D] = tank[D] * rate[D]` with rates 0.1 and 0.2. Discovery +/// mode scores every causal edge, so no loop is needed. Hand calculation at +/// t = 2 (dt = 1): +/// +/// ```text +/// element a (rate 0.1): tank 100, 95, 90.5 drain 10, 9.5, 9.05 +/// net = fill - drain: -4.5 at t = 1, 0.95 at t = 2, Δnet = 5.45 +/// LS(fill -> tank[a]) = +|5 / 5.45| = 0.91743... +/// LS(drain -> tank[a]) = -|-0.45 / 5.45| = -0.08257... +/// element b (rate 0.2): tank 100, 85, 73 drain 20, 17, 14.6 +/// net = fill - drain: -12 at t = 1, -4.6 at t = 2, Δnet = 7.4 +/// LS(fill -> tank[b]) = +|5 / 7.4| = 0.67568... +/// LS(drain -> tank[b]) = -|-2.4 / 7.4| = -0.32432... +/// ``` +/// +/// and at every step each element's `fill` score is `|Δfill / Δnet[e]|` +/// computed from the run's own `fill` and `drain[e]` series. +#[test] +fn a_scalar_flow_into_an_arrayed_stock_is_scored_per_element() { + let project = TestProject::new("scalar_flow_arrayed_stock") + .with_sim_time(0.0, 4.0, 1.0) + .named_dimension("D", &["a", "b"]) + .array_stock("tank[D]", "100", &["fill"], &["drain"], None) + .flow("fill", "5 + STEP(5, 2)", None) + .array_flow("drain[D]", "tank[D] * rate[D]", None) + .array_with_ranges("rate[D]", vec![("a", "0.1"), ("b", "0.2")]) + .build_datamodel(); + let results = ltm_run(&project, true).results; + + let fill_name = link_score("fill", "tank"); + assert!( + results + .offsets + .keys() + .all(|k| !k.as_str().contains("fill\u{2192}tank[")), + "the scalar flow's score is one arrayed variable, not per-element scalars" + ); + let fill = ltm_series(&results, "fill", 0); + for (slot, (elem, fill_at_2, drain_at_2)) in [ + ("a", 5.0 / 5.45, -0.45 / 5.45), + ("b", 5.0 / 7.4, -2.4 / 7.4), + ] + .into_iter() + .enumerate() + { + let fill_score = ltm_series(&results, &fill_name, slot); + let drain_score = ltm_series(&results, &link_score("drain", "tank"), slot); + assert!( + (fill_score[2] - fill_at_2).abs() < 1e-9, + "LS(fill -> tank[{elem}]) at t = 2: got {}, expected {fill_at_2}", + fill_score[2] + ); + assert!( + (drain_score[2] - drain_at_2).abs() < 1e-9, + "LS(drain -> tank[{elem}]) at t = 2: got {}, expected {drain_at_2}", + drain_score[2] + ); + // Model variables are keyed per element; LTM variables once, bare. + let drain = ltm_series(&results, &format!("drain[{elem}]"), 0); + for t in 1..results.step_count { + let d_fill = fill[t] - fill[t - 1]; + let d_net = (fill[t] - drain[t]) - (fill[t - 1] - drain[t - 1]); + let expected = if d_fill == 0.0 || d_net == 0.0 { + 0.0 + } else { + (d_fill / d_net).abs() + }; + assert!( + (fill_score[t] - expected).abs() < 1e-9, + "LS(fill -> tank[{elem}]) at t = {t}: got {}, expected {expected}", + fill_score[t] + ); + } + } +} + +/// A scalar inflow that reads one element of the arrayed stock it feeds: +/// `tank[D]` (a: 50, b: 100), `fill = 2 + 0.05 * tank[a]`, +/// `drain[D] = tank[D] * rate[D]` with rates 0.1 and 0.2, dt 1. +fn scalar_flow_loop_project() -> simlin_engine::datamodel::Project { + TestProject::new("scalar_flow_loop") + .with_sim_time(0.0, 6.0, 1.0) + .named_dimension("D", &["a", "b"]) + .array_with_ranges("init[D]", vec![("a", "50"), ("b", "100")]) + .array_stock("tank[D]", "init[D]", &["fill"], &["drain"], None) + .flow("fill", "2 + 0.05 * tank[a]", None) + .array_flow("drain[D]", "tank[D] * rate[D]", None) + .array_with_ranges("rate[D]", vec![("a", "0.1"), ("b", "0.2")]) + .build_datamodel() +} + +/// The loop through a scalar flow into one element of an arrayed stock, in +/// exhaustive mode, on [`scalar_flow_loop_project`]. +/// +/// Hand calculation, valid at every step after the first (`Δtank[a] != 0`): +/// +/// ```text +/// net[a] = fill - drain[a] = 2 + 0.05 tank[a] - 0.1 tank[a] = 2 - 0.05 tank[a] +/// Δfill = 0.05 Δtank[a], Δnet[a] = -0.05 Δtank[a] +/// LS(fill -> tank[a]) = +|Δfill / Δnet[a]| = 1 +/// LS(drain -> tank[a]) = -|Δdrain[a] / Δnet[a]| = -|0.1 / -0.05| = -2 +/// LS(tank[a] -> fill) = +1 (fill is linear in its one moving input) +/// LS(tank[a] -> drain[a]) = +1 +/// loop tank[a] -> fill -> tank[a] = +1 * +1 = +1 +/// loop tank[a] -> drain -> tank[a] = +1 * -2 = -2 +/// ``` +/// +/// and for the b slot, which no loop through `fill` reaches, +/// `LS(fill -> tank[b]) = |Δfill / Δnet[b]|` from the run's own series with +/// `net[b] = fill - drain[b]`. +#[test] +fn a_loop_through_a_scalar_flow_into_an_arrayed_stock_scores_its_slot() { + let run = ltm_run(&scalar_flow_loop_project(), false); + let results = &run.results; + let fill_score = |slot: usize| ltm_series(results, &link_score("fill", "tank"), slot); + let drain_score = |slot: usize| ltm_series(results, &link_score("drain", "tank"), slot); + let from_step_one = |series: &[f64], expected: f64, what: &str| { + assert_eq!(series[0], 0.0, "{what} is guarded to 0 at the start"); + for (t, v) in series.iter().enumerate().skip(1) { + assert!( + (v - expected).abs() < 1e-12, + "{what} at t = {t}: got {v}, expected {expected}" + ); + } + }; + from_step_one(&fill_score(0), 1.0, "LS(fill -> tank[a])"); + from_step_one(&drain_score(0), -2.0, "LS(drain -> tank[a])"); + let fill = ltm_series(results, "fill", 0); + let drain_b = ltm_series(results, "drain[b]", 0); + let fill_b = fill_score(1); + for t in 1..results.step_count { + let d_fill = fill[t] - fill[t - 1]; + let d_net = (fill[t] - drain_b[t]) - (fill[t - 1] - drain_b[t - 1]); + let expected = if d_fill == 0.0 || d_net == 0.0 { + 0.0 + } else { + (d_fill / d_net).abs() + }; + assert!( + (fill_b[t] - expected).abs() < 1e-12, + "LS(fill -> tank[b]) at t = {t}: got {}, expected {expected}", + fill_b[t] + ); + } + + // The loop through `fill` is one loop, visiting `tank` at `a`; its raw + // score is the product +1 * +1 at every step after the first. + let through_fill = run.loop_through("fill"); + let fill_loop = ltm_series( + results, + &format!("$\u{205A}ltm\u{205A}loop_score\u{205A}{through_fill}"), + 0, + ); + from_step_one(&fill_loop, 1.0, "the loop through fill"); + // The drain loop is one loop over D; its `a` slot is +1 * -2. + let through_drain: Vec<&simlin_engine::db::DetectedLoop> = run + .loops + .iter() + .filter(|l| l.variables.iter().any(|v| v == "drain")) + .collect(); + assert_eq!( + through_drain.len(), + 1, + "one loop through drain (arrayed over D); got {:?}", + run.loops + .iter() + .map(|l| (l.id.clone(), l.variables.clone())) + .collect::>() + ); + let drain_loop_a = ltm_series( + results, + &format!( + "$\u{205A}ltm\u{205A}loop_score\u{205A}{}", + through_drain[0].id + ), + 0, + ); + from_step_one(&drain_loop_a, -2.0, "the loop through drain at a"); +} + +/// A stock with only an outflow: the net aux is `(0) - (decay)`, so the +/// flow-to-stock score is `-|Δdecay / -Δdecay| = -1` exactly, the aux is the +/// negated flow exactly, and the balancing loop `s -> decay -> s` is +/// `+1 * -1 = -1` exactly (`decay = s * 0.1` is linear in its one input). +#[test] +fn an_outflow_only_stock_scores_minus_one() { + let project = TestProject::new("outflow_only") + .with_sim_time(0.0, 5.0, 1.0) + .stock("s", "100", &[], &["decay"], None) + .flow("decay", "s * 0.1", None) + .build_datamodel(); + let run = ltm_run(&project, false); + let results = &run.results; + + let decay_score = ltm_series(results, &link_score("decay", "s"), 0); + assert_eq!(decay_score, vec![0.0, -1.0, -1.0, -1.0, -1.0, -1.0]); + let net = ltm_series(results, "$\u{205A}ltm\u{205A}net\u{205A}s", 0); + let decay = ltm_series(results, "decay", 0); + assert_eq!( + net, + decay.iter().map(|d| -d).collect::>(), + "the net-flow aux of an outflow-only stock is the negated outflow" + ); + let loop_id = run.loop_through("decay"); + let loop_score = ltm_series( + results, + &format!("$\u{205A}ltm\u{205A}loop_score\u{205A}{loop_id}"), + 0, + ); + assert_eq!(loop_score, vec![0.0, -1.0, -1.0, -1.0, -1.0, -1.0]); +} + +/// Replace stock `name`'s initial-value equation in `project` with a +/// per-element one (`Equation::Arrayed` over `dims`). +fn set_stock_initial_per_element( + project: &mut simlin_engine::datamodel::Project, + name: &str, + dims: &[&str], + elements: &[(&str, &str)], +) { + let stock = project.models[0] + .variables + .iter_mut() + .find_map(|v| match v { + Variable::Stock(s) if s.ident == name => Some(s), + _ => None, + }) + .unwrap_or_else(|| panic!("no stock named {name}")); + stock.equation = Equation::Arrayed( + dims.iter().map(|d| d.to_string()).collect(), + elements + .iter() + .map(|(e, eqn)| (e.to_string(), eqn.to_string(), None, None)) + .collect(), + None, + false, + ); +} + +/// A stock's initial value is the one equation it has, so the reference-site +/// IR classifies a flow the initial reads as that stock's site for the +/// `flow -> stock` edge -- but the edge is the wiring, not that read, and its +/// score has exactly one shape. A per-element initial reading its own flow +/// (`s[a] = f[a] * 5`, `s[b] = f[b] * 7`) must yield ONE arrayed `f -> s` +/// score, not one identical score per element; a pinned-element read in a +/// two-dimensional initial (`s[D,E] = f[a, E] * 5`) must yield one arrayed +/// score too, with no warning, rather than being diverted to the +/// per-element arm. In both, `f = 3 + STEP(2, 2)` is the stock's only flow, +/// so every slot's score is `|Δf / Δf| = 1` at t = 2 and 0 elsewhere. +/// Discovery mode scores the edge without a loop. +#[test] +fn a_per_element_initial_value_reading_the_flow_does_not_split_the_score() { + let one_dim = { + let mut project = TestProject::new("per_element_init_1d") + .with_sim_time(0.0, 4.0, 1.0) + .named_dimension("D", &["a", "b"]) + .array_stock("s[D]", "0", &["f"], &[], None) + .array_flow("f[D]", "3 + STEP(2, 2)", None) + .build_datamodel(); + set_stock_initial_per_element( + &mut project, + "s", + &["D"], + &[("a", "f[a] * 5"), ("b", "f[b] * 7")], + ); + (project, 2usize) + }; + let two_dims = { + let mut project = TestProject::new("per_element_init_2d") + .with_sim_time(0.0, 4.0, 1.0) + .named_dimension("D", &["a", "b"]) + .named_dimension("E", &["x", "y"]) + .array_stock("s[D,E]", "0", &["f"], &[], None) + .array_flow("f[D,E]", "3 + STEP(2, 2)", None) + .build_datamodel(); + // Every element's initial reads the flow's `a` row: a pinned element + // on D, iterated on E. + let Some(Variable::Stock(stock)) = project.models[0] + .variables + .iter_mut() + .find(|v| matches!(v, Variable::Stock(s) if s.ident == "s")) + else { + panic!("no stock named s"); + }; + stock.equation = Equation::ApplyToAll( + vec!["D".to_string(), "E".to_string()], + "f[a, E] * 5".to_string(), + ); + (project, 4usize) + }; + for (project, slots) in [one_dim, two_dims] { + let run = ltm_run(&project, true); + let name = link_score("f", "s"); + let mut scores: Vec<&str> = run + .results + .offsets + .keys() + .map(|k| k.as_str()) + .filter(|k| k.starts_with("$\u{205A}ltm\u{205A}link_score\u{205A}")) + .collect(); + scores.sort_unstable(); + assert_eq!( + scores, + vec![name.as_str()], + "{}: one arrayed f -> s score and no per-element score", + project.name + ); + assert!( + run.diagnostics.is_empty(), + "{}: the structural edge raises no LTM warning; got {:?}", + project.name, + run.diagnostics + ); + for slot in 0..slots { + assert_eq!( + ltm_series(&run.results, &name, slot), + vec![0.0, 0.0, 1.0, 0.0, 0.0], + "{}: slot {slot} of f -> s", + project.name + ); + } + } +} + +/// The net-flow aux is LTM machinery, not a causal node: it is a results +/// series, but it appears in no detected loop's node sequence, no causal +/// edge (the graph the FFI's link listing is built over), and no discovered +/// loop's links. +#[test] +fn the_net_flow_aux_is_no_loop_node_and_no_link() { + let project = scalar_flow_loop_project(); + let net_marker = "\u{205A}net\u{205A}"; + let run = ltm_run(&project, false); + assert!( + run.results + .offsets + .keys() + .any(|k| k.as_str() == "$\u{205A}ltm\u{205A}net\u{205A}tank"), + "the net aux is a results series (so this test is not vacuous)" + ); + for l in &run.loops { + assert!( + l.variables.iter().all(|v| !v.contains(net_marker)), + "loop {} names the net aux: {:?}", + l.id, + l.variables + ); + } + assert!( + run.edges + .iter() + .all(|(from, to)| !from.contains(net_marker) && !to.contains(net_marker)), + "a causal edge names the net aux: {:?}", + run.edges + ); + + let inputs = ltm_discovery_inputs(&project, "main"); + let found = ltm_finding::discover_loops_with_graph( + &inputs.vm_results, + &inputs.causal_graph, + &inputs.stocks, + &inputs.ltm_vars, + &inputs.dims, + &inputs.expansion, + &inputs.sub_model_output_ports, + None, + ) + .expect("discovery runs") + .loops; + assert!(!found.is_empty(), "discovery finds the loops"); + for fl in &found { + assert!( + fl.loop_info + .links + .iter() + .all(|l| !l.from.as_str().contains(net_marker) + && !l.to.as_str().contains(net_marker)), + "discovered loop {} names the net aux: {}", + fl.loop_info.id, + fl.loop_info.format_path() + ); + } +} diff --git a/src/simlin-engine/tests/integration/main.rs b/src/simlin-engine/tests/integration/main.rs index 725447ed1..c27a66fe0 100644 --- a/src/simlin-engine/tests/integration/main.rs +++ b/src/simlin-engine/tests/integration/main.rs @@ -33,6 +33,7 @@ mod layout; mod ltm_array_agg; mod ltm_discovery_large_models; mod ltm_dt_invariance; +mod ltm_flow_to_stock; // Compares xmutil-based MDL parsing against the native Rust parser, so it // needs the optional xmutil C++ converter compiled in. #[cfg(feature = "xmutil")] diff --git a/src/simlin-engine/tests/integration/simulate.rs b/src/simlin-engine/tests/integration/simulate.rs index 205fcfa22..04358ddac 100644 --- a/src/simlin-engine/tests/integration/simulate.rs +++ b/src/simlin-engine/tests/integration/simulate.rs @@ -6789,8 +6789,8 @@ fn corpus_clearn_macros_import() { /// /// Layout impact (the resource this gate protects -- #654's VM limit of 65,536 /// u16 result slots, NOT `wasmgen::lower`'s unrelated `MAX_UNROLL_UNITS`): the -/// per-step result-row width is **30,123 slots**, 46% of the ceiling, with -/// 35,413 free. Both numbers come from +/// per-step result-row width is **28,725 slots**, 44% of the ceiling, with +/// 36,811 free. Both numbers come from /// `examples/ltm_slot_width.rs`, so re-deriving them is a command rather than a /// reconstruction -- and they are the CURRENT totals: the transition records /// below quote earlier values as the left-hand side of a move, which is what @@ -6898,6 +6898,29 @@ fn corpus_clearn_macros_import() { /// arrayed helpers of three scores, so the width is 29,398 -> 29,447 slots /// and the margin 36,089 free against the 65,536-slot ceiling. /// +/// The flow-to-stock score's net-flow form moved the count UP, 6,193 -> +/// 6,224 (+31), and the width DOWN, 29,447 -> 28,725 (-722). The count is +/// arithmetic over `examples/ltm_var_dump.rs`: +35 net-flow auxes +/// (`$⁚ltm⁚net⁚{stock}`, one per stock with a scored flow-to-stock edge -- +/// 24 in `main`, 11 across the stdlib templates and `sample_until`), +2 +/// arrayed scores for the two scalar flows into arrayed stocks +/// (`global_anthropogenic_ch4_emissions -> ch4_in_atm` and +/// `global_total_c_emissions -> c_in_atmosphere`, over the 3-element +/// sensitivity dimension), and -6 for the per-element scalars those two +/// edges carried before, which were partials of the stocks' INITIAL-VALUE +/// equations rather than of anything the flow moves. The width is read off +/// the result-column diff of a C-LEARN `simlin simulate --ltm` run on the +/// previous and the new CLI: -888 nested-lag capture helper columns +/// (`$⁚$⁚ltm⁚link_score⁚{flow}→{stock}⁚{n}⁚arg0[..]`, one slot each; the +/// `PREVIOUS(PREVIOUS(..))` reads of the retired stock-history numerator), +/// -6 per-element scalars, +6 for the two arrayed scores, and the remaining +/// +166 slots are the 123 net-aux instances (a stdlib template's aux is +/// instantiated once per call site; the 166 is the remainder of this +/// arithmetic, not a separate measurement). Every added column is a net aux +/// or one of those two scores and every removed one is a nested-lag helper +/// or one of those six scalars; the margin is 36,811 free against the +/// 65,536-slot ceiling. +/// /// The pin below catches emission changes in EITHER direction, and re-deriving /// it means re-measuring BOTH numbers, not just the count. #[test] @@ -6922,7 +6945,7 @@ fn clearn_ltm_var_count_guardrail() { }) .sum(); assert_eq!( - total, 6193, + total, 6224, "C-LEARN's emitted LTM var count moved; if this is an intentional \ emission change, re-derive the layout-slot impact (the #654 \ ceiling) and update this pin with the new numbers" diff --git a/src/simlin-engine/tests/integration/simulate_ltm.rs b/src/simlin-engine/tests/integration/simulate_ltm.rs index bb23a5c04..4aea4b9b8 100644 --- a/src/simlin-engine/tests/integration/simulate_ltm.rs +++ b/src/simlin-engine/tests/integration/simulate_ltm.rs @@ -86,9 +86,18 @@ fn load_ltm_results(file_path: &str) -> StdResult> { let header = rdr.headers()?; - // The reference data appears to be shifted by 1 DT to the left compared to our output. - // Values at reference t=N match our calculations at t=N+1. - // We shift the reference timestamps forward by 1 when loading. + // A time-labeling convention, applied to EVERY link score: the reference + // tool labels the score computed over `[t, t + dt]` with `t`, while + // Simlin labels the same computation, over `[t - dt, t]`, with `t` (its + // reading of `PREVIOUS`). So a reference value at `t = N` is Simlin's at + // `t = N + 1`, and the reference timestamps are shifted forward by one + // dt when loading. This is not a flow-to-stock timing effect: the + // fixture has one flow per stock, so its flow-to-stock scores are + // identically 1 in any convention, and the whole shift sits in the + // instantaneous links -- with the +dt relabel the relative loop scores + // agree to the file's rounding, and inverting |b1|/|r1| = + // (pop/1000)/(1 - pop/1000) on each reference column reproduces pop at + // the column's own t. let dt = 1.0; // DT from the logistic growth model let mut times: Vec = Vec::new(); @@ -174,8 +183,10 @@ fn ensure_ltm_results( let actual_value = *actual_value; actual_series.push((time, actual_value)); - // Skip t=1 comparison - at initialization we don't have enough history - // for meaningful link scores (need PREVIOUS values) + // The reference's first column (its t = 0, relabelled + // t = 1) is the reference tool's own startup zero; Simlin's + // t = 1 score is its first computed value. Skip that one + // column. if (time - 1.0).abs() < 1e-9 { break; } @@ -2244,16 +2255,12 @@ fn test_a2a_flow_to_stock_link_score() { /// slot is exactly `+1` at `dt = 1` regardless of that element's gain /// (Schoenberg, Davidsen & Eberlein 2020, sec. 4.1 / Appendix B). /// -/// The bug: `generate_flow_to_stock_equation` emitted the flow-to-stock -/// link score with *bare* arrayed names. Inside the resulting -/// `Equation::ApplyToAll` the `PREVIOUS(PREVIOUS(...))` terms route their -/// inner `PREVIOUS(name)` through a synthesized *scalar* helper aux (see -/// `builtins_visitor`), which cannot hold an arrayed value -- so the -/// helper fragment failed to compile and the LTM compiler silently -/// stubbed it to 0. With the nested-PREVIOUS terms zeroed the score -/// collapsed to `1/9` (`= 0.111...`) instead of `1`. `dt = 1` is chosen -/// so the Finding-1 `dt` factor is a no-op and any deviation from `+1` is -/// purely Finding 2. +/// The regression this guards against is an arrayed flow-to-stock score +/// whose fragment fails to compile per element -- a bare arrayed name in a +/// position the apply-to-all expansion cannot hold, say -- and is silently +/// stubbed to 0, collapsing the loop score to a wrong constant (`1/9` was +/// the observed value) instead of `1`. `dt = 1` is chosen so any deviation +/// from `+1` is a per-element compile failure, not a step-size effect. #[test] fn arrayed_isolated_loop_raw_score_is_one_per_element() { let project = TestProject::new("arrayed_isolated_loop") @@ -2289,10 +2296,9 @@ fn arrayed_isolated_loop_raw_score_is_one_per_element() { // Two regions -> two per-element loop-score slots (slot 0 = north, // slot 1 = south). Each is an isolated reinforcing one-stock loop, so - // each slot is exactly +1 after the two-step startup guard (the - // flow-to-stock score's second-order denominator needs two steps of - // history before it is defined). - const STARTUP_STEPS: usize = 2; + // each slot is exactly +1 after the one-step startup guard (every score + // reads one step of history). + const STARTUP_STEPS: usize = 1; for (elem, region) in ["north", "south"].iter().enumerate() { let slot = base_offset + elem; for step in 0..results.step_count { @@ -2307,8 +2313,8 @@ fn arrayed_isolated_loop_raw_score_is_one_per_element() { assert!( (value - 1.0).abs() < 1e-6, "region {region}: arrayed isolated loop score at step {step} is {value}, \ - expected exactly +1. A value near 1/9 means the flow-to-stock link \ - score's nested PREVIOUS terms were stubbed to 0 (LTM review Finding 2)." + expected exactly +1. A constant other than 1 means part of a \ + per-element score fragment was silently stubbed to 0." ); } } @@ -8071,7 +8077,7 @@ fn test_no_duplicate_ltm_vars_with_agg_routed_and_direct_edge() { /// Build a 2-region `share[r] = pop[r] / SUM(pop[*])` model with heterogeneous /// stock initial values (`pop[big] >> pop[small]`), `pop` fed back by /// `update[r] = share[r] * pop[r] * c` -- the `* pop[r]` makes growth curved -/// (a near-constant feedback flow has ~zero second-order differences, so the +/// (a near-constant feedback flow has ~zero step-to-step change, so the /// flow→stock link score -- and thus every loop score -- would vanish; the /// curvature keeps discovery's loop scores non-degenerate). The /// reducer `SUM(pop[*])` is a subexpression, so Phase 5 hoists it into diff --git a/src/simlin-engine/tests/integration/simulate_ltm_pinned.rs b/src/simlin-engine/tests/integration/simulate_ltm_pinned.rs index d04fd73ac..3488e986a 100644 --- a/src/simlin-engine/tests/integration/simulate_ltm_pinned.rs +++ b/src/simlin-engine/tests/integration/simulate_ltm_pinned.rs @@ -1055,7 +1055,7 @@ fn pinned_scalar_feeder_agg_loop_scored_in_discovery_mode() { score; got: {eq}" ); - // And the simulated pin score is sustained non-zero past the two-step + // And the simulated pin score is sustained non-zero past the one-step // startup guard -- the agg halves compile, so the score is no longer // silently 0. let mut vm = Vm::new(compiled).unwrap(); @@ -1065,7 +1065,7 @@ fn pinned_scalar_feeder_agg_loop_scored_in_discovery_mode() { .offsets .get(pin_score.name.as_str()) .expect("pin1 loop_score must have a results slot"); - const STARTUP_STEPS: usize = 2; + const STARTUP_STEPS: usize = 1; let series: Vec = results.iter().map(|row| row[off]).collect(); assert!( series.len() > STARTUP_STEPS, diff --git a/src/simlin-engine/tests/integration/test_helpers.rs b/src/simlin-engine/tests/integration/test_helpers.rs index 4b533ab7c..edf771167 100644 --- a/src/simlin-engine/tests/integration/test_helpers.rs +++ b/src/simlin-engine/tests/integration/test_helpers.rs @@ -876,3 +876,99 @@ pub fn ltm_discovery_inputs( sub_model_output_ports, } } + +/// One LTM-instrumented run of `project`'s `main` model: the results, the +/// loop-id -> per-slot cycle-partition map that +/// `ltm_post::compute_rel_loop_scores` consumes, and the structurally +/// detected loops (id, node sequence, static polarity) the ids refer to. +#[allow(dead_code)] +pub struct LtmRun { + pub results: Results, + pub loop_partitions: simlin_engine::indexmap::IndexMap>>, + pub loops: Vec, + /// The variable-level causal edges (`model_causal_edges`), as + /// `(from, to)` pairs -- the graph the FFI's link listing is built over. + pub edges: Vec<(String, String)>, + /// The LTM derivation's own warnings, rendered. + pub diagnostics: Vec, +} + +impl LtmRun { + /// The id of the one detected loop whose node sequence names `variable`. + #[allow(dead_code)] + pub fn loop_through(&self, variable: &str) -> &str { + let mut hits = self + .loops + .iter() + .filter(|l| l.variables.iter().any(|v| v == variable)); + let found = hits + .next() + .unwrap_or_else(|| panic!("no detected loop passes through {variable:?}")); + assert!( + hits.next().is_none(), + "more than one detected loop passes through {variable:?}" + ); + &found.id + } +} + +/// Compile `project`'s `main` model with the LTM overlay on -- exhaustive +/// mode, or discovery (every causal edge scored) when `discovery` is set -- +/// and run it to completion in the bytecode VM. +/// +/// Imperative Shell: drives the salsa compile pipeline and the VM. +#[allow(dead_code)] +pub fn ltm_run(project: &datamodel::Project, discovery: bool) -> LtmRun { + use simlin_engine::db::{ + SimlinDb, compile_project_incremental, model_causal_edges, model_detected_loops, + model_ltm_variables, set_project_ltm_discovery_mode, sync_from_datamodel_incremental, + }; + + let mut db = SimlinDb::default(); + let sync = sync_from_datamodel_incremental(&mut db, project, None); + if discovery { + set_project_ltm_discovery_mode(&mut db, sync.project, true); + } + let compiled = + compile_project_incremental(&db, sync.project, "main", simlin_engine::db::LtmOverlay::On) + .expect("LTM-enabled compilation should succeed"); + let source_model = sync.models["main"].source_model; + let ltm = model_ltm_variables(&db, source_model, sync.project); + let loop_partitions = ltm.loop_partitions.clone(); + let diagnostics = ltm.diagnostics.iter().map(|d| format!("{d:?}")).collect(); + let loops = model_detected_loops(&db, source_model, sync.project).loops; + let mut edges: Vec<(String, String)> = model_causal_edges(&db, source_model, sync.project) + .edges + .iter() + .flat_map(|(from, tos)| tos.iter().map(move |to| (from.clone(), to.clone()))) + .collect(); + edges.sort(); + let mut vm = Vm::new(compiled).expect("Vm::new should succeed"); + vm.run_to_end().expect("Vm::run_to_end should succeed"); + LtmRun { + results: vm.into_results(), + loop_partitions, + loops, + edges, + diagnostics, + } +} + +/// The saved series of results key `name` at element slot `slot` (`0` for a +/// scalar key; an arrayed key is registered once, by its bare name, at its +/// base slot), bounded to the run's real step count. Panics naming the +/// available keys when `name` is not a results key. +#[allow(dead_code)] +pub fn ltm_series(results: &Results, name: &str, slot: usize) -> Vec { + let ident = results + .offsets + .keys() + .find(|k| k.as_str() == name) + .unwrap_or_else(|| { + let mut keys: Vec<&str> = results.offsets.keys().map(|k| k.as_str()).collect(); + keys.sort_unstable(); + panic!("{name:?} is not a results key; have: {keys:?}") + }); + let offset = results.offsets[ident] + slot; + results.iter().map(|row| row[offset]).collect() +} From 6a90255e6b19c6b38c958db6a92a053104dabe78 Mon Sep 17 00:00:00 2001 From: Bobby Powers Date: Tue, 8 Sep 2026 21:37:21 -0700 Subject: [PATCH 03/10] engine: sign input-to-module links from the sub-model's pathways The static polarity of an edge into a module instance was Positive for every input port, so a loop closed through a DELAY3's delay-time port was labelled r1 while its runtime score is negative at every step, and every structural-only surface (libsimlin and WASM get_loops, the layout, the r/b/u ids themselves) carried the wrong sign. A module's input port has no sign of its own: the sign the parent sees is the sign of the sub-model's internal pathways from that port to the output the parent reads, exactly the collapse the macro treatment applies to magnitudes. CausalGraph::module_input_polarity composes those pathways. Each link on a pathway is signed by the same get_link_polarity rules a top-level link gets (recursively, so a hop into a nested instance composes the same way), so there is one owner of what a link's sign is; the helper only multiplies signs along each pathway and compares. The edge is Positive or Negative when every pathway from every entry port the source feeds, to every output port the parent reads (the sub-model's sinks when it reads none), agrees; it is Unknown when any pathway carries an Unknown link, two pathways or two read ports disagree, the port reaches no read output, or the enumeration was truncated (a dropped pathway could disagree). Read ports are the union over every parent reader, loop or not, because the sign is a property of the edge (the per-exit-port link-score override already scores each loop against the pathway it traverses); a reporting aux that reads a second, opposite-signed port therefore turns the edge and every loop through it Unknown. A fed port that reaches no read output cannot carry a loop and is ignored. Matching the entry port through normalize_module_ref also signs a module-to-module edge wired from an output reference, which fell through to Unknown before, and the stock enrichment uses the same match. Sub-graphs now carry their own instances' graphs, recursively, so a user module wrapping a SMOOTH composes one level down; the recursion is bounded by the module-cycle gate model_detected_loops applies, and compute_link_polarities carries that gate too. Cost: signing a module edge re-enumerates the sub-model's pathways per call, so a graph build pays O(module edges) enumerations instead of one per instance (+28% on a 41-loop DELAY3 model in the reviewer's measurement). A per-(instance, exit set) memo would make it O(instances); it is a known cost, not fixed here. Pinned through model_detected_loops: DELAY3(inp, tau) with tau = 2 + 0.02*s is b1 (every delay_time -> output pathway is Negative: stock/(delay_time/3) on each); SMTH1(0.1*s, 4) stays r1; SMTH1(inp, tau) is b1, following the analyzer's division convention for (input - output)/delay_time (the true sign depends on the gap's sign; the runtime series decides Rux/Bux/U, and narrowing that convention would relabel every gap/adjustment_time link, so it is left alone). One test per arm of the rule in ltm/module_polarity_tests.rs, and pysimlin checks Model.loops and Run.loops agree on b1 for the DELAY3 loop. The module-tests expectation of u1 for a level -> m -> SMTH1 -> inflow loop encoded the old Unknown; every hop of that loop is positive and it is now r1. --- docs/design/ltm--loops-that-matter.md | 35 +- docs/reference/ltm--loops-that-matter.md | 9 + src/pysimlin/tests/test_module_polarity.py | 56 +++ src/simlin-engine/src/db/analysis.rs | 50 ++- src/simlin-engine/src/db/ltm_module_tests.rs | 7 +- src/simlin-engine/src/ltm/graph.rs | 176 +++++++-- .../src/ltm/module_polarity_tests.rs | 367 ++++++++++++++++++ src/simlin-engine/src/ltm/tests.rs | 5 + 8 files changed, 661 insertions(+), 44 deletions(-) create mode 100644 src/pysimlin/tests/test_module_polarity.py create mode 100644 src/simlin-engine/src/ltm/module_polarity_tests.rs diff --git a/docs/design/ltm--loops-that-matter.md b/docs/design/ltm--loops-that-matter.md index c3e540604..5f1a2c1ca 100644 --- a/docs/design/ltm--loops-that-matter.md +++ b/docs/design/ltm--loops-that-matter.md @@ -777,6 +777,28 @@ AST (`Ast`) at compile time. The recursive analysis independent expressions from truly non-monotonic ones - **Flow-to-stock**: Inflows are `Positive`, outflows are `Negative` (fixed structural polarity) +- **Input-to-module** (`CausalGraph::module_input_polarity`): the sign of + the sub-model's own pathways from the entry port(s) the source feeds to + the output port(s) the parent reads (`module_outputs_read`; the + sub-model's sinks when the parent reads nothing). Each pathway's links + are signed by these same rules -- recursively for a hop into a nested + instance, whose graph the sub-graph carries + (`model_variables_and_module_graphs`) -- and multiplied; the edge is + `Positive` / `Negative` when every pathway agrees and `Unknown` when any + pathway carries an `Unknown` link, two pathways or two read ports + disagree, no fed port reaches a read output, or the pathway enumeration + was truncated (a fed port that reaches no read output cannot carry a + loop and is ignored). The read ports are the union over EVERY parent + reader, loop or not, because the sign is a property of the edge: a + reporting aux that reads a second, opposite-signed output turns the + edge -- and the label of every loop through it -- to `u`, even though + the loop exits by the other port and the runtime per-exit-port override + scores it correctly. So a DELAY3's delay-time port is `Negative` + (`stock/(delay_time/3)` on every pathway), its `input` port `Positive`, + and a SMTH1's delay-time port `Negative` by the division convention + above (`(input - output)/delay_time`). The `module -> variable` edge + needs no special arm: the reader's equation names the output + (`module·port`) and the ordinary analysis applies. - **Arrayed equations**: Checks all elements; returns `Unknown` if any two elements disagree @@ -801,12 +823,13 @@ simulation (e.g., the yeast alcohol model from the papers). #### Which surfaces reclassify, and which do not (GH #679) `model_detected_loops` is a *pre-simulation* salsa query, so it can only report -*structural* polarity. Pervasively for module-heavy models the static polarity -of a `variable -> module` / `module -> variable` black-box link is `Unknown`, -so a loop through a module boundary is labelled `Undetermined` (confidence 0.0) -even when its simulated loop score is single-signed at every active step. -Runtime reclassification is therefore a *post-simulation* concern, and the -surfaces handle it differently: +*structural* polarity. A `variable -> module` link is signed from the +sub-model's pathways (see "Static Polarity"), so it is `Unknown` whenever +those pathways disagree or contain an unsigned link -- common in module-heavy +models -- and a loop through such a boundary is labelled `Undetermined` +(confidence 0.0) even when its simulated loop score is single-signed at every +active step. Runtime reclassification is therefore a *post-simulation* +concern, and the surfaces handle it differently: - **Discovery (`analyze_model` / MCP / `simlin_analyze_discover_loops`)**: the `FoundLoop` path in `ltm_finding.rs` derives each loop's polarity directly diff --git a/docs/reference/ltm--loops-that-matter.md b/docs/reference/ltm--loops-that-matter.md index f3178851c..28cc37cf8 100644 --- a/docs/reference/ltm--loops-that-matter.md +++ b/docs/reference/ltm--loops-that-matter.md @@ -645,6 +645,15 @@ Determined from model structure: - **Balancing (B):** Odd number of negative links -> negative loop score - **Undetermined (U):** Any link has unknown polarity (a conservative classification) +> **Simlin implementation note: links into a module.** A link that feeds a module +> instance's input port (a DELAY3's delay time, a user module's input) is signed by +> composing the sub-model's own link polarities along every internal pathway from +> that port to the output(s) the parent reads: every pathway agreeing gives that +> sign; a disagreement, an unsigned link, or a truncated enumeration gives Unknown. +> The papers say nothing about this; it is the macro-collapse principle of Section 6 +> applied to the sign. The link into a DELAY3's delay-time port is therefore negative +> (`stock / (delay_time / 3)` on every pathway) and the link into its input port positive. + #### Runtime Polarity Some models contain links (and therefore loops) that change polarity during simulation. diff --git a/src/pysimlin/tests/test_module_polarity.py b/src/pysimlin/tests/test_module_polarity.py new file mode 100644 index 000000000..24b314bfa --- /dev/null +++ b/src/pysimlin/tests/test_module_polarity.py @@ -0,0 +1,56 @@ +"""Static polarity of a loop through a module's input port. + +The engine composes the sub-model's own link polarities along every pathway +from the entry port to the output the parent reads, so the pre-simulation +``Model.loops`` label of a loop through a delay-time port is the sign the +runtime series shows -- a DELAY3's delay time reaches its output through +``stock / (delay_time / 3)``, a balancing link. +""" + +from __future__ import annotations + +import pytest + +import simlin +from simlin import LoopPolarity +from simlin.types import Aux, Flow, Stock + + +def _delay_time_loop_model(flow_equation: str) -> simlin.Model: + """``s -> tau -> module -> f -> s`` with the module's delay time closing + the loop.""" + project = simlin.Project.new(name="delay_port", sim_start=0.0, sim_stop=30.0, dt=0.5) + model = project.main_model + with model.edit() as (_current, patch): + patch.upsert(Stock(name="s", initial_equation="50", inflows=["f"], outflows=[])) + patch.upsert(Aux(name="tau", equation="2 + 0.02 * s")) + patch.upsert(Aux(name="inp", equation="10")) + patch.upsert(Flow(name="f", equation=flow_equation)) + return model + + +class TestModuleInputPortPolarity: + def test_delay3_delay_time_loop_is_balancing_before_and_after_simulation(self) -> None: + model = _delay_time_loop_model("DELAY3(inp, tau)") + structural = model.loops + assert len(structural) == 1 + assert structural[0].id == "b1" + assert structural[0].polarity == LoopPolarity.BALANCING + assert structural[0].polarity_confidence == pytest.approx(1.0) + + run = model.run(analyze_loops=True) + assert run.ltm_mode == "exhaustive" + assert len(run.loops) == 1 + assert run.loops[0].id == "b1" + assert run.loops[0].polarity == LoopPolarity.BALANCING + + def test_smth1_input_port_loop_stays_reinforcing(self) -> None: + project = simlin.Project.new(name="smth_input", sim_start=0.0, sim_stop=30.0, dt=0.5) + model = project.main_model + with model.edit() as (_current, patch): + patch.upsert(Stock(name="s", initial_equation="50", inflows=["f"], outflows=[])) + patch.upsert(Flow(name="f", equation="SMTH1(0.1 * s, 4)")) + structural = model.loops + assert [(lp.id, lp.polarity) for lp in structural] == [("r1", LoopPolarity.REINFORCING)] + run = model.run(analyze_loops=True) + assert [(lp.id, lp.polarity) for lp in run.loops] == [("r1", LoopPolarity.REINFORCING)] diff --git a/src/simlin-engine/src/db/analysis.rs b/src/simlin-engine/src/db/analysis.rs index 935b861e4..d551fedf0 100644 --- a/src/simlin-engine/src/db/analysis.rs +++ b/src/simlin-engine/src/db/analysis.rs @@ -1465,6 +1465,18 @@ type CausalGraphModuleData = ( /// ports from this sub-graph. A pathless module's sub-graph enumerates no /// pathways, so it is harmless; stock enrichment over a stockless sub-graph /// finds no stocks. +/// +/// Sub-graphs are built recursively -- each carries the graphs of ITS +/// instances -- so a pathway through a nested instance (a user module +/// wrapping a SMOOTH) is signed one level down by the same rule +/// (`CausalGraph::module_input_polarity`). The recursion is bounded by the +/// same gate `model_detected_loops` applies: a model the project's module +/// graph reaches a cycle from gets no sub-graphs at all (its cycle is the +/// model error the diagnostics pass reports, GH #806). That gate reads only +/// explicit instances, which is sound here for the reason +/// `project_module_graph` documents: the implicit instances the recursion +/// also follows target stdlib and macro models, which never instantiate a +/// user model, so no cycle closes through them. fn model_variables_and_module_graphs( db: &dyn Db, model: SourceModel, @@ -1473,16 +1485,23 @@ fn model_variables_and_module_graphs( let edges_result = model_causal_edges(db, model, project); let variables = model_lowered_variables(db, model, project); - let project_models = project.models(db); let mut module_graphs: HashMap, Box> = HashMap::new(); - - for (module_var_name, sub_model_name) in &edges_result.dynamic_modules { - if let Some(sub_source_model) = project_models.get(sub_model_name.as_str()) { - let sub_edges_result = model_causal_edges(db, *sub_source_model, project); - let mut sub_graph = causal_graph_from_edges(sub_edges_result); - sub_graph.variables = model_lowered_variables(db, *sub_source_model, project); - sub_graph.module_outputs_read = Arc::clone(&sub_edges_result.module_outputs_read); - module_graphs.insert(Ident::new(module_var_name), Box::new(sub_graph)); + let reaches_a_cycle = crate::db::project_module_graph(db, project) + .cycle_error_from(model.name(db)) + .is_some(); + if !reaches_a_cycle { + let project_models = project.models(db); + for (module_var_name, sub_model_name) in &edges_result.dynamic_modules { + if let Some(sub_source_model) = project_models.get(sub_model_name.as_str()) { + let sub_edges_result = model_causal_edges(db, *sub_source_model, project); + let mut sub_graph = causal_graph_from_edges(sub_edges_result); + let (sub_variables, sub_outputs_read, nested) = + model_variables_and_module_graphs(db, *sub_source_model, project); + sub_graph.variables = sub_variables; + sub_graph.module_outputs_read = sub_outputs_read; + sub_graph.module_graphs = nested; + module_graphs.insert(Ident::new(module_var_name), Box::new(sub_graph)); + } } } @@ -2919,11 +2938,24 @@ fn detected_polarity_from_ltm(polarity: &crate::ltm::LoopPolarity) -> DetectedLo /// reading each variable's lowered form (`model_lowered_variables`) /// and analyzing how each source variable appears in the target's /// equation. +/// +/// An analysis entry point, so it carries the module-cycle gate +/// `model_detected_loops` applies: a model the project's module graph +/// reaches a cycle from has no link polarities -- the empty map -- because +/// the cycle is the model error the diagnostics pass reports, and signing +/// a module edge walks the instance's sub-graph, which a module cycle makes +/// unbounded (GH #806). pub fn compute_link_polarities( db: &dyn Db, model: SourceModel, project: SourceProject, ) -> HashMap<(String, String), crate::ltm::LinkPolarity> { + if crate::db::project_module_graph(db, project) + .cycle_error_from(model.name(db)) + .is_some() + { + return HashMap::new(); + } let graph = causal_graph_with_modules(db, model, project); graph.all_link_polarities() } diff --git a/src/simlin-engine/src/db/ltm_module_tests.rs b/src/simlin-engine/src/db/ltm_module_tests.rs index 263ee820c..5d091eb2d 100644 --- a/src/simlin-engine/src/db/ltm_module_tests.rs +++ b/src/simlin-engine/src/db/ltm_module_tests.rs @@ -1843,7 +1843,10 @@ fn test_multi_output_module_link_score_holds_document_order_first_live() { /// argument, so the instance is wired straight from `m·a`, and /// `find_model_output_ports` reads every helper's `Dt` reads ("Phase 8.5 /// semantic divergences" 6). The sub-model is instrumented, and a loop -/// through it selects the `via⁚a` exit override at the instance. +/// through it selects the `via⁚a` exit override at the instance. The loop +/// is `r1`: every hop is positive, the two module entries included -- +/// `level -> m` through `a = inp * 0.5` and `m -> smth1` through the +/// smooth's input port, both composed from the sub-models' pathways. #[test] fn a_stdlib_instance_with_bare_arguments_reads_an_output_port() { let child = || { @@ -1902,7 +1905,7 @@ fn a_stdlib_instance_with_bare_arguments_reads_an_output_port() { .push(x_module_named("m", "child", &[(".level", "m.inp")], None)); looped.models.push(child()); let main = ltm_names(&looped, "main"); - for key in ["$|ltm|link_score|level>m|via|a", "$|ltm|loop_score|u1"] { + for key in ["$|ltm|link_score|level>m|via|a", "$|ltm|loop_score|r1"] { assert!(main.contains(&name(key)), "{key} in {main:?}"); } assert_eq!(ltm_names(&looped, "child").len(), 4); diff --git a/src/simlin-engine/src/ltm/graph.rs b/src/simlin-engine/src/ltm/graph.rs index a229d9e16..0a638a1e3 100644 --- a/src/simlin-engine/src/ltm/graph.rs +++ b/src/simlin-engine/src/ltm/graph.rs @@ -28,7 +28,10 @@ use super::partitions::{CyclePartitions, tarjan_scc}; use super::polarity::{ analyze_agg_consumer_polarity, analyze_link_polarity, compose_with_lookup_polarity, }; -use super::types::{Link, LinkPolarity, Loop, LoopPolarity, TruncatedByBudget}; +use super::types::{ + Link, LinkPolarity, Loop, LoopPolarity, TruncatedByBudget, is_synthetic_node_name, + normalize_module_ref, +}; /// Internal module pathways keyed by input port: each port maps to its list of /// open `input -> ... -> output` link-paths. @@ -492,34 +495,43 @@ impl CausalGraph { let pred_idx = if i == 0 { circuit.len() - 1 } else { i - 1 }; let predecessor = &circuit[pred_idx]; - // Find which input port the predecessor maps to. - let internal_port = module_var - .iter() - .find(|inp| &inp.src == predecessor) - .map(|inp| &inp.dst); + // The input port(s) the predecessor feeds -- the same match the + // polarity composition makes (`entry_ports`), so a hop wired + // from another instance's output resolves here too. + let ports: Vec<&Ident> = entry_ports(module_var, predecessor).collect(); let (pathways, truncated_ports) = module_graph.enumerate_pathways_to_outputs_with_truncation(&[]); - let internal_stocks: Vec> = if let Some(port) = internal_port { - // Collect stocks from all pathways for the matched input port. - // When that port's pathway enumeration hit the budget (GH #649) + let internal_stocks: Vec> = if ports.is_empty() { + // Predecessor doesn't match any module input (shouldn't happen + // with a well-formed graph). Conservative fallback. + all_module_stocks(module_graph, node) + } else { + // Collect stocks from all pathways of every matched input port. + // When a port's pathway enumeration hit the budget (GH #649) // the kept paths are a prefix and could miss a stock that lives // only on a dropped pathway, so degrade to the conservative // "all module-internal stocks" fallback rather than silently - // dropping stocks from the enriched loop. - match pathways.get(port) { - Some(paths) if !truncated_ports.contains(port) => { - collect_stocks_from_pathways(module_graph, paths, node) + // dropping stocks from the enriched loop; a port with no + // pathway at all falls back the same way. + let mut collected: Vec> = Vec::new(); + for port in ports { + match pathways.get(port) { + Some(paths) if !truncated_ports.contains(port) => { + for stock in collect_stocks_from_pathways(module_graph, paths, node) { + if !collected.contains(&stock) { + collected.push(stock); + } + } + } + _ => { + collected = all_module_stocks(module_graph, node); + break; + } } - // Port has no pathway, or its enumeration was truncated: - // fall back to all module-internal stocks. - _ => all_module_stocks(module_graph, node), } - } else { - // Predecessor doesn't match any module input (shouldn't happen - // with a well-formed graph). Conservative fallback. - all_module_stocks(module_graph, node) + collected }; for s in internal_stocks { @@ -1166,15 +1178,22 @@ impl CausalGraph { // If 'from' is not a flow for this stock, fall through to AST analysis } - // When the target is a module, the edge represents an input - // feeding into the module. Module inputs are direct bindings - // (positive relationship). If the module has an internal graph, - // we could trace through it, but for the input->module edge - // itself the polarity is positive. + // When the target is a module instance the edge feeds one (or + // more) of its input ports, and the sign the parent sees is the + // sign of the sub-model's internal pathways from that port to + // the output(s) the parent reads -- a DELAY3's delay-time port + // reaches `output` through `stock/(delay_time/3)` and is + // Negative, its `input` port Positive. Compose those pathways + // rather than labelling every input edge Positive: the loop id + // (`r`/`b`/`u` prefix) and every structural-only surface read + // this sign, and a wrong one is a loop reported reinforcing that + // scores negative at every step. if let VarKind::Module { inputs, .. } = &to_var.kind - && inputs.iter().any(|inp| &inp.src == from) + && inputs + .iter() + .any(|inp| &normalize_module_ref(&inp.src) == from) { - return LinkPolarity::Positive; + return self.module_input_polarity(from, to, inputs); } // General case: analyze the equation AST. The AST is the RAW @@ -1191,6 +1210,93 @@ impl CausalGraph { LinkPolarity::Unknown } + /// Static polarity of an `input -> module` edge: the sign of the + /// sub-model's own pathways from the entry port(s) `from` feeds to the + /// output port(s) the parent reads. + /// + /// Every pathway is a chain of the sub-model's links, each signed by the + /// same [`get_link_polarity`](Self::get_link_polarity) rules a top-level + /// link gets -- recursively: the sub-graph carries its own nested + /// instances' graphs (`db::analysis::model_variables_and_module_graphs`), + /// so a hop into a nested instance composes the same way one level down + /// -- so there is one owner of what a link's sign is; this only + /// multiplies the signs along each pathway and compares. The edge is + /// `Positive` or `Negative` when every pathway agrees on that sign, and + /// `Unknown` when any pathway carries an `Unknown` link, two pathways + /// (or two read ports, or two entry ports fed by `from`) disagree, no + /// fed entry port reaches a read output, or the pathway enumeration was + /// truncated (a dropped pathway could disagree). A fed port that + /// reaches no read output cannot carry a loop, so it is ignored rather + /// than vetoing the sign the other ports give. + /// + /// The exit ports are the union of what EVERY parent reader reads from + /// the instance (`module_outputs_read`), loop or not: a link's sign is + /// a property of the edge, not of the port one particular loop exits by + /// (the runtime per-exit-port override, `compute_module_link_overrides`, + /// scores each loop against the pathway it traverses), so a reporting + /// variable reading a second, opposite-signed port turns the edge -- and + /// every loop through it -- `Unknown`. When the parent reads nothing + /// from the instance the sub-model's own output convention decides + /// (`enumerate_pathways_to_outputs_with_truncation` with no ports: its + /// sinks), the same fallback the composite path takes. + fn module_input_polarity( + &self, + from: &Ident, + module: &Ident, + inputs: &[crate::variable::ModuleInput], + ) -> LinkPolarity { + let Some(module_graph) = self.module_graphs.get(module) else { + return LinkPolarity::Unknown; + }; + // The nodes of the sub-graph the parent reads: `x` for a read of + // `m·x`, the nested instance `n` for a read of `m·n·x`. A read of + // the instance's LTM internals (`m·$⁚ltm⁚…`) names no model output + // -- the rule `db::unique_module_output` applies when it picks a + // loop's exit port; no source variable reads one today (the LTM + // synthetics are emitted after the dependency walk), so the filter is + // defensive. + let mut exit_nodes: Vec> = self + .module_outputs_read + .values() + .flatten() + .filter(|t| t.module_path.first() == Some(module)) + .filter(|t| !is_synthetic_node_name(t.variable.as_str())) + .map(|t| { + t.module_path + .get(1) + .cloned() + .unwrap_or_else(|| t.variable.clone()) + }) + .collect(); + exit_nodes.sort(); + exit_nodes.dedup(); + let (pathways, truncated_ports) = + module_graph.enumerate_pathways_to_outputs_with_truncation(&exit_nodes); + + let mut sign: Option = None; + for port in entry_ports(inputs, from) { + if truncated_ports.contains(port) { + return LinkPolarity::Unknown; + } + // A port that reaches no read output cannot carry a loop. + let Some(paths) = pathways.get(port) else { + continue; + }; + for path in paths { + let path_sign = path.iter().fold(LinkPolarity::Positive, |acc, link| { + acc.compose(link.polarity) + }); + match (sign, path_sign) { + (_, LinkPolarity::Unknown) => return LinkPolarity::Unknown, + (Some(prev), next) if prev != next => return LinkPolarity::Unknown, + (None, next) => sign = Some(next), + (Some(_), _) => {} + } + } + } + sign.unwrap_or(LinkPolarity::Unknown) + } + /// Polarity of `consumer`'s equation with respect to a reducer /// subexpression `reducer_subexpr_text` -- the polarity of a synthetic /// aggregate-node hop `$⁚ltm⁚agg → consumer` (GH #516). @@ -1281,6 +1387,22 @@ impl CausalGraph { } } +/// The input ports of a module instance that `from` feeds: the `dst` of +/// every input whose source, with a `·output` suffix stripped +/// (`normalize_module_ref`, so an input wired from another instance's +/// output matches that instance's node), is `from`. The one owner of the +/// entry-port match, read by the polarity composition and the stock +/// enrichment alike. +fn entry_ports<'a>( + inputs: &'a [crate::variable::ModuleInput], + from: &'a Ident, +) -> impl Iterator> + 'a { + inputs + .iter() + .filter(move |inp| &normalize_module_ref(&inp.src) == from) + .map(|inp| &inp.dst) +} + /// Collect stocks from a set of internal pathways, namespaced with the /// module instance name (using interpunct separator). fn collect_stocks_from_pathways( diff --git a/src/simlin-engine/src/ltm/module_polarity_tests.rs b/src/simlin-engine/src/ltm/module_polarity_tests.rs new file mode 100644 index 000000000..876eef05e --- /dev/null +++ b/src/simlin-engine/src/ltm/module_polarity_tests.rs @@ -0,0 +1,367 @@ +// Copyright 2026 The Simlin Authors. All rights reserved. +// Use of this source code is governed by the Apache License, +// Version 2.0, that can be found in the LICENSE file. + +//! Static polarity of an `input -> module` edge: the composition of the +//! sub-model's own link polarities along every pathway from the entry port +//! `from` feeds to the output port(s) the parent reads. One test per arm of +//! the rule, each through the real pipeline (`sync_from_datamodel` -> +//! `model_detected_loops`) so the loop id and polarity a user sees are what +//! is pinned; the edge-level tests read the same `get_link_polarity` the loop +//! builders call. +//! +//! Arms: every pathway Negative (DELAY3's delay-time port); every pathway +//! Positive (SMTH1's input port); Negative by the Div convention (SMTH1's +//! delay-time port); pathways of one port disagreeing; two read output ports +//! disagreeing; an entry port with no pathway to a read output; a truncated +//! pathway enumeration; one source feeding two agreeing entry ports; an +//! Unknown link inside a pathway; a module the parent reads nothing from (the +//! sub-model's own output convention decides); and the `module -> variable` +//! edge, unchanged, in `test_module_polarity_through_output_ref`. + +use super::*; +use crate::db::DetectedLoopsResult; +use crate::ltm::ModulePathwayBudgetGuard; +use crate::test_common::TestProject; +use crate::testutils::x_module; + +/// The detected loops of `project`'s `main` model. +fn detected(project: &crate::datamodel::Project) -> DetectedLoopsResult { + let db = SimlinDb::default(); + let sync = sync_from_datamodel(&db, project); + model_detected_loops(&db, sync.models["main"].source, sync.project).clone() +} + +/// `(id, polarity)` of the single detected loop of `project`'s `main` model. +fn the_loop(project: &crate::datamodel::Project) -> (String, DetectedLoopPolarity) { + let detected = detected(project); + assert_eq!( + detected.loops.len(), + 1, + "expected exactly one loop, got {:?}", + detected + .loops + .iter() + .map(|l| (&l.id, &l.variables)) + .collect::>() + ); + (detected.loops[0].id.clone(), detected.loops[0].polarity) +} + +/// The static polarity of the parent-model edge `from -> to` as the loop +/// builders read it. +fn edge_polarity(project: &crate::datamodel::Project, from: &str, to: &str) -> LinkPolarity { + let db = SimlinDb::default(); + let sync = sync_from_datamodel(&db, project); + let graph = causal_graph_with_modules(&db, sync.models["main"].source, sync.project); + graph.get_link_polarity(&Ident::new(from), &Ident::new(to)) +} + +/// A stock `s` whose inflow `f` is a stdlib module call closed through +/// `tau = 2 + 0.02 * s`: `s -> tau -> module -> f -> s`. +fn stdlib_delay_port_project(flow_eqn: &str) -> crate::datamodel::Project { + TestProject::new("delay_port") + .with_sim_time(0.0, 30.0, 0.5) + .stock("s", "50", &["f"], &[], None) + .aux("tau", "2 + 0.02 * s", None) + .aux("inp", "10", None) + .flow("f", flow_eqn, None) + .build_datamodel() +} + +/// A parent model `s -> x -> m -> f -> s` around the user sub-model `m` +/// (`x_module` instantiates the model of the same name) whose read output is +/// `f_eqn`'s reference. +fn user_module_project( + sub_model: crate::datamodel::Model, + refs: &[(&str, &str)], + f_eqn: &str, +) -> crate::datamodel::Project { + let main = x_model( + "main", + vec![ + x_stock("s", "100", &["f"], &[], None), + x_aux("x", "0.1 * s", None), + x_module("m", refs, None), + x_flow("f", f_eqn, None), + ], + ); + x_project(sim_specs_with_units("years"), &[main, sub_model]) +} + +#[test] +fn delay_time_port_of_delay3_makes_the_loop_balancing() { + // Inside delay3 every pathway from `delay_time` to `output` is Negative: + // the direct `output = stock_3/(delay_time/3)`, and the ones through + // `flow_2` and `flow_1` into the stock chain (each `stock_n/(delay_time/3)` + // flow is Negative in the delay time, the stock hops Positive). The + // edge is Negative, the loop has one negative link, and its id says so; + // its runtime score is negative at every step. + let project = stdlib_delay_port_project("DELAY3(inp, tau)"); + let module = "$\u{205A}f\u{205A}0\u{205A}delay3"; + assert_eq!( + edge_polarity(&project, "tau", module), + LinkPolarity::Negative + ); + assert_eq!( + the_loop(&project), + ("b1".to_string(), DetectedLoopPolarity::Balancing) + ); +} + +#[test] +fn input_port_of_smth1_keeps_the_loop_reinforcing() { + // `s -> SMTH1 at input (through the hoisted `0.1 * s` helper) -> f -> s`: + // the one pathway `input -> flow -> output` is Positive. + let project = TestProject::new("smth_input") + .with_sim_time(0.0, 30.0, 0.5) + .stock("s", "50", &["f"], &[], None) + .flow("f", "SMTH1(0.1 * s, 4)", None) + .build_datamodel(); + let module = "$\u{205A}f\u{205A}0\u{205A}smth1"; + let helper = "$\u{205A}f\u{205A}0\u{205A}arg0"; + assert_eq!( + edge_polarity(&project, helper, module), + LinkPolarity::Positive + ); + assert_eq!( + the_loop(&project), + ("r1".to_string(), DetectedLoopPolarity::Reinforcing) + ); +} + +#[test] +fn delay_time_port_of_smth1_follows_the_division_convention() { + // `s -> tau -> SMTH1 at delay_time -> f -> s`. The one pathway is + // `delay_time -> flow -> output` with `flow = (input - Output)/delay_time`: + // the analyzer's Div arm labels a non-constant divisor Negative under its + // positive-numerator convention (`(a - b)/y` flips, the same reading + // `gap / adjustment_time` gets), so the loop is reported balancing. The + // true sign depends on the sign of the gap `input - Output`; the runtime + // series settles Rux/Bux/U, while the static label follows the convention + // every other `x / y` link in the model follows. + let project = stdlib_delay_port_project("SMTH1(inp, tau)"); + let module = "$\u{205A}f\u{205A}0\u{205A}smth1"; + assert_eq!( + edge_polarity(&project, "tau", module), + LinkPolarity::Negative + ); + assert_eq!( + the_loop(&project), + ("b1".to_string(), DetectedLoopPolarity::Balancing) + ); +} + +#[test] +fn pathways_of_one_port_that_disagree_make_the_edge_unknown() { + // `input -> pos -> output` is Positive and `input -> neg -> output` is + // Negative: the port's sign is not a sign. + let sub = x_model( + "m", + vec![ + x_aux("input", "0", None), + x_aux("pos", "input", None), + x_aux("neg", "0 - input", None), + x_aux("output", "pos + neg", None), + ], + ); + let project = user_module_project(sub, &[("x", "m.input")], "m.output"); + assert_eq!(edge_polarity(&project, "x", "m"), LinkPolarity::Unknown); + assert_eq!( + the_loop(&project), + ("u1".to_string(), DetectedLoopPolarity::Undetermined) + ); +} + +#[test] +fn two_read_output_ports_that_disagree_make_the_edge_unknown() { + // The parent reads both `out_pos = input` and `out_neg = 0 - input`; the + // edge's sign is a property of the edge, not of the port one loop exits + // by, so two read ports of opposite sign leave it Unknown. + let sub = x_model( + "m", + vec![ + x_aux("input", "0", None), + x_aux("out_pos", "input", None), + x_aux("out_neg", "0 - input", None), + ], + ); + let project = user_module_project(sub, &[("x", "m.input")], "m.out_pos + m.out_neg"); + assert_eq!(edge_polarity(&project, "x", "m"), LinkPolarity::Unknown); + assert_eq!( + the_loop(&project), + ("u1".to_string(), DetectedLoopPolarity::Undetermined) + ); +} + +#[test] +fn an_entry_port_with_no_pathway_to_a_read_output_is_unknown() { + // `output` does not depend on `input` at all (only `sink` reads it). The + // causal graph still carries `x -> m -> f` (the parent feeds the port and + // reads the output), so the loop is enumerated; its sign cannot be + // composed from anything. + let sub = x_model( + "m", + vec![ + x_aux("input", "0", None), + x_aux("sink", "input", None), + x_aux("output", "5", None), + ], + ); + let project = user_module_project(sub, &[("x", "m.input")], "m.output"); + assert_eq!(edge_polarity(&project, "x", "m"), LinkPolarity::Unknown); + assert_eq!( + the_loop(&project), + ("u1".to_string(), DetectedLoopPolarity::Undetermined) + ); +} + +#[test] +fn a_truncated_pathway_enumeration_makes_the_edge_unknown() { + // A diamond (`input -> a`, `input -> b`, `output = a + b`) has two + // Positive pathways. Under a budget of one the enumeration keeps one + // pathway and reports truncation; a dropped pathway could disagree, so + // the edge is Unknown -- and with the full enumeration it is Positive. + let sub = x_model( + "m", + vec![ + x_aux("input", "0", None), + x_aux("a", "input * 2", None), + x_aux("b", "input * 3", None), + x_aux("output", "a + b", None), + ], + ); + let project = user_module_project(sub, &[("x", "m.input")], "m.output"); + assert_eq!(edge_polarity(&project, "x", "m"), LinkPolarity::Positive); + let _guard = ModulePathwayBudgetGuard::new(1); + assert_eq!(edge_polarity(&project, "x", "m"), LinkPolarity::Unknown); +} + +#[test] +fn a_source_feeding_two_entry_ports_with_agreeing_pathways_keeps_the_sign() { + // `x` is wired to both `in1` and `in2`; `output = in1 + in2` reaches the + // output from each port with the same sign. + let sub = x_model( + "m", + vec![ + x_aux("in1", "0", None), + x_aux("in2", "0", None), + x_aux("output", "in1 + in2", None), + ], + ); + let project = user_module_project(sub, &[("x", "m.in1"), ("x", "m.in2")], "m.output"); + assert_eq!(edge_polarity(&project, "x", "m"), LinkPolarity::Positive); + assert_eq!( + the_loop(&project), + ("r1".to_string(), DetectedLoopPolarity::Reinforcing) + ); +} + +#[test] +fn an_unknown_link_inside_the_pathway_makes_the_edge_unknown() { + // `output = ABS(input)` is non-monotone, so the only pathway carries an + // Unknown link. + let sub = x_model( + "m", + vec![ + x_aux("input", "0", None), + x_aux("output", "ABS(input)", None), + ], + ); + let project = user_module_project(sub, &[("x", "m.input")], "m.output"); + assert_eq!(edge_polarity(&project, "x", "m"), LinkPolarity::Unknown); + assert_eq!( + the_loop(&project), + ("u1".to_string(), DetectedLoopPolarity::Undetermined) + ); +} + +#[test] +fn a_module_the_parent_reads_nothing_from_uses_its_own_output_convention() { + // Nothing in the parent reads `m`, so no loop runs through it; the + // `x -> m` edge still has a sign, composed to the sub-model's own output + // convention (its sinks): `output = 0 - input` is Negative. + let sub = x_model( + "m", + vec![ + x_aux("input", "0", None), + x_aux("output", "0 - input", None), + ], + ); + let main = x_model( + "main", + vec![ + x_aux("x", "3", None), + x_module("m", &[("x", "m.input")], None), + ], + ); + let project = x_project(sim_specs_with_units("years"), &[main, sub]); + assert_eq!(edge_polarity(&project, "x", "m"), LinkPolarity::Negative); + assert!(detected(&project).loops.is_empty()); +} + +#[test] +fn a_nested_instance_composes_recursively() { + // A user module wrapping a stdlib smooth: `m.output = SMTH1(m.input, 3)`. + // The pathway `input -> $smth1 -> output` inside `m` crosses a nested + // instance, whose hop is signed by the same rule one level down (the + // smooth's `input -> flow -> output` is Positive), so the parent's + // `x -> m` edge is Positive and the loop keeps its reinforcing id. + let sub = x_model( + "m", + vec![ + x_aux("input", "0", None), + x_aux("output", "SMTH1(input, 3)", None), + ], + ); + let project = user_module_project(sub, &[("x", "m.input")], "m.output"); + assert_eq!(edge_polarity(&project, "x", "m"), LinkPolarity::Positive); + assert_eq!( + the_loop(&project), + ("r1".to_string(), DetectedLoopPolarity::Reinforcing) + ); +} + +#[test] +fn a_source_feeding_two_entry_ports_with_disagreeing_pathways_is_unknown() { + // `x` is wired to both `in1` and `in2`; `output = in1 - in2` reaches the + // output Positive from one port and Negative from the other, so the + // edge has no sign -- the second port is consulted, not just the first. + let sub = x_model( + "m", + vec![ + x_aux("in1", "0", None), + x_aux("in2", "0", None), + x_aux("output", "in1 - in2", None), + ], + ); + let project = user_module_project(sub, &[("x", "m.in1"), ("x", "m.in2")], "m.output"); + assert_eq!(edge_polarity(&project, "x", "m"), LinkPolarity::Unknown); + assert_eq!( + the_loop(&project), + ("u1".to_string(), DetectedLoopPolarity::Undetermined) + ); +} + +#[test] +fn a_fed_port_that_reaches_no_read_output_is_ignored() { + // `x` is wired to both `in1` and `in2`, but only `in1` reaches the read + // output (`in2` feeds a dead-end `sink`). A port that reaches no read + // output cannot carry the loop, so it does not veto the sign the other + // port gives. + let sub = x_model( + "m", + vec![ + x_aux("in1", "0", None), + x_aux("in2", "0", None), + x_aux("sink", "in2", None), + x_aux("output", "in1 * 2", None), + ], + ); + let project = user_module_project(sub, &[("x", "m.in1"), ("x", "m.in2")], "m.output"); + assert_eq!(edge_polarity(&project, "x", "m"), LinkPolarity::Positive); + assert_eq!( + the_loop(&project), + ("r1".to_string(), DetectedLoopPolarity::Reinforcing) + ); +} diff --git a/src/simlin-engine/src/ltm/tests.rs b/src/simlin-engine/src/ltm/tests.rs index 590b709b1..e265a4ee3 100644 --- a/src/simlin-engine/src/ltm/tests.rs +++ b/src/simlin-engine/src/ltm/tests.rs @@ -4859,3 +4859,8 @@ mod polarity; /// `use super::*` reaches this file's private helpers. #[path = "with_lookup_tests.rs"] mod with_lookup; + +/// Static polarity of `input -> module` edges (pathway composition), in a +/// sibling file for the same line-count reason. +#[path = "module_polarity_tests.rs"] +mod module_polarity; From 3afd63764ce48259f9d9a3a0542771637300a71b Mon Sep 17 00:00:00 2001 From: Bobby Powers Date: Tue, 8 Sep 2026 22:07:42 -0700 Subject: [PATCH 04/10] engine: run LTM under every integration method The overlay was refused under RK2/RK4 whenever a model emitted a flow-to-stock link score (GH #486). The only justification was the old score's stock-history numerator, which aligned to the Euler update; the score is now a ratio of flow deltas over the stock's net-flow aux and reads no stock history, so the refusal has no premise left. A link score is a ratio of integration-step (dt) deltas reported at the saved steps: `PREVIOUS` reads the state the previous dt step ended in, because the VM snapshots `prev_values` on every dt iteration before the save/advance logic decides whether the row is recorded, so with save_step > dt a recorded score is the last dt step's ratio and not a re-differencing over the saved interval. That is the 2020 paper's form (Schoenberg, Davidsen and Eberlein, section 6.1: "computed at each dt"). Under RK2/RK4 the VM re-evaluates the flows at the restored end-of-step state before snapshotting it (the RK stages' trial-point evaluations are overwritten; wasm mirrors this), so the ratio is taken over the method's own trajectory. The paper puts Runge-Kutta compatibility as "in principle"; the design doc's note states this form and its boundary: cross-method equality holds for flows proportional to their stock, and in general the scores follow the method's trajectory. Pinned (tests/integration/ltm_integration_method.rs): an isolated single-flow loop and the births/deaths model score identically under Euler, RK2 and RK4 -- every link and loop score and every relative loop score equal to 1e-9 at every saved step -- while their stock trajectories differ; a nonlinear outflow (0.02 * s ^ 1.3 against births 0.1 * s) scores -22.744 under Euler and -22.749 under RK4 at t = 2; with save_step = 2 * dt the recorded scores are the dt-resolution run's at the saved times, and the ratio over the whole saved interval is a different number by more than 0.1 (a change to when prev_values is snapshotted fails this). Elsewhere: an RK4 model's LTM columns agree between wasm and the VM; mark2.mdl (a Vensim RK4 model) compiles and runs with the overlay and every model series is the LTM-free run's; through libsimlin an LTM sim on an RK4 model reports exhaustive mode and leaves the trajectory untouched; the MCP read_model reports an RK4 model's loop with no analysisError; pysimlin's run() on an RK4 model neither warns nor falls back to a run without loop analysis, and analyze() reports its loop. Removed with the guard: `effective_non_euler_method`, `model_emits_flow_to_stock_score`, `ltm_non_euler_diagnostic_message`, the GH #486/#663 test block in ltm_unified_tests (its per-element decorated-name probes included), and the refusal-expecting tests in libsimlin, MCP core, pysimlin and the wasm harness. The GH #660 "compile failure surfaces analysis_error" arm keeps its mechanism on a model whose flow reads an undefined variable at every surface (engine, libsimlin, MCP read_model on the wire, pysimlin); the message names the variable that failed to compile, never the missing reference. TestProject gains the `with_save_step` its `with_sim_time` comment already pointed at. The C header is regenerated from the FFI docs; tech-debt item 31 is resolved. --- docs/design/ltm--loops-that-matter.md | 34 +- docs/tech-debt.md | 4 +- src/libsimlin/CLAUDE.md | 2 +- src/libsimlin/simlin.h | 13 +- src/libsimlin/src/ffi.rs | 4 +- src/libsimlin/src/lib.rs | 2 +- src/libsimlin/src/project.rs | 9 +- src/libsimlin/src/simulation.rs | 16 +- src/libsimlin/tests/integration/analysis.rs | 19 +- src/libsimlin/tests/integration/errors.rs | 28 +- src/libsimlin/tests/integration/simulation.rs | 55 +-- src/pysimlin/simlin/analysis.py | 6 +- src/pysimlin/simlin/model.py | 5 +- src/pysimlin/tests/test_discovery.py | 52 ++- src/pysimlin/tests/test_ltm_mode_and_links.py | 43 +- src/simlin-engine/CLAUDE.md | 2 +- src/simlin-engine/src/analysis.rs | 92 ++-- src/simlin-engine/src/db/assemble.rs | 62 --- src/simlin-engine/src/db/ltm/mod.rs | 112 ----- src/simlin-engine/src/db/ltm_unified_tests.rs | 437 ------------------ src/simlin-engine/src/test_common.rs | 8 + .../integration/ltm_integration_method.rs | 243 ++++++++++ src/simlin-engine/tests/integration/main.rs | 1 + .../tests/integration/simulate.rs | 50 +- .../tests/integration/simulate_ltm_wasm.rs | 41 +- src/simlin-mcp-core/CLAUDE.md | 4 +- src/simlin-mcp-core/src/tools/edit_model.rs | 7 +- src/simlin-mcp-core/src/tools/read_model.rs | 5 +- .../tests/integration/read_model_e2e.rs | 81 +++- 29 files changed, 558 insertions(+), 879 deletions(-) create mode 100644 src/simlin-engine/tests/integration/ltm_integration_method.rs diff --git a/docs/design/ltm--loops-that-matter.md b/docs/design/ltm--loops-that-matter.md index e8a6156f3..1b747bb9a 100644 --- a/docs/design/ltm--loops-that-matter.md +++ b/docs/design/ltm--loops-that-matter.md @@ -1935,13 +1935,33 @@ enumerated loops. ## Current Limitations -### Euler Integration Only - -`assemble_simulation` refuses the overlay under RK2/RK4 (GH #486) when any -instantiated model emits a flow-to-stock score. The scores are differences of -saved-step values, so the guard keeps them on Euler-stepped trajectories; the -2020 paper (section 6.1) says the method is compatible with Runge-Kutta "in -principle", and Simlin has not established that for RK-stepped runs. +### Integration Methods and Save Step + +LTM runs under Euler, RK2 and RK4 alike (GH #486). A link score is a ratio of +integration-step (dt) deltas, reported at the saved steps: `PREVIOUS` reads the +state the previous dt step ended in, because the VM snapshots `prev_values` on +every dt iteration before the save/advance logic decides whether the row is +recorded. With `save_step > dt` a recorded score is therefore the ratio over +the last dt step ending at that time -- the same number a `save_step = dt` run +records there -- and not a re-differencing of the flows over the saved +interval (`tests/integration/ltm_integration_method.rs` pins the two apart on +a nonlinear flow, where they differ by more than a unit of score). This is the +2020 paper's form (Schoenberg, Davidsen and Eberlein, section 6.1: the scores +are "computed at each dt"). + +Under RK2/RK4 the VM re-evaluates the flows at the restored end-of-step state +before snapshotting it (the RK stages' trial-point evaluations are +overwritten; wasm mirrors this), so the dt-step ratio is taken over the +method's own trajectory and never over an intra-step stage evaluation. The +paper puts Runge-Kutta compatibility as "in principle" ("could in principle +work ... with Runge-Kutta integration"); this is the form it takes here. Note +the boundary of what that buys: a model whose flows are proportional to its +stock scores identically under all three methods while its stock trajectories +differ, because the ratios of flow deltas cancel the stock's step, but that is +a property of proportional flows, not of the method. In general the scores +follow the method's trajectory: on the test's nonlinear model (`deaths = +0.02 * s ^ 1.3` against `births = 0.1 * s`) the deaths-to-stock score at +`t = 2` is -22.744 under Euler and -22.749 under RK4. ### Performance on Very Large Models diff --git a/docs/tech-debt.md b/docs/tech-debt.md index 94c57aa17..4d0237441 100644 --- a/docs/tech-debt.md +++ b/docs/tech-debt.md @@ -279,8 +279,8 @@ Known debt items consolidated from CLAUDE.md files and codebase analysis. Each e ### 31. RK4 + LTM Combination Has No Hard-Error Guard - **Component**: simlin-engine (src/simlin-engine/src/ltm_augment.rs flow-to-stock path) -- **Severity**: medium -- **Description**: LTM link scores are differences of saved-step values (a flow-to-stock score is `|Δflow / Δnet|` over the stock's net-flow aux), and whether they describe an RK-stepped trajectory faithfully has not been established, so `assemble_simulation` refuses LTM under RK2/RK4 (GH #486) when any instantiated model emits a flow-to-stock score. Open question: lift the refusal once the RK case is validated (the 2020 paper's section 6.1 calls the method compatible with Runge-Kutta "in principle"). +- **Severity**: RESOLVED (2026-09-07) +- **Description**: (**Resolved**.) The flow-to-stock link score is `|Δflow / Δnet|` over the stock's synthetic net-flow aux, a ratio of integration-step (dt) deltas with no stock history behind it, and the VM and wasm both re-evaluate the flows at the restored end-of-step state before snapshotting it under RK2/RK4, so LTM runs under every integration method and needs no guard (`tests/integration/ltm_integration_method.rs` pins Euler, RK2 and RK4 identical on an isolated loop and on the births/deaths model, whose flows are proportional to the stock, and pins a nonlinear flow following each method's own trajectory). Keeping the entry as a pointer to the design doc's "Integration Methods and Save Step" note. - **Tracked in**: #486 (LTM tracking epic: #488) - **Owner**: unassigned - **Last reviewed**: 2026-04-29 diff --git a/src/libsimlin/CLAUDE.md b/src/libsimlin/CLAUDE.md index 16e3d3058..598312916 100644 --- a/src/libsimlin/CLAUDE.md +++ b/src/libsimlin/CLAUDE.md @@ -27,7 +27,7 @@ Error formatting has no module here: `src/patch.rs` imports `simlin_engine::erro - `simlin_project_replace_contents(dst, src, out_error)` - The in-place reload primitive: `dst`'s `datamodel::Project` becomes a deep clone of `src`'s and `dst`'s salsa db is re-synced incrementally (`db.sync`), so unchanged variables keep their compile fragments. A caller reloading a project from disk opens the new bytes with the matching `simlin_project_open_*` into a scratch project and replaces from it -- every format is covered by composition, so there are deliberately no per-format replace variants. Live-handle contract: a `SimlinModel` holds `*const SimlinProject` + a model NAME, so existing handles on `dst` stay valid and observe the new contents (variables, `get_errors`, a fresh `simlin_sim_new`); a handle whose model name is absent from the replacement returns `BadModelName` from queries and its sim fails on first run with `NotSimulatable` naming the model (never UB), and works again once a model of that name reappears; a `SimlinSim` created BEFORE the replace is a stale snapshot for the `simlin_sim_*` entry points (its results and `reset`/re-run stay bound to the program it compiled -- the same posture `apply_patch` takes toward existing sims), but NOT for the sim-bearing analysis FFIs (`simlin_analyze_get_loops_runtime`, `simlin_analyze_get_links`, ...), which enumerate from the CURRENT db and read the stale results by positional loop id, so they mix old results with the new model; callers re-run after a replace before analyzing. Locking: `src`'s datamodel lock is taken alone and released before `dst`'s datamodel-then-db locks are acquired, so opposite-direction replaces cannot deadlock and `dst == src` is a permitted no-op re-sync; the db re-sync and the datamodel swap happen under both `dst` locks so no reader sees them disagree (`tests_concurrency.rs::test_replace_contents_lock_order_matches_readers` pins the order: an inversion deadlocks against the readers and times out). Refcounts of both projects are untouched; nothing crosses the boundary that a caller must free - `simlin_project_get_model()` - Get model handle by name (or default) - `simlin_project_is_simulatable()` - Check compilability - - `simlin_project_get_errors()` - Collect all diagnostics. Project compilability is an INTRINSIC property, assessed with LTM OFF: the compile/VM-validation channel (`vm_error`) is always computed with the LTM overlay `Off` (`engine::db::LtmOverlay`), so an LTM-only rejection (the GH #486 non-Euler hard `Err` from `assemble_simulation`, which rides the compile path) never masquerades as a project error on a model that simulates fine without LTM. LTM is an analysis overlay, not part of whether the model is a valid simulation. Separately, the LTM *diagnostics* (auto-flip-to-discovery advisory, synthetic-fragment compile failures, the GH #311 partial-equation warnings) accumulate via `model_all_diagnostics` -> `model_ltm_variables` independent of assembly success, so when the project's `ltm_requested` latch is set (any prior `simlin_sim_new` with `enable_ltm=true`) `get_errors` runs the `collect_all_diagnostics` harvest with the overlay `On` -- NOT a recompile that feeds `vm_error` -- so those advisories reach the caller (GH #466; `simlin-mcp-core` harvests the same way, GH #662). The overlay is an argument of every query rather than a salsa input, so the `Off` and `On` harvests are memoized side by side and a concurrent sim creation never observes a partial LTM state; salsa memoizes the LTM synthesis from the `sim_new` compile, so the `On` harvest revalidates rather than recomputes. A project that never requested LTM harvests `Off` and pays no LTM synthesis cost (preserving the scoping the engine's `test_ltm_disabled_gate_suppresses_auto_flip_warning` pins). The `ltm_requested` latch is a set-once `AtomicBool` on `SimlinProject`. Each `SimlinErrorDetail` carries a `severity` field (`SimlinErrorSeverity { Error, Warning }`) read off `errors::FormattedError.severity`, which the engine populates from the originating `Diagnostic.severity` (a `format_simulation_error` detail -- a compile/VM-validation failure with no diagnostic behind it -- is `Error` by construction). Threading it on the `FormattedError` rather than alongside it (GH #919) means the severity can neither be forgotten at a call site nor drift from the wording of `message`, which the engine now words from the same field. Callers (pysimlin `check()`, the TS engine's `model.ts`) present the LTM auto-flip advisory as a warning rather than claiming the model is broken + - `simlin_project_get_errors()` - Collect all diagnostics. Project compilability is an INTRINSIC property, assessed with LTM OFF: the compile/VM-validation channel (`vm_error`) is always computed with the LTM overlay `Off` (`engine::db::LtmOverlay`), so an LTM-only failure (a synthetic fragment the compiler refuses) never masquerades as a project error on a model that simulates fine without LTM. LTM is an analysis overlay, not part of whether the model is a valid simulation. Separately, the LTM *diagnostics* (auto-flip-to-discovery advisory, synthetic-fragment compile failures, the GH #311 partial-equation warnings) accumulate via `model_all_diagnostics` -> `model_ltm_variables` independent of assembly success, so when the project's `ltm_requested` latch is set (any prior `simlin_sim_new` with `enable_ltm=true`) `get_errors` runs the `collect_all_diagnostics` harvest with the overlay `On` -- NOT a recompile that feeds `vm_error` -- so those advisories reach the caller (GH #466; `simlin-mcp-core` harvests the same way, GH #662). The overlay is an argument of every query rather than a salsa input, so the `Off` and `On` harvests are memoized side by side and a concurrent sim creation never observes a partial LTM state; salsa memoizes the LTM synthesis from the `sim_new` compile, so the `On` harvest revalidates rather than recomputes. A project that never requested LTM harvests `Off` and pays no LTM synthesis cost (preserving the scoping the engine's `test_ltm_disabled_gate_suppresses_auto_flip_warning` pins). The `ltm_requested` latch is a set-once `AtomicBool` on `SimlinProject`. Each `SimlinErrorDetail` carries a `severity` field (`SimlinErrorSeverity { Error, Warning }`) read off `errors::FormattedError.severity`, which the engine populates from the originating `Diagnostic.severity` (a `format_simulation_error` detail -- a compile/VM-validation failure with no diagnostic behind it -- is `Error` by construction). Threading it on the `FormattedError` rather than alongside it (GH #919) means the severity can neither be forgotten at a call site nor drift from the wording of `message`, which the engine now words from the same field. Callers (pysimlin `check()`, the TS engine's `model.ts`) present the LTM auto-flip advisory as a warning rather than claiming the model is broken ### Simulation lifecycle diff --git a/src/libsimlin/simlin.h b/src/libsimlin/simlin.h index 071d1d327..d74a4bb03 100644 --- a/src/libsimlin/simlin.h +++ b/src/libsimlin/simlin.h @@ -376,9 +376,7 @@ typedef struct { int64_t universe_loops; // Non-NULL when the model could not be compiled or analyzed for LTM at // all -- a malformed equation, an unresolved reference, or a hard - // compile failure such as the non-Euler-integration-with-a-stock-loop - // rejection (GH #486, which needs Euler stepping for its flow-to-stock - // link-score formula). When set, every OTHER field describes an + // compile failure. When set, every OTHER field describes an // analysis that never started: `loops`/`periods`/`partitions` are // empty, `loop_count`/`period_count`/`partition_count`/`retained_loops` // are `0`, `enumeration_complete` is `false`, and `universe_loops` is @@ -1657,10 +1655,11 @@ void simlin_project_render_png(SimlinProject *project, // `enable_ltm` requests Loops That Matter instrumentation. For an ordinary // model this produces a sim whose results carry the LTM link/loop-score // series. For a model containing a conveyor or queue stock, LTM is a -// documented degradation: the flow-to-stock link-score formula assumes plain -// INTEG under Euler, which neither special stock is, so the sim is created -// WITHOUT LTM instrumentation and `simlin_sim_get_ltm_mode` reports -// `Disabled`. `enable_ltm = true` is still honored as a request in that case: +// documented degradation: the flow-to-stock link score treats a stock's net +// flow as its rate of change (plain INTEG), which neither special stock is, +// so the sim is created WITHOUT LTM instrumentation and +// `simlin_sim_get_ltm_mode` reports `Disabled`. `enable_ltm = true` is still +// honored as a request in that case: // `simlin_project_get_errors` will surface a `ConveyorLtmDegraded` / // `QueueLtmDegraded` `Warning` naming the offending stock, so the caller learns // why scores are absent instead of the request being silently dropped. diff --git a/src/libsimlin/src/ffi.rs b/src/libsimlin/src/ffi.rs index 31d205b6a..60a61e9be 100644 --- a/src/libsimlin/src/ffi.rs +++ b/src/libsimlin/src/ffi.rs @@ -373,9 +373,7 @@ pub struct SimlinDiscoveryResult { pub universe_loops: i64, /// Non-NULL when the model could not be compiled or analyzed for LTM at /// all -- a malformed equation, an unresolved reference, or a hard - /// compile failure such as the non-Euler-integration-with-a-stock-loop - /// rejection (GH #486, which needs Euler stepping for its flow-to-stock - /// link-score formula). When set, every OTHER field describes an + /// compile failure. When set, every OTHER field describes an /// analysis that never started: `loops`/`periods`/`partitions` are /// empty, `loop_count`/`period_count`/`partition_count`/`retained_loops` /// are `0`, `enumeration_complete` is `false`, and `universe_loops` is diff --git a/src/libsimlin/src/lib.rs b/src/libsimlin/src/lib.rs index 2bb9d68ed..4864bdcd1 100644 --- a/src/libsimlin/src/lib.rs +++ b/src/libsimlin/src/lib.rs @@ -437,7 +437,7 @@ pub struct SimlinProject { /// decide whether to collect diagnostics under the LTM overlay, so the /// LTM diagnostic pipeline (the auto-flip-to-discovery warning, /// synthetic-fragment compile failures, the GH #311 partial-equation - /// warnings, and the GH #486 non-Euler Error) reaches the caller of a + /// warnings) reaches the caller of a /// project that simulated with LTM (GH #466), while a project that never /// requested LTM pays no LTM synthesis cost in `get_errors`. The overlay /// is an argument of the queries, not a flag on the project, so no diff --git a/src/libsimlin/src/project.rs b/src/libsimlin/src/project.rs index 8ca82d439..1d769af87 100644 --- a/src/libsimlin/src/project.rs +++ b/src/libsimlin/src/project.rs @@ -879,12 +879,11 @@ pub unsafe extern "C" fn simlin_project_get_errors( }; // Compilability is an INTRINSIC property of the project, assessed with LTM - // OFF. LTM is an analysis overlay (the flow-to-stock link-score formula - // assumes Euler integration, etc.), not part of whether the model is a + // OFF. LTM is an analysis overlay, not part of whether the model is a // valid, runnable simulation -- so the compile/VM-validation channel must - // never be computed under the overlay, or an LTM-only rejection (the GH - // #486 non-Euler hard `Err` from `assemble_simulation`) would masquerade - // as a project error on a model that simulates fine. `build_sim` + // never be computed under the overlay, or an LTM-only failure (a synthetic + // fragment the compiler refuses, say) would masquerade as a project error + // on a model that simulates fine. `build_sim` // additionally routes a conveyor/queue model through its special expansion // build path (also LTM-off), so a valid special-stock model is not // mis-reported as a project error by the ordinary path's NotExpanded guard. diff --git a/src/libsimlin/src/simulation.rs b/src/libsimlin/src/simulation.rs index 2fab86c46..725090285 100644 --- a/src/libsimlin/src/simulation.rs +++ b/src/libsimlin/src/simulation.rs @@ -35,10 +35,11 @@ fn is_internal_var(name: &str) -> bool { /// `enable_ltm` requests Loops That Matter instrumentation. For an ordinary /// model this produces a sim whose results carry the LTM link/loop-score /// series. For a model containing a conveyor or queue stock, LTM is a -/// documented degradation: the flow-to-stock link-score formula assumes plain -/// INTEG under Euler, which neither special stock is, so the sim is created -/// WITHOUT LTM instrumentation and `simlin_sim_get_ltm_mode` reports -/// `Disabled`. `enable_ltm = true` is still honored as a request in that case: +/// documented degradation: the flow-to-stock link score treats a stock's net +/// flow as its rate of change (plain INTEG), which neither special stock is, +/// so the sim is created WITHOUT LTM instrumentation and +/// `simlin_sim_get_ltm_mode` reports `Disabled`. `enable_ltm = true` is still +/// honored as a request in that case: /// `simlin_project_get_errors` will surface a `ConveyorLtmDegraded` / /// `QueueLtmDegraded` `Warning` naming the offending stock, so the caller learns /// why scores are absent instead of the request being silently dropped. @@ -127,9 +128,10 @@ pub unsafe extern "C" fn simlin_sim_new( // overlay), so the sim is created WITHOUT instrumentation and // `get_ltm_mode` reports Disabled. LTM over a conveyor/queue is a // documented degradation (docs/design/conveyors.md §9.6, queues.md - // §10.5): the flow-to-stock link-score formula assumes plain INTEG - // under Euler, and neither special stock is INTEG. `enable_ltm=true` is - // still HONORED as a request via the `ltm_requested` latch above, so + // §10.5): the flow-to-stock link score treats a stock's net flow as + // its rate of change (plain INTEG), and neither special stock is + // INTEG. `enable_ltm=true` is still HONORED as a request via the + // `ltm_requested` latch above, so // `simlin_project_get_errors` surfaces the ConveyorLtmDegraded / // QueueLtmDegraded warning that explains why scores are absent. let special = result.as_ref().map(|b| b.special).unwrap_or(false); diff --git a/src/libsimlin/tests/integration/analysis.rs b/src/libsimlin/tests/integration/analysis.rs index 6fb7f6e39..db7e1b98c 100644 --- a/src/libsimlin/tests/integration/analysis.rs +++ b/src/libsimlin/tests/integration/analysis.rs @@ -2649,10 +2649,10 @@ fn discover_loops_null_model_errors_without_panic() { } } -/// A model that CANNOT be analyzed for LTM at all (GH #486: a non-Euler -/// integration method with a stock in a feedback loop -- the flow-to-stock -/// link-score formula assumes Euler stepping) must surface `analysis_error` -/// non-NULL rather than returning a successful-looking empty/sampled result: +/// A model that CANNOT be analyzed for LTM at all (here a flow reading a +/// variable that does not exist, so nothing compiles) must surface +/// `analysis_error` non-NULL rather than returning a successful-looking +/// empty/sampled result: /// `simlin_analyze_discover_loops` itself still succeeds (this is a /// STRUCTURAL fact about the model, not an FFI error), but the discovery /// result reports "analysis never ran" -- `enumeration_complete == false` @@ -2664,9 +2664,8 @@ fn discover_loops_reports_analysis_error_when_ltm_never_ran() { unsafe { let test_project = TestProject::new("main") .with_sim_time(0.0, 10.0, 1.0) - .with_sim_method(simlin_engine::datamodel::SimMethod::RungeKutta4) .stock("population", "100", &["births"], &[], None) - .flow("births", "population * 0.02", None); + .flow("births", "population * nonexistent_variable", None); let datamodel_project = test_project.build_datamodel(); let project = engine_serde::serialize(&datamodel_project).unwrap(); let mut buf = Vec::new(); @@ -2692,14 +2691,16 @@ fn discover_loops_reports_analysis_error_when_ltm_never_ran() { let res = &*result; assert!( !res.analysis_error.is_null(), - "RK4 + a stock in a loop cannot be compiled for LTM analysis" + "an unresolved reference cannot be compiled for LTM analysis" ); let msg = CStr::from_ptr(res.analysis_error) .to_string_lossy() .into_owned(); + // The message names the variable that failed to compile, never the + // reference it could not resolve. assert!( - msg.contains("Euler"), - "analysis_error must reference the Euler assumption, got: {msg}" + msg.contains("births"), + "analysis_error must name the failing variable (births), got: {msg}" ); assert_eq!(res.loop_count, 0, "analysis never reached candidates"); assert_eq!(res.period_count, 0); diff --git a/src/libsimlin/tests/integration/errors.rs b/src/libsimlin/tests/integration/errors.rs index 6566f488f..dfc49e28e 100644 --- a/src/libsimlin/tests/integration/errors.rs +++ b/src/libsimlin/tests/integration/errors.rs @@ -709,9 +709,7 @@ fn test_get_errors_ltm_warning_survives_subsequent_non_ltm_sim() { } /// Build a single-stock feedback-loop project (population/births/birth_rate) -/// with the requested integration method. Under RK4 an LTM-enabled compile is -/// rejected (the flow-to-stock link-score formula assumes Euler -- GH #486), -/// but the project itself simulates fine without LTM. +/// with the requested integration method. fn build_feedback_loop_datamodel( name: &str, method: engine::datamodel::SimMethod, @@ -725,15 +723,14 @@ fn build_feedback_loop_datamodel( } /// GH #466 follow-up regression: a project that uses RK4 is intrinsically -/// valid -- it simulates fine without LTM. Creating an LTM sim on it (which the -/// GH #486 guard rejects at compile time) must NOT make `get_errors` report the -/// non-Euler rejection as a project error. LTM is an analysis overlay, not part -/// of the project's intrinsic compilability: `get_errors` assesses the -/// compile/vm_error channel with LTM OFF, and uses the latched re-enable only to -/// harvest the additional LTM diagnostics (which here are none, since the model -/// does not auto-flip and has no failing fragments). +/// valid, and creating an LTM sim on it must NOT make `get_errors` report +/// anything. LTM is an analysis overlay, not part of the project's intrinsic +/// compilability: `get_errors` assesses the compile/vm_error channel with LTM +/// OFF, and uses the latched re-enable only to harvest the additional LTM +/// diagnostics (which here are none, since the model does not auto-flip and +/// has no failing fragments). #[test] -fn test_get_errors_rk4_ltm_compile_failure_is_not_a_project_error() { +fn test_get_errors_after_an_rk4_ltm_sim_stays_clean() { let datamodel = build_feedback_loop_datamodel("rk4_ltm_overlay", engine::datamodel::SimMethod::RungeKutta4); let proj = open_project_from_datamodel(&datamodel); @@ -753,9 +750,8 @@ fn test_get_errors_rk4_ltm_compile_failure_is_not_a_project_error() { assert!(me.is_null()); assert!(!model.is_null()); - // Create an LTM-enabled sim; the GH #486 rejection rides the compile/VM - // path. The sim object is still created (the error defers to run time), - // and the project's ltm_requested latch is now set. + // Create an LTM-enabled sim; the project's ltm_requested latch is now + // set. let mut se: *mut SimlinError = ptr::null_mut(); let sim = simlin_sim_new(model, true, &mut se as *mut *mut SimlinError); if !se.is_null() { @@ -763,14 +759,12 @@ fn test_get_errors_rk4_ltm_compile_failure_is_not_a_project_error() { } // The regression: get_errors must still report the RK4 model as clean. - // The non-Euler rejection is an LTM-overlay concern, not a project error. let mut e1: *mut SimlinError = ptr::null_mut(); let post = simlin_project_get_errors(proj, &mut e1 as *mut *mut SimlinError); assert!(e1.is_null()); assert!( post.is_null(), - "RK4 model that simulates fine without LTM must not report errors \ - after a latched LTM sim; the non-Euler rejection is an analysis overlay" + "an RK4 model must not report errors after a latched LTM sim" ); if !sim.is_null() { diff --git a/src/libsimlin/tests/integration/simulation.rs b/src/libsimlin/tests/integration/simulation.rs index 2d882f092..7f4c213fb 100644 --- a/src/libsimlin/tests/integration/simulation.rs +++ b/src/libsimlin/tests/integration/simulation.rs @@ -1207,12 +1207,13 @@ fn test_ltm_enabled_sim() { } } -/// GH #486: enabling LTM on a model with a non-Euler integration method must -/// fail `simlin_sim_new` cleanly with a readable error referencing the Euler -/// assumption -- not silently produce mathematically-wrong link scores. The -/// same model with LTM disabled must still compile and simulate. +/// Enabling LTM on a model with a non-Euler integration method compiles and +/// runs: the scores are dt-step ratios reported at the saved steps (the +/// engine's `ltm_integration_method` tests pin them equal to Euler's for a +/// flow proportional to its stock, as this one is), the sim reports an LTM +/// mode, and the overlay leaves the model's own series untouched. #[test] -fn test_ltm_non_euler_sim_fails_cleanly() { +fn test_ltm_non_euler_sim_runs() { let datamodel = TestProject::new("ltm_rk4") .with_sim_time(0.0, 10.0, 1.0) .with_sim_method(engine::datamodel::SimMethod::RungeKutta4) @@ -1227,43 +1228,29 @@ fn test_ltm_non_euler_sim_fails_cleanly() { assert!(err.is_null()); assert!(!model.is_null()); - // LTM enabled on an RK4 model: the compile failure is deferred to run - // time (the established `simlin_sim_new` contract returns a non-null - // handle carrying the compile error and surfaces it on run), and the - // surfaced error references the Euler assumption. err = ptr::null_mut(); let sim_ltm = simlin_sim_new(model, true, &mut err as *mut *mut SimlinError); - assert!( - err.is_null(), - "simlin_sim_new defers the compile error to run" - ); + assert!(err.is_null(), "an LTM + RK4 sim compiles"); assert!(!sim_ltm.is_null()); + run_to_end(sim_ltm); err = ptr::null_mut(); - simlin_sim_run_to_end(sim_ltm, &mut err as *mut *mut SimlinError); - assert!( - !err.is_null(), - "running an LTM + RK4 sim must surface an error" - ); - let msg_ptr = simlin_error_get_message(err); - assert!(!msg_ptr.is_null(), "the error must carry a message"); - let msg = CStr::from_ptr(msg_ptr).to_str().unwrap(); - assert!( - msg.contains("Euler"), - "the error must reference the Euler assumption, got: {msg}" - ); - simlin_error_free(err); - simlin_sim_unref(sim_ltm); + let mode = simlin_sim_get_ltm_mode(sim_ltm, &mut err); + assert!(err.is_null()); + assert_eq!(mode, SimlinLtmMode::Exhaustive, "LTM ran under RK4"); - // The same model without LTM compiles and runs as before. err = ptr::null_mut(); - let sim_no_ltm = simlin_sim_new(model, false, &mut err as *mut *mut SimlinError); + let sim_plain = simlin_sim_new(model, false, &mut err as *mut *mut SimlinError); assert!(err.is_null(), "RK4 without LTM must compile"); - assert!(!sim_no_ltm.is_null()); - err = ptr::null_mut(); - simlin_sim_run_to_end(sim_no_ltm, &mut err as *mut *mut SimlinError); - assert!(err.is_null(), "RK4 without LTM must simulate"); + assert!(!sim_plain.is_null()); + run_to_end(sim_plain); + assert_eq!( + get_series_vec(sim_ltm, "population", 64), + get_series_vec(sim_plain, "population", 64), + "the overlay leaves the RK4 trajectory untouched" + ); - simlin_sim_unref(sim_no_ltm); + simlin_sim_unref(sim_ltm); + simlin_sim_unref(sim_plain); simlin_model_unref(model); simlin_project_unref(proj); } diff --git a/src/pysimlin/simlin/analysis.py b/src/pysimlin/simlin/analysis.py index 2a1afc74a..4f87868c2 100644 --- a/src/pysimlin/simlin/analysis.py +++ b/src/pysimlin/simlin/analysis.py @@ -539,10 +539,8 @@ class Analysis: analysis_error: str | None = None """Set when the model could not be compiled or analyzed for LTM AT ALL -- - a malformed equation, an unresolved reference, or a hard compile failure - such as picking a non-Euler integration method on a model with a stock in - a feedback loop (that link-score formula assumes Euler stepping). When - set, every other field describes an analysis that never STARTED: `loops`, + a malformed equation, an unresolved reference, or a hard compile failure. + When set, every other field describes an analysis that never STARTED: `loops`, `dominant_periods`, and `partitions` are empty, `retained_loops` is 0, `enumeration_complete` is False, and `universe_loops` is None -- the SAME shape a genuinely SAMPLED analysis that happened to find zero loops would diff --git a/src/pysimlin/simlin/model.py b/src/pysimlin/simlin/model.py index a4f83f300..2f316daab 100644 --- a/src/pysimlin/simlin/model.py +++ b/src/pysimlin/simlin/model.py @@ -816,9 +816,8 @@ def analyze(self, timeout: float | None = None) -> Analysis: * ``analysis_error`` is non-``None`` when the model could not be compiled/analyzed for LTM AT ALL (a malformed equation, an - unresolved reference, or a hard compile failure such as choosing a - non-Euler integration method on a model with a stock in a feedback - loop). When set, ``loops``/``dominant_periods``/``partitions`` are + unresolved reference, or a hard compile failure). When set, + ``loops``/``dominant_periods``/``partitions`` are empty, ``enumeration_complete`` is False, and ``universe_loops`` is ``None`` -- the SAME shape a genuinely sampled analysis with zero discovered loops would report, which is exactly why this is the diff --git a/src/pysimlin/tests/test_discovery.py b/src/pysimlin/tests/test_discovery.py index 1debd9389..6fd96262c 100644 --- a/src/pysimlin/tests/test_discovery.py +++ b/src/pysimlin/tests/test_discovery.py @@ -283,15 +283,16 @@ class TestAnalyzeAnalysisError: first.""" def test_unanalyzable_model_reports_analysis_error( - self, rk4_feedback_model: simlin.Model + self, unanalyzable_model: simlin.Model ) -> None: - # RK4 + a stock in a feedback loop cannot be compiled for LTM (the - # flow-to-stock link-score formula assumes Euler stepping -- GH #486). - # `Model.analyze()` itself must not raise: this is a structural fact - # about the model, surfaced as data on the result. - analysis = rk4_feedback_model.analyze() + # A flow reading a variable that does not exist cannot be compiled, so + # LTM analysis never starts. `Model.analyze()` itself must not raise: + # this is a fact about the model, surfaced as data on the result. + analysis = unanalyzable_model.analyze() assert analysis.analysis_error is not None - assert "Euler" in analysis.analysis_error + # The message names the variable that failed to compile, never the + # reference it could not resolve. + assert "births" in analysis.analysis_error, "names the failing variable (births)" assert analysis.loops == () assert analysis.dominant_periods == () assert analysis.partitions == () @@ -305,16 +306,17 @@ def test_healthy_model_has_no_analysis_error(self, logistic_model: simlin.Model) analysis = logistic_model.analyze() assert analysis.analysis_error is None + def test_rk4_model_is_analyzed(self, rk4_feedback_model: simlin.Model) -> None: + # LTM runs under RK4 (its scores dt-step ratios reported at the saved + # steps), so a feedback loop on an RK4 model is analyzed like any other. + analysis = rk4_feedback_model.analyze() + assert analysis.analysis_error is None + assert [loop.id for loop in analysis.loops] == ["r1"] + @pytest.fixture def rk4_feedback_model() -> simlin.Model: - """A single-stock feedback-loop model using RK4 integration. - - Simulates fine without LTM, but an LTM-enabled compile is rejected (the - flow-to-stock link-score formula assumes Euler -- GH #486), which is what - makes `Model.analyze()` report `analysis_error` instead of running - discovery at all. - """ + """A single-stock feedback-loop model using RK4 integration.""" from simlin.types import Aux, Flow, Stock project = simlin.Project.new(name="rk4_feedback", sim_start=0.0, sim_stop=10.0, dt=1.0) @@ -329,6 +331,28 @@ def rk4_feedback_model() -> simlin.Model: return model +@pytest.fixture +def unanalyzable_model(tmp_path: Path) -> simlin.Model: + """A feedback-loop model whose flow reads a variable that does not exist, + so nothing compiles and `Model.analyze()` reports `analysis_error` + instead of running discovery at all. Loaded from XMILE, because the edit + API refuses a patch that introduces a compile error.""" + xmile = b""" + +
testtest
+ 010
1
+ + + 100births + population * nonexistent_variable + + +
""" + path = tmp_path / "unanalyzable.stmx" + path.write_bytes(xmile) + return simlin.load(path) + + @pytest.fixture def large_horizon_model() -> simlin.Model: """Build a large-horizon balancing model in memory for truncation tests. diff --git a/src/pysimlin/tests/test_ltm_mode_and_links.py b/src/pysimlin/tests/test_ltm_mode_and_links.py index 7c2ebabbc..ff775fa13 100644 --- a/src/pysimlin/tests/test_ltm_mode_and_links.py +++ b/src/pysimlin/tests/test_ltm_mode_and_links.py @@ -12,6 +12,7 @@ from __future__ import annotations +import warnings from typing import TYPE_CHECKING import pytest @@ -566,11 +567,7 @@ def test_warning_surfaces_via_run_with_loop_analysis(self) -> None: def _rk4_feedback_model() -> simlin.Model: - """Build a single-stock feedback-loop model that uses RK4 integration. - - The model simulates fine without LTM, but an LTM-enabled compile is rejected - (the flow-to-stock link-score formula assumes Euler -- GH #486). - """ + """Build a single-stock feedback-loop model that uses RK4 integration.""" project = simlin.Project.new(name="rk4_feedback", sim_start=0.0, sim_stop=10.0, dt=1.0) project.set_sim_specs(sim_method="rk4") model = project.main_model @@ -585,39 +582,37 @@ def _rk4_feedback_model() -> simlin.Model: return model -class TestLtmOverlayNotProjectError: - """GH #466 follow-up: LTM is an analysis overlay, not part of the project's - intrinsic validity. A latched LTM run on an RK4 model (whose LTM compile the - GH #486 guard rejects) must NOT make get_errors()/check() report the model as - broken -- it simulates fine without LTM. This is the reviewer's exact repro. +class TestLtmUnderRk4: + """LTM runs under RK4 like under Euler: the scores are dt-step ratios + reported at the saved steps (the engine's `ltm_integration_method` tests + pin them equal to Euler's for a flow proportional to its stock, as this + one is), so `run()` neither warns nor falls back to a run without loop + analysis, and the project reports no errors afterwards (GH #466: LTM is an + analysis overlay, not part of the project's intrinsic validity). """ - def test_rk4_run_then_get_errors_is_clean(self) -> None: + def test_rk4_run_analyzes_loops_without_warning(self) -> None: model = _rk4_feedback_model() project = model.project assert project is not None - - # Baseline: a fresh RK4 model has no errors. assert project.get_errors() == [] - # run() defaults to analyze_loops=True -> LTM sim -> ltm_requested latch. - # The LTM compile fails under RK4, so run() degrades to a non-LTM run - # (emitting a warning); the run itself succeeds. - with pytest.warns(RuntimeWarning): + # run() defaults to analyze_loops=True -> an LTM sim; no fallback, no warning. + with warnings.catch_warnings(): + warnings.simplefilter("error") run = model.run() assert not run.results.empty + assert run.ltm_mode == "exhaustive" + assert [loop.id for loop in run.loops] == ["r1"], "the births loop is scored" - # The regression: the RK4 model still reports no errors. The non-Euler - # rejection is an LTM-overlay concern, not a project error. - assert project.get_errors() == [], ( - "a latched LTM run on an RK4 model that simulates fine must not make " - "get_errors report the non-Euler rejection as a project error" - ) + # The RK4 model still reports no errors after the latched LTM run. + assert project.get_errors() == [] def test_rk4_run_then_check_reports_no_errors(self) -> None: model = _rk4_feedback_model() - with pytest.warns(RuntimeWarning): + with warnings.catch_warnings(): + warnings.simplefilter("error") model.run() issues = model.check() diff --git a/src/simlin-engine/CLAUDE.md b/src/simlin-engine/CLAUDE.md index e4b96e4a0..0b5c85fb0 100644 --- a/src/simlin-engine/CLAUDE.md +++ b/src/simlin-engine/CLAUDE.md @@ -151,7 +151,7 @@ Rules: - A generated equation's dependency shapes come from `ltm_dep_shape` (`model_dep_shape` plus LTM helpers and LTM synthetic variables), for a synthetic variable's fragment and a generated helper's alike, and its `Expr2` lowering runs under those shapes. - The generated-text boundary is one place: the guard form a generator prints around its wrapped tree is parsed once, at the emitter (`LtmArm::new`); nothing upstream prints engine text to parse it again, and a helper carries no `Variable::eqn` (the generators read a target's axes off its `ast()`, keeping the `eqn` arm for user-authored source text). - No silent wrong numbers in LTM: a link-score fragment that fails to compile is reported by `model_ltm_fragment_diagnostics` as a `Warning`, never silently stubbed to zero, and a target whose lowering failed or a read the describers cannot name is declined with a warning (`ltm_unified_tests::a_target_whose_lowering_failed_is_declined_not_scored`, `ltm_array_agg::a_per_element_slots_bare_reducer_argument_is_declined_not_scored`). Every LTM warning is a fact on the derivation's value (`LtmVariablesResult::diagnostics`, the fragment pass's `Vec`, `ShapedLinkScore::Unscoreable`'s payload) that `db::model_all_diagnostics` emits once per model, so a declined edge, a mode flip or a failed fragment in a sub-model is reported once however many parents instantiate it. -- Under RK2/RK4 a flow-to-stock link score is meaningless, so `assemble_simulation` rejects that combination. Conveyor and queue models do not participate in LTM. Open LTM work is organized under the `ltm` label on GitHub (epic #488). +- LTM runs under Euler, RK2 and RK4 alike. A link score is a ratio of integration-step (dt) deltas reported at the saved steps: `PREVIOUS` reads the state the previous dt step ended in (the VM snapshots `prev_values` on every dt iteration, before the save/advance decision, so with save_step > dt a recorded score is the last dt step's ratio, not the saved interval's), and the VM and wasm both re-evaluate the flows at the restored end-of-step state before snapshotting it under RK, so the scores follow the method's own trajectory (`tests/integration/ltm_integration_method.rs`: a flow proportional to its stock scores identically under all three methods; a nonlinear one does not). Conveyor and queue models do not participate in LTM. Open LTM work is organized under the `ltm` label on GitHub (epic #488). ## Special stocks: conveyors and queues diff --git a/src/simlin-engine/src/analysis.rs b/src/simlin-engine/src/analysis.rs index c8bd91159..d80ef5882 100644 --- a/src/simlin-engine/src/analysis.rs +++ b/src/simlin-engine/src/analysis.rs @@ -61,10 +61,9 @@ pub struct ModelAnalysis { /// `Some(message)` when the model could **not** be compiled or simulated /// for LTM analysis, so the loop fields are empty *because of a failure* /// rather than because the model genuinely has no loops. This is the - /// actionable diagnostic the caller should surface: most notably the GH #486 - /// Euler guidance when a model with stocks in a loop selects a non-Euler - /// integrator with LTM enabled, but also any other compile error - /// (unresolved references, etc.) or a `Vm::new`/`run_to_end` failure. + /// actionable diagnostic the caller should surface: a compile error + /// (an unresolved reference, a malformed equation) or a + /// `Vm::new`/`run_to_end` failure. /// `None` means the pipeline either /// succeeded or degraded gracefully for a non-compile reason (a structural /// edge case), so an empty `loop_dominance` with `analysis_error == None` @@ -137,8 +136,8 @@ fn model_snapshot(project: &datamodel::Project, model_name: &str) -> Option (Some(result), None), Ok(None) => (None, None), @@ -355,11 +354,10 @@ struct PipelineResult { /// bailed). Empty loops, but NOT a compile error -- the caller reports "no /// loops", not "could not analyse". /// * `Err(message)` -- the model could not be *compiled* for LTM analysis (a -/// malformed equation, an unresolved reference, or the GH #486 non-Euler -/// hard-fail). The message is the actionable compile error the caller -/// surfaces via `ModelAnalysis::analysis_error` (GH #660); before this it was -/// swallowed by an `.ok()?` and the empty result looked identical to "no -/// loops". +/// malformed equation or an unresolved reference). The message is the +/// actionable compile error the caller surfaces via +/// `ModelAnalysis::analysis_error` (GH #660): never swallow it into the +/// `Ok(None)` arm, whose empty result is indistinguishable from "no loops". /// /// Uses the caller-provided salsa `SimlinDb` and `SourceProject` for /// both compilation/simulation and structural loop analysis (via @@ -398,12 +396,12 @@ fn run_ltm_pipeline( // degrades to empty loops (conveyor/queue + LTM is a documented // degradation) and reports no false error. // - // A compile failure here is still the actionable GH #660 case: the GH #486 - // non-Euler hard-fail (and any other compile/`Vm::new`/`run_to_end` error) - // carries a specific, user-facing message that must reach the caller rather - // than collapsing into an empty "no loops" result. Format it the same way - // regardless of origin: prefer the rich `details` (e.g. the Euler guidance), - // fall back to the code's Display when a bare error carries none. + // A compile failure here is the actionable GH #660 case: a compile, + // `Vm::new` or `run_to_end` error carries a specific, user-facing message + // that must reach the caller rather than collapsing into an empty "no + // loops" result. Format it the same way regardless of origin: prefer the + // rich `details` (the failing variable's name, say), fall back to the + // code's Display when a bare error carries none. let mut vm = crate::build_sim( db, source_project, @@ -1254,33 +1252,29 @@ mod tests { // ---- GH #660: a compile failure surfaces an actionable analysis_error ---- - /// An RK4 model with a stock in a loop cannot be compiled for LTM analysis - /// (the flow-to-stock link-score formula assumes Euler; GH #486). Before - /// GH #660 the compile `Err` was swallowed by `run_ltm_pipeline`'s `.ok()?` - /// and `analyze_model` returned an empty `ModelAnalysis` indistinguishable - /// from "this model genuinely has no loops". Now the Euler guidance must - /// reach the caller through `ModelAnalysis::analysis_error`, with the loop - /// fields empty. + /// A model that cannot be compiled -- here a flow reading a variable that + /// does not exist -- populates `ModelAnalysis::analysis_error` with the + /// compile failure. Before GH #660 the compile `Err` was swallowed by + /// `run_ltm_pipeline`'s `.ok()?` and `analyze_model` returned an empty + /// `ModelAnalysis` indistinguishable from "this model genuinely has no + /// loops"; now the failure reaches the caller, with the loop fields empty. #[test] - fn ac660_rk4_ltm_surfaces_euler_analysis_error() { - let project = crate::test_common::TestProject::new("main") - .with_sim_time(0.0, 10.0, 1.0) - .with_sim_method(datamodel::SimMethod::RungeKutta4) - .stock("population", "100", &["births"], &[], None) - .flow("births", "population * 0.02", None) - .build_datamodel(); + fn ac660_a_compile_failure_surfaces_analysis_error() { + let project = broken_project(); let (mut db, sp) = synced_db(&project); let analysis = analyze_model(&project, &mut db, sp, "main", None) - .expect("analyze_model should not return Err on a compilable-without-LTM model"); + .expect("analyze_model should not return Err on an uncompilable model"); let msg = analysis .analysis_error .as_deref() - .expect("RK4 + LTM compile failure must populate analysis_error"); + .expect("a compile failure must populate analysis_error"); + // The message names the variable that failed to compile, never the + // reference it could not resolve. assert!( - msg.contains("Euler"), - "the analysis_error must reference the Euler assumption, got: {msg}" + msg.contains("births"), + "the analysis_error must name the failing variable (births), got: {msg}" ); // The model snapshot is still intact and the loop fields stay empty. @@ -1290,11 +1284,11 @@ mod tests { ); assert!( analysis.loop_dominance.is_empty(), - "loop_dominance must be empty when the model can't be compiled for LTM" + "loop_dominance must be empty when the model can't be compiled" ); assert!( analysis.time.is_empty(), - "time must be empty when the model can't be compiled for LTM" + "time must be empty when the model can't be compiled" ); } @@ -1323,26 +1317,6 @@ mod tests { ); } - /// A structural/equation error that prevents compilation (not the #486 - /// Euler case) is still a compile failure, so it too must surface through - /// `analysis_error` rather than vanishing into an empty result. - #[test] - fn ac660_broken_model_surfaces_analysis_error() { - let project = broken_project(); - let (mut db, sp) = synced_db(&project); - let analysis = analyze_model(&project, &mut db, sp, "main", None) - .expect("analyze_model should not return Err"); - - assert!( - analysis.analysis_error.is_some(), - "a model that fails to compile must populate analysis_error" - ); - assert!( - analysis.loop_dominance.is_empty(), - "loop_dominance must be empty when compilation fails" - ); - } - /// A stockless passthrough sub-model exposing two outputs of opposing /// sign from one input: `pos = input_val * 0.02`, `neg = 0 - input_val`. /// `input_val` is a real input port; both `pos` and `neg` are read by the diff --git a/src/simlin-engine/src/db/assemble.rs b/src/simlin-engine/src/db/assemble.rs index 5f4114ad0..6abbd3f3d 100644 --- a/src/simlin-engine/src/db/assemble.rs +++ b/src/simlin-engine/src/db/assemble.rs @@ -1238,11 +1238,6 @@ fn collect_fragments<'db>( let mut ltm_tail: Vec = Vec::new(); if overlay == LtmOverlay::On { - // GH #486's non-Euler rejection is NOT enforced per module here: the - // integration method that actually runs is a single, main-model- - // governed property of the whole assembled simulation, so it is - // resolved once, against the main-governed method, in - // `assemble_simulation` (`ltm_non_euler_guard`). let ltm_vars = model_ltm_variables(db, model, project); for (index, ltm_var) in ltm_vars.vars.iter().enumerate() { let name = canonicalize(<m_var.name).into_owned(); @@ -1755,63 +1750,6 @@ pub fn assemble_simulation( // Each unique (model_name, input_set) pair gets its own CompiledModule. let module_instances = enumerate_module_instances(db, project, &main_model_name)?; - // GH #486: the LTM flow-to-stock link-score formula is only valid under - // Euler integration; under RK2/RK4 the scores are mathematically - // meaningless, and non-Euler IS honored at runtime (the VM and wasm - // backends both have distinct RK2/RK4 stepping loops), so it would - // silently produce plausible-but-wrong scores. The method that actually - // runs is the SINGLE main-model-governed `Specs.method` resolved below - // (root override else project specs); a submodel's own override is dead. - // Resolve it once and reject only when the assembled sim actually produces - // a flow-to-stock score against that main-governed method. - // - // GH #663 refinement: the old guard rejected on the mere presence of a - // stock in any instantiated model. That is a false positive for a loop-free - // model (an open-loop accumulation -- a constant inflow that never reads the - // stock back): in exhaustive mode LTM scores only the edges of detected - // feedback loops, so such a stock emits NO flow-to-stock score and the - // non-Euler method has nothing to corrupt. The refined precondition is the - // EXACT thing #486 protects against: "an instantiated model actually emits a - // flow-to-stock link score" (`model_emits_flow_to_stock_score`). - // - // The check iterates the instantiated set (root + every transitively- - // instantiated submodule) and asks each model's own - // `model_emits_flow_to_stock_score`, so a flow-to-stock score produced - // entirely inside a submodel instance is caught -- the assembly emits that - // submodel's LTM vars too. The root may be stock-free while a submodel under - // the main-governed method scores a flow-to-stock link, exactly the hazard a - // per-submodel-specs check missed. Unused model definitions sitting in the - // project are never instantiated, so they are irrelevant. - // - // Reading the emitted var set (rather than a loop-presence proxy) is what - // makes the refinement SOUND across modes: in discovery mode (user-forced or - // auto-flipped) and in any model with input ports, `model_ltm_variables` - // scores ALL causal edges, so it emits a flow-to-stock score for an - // open-loop stock's `flow → stock` edge even though no loop contains that - // stock. A "has any feedback loop" proxy would under-reject those models; the - // direct var-set test cannot. `model_ltm_variables` is already salsa-computed - // on this LTM-enabled assembly path, so this is a cache hit plus a linear - // scan. - // - // This rejection rides the `assemble_simulation` `Err`, so it reaches - // `simlin_sim_new`, `simlin_project_get_errors` (the `vm_error` channel), - // and the wasm backend -- the sim-compile path that, unlike - // `collect_all_diagnostics`, is what every runnable consumer goes through. - if overlay == LtmOverlay::On - && let Some(root_model) = project_models.get(main_model_canonical.as_ref()) - && let Some(method) = ltm::effective_non_euler_method(db, *root_model, project) - { - let any_flow_to_stock_score = module_instances.keys().any(|name| { - let canonical = canonicalize(name.as_str()); - project_models - .get(canonical.as_ref()) - .is_some_and(|sm| ltm::model_emits_flow_to_stock_score(db, *sm, project)) - }); - if any_flow_to_stock_score { - return Err(ltm::ltm_non_euler_diagnostic_message(method)); - } - } - // Sort module names: main first, then all others alphabetically let main_ident = Ident::::new(&main_model_name); let mut module_names: Vec<&Ident> = module_instances.keys().collect(); diff --git a/src/simlin-engine/src/db/ltm/mod.rs b/src/simlin-engine/src/db/ltm/mod.rs index 552f662e0..bce79439d 100644 --- a/src/simlin-engine/src/db/ltm/mod.rs +++ b/src/simlin-engine/src/db/ltm/mod.rs @@ -23,7 +23,6 @@ use std::collections::{HashMap, HashSet}; use crate::canonicalize; use crate::common::{Canonical, Ident}; -use crate::datamodel; use crate::ltm::strip_subscript; use super::{ @@ -187,103 +186,6 @@ pub(crate) fn endpoint_dimensions( .map(|shape| shape.dims) } -/// The single integration method the assembled simulation actually runs, when -/// it is NOT Euler -- the method the GH #486 guard keeps LTM off. -/// -/// A `CompiledSimulation` has exactly ONE `Specs.method`, resolved by -/// `assemble_simulation` from the MAIN (root) model's `model_sim_specs` -/// override else the project specs. A submodel's own `model_sim_specs` is -/// never consulted by the VM, so the GH #486 guard must resolve and apply that -/// single main-governed method -- NOT each model's own specs (which is the -/// blocker the per-model resolution had). `root_model` is the model named in -/// `assemble_simulation(.., main_model_name)`. Returns `None` for Euler (the -/// supported case), `Some(method)` otherwise. -pub(super) fn effective_non_euler_method( - db: &dyn Db, - root_model: SourceModel, - project: SourceProject, -) -> Option { - let method = match root_model.model_sim_specs(db) { - Some(specs) => specs.sim_method, - None => project.sim_specs(db).sim_method, - }; - match method { - datamodel::SimMethod::Euler => None, - other => Some(other), - } -} - -/// Whether `model_ltm_variables` emits at least one flow-to-stock link score -/// for this model -- the EXACT precondition the GH #486/#663 non-Euler guard -/// gates on. -/// -/// The guard keys on a flow-to-stock score rather than on "the model has a -/// stock" because the latter over-rejects a loop-free model (GH #663): in -/// exhaustive mode LTM scores only the edges of detected feedback loops, so an -/// open-loop stock (a constant inflow that never reads the stock back) emits -/// NO flow-to-stock score and there is nothing to guard. -/// -/// Crucially this is mode-aware where a loop-presence proxy is NOT: in -/// DISCOVERY mode (user-forced or auto-flipped) and in any model with input -/// ports, `model_ltm_variables` scores ALL causal edges, so it DOES emit a -/// flow-to-stock score for an open-loop stock's `flow → stock` edge even -/// though that stock is in no loop. Reading the emitted var set directly is -/// therefore the only sound test: it cannot under-reject a discovery-mode -/// model with a stock the way "has any loop" would. -/// -/// A causal edge into a stock can only originate from one of its flows -/// (`model_causal_edges` adds `flow → stock` edges and nothing else points at -/// a stock -- a stock's equation is its initial value, not `inflow-outflow`), -/// so "a link-score var whose `to` endpoint is a stock" is exactly "a -/// flow-to-stock score". `link_score_edge_endpoints` strips any element -/// subscript, so an arrayed stock's per-element score matches its base name. -/// -/// Bounded on the guard's path: `model_ltm_variables` is the same query the -/// Euler assembly path runs unconditionally (its cost is capped by the -/// auto-flip-to-discovery gate and the circuit budget), so the guard is not -/// adding an unbounded computation. On the rejection path the guard runs BEFORE -/// `assemble_module`, so it computes the query rather than hitting a cache; but -/// it pays that at most once per instantiated stock-bearing model, and any -/// later assembly of the same model gets the salsa cache hit. -pub(super) fn model_emits_flow_to_stock_score( - db: &dyn Db, - model: SourceModel, - project: SourceProject, -) -> bool { - let stocks = &crate::db::model_causal_edges(db, model, project).stocks; - if stocks.is_empty() { - return false; - } - let ltm = model_ltm_variables(db, model, project); - ltm.vars - .iter() - .any(|v| link_score_edge_endpoints(&v.name).is_some_and(|(_from, to)| stocks.contains(&to))) -} - -/// Human-readable name of a non-Euler integration method, used in the GH #486 -/// diagnostic so the message names the offending method concretely. -fn sim_method_display_name(method: datamodel::SimMethod) -> &'static str { - match method { - datamodel::SimMethod::Euler => "Euler", - datamodel::SimMethod::RungeKutta2 => "RK2 (2nd-order Runge-Kutta)", - datamodel::SimMethod::RungeKutta4 => "RK4 (4th-order Runge-Kutta)", - } -} - -/// The GH #486 rejection message: LTM was requested on a simulation whose -/// (main-model-governed) integration method is non-Euler. Returned as the -/// `assemble_simulation` `Err` so it reaches `simlin_sim_new`, -/// `simlin_project_get_errors` (the `vm_error` channel), and the wasm backend -/// (`WasmGenError::Unsupported`) -- the sim-compile path never produces -/// silently-wrong scores. -pub(super) fn ltm_non_euler_diagnostic_message(method: datamodel::SimMethod) -> String { - format!( - "LTM (Loops That Matter) analysis requires Euler integration, but this model uses \ - {}. Switch the integration method to Euler, or disable LTM analysis.", - sim_method_display_name(method), - ) -} - /// THE shared stateless predicate for the LTM early-return gates (GH #748, /// GH #749): a model is stateless when it has no parent-level stocks, no /// PREVIOUS-lagged dt dependency, AND no (transitively) stock-carrying @@ -1126,20 +1028,6 @@ pub fn model_ltm_variables( }; } - // GH #486's non-Euler rejection is NOT emitted here. The integration - // method that the VM actually honors is a single, main-model-governed - // property of the assembled simulation (`assemble_simulation`'s `Specs` - // selection: the root model's `model_sim_specs` override else the project - // specs -- a submodel's own override is dead), and this per-model query has - // no main-model context. Emitting per-model with each model's own specs is - // both wrong (a stock-free main that overrides to RK4 would be missed while - // its stock-bearing submodel falls back to the project's Euler; conversely - // a submodel that overrides to RK4 under a Euler main would be wrongly - // rejected) and duplicative. The check lives once in `assemble_simulation` - // against the resolved method; it surfaces through `compile_project_ - // incremental` to `simlin_sim_new`, `simlin_project_get_errors` (the - // `vm_error` channel), and the wasm backend. - // When the user explicitly requested discovery mode, honor it // directly. Otherwise auto-flip if either: // 1. The variable-level SCC exceeds `MAX_LTM_SCC_NODES`, or diff --git a/src/simlin-engine/src/db/ltm_unified_tests.rs b/src/simlin-engine/src/db/ltm_unified_tests.rs index 805f4a18d..e723c5771 100644 --- a/src/simlin-engine/src/db/ltm_unified_tests.rs +++ b/src/simlin-engine/src/db/ltm_unified_tests.rs @@ -3893,443 +3893,6 @@ fn test_ltm_var_name_index_matches_vars() { } } -// ── GH #486: LTM requires Euler integration ───────────────────────────── -// -// The engine keeps LTM on Euler-stepped runs and rejects the combination of -// the overlay with RK2/RK4 at sim-compile time, gated on a flow-to-stock -// score actually being emitted (`model_emits_flow_to_stock_score`). -// -// CRITICAL granularity (the multi-model probes below): the VM has a SINGLE -// global integration method, taken from the MAIN (root) model's -// `model_sim_specs` override else the project specs -- a submodel's own -// `model_sim_specs` is NEVER consulted by the VM (`assemble_simulation`'s -// `Specs` selection). So the guard must resolve and apply that one -// main-governed method, not each model's own specs. - -/// Build a single-stock feedback-loop project (population/births/birth_rate) -/// with the requested integration method, so the GH #486 guard can be -/// exercised under each non-Euler method. -#[cfg(test)] -fn feedback_loop_project_with_method(method: datamodel::SimMethod) -> datamodel::Project { - let mut project = feedback_loop_project(); - project.sim_specs.sim_method = method; - project -} - -/// The substring every GH #486 rejection must carry so users understand the -/// Euler assumption behind it. -#[cfg(test)] -const LTM_EULER_DIAGNOSTIC_MARKER: &str = "Euler"; - -/// Whether compiling `main` with LTM enabled fails with the GH #486 -/// Euler-assumption error. Returns `Some(details)` on the expected rejection, -/// `None` when the compile succeeds (so the same probe pins both directions). -#[cfg(test)] -fn ltm_euler_rejection(db: &SimlinDb, source_project: SourceProject) -> Option { - match compile_project_incremental(db, source_project, "main", crate::db::LtmOverlay::On) { - Ok(_) => None, - Err(err) => { - let details = err.details.unwrap_or_default(); - assert!( - details.contains(LTM_EULER_DIAGNOSTIC_MARKER), - "an LTM compile failure here must be the Euler rejection: {details}" - ); - Some(details) - } - } -} - -#[test] -fn test_ltm_with_rk4_fails_sim_compilation() { - // `simlin_sim_new` and `simlin_project_get_errors` both build the sim via - // `compile_project_incremental`; the Euler rejection rides its `Err`, so - // asserting against the compile `Err` covers the production diagnostics - // surface. (The per-model accumulator path can't carry this diagnostic - // correctly: `model_all_diagnostics` has no main-model concept, and the - // method is main-governed.) - - let db = SimlinDb::default(); - let project = feedback_loop_project_with_method(datamodel::SimMethod::RungeKutta4); - let source_project = sync_from_datamodel(&db, &project).project; - - let details = - ltm_euler_rejection(&db, source_project).expect("LTM + RK4 must fail sim compilation"); - assert!( - details.contains("RK4") || details.contains("Runge"), - "the error should name the offending integration method: {details}" - ); -} - -#[test] -fn test_ltm_with_rk2_fails_sim_compilation() { - let db = SimlinDb::default(); - let project = feedback_loop_project_with_method(datamodel::SimMethod::RungeKutta2); - let source_project = sync_from_datamodel(&db, &project).project; - - assert!( - ltm_euler_rejection(&db, source_project).is_some(), - "LTM + RK2 must fail sim compilation" - ); -} - -#[test] -fn test_ltm_with_euler_compiles_clean() { - let db = SimlinDb::default(); - let project = feedback_loop_project_with_method(datamodel::SimMethod::Euler); - let source_project = sync_from_datamodel(&db, &project).project; - - assert!( - ltm_euler_rejection(&db, source_project).is_none(), - "LTM + Euler is the supported combination and must compile" - ); -} - -#[test] -fn test_rk4_without_ltm_compiles_clean() { - // The guard fires ONLY when LTM is actually requested. The same non-Euler - // model with LTM disabled compiles and simulates as before. - let db = SimlinDb::default(); - let project = feedback_loop_project_with_method(datamodel::SimMethod::RungeKutta4); - let source_project = sync_from_datamodel(&db, &project).project; - // A plain compile: the LTM overlay stays off. - - assert!( - compile_project_incremental(&db, source_project, "main", crate::db::LtmOverlay::Off) - .is_ok(), - "RK4 model without LTM should compile" - ); -} - -/// A user-defined submodel with an internal stock + flow feedback loop and an -/// output port. Used by the multi-model probes: the submodel's own -/// `model_sim_specs` is set by the caller to assert the VM ignores it. -#[cfg(test)] -fn feedback_submodel(sim_specs: Option) -> datamodel::Model { - let mut m = x_model( - "sub", - vec![ - x_aux("input", "0", None), - x_stock("level", "50", &["adjust"], &[], None), - x_flow("adjust", "(input - level) / 5", None), - x_aux("output", "level", None), - ], - ); - m.sim_specs = sim_specs; - m -} - -/// GH #486 false-negative probe (the silent #486 hazard): the MAIN model has -/// no stock and overrides the integration method to RK4, while the submodel it -/// instantiates holds the stock+flow feedback loop and has no override. The VM -/// integrates EVERYTHING under the main-governed RK4, so the submodel's -/// flow-to-stock link score is produced and run -- and must be rejected. A -/// per-submodel check (the bug) would miss it: the submodel falls back to the -/// project's Euler. -#[test] -fn test_ltm_main_rk4_submodel_stock_is_rejected() { - let mut main = x_model( - "main", - vec![ - x_module("sub", &[("driver", "input")], None), - x_aux("driver", "1", None), - x_aux("observed", "sub.output", None), - ], - ); - // Main overrides the integration method to RK4; this is what the VM honors. - main.sim_specs = Some(datamodel::SimSpecs { - sim_method: datamodel::SimMethod::RungeKutta4, - ..feedback_loop_project().sim_specs - }); - - let mut project = feedback_loop_project(); - // `feedback_submodel` carries the internal `level <-> adjust` feedback loop. - project.models = vec![main, feedback_submodel(None)]; - // Project default stays Euler -- the trap the buggy per-submodel resolution - // falls into. - project.sim_specs.sim_method = datamodel::SimMethod::Euler; - - let db = SimlinDb::default(); - let source_project = sync_from_datamodel(&db, &project).project; - - let details = ltm_euler_rejection(&db, source_project) - .expect("main-RK4 + submodel-stock LTM must be rejected (the #486 hazard)"); - assert!( - details.contains("RK4") || details.contains("Runge"), - "the rejection must name the main-governed method: {details}" - ); -} - -/// GH #486 false-positive probe: the submodel overrides the integration method -/// to RK4, but the MAIN model (and project) are Euler. The VM runs Euler (main -/// governs); the submodel's RK4 override is dead. The scores are valid, so the -/// compile must SUCCEED. A per-submodel check (the bug) would wrongly reject. -#[test] -fn test_ltm_submodel_rk4_override_main_euler_is_accepted() { - // Main holds its own stock+flow feedback loop (so LTM produces scores) and - // also instantiates the RK4-overriding submodel. - let main = x_model( - "main", - vec![ - x_stock("population", "100", &["births"], &[], None), - x_flow("births", "population * birth_rate", None), - x_aux("birth_rate", "0.1", None), - x_module("sub", &[("births", "input")], None), - x_aux("observed", "sub.output", None), - ], - ); - // The submodel overrides to RK4 -- which the VM IGNORES (main governs). - let sub_specs = datamodel::SimSpecs { - sim_method: datamodel::SimMethod::RungeKutta4, - ..feedback_loop_project().sim_specs - }; - - let mut project = feedback_loop_project(); - project.models = vec![main, feedback_submodel(Some(sub_specs))]; - project.sim_specs.sim_method = datamodel::SimMethod::Euler; - - let db = SimlinDb::default(); - let source_project = sync_from_datamodel(&db, &project).project; - - let compiled = - compile_project_incremental(&db, source_project, "main", crate::db::LtmOverlay::On); - assert!( - compiled.is_ok(), - "main-Euler + submodel-RK4-override must compile (VM runs Euler): {:?}", - compiled.err() - ); - // And it actually simulates under Euler. - let mut vm = crate::vm::Vm::new(compiled.unwrap()).expect("VM creation should succeed"); - vm.run_to_end().expect("Euler sim should run to completion"); -} - -// ── GH #663: the non-Euler guard must not over-reject loop-free models ─── -// -// #486's guard rejected RK2/RK4 + LTM whenever ANY instantiated model had a -// stock. But LTM only emits a flow-to-stock link score for a stock that -// participates in a feedback loop; a stock in NO loop (an open-loop -// accumulation -- a constant inflow that never reads the stock back) produces -// zero flow-to-stock scores, so the non-Euler method has nothing to corrupt -// and the rejection is a false positive. The refined guard fires only when an -// instantiated model has a feedback loop (which, in a well-formed SD model, -// necessarily passes through a stock). - -/// An open-loop accumulation model: a stock fed by a CONSTANT inflow. The -/// inflow `rate` does not read the stock back, so there is no feedback loop -/// and LTM emits no flow-to-stock link score for `tank`. With the requested -/// integration method on the project specs. -#[cfg(test)] -fn open_loop_project_with_method(method: datamodel::SimMethod) -> datamodel::Project { - let mut project = feedback_loop_project(); - project.sim_specs.sim_method = method; - project.models = vec![x_model( - "main", - vec![ - // `tank` accumulates `rate`, but nothing reads `tank` -> no cycle. - x_stock("tank", "0", &["rate"], &[], None), - x_flow("rate", "fill_rate", None), - x_aux("fill_rate", "5", None), - ], - )]; - project -} - -#[test] -fn test_ltm_rk4_open_loop_stock_compiles() { - // The RED case for GH #663: a stock with no feedback loop under RK4 + LTM - // emits no flow-to-stock scores, so the non-Euler guard must NOT reject it. - - let db = SimlinDb::default(); - let project = open_loop_project_with_method(datamodel::SimMethod::RungeKutta4); - let source_project = sync_from_datamodel(&db, &project).project; - - let compiled = - compile_project_incremental(&db, source_project, "main", crate::db::LtmOverlay::On); - assert!( - compiled.is_ok(), - "RK4 + LTM on a loop-free (open-loop accumulation) model must compile: {:?}", - compiled.err() - ); - // And it simulates -- the stock genuinely accumulates under RK4. - let mut vm = crate::vm::Vm::new(compiled.unwrap()).expect("VM creation should succeed"); - vm.run_to_end().expect("RK4 sim should run to completion"); -} - -#[test] -fn test_ltm_rk2_open_loop_stock_compiles() { - let db = SimlinDb::default(); - let project = open_loop_project_with_method(datamodel::SimMethod::RungeKutta2); - let source_project = sync_from_datamodel(&db, &project).project; - - assert!( - compile_project_incremental(&db, source_project, "main", crate::db::LtmOverlay::On).is_ok(), - "RK2 + LTM on a loop-free model must compile" - ); -} - -#[test] -fn test_ltm_euler_open_loop_stock_compiles() { - // The Euler control: a loop-free model under Euler always compiled and - // still must -- this pins that the refined guard didn't change the Euler - // path or break loop-free models generally. - - let db = SimlinDb::default(); - let project = open_loop_project_with_method(datamodel::SimMethod::Euler); - let source_project = sync_from_datamodel(&db, &project).project; - - assert!( - compile_project_incremental(&db, source_project, "main", crate::db::LtmOverlay::On).is_ok(), - "Euler + LTM on a loop-free model must compile" - ); -} - -/// The soundness corner of GH #663: an open-loop stock has NO feedback loop, -/// but DISCOVERY mode (here user-forced) scores EVERY causal edge -- including -/// the open-loop `rate → tank` flow-to-stock edge. That flow-to-stock score is -/// exactly the Euler-only formula, so under RK4 the model must STILL be -/// rejected even though it has no loop. A guard that keyed on "has any feedback -/// loop" would wrongly accept this (the stock is in no loop); the guard must -/// key on "actually emits a flow-to-stock score". -#[test] -fn test_ltm_rk4_open_loop_stock_in_discovery_mode_is_rejected() { - use salsa::Setter; - - let mut db = SimlinDb::default(); - let project = open_loop_project_with_method(datamodel::SimMethod::RungeKutta4); - let source_project = sync_from_datamodel(&db, &project).project; - // Force discovery mode: now ALL edges get link scores, so the open-loop - // stock's flow-to-stock edge IS scored. - source_project.set_ltm_discovery_mode(&mut db).to(true); - - let details = ltm_euler_rejection(&db, source_project).expect( - "RK4 + LTM + forced-discovery on an open-loop stock must be rejected: discovery \ - scores the flow-to-stock edge", - ); - assert!( - details.contains("RK4") || details.contains("Runge"), - "the rejection must name the offending method: {details}" - ); -} - -/// The Euler control for the forced-discovery open-loop case: Euler is the -/// supported integration, so even though a flow-to-stock score IS emitted, the -/// combination compiles. Pins that the discovery-mode rejection is gated on the -/// non-Euler method, not on the mere emission of a flow-to-stock score. -#[test] -fn test_ltm_euler_open_loop_stock_in_discovery_mode_compiles() { - use salsa::Setter; - - let mut db = SimlinDb::default(); - let project = open_loop_project_with_method(datamodel::SimMethod::Euler); - let source_project = sync_from_datamodel(&db, &project).project; - source_project.set_ltm_discovery_mode(&mut db).to(true); - - assert!( - ltm_euler_rejection(&db, source_project).is_none(), - "Euler + LTM + forced-discovery on an open-loop stock must compile" - ); -} - -/// An ARRAYED open-loop stock with a DECORATED flow-to-stock `to` endpoint. A -/// scalar flow `rate` (a constant, never reading the stock back) broadcasts into -/// an arrayed stock `tank[D]`, so there is no feedback loop through any element. -/// Under RK4 + LTM with FORCED DISCOVERY, discovery scores every causal edge, so -/// a flow-to-stock score IS emitted for the `rate → tank` edge -- and because -/// the target is arrayed, the score is emitted PER ELEMENT with a decorated `to` -/// endpoint: `$⁚ltm⁚link_score⁚rate→tank[a]`, `…→tank[b]`, `…→tank[c]`. The bare -/// stock set is `{"tank"}`, so the guard must strip the `[a]` element decoration -/// off the `to` endpoint before matching. This pins the full decode chain on a -/// decorated `to` endpoint: `link_score_edge_endpoints` parses the name, -/// `strip_subscript` strips the element decoration, and `stocks.contains` -/// matches the bare stock base name. A guard that compared the decorated `to` -/// endpoint (`"tank[a]"`) against the bare stock set (`{"tank"}`) directly would -/// miss the arrayed score entirely and wrongly accept the non-Euler model. -#[test] -fn test_ltm_rk4_arrayed_open_loop_stock_in_discovery_mode_is_rejected() { - use salsa::Setter; - - let project = TestProject::new("arrayed_open_loop") - .with_sim_method(datamodel::SimMethod::RungeKutta4) - .named_dimension("D", &["a", "b", "c"]) - // `tank[D]` accumulates the scalar `rate` (broadcast across D), but - // `rate` is a constant that never reads `tank` -> no cycle. The arrayed - // target makes the flow-to-stock score per-element, so the `to` endpoint - // carries an `[a]`/`[b]`/`[c]` decoration the guard must strip. - .array_stock("tank[D]", "0", &["rate"], &[], None) - .flow("rate", "fill_rate", None) - .aux("fill_rate", "5", None) - .build_datamodel(); - - let mut db = SimlinDb::default(); - let source_project = sync_from_datamodel(&db, &project).project; - // Force discovery mode: ALL edges get link scores, so the arrayed open-loop - // stock's `rate → tank[d]` flow-to-stock edges ARE scored (per element). - source_project.set_ltm_discovery_mode(&mut db).to(true); - - let details = ltm_euler_rejection(&db, source_project).expect( - "RK4 + LTM + forced-discovery on an ARRAYED open-loop stock must be rejected: \ - discovery scores the per-element flow-to-stock edge with a decorated `to` endpoint", - ); - assert!( - details.contains("RK4") || details.contains("Runge"), - "the rejection must name the offending method: {details}" - ); -} - -/// A stock-free RK4 main that instantiates a submodel whose stock is in NO -/// internal loop, but which the parent reads through an input/output port. -/// A module with input ports scores ALL its causal edges (the pathway/composite -/// machinery needs them), so `model_ltm_variables` emits the submodel's -/// `fill → level` flow-to-stock score even though `level` is in no loop. That -/// score is the Euler-only formula, so under the main-governed RK4 the model -/// must be rejected. This pins that the refined guard reads the EMITTED scores -/// (a loop-presence proxy would wrongly accept this -- the submodel stock has -/// no loop). -#[test] -fn test_ltm_main_rk4_input_port_submodel_stock_is_rejected() { - let mut main = x_model( - "main", - vec![ - x_module("sub", &[("driver", "input")], None), - x_aux("driver", "1", None), - x_aux("observed", "sub.output", None), - ], - ); - main.sim_specs = Some(datamodel::SimSpecs { - sim_method: datamodel::SimMethod::RungeKutta4, - ..feedback_loop_project().sim_specs - }); - - // `level` accumulates `fill` (= the `input` port), never read back, so the - // submodel has no internal feedback loop -- but its input port forces - // all-edge scoring, which emits the `fill → level` flow-to-stock score. - let mut sub = x_model( - "sub", - vec![ - x_aux("input", "0", None), - x_stock("level", "0", &["fill"], &[], None), - x_flow("fill", "input", None), - x_aux("output", "level", None), - ], - ); - sub.sim_specs = None; - - let mut project = feedback_loop_project(); - project.models = vec![main, sub]; - project.sim_specs.sim_method = datamodel::SimMethod::Euler; - - let db = SimlinDb::default(); - let source_project = sync_from_datamodel(&db, &project).project; - - let details = ltm_euler_rejection(&db, source_project).expect( - "main-RK4 + input-port submodel-stock LTM must be rejected: an input-port submodel \ - scores its flow-to-stock edge even without a loop", - ); - assert!( - details.contains("RK4") || details.contains("Runge"), - "the rejection must name the main-governed method: {details}" - ); -} - // --------------------------------------------------------------------------- // GH #738: scalar-target agg over an array *expression* (`SUM(pop[*] * scale)`) // --------------------------------------------------------------------------- diff --git a/src/simlin-engine/src/test_common.rs b/src/simlin-engine/src/test_common.rs index 23990e3a1..f2b0dfdf2 100644 --- a/src/simlin-engine/src/test_common.rs +++ b/src/simlin-engine/src/test_common.rs @@ -191,6 +191,14 @@ impl TestProject { self } + /// Set the interval between recorded rows. `with_sim_time` resets it to + /// dt, so call this after it. + #[allow(dead_code)] + pub fn with_save_step(mut self, save_step: f64) -> Self { + self.sim_specs.save_step = Some(datamodel::Dt::Dt(save_step)); + self + } + /// Set the integration method pub fn with_sim_method(mut self, method: datamodel::SimMethod) -> Self { self.sim_specs.sim_method = method; diff --git a/src/simlin-engine/tests/integration/ltm_integration_method.rs b/src/simlin-engine/tests/integration/ltm_integration_method.rs new file mode 100644 index 000000000..d5725705d --- /dev/null +++ b/src/simlin-engine/tests/integration/ltm_integration_method.rs @@ -0,0 +1,243 @@ +// Copyright 2026 The Simlin Authors. All rights reserved. +// Use of this source code is governed by the Apache License, +// Version 2.0, that can be found in the LICENSE file. + +//! LTM under every integration method and every save step. +//! +//! A link score is a ratio of integration-step (dt) deltas, reported at the +//! saved steps: `PREVIOUS` reads the state the previous dt step ended in +//! (the VM snapshots `prev_values` on every dt iteration, before the +//! save/advance logic decides whether the row is recorded), and under +//! RK2/RK4 the VM re-evaluates the flows at the restored end-of-step state +//! before it snapshots that state (the RK stages' trial-point evaluations +//! are overwritten). This is the 2020 paper's form (Schoenberg, Davidsen and +//! Eberlein, section 6.1: the scores are "computed at each dt"); the paper +//! puts Runge-Kutta compatibility as "in principle", and the dt-step ratio +//! over the method's own trajectory is the form it takes here. +//! +//! Two pins: +//! +//! * A model whose flows are proportional to its stock scores identically +//! under Euler, RK2 and RK4 at every saved step -- the scores are ratios of +//! flow deltas, and those cancel the stock's step -- even though the stock +//! trajectories, and with them the flows and the net-flow auxes, differ. +//! That equality is a property of PROPORTIONAL flows, not of the method: +//! in general a score follows the method's trajectory, and the nonlinear +//! model below scores differently under Euler and RK4. +//! * A saved step that spans two dt steps reports the ratio over the LAST dt +//! step, the same number the dt-resolution run reports at that time; it +//! does not re-difference the flows over the saved interval. + +use simlin_engine::datamodel::{Project, SimMethod}; +use simlin_engine::ltm_post; +use simlin_engine::test_common::TestProject; + +use crate::test_helpers::{ltm_run, ltm_series}; + +/// The isolated reinforcing loop `s -> births -> s`. +fn isolated_loop(method: SimMethod) -> Project { + TestProject::new("iso_method") + .with_sim_time(0.0, 8.0, 1.0) + .with_sim_method(method) + .stock("s", "100", &["births"], &[], None) + .flow("births", "s * 0.1", None) + .build_datamodel() +} + +/// Births `0.1 * a` against deaths `a / 20`: the 67/33 split. +fn births_deaths(method: SimMethod) -> Project { + TestProject::new("bd_method") + .with_sim_time(0.0, 8.0, 1.0) + .with_sim_method(method) + .stock("a", "100", &["births"], &["deaths"], None) + .flow("births", "0.1 * a", None) + .flow("deaths", "a / 20", None) + .build_datamodel() +} + +/// Births `0.1 * s` against deaths `0.02 * s ^ 1.3`: the outflow is not +/// proportional to the stock, so a score depends on which two points of the +/// trajectory it differences, and the two flows nearly balance, so that +/// dependence is large. +fn nonlinear_deaths(method: SimMethod, dt: f64, save_step: f64) -> Project { + TestProject::new("nonlinear_method") + .with_sim_time(0.0, 4.0, dt) + .with_save_step(save_step) + .with_sim_method(method) + .stock("s", "100", &["births"], &["deaths"], None) + .flow("births", "0.1 * s", None) + .flow("deaths", "0.02 * s ^ 1.3", None) + .build_datamodel() +} + +/// `$⁚ltm⁚link_score⁚{from}→{to}`. +fn link_score(from: &str, to: &str) -> String { + format!("$\u{205A}ltm\u{205A}link_score\u{205A}{from}\u{2192}{to}") +} + +fn max_abs_diff(a: &[f64], b: &[f64]) -> f64 { + assert_eq!(a.len(), b.len(), "series lengths differ"); + a.iter() + .zip(b) + .map(|(x, y)| (x - y).abs()) + .fold(0.0, f64::max) +} + +/// One model of the parity check: its label, its stock, and its builder. +type ModelCase = (&'static str, &'static str, fn(SimMethod) -> Project); + +#[test] +fn every_integration_method_scores_a_proportional_flow_identically() { + let models: [ModelCase; 2] = [ + ("isolated loop", "s", isolated_loop), + ("births/deaths", "a", births_deaths), + ]; + for (label, stock, build) in models { + let euler = ltm_run(&build(SimMethod::Euler), false); + let euler_rel = ltm_post::compute_rel_loop_scores(&euler.results, &euler.loop_partitions); + let mut ltm_keys: Vec = euler + .results + .offsets + .keys() + .map(|k| k.as_str().to_string()) + .filter(|k| k.starts_with("$\u{205A}ltm\u{205A}")) + .collect(); + ltm_keys.sort(); + assert!( + ltm_keys.iter().any(|k| k.contains("loop_score")), + "{label}: the Euler run carries LTM scores" + ); + + for method in [SimMethod::RungeKutta2, SimMethod::RungeKutta4] { + let rk = ltm_run(&build(method), false); + // The trajectories genuinely differ, so the equal scores below + // are not the same numbers scored twice. + let stock_gap = max_abs_diff( + <m_series(&euler.results, stock, 0), + <m_series(&rk.results, stock, 0), + ); + assert!( + stock_gap > 1e-3, + "{label} under {method:?}: the stock trajectory should differ from Euler's, \ + max gap {stock_gap:e}" + ); + let mut rk_keys: Vec = rk + .results + .offsets + .keys() + .map(|k| k.as_str().to_string()) + .filter(|k| k.starts_with("$\u{205A}ltm\u{205A}")) + .collect(); + rk_keys.sort(); + assert_eq!( + rk_keys, ltm_keys, + "{label} under {method:?}: the same LTM variables" + ); + // The SCORES agree; the net-flow aux is a flow value and follows + // the trajectory like the flows do. + let mut scores_compared = 0; + for key in <m_keys { + let is_score = key.contains("\u{205A}link_score\u{205A}") + || key.contains("\u{205A}loop_score\u{205A}"); + if !is_score { + continue; + } + scores_compared += 1; + let diff = max_abs_diff( + <m_series(&euler.results, key, 0), + <m_series(&rk.results, key, 0), + ); + assert!( + diff < 1e-9, + "{label} under {method:?}: {key} differs from Euler by {diff:e}" + ); + } + assert!( + scores_compared >= 3, + "{label}: link and loop scores were compared" + ); + let rk_rel = ltm_post::compute_rel_loop_scores(&rk.results, &rk.loop_partitions); + assert_eq!(rk_rel.len(), euler_rel.len()); + for (id, series) in &euler_rel { + let diff = max_abs_diff(series, &rk_rel[id]); + assert!( + diff < 1e-9, + "{label} under {method:?}: relative score of {id} differs from Euler by {diff:e}" + ); + } + } + } +} + +/// The boundary of the parity above: with a nonlinear outflow the scores +/// follow the method's trajectory, so Euler and RK4 disagree. +#[test] +fn a_nonlinear_flow_scores_differently_under_euler_and_rk4() { + let euler = ltm_run(&nonlinear_deaths(SimMethod::Euler, 0.5, 0.5), false); + let rk4 = ltm_run(&nonlinear_deaths(SimMethod::RungeKutta4, 0.5, 0.5), false); + let key = link_score("deaths", "s"); + let euler_score = ltm_series(&euler.results, &key, 0); + let rk4_score = ltm_series(&rk4.results, &key, 0); + // t = 2 is index 4 at save_step 0.5. + let gap = (euler_score[4] - rk4_score[4]).abs(); + assert!( + gap > 1e-3, + "deaths -> s at t = 2: Euler {} vs RK4 {} should differ", + euler_score[4], + rk4_score[4] + ); +} + +/// With save_step = 2 * dt the recorded score at a saved time is the ratio +/// over the last dt step ending there -- the number the dt-resolution run +/// records at that time -- and not the ratio over the whole saved interval. +/// The two forms differ by more than a unit of score on this model, so a +/// change to when `prev_values` is snapshotted would fail here. +#[test] +fn a_saved_step_spanning_two_dt_steps_reports_the_last_dt_step_ratio() { + let key = link_score("deaths", "s"); + let net = "$\u{205A}ltm\u{205A}net\u{205A}s"; + // The Euler ratio at t = 2 over [1.5, 2] and, for the record, the same + // ratio under RK4: the two methods' trajectories differ, so the scores + // do too (the parity boundary). + let pinned_at_t2 = [ + (SimMethod::Euler, -22.744), + (SimMethod::RungeKutta4, -22.749), + ]; + for (method, pinned) in pinned_at_t2 { + let fine = ltm_run(&nonlinear_deaths(method, 0.5, 0.5), false); + let coarse = ltm_run(&nonlinear_deaths(method, 0.5, 1.0), false); + let fine_score = ltm_series(&fine.results, &key, 0); + let coarse_score = ltm_series(&coarse.results, &key, 0); + let fine_deaths = ltm_series(&fine.results, "deaths", 0); + let fine_net = ltm_series(&fine.results, net, 0); + assert_eq!(fine_score.len(), 9, "{method:?}: 0..=4 at save_step 0.5"); + assert_eq!(coarse_score.len(), 5, "{method:?}: 0..=4 at save_step 1"); + + // Same dt, so the same trajectory: the coarse rows are the fine rows + // at the saved times, scores included. + for t in 1..=4usize { + let fine_at_t = fine_score[2 * t]; + let coarse_at_t = coarse_score[t]; + assert!( + (fine_at_t - coarse_at_t).abs() < 1e-12, + "{method:?} at t = {t}: coarse {coarse_at_t} vs fine {fine_at_t}" + ); + // The ratio over the whole saved interval [t - 1, t] is a + // different number. + let d_deaths = fine_deaths[2 * t] - fine_deaths[2 * t - 2]; + let d_net = fine_net[2 * t] - fine_net[2 * t - 2]; + let over_saved_interval = -(d_deaths / d_net).abs(); + assert!( + (over_saved_interval - coarse_at_t).abs() > 0.1, + "{method:?} at t = {t}: the saved-interval ratio {over_saved_interval} \ + should be distinguishable from the recorded {coarse_at_t}" + ); + } + assert!( + (coarse_score[2] - pinned).abs() < 5e-4, + "{method:?} at t = 2: deaths -> s recorded {} (pinned {pinned})", + coarse_score[2] + ); + } +} diff --git a/src/simlin-engine/tests/integration/main.rs b/src/simlin-engine/tests/integration/main.rs index c27a66fe0..162babd16 100644 --- a/src/simlin-engine/tests/integration/main.rs +++ b/src/simlin-engine/tests/integration/main.rs @@ -34,6 +34,7 @@ mod ltm_array_agg; mod ltm_discovery_large_models; mod ltm_dt_invariance; mod ltm_flow_to_stock; +mod ltm_integration_method; // Compares xmutil-based MDL parsing against the native Rust parser, so it // needs the optional xmutil C++ converter compiled in. #[cfg(feature = "xmutil")] diff --git a/src/simlin-engine/tests/integration/simulate.rs b/src/simlin-engine/tests/integration/simulate.rs index 04358ddac..11c40ef5d 100644 --- a/src/simlin-engine/tests/integration/simulate.rs +++ b/src/simlin-engine/tests/integration/simulate.rs @@ -5308,42 +5308,36 @@ fn mark2_mdl_compiles_after_protobuf_roundtrip() { } /// The browser's model.run() defaults analyzeLtm=true, so simNew is called -/// with enable_ltm=true. mark2.mdl declares RK4 integration, and the LTM -/// flow-to-stock link-score formula is only valid under Euler (GH #486), so -/// enabling LTM on this real-world model must now be rejected with the -/// Euler-assumption error rather than silently producing wrong scores. The -/// same model still compiles and runs with LTM disabled. (SMOOTH/DELAY-in-a- -/// feedback-loop compiling WITH LTM is covered by the Euler-based LTM tests in -/// `db::ltm_module_tests`.) -#[test] -fn mark2_mdl_rejects_ltm_under_rk4() { +/// with enable_ltm=true. mark2.mdl declares RK4 integration: the LTM overlay +/// compiles and runs on this real-world model, its scores the dt-step ratios +/// reported at the saved steps (`ltm_integration_method.rs`), and the overlay +/// leaves the simulation itself untouched -- every model variable's series is +/// the same with and without it. +#[test] +fn mark2_mdl_simulates_with_ltm_under_rk4() { let contents = std::fs::read_to_string("../../test/bobby/vdf/econ/mark2.mdl").expect("read mark2.mdl"); let project = open_vensim(&contents).expect("parse mark2.mdl"); let mut db = SimlinDb::default(); let sync = sync_from_datamodel_incremental(&mut db, &project, None); - // The LTM overlay on an RK4 model: the compile is rejected with the Euler - // assumption explained. - let err = - compile_project_incremental(&db, sync.project, "main", simlin_engine::db::LtmOverlay::On) - .expect_err("LTM + RK4 must be rejected"); - let details = err.details.unwrap_or_default(); + let run = |overlay: simlin_engine::db::LtmOverlay| -> Results { + let compiled = compile_project_incremental(&db, sync.project, "main", overlay) + .unwrap_or_else(|e| panic!("mark2.mdl should compile with overlay {overlay:?}: {e}")); + let mut vm = Vm::new(compiled).expect("VM creation should succeed"); + vm.run_to_end().expect("VM should run to completion"); + vm.into_results() + }; + let plain = run(simlin_engine::db::LtmOverlay::Off); + let with_ltm = run(simlin_engine::db::LtmOverlay::On); assert!( - details.contains("Euler"), - "the rejection must reference the Euler assumption: {details}" + with_ltm.offsets.keys().any(|k| k + .as_str() + .starts_with("$\u{205A}ltm\u{205A}loop_score\u{205A}")), + "the LTM run carries loop scores" ); - - // Without the overlay, the same RK4 model compiles and simulates as before. - let compiled = compile_project_incremental( - &db, - sync.project, - "main", - simlin_engine::db::LtmOverlay::Off, - ) - .expect("mark2.mdl should compile without LTM"); - let mut vm = Vm::new(compiled).expect("VM creation should succeed"); - vm.run_to_end().expect("VM should run to completion"); + // Every model variable's series is the LTM-free run's, exactly. + ensure_results_excluding(&plain, &with_ltm, &[]); } // =========================================================================== diff --git a/src/simlin-engine/tests/integration/simulate_ltm_wasm.rs b/src/simlin-engine/tests/integration/simulate_ltm_wasm.rs index aee3fcf95..e763b0440 100644 --- a/src/simlin-engine/tests/integration/simulate_ltm_wasm.rs +++ b/src/simlin-engine/tests/integration/simulate_ltm_wasm.rs @@ -474,32 +474,33 @@ fn unsupported_ltm_model_returns_wasmgen_error() { ); } -/// GH #486: the wasm LTM compile path shares the salsa pipeline, so a non-Euler -/// integration method with LTM enabled surfaces the same Euler-assumption guard -/// as a clean `WasmGenError` -- no silently-wrong LTM slab. +/// LTM under RK4 lowers to wasm and agrees with the VM: both backends +/// re-evaluate the flows at the restored end-of-step state before +/// snapshotting it, so the LTM columns are the dt-step ratios reported at the +/// saved steps on both (`tests/integration/ltm_integration_method.rs` pins +/// those scores equal to Euler's for this proportional-flow model; this pins +/// the two backends equal to each other under RK4). #[test] -fn ltm_non_euler_returns_wasmgen_error() { +fn series_rk4_births_deaths_matches_vm() { use simlin_engine::test_common::TestProject; let project = TestProject::new("ltm_rk4_wasm") - .with_sim_time(0.0, 10.0, 1.0) + .with_sim_time(0.0, 8.0, 1.0) .with_sim_method(datamodel::SimMethod::RungeKutta4) - .stock("population", "100", &["births"], &[], None) - .flow("births", "population * 0.02", None) + .stock("a", "100", &["births"], &["deaths"], None) + .flow("births", "0.1 * a", None) + .flow("deaths", "a / 20", None) .build_datamodel(); - - match compile_datamodel_to_artifact(&project, "main", true, false) { - Ok(_) => panic!("wasm compile of an RK4 LTM model must fail, not succeed"), - Err(WasmGenError::Unsupported(msg)) => assert!( - msg.contains("Euler"), - "the error must reference the Euler assumption, got: {msg}" - ), - } - - // The same model lowers cleanly with LTM disabled (the guard fires only - // when LTM is requested). - compile_datamodel_to_artifact(&project, "main", false, false) - .expect("RK4 model without LTM must lower to wasm"); + let vm = vm_results_for_ltm(&project, "main"); + let wasm = wasm_results_for_ltm(&project, "main") + .unwrap_or_else(|msg| panic!("an RK4 LTM model should lower to wasm: {msg}")); + assert!( + wasm.offsets + .keys() + .any(|k| k.as_str().starts_with(LTM_LOOP_SCORE_PREFIX)), + "the RK4 wasm run carries loop scores" + ); + assert_ltm_slabs_match(&vm, &wasm); } // --------------------------------------------------------------------------- diff --git a/src/simlin-mcp-core/CLAUDE.md b/src/simlin-mcp-core/CLAUDE.md index 01684c3d8..fcf183951 100644 --- a/src/simlin-mcp-core/CLAUDE.md +++ b/src/simlin-mcp-core/CLAUDE.md @@ -19,7 +19,7 @@ The library is generic over a concrete `A: ProjectAccess` (not `dyn`) so rmcp's - `src/types.rs` -- Wire-format types preserved verbatim from the pre-rmcp binary (`SourceFormat`, `LoopDominanceSummary`, `DominantPeriodOutput`, `ErrorOutput`) plus `build_empty_project` and `build_empty_project_with_specs` shared with the new-project HTTP route in `simlin-serve` for byte-identical create output, and the MDL export wire helpers: `mdl_export_warnings_to_outputs` / `mdl_export_error_to_output` are the one place the MDL writer's warnings and hard errors become wire `ErrorOutput`s (`generic` code, `model` kind, `MDL export:` message prefix, scoped to the single non-macro model) -- both `FileSystemAccess` and `simlin-serve`'s writer go through them so pysimlin, stdio MCP, and serve report the same thing -- and `preflight_export` dry-runs the writer for a `SourceFormat` (warnings a save would report, or the `Validation` error it would fail with). - `src/open.rs` -- Format-detection + parsing helpers (`format_for_extension`, `open_project`, `resolve_model_name`). I/O-free: callers pass already-loaded bytes. `format_for_extension` is the single extension dispatcher for the MCP surface: `open_project` picks its parser with it and `create_model` picks the writer for a fresh file with it, so a path's extension always names the format on disk (`.stmx`/`.xmile`/`.xml` XMILE, `.mdl` Vensim, everything else content-detected JSON). `ensure_variable_uids` is private: `open_project` is the only caller, and running it is part of what opening a project means rather than a step a caller may skip. - `src/fs_access.rs` -- `FileSystemAccess`, the stateless filesystem `ProjectAccess` impl. Used by the `simlin-mcp` binary (re-exported there as `simlin_mcp::access::FileSystemAccess`) AND by this crate's integration suites (`test_support::TestFileSystemAccess` is a type alias for it). It lives here rather than in the binary so the tests exercise the shipping impl -- a hand-maintained near-copy drifts at exactly the points where this file is non-trivial (the MDL lossiness-warning channel and the SD-AI `relationships` regeneration on save), so a test saving through a copy proves something about a simpler function than the one that ships. Every `SourceFormat` is written back in place in its own format; `Mdl` goes through `simlin_engine::to_mdl_with_warnings`, whose hard errors (more than one non-macro model, an ordinary Module variable) fail the save and whose lossiness warnings ride `SaveOutcome::warnings`. -- `src/tools/` -- The three reused tools (`read_model.rs`, `edit_model.rs`, `create_model.rs`) as async free functions taking `&impl ProjectAccess`. Exposed types use `#[serde(rename_all = "camelCase")]`; the curated *input* types deliberately exclude engine-internal fields (`uid`, `compat`, `aiState`), while both tool outputs embed the full engine `json::Model`, which serializes `uid`/`compat` when populated. `read_model`/`edit_model` both unconditionally run LTM loop analysis (`analysis::analyze_model`), so they (1) carry an `analysisError` field that surfaces the actionable compile error when a model can't be compiled for LTM -- most notably the GH #486 Euler guidance -- instead of returning a silent empty `loopDominance` (GH #660), and (2) collect their `collect_all_diagnostics` passes with the LTM overlay on (`simlin_engine::db::LtmOverlay::On`, the same harvest libsimlin's `simlin_project_get_errors` runs for a project that requested LTM, GH #466), surfacing the LTM auto-flip-to-discovery advisory and synthetic-fragment compile-failure warnings in a model-scoped `warnings` field that a collection under `Off` never carries (GH #662). `edit_model` enables LTM on both its pre- and post-edit diagnostic passes so the new-error gate compares like-with-like (the LTM advisories are Warnings, and the #486 rejection rides the assemble path, so neither affects the Error-severity gate). Both outputs also carry the discovery-completeness triple `enumerationComplete` / `retainedLoops` / `universeLoops` from `analysis::ModelAnalysis`, plus `truncated` (elided when false, like `aggRecoveryTruncated`: candidate generation stopped early, which the fallback's candidate bound can cause even without a wall-clock budget). `enumerationComplete` is the one result-level flag that is ALWAYS serialized: `aggRecoveryTruncated` elides its `false` because the interesting value there is `true`, whereas here the interesting value IS `false` (a sampled analysis), and a client that cannot see the field reads a sample as exhaustive -- the exact mistake the flag exists to prevent. `retainedLoops` is likewise unconditional (`0` is a real statement, so no absence would mean anything a `0` does not); `universeLoops` is the only optional one, elided when `enumerationComplete` is false because a sample has no universe to report, which is a different claim from a universe of zero. `edit_model` pre-flights the writer (`preflight_export`) before touching the store -- a project the format cannot hold is a `Validation` error and never lands, so a registry-backed store is not left holding a merged doc it cannot persist -- and appends the writer's lossiness warnings to that same `warnings` field: the store's `SaveOutcome::warnings` after a real write, the pre-flight's on a `dryRun`, so an agent can preview what a save would degrade. +- `src/tools/` -- The three reused tools (`read_model.rs`, `edit_model.rs`, `create_model.rs`) as async free functions taking `&impl ProjectAccess`. Exposed types use `#[serde(rename_all = "camelCase")]`; the curated *input* types deliberately exclude engine-internal fields (`uid`, `compat`, `aiState`), while both tool outputs embed the full engine `json::Model`, which serializes `uid`/`compat` when populated. `read_model`/`edit_model` both unconditionally run LTM loop analysis (`analysis::analyze_model`), so they (1) carry an `analysisError` field that surfaces the actionable compile error (naming the variable that failed to compile) when a model can't be compiled for LTM instead of returning a silent empty `loopDominance` (GH #660), and (2) collect their `collect_all_diagnostics` passes with the LTM overlay on (`simlin_engine::db::LtmOverlay::On`, the same harvest libsimlin's `simlin_project_get_errors` runs for a project that requested LTM, GH #466), surfacing the LTM auto-flip-to-discovery advisory and synthetic-fragment compile-failure warnings in a model-scoped `warnings` field that a collection under `Off` never carries (GH #662). `edit_model` enables LTM on both its pre- and post-edit diagnostic passes so the new-error gate compares like-with-like (the LTM advisories are Warnings, so they do not affect the Error-severity gate). Both outputs also carry the discovery-completeness triple `enumerationComplete` / `retainedLoops` / `universeLoops` from `analysis::ModelAnalysis`, plus `truncated` (elided when false, like `aggRecoveryTruncated`: candidate generation stopped early, which the fallback's candidate bound can cause even without a wall-clock budget). `enumerationComplete` is the one result-level flag that is ALWAYS serialized: `aggRecoveryTruncated` elides its `false` because the interesting value there is `true`, whereas here the interesting value IS `false` (a sampled analysis), and a client that cannot see the field reads a sample as exhaustive -- the exact mistake the flag exists to prevent. `retainedLoops` is likewise unconditional (`0` is a real statement, so no absence would mean anything a `0` does not); `universeLoops` is the only optional one, elided when `enumerationComplete` is false because a sample has no universe to report, which is a different claim from a universe of zero. `edit_model` pre-flights the writer (`preflight_export`) before touching the store -- a project the format cannot hold is a `Validation` error and never lands, so a registry-backed store is not left holding a merged doc it cannot persist -- and appends the writer's lossiness warnings to that same `warnings` field: the store's `SaveOutcome::warnings` after a real write, the pre-flight's on a `dryRun`, so an agent can preview what a save would degrade. - `src/server.rs` -- `SimlinMcpServer` rmcp `ServerHandler` impl with the three `#[tool]` macros plus `list_resources` and `read_resource`. `version` is plumbed in by the binary so `serverInfo.version` reflects the binary's `CARGO_PKG_VERSION`, not the library's. - `src/test_support.rs` -- `#[doc(hidden)]` integration-test fixtures, gated behind the `test-support` feature so they are not compiled into shipped binaries (the crate takes a self dev-dependency enabling the feature so `tests/` still resolves them). `TestFileSystemAccess` is a type alias for the production `fs_access::FileSystemAccess`, never a second implementation (see its rustdoc for why); `chain_scc_project_json` builds the oversized-SCC model the LTM auto-flip warning tests need. @@ -28,7 +28,7 @@ The library is generic over a concrete `A: ProjectAccess` (not `dyn`) so rmcp's The consumers of this tool surface are agents, and a tool result is context for the consuming agent's next decision -- design every result for that role. A capability is only usable when the agent can traverse the complete loop: discover the tool, recognize when it applies, invoke it correctly, interpret the result, recover from failure, and verify the effect. Concretely: - **Quiet success, bounded results.** A successful call returns the data needed for the next decision, not a transcript of the work. Large payloads (full model JSON, long loop lists) should be curated -- the edit/create *input* types deliberately exclude engine-internal fields (`uid`, `compat`, `aiState`) so agents never have to supply them. The outputs currently embed the full engine `json::Model` (which serializes `uid`/`compat` when populated), so a new output payload does not inherit that curation for free -- apply it deliberately. -- **Errors name the violated invariant and the repair.** "edit introduces compilation errors" plus the structured `ErrorOutput` list tells the agent what rule was broken and where to fix it; a bare failure would leave the agent guessing at hidden state. New error paths should meet that bar (the GH #486 Euler guidance in `analysisError` is the model to follow: it converts a silent empty result into an actionable explanation). +- **Errors name the violated invariant and the repair.** "edit introduces compilation errors" plus the structured `ErrorOutput` list tells the agent what rule was broken and where to fix it; a bare failure would leave the agent guessing at hidden state. New error paths should meet that bar (`analysisError` is the model to follow: it converts a silent empty result into an actionable explanation naming the variable that could not be compiled). - **Advisory context rides the result.** Warnings and analysis errors surface conditions the agent cannot otherwise observe (LTM auto-flip, synthetic-fragment failures). When adding a tool, ask what the agent would need to relay through a human today and put that in the result instead. - **Tool descriptions advertise what and why.** The schema-visible name and description are how an agent selects the tool at the moment of need; detail belongs in the result, not the catalog. diff --git a/src/simlin-mcp-core/src/tools/edit_model.rs b/src/simlin-mcp-core/src/tools/edit_model.rs index 480e9a013..97772792d 100644 --- a/src/simlin-mcp-core/src/tools/edit_model.rs +++ b/src/simlin-mcp-core/src/tools/edit_model.rs @@ -299,10 +299,9 @@ pub async fn edit_model( // diagnostic passes collect under the LTM overlay (GH #662), as libsimlin // does for a project that simulated with LTM (GH #466). Collecting under // the overlay on the pre- AND post-edit passes keeps the new-error delta - // symmetric: the LTM-only diagnostics (advisory Warnings; the GH #486 - // non-Euler rejection rides the assemble path, not this accumulator) are - // computed the same way on both sides, so they can never spuriously read - // as a "new error". + // symmetric: the LTM-only diagnostics (advisory Warnings) are computed + // the same way on both sides, so they can never spuriously read as a + // "new error". let pre_edit_error_keys: std::collections::HashSet<_> = { let pre_db = simlin_engine::db::SimlinDb::default(); let pre_sync = simlin_engine::db::sync_from_datamodel(&pre_db, &project); diff --git a/src/simlin-mcp-core/src/tools/read_model.rs b/src/simlin-mcp-core/src/tools/read_model.rs index 16ff19c44..9ad58aa54 100644 --- a/src/simlin-mcp-core/src/tools/read_model.rs +++ b/src/simlin-mcp-core/src/tools/read_model.rs @@ -119,9 +119,8 @@ pub struct ReadModelOutput { pub warnings: Vec, /// `Some(message)` when the model could not be compiled for LTM loop /// analysis, so `loop_dominance` is empty *because of a failure*, not - /// because the model has no loops (GH #660). The message is actionable -- - /// most notably the GH #486 Euler guidance for a non-Euler model with LTM - /// enabled. Elided from the wire shape when `None`. + /// because the model has no loops (GH #660). Elided from the wire shape + /// when `None`. #[serde(skip_serializing_if = "Option::is_none")] pub analysis_error: Option, } diff --git a/src/simlin-mcp-core/tests/integration/read_model_e2e.rs b/src/simlin-mcp-core/tests/integration/read_model_e2e.rs index c7735207d..c0b48f0a9 100644 --- a/src/simlin-mcp-core/tests/integration/read_model_e2e.rs +++ b/src/simlin-mcp-core/tests/integration/read_model_e2e.rs @@ -245,14 +245,11 @@ async fn read_model_surfaces_cycle_partitions() { ); } -/// GH #660: an RK4 model with a stock in a loop cannot be compiled for LTM -/// analysis (the flow-to-stock link-score formula assumes Euler; GH #486). -/// Before #660 the read_model surface returned an empty `loopDominance` with -/// no hint why; now the actionable Euler guidance must reach the caller via -/// the `analysisError` field so an agent asking "what loops?" understands the -/// model needs Euler (or LTM disabled). +/// An RK4 model with a stock in a loop is analyzed like any other: LTM runs +/// under every integration method (its scores dt-step ratios reported at the +/// saved steps), so `read_model` reports the loop and no `analysisError`. #[tokio::test] -async fn read_model_rk4_loop_surfaces_euler_analysis_error() { +async fn read_model_rk4_loop_is_analyzed() { let rk4_model = serde_json::json!({ "name": "rk4_loop", "simSpecs": { @@ -283,26 +280,84 @@ async fn read_model_rk4_loop_surfaces_euler_analysis_error() { }; let output = read_model(&TestFileSystemAccess, input).await.unwrap(); + assert!( + output.analysis_error.is_none(), + "RK4 + LTM read_model analyzes the model: {:?}", + output.analysis_error + ); + assert!( + !output.loop_dominance.is_empty(), + "the reinforcing loop through births is reported under RK4" + ); + + // Nothing reaches the wire (serialized) shape either. + let value = serde_json::to_value(&output).unwrap(); + assert!( + value.get("analysisError").is_none(), + "no analysisError is serialized for an analyzable model" + ); +} + +/// GH #660: a model that cannot be compiled for LTM analysis -- here a flow +/// reading a variable that does not exist -- surfaces the actionable compile +/// error through `analysisError`, on the struct and on the wire under +/// camelCase, instead of an empty `loopDominance` that an agent asking "what +/// loops?" could not tell from "none". +#[tokio::test] +async fn read_model_uncompilable_loop_surfaces_analysis_error() { + let broken_model = serde_json::json!({ + "name": "broken_loop", + "simSpecs": { + "startTime": 0.0, + "endTime": 10.0, + "dt": "1" + }, + "models": [{ + "name": "main", + "stocks": [ + {"uid": 1, "name": "population", "initialEquation": "100", + "inflows": ["births"], "outflows": []} + ], + "flows": [ + {"uid": 2, "name": "births", + "equation": "population * nonexistent_variable"} + ] + }] + }); + + let dir = tempfile::tempdir().unwrap(); + let path = dir.path().join("broken_loop.sd.json"); + std::fs::write(&path, broken_model.to_string()).unwrap(); + + let input = ReadModelInput { + project_path: path.to_str().unwrap().to_string(), + model_name: None, + }; + let output = read_model(&TestFileSystemAccess, input).await.unwrap(); + let msg = output .analysis_error .as_deref() - .expect("RK4 + LTM read_model must surface an analysisError"); + .expect("an uncompilable model must surface an analysisError"); + // The message names the variable that failed to compile, never the + // reference it could not resolve. assert!( - msg.contains("Euler"), - "analysisError must reference the Euler assumption, got: {msg}" + msg.contains("births"), + "analysisError must name the failing variable (births), got: {msg}" ); assert!( output.loop_dominance.is_empty(), "loop_dominance must be empty when the model can't be compiled for LTM" ); - // It must also reach the wire (serialized) shape under camelCase. + // It reaches the wire (serialized) shape under camelCase. let value = serde_json::to_value(&output).unwrap(); assert!( value["analysisError"] .as_str() - .is_some_and(|s| s.contains("Euler")), - "serialized analysisError must carry the Euler guidance" + .is_some_and(|s| s.contains("births")), + "serialized analysisError must carry the compile error, got: {}", + value["analysisError"] ); } From 95f848ffd6c3de93f8dd3f7eb509b3fd86d9708a Mon Sep 17 00:00:00 2001 From: Bobby Powers Date: Tue, 8 Sep 2026 22:14:35 -0700 Subject: [PATCH 05/10] engine: classify runtime polarity on the relative loop scores The Rux/Bux/U classification summed a loop's RAW loop scores over time. Raw scores are unbounded: at a dominance inflection every raw score in a partition diverges together, so a handful of inflection steps can outweigh the whole run, and an exogenous change that swamps a target's change shrinks one phase's raw scores to nothing. A series balancing for 200 steps and reinforcing for three inflection steps read confidence 0.23 on the raw base; a lone loop balancing for 40 of its 439 active steps read Rux at 0.9999 only because a ramp had made those 40 raw magnitudes tiny. Both callers -- reclassify_loops_from_results on the exhaustive path and rank_truncate_and_id on the discovery path -- now feed the classifier the loop's partition-relative series from the one normalization owner (compute_rel_loop_scores; the rel_scores discovery already attaches). Each sample is bounded to [-1, 1] and is the loop's share of its partition at that step, so the confidence is the dominance-weighted time share of each sign, and for a loop alone in its partition (relative score exactly +1/-1 while active) the plain time share of its sign: Mostly* then needs the minority sign on at most half a percent of the active steps. The papers define the ratio on instantaneous pathway scores and the base a reference tool uses for loops is undocumented, so this is a recorded judgment (types.rs and the design doc), not a reproduction. Discovery classifies inside rank_truncate_and_id, once the partition totals exist and before ids are assigned. A never-active loop is not a fallback case there: its zero/NaN series never passes retention, so discovery does not report it at all (the exhaustive surface keeps its structural label). Numbers pinned through the real pipeline (exhaustive reclassification, discovery parity, libsimlin Rux and Bux, pysimlin): the GH #679 lone-loop fixture reads Undetermined at its time share, 0.8178 with 399 reinforcing and 40 balancing steps (it was Rux 0.9999 on the raw base; three tests re-baselined with the reason, and the tests read the step counts off the series rather than pinning the number of startup steps); a competing-sibling fixture -- an inflow loop with gain +0.02 flipped to -0.02 for eight steps while the outflow loop's gain rises from 0.01 to 2 and carries the partition -- is MostlyReinforcing at 0.9994 (share 2/3 at about 430 steps, about 0.0099 at 8), and its outflow mirror MostlyBalancing at 0.9994; the yeast alcohol model's growth loop is Undetermined at 0.9655 with 146 positive and 53 negative steps; the synthetic 200-plus-3-step series reads 0.9706 on the relative base against 0.2308 on the raw one; an A2A loop with one reinforcing and one balancing element reads Undetermined at exactly 0 under the all-slots reading. --- docs/design/ltm--loops-that-matter.md | 67 ++- docs/reference/ltm--loops-that-matter.md | 6 +- src/libsimlin/simlin.h | 7 +- src/libsimlin/src/analysis.rs | 7 +- src/libsimlin/tests/integration/analysis.rs | 151 ++++-- src/pysimlin/simlin/analysis.py | 5 +- src/pysimlin/simlin/sim.py | 7 +- .../tests/test_runtime_polarity_base.py | 97 ++++ src/simlin-engine/src/db/analysis.rs | 70 ++- src/simlin-engine/src/ltm/types.rs | 30 +- src/simlin-engine/src/ltm_finding.rs | 54 +- src/simlin-engine/src/ltm_finding_tests.rs | 47 ++ src/simlin-engine/src/ltm_post.rs | 44 ++ .../tests/integration/simulate_ltm.rs | 492 +++++++++++++++--- 14 files changed, 889 insertions(+), 195 deletions(-) create mode 100644 src/pysimlin/tests/test_runtime_polarity_base.py diff --git a/docs/design/ltm--loops-that-matter.md b/docs/design/ltm--loops-that-matter.md index 5f1a2c1ca..840f430ac 100644 --- a/docs/design/ltm--loops-that-matter.md +++ b/docs/design/ltm--loops-that-matter.md @@ -808,7 +808,30 @@ classified as `Undetermined` (`calculate_polarity`). ### Runtime Polarity `LoopPolarity::from_runtime_scores()` in `ltm/types.rs` classifies polarity -based on actual simulation results. It filters out NaN and zero values, then: +from the loop's **partition-relative** score series -- the one owner's output +(`ltm_post::compute_rel_loop_scores` on the exhaustive path, the `rel_scores` +`rank_truncate_and_id` attaches on the discovery path), never the raw +`loop_score`. Each relative sample is bounded to `[-1, 1]` and weighted by the +loop's share of its partition at that step, so the confidence ratio is the +dominance-weighted time share of each sign; a raw base is unbounded and lets +the few steps around a dominance inflection, where every raw score in the +partition diverges, decide the label by themselves. For a loop alone in its +partition the relative sample is exactly `+1`/`-1`/`0`, so its confidence is +the plain time share of its sign, and `Mostly*` requires the minority sign on +at most half a percent of the active steps. + +The base is a judgment, not a reproduction. The papers define the confidence +ratio on instantaneous *pathway* scores (reference section 13.7); the base a +reference tool uses when it labels a *loop* Rux/Bux is undocumented and +cannot be checked. The relative base is chosen because it is bounded and +dominance-weighted: on the raw base a lone loop that spends a tenth of its +run balancing can still clear the 0.99 gate whenever an exogenous change +swamps the change in its target and shrinks that phase's raw scores to +nothing, which is the raw-magnitude incomparability relative scores exist to +remove. `exhaustive_lone_loop_sign_flip_confidence_is_its_time_share` pins +that case as `Undetermined`. + +The classifier filters out NaN and zero values, then: - All remaining scores positive -> `Reinforcing` - All remaining scores negative -> `Balancing` - Mixed signs, one polarity dominant with confidence >= @@ -832,10 +855,12 @@ active step. Runtime reclassification is therefore a *post-simulation* concern, and the surfaces handle it differently: - **Discovery (`analyze_model` / MCP / `simlin_analyze_discover_loops`)**: the - `FoundLoop` path in `ltm_finding.rs` derives each loop's polarity directly - from `from_runtime_scores` over the loop's own per-step score series - (falling back to the trimmed-chain structural polarity for an all-zero/NaN - series). Fully reclassified. + `FoundLoop` path in `ltm_finding.rs` derives each loop's polarity from + `from_runtime_scores` over the loop's partition-relative series once + `rank_truncate_and_id` has the partition totals. A never-active loop + (all-zero/NaN series) is not reported at all: retention drops it before + classification, so every discovered loop carries a runtime label. Fully + reclassified. - **pysimlin `Run.loops`**: sources polarity / confidence / partition straight from the engine primitive (bound as `Sim.get_loops_runtime` -> `reclassify_loops_from_results`, GH #679/#685, the all-slots Rust source of @@ -852,12 +877,12 @@ concern, and the surfaces handle it differently: runtime one. `db::analysis::reclassify_loops_from_results(loops, results, loop_partitions)` -is the **canonical in-engine reclassification primitive** -- it reads each -loop's `$⁚ltm⁚loop_score⁚{id}` slot(s) from a `Results` and applies -`from_runtime_scores` to overwrite `polarity`/`polarity_confidence`. As of this -writing it has **no production caller**: it exists so a future sim-bearing Rust -consumer (e.g. when GH #495's FFI lands) has one correct place to call rather -than re-deriving the loop-score read. It is exercised by engine tests. +is the **canonical in-engine reclassification primitive** -- it normalizes +every loop's `$⁚ltm⁚loop_score⁚{id}` slot(s) in a `Results` through +`ltm_post::compute_rel_loop_scores` and applies `from_runtime_scores` to the +relative series to overwrite `polarity`/`polarity_confidence`. Its production +caller is libsimlin's `simlin_analyze_get_loops_runtime` (and through it +pysimlin's `Run.loops`). The **loop id never changes** under reclassification. Loop detection and the deterministic `r{n}`/`b{n}`/`u{n}` id assignment happen at compile time before @@ -869,14 +894,18 @@ zero or non-finite) keeps its structural polarity -- there is no runtime evidence to override it. **A2A semantics across the sites.** The Rust `reclassify_loops_from_results` -helper concatenates *all* element slots of an A2A loop into one sample set (so a -loop that is reinforcing in one element and balancing in another classifies -`Undetermined`). pysimlin `Run.loops` is built on this primitive, so it reports -exactly this all-slots classification. Discovery uses one scalar score series -per `FoundLoop` (its links are element-level, so a discovered loop is always -scalar). The exhaustive (sim-bearing) and discovery surfaces thus -agree on scalar loops and differ only in how an A2A loop's element slots are -reduced -- the exhaustive path now uses the all-slots reading rather than slot 0. +helper concatenates *all* element slots of an A2A loop into one sample set, +each slot normalized in its own partition, so a loop that is reinforcing in +one element and balancing in another classifies `Undetermined` +(`exhaustive_a2a_loop_with_opposite_signed_elements_is_undetermined` pins +this at confidence 0 for two isolated one-stock elements of opposite sign). +pysimlin `Run.loops` is built on this primitive, so it reports exactly this +all-slots classification. Discovery uses one scalar score series per +`FoundLoop` (its links are element-level, so a discovered loop is always +scalar). The exhaustive (sim-bearing) and discovery surfaces thus agree on +scalar loops and differ only in how an A2A loop's element slots are reduced: +all slots read together on the exhaustive path, one element-level loop per +slot on the discovery path. ## Post-Simulation Loop Discovery diff --git a/docs/reference/ltm--loops-that-matter.md b/docs/reference/ltm--loops-that-matter.md index 28cc37cf8..f6c9242dc 100644 --- a/docs/reference/ltm--loops-that-matter.md +++ b/docs/reference/ltm--loops-that-matter.md @@ -680,8 +680,10 @@ nature of links is not important over the course of the simulation. > **Simlin implementation note.** Loop detection (and the deterministic loop-id > assignment) happens before simulation, so the structural label is what a pre-simulation > surface reports. Whether the *runtime* polarity is surfaced depends on the consumer: -> discovery (`analyze_model` / MCP) and pysimlin `Run.loops` reclassify from the runtime -> `loop_score` series while keeping the loop id stable, whereas the libsimlin / WASM / TS +> discovery (`analyze_model` / MCP) and pysimlin `Run.loops` reclassify from the loop's +> partition-relative score series (Section 4.4; bounded per step, so the confidence is the +> dominance-weighted time share of each sign, and for a loop alone in its partition the +> plain time share) while keeping the loop id stable, whereas the libsimlin / WASM / TS > `get_loops` surface is structural-only (it has no simulation results in hand and folds > Rux/Bux to R/B at the FFI boundary -- surfacing runtime polarity there is tracked under > GH #495). See the "Runtime Polarity" section of diff --git a/src/libsimlin/simlin.h b/src/libsimlin/simlin.h index 956309f18..8a755a8e9 100644 --- a/src/libsimlin/simlin.h +++ b/src/libsimlin/simlin.h @@ -504,8 +504,11 @@ SimlinLoops *simlin_analyze_get_loops(SimlinModel *model, SimlinError **out_erro // `reclassify_loops_from_results` primitive (GH #679) over it: for every loop // whose `$⁚ltm⁚loop_score⁚{id}` series exists in the results, the loop's // polarity and confidence are overwritten by -// `crate::ltm::LoopPolarity::from_runtime_scores` (the LTM papers' Rux/Bux/U -// classification with the 0.99 confidence gate). A loop whose runtime score +// `crate::ltm::LoopPolarity::from_runtime_scores` over the loop's +// partition-RELATIVE series (its per-step share of its cycle partition, the +// same series `simlin_analyze_get_relative_loop_score` returns): the LTM +// papers' Rux/Bux/U classification with the 0.99 confidence gate, on a +// bounded, dominance-weighted base. A loop whose runtime score // is never active keeps its structural classification. Loop IDs are stable // (a `u1` stays `u1` even after its polarity flips to Reinforcing). // diff --git a/src/libsimlin/src/analysis.rs b/src/libsimlin/src/analysis.rs index 01657233e..ed3e9a53e 100644 --- a/src/libsimlin/src/analysis.rs +++ b/src/libsimlin/src/analysis.rs @@ -537,8 +537,11 @@ pub unsafe extern "C" fn simlin_analyze_get_loops( /// `reclassify_loops_from_results` primitive (GH #679) over it: for every loop /// whose `$⁚ltm⁚loop_score⁚{id}` series exists in the results, the loop's /// polarity and confidence are overwritten by -/// `crate::ltm::LoopPolarity::from_runtime_scores` (the LTM papers' Rux/Bux/U -/// classification with the 0.99 confidence gate). A loop whose runtime score +/// `crate::ltm::LoopPolarity::from_runtime_scores` over the loop's +/// partition-RELATIVE series (its per-step share of its cycle partition, the +/// same series `simlin_analyze_get_relative_loop_score` returns): the LTM +/// papers' Rux/Bux/U classification with the 0.99 confidence gate, on a +/// bounded, dominance-weighted base. A loop whose runtime score /// is never active keeps its structural classification. Loop IDs are stable /// (a `u1` stays `u1` even after its polarity flips to Reinforcing). /// diff --git a/src/libsimlin/tests/integration/analysis.rs b/src/libsimlin/tests/integration/analysis.rs index 7f11ddc60..ae5d0bf83 100644 --- a/src/libsimlin/tests/integration/analysis.rs +++ b/src/libsimlin/tests/integration/analysis.rs @@ -3368,31 +3368,51 @@ fn runtime_loops_reclassify_sign_flipping_loop() { } /// #679: the runtime loops surface must be able to report a DEFINITE -/// `MostlyReinforcing` (Rux) with its real dominance-ratio confidence -- the -/// sibling test above (`runtime_loops_reclassify_sign_flipping_loop`) only -/// pins that a sign-straddling series leaves the {Rux, Bux, U} family, while -/// this one pins the Rux outcome itself through the C ABI. +/// `MostlyReinforcing` (Rux) / `MostlyBalancing` (Bux) with its real +/// confidence -- the sibling test above +/// (`runtime_loops_reclassify_sign_flipping_loop`) only pins that a +/// sign-straddling series leaves the {Rux, Bux, U} family, while the two +/// tests below pin each outcome itself through the C ABI. /// /// The fixture (shared shape with the engine's -/// `exhaustive_mixed_sign_dominant_loop_reclassifies_to_rux`): `f = s*g + d` -/// where the marginal gain `g` is +0.02 for t < 100 (loop score +1 per step, -/// ~400 steps) then -0.0001 while an exogenous ramp `d` dominates the -/// per-step change in `f` (40 negative samples of magnitude ~1e-4..1e-3). -/// The dominance ratio lands at ~0.9999: above the 0.99 Rux gate, strictly -/// below 1.0. Structurally the loop is Reinforcing (the bare co-factor `g` -/// is positive by the SD labeling convention -- and is exactly the kind of -/// quantity the convention can be wrong about, since its value flips sign -/// mid-run), so Rux through this surface both remains the label the issue -/// says was unreachable AND demonstrates runtime evidence downgrading a +/// `exhaustive_competing_sign_flip_reclassifies_to_rux` / `_bux`): a stock +/// with an inflow loop `f_in = s*` and an outflow loop +/// `f_out = s*`, where the `flipping` loop's gain is `g` +/// (+0.02, flipped to -0.02 for t in [100, 102)) and the other loop's is `d` +/// (0.01, raised to 2 in that window so it carries the partition while the +/// flipping loop's sign is wrong). The classifier reads the partition-RELATIVE +/// series: the flipping loop's share is 2/3 at each of its roughly 430 +/// same-sign steps and ~0.0099 at its 8 opposite-sign steps, so the +/// confidence is ~0.9994 -- above the 0.99 gate, strictly below 1.0. +/// Structurally the flipping loop is single-signed (the bare co-factor `g` +/// is positive by the SD labeling convention -- exactly the kind of quantity +/// the convention can be wrong about, since its value flips sign mid-run), +/// so `Mostly*` through this surface both remains the label the issue says +/// was unreachable AND demonstrates runtime evidence downgrading a /// conventional structural sign. -#[test] -fn runtime_loops_report_definite_rux() { - let test_project = TestProject::new("rux_mixed_sign") +/// +/// Returns the flipping loop's (structural polarity, structural confidence, +/// runtime polarity, runtime confidence) and asserts its id is the same on +/// both surfaces. +fn competing_sign_flip_through_the_c_abi( + flipping: &str, +) -> (SimlinLoopPolarity, f64, SimlinLoopPolarity, f64) { + let (inflow_gain, outflow_gain, flow) = if flipping == "inflow" { + ("g", "d", "f_in") + } else { + ("d", "g", "f_out") + }; + let test_project = TestProject::new("competing_sign_flip") .with_sim_time(0.0, 110.0, 0.25) - .aux("g", "IF TIME < 100 THEN 0.02 ELSE -0.0001", None) - .aux("d", "IF TIME < 100 THEN 0 ELSE 50 * (TIME - 100)", None) - .stock("s", "100", &["f"], &[], None) - .flow("f", "s * g + d", None); + .aux( + "g", + "IF TIME < 100 OR TIME >= 102 THEN 0.02 ELSE -0.02", + None, + ) + .aux("d", "IF TIME < 100 OR TIME >= 102 THEN 0.01 ELSE 2", None) + .stock("s", "100", &["f_in"], &["f_out"], None) + .flow("f_in", &format!("s * {inflow_gain}"), None) + .flow("f_out", &format!("s * {outflow_gain}"), None); let datamodel_project = test_project.build_datamodel(); let project = engine_serde::serialize(&datamodel_project).unwrap(); let mut buf = Vec::new(); @@ -3412,49 +3432,94 @@ fn runtime_loops_report_definite_rux() { simlin_sim_run_to_end(sim, &mut err); assert!(err.is_null()); - // Structural baseline: one loop, Reinforcing by the bare-co-factor - // convention at the binary structural confidence 1.0. + let is_flipping_loop = |l: &SimlinLoop| { + std::slice::from_raw_parts(l.variables, l.var_count) + .iter() + .any(|v| CStr::from_ptr(*v).to_str().unwrap() == flow) + }; + let structural = simlin_analyze_get_loops(model, &mut err); assert!(err.is_null()); assert!(!structural.is_null()); - assert_eq!((*structural).count, 1, "the fixture has exactly one loop"); - let struct_loop = &std::slice::from_raw_parts((*structural).loops, 1)[0]; - assert!(struct_loop.polarity == SimlinLoopPolarity::Reinforcing); - assert_eq!(struct_loop.polarity_confidence, 1.0); + assert_eq!((*structural).count, 2, "an inflow loop and an outflow loop"); + let struct_loop = std::slice::from_raw_parts((*structural).loops, 2) + .iter() + .find(|l| is_flipping_loop(l)) + .expect("the flipping loop"); - // Runtime surface: the same loop reports the definite Rux variant - // with the real (>= 0.99, < 1.0) dominance ratio. err = ptr::null_mut(); let runtime = simlin_analyze_get_loops_runtime(sim, &mut err); assert!(err.is_null()); assert!(!runtime.is_null()); - assert_eq!((*runtime).count, 1); - let rt_loop = &std::slice::from_raw_parts((*runtime).loops, 1)[0]; - assert!( - rt_loop.polarity == SimlinLoopPolarity::MostlyReinforcing, - "an overwhelmingly-positive mixed-sign loop_score must surface as \ - MostlyReinforcing through the C ABI (confidence {})", - rt_loop.polarity_confidence - ); - assert!( - rt_loop.polarity_confidence >= 0.99 && rt_loop.polarity_confidence < 1.0, - "the Rux confidence is the real dominance ratio (>= the 0.99 gate, \ - strictly < 1.0 because both signs are present); got {}", - rt_loop.polarity_confidence - ); + assert_eq!((*runtime).count, 2); + let rt_loop = std::slice::from_raw_parts((*runtime).loops, 2) + .iter() + .find(|l| is_flipping_loop(l)) + .expect("the flipping loop"); + // The loop id is stable across the two surfaces. let struct_id = CStr::from_ptr(struct_loop.id).to_str().unwrap(); let rt_id = CStr::from_ptr(rt_loop.id).to_str().unwrap(); assert_eq!(struct_id, rt_id); + let out = ( + struct_loop.polarity, + struct_loop.polarity_confidence, + rt_loop.polarity, + rt_loop.polarity_confidence, + ); + simlin_free_loops(structural); simlin_free_loops(runtime); simlin_sim_unref(sim); simlin_model_unref(model); simlin_project_unref(proj); + out } } +/// The Rux arm through the C ABI: the inflow loop flips, and is +/// `MostlyReinforcing` at the real (>= 0.99, < 1.0) confidence against a +/// structural Reinforcing/1.0 baseline. +#[test] +fn runtime_loops_report_definite_rux() { + let (structural, structural_confidence, runtime, confidence) = + competing_sign_flip_through_the_c_abi("inflow"); + assert!(structural == SimlinLoopPolarity::Reinforcing); + assert_eq!(structural_confidence, 1.0); + assert!( + runtime == SimlinLoopPolarity::MostlyReinforcing, + "an overwhelmingly-positive mixed-sign relative series must surface as \ + MostlyReinforcing through the C ABI (confidence {confidence})" + ); + assert!( + (0.99..1.0).contains(&confidence), + "the Rux confidence is the real dominance ratio (>= the 0.99 gate, \ + strictly < 1.0 because both signs are present); got {confidence}" + ); +} + +/// The Bux arm through the C ABI: the outflow loop flips, and is +/// `MostlyBalancing` at the same confidence against a structural +/// Balancing/1.0 baseline (an outflow's link to its stock is Negative). +#[test] +fn runtime_loops_report_definite_bux() { + let (structural, structural_confidence, runtime, confidence) = + competing_sign_flip_through_the_c_abi("outflow"); + assert!(structural == SimlinLoopPolarity::Balancing); + assert_eq!(structural_confidence, 1.0); + assert!( + runtime == SimlinLoopPolarity::MostlyBalancing, + "an overwhelmingly-negative mixed-sign relative series must surface as \ + MostlyBalancing through the C ABI (confidence {confidence})" + ); + assert!( + (0.99..1.0).contains(&confidence), + "the Bux confidence is the real dominance ratio (>= the 0.99 gate, \ + strictly < 1.0 because both signs are present); got {confidence}" + ); +} + /// The runtime loops FFI requires a run sim: calling it on a freshly-created /// (un-run) sim reports an error and returns NULL rather than panicking. #[test] diff --git a/src/pysimlin/simlin/analysis.py b/src/pysimlin/simlin/analysis.py index 2a1afc74a..9698535c4 100644 --- a/src/pysimlin/simlin/analysis.py +++ b/src/pysimlin/simlin/analysis.py @@ -328,8 +328,9 @@ class Loop: FFI verbatim (GH #495) -- the MOSTLY_* ("Rux"/"Bux") variants are no longer coalesced onto REINFORCING/BALANCING. They occur on the runtime surfaces (``Run.loops``, ``Sim.get_loops_runtime``, and discovery), where the - polarity is classified from runtime score series; see - :attr:`polarity_confidence`.""" + polarity is classified from the loop's partition-relative runtime score + series (its per-step share of its cycle partition, the series + :attr:`behavior_time_series` carries); see :attr:`polarity_confidence`.""" polarity_confidence: float = 1.0 """Polarity-confidence ratio in ``[0.0, 1.0]`` behind :attr:`polarity` diff --git a/src/pysimlin/simlin/sim.py b/src/pysimlin/simlin/sim.py index a84405e57..8c6b46d45 100644 --- a/src/pysimlin/simlin/sim.py +++ b/src/pysimlin/simlin/sim.py @@ -343,9 +343,12 @@ def get_loops_runtime(self) -> list[Loop]: only a model and reports STRUCTURAL polarity). It builds the same exhaustive structural loop set, then reclassifies each loop's polarity and ``polarity_confidence`` from its post-simulation - ``$:ltm:loop_score:{id}`` series via the engine's + ``$:ltm:loop_score:{id}`` series, normalized to the loop's share of its + cycle partition at each step, via the engine's ``reclassify_loops_from_results`` primitive (GH #679): the LTM papers' - Rux/Bux/U runtime classification with the 0.99 confidence gate. A loop + Rux/Bux/U runtime classification with the 0.99 confidence gate, read on + the partition-relative series (the same series + :attr:`Loop.behavior_time_series` carries). A loop whose runtime score is never active keeps its structural classification; loop ids are stable across the two surfaces. diff --git a/src/pysimlin/tests/test_runtime_polarity_base.py b/src/pysimlin/tests/test_runtime_polarity_base.py new file mode 100644 index 000000000..b3f606f22 --- /dev/null +++ b/src/pysimlin/tests/test_runtime_polarity_base.py @@ -0,0 +1,97 @@ +"""Runtime polarity is classified on the partition-relative loop-score series. + +The engine primitive behind ``Run.loops`` (``reclassify_loops_from_results``) +feeds the classifier each loop's share of its cycle partition, bounded per +step, so the confidence is the dominance-weighted time share of each sign. +``exhaustive_competing_sign_flip_reclassifies_to_rux`` in the engine holds the +hand calculation for this fixture; here the same numbers arrive through the +FFI and ``Run.loops``. +""" + +from __future__ import annotations + +import numpy as np +import pytest + +import simlin +from simlin import LoopPolarity +from simlin.types import Aux, Flow, Stock + + +def _competing_sign_flip_model(flipping: str) -> simlin.Model: + """A stock with an inflow loop and an outflow loop; the ``flipping`` loop's + gain g is +0.02, flipped to -0.02 for t in [100, 102), while the other + loop's gain d rises from 0.01 to 2 in that window and carries the + partition.""" + inflow_gain, outflow_gain = ("g", "d") if flipping == "inflow" else ("d", "g") + project = simlin.Project.new(name="competing_flip", sim_start=0.0, sim_stop=110.0, dt=0.25) + model = project.main_model + with model.edit() as (_current, patch): + patch.upsert(Aux(name="g", equation="IF TIME < 100 OR TIME >= 102 THEN 0.02 ELSE -0.02")) + patch.upsert(Aux(name="d", equation="IF TIME < 100 OR TIME >= 102 THEN 0.01 ELSE 2")) + patch.upsert(Stock(name="s", initial_equation="100", inflows=["f_in"], outflows=["f_out"])) + patch.upsert(Flow(name="f_in", equation=f"s * {inflow_gain}")) + patch.upsert(Flow(name="f_out", equation=f"s * {outflow_gain}")) + return model + + +class TestRuntimePolarityOnTheRelativeBase: + @pytest.mark.parametrize( + ("flipping", "expected", "minority_sign"), + [ + ("inflow", LoopPolarity.MOSTLY_REINFORCING, -1), + ("outflow", LoopPolarity.MOSTLY_BALANCING, 1), + ], + ) + def test_competing_sign_flip_is_mostly_its_dominant_sign( + self, flipping: str, expected: LoopPolarity, minority_sign: int + ) -> None: + model = _competing_sign_flip_model(flipping) + run = model.run(analyze_loops=True) + assert run.ltm_mode == "exhaustive" + flow = "f_in" if flipping == "inflow" else "f_out" + loop = next(lp for lp in run.loops if flow in lp.variables) + assert loop.polarity == expected + assert 0.99 <= loop.polarity_confidence < 1.0 + assert loop.polarity_confidence == pytest.approx(0.9994, abs=1e-3) + # The behavior series is the relative series the label was read from: + # eight steps carry the minority sign, at a share of about one percent. + series = loop.behavior_time_series + assert series is not None + minority = series[np.sign(series) == minority_sign] + assert minority.size == 8 + assert np.all(np.abs(minority) < 0.02) + + def test_lone_loop_confidence_is_the_time_share_of_its_sign(self) -> None: + """A loop alone in its partition has a relative score of exactly +1/-1 + while active, so the confidence is the time share of its sign: this + loop is reinforcing for the ~400 steps before t = 100 and balancing + for the 40 after, roughly nine percent of its active life, which is + Undetermined at a confidence of about 0.82. The counts are read off + the series the engine classified rather than pinned, because the + number of active startup steps is an engine detail this test does not + own.""" + project = simlin.Project.new(name="lone_flip", sim_start=0.0, sim_stop=110.0, dt=0.25) + model = project.main_model + with model.edit() as (_current, patch): + patch.upsert(Aux(name="g", equation="IF TIME < 100 THEN 0.02 ELSE -0.0001")) + patch.upsert(Aux(name="d", equation="IF TIME < 100 THEN 0 ELSE 50 * (TIME - 100)")) + patch.upsert(Stock(name="s", initial_equation="100", inflows=["f"], outflows=[])) + patch.upsert(Flow(name="f", equation="s * g + d")) + run = model.run(analyze_loops=True) + (loop,) = run.loops + series = loop.behavior_time_series + assert series is not None + assert np.all(np.isin(series, [-1.0, 0.0, 1.0])) + positive = int((series > 0).sum()) + negative = int((series < 0).sum()) + # Balancing at each of the 40 steps in (100, 110]; reinforcing at every + # other active step. The run saves 441 steps and the loop is active at + # all of them but the startup step or two the flow-to-stock score needs. + assert negative == 40 + active = positive + negative + assert 439 <= active <= 441 + time_share = abs(positive - negative) / active + assert 0.8 < time_share < 0.83 + assert loop.polarity == LoopPolarity.UNDETERMINED + assert loop.polarity_confidence == pytest.approx(time_share) diff --git a/src/simlin-engine/src/db/analysis.rs b/src/simlin-engine/src/db/analysis.rs index d551fedf0..0ddc0ce26 100644 --- a/src/simlin-engine/src/db/analysis.rs +++ b/src/simlin-engine/src/db/analysis.rs @@ -2826,10 +2826,12 @@ fn detected_loop_from_loop(l: &crate::ltm::Loop, pin_name: &str) -> DetectedLoop /// active step. /// /// After simulation the per-loop `loop_score` series exists; this helper -/// reads each loop's `$⁚ltm⁚loop_score⁚{id}` slot(s) out of `results` and -/// runs [`crate::ltm::LoopPolarity::from_runtime_scores`] over them, -/// overwriting the loop's `polarity` and `polarity_confidence` with the -/// runtime classification. The loop **id stays unchanged**: loop detection +/// normalizes every loop's `$⁚ltm⁚loop_score⁚{id}` slot(s) in `results` +/// through `ltm_post::compute_rel_loop_scores` and runs +/// [`crate::ltm::LoopPolarity::from_runtime_scores`] over each loop's +/// partition-relative series, overwriting the loop's `polarity` and +/// `polarity_confidence` with the runtime classification. The loop **id +/// stays unchanged**: loop detection /// and the deterministic `r{n}`/`b{n}`/`u{n}` id assignment happen at /// compile time before any simulation, and the FFI id->score correspondence /// plus salsa caching depend on the id never changing retroactively. A loop @@ -2860,18 +2862,31 @@ fn detected_loop_from_loop(l: &crate::ltm::Loop, pin_name: &str) -> DetectedLoop /// `analyze_model` / MCP surface is discovery-based and reclassifies through /// the `FoundLoop` path. /// +/// # The base is the partition-relative series +/// +/// The samples fed to the classifier are the loop's **partition-relative** +/// scores from the one owner (`ltm_post::compute_rel_loop_scores`), not its +/// raw `loop_score`: each sample is bounded to `[-1, 1]` and weighted by the +/// loop's share of its partition at that step, so the confidence reads as +/// the dominance-weighted time share of each sign. A raw base lets a handful +/// of inflection steps, where every raw score in the partition diverges, +/// decide the label on their own. For a loop alone in its partition the +/// relative sample is exactly `+1`/`-1`/`0`, so its confidence is the plain +/// time share of its sign. See [`crate::ltm::LoopPolarity::from_runtime_scores`]. +/// /// # A2A semantics differ between the two reclassification sites /// /// `loop_partitions` is the per-loop slot->partition map carried on /// `LtmVariablesResult::loop_partitions`; its slot-vector length is the /// `loop_score` series' slot count. For an A2A (per-element) loop this helper -/// **concatenates every element slot's series into one sample set** and -/// classifies the mixed result: if any element of the loop is balancing while -/// another is reinforcing the loop classifies `Undetermined` (a deliberate -/// "the loop's sign is not uniform across the array" reading). This is NOT -/// the input construction discovery uses, so do not claim they agree: -/// **discovery** (`ltm_finding`) classifies each `FoundLoop` from its own -/// single scalar score series. +/// **concatenates every element slot's series into one sample set** (the +/// owner's step-major per-slot layout, each slot normalized in its own +/// partition) and classifies the mixed result: if any element of the loop +/// is balancing while another is reinforcing the loop classifies +/// `Undetermined` (a deliberate "the loop's sign is not uniform across the +/// array" reading). This is NOT the input construction discovery uses, so do +/// not claim they agree: **discovery** (`ltm_finding`) classifies each +/// `FoundLoop` from its own single scalar relative series. /// /// Both sites share the *scalar* semantics (`from_runtime_scores`'s NaN/zero /// filter; all-positive -> Reinforcing, all-negative -> Balancing, mixed @@ -2883,36 +2898,15 @@ pub fn reclassify_loops_from_results( results: &crate::Results, loop_partitions: &indexmap::IndexMap>>, ) { + // One normalization for every loop at once; a loop with no emitted + // `loop_score` column (discovery mode scores only pinned loops) is + // absent from the map and keeps its structural label. + let relative = crate::ltm_post::compute_rel_loop_scores(results, loop_partitions); for loop_item in loops.iter_mut() { - let Some(&base_off) = results - .offsets - .get(&crate::ltm_post::loop_score_ident(&loop_item.id)) - else { - // No emitted loop_score series (e.g. discovery mode emits scores - // only for pinned loops): nothing to reclassify against. + let Some(series) = relative.get(&loop_item.id) else { continue; }; - - // An A2A loop's loop_score occupies `n_slots` consecutive offsets; - // a scalar/cross-element/mixed loop has exactly one. The slot count - // comes from the loop's partition vector (1 when absent), the slot - // count `ltm_post::compute_rel_loop_scores` lays the loop out with. - let n_slots = loop_partitions - .get(&loop_item.id) - .map(|p| p.len().max(1)) - .unwrap_or(1); - - let mut scores: Vec = Vec::with_capacity(results.step_count * n_slots); - for row in results.iter() { - for slot in 0..n_slots { - let off = base_off + slot; - if off < results.step_size { - scores.push(row[off]); - } - } - } - - if let Some((polarity, confidence)) = crate::ltm::LoopPolarity::from_runtime_scores(&scores) + if let Some((polarity, confidence)) = crate::ltm::LoopPolarity::from_runtime_scores(series) { loop_item.polarity = detected_polarity_from_ltm(&polarity); loop_item.polarity_confidence = confidence; diff --git a/src/simlin-engine/src/ltm/types.rs b/src/simlin-engine/src/ltm/types.rs index 33353c2e3..44dfc851b 100644 --- a/src/simlin-engine/src/ltm/types.rs +++ b/src/simlin-engine/src/ltm/types.rs @@ -168,9 +168,9 @@ impl Loop { /// - Odd number of negative links → Balancing /// - ANY link with unknown polarity → Undetermined /// -/// At runtime, the loop score series is also classified according to the -/// signed-sum confidence ratio `|r - |b|| / (r + |b|)` (Schoenberg and -/// Eberlein, 2020; see `docs/reference/ltm--loops-that-matter.md`): +/// At runtime, the loop's partition-relative score series is also classified +/// according to the signed-sum confidence ratio `|r - |b|| / (r + |b|)` +/// (Schoenberg and Eberlein, 2020; see `docs/reference/ltm--loops-that-matter.md`): /// - r is the sum of positive scores, |b| the absolute sum of negative scores /// - When all valid scores share a sign the ratio is exactly 1 and the loop /// is classified Reinforcing or Balancing. @@ -225,8 +225,28 @@ pub enum LoopPolarity { pub const POLARITY_CONFIDENCE_THRESHOLD: f64 = 0.99; impl LoopPolarity { - /// Classify loop polarity based on actual runtime loop score values - /// and return the polarity-confidence ratio in `[0.0, 1.0]`. + /// Classify loop polarity from a runtime loop-score series and return the + /// polarity-confidence ratio in `[0.0, 1.0]`. + /// + /// `scores` is the loop's **partition-relative** series (the one owner, + /// `ltm_post::compute_rel_loop_scores` / `relative_series`), never its + /// raw `loop_score`. Each relative sample is bounded to `[-1, 1]` and is + /// the loop's share of its partition at that step, so the ratio below is + /// the dominance-weighted time share of each sign: a loop that carries + /// the partition while reinforcing and is negligible while balancing is + /// `MostlyReinforcing`, and a loop that spends a tenth of its active + /// steps balancing at full share is `Undetermined`. The raw series is + /// the wrong base because it is unbounded: at a dominance inflection + /// every raw score in the partition diverges together, so a few + /// inflection steps can outweigh the whole run, and an exogenous change + /// that swamps a target's change shrinks the raw scores of one phase to + /// nothing. For a loop alone in its partition the relative sample is + /// exactly `+1`, `-1` or `0`, so its confidence is the plain time share of + /// its sign and `Mostly*` needs the minority sign on at most half a + /// percent of its active steps. The papers define the ratio on + /// instantaneous pathway scores; the base a reference tool uses for loops + /// is not documented, so this is a judgment recorded here, not a + /// reproduction. /// /// The confidence is `|r - |b|| / (r + |b|)` over the valid (finite, /// non-zero) entries, where `r` and `|b|` are the sum of positive and diff --git a/src/simlin-engine/src/ltm_finding.rs b/src/simlin-engine/src/ltm_finding.rs index 5d0408422..3c2406358 100644 --- a/src/simlin-engine/src/ltm_finding.rs +++ b/src/simlin-engine/src/ltm_finding.rs @@ -3325,29 +3325,26 @@ pub(crate) fn discover_loops_with_deadlines( } let polarity_structural = causal_graph.calculate_polarity(&reported_links); - // Determine runtime polarity from scores, capturing the confidence - // ratio alongside it (GH #495). When the loop has no valid runtime - // scores we fall back to the structural polarity; the matching - // confidence mirrors the structural pipeline's convention in - // `db::analysis` (1.0 when the polarity is determined, 0.0 when it is - // Undetermined) so the discovery and structural surfaces agree on what - // a "fully confident" loop reports. - let runtime_scores: Vec = scores.iter().map(|(_, s)| *s).collect(); - let (polarity, polarity_confidence) = LoopPolarity::from_runtime_scores(&runtime_scores) - .unwrap_or_else(|| { - let confidence = if polarity_structural == LoopPolarity::Undetermined { - 0.0 - } else { - 1.0 - }; - (polarity_structural, confidence) - }); + // The structural polarity, with the structural pipeline's binary + // confidence (1.0 when every link is signed, 0.0 when one is Unknown), + // is a placeholder until the partition-relative series exist: + // `rank_truncate_and_id` reclassifies every surviving loop from those + // (GH #495). A loop with no active step never survives retention, so + // no discovered loop reports this placeholder except on the + // no-score-data path (a run with no saved steps), where there is + // nothing to classify; the binary convention keeps that path's + // meaning of a 1.0/0.0 confidence the structural surface's. + let polarity_confidence = if polarity_structural == LoopPolarity::Undetermined { + 0.0 + } else { + 1.0 + }; let loop_info = Loop { id: String::new(), // Will be assigned below links: reported_links, stocks: loop_stocks, - polarity, + polarity: polarity_structural, dimensions: vec![], slot_links: vec![], }; @@ -4297,7 +4294,11 @@ fn rank_truncate_and_id( // series get in `ltm_post::compute_rel_loop_scores`. This is the [-1, 1] // importance series `analysis::to_loop_summary` / `to_feedback_loop` // surface, so dominance/ranking is partition-relative (comparable across - // partitions) rather than raw-magnitude-biased. + // partitions) rather than raw-magnitude-biased -- and it is the base the + // runtime polarity is classified on (GH #495), so the label a loop gets + // here is the one the exhaustive reclassification gives the same series + // (`db::analysis::reclassify_loops_from_results`). The ids assigned + // below read the classified polarity, so this runs first. let mut keyed: Vec<(RelativeImportance, FoundLoop)> = std::mem::take(found_loops) .into_iter() .enumerate() @@ -4306,6 +4307,21 @@ fn rank_truncate_and_id( let mean_rel = mean_relative_contribution(&fl, totals); let key = loop_sort_key(&fl.loop_info); fl.rel_scores = relative_series(fl.scores.iter().map(|&(_, score)| score), totals); + // Retention kept this loop for a step at which its finite score + // is at least MIN_CONTRIBUTION of a positive finite total, and + // that step is a finite non-zero relative sample, so the + // classifier always has evidence here: its `None` arm is the API + // shape, not a reachable fallback (pinned by + // `a_never_active_loop_is_dropped_rather_than_reported_with_its_structural_label`). + let classified = LoopPolarity::from_runtime_scores(&fl.rel_scores); + debug_assert!( + classified.is_some(), + "a retained loop has a finite non-zero relative sample" + ); + if let Some((polarity, confidence)) = classified { + fl.loop_info.polarity = polarity; + fl.polarity_confidence = confidence; + } ( RelativeImportance { mean_rel, diff --git a/src/simlin-engine/src/ltm_finding_tests.rs b/src/simlin-engine/src/ltm_finding_tests.rs index 52df8cb52..91957b4da 100644 --- a/src/simlin-engine/src/ltm_finding_tests.rs +++ b/src/simlin-engine/src/ltm_finding_tests.rs @@ -1044,6 +1044,53 @@ fn test_rank_and_filter_truncates_to_max_loops() { assert_eq!(loops.len(), CAP, "Should truncate to the cap ({CAP})"); } +/// A loop with no active step -- every score zero or NaN -- never survives +/// `rank_and_filter`, whether it shares a partition with an active loop or +/// sits alone in its own Solo group: retention keeps a loop only for a step +/// at which its finite score is at least `MIN_CONTRIBUTION` of a positive +/// finite total, and that same step is a finite non-zero relative sample. +/// So every loop `rank_truncate_and_id` classifies has a valid sample and +/// `from_runtime_scores` is `Some` for it: a discovered loop never reports +/// the structural placeholder label it was built with. +#[test] +fn a_never_active_loop_is_dropped_rather_than_reported_with_its_structural_label() { + let mut loops = vec![ + make_found_loop_with_scores( + &[("a", "b"), ("b", "a")], + &["stock_x"], + LoopPolarity::Reinforcing, + 1.0, + vec![(0.0, 1.0), (1.0, 1.0), (2.0, 1.0)], + ), + // Shares stock_x's partition with the active loop: 0 / total = 0. + make_found_loop_with_scores( + &[("c", "d"), ("d", "c")], + &["stock_x"], + LoopPolarity::Balancing, + 0.0, + vec![(0.0, 0.0), (1.0, f64::NAN), (2.0, 0.0)], + ), + // Unpartitioned, so its own |score| series is its total: never > 0. + make_found_loop_with_scores( + &[("e", "f"), ("f", "e")], + &["stock_y"], + LoopPolarity::Balancing, + 0.0, + vec![(0.0, f64::NAN), (1.0, 0.0), (2.0, 0.0)], + ), + ]; + let partitions = single_partition(&["stock_x"]); + rank_and_filter(&mut loops, &partitions, None); + assert_eq!(loops.len(), 1, "only the active loop survives retention"); + assert_eq!(loops[0].loop_info.polarity, LoopPolarity::Reinforcing); + assert_eq!(loops[0].polarity_confidence, 1.0); + assert!( + loops[0].rel_scores.iter().all(|v| *v == 1.0), + "the survivor was classified from its relative series; got {:?}", + loops[0].rel_scores + ); +} + #[test] fn test_rank_and_filter_removes_low_contribution() { // Create loops where one dominates and others have negligible contribution. diff --git a/src/simlin-engine/src/ltm_post.rs b/src/simlin-engine/src/ltm_post.rs index 263669d59..8e62d011d 100644 --- a/src/simlin-engine/src/ltm_post.rs +++ b/src/simlin-engine/src/ltm_post.rs @@ -1440,6 +1440,50 @@ mod tests { assert_eq!(rel_b, vec![0.75, 0.5, 0.0]); } + /// The runtime polarity classifier reads the owner's relative series, + /// and the choice of base is what decides a loop with a few violent + /// inflection steps. A synthetic series: a loop balancing at raw -0.5 + /// for 200 steps and reinforcing at 40/80/40 for three inflection steps, + /// beside a sibling at 0.7 that explodes to 60/100/60 at the same + /// inflection. On the raw base the three steps carry 160 of 260 units of + /// mass (confidence 0.23); on the relative base they are 0.4, 0.44, 0.4 + /// against 200 steps at -5/12 (confidence 0.97). Both are Undetermined + /// at the 0.99 gate; the number the classifier reports is the relative + /// one. + #[test] + fn runtime_polarity_confidence_reads_the_relative_series() { + let mut raw = vec![-0.5_f64; 203]; + raw[100..103].copy_from_slice(&[40.0, 80.0, 40.0]); + let mut sibling = vec![0.7_f64; 203]; + sibling[100..103].copy_from_slice(&[60.0, 100.0, 60.0]); + let results = make_results_for_loops(&[("L", &raw), ("S", &sibling)]); + let partitions = mapping(&[("L", Some(0)), ("S", Some(0))]); + + let (raw_polarity, raw_confidence) = + crate::ltm::LoopPolarity::from_runtime_scores(&raw).unwrap(); + assert_eq!(raw_polarity, crate::ltm::LoopPolarity::Undetermined); + assert!( + (raw_confidence - 60.0 / 260.0).abs() < 1e-12, + "got {raw_confidence}" + ); + + let relative = compute_rel_loop_scores(&results, &partitions); + let series = &relative["L"]; + assert!((series[0] - (-5.0 / 12.0)).abs() < 1e-12); + assert!((series[101] - 80.0 / 180.0).abs() < 1e-12); + let (polarity, confidence) = crate::ltm::LoopPolarity::from_runtime_scores(series).unwrap(); + assert_eq!(polarity, crate::ltm::LoopPolarity::Undetermined); + // r = 0.4 + 4/9 + 0.4, |b| = 200 * 5/12. + let r = 0.4 + 4.0 / 9.0 + 0.4; + let b = 200.0 * 5.0 / 12.0; + let expected = (b - r) / (r + b); + assert!( + (confidence - expected).abs() < 1e-12, + "got {confidence}, expected {expected}" + ); + assert!((confidence - 0.9706).abs() < 1e-4); + } + /// `argmax_abs_by_step` on a 3-slot series: the slot with the largest /// magnitude wins each step, sign preserved. #[test] diff --git a/src/simlin-engine/tests/integration/simulate_ltm.rs b/src/simlin-engine/tests/integration/simulate_ltm.rs index 78f6af663..1423c12d1 100644 --- a/src/simlin-engine/tests/integration/simulate_ltm.rs +++ b/src/simlin-engine/tests/integration/simulate_ltm.rs @@ -10453,11 +10453,10 @@ fn exhaustive_never_active_loop_keeps_structural_polarity() { ); } -/// GH #679 fixture: a single-stock loop whose runtime loop score genuinely -/// MIXES signs with overwhelming positive dominance, so runtime -/// reclassification must produce `MostlyReinforcing` (the LTM papers' "Rux") -/// -- the variant that is structurally unreachable on the exhaustive -/// pre-simulation surface. +/// GH #679 fixture: a single-stock loop, ALONE in its partition, whose +/// runtime loop score genuinely MIXES signs -- reinforcing at every active +/// step before t = 100 (about 400 of them at dt = 0.25), balancing at the +/// 40 after. /// /// Construction: `f = s * g + d` with a sign-flipping marginal gain `g` and /// an exogenous drive `d`. @@ -10472,15 +10471,24 @@ fn exhaustive_never_active_loop_keeps_structural_polarity() { /// so each of the 40 negative loop-score samples has magnitude /// `|g * delta s / delta f|` of only ~1e-4..1e-3. /// -/// The dominance ratio `|r - |b|| / (r + |b|)` lands at ~0.9999 -- above the -/// 0.99 Rux gate but strictly below 1.0 (both signs are present). The -/// `s -> f` link is statically signed Positive by the bare-named-co-factor +/// On the raw base the dominance ratio `|r - |b|| / (r + |b|)` would land at +/// ~0.9999, above the 0.99 Rux gate -- only because the exogenous ramp +/// swamped `delta f` and shrank the 40 balancing samples to nothing, the +/// raw-magnitude incomparability the relative base exists to remove. The +/// classifier reads the partition-RELATIVE series, and a loop alone in its +/// partition has a relative score of exactly +1/-1 while active, so the +/// ratio is the time share of its sign, `(400 - 40) / 440` or about 0.82, +/// below the gate. (The exact count depends on how many startup steps the +/// flow-to-stock score leaves inactive; the test reads it off the series.) +/// The loop spent nine percent of its active life balancing, and the label +/// says so: `Undetermined` at that confidence. +/// +/// The `s -> f` link is statically signed Positive by the bare-named-co-factor /// convention (`g` is a named quantity, conventionally positive-valued), /// so the structural label is Reinforcing/1.0 -- and `g` is exactly the -/// kind of quantity that convention can be WRONG about, which is what -/// makes this fixture doubly useful: the runtime reclassification both -/// upgrades the label to the honest mixed-sign Rux AND demonstrates that -/// runtime evidence overrides a conventional structural sign. +/// kind of quantity that convention can be WRONG about: the runtime +/// reclassification demonstrates runtime evidence overriding a conventional +/// structural sign. fn mixed_sign_dominant_loop_project() -> simlin_engine::datamodel::Project { TestProject::new("rux_mixed_sign") .with_sim_time(0.0, 110.0, 0.25) @@ -10491,15 +10499,16 @@ fn mixed_sign_dominant_loop_project() -> simlin_engine::datamodel::Project { .build_datamodel() } -/// GH #679: the exhaustive surface must be able to report the mixed-sign -/// `MostlyReinforcing` (Rux) classification with its REAL dominance-ratio -/// confidence -- not just the single-signed R/B flip that -/// `exhaustive_module_loop_polarity_reclassified_from_runtime` pins. -/// This is the canonical case from the issue: a loop whose polarity -/// genuinely changes sign during simulation, where the honest post-sim -/// answer is Rux with a concrete confidence, not a blanket Undetermined. +/// GH #679: the exhaustive surface reports a mixed-sign loop's runtime +/// classification with its REAL confidence -- not just the single-signed +/// R/B flip that `exhaustive_module_loop_polarity_reclassified_from_runtime` +/// pins. For a loop ALONE in its partition the relative series is +1/-1 +/// while active, so the confidence is the time share of its sign, here +/// about 0.82 (see the fixture doc): sub-threshold, `Undetermined`. +/// The Rux and Bux outcomes need a competing sibling and are pinned by +/// `exhaustive_competing_sign_flip_reclassifies_to_rux` / `_bux`. #[test] -fn exhaustive_mixed_sign_dominant_loop_reclassifies_to_rux() { +fn exhaustive_lone_loop_sign_flip_confidence_is_its_time_share() { use simlin_engine::ltm::POLARITY_CONFIDENCE_THRESHOLD; let project = mixed_sign_dominant_loop_project(); @@ -10531,24 +10540,228 @@ fn exhaustive_mixed_sign_dominant_loop_reclassifies_to_rux() { let mut loops = detected.loops.clone(); reclassify_loops_from_results(&mut loops, &results, &loop_partitions); + // The relative series of a lone loop is +1/-1/0 exactly, so the + // confidence is the time share of its sign: count the signs from the + // owner's series rather than pinning the fixture doc's approximate + // counts. + let relative = ltm_post::compute_rel_loop_scores(&results, &loop_partitions); + let series = &relative[&loops[0].id]; + assert!( + series.iter().all(|v| *v == 0.0 || v.abs() == 1.0), + "a loop alone in its partition has a relative score of exactly 0 or +/-1" + ); + let positive = series.iter().filter(|v| **v > 0.0).count() as f64; + let negative = series.iter().filter(|v| **v < 0.0).count() as f64; + assert!(positive > 0.0 && negative > 0.0, "both signs occur"); + let time_share = (positive - negative).abs() / (positive + negative); + assert!( + time_share < POLARITY_CONFIDENCE_THRESHOLD, + "nine percent of the active steps balancing is far below the gate: {time_share}" + ); assert_eq!( loops[0].polarity, - DetectedLoopPolarity::MostlyReinforcing, - "a mixed-sign series with >=0.99 positive dominance must classify Rux \ + DetectedLoopPolarity::Undetermined, + "a lone loop balancing for a tenth of its active life is Undetermined \ (got {:?} at confidence {})", loops[0].polarity, loops[0].polarity_confidence ); assert!( - loops[0].polarity_confidence >= POLARITY_CONFIDENCE_THRESHOLD - && loops[0].polarity_confidence < 1.0, - "the Rux confidence is the REAL dominance ratio: at or above the {} gate but \ - strictly below 1.0 (both signs are present in the series); got {}", - POLARITY_CONFIDENCE_THRESHOLD, + (loops[0].polarity_confidence - time_share).abs() < 1e-12, + "the confidence is the time share of the sign, {time_share}; got {}", loops[0].polarity_confidence ); } +/// A stock `s` with an inflow loop and an outflow loop; the loop through +/// `flipping` has gain `g` (+0.02, flipped to -0.02 for t in [100, 102)), +/// the other loop gain `d` (0.01, raised to 2 in that window so it carries +/// the partition while the first loop's sign is wrong). With the inflow +/// flipping the inflow loop is reinforcing-then-briefly-balancing (Rux); +/// with the outflow flipping the outflow loop is the mirror (Bux). +/// +/// Hand calculation, both loops through `s` in one partition, dt = 0.25 +/// over [0, 110] (441 saved steps, N of them active -- all but the startup +/// step or two the flow-to-stock score leaves at zero): outside the window +/// the flipping loop's raw score is 2 and the other's is 1 in magnitude, so +/// the flipping loop's share is 2/3 at every one of its N - 8 same-sign +/// steps; inside the window its raw score is ~0.0099 against the other +/// loop's ~0.99, a share of ~0.0099 at each of its 8 opposite-sign steps. +/// r ~= (N - 8) * 2/3 ~= 288, |b| ~= 8 * 0.0099 = 0.08, confidence +/// (288 - 0.08) / 288.1 = 0.9994 for any N near 440: above the gate and +/// below 1, `Mostly*`. +fn competing_sign_flip_project(flipping: &str) -> simlin_engine::datamodel::Project { + let (inflow_gain, outflow_gain) = if flipping == "inflow" { + ("g", "d") + } else { + ("d", "g") + }; + TestProject::new("competing_sign_flip") + .with_sim_time(0.0, 110.0, 0.25) + .aux( + "g", + "IF TIME < 100 OR TIME >= 102 THEN 0.02 ELSE -0.02", + None, + ) + .aux("d", "IF TIME < 100 OR TIME >= 102 THEN 0.01 ELSE 2", None) + .stock("s", "100", &["f_in"], &["f_out"], None) + .flow("f_in", &format!("s * {inflow_gain}"), None) + .flow("f_out", &format!("s * {outflow_gain}"), None) + .build_datamodel() +} + +/// Reclassify `competing_sign_flip_project(flipping)` on the exhaustive path +/// and return the flipping loop (the one through `f_{flipping}`) with the +/// sign counts of its relative series. +fn reclassified_flipping_loop(flipping: &str) -> (DetectedLoop, usize, usize) { + let project = competing_sign_flip_project(flipping); + let mut db = SimlinDb::default(); + let sync = sync_from_datamodel_incremental(&mut db, &project, None); + let source_model = sync.models["main"].source_model; + let detected = model_detected_loops(&db, source_model, sync.project); + assert_eq!( + detected.loops.len(), + 2, + "one inflow loop and one outflow loop" + ); + let (compiled, loop_partitions) = compile_ltm_incremental_with_partitions(&project); + let mut vm = Vm::new(compiled).unwrap(); + vm.run_to_end().expect("simulation should run"); + let results = vm.into_results(); + let mut loops = detected.loops.clone(); + reclassify_loops_from_results(&mut loops, &results, &loop_partitions); + let relative = ltm_post::compute_rel_loop_scores(&results, &loop_partitions); + let flow = if flipping == "inflow" { + "f_in" + } else { + "f_out" + }; + let l = loops + .iter() + .find(|l| l.variables.iter().any(|v| v == flow)) + .expect("the flipping loop") + .clone(); + let series = &relative[&l.id]; + let positive = series.iter().filter(|v| **v > 0.0).count(); + let negative = series.iter().filter(|v| **v < 0.0).count(); + (l, positive, negative) +} + +/// GH #679, the Rux arm on the relative base: a loop whose wrong-signed +/// steps coincide with a sibling carrying the partition is +/// `MostlyReinforcing` at the hand-calculated ~0.9994 (see +/// `competing_sign_flip_project`). +#[test] +fn exhaustive_competing_sign_flip_reclassifies_to_rux() { + use simlin_engine::ltm::POLARITY_CONFIDENCE_THRESHOLD; + let (l, positive, negative) = reclassified_flipping_loop("inflow"); + assert!( + positive > 400 && negative == 8, + "got {positive} positive, {negative} negative steps" + ); + assert_eq!( + l.polarity, + DetectedLoopPolarity::MostlyReinforcing, + "got {:?} at confidence {}", + l.polarity, + l.polarity_confidence + ); + assert!( + l.polarity_confidence >= POLARITY_CONFIDENCE_THRESHOLD + && l.polarity_confidence < 1.0 + && (l.polarity_confidence - 0.9994).abs() < 1e-3, + "the Rux confidence is the hand-calculated ~0.9994; got {}", + l.polarity_confidence + ); +} + +/// The Bux arm: the mirror fixture (the OUTFLOW loop flips) is +/// `MostlyBalancing` at the same ~0.9994. +#[test] +fn exhaustive_competing_sign_flip_reclassifies_to_bux() { + use simlin_engine::ltm::POLARITY_CONFIDENCE_THRESHOLD; + let (l, positive, negative) = reclassified_flipping_loop("outflow"); + assert!( + negative > 400 && positive == 8, + "got {positive} positive, {negative} negative steps" + ); + assert_eq!( + l.polarity, + DetectedLoopPolarity::MostlyBalancing, + "got {:?} at confidence {}", + l.polarity, + l.polarity_confidence + ); + assert!( + l.polarity_confidence >= POLARITY_CONFIDENCE_THRESHOLD + && l.polarity_confidence < 1.0 + && (l.polarity_confidence - 0.9994).abs() < 1e-3, + "the Bux confidence is the hand-calculated ~0.9994; got {}", + l.polarity_confidence + ); +} + +/// The yeast alcohol model (Schoenberg et al. 2020, section "yeast alcohol +/// model": `b = c*(1.1 - 0.1*a)/b1`, `d = c*EXP(a - 11)/d1`, `da/dt = p*c`, +/// c0 = 1, a0 = 0, b1 = 16, d1 = 30, p = 0.01, dt 0.5 over [0, 100]). +/// Its growth loop `c -> b -> c` is reinforcing for 146 active steps and +/// balancing for 53 once the alcohol level pushes `1.1 - 0.1*a` negative, +/// so it is `Undetermined` on either base; on the relative base the +/// confidence is 0.9655 (r = 99.9, |b| = 1.76 over the partition-relative +/// shares), the reading a dominance-weighted time share gives it. +#[test] +fn yeast_growth_loop_stays_undetermined_on_the_relative_base() { + use simlin_engine::ltm::POLARITY_CONFIDENCE_THRESHOLD; + let project = TestProject::new("yeast") + .with_sim_time(0.0, 100.0, 0.5) + .stock("c", "1", &["b"], &["d"], None) + .stock("a", "0", &["prod"], &[], None) + .aux("b1", "16", None) + .aux("d1", "30", None) + .aux("p", "0.01", None) + .flow("b", "c * (1.1 - 0.1 * a) / b1", None) + .flow("d", "c * EXP(a - 11) / d1", None) + .flow("prod", "p * c", None) + .build_datamodel(); + let mut db = SimlinDb::default(); + let sync = sync_from_datamodel_incremental(&mut db, &project, None); + let source_model = sync.models["main"].source_model; + let detected = model_detected_loops(&db, source_model, sync.project); + let (compiled, loop_partitions) = compile_ltm_incremental_with_partitions(&project); + let mut vm = Vm::new(compiled).unwrap(); + vm.run_to_end().expect("the yeast model simulates"); + let results = vm.into_results(); + let mut loops = detected.loops.clone(); + reclassify_loops_from_results(&mut loops, &results, &loop_partitions); + let relative = ltm_post::compute_rel_loop_scores(&results, &loop_partitions); + + let growth = loops + .iter() + .find(|l| { + let mut vars = l.variables.clone(); + vars.sort(); + vars == ["b", "c"] + }) + .expect("the c -> b -> c growth loop"); + let series = &relative[&growth.id]; + let positive = series.iter().filter(|v| **v > 0.0).count(); + let negative = series.iter().filter(|v| **v < 0.0).count(); + assert_eq!((positive, negative), (146, 53)); + assert_eq!( + growth.polarity, + DetectedLoopPolarity::Undetermined, + "got {:?} at confidence {}", + growth.polarity, + growth.polarity_confidence + ); + assert!( + (growth.polarity_confidence - 0.9655).abs() < 1e-3 + && growth.polarity_confidence < POLARITY_CONFIDENCE_THRESHOLD, + "the relative-base confidence is 0.9655; got {}", + growth.polarity_confidence + ); +} + /// GH #679, the sub-threshold side of the Rux gate: a loop whose score /// spends comparable magnitude on both signs stays `Undetermined` after /// runtime reclassification -- but with a real (non-zero, sub-0.99) @@ -10598,58 +10811,215 @@ fn exhaustive_mixed_sign_balanced_loop_stays_undetermined() { } /// GH #679 parity: discovery mode and the exhaustive runtime-reclassification -/// path must agree on the Rux fixture -- same `MostlyReinforcing` label and -/// the same dominance-ratio confidence. Both classify via -/// `LoopPolarity::from_runtime_scores`; for this single-loop fixture the -/// discovery score series is the same per-step product of -/// link scores the exhaustive `loop_score` variable computes, so the two -/// modes must not disagree about the same model. +/// path classify the same partition-relative series, so they agree on both +/// fixtures -- the lone loop (`Undetermined` at its time share) and the +/// competing Rux loop -- label and confidence alike. The discovery series +/// is the same per-step product of link scores the exhaustive `loop_score` +/// variable computes, normalized against the same partition. #[test] -fn discovery_rux_classification_matches_exhaustive() { - use simlin_engine::ltm::{LoopPolarity, POLARITY_CONFIDENCE_THRESHOLD}; +fn discovery_classification_matches_exhaustive_on_the_relative_base() { + use simlin_engine::ltm::LoopPolarity; - let project = mixed_sign_dominant_loop_project(); + let exhaustive = |project: &simlin_engine::datamodel::Project, flow: &str| { + let mut db = SimlinDb::default(); + let sync = sync_from_datamodel_incremental(&mut db, project, None); + let source_model = sync.models["main"].source_model; + let detected = model_detected_loops(&db, source_model, sync.project); + let (compiled, loop_partitions) = compile_ltm_incremental_with_partitions(project); + let mut vm = Vm::new(compiled).unwrap(); + vm.run_to_end().expect("exhaustive simulation should run"); + let results = vm.into_results(); + let mut loops = detected.loops.clone(); + reclassify_loops_from_results(&mut loops, &results, &loop_partitions); + loops + .into_iter() + .find(|l| l.variables.iter().any(|v| v == flow)) + .expect("the loop through the flow") + }; + let discovered = |project: &simlin_engine::datamodel::Project, flow: &str| { + let (_results, found) = run_discovery(project); + found + .into_iter() + .find(|fl| fl.loop_info.links.iter().any(|l| l.from.as_str() == flow)) + .expect("discovery finds the loop through the flow") + }; + let to_ltm = |p: DetectedLoopPolarity| match p { + DetectedLoopPolarity::Reinforcing => LoopPolarity::Reinforcing, + DetectedLoopPolarity::Balancing => LoopPolarity::Balancing, + DetectedLoopPolarity::MostlyReinforcing => LoopPolarity::MostlyReinforcing, + DetectedLoopPolarity::MostlyBalancing => LoopPolarity::MostlyBalancing, + DetectedLoopPolarity::Undetermined => LoopPolarity::Undetermined, + }; + + for (project, flow, expected) in [ + ( + mixed_sign_dominant_loop_project(), + "f", + LoopPolarity::Undetermined, + ), + ( + competing_sign_flip_project("inflow"), + "f_in", + LoopPolarity::MostlyReinforcing, + ), + ] { + let e = exhaustive(&project, flow); + let d = discovered(&project, flow); + assert_eq!(to_ltm(e.polarity), expected, "exhaustive label on {flow}"); + assert_eq!(d.loop_info.polarity, expected, "discovery label on {flow}"); + // Both modes classify the same per-step product of the same link-score + // series over the same partition total, so the confidences agree to + // float precision -- a loose 1e-6 tolerance allows for the two paths' + // different accumulation order. + assert!( + (d.polarity_confidence - e.polarity_confidence).abs() < 1e-6, + "{flow}: discovery confidence {} must match exhaustive confidence {}", + d.polarity_confidence, + e.polarity_confidence + ); + } +} + +/// The all-slots reading of an A2A loop: `reclassify_loops_from_results` +/// concatenates every element slot's relative series into one sample set, +/// so a loop reinforcing in one element and balancing in another is +/// `Undetermined` -- the loop's sign is not uniform across the array -- +/// even though each element on its own is single-signed and the structural +/// label (the bare co-factor `g` signs `s -> f` Positive) is Reinforcing. +/// Each region is an isolated one-stock loop, so its relative score is +/// exactly `+1` (region `a`, `g = 0.02`) or `-1` (region `b`, `g = -0.02`) +/// at every active step, and the two slots cancel: confidence exactly 0. +#[test] +fn exhaustive_a2a_loop_with_opposite_signed_elements_is_undetermined() { + let project = TestProject::new("a2a_mixed_elements") + .with_sim_time(0.0, 8.0, 1.0) + .named_dimension("region", &["a", "b"]) + .array_stock("s[region]", "100", &["f"], &[], None) + .array_flow("f[region]", "s[region] * g[region]", None) + .array_with_ranges("g[region]", vec![("a", "0.02"), ("b", "-0.02")]) + .build_datamodel(); - // Exhaustive: reclassify the detected loop from the simulated series. let mut db = SimlinDb::default(); let sync = sync_from_datamodel_incremental(&mut db, &project, None); let source_model = sync.models["main"].source_model; let detected = model_detected_loops(&db, source_model, sync.project); + assert_eq!(detected.loops.len(), 1, "one A2A loop over both regions"); + assert_eq!( + detected.loops[0].polarity, + DetectedLoopPolarity::Reinforcing, + "the bare co-factor g signs s -> f Positive, so the structural label is Reinforcing" + ); + let id = detected.loops[0].id.clone(); + let (compiled, loop_partitions) = compile_ltm_incremental_with_partitions(&project); + assert_eq!(loop_partitions[&id].len(), 2, "one slot per region"); let mut vm = Vm::new(compiled).unwrap(); - vm.run_to_end().expect("exhaustive simulation should run"); + vm.run_to_end().expect("simulation should run"); let results = vm.into_results(); - let mut exhaustive_loops = detected.loops.clone(); - reclassify_loops_from_results(&mut exhaustive_loops, &results, &loop_partitions); + + // The owner's series is step-major, `series[t * 2 + k]`: region a in + // slot 0, region b in slot 1. + let relative = ltm_post::compute_rel_loop_scores(&results, &loop_partitions); + let series = &relative[&id]; + assert_eq!(series.len(), 2 * results.step_count); + let mut active_steps = 0; + for t in 0..results.step_count { + let (a, b) = (series[2 * t], series[2 * t + 1]); + assert_eq!( + a == 0.0, + b == 0.0, + "both regions are active at the same steps" + ); + if a != 0.0 { + active_steps += 1; + assert_eq!( + a, 1.0, + "region a is an isolated reinforcing loop at step {t}" + ); + assert_eq!( + b, -1.0, + "region b is an isolated balancing loop at step {t}" + ); + } + } + assert!(active_steps > 0, "the loop is active"); + + let mut loops = detected.loops.clone(); + reclassify_loops_from_results(&mut loops, &results, &loop_partitions); assert_eq!( - exhaustive_loops[0].polarity, - DetectedLoopPolarity::MostlyReinforcing + loops[0].polarity, + DetectedLoopPolarity::Undetermined, + "one reinforcing element and one balancing element read together are \ + Undetermined; got {:?} at confidence {}", + loops[0].polarity, + loops[0].polarity_confidence ); - let exhaustive_confidence = exhaustive_loops[0].polarity_confidence; + assert_eq!( + loops[0].polarity_confidence, 0.0, + "the two slots carry equal mass of opposite sign" + ); +} - // Discovery: run loop discovery over the discovery-mode sim. - let (_discovery_results, found) = run_discovery(&project); - assert_eq!(found.len(), 1, "discovery finds the single feedback loop"); +/// A loop that is never active -- `f = s * g` with `g = 0`, so neither `f` +/// nor `s` ever changes and every link score is exactly zero -- is not +/// reported by discovery at all: enumeration considers only edges active at +/// some step, and `rank_and_filter`'s retention keeps a loop only for a step +/// at which it holds `MIN_CONTRIBUTION` of its partition. So discovery never +/// asks the classifier about a series with no valid sample. The exhaustive +/// surface does report the loop, from the structural analysis, and +/// reclassification leaves that label alone: with no runtime evidence there +/// is nothing to override it with. +#[test] +fn never_active_loop_is_structural_on_the_exhaustive_surface_and_absent_from_discovery() { + let project = TestProject::new("never_active") + .with_sim_time(0.0, 10.0, 1.0) + .aux("g", "0", None) + .stock("s", "100", &["f"], &[], None) + .flow("f", "s * g", None) + .build_datamodel(); - assert_eq!( - found[0].loop_info.polarity, - LoopPolarity::MostlyReinforcing, - "discovery must report the same Rux label exhaustive reclassification does" + let mut db = SimlinDb::default(); + let sync = sync_from_datamodel_incremental(&mut db, &project, None); + let source_model = sync.models["main"].source_model; + let detected = model_detected_loops(&db, source_model, sync.project); + assert_eq!(detected.loops.len(), 1, "the loop exists structurally"); + let structural = ( + detected.loops[0].polarity, + detected.loops[0].polarity_confidence, ); + + let (compiled, loop_partitions) = compile_ltm_incremental_with_partitions(&project); + let mut vm = Vm::new(compiled).unwrap(); + vm.run_to_end().expect("simulation should run"); + let results = vm.into_results(); + let relative = ltm_post::compute_rel_loop_scores(&results, &loop_partitions); assert!( - found[0].polarity_confidence >= POLARITY_CONFIDENCE_THRESHOLD - && found[0].polarity_confidence < 1.0, - "discovery Rux confidence must clear the gate and stay below 1.0; got {}", - found[0].polarity_confidence + relative[&detected.loops[0].id] + .iter() + .all(|v| *v == 0.0 || v.is_nan()), + "the loop's relative series has no valid sample" + ); + let mut loops = detected.loops.clone(); + reclassify_loops_from_results(&mut loops, &results, &loop_partitions); + assert_eq!( + (loops[0].polarity, loops[0].polarity_confidence), + structural, + "no runtime evidence, so the structural label stands" ); - // Both modes classify the same per-step product of the same link-score - // series, so the confidences agree to float precision -- a loose 1e-6 - // tolerance allows for the two paths' different accumulation order. + + let (_results, found) = run_discovery(&project); assert!( - (found[0].polarity_confidence - exhaustive_confidence).abs() < 1e-6, - "discovery confidence {} must match exhaustive confidence {}", - found[0].polarity_confidence, - exhaustive_confidence + found.is_empty(), + "discovery does not report a never-active loop; got {:?}", + found + .iter() + .map(|fl| fl + .loop_info + .links + .iter() + .map(|l| l.from.as_str()) + .collect::>()) + .collect::>() ); } From e5319fbd8879aa8b5ece0cf5d415131115be0b17 Mon Sep 17 00:00:00 2001 From: Bobby Powers Date: Tue, 8 Sep 2026 23:05:26 -0700 Subject: [PATCH 06/10] engine: one owner for what an element subscript names A per-element arm spelled "nyc, young" (or "NYC,Young") simulated every element of its variable as 0 with no diagnostic. The compiler keyed the arm by canonicalizing the whole subscript, which turns the space after the comma into an underscore ("nyc,_young"), a key no combination of canonical element names ever equals, so the arm matched nothing and the element took the arrayed expansion's fabricated zero; the GH #905 advisory accepted the union of a per-part rule and the compiler's whole-string rule, so the spelling one side resolved and the other dropped was reported by neither. CanonicalElementName::from_subscript (and its from_parts twin for a subscript already split by SubscriptIterator) is now the one owner of the key: split on the commas outside double quotes, trim and canonicalize each part, join with ",". Every side of the match uses it -- the arm keys parse_equation builds, the combination keys expand_per_element, lower_arrayed_arms, unfilled_arms, the builtins visitor's per-element instance expansion, the LTM lookup slot classifier and the per-element GF table layout expand against, the conveyor init-list key (delegates), the two advisories -- and the JSON, protobuf and XMILE readers store the canonical spelling. The MDL reader stores the source spelling; the datamodel invariant is only that every consumer re-keys through the owner, stated on it, along with the standing limitation that an element name containing a comma keys like the combination of its pieces. Severity follows the consequence, uniformly across sources. An arm whose subscript names no element of the variable's dimensions is an unused arm: the existing UnknownElementSubscript warning, now by the one key. A declared element that no arm names and no EXCEPT default applies to evaluates to the fabricated 0: a new MissingElementEquation warning on the variable names every such element in row-major order and says they evaluate to 0. Coverage is read off the parsed shape the compiler expands (variable::elements_without_an_arm), never the entry text: an entry with an empty equation is no arm, an empty EXCEPT default covers nothing, and a gf-only entry is an arm because the expansion wraps its table around the fabricated input. It is a warning, not an error, on Vensim ground truth rather than reasoning: Vensim defines a subscripted variable on part of its range as a matter of course (h[DimA] :EXCEPT: [SubA] = 8 defines h[A1] only), and its own export for the sdeverywhere except model (test/sdeverywhere/models/except/except.dat) lists no h[A2] at all -- the element does not exist there, so the 0 the MDL import gives it is a value nothing in a valid Vensim model reads. As an error it refused five corpus tests (except, except2, subscript, directconst and the inline except_basic, whose test now pins the warning naming a2 and a3 on h). The fidelity fix for imported models, the MDL reader declaring a partially defined variable over the subrange its arms cover, is GH #1059. Pinned through the real pipeline: a proptest that any spacing and capitalization of a one- to three-dimensional combination keys the canonical join (and the join is a fixed point); the sliced_sum fixture through Simlin JSON, where "nyc, young" and " Nyc , Young " now store the canonical spelling and simulate bit-exact equal to the canonical run across every saved series; the three arms -- a typo'd arm (one UnknownElementSubscript warning naming nyc,yung and one MissingElementEquation warning naming nyc,young on pop), a dropped la,old arm (the warning names it and pop[la,old] is 0 at every step while its siblings are untouched), an EXCEPT default covering the dropped arm (no finding, simulates), and an arm for an element of another dimension (warning only, simulates); the empty-equation entry, the empty EXCEPT default and the XMILE element with neither equation nor gf (each reported, each simulating as 0), and the gf-only element (unreported, gf(0)); a protobuf-reader pin for a spaced, capitalized subscript. The unfilled-equation test for a sparse array now also asserts that the new warning names its armless slot. RealBeer4 and the rest of the metasd and sdeverywhere corpora load and simulate unchanged. --- src/libsimlin/src/lib.rs | 3 + src/simlin-engine/CLAUDE.md | 1 + src/simlin-engine/src/ast/mod.rs | 2 +- src/simlin-engine/src/builtins_visitor.rs | 18 +- src/simlin-engine/src/common.rs | 176 ++++++++ src/simlin-engine/src/compiler/mod.rs | 12 +- src/simlin-engine/src/conveyor_compile.rs | 14 +- src/simlin-engine/src/db/diagnostic.rs | 207 +++++---- src/simlin-engine/src/db/diagnostic_tests.rs | 25 +- src/simlin-engine/src/json.rs | 27 +- .../src/ltm_augment_with_lookup.rs | 4 +- src/simlin-engine/src/mdl/convert/helpers.rs | 5 +- src/simlin-engine/src/per_element_gf_tests.rs | 21 +- src/simlin-engine/src/serde.rs | 44 +- .../src/unfilled_equation_tests.rs | 32 +- src/simlin-engine/src/variable.rs | 135 ++++-- src/simlin-engine/src/xmile/variables.rs | 6 +- .../tests/integration/element_subscripts.rs | 415 ++++++++++++++++++ src/simlin-engine/tests/integration/main.rs | 1 + .../tests/integration/simulate.rs | 36 +- 20 files changed, 983 insertions(+), 201 deletions(-) create mode 100644 src/simlin-engine/tests/integration/element_subscripts.rs diff --git a/src/libsimlin/src/lib.rs b/src/libsimlin/src/lib.rs index 068b41a7e..114bacb87 100644 --- a/src/libsimlin/src/lib.rs +++ b/src/libsimlin/src/lib.rs @@ -342,6 +342,9 @@ impl From for SimlinErrorCode { // NaN literal) likewise collapses to the wire Generic code; the // message names the variable, which is the whole point of it. engine::ErrorCode::UnfilledEquation => SimlinErrorCode::Generic, + // A declared element with no equation likewise collapses to the + // wire Generic code; the message names the variable and elements. + engine::ErrorCode::MissingElementEquation => SimlinErrorCode::Generic, } } } diff --git a/src/simlin-engine/CLAUDE.md b/src/simlin-engine/CLAUDE.md index ba4caeee8..39c00839c 100644 --- a/src/simlin-engine/CLAUDE.md +++ b/src/simlin-engine/CLAUDE.md @@ -165,6 +165,7 @@ A conveyor or queue stock is not compiled directly. `conveyor_compile::expand_co - `common.rs` -- `ErrorCode`, `Error`, `EquationError`, `Result`; the identifier types. An `EquationError` is `{start, end, code, details}`: the code names the CLASS of failure, the span points at the offending text (which `errors.rs` renders as a source snippet), and `details` carries the reason when it is not visible in the span -- the name that did not resolve, the arity a call missed, the identifier a lowering could not shape. A parse error therefore writes no `details`: the snippet is its reason, and every Rust surface renders that snippet. - Nothing between a raising site and `collect_all_diagnostics` drops the reason or puts something else in its place: `From for EquationError` carries `Error::details` across, the compiler's lowering entry points hand a rejection's `details` to the `Error` they build -- never its span, which is a byte offset into a string the user never sees -- and `UnitError::DefinitionError` wraps an `EquationError` rather than keeping a second reason field beside it. - `canonicalize` lowercases and normalizes a raw name; `Ident`, `CanonicalDimensionName`, and `CanonicalElementName` are interned, so clone is a refcount bump and equality is pointer equality. Sub-model variables are addressed as `module·var` (U+00B7). + - `CanonicalElementName::from_subscript` / `from_parts` is the one owner of what a per-element subscript names (split on commas outside quotes, each part canonicalized, joined with `,`): every side of an element-arm match -- the arm keys, the declared-combination keys, the per-element gf tables, the conveyor init lists, the two arrayed-arm advisories, the readers -- goes through it. Never key one side by canonicalizing the whole subscript. `variable::elements_without_an_arm` is likewise the one owner of which declared elements have no arm (parsed arms plus per-element tables plus a live EXCEPT default). - `diagnostic.rs` -- the one diagnostic payload. `DiagnosticError::{Equation(EquationError), Model(Error), Unit(UnitError), Assembly(String)}` is the typed error exactly as its raising site produced it, and `Diagnostic { model, variable, owner, severity, error }` is that error plus the context it is reported under -- the salsa accumulator, built at the salsa layer's raising sites (the fragment constructors, the per-model advisories, the unit pass, the LTM `ltm_warning`), where the model, the variable and the severity are known; a `Variable` carries only context-free errors and nothing between a site and the drain re-attaches context, so nothing asserts it. - `owner` is the variable a generated helper's row is presented under (the one it was synthesized for: a user variable, or for an LTM helper the synthetic link score); `variable` stays the helper's physical name. - Consumers read `code()`, `category()` (`DiagnosticCategory`: the four arms with `Unit` split by `UnitError`'s three), `location()`, `reason()` and `is(category, code)` rather than matching arms. The unit-inference umbrella is a `Unit(InferenceError)` filed under no variable, like every other unit diagnostic. When several arms of an arrayed equation fail to lower, the arm reported is the first in the dimensions' declared element order (`ast::lower_arrayed_arms`). diff --git a/src/simlin-engine/src/ast/mod.rs b/src/simlin-engine/src/ast/mod.rs index 322ba564e..06afbb527 100644 --- a/src/simlin-engine/src/ast/mod.rs +++ b/src/simlin-engine/src/ast/mod.rs @@ -235,7 +235,7 @@ fn lower_arrayed_arms( return Ok(lowered); } for combination in crate::dimensions::SubscriptIterator::new(dims) { - let key = CanonicalElementName::from_raw(&combination.join(",")); + let key = CanonicalElementName::from_parts(&combination); if let Some(first) = failures.iter().position(|(id, _)| *id == key) { return Err(failures.swap_remove(first).1); } diff --git a/src/simlin-engine/src/builtins_visitor.rs b/src/simlin-engine/src/builtins_visitor.rs index 3e5d30251..93a0eacb1 100644 --- a/src/simlin-engine/src/builtins_visitor.rs +++ b/src/simlin-engine/src/builtins_visitor.rs @@ -1043,7 +1043,7 @@ pub fn instantiate_implicit_modules( if requirements(&ast) == PerElement::ModuleInstance && !dimensions.is_empty() { let mut elements = HashMap::new(); for subscript in SubscriptIterator::new(&dimensions) { - let subscript_key = CanonicalElementName::from_raw(&subscript.join(",")); + let subscript_key = CanonicalElementName::from_parts(&subscript); let mut walker = visitor().with_active_element(&dimensions, &subscript); let transformed = walker.walk(ast.clone())?; collect(walker)?; @@ -1087,11 +1087,8 @@ pub fn instantiate_implicit_modules( // distinct slots never claim one name (PR #668). let mut new_elements = HashMap::new(); for (subscript_key, equation) in elements_in_stable_order(elements) { - let subscript_parts: Vec = subscript_key - .as_str() - .split(',') - .map(|s| s.to_string()) - .collect(); + let subscript_parts: Vec = + subscript_key.parts().map(str::to_string).collect(); let mut walker = visitor().with_active_element(&dimensions, &subscript_parts); let transformed = walker.walk(equation)?; collect(walker)?; @@ -1102,8 +1099,7 @@ pub fn instantiate_implicit_modules( if let Some(default_expr) = default_expr { let missing: Vec> = SubscriptIterator::new(&dimensions) .filter(|subscript| { - !new_elements - .contains_key(&CanonicalElementName::from_raw(&subscript.join(","))) + !new_elements.contains_key(&CanonicalElementName::from_parts(subscript)) }) .collect(); if apply_default_to_missing @@ -1117,10 +1113,8 @@ pub fn instantiate_implicit_modules( let mut walker = visitor().with_active_element(&dimensions, &subscript); let transformed = walker.walk(default_expr.clone())?; collect(walker)?; - new_elements.insert( - CanonicalElementName::from_raw(&subscript.join(",")), - transformed, - ); + new_elements + .insert(CanonicalElementName::from_parts(&subscript), transformed); } apply_default = false; } else { diff --git a/src/simlin-engine/src/common.rs b/src/simlin-engine/src/common.rs index 579b8a1f2..ed81aa2fd 100644 --- a/src/simlin-engine/src/common.rs +++ b/src/simlin-engine/src/common.rs @@ -620,6 +620,23 @@ pub enum ErrorCode { /// diagnostic that asserted otherwise would send the modeller looking in the /// wrong place. UnfilledEquation, + /// A declared element of a non-apply-to-all arrayed variable has no + /// equation: no `` arm names it and no EXCEPT default applies, + /// so the compiler's arrayed expansion assigns it a fabricated zero (the + /// silent wrong number GH #905 left open). Warning-level, naming the + /// elements: Vensim legitimately defines a subscripted variable on part + /// of its range (`h[DimA] :EXCEPT: [SubA] = 8` defines `h[A1]` only, and + /// Vensim's own output -- `test/sdeverywhere/models/except/except.dat` + /// -- has no `h[A2]` at all), so an Error would refuse real models whose + /// fabricated zeros nothing reads, while an API-built model that dropped + /// an arm is told exactly which elements it lost. An arm whose subscript + /// names nothing is the sibling shape, reported by + /// [`ErrorCode::UnknownElementSubscript`]; a typo'd arm leaves its + /// element uncovered, so both fire. The fidelity fix for imported models + /// is for the MDL reader to declare a partially defined variable over the + /// subrange its arms cover, so that no fabricated element exists to warn + /// about (GH #1059). + MissingElementEquation, } impl fmt::Display for ErrorCode { @@ -708,6 +725,7 @@ impl fmt::Display for ErrorCode { UnknownElementSubscript => "unknown_element_subscript", MacroContainsModule => "macro_contains_module", UnfilledEquation => "unfilled_equation", + MissingElementEquation => "missing_element_equation", }; write!(f, "{name}") @@ -2360,6 +2378,77 @@ impl CanonicalElementName { CanonicalElementName(CanonicalStorage::intern(&canonicalize(s))) } + /// The key of a per-element equation's subscript -- `nyc`, `nyc,young`, + /// `"a1,b1"` -- as one comma-joined string of canonical element names: + /// the subscript is split on the commas outside double quotes, each part + /// trimmed and canonicalized, and the parts joined with `,`. + /// + /// This is the ONE owner of what a subscript names. Every side of the + /// match goes through it -- the arm keys the compiler expands + /// (`variable::parse_equation`), the combination keys it expands them + /// against (`compiler::expand_per_element`, `ast::lower_arrayed_arms`, + /// `variable::unfilled_arms`), the per-element graphical-function + /// tables, the conveyor init lists, the unknown-subscript advisory, and + /// every reader that stores a subscript (XMILE, JSON, protobuf) -- so + /// `nyc, young`, `NYC,Young` and `nyc,young` are the same element and a + /// spelling that matches nowhere is reported rather than dropped. + /// Never key one side by canonicalizing the whole string instead: that + /// turns the space after a comma into an underscore (`nyc,_young`), a + /// key no combination of canonical element names ever equals, and the + /// arm is silently ignored. A quoted whole subscript (`"a1,b1"`) is + /// one part whose quotes `canonicalize` strips, so it keys the same + /// two-dimensional combination as `a1,b1`; a comma inside quotes is part + /// of an element name. + /// + /// Readers may store any spelling: the JSON, protobuf and XMILE readers + /// store the canonical key, the MDL reader stores the source spelling + /// (`mdl/convert/helpers.rs`), and every consumer re-keys through this + /// owner, so the datamodel invariant is only that a subscript names its + /// element under this rule. + /// + /// Standing limitation: the key is the parts joined with `,`, so an + /// element whose NAME contains a comma is indistinguishable from the + /// combination of its pieces -- `"ny,c"` on a one-dimensional variable + /// keys as `ny,c`, the same key the pair (`ny`, `c`) would have on a + /// two-dimensional one; `parts` splits it back the same way. + pub fn from_subscript(subscript: &str) -> Self { + let mut parts: Vec = Vec::new(); + let mut current = String::new(); + let mut in_quotes = false; + for ch in subscript.chars() { + match ch { + '"' => { + in_quotes = !in_quotes; + current.push(ch); + } + ',' if !in_quotes => parts.push(std::mem::take(&mut current)), + _ => current.push(ch), + } + } + parts.push(current); + Self::from_parts(&parts) + } + + /// The same key as [`Self::from_subscript`], from a subscript already + /// split into its parts -- one element name per dimension, as + /// `dimensions::SubscriptIterator` enumerates the declared combinations. + /// Each part is trimmed and canonicalized, and the parts joined with `,`. + pub fn from_parts>(parts: &[S]) -> Self { + let key = parts + .iter() + .map(|part| canonicalize(part.as_ref().trim())) + .collect::>() + .join(","); + CanonicalElementName(CanonicalStorage::intern(&key)) + } + + /// The key's parts, one canonical element name per dimension: the + /// inverse of [`Self::from_parts`] under the comma limitation noted on + /// [`Self::from_subscript`]. + pub fn parts(&self) -> impl Iterator { + self.as_str().split(',') + } + /// Get the underlying canonical string pub fn as_str(&self) -> &str { self.0.as_str() @@ -2892,6 +2981,93 @@ mod whitespace_replacement_tests { } } +#[cfg(test)] +mod element_subscript_key_tests { + //! `CanonicalElementName::from_subscript` is the one owner of what a + //! per-element subscript names, so every spelling a reader or an API + //! caller can produce for one element combination must key the same + //! string as the combination's canonical comma-joined element names. + + use super::CanonicalElementName; + use proptest::prelude::*; + + /// A canonical element name as the dimension declares it. + fn element_name() -> impl Strategy { + "[a-z][a-z0-9_]{0,7}".prop_map(|s| s.to_string()) + } + + /// Spell one canonical element name the way a modeller or a tool might: + /// optional spaces on either side, and an upper-cased first letter. + fn spelling(name: &str, pad_left: usize, pad_right: usize, upper: bool) -> String { + let mut spelled = String::new(); + spelled.push_str(&" ".repeat(pad_left)); + if upper { + let mut chars = name.chars(); + if let Some(first) = chars.next() { + spelled.extend(first.to_uppercase()); + spelled.push_str(chars.as_str()); + } + } else { + spelled.push_str(name); + } + spelled.push_str(&" ".repeat(pad_right)); + spelled + } + + proptest! { + /// Any spacing and capitalization of a one-, two- or three-dimensional + /// combination keys the same string as the canonical join -- and the + /// canonical join is a fixed point of the owner. + #[test] + fn every_spelling_of_a_combination_keys_the_canonical_join( + names in prop::collection::vec(element_name(), 1..=3), + pads in prop::collection::vec((0usize..=2, 0usize..=2, any::()), 3), + ) { + let canonical = names.join(","); + let spelled: Vec = names + .iter() + .zip(pads.iter()) + .map(|(name, &(l, r, upper))| spelling(name, l, r, upper)) + .collect::>(); + let spelled = spelled.join(","); + let key = CanonicalElementName::from_subscript(&spelled); + prop_assert_eq!(key.as_str(), canonical.as_str(), "spelling {:?}", spelled); + let fixed_point = CanonicalElementName::from_subscript(&canonical); + prop_assert_eq!(fixed_point.as_str(), canonical.as_str()); + } + } + + /// A quoted whole subscript is one part whose quotes canonicalization + /// strips, so it keys the two-dimensional combination `a1,b1`; a comma + /// inside quotes stays part of the element name. + #[test] + fn quotes_keep_a_comma_inside_an_element_name() { + assert_eq!( + CanonicalElementName::from_subscript("\"a1,b1\"").as_str(), + "a1,b1" + ); + assert_eq!( + CanonicalElementName::from_subscript("\"a,b\", c").as_str(), + "a,b,c" + ); + } + + /// Whole-string canonicalization is the key no combination ever equals: + /// the space after the comma becomes an underscore. Pinned so the owner + /// is never replaced by it. + #[test] + fn whole_string_canonicalization_is_not_the_key() { + assert_eq!( + CanonicalElementName::from_subscript("nyc, young").as_str(), + "nyc,young" + ); + assert_ne!( + CanonicalElementName::from_raw("nyc, young").as_str(), + "nyc,young" + ); + } +} + #[cfg(test)] mod identifier_part_iterator_tests { use super::*; diff --git a/src/simlin-engine/src/compiler/mod.rs b/src/simlin-engine/src/compiler/mod.rs index c0a053406..f8eb9efb2 100644 --- a/src/simlin-engine/src/compiler/mod.rs +++ b/src/simlin-engine/src/compiler/mod.rs @@ -889,8 +889,10 @@ fn check_stock_updates_are_emittable(ast: &[Expr], ident: &str) -> Result<()> { /// expansion's fabricated `Const(0.0)` input and therefore evaluates /// `gf(0)`: the wrap applies uniformly to whatever input the expansion /// produced. Both the old `0` and the new `gf(0)` are silent fabrications -/// (Stella rejects that element shape outright); the zero-fill itself, and -/// a diagnostic for this class, are the open remainder of GH #905. +/// (Stella rejects that element shape outright). Standing constraint: a +/// gf-only element's arm IS its table (`variable::elements_without_an_arm`), +/// so the `MissingElementEquation` advisory that names an ARMLESS element's +/// zero-fill never names it; keep the two in step if either rule moves. /// /// Only the per-element `AssignCurr` nodes the expansion paths emit are /// rewritten; hoisted pre-computations (`AssignTemp`) feed those assignments @@ -1431,8 +1433,8 @@ fn lower_element(ctx: &Context, elem_ctx: &Context, ast: &crate::ast::Expr2) -> /// row-major order, lower the arm it evaluates under that element's active /// subscripts and assign the result to the element's slot. An element with no /// arm at all -- no explicit equation, and an EXCEPT default that does not -/// apply to it -- is assigned a fabricated zero (the open remainder of -/// GH #905). +/// apply to it -- is assigned a fabricated zero, which the +/// `MissingElementEquation` advisory names (GH #905). /// /// Nothing is hoisted here. An element whose expression is or contains an array /// value codegen cannot express in place is left as it lowered, and @@ -1450,7 +1452,7 @@ fn expand_per_element( let active_dims = Arc::<[Dimension]>::from(dims.to_vec()); let mut exprs: Vec = Vec::new(); for (i, subscripts) in SubscriptIterator::new(dims).enumerate() { - let key = CanonicalElementName::from_raw(&subscripts.join(",")); + let key = CanonicalElementName::from_parts(&subscripts); let Some(ast) = arrayed_arm(elements, default_ast, apply_default_for_missing, &key) else { exprs.push(Expr::AssignCurr( base.offset_by(i), diff --git a/src/simlin-engine/src/conveyor_compile.rs b/src/simlin-engine/src/conveyor_compile.rs index 990fe6720..b60b5c19d 100644 --- a/src/simlin-engine/src/conveyor_compile.rs +++ b/src/simlin-engine/src/conveyor_compile.rs @@ -811,16 +811,12 @@ type InitListValues = (Vec, String); /// Canonical comma-joined subscript key for matching an `Equation::Arrayed` /// element entry against the row-major [`element_subscripts_for_dims`] -/// suffixes. Each comma-separated part is canonicalized independently (the -/// XMILE reader already stores element keys this way -- `xmile/variables.rs` -/// `convert_equation` -- but MDL-sourced or hand-built datamodels may not), -/// so both sides normalize identically regardless of case or whitespace. +/// suffixes: the one owner of that key, +/// [`crate::common::CanonicalElementName::from_subscript`], as a `String`. pub(crate) fn canonical_subscript_key(subscript: &str) -> String { - subscript - .split(',') - .map(canon) - .collect::>() - .join(",") + crate::common::CanonicalElementName::from_subscript(subscript) + .as_str() + .to_string() } /// Resolve a conveyor stock's initial equation against the §7.2 explicit-list diff --git a/src/simlin-engine/src/db/diagnostic.rs b/src/simlin-engine/src/db/diagnostic.rs index ee6a4b105..b015235a2 100644 --- a/src/simlin-engine/src/db/diagnostic.rs +++ b/src/simlin-engine/src/db/diagnostic.rs @@ -83,11 +83,12 @@ use crate::common::{Error, UnitError}; /// time that is not an integer multiple of dt) and /// `ConveyorLeakFractionsExceedOne` (constant linear leak fractions /// summing above 1), docs/design/conveyors.md §4.1 / §5.1. -/// 4b. `emit_unknown_element_subscript_warnings` -- the unconditional -/// `UnknownElementSubscript` advisory: a non-apply-to-all arrayed -/// variable's `` entry whose subscript names no declared -/// element combination is silently dropped everywhere downstream, so a -/// typo'd subscript simulates plausibly-but-wrong with no signal (GH #905). +/// 4b. `emit_element_subscript_warnings` -- the unconditional arrayed-arm +/// advisories (GH #905): `UnknownElementSubscript`, a non-apply-to-all +/// arrayed variable's `` entry whose subscript names no declared +/// element combination and is dropped everywhere downstream; and +/// `MissingElementEquation`, a declared element no entry names and no +/// EXCEPT default applies to, which the compiler assigns a fabricated 0. /// 4c. `emit_unfilled_equation_warnings` -- the unconditional /// `UnfilledEquation` advisory: a variable whose equation is nothing but /// the NaN literal has no usable equation, so it simulates as NaN and that @@ -283,7 +284,7 @@ pub fn model_all_diagnostics( // by every downstream consumer; surface it as a Warning here so the // typo'd equation does not vanish without a trace. Unconditional, like // the conveyor spec advisories above. - emit_unknown_element_subscript_warnings(db, model, project); + emit_element_subscript_warnings(db, model, project); // Unfilled equations: a variable whose equation is nothing but the NaN // literal, which is where Vensim's `A FUNCTION OF(...)` placeholder lands. @@ -692,7 +693,9 @@ fn emit_conveyor_spec_warnings(db: &dyn Db, model: SourceModel, project: SourceP } } -/// Emit one Warning-severity [`crate::common::ErrorCode::UnknownElementSubscript`] +/// Emit the two arrayed-arm advisories for every non-apply-to-all arrayed +/// variable: one Warning-severity +/// [`crate::common::ErrorCode::UnknownElementSubscript`] /// diagnostic per distinct non-apply-to-all `` entry in `model` whose /// subscript names no declared element combination of the variable's /// dimensions (GH #905). @@ -709,64 +712,49 @@ fn emit_conveyor_spec_warnings(db: &dyn Db, model: SourceModel, project: SourceP /// one-character typo therefore produces a plausible-looking but wrong /// simulation with no signal anywhere. /// -/// An entry is flagged only when BOTH production matching rules fail, i.e. -/// this check accepts the UNION of what any consumer accepts: +/// An entry is flagged exactly when the ONE key every consumer matches by +/// -- [`CanonicalElementName::from_subscript`], the compiler's arm keys, the +/// GF-table layout, the conveyor init lists -- equals no declared +/// combination's key. There is no second rule to disagree with the +/// consumers: when this warns, every consumer really does drop the entry +/// and "the equation is ignored" is accurate, and an entry it accepts is +/// one the compiler expands. The owner splits on commas outside quotes +/// and canonicalizes each part, so whitespace and case variants +/// (`a1, b1`), quoted whole subscripts (`"a1,b1"`) and comma-containing +/// element names all resolve. /// -/// 1. **Per-part** (the conveyor init-list matcher, -/// [`crate::conveyor_compile::canonical_subscript_key`] semantics): split -/// on `,`, canonicalize each part, and resolve it against the -/// corresponding dimension via `Dimension::get_offset` (named elements by -/// canonical name; indexed dimensions by parsed `1..=size`). This accepts -/// whitespace-around-comma variants ("a1, b1") that rule 2 mangles. -/// 2. **Whole-string** (the plain compiler's rule -- `variable.rs` -/// `parse_equation` keys entries by -/// `CanonicalElementName::from_raw(subscript)` and -/// `compiler::expand_per_element` looks up -/// `from_raw(combination.join(","))` -- and equally the per-element -/// GF-table layout, `variable::build_tables` / -/// `reorder_arrayed_element_tables`): the entry's whole canonicalized -/// subscript equals some declared combination's comma-joined key. This -/// accepts comma-CONTAINING element names (`board{"a,b", "c"}`, entry -/// `a,b`) and quoted whole subscripts (`"a1,b1"`, whose balanced quotes -/// `canonicalize` strips) that rule 1 mis-splits -- entries the compiler -/// genuinely resolves and simulates. +/// Indexed-dimension nuance: the combination keys are the exact +/// `SubscriptIterator` strings (`"1"`), so an alternate numeric spelling +/// like `"01"` matches nothing and IS warned -- which is also what every +/// consumer does with it. /// -/// So an entry that ANY consumer resolves is never flagged; when this warns, -/// every consumer really does drop the entry and "the equation is ignored" -/// is accurate. (The rule-1-vs-rule-2 split is a pre-existing inconsistency -/// between the conveyor path and the compiler/GF paths; XMILE/MDL-sourced -/// keys are normalized per-part by their readers, so the divergent spellings -/// are only reachable from API-built datamodels.) +/// The complementary GH #905 shape is the second advisory: one Warning- +/// severity [`crate::common::ErrorCode::MissingElementEquation`] per variable +/// naming, in row-major order, every declared element combination no arm +/// covers -- the elements `compiler::expand_per_element` assigns a fabricated +/// 0. Coverage is `variable::elements_without_an_arm` over the PARSED +/// equation the compiler expands, not the entry list: an entry with an empty +/// equation is no arm, an empty EXCEPT default covers nothing, and a gf-only +/// entry is an arm. A variable carrying a fatal diagnostic (an arm that does +/// not parse, say) is skipped: it never evaluates, so its own error is the +/// report. It reads the same keys as the first advisory, so an entry that +/// one accepts covers its element and a typo'd entry fires both. /// -/// Indexed-dimension nuance: rule 1's `get_offset` PARSES the part as an -/// integer, so alternate numeric spellings like `"01"` are accepted here -/// while every consumer matches the exact `SubscriptIterator` string key -/// (`"1"`) and silently drops them. On indexed dimensions this diagnostic is -/// therefore deliberately MORE forgiving than all consumers -- an `"01"` -/// entry is silently dropped and NOT warned. That is the accepted lenient -/// direction: a missed warning (false negative), never a false claim that a -/// used equation is ignored. -/// -/// Warning -- not Error -- severity, deliberately: the vendored corpus -/// contains real imported models that currently rely on the silent drop -/// (e.g. `test/metasd/beer-game/RealBeer4-Sterman13.mdl`, where the MDL -/// importer's synthesized net-flow variable for a stock defined piecewise -/// over subranges carries parent-range element entries against -/// subrange-typed dimensions), so an Error would newly reject previously -/// loadable projects. The complementary GH #905 shape -- a DECLARED element -/// with no entry and no default, which silently evaluates to 0 -- remains -/// undiagnosed here. +/// Both are Warnings, not Errors, deliberately. Vensim defines a +/// subscripted variable on part of its range as a matter of course +/// (`h[DimA] :EXCEPT: [SubA] = 8` defines `h[A1]` only), and its own output +/// (`test/sdeverywhere/models/except/except.dat`) lists no `h[A2]`: the +/// element does not exist there, so the fabricated 0 the MDL import gives it +/// is a value nothing in a valid Vensim model reads. An Error would refuse +/// those models; a Warning tells an API-built model exactly which elements +/// it left without an equation. /// /// An unresolvable dimension NAME is skipped entirely: the equation already /// surfaces `BadDimensionName` through `compile_var_fragment`, and cascading /// one warning per entry on top of it would be noise. Variables are visited /// in sorted-name order and duplicate unknown keys deduplicated so /// accumulation is deterministic. -fn emit_unknown_element_subscript_warnings( - db: &dyn Db, - model: SourceModel, - project: SourceProject, -) { +fn emit_element_subscript_warnings(db: &dyn Db, model: SourceModel, project: SourceProject) { use crate::common::{CanonicalElementName, ErrorCode, ErrorKind}; use std::collections::HashSet; @@ -779,11 +767,11 @@ fn emit_unknown_element_subscript_warnings( for var_name in var_names { let svar = &source_vars[var_name]; - let crate::datamodel::Equation::Arrayed(dim_names, elements, _, _) = svar.equation(db) - else { + let equation = svar.equation(db); + let crate::datamodel::Equation::Arrayed(dim_names, elements, _, _) = equation else { continue; }; - if dim_names.is_empty() || elements.is_empty() { + if dim_names.is_empty() { continue; } @@ -793,53 +781,28 @@ fn emit_unknown_element_subscript_warnings( continue; }; - // Rule-2 declared-combination key set (see the rustdoc), materialized - // lazily: only a variable with at least one rule-1 failure pays the - // cross-product cost, and the compiler expands the same product for - // every arrayed variable anyway. - let mut whole_string_keys: Option> = None; + // The declared combinations' keys -- the same keys + // `compiler::expand_per_element` expands the arms against -- + // materialized once per arrayed variable. + let combination_keys: HashSet = + crate::dimensions::SubscriptIterator::new(&dims) + .map(|combination| CanonicalElementName::from_parts(&combination)) + .collect(); - let mut warned: HashSet = HashSet::new(); + let mut warned: HashSet = HashSet::new(); for (subscript, _, _, _) in elements { - // Rule 1 (per-part, the conveyor matcher): membership in the - // cross product of the dimensions' elements is per-position - // membership, so no combination set needs materializing. - let parts: Vec = subscript - .split(',') - .map(CanonicalElementName::from_raw) - .collect(); - let per_part_matches = parts.len() == dims.len() - && parts - .iter() - .zip(dims.iter()) - .all(|(part, dim)| dim.get_offset(part).is_some()); - if per_part_matches { + let key = CanonicalElementName::from_subscript(subscript); + if combination_keys.contains(&key) { continue; } - // Rule 2 (whole-string, the compiler/GF matcher): the entry's - // whole canonicalized subscript equals some declared - // combination's comma-joined key -- exactly how - // `compiler::expand_per_element` resolves entries, so - // e.g. a comma-containing element name or a quoted whole - // subscript that rule 1 mis-splits is recognized as resolved. - let whole_string_keys = whole_string_keys.get_or_insert_with(|| { - crate::dimensions::SubscriptIterator::new(&dims) - .map(|combination| CanonicalElementName::from_raw(&combination.join(","))) - .collect() - }); - if whole_string_keys.contains(&CanonicalElementName::from_raw(subscript)) { - continue; - } - let canonical_key = crate::conveyor_compile::canonical_subscript_key(subscript); - if !warned.insert(canonical_key) { + if !warned.insert(key) { continue; } let dims_display = dim_names.join(", "); let msg = format!( "array variable '{var_name}' has an equation for element '{subscript}', but \ '{subscript}' does not name an element of its dimension(s) [{dims_display}]; \ - the equation is ignored, and any declared element without its own equation \ - silently evaluates to 0" + the equation is ignored" ); Diagnostic { model: model_name.clone(), @@ -854,6 +817,58 @@ fn emit_unknown_element_subscript_warnings( } .accumulate(db); } + + // The declared elements no arm covers: `compiler::arrayed_arm` gives + // each of these the fabricated 0. Read off the parsed equation (a memo + // read, see `parsed_equation_ast`) and the per-element table layout, + // the two shapes the compiler expands. One Warning per variable, + // elements in row-major order, so a run reports the same text every + // time. + let parsed = parse_source_variable(db, *svar, project); + let Some(ast) = parsed.variable.ast() else { + continue; + }; + // A variable that will not compile never evaluates, so there is no + // fabricated 0 to name: an arm that does not parse is dropped from the + // map like an empty one, but it is reported by its own parse error, + // and "evaluates to 0" would be false beside it. + if parsed.variable.fatal_diagnostics().next().is_some() { + continue; + } + let tables = parsed.variable.tables(); + let per_element_tables = crate::variable::has_per_element_tables(equation); + let missing: Vec = crate::variable::elements_without_an_arm(ast, |offset| { + per_element_tables && tables.get(offset).is_some_and(|t| !t.x.is_empty()) + }) + .iter() + .map(|key| format!("'{}'", key.as_str())) + .collect(); + if missing.is_empty() { + continue; + } + let msg = format!( + "array variable '{var_name}' has no equation for {}: no element entry names \ + {} and no default equation applies, so {} to 0", + missing.join(", "), + if missing.len() == 1 { "it" } else { "them" }, + if missing.len() == 1 { + "it evaluates" + } else { + "they evaluate" + }, + ); + Diagnostic { + model: model_name.clone(), + variable: Some(var_name.clone()), + owner: None, + severity: DiagnosticSeverity::Warning, + error: DiagnosticError::Model(Error::new( + ErrorKind::Model, + ErrorCode::MissingElementEquation, + Some(msg), + )), + } + .accumulate(db); } } diff --git a/src/simlin-engine/src/db/diagnostic_tests.rs b/src/simlin-engine/src/db/diagnostic_tests.rs index e59003b17..b4f19520d 100644 --- a/src/simlin-engine/src/db/diagnostic_tests.rs +++ b/src/simlin-engine/src/db/diagnostic_tests.rs @@ -2420,13 +2420,12 @@ fn test_unknown_element_subscript_no_warning_for_canonical_variants() { ); } -/// Adversarial counterexample 1: an element NAME containing a comma. The -/// compiler matches the whole canonicalized subscript string against the -/// declared combination's comma-joined key, so an entry for the literal -/// element "a,b" of a one-dimensional variable RESOLVES and its equation is -/// used by the simulation. The per-part matcher alone would mis-split it -/// into two parts and flag it -- with a message falsely claiming the -/// equation is ignored. Must not warn. +/// An element NAME containing a comma: the key owner +/// (`CanonicalElementName::from_subscript`) joins canonical parts with `,`, +/// so the entry `a,b` and the declared element `a,b` key identically, the +/// compiler resolves the entry, and the advisory -- which matches by the same +/// key -- must not warn (the comma ambiguity is the owner's documented +/// limitation, not a matching rule of its own). #[test] fn test_unknown_element_subscript_comma_element_name_not_flagged() { let db = SimlinDb::default(); @@ -2448,13 +2447,11 @@ fn test_unknown_element_subscript_comma_element_name_not_flagged() { ); } -/// Adversarial counterexample 2: a QUOTED whole subscript on a -/// two-dimensional variable. `canonicalize` strips balanced quotes, so the -/// whole-string key of `"a1,b1"` equals the declared combination `a1,b1` -/// and the compiler resolves the entry. The per-part split would leave -/// unbalanced quote characters on each half and flag it. Must not warn. -/// (Only reachable from API-built datamodels -- both file readers normalize -/// per-part on import.) +/// A QUOTED whole subscript on a two-dimensional variable: the owner splits +/// only on commas outside quotes, so `"a1,b1"` is one part whose quotes +/// `canonicalize` strips, keying the declared combination `a1,b1`; the +/// compiler resolves the entry and the advisory must not warn. (Reachable +/// from API-built datamodels; the file readers store the canonical key.) #[test] fn test_unknown_element_subscript_quoted_whole_subscript_not_flagged() { let db = SimlinDb::default(); diff --git a/src/simlin-engine/src/json.rs b/src/simlin-engine/src/json.rs index bb6523dee..e51119d95 100644 --- a/src/simlin-engine/src/json.rs +++ b/src/simlin-engine/src/json.rs @@ -878,7 +878,14 @@ impl From for datamodel::Stock { .into_iter() .map(|ee| { ( - ee.subscript, + // The datamodel stores the element key's + // canonical spelling, the one owner of + // what a subscript names (`from_subscript`). + crate::common::CanonicalElementName::from_subscript( + &ee.subscript, + ) + .as_str() + .to_string(), ee.equation, ee.compat .and_then(|c| c.active_initial) @@ -966,7 +973,14 @@ impl From for datamodel::Flow { .into_iter() .map(|ee| { ( - ee.subscript, + // The datamodel stores the element key's + // canonical spelling, the one owner of + // what a subscript names (`from_subscript`). + crate::common::CanonicalElementName::from_subscript( + &ee.subscript, + ) + .as_str() + .to_string(), ee.equation, ee.compat .and_then(|c| c.active_initial) @@ -1049,7 +1063,14 @@ impl From for datamodel::Aux { .into_iter() .map(|ee| { ( - ee.subscript, + // The datamodel stores the element key's + // canonical spelling, the one owner of + // what a subscript names (`from_subscript`). + crate::common::CanonicalElementName::from_subscript( + &ee.subscript, + ) + .as_str() + .to_string(), ee.equation, ee.compat .and_then(|c| c.active_initial) diff --git a/src/simlin-engine/src/ltm_augment_with_lookup.rs b/src/simlin-engine/src/ltm_augment_with_lookup.rs index 33361c3ad..d99b3b955 100644 --- a/src/simlin-engine/src/ltm_augment_with_lookup.rs +++ b/src/simlin-engine/src/ltm_augment_with_lookup.rs @@ -302,9 +302,7 @@ impl WithLookupSlotRefs { crate::dimensions::SubscriptIterator::new(target_ast_dims) .enumerate() .filter(|(offset, _)| tables.get(*offset).is_some_and(|t| !t.x.is_empty())) - .map(|(_, subscripts)| { - crate::common::CanonicalElementName::from_raw(&subscripts.join(",")) - }) + .map(|(_, subscripts)| crate::common::CanonicalElementName::from_parts(&subscripts)) .collect(); SlotRefKind::PerElement { quoted, with_table } } diff --git a/src/simlin-engine/src/mdl/convert/helpers.rs b/src/simlin-engine/src/mdl/convert/helpers.rs index 94ad96413..4a4702389 100644 --- a/src/simlin-engine/src/mdl/convert/helpers.rs +++ b/src/simlin-engine/src/mdl/convert/helpers.rs @@ -194,7 +194,10 @@ pub(super) fn cartesian_product(dim_elements: &[Vec]) -> Vec { result = new_result; } - // Join each combination into a comma-separated string (no spaces - compiler expects "a,b" not "a, b") + // Join each combination into a comma-separated string, in the source's + // spelling: every consumer re-keys a stored subscript through + // `CanonicalElementName::from_subscript`, so the reader owes it only a + // subscript that names its element under that rule. result.into_iter().map(|combo| combo.join(",")).collect() } diff --git a/src/simlin-engine/src/per_element_gf_tests.rs b/src/simlin-engine/src/per_element_gf_tests.rs index 86d518e0e..300b0644c 100644 --- a/src/simlin-engine/src/per_element_gf_tests.rs +++ b/src/simlin-engine/src/per_element_gf_tests.rs @@ -1022,8 +1022,9 @@ fn zero_point_gf_on_scalar_keeps_raw_input_equation() { /// expansion's fabricated `Const(0.0)`, and the wrap applies uniformly, so /// the element evaluates `gf(0)` (pre-#909 it evaluated the bare fabricated /// 0). Both are silent fabrications -- Stella rejects this element shape -- -/// and the zero-fill diagnostic is the open remainder of GH #905; this pins -/// the uniform-wrap choice. +/// and a gf-only element's arm is its table, so the `MissingElementEquation` +/// advisory never names its zero-fill (`variable::elements_without_an_arm`); +/// this pins the uniform-wrap choice and that silence. #[test] fn gf_only_element_without_default_evaluates_gf_of_fabricated_zero() { // Z has a real input equation (time); A is gf-only (empty eqn) with NO @@ -1058,6 +1059,22 @@ fn gf_only_element_without_default_evaluates_gf_of_fabricated_zero() { would mean an empty placeholder was consulted)", get("a") ); + + // The gf IS the element's arm: the entry names `A` and the table decides + // its value, so the missing-element advisory that names an armless slot + // must not name it. + let mut db = SimlinDb::default(); + let sync = sync_from_datamodel_incremental(&mut db, &project, None); + let diagnostics = + crate::db::collect_all_diagnostics(&db, sync.project, crate::db::LtmOverlay::Off); + assert!( + !diagnostics.iter().any(|d| matches!( + &d.error, + crate::db::DiagnosticError::Model(e) + if e.code == crate::common::ErrorCode::MissingElementEquation + )), + "a gf-only element has an arm and is not reported as missing: {diagnostics:?}" + ); } /// Like [`arrayed_gf_project_with_elem_eqn`] but with a PER-element equation diff --git a/src/simlin-engine/src/serde.rs b/src/simlin-engine/src/serde.rs index 17357c5e3..b20d917f2 100644 --- a/src/simlin-engine/src/serde.rs +++ b/src/simlin-engine/src/serde.rs @@ -587,7 +587,14 @@ impl From for Equation { .into_iter() .map(|e| { ( - migrate_stored_ident(e.subscript), + // Migrated like the declaration (#690), then + // stored as the element key's canonical + // spelling (`from_subscript`, the one owner). + crate::common::CanonicalElementName::from_subscript( + &migrate_stored_ident(e.subscript), + ) + .as_str() + .to_string(), e.equation, e.initial_equation, e.gf.map(GraphicalFunction::from), @@ -648,6 +655,41 @@ fn test_has_except_default_proto_roundtrip() { assert_eq!(roundtripped, eq); } +/// The protobuf reader stores a per-element subscript as the element key's +/// canonical spelling (`CanonicalElementName::from_subscript`), like the JSON +/// and XMILE readers: a stored `"NYC, Young"` reads back as `nyc,young`. +#[test] +fn arrayed_element_subscripts_deserialize_to_the_canonical_key() { + let proto = project_io::variable::Equation { + equation: Some(project_io::variable::equation::Equation::Arrayed( + project_io::variable::ArrayedEquation { + dimension_names: vec!["Region".to_string(), "Age".to_string()], + elements: vec![ + project_io::variable::arrayed_equation::Element { + subscript: "NYC, Young".to_string(), + equation: "100".to_string(), + initial_equation: None, + gf: None, + }, + project_io::variable::arrayed_equation::Element { + subscript: " boston ,old".to_string(), + equation: "80".to_string(), + initial_equation: None, + gf: None, + }, + ], + default_equation: None, + has_except_default: None, + }, + )), + }; + let Equation::Arrayed(_, elements, _, _) = Equation::from(proto) else { + panic!("expected Arrayed"); + }; + let keys: Vec<&str> = elements.iter().map(|(k, _, _, _)| k.as_str()).collect(); + assert_eq!(keys, ["nyc,young", "boston,old"]); +} + #[test] fn test_has_except_default_absent_defaults_to_false() { let proto = project_io::variable::Equation { diff --git a/src/simlin-engine/src/unfilled_equation_tests.rs b/src/simlin-engine/src/unfilled_equation_tests.rs index ccfadd9dd..aae161dfb 100644 --- a/src/simlin-engine/src/unfilled_equation_tests.rs +++ b/src/simlin-engine/src/unfilled_equation_tests.rs @@ -209,7 +209,7 @@ fn decision_cell(ast: &Ast) -> String { let mut unfilled = 0usize; let mut uncovered = false; for combination in crate::dimensions::SubscriptIterator::new(dims) { - let key = CanonicalElementName::from_raw(&combination.join(",")); + let key = CanonicalElementName::from_parts(&combination); match elements.get(&key) { Some(e) => { with_arm += 1; @@ -839,15 +839,13 @@ fn an_unreachable_nan_default_is_not_reported() { } /// A SPARSE array whose every listed arm is unfilled is not WHOLLY unfilled: -/// the slots with no arm compile to a finite `0`, not to NaN. +/// the slot with no arm compiles to a finite `0`, not to NaN. /// /// `sparse[c]` has no entry and no default, so it is 0 while `[a]` and `[b]` are /// NaN. Before the coverage axis this said "variable 'sparse' has no equation", /// which reads as every slot being NaN. It now names the two arms that are. -/// -/// (That silent `0` for an armless slot is its own reportable shape and a -/// deliberately separate one -- GH #905 covers it -- so this test asserts the -/// value rather than expecting a second finding.) +/// The armless slot is a separate finding of its own: the +/// `MissingElementEquation` advisory names `c` as evaluating to 0. #[test] fn a_sparse_array_of_unfilled_arms_names_the_arms_not_the_variable() { let project = read_xmile( @@ -869,6 +867,28 @@ fn a_sparse_array_of_unfilled_arms_names_the_arms_not_the_variable() { has no equation: {:?}", findings[0].2 ); + let missing: Vec<(String, String)> = diagnostics(&project) + .into_iter() + .filter_map(|d| match &d.error { + DiagnosticError::Model(e) if e.code == ErrorCode::MissingElementEquation => { + assert_eq!(DiagnosticSeverity::Warning, d.severity); + Some(( + d.variable.clone().unwrap_or_default(), + e.get_details().unwrap_or_default().to_string(), + )) + } + _ => None, + }) + .collect(); + assert_eq!(1, missing.len(), "one armless slot: {missing:#?}"); + assert_eq!("sparse", missing[0].0); + assert!( + missing[0] + .1 + .starts_with("array variable 'sparse' has no equation for 'c'"), + "the warning names the armless element: {:?}", + missing[0].1 + ); assert!(final_value(&project, "sparse[a]").is_nan()); assert!(final_value(&project, "sparse[b]").is_nan()); assert_eq!( diff --git a/src/simlin-engine/src/variable.rs b/src/simlin-engine/src/variable.rs index fef566f02..3340262c0 100644 --- a/src/simlin-engine/src/variable.rs +++ b/src/simlin-engine/src/variable.rs @@ -334,12 +334,27 @@ pub(crate) fn reorder_arrayed_element_tables( ) -> Vec { crate::dimensions::SubscriptIterator::new(dims) .map(|subscripts| { - let key = CanonicalElementName::from_raw(&subscripts.join(",")); + let key = CanonicalElementName::from_parts(&subscripts); present.get(&key).map(&clone_table).unwrap_or_else(&empty) }) .collect() } +/// Whether an equation's graphical-function tables are laid out PER ELEMENT +/// -- one slot per declared combination in row-major order, an empty +/// placeholder where the element has no gf (`reorder_arrayed_element_tables`) +/// -- rather than one variable-level table. The layout [`build_tables`] +/// chooses when any entry of an arrayed equation carries its own gf; the one +/// rule, read by the missing-element advisory to tell a gf-only element's +/// table from a variable-level one. +pub(crate) fn has_per_element_tables(equation: &datamodel::Equation) -> bool { + matches!( + equation, + datamodel::Equation::Arrayed(_, elements, _, _) + if elements.iter().any(|(_, _, _, gf)| gf.is_some()) + ) +} + /// Build the tables vector from equation and variable-level gf. /// For arrayed variables with per-element gfs, tables are built from each /// element and laid out by the element's declared dimension index (via @@ -357,46 +372,43 @@ fn build_tables( let mut errors = Vec::new(); // Check for per-element gfs in arrayed equation - if let datamodel::Equation::Arrayed(dim_names, elements, _, _) = equation { - let has_element_gfs = elements.iter().any(|(_, _, _, gf)| gf.is_some()); - if has_element_gfs { - // Parse each element's table, keyed by the element's canonical - // (comma-joined) subscript name. Elements without a GF are simply - // absent from the map and get an empty placeholder at their slot. - let mut present: HashMap = HashMap::new(); - for (subscript, _, _, elem_gf) in elements { - match parse_table(elem_gf.as_ref()) { - Ok(Some(table)) => { - present.insert(CanonicalElementName::from_raw(subscript), table); - } - Ok(None) => {} - Err(err) => errors.push(err), + if let datamodel::Equation::Arrayed(dim_names, elements, _, _) = equation + && has_per_element_tables(equation) + { + // Parse each element's table, keyed by the element's canonical + // (comma-joined) subscript name. Elements without a GF are simply + // absent from the map and get an empty placeholder at their slot. + let mut present: HashMap = HashMap::new(); + for (subscript, _, _, elem_gf) in elements { + match parse_table(elem_gf.as_ref()) { + Ok(Some(table)) => { + present.insert(CanonicalElementName::from_subscript(subscript), table); } + Ok(None) => {} + Err(err) => errors.push(err), } - - // Resolve the equation's dimensions so the reorder maps each - // element name to its row-major declared-order flat offset. If the - // dimensions cannot be resolved (a separate BadDimensionName error - // the model already surfaces), fall back to the original - // Vec-positional layout rather than dropping tables. - let tables = match get_dimensions(dimensions, dim_names) { - Ok(dims) => { - reorder_arrayed_element_tables(&dims, &present, Table::empty, |t: &Table| { - t.clone() - }) - } - Err(_) => elements - .iter() - .map(|(subscript, _, _, _)| { - present - .get(&CanonicalElementName::from_raw(subscript)) - .cloned() - .unwrap_or_else(Table::empty) - }) - .collect(), - }; - return (tables, errors); } + + // Resolve the equation's dimensions so the reorder maps each + // element name to its row-major declared-order flat offset. If the + // dimensions cannot be resolved (a separate BadDimensionName error + // the model already surfaces), fall back to the original + // Vec-positional layout rather than dropping tables. + let tables = match get_dimensions(dimensions, dim_names) { + Ok(dims) => { + reorder_arrayed_element_tables(&dims, &present, Table::empty, |t: &Table| t.clone()) + } + Err(_) => elements + .iter() + .map(|(subscript, _, _, _)| { + present + .get(&CanonicalElementName::from_subscript(subscript)) + .cloned() + .unwrap_or_else(Table::empty) + }) + .collect(), + }; + return (tables, errors); } // Fall back to variable-level gf @@ -550,7 +562,7 @@ pub(crate) fn unfilled_arms(ast: &Ast) -> Option { let mut slots_with_an_arm = 0usize; let mut default_is_selected = false; for combination in crate::dimensions::SubscriptIterator::new(dims) { - let key = CanonicalElementName::from_raw(&combination.join(",")); + let key = CanonicalElementName::from_parts(&combination); match elements.get(&key) { Some(expr) => { slots_with_an_arm += 1; @@ -567,9 +579,9 @@ pub(crate) fn unfilled_arms(ast: &Ast) -> Option { let default_unfilled = default_is_selected && *apply_default_to_missing && default.as_ref().is_some_and(is_nan_constant); - // The silent `0` is finite, so a slot that falls to it is NOT an - // unfilled equation. (It is its own reportable shape, and a - // deliberately separate one: GH #905.) + // The fabricated `0` is finite, so a slot that falls to it is NOT an + // unfilled equation. (It is reported separately, by the + // `MissingElementEquation` advisory.) let slots_past_the_arms_are_nan = !default_is_selected || default_unfilled; if unfilled.is_empty() && !default_unfilled { @@ -587,6 +599,43 @@ pub(crate) fn unfilled_arms(ast: &Ast) -> Option { } } +/// The declared element combinations of a PARSED arrayed equation that no arm +/// covers -- the slots `compiler::expand_per_element` assigns the fabricated +/// 0 -- in row-major declared order; empty for a scalar or apply-to-all +/// equation, and for an arrayed one whose EXCEPT default is live. +/// +/// Like [`unfilled_arms`] this reads the parsed `Ast`, never the datamodel's +/// entry list, for the same reason: an entry whose equation is empty is +/// dropped by [`parse_equation`] and takes the fabricated 0 exactly as an +/// absent entry does, and an EXCEPT default that is the empty string parses +/// to `None` and covers nothing. (An entry that does NOT parse is dropped the +/// same way, but its variable carries a fatal parse error and never +/// evaluates; the caller skips such a variable rather than naming a 0 it +/// will not produce.) The one arm the `Ast` does not +/// carry is a graphical function: an element whose entry has its own gf has +/// an arm even when its equation is empty -- the expansion wraps the table +/// around the fabricated 0 input (`gf(0)`, pinned by +/// `gf_only_element_without_default_evaluates_gf_of_fabricated_zero`) -- so +/// `per_element_table_at(offset)`, the per-element table layout +/// [`build_tables`] built ([`has_per_element_tables`]), counts as coverage. +pub(crate) fn elements_without_an_arm( + ast: &Ast, + per_element_table_at: impl Fn(usize) -> bool, +) -> Vec { + let Ast::Arrayed(dims, elements, default, apply_default_to_missing) = ast else { + return vec![]; + }; + if *apply_default_to_missing && default.is_some() { + return vec![]; + } + crate::dimensions::SubscriptIterator::new(dims) + .enumerate() + .map(|(offset, combination)| (offset, CanonicalElementName::from_parts(&combination))) + .filter(|(offset, key)| !elements.contains_key(key) && !per_element_table_at(*offset)) + .map(|(_, key)| key) + .collect() +} + /// Is `expr` exactly a NaN constant -- the whole formula, not a NaN inside one? /// /// A root-level `Expr0::Const` IS the whole equation: the parser builds no node @@ -912,7 +961,7 @@ fn parse_equation( parse_inner(eqn) }; errors.extend(single_errors); - (CanonicalElementName::from_raw(subscript), ast) + (CanonicalElementName::from_subscript(subscript), ast) }) .filter(|(_, ast)| ast.is_some()) .map(|(subscript, ast)| (subscript, ast.unwrap())) diff --git a/src/simlin-engine/src/xmile/variables.rs b/src/simlin-engine/src/xmile/variables.rs index 19924d90f..1691af078 100644 --- a/src/simlin-engine/src/xmile/variables.rs +++ b/src/simlin-engine/src/xmile/variables.rs @@ -351,11 +351,13 @@ macro_rules! convert_equation( None => vec![], }; let elements = elements.into_iter().map(|e| { - let canonical_subscripts: Vec<_> = e.subscript.split(",").map(|s| canonicalize(s.trim()).into_owned()).collect(); + // The element key's canonical spelling, the one owner of what a + // subscript names (`CanonicalElementName::from_subscript`). + let canonical_subscript = crate::common::CanonicalElementName::from_subscript(&e.subscript).as_str().to_string(); // An eqn-less element (e.g. gf-only, GH #907) gets an empty // equation: with a gf that makes the element a per-element // lookup table; the writer re-emits it without an tag. - (canonical_subscripts.join(","), e.eqn.unwrap_or_default(), e.initial_eqn, e.gf.map(datamodel::GraphicalFunction::from)) + (canonical_subscript, e.eqn.unwrap_or_default(), e.initial_eqn, e.gf.map(datamodel::GraphicalFunction::from)) }).collect(); // When a top-level coexists with entries, the // top-level eqn is the EXCEPT default equation. diff --git a/src/simlin-engine/tests/integration/element_subscripts.rs b/src/simlin-engine/tests/integration/element_subscripts.rs new file mode 100644 index 000000000..4824539ff --- /dev/null +++ b/src/simlin-engine/tests/integration/element_subscripts.rs @@ -0,0 +1,415 @@ +// Copyright 2026 The Simlin Authors. All rights reserved. +// Use of this source code is governed by the Apache License, +// Version 2.0, that can be found in the LICENSE file. + +//! A per-element subscript names the same element however it is spelled, and +//! a subscript naming nothing is reported: the one key owner +//! (`CanonicalElementName::from_subscript`) applied at the JSON boundary and +//! in every consumer. The fixture is the arrays investigation's `sliced_sum` +//! model, whose spaced spelling ("nyc, young") once simulated every element +//! as 0 with no diagnostic. +//! +//! Two advisories, both Warnings: an arm whose subscript names nothing, or +//! names an element outside the variable's dimensions, is an unused arm +//! (`UnknownElementSubscript`); a declared element left with no arm and no +//! applicable default evaluates to the compiler's fabricated zero, and +//! `MissingElementEquation` names it. The zero stays a Warning because +//! Vensim defines subscripted variables on part of their range as a matter +//! of course (the elements do not exist there; see the sdeverywhere `except` +//! corpus), so an Error would refuse real models. + +use std::collections::HashMap; + +use simlin_engine::common::ErrorCode; +use simlin_engine::db::{ + DiagnosticError, LtmOverlay, SimlinDb, collect_all_diagnostics, compile_project_incremental, + sync_from_datamodel_incremental, +}; +use simlin_engine::{Results, Vm, json}; + +/// The `sliced_sum` model as Simlin JSON, with the stock's six per-element +/// subscripts spelled by `subscript`. +fn sliced_sum_json(subscript: impl Fn(&str, &str) -> String) -> String { + let elements: Vec = [ + ("nyc", "young", "100"), + ("nyc", "old", "50"), + ("boston", "young", "200"), + ("boston", "old", "80"), + ("la", "young", "60"), + ("la", "old", "40"), + ] + .iter() + .map(|(region, age, init)| { + format!( + r#"{{"subscript": "{}", "equation": "{init}"}}"#, + subscript(region, age) + ) + }) + .collect(); + format!( + r#"{{ + "name": "sliced_sum", + "simSpecs": {{"startTime": 0.0, "endTime": 20.0, "dt": "1", "method": "euler"}}, + "models": [{{ + "name": "main", + "stocks": [{{"name": "pop", "inflows": ["growth"], "outflows": [], + "arrayedEquation": {{"dimensions": ["Region", "Age"], "elements": [{}]}}}}], + "flows": [{{"name": "growth", "arrayedEquation": {{"dimensions": ["Region", "Age"], + "equation": "pop * rate * (1 - region_total / 2000)"}}}}], + "auxiliaries": [ + {{"name": "region_total", "arrayedEquation": {{"dimensions": ["Region"], "equation": "SUM(pop[Region, *])"}}}}, + {{"name": "rate", "arrayedEquation": {{"dimensions": ["Age"], "elements": [ + {{"subscript": "young", "equation": "0.08"}}, {{"subscript": "old", "equation": "0.02"}}]}}}} + ], + "views": [] + }}], + "dimensions": [ + {{"name": "Region", "elements": ["nyc", "boston", "la"]}}, + {{"name": "Age", "elements": ["young", "old"]}} + ], + "units": [] +}}"#, + elements.join(", ") + ) +} + +fn load(json_text: &str) -> simlin_engine::datamodel::Project { + let project: json::Project = serde_json::from_str(json_text).expect("valid Simlin JSON"); + project.into() +} + +/// Simulate `project` (no LTM) and return every saved series by name. +fn simulate(project: &simlin_engine::datamodel::Project) -> (Results, HashMap>) { + let mut db = SimlinDb::default(); + let sync = sync_from_datamodel_incremental(&mut db, project, None); + let compiled = compile_project_incremental(&db, sync.project, "main", LtmOverlay::Off) + .expect("the model compiles"); + let mut vm = Vm::new(compiled).unwrap(); + vm.run_to_end().expect("the model simulates"); + let results = vm.into_results(); + let series: HashMap> = results + .offsets + .iter() + .map(|(name, &off)| { + ( + name.as_str().to_string(), + results.iter().map(|row| row[off]).collect(), + ) + }) + .collect(); + (results, series) +} + +/// `"nyc, young"` and `"NYC,Young"` are the element `nyc,young`: the spaced +/// and capitalized spellings simulate bit-for-bit like the canonical one, and +/// the datamodel stores the canonical spelling for all three. +#[test] +fn spaced_and_capitalized_subscripts_simulate_like_the_canonical_spelling() { + let canonical = load(&sliced_sum_json(|r, a| format!("{r},{a}"))); + let spaced = load(&sliced_sum_json(|r, a| format!("{r}, {a}"))); + let capitalized = load(&sliced_sum_json(|r, a| { + let cap = |s: &str| { + let mut c = s.chars(); + c.next() + .map(|f| f.to_uppercase().collect::() + c.as_str()) + .unwrap_or_default() + }; + format!(" {} , {} ", cap(r), cap(a)) + })); + + let stored_subscripts = |p: &simlin_engine::datamodel::Project| -> Vec { + let stock = p.models[0] + .variables + .iter() + .find(|v| v.get_ident() == "pop") + .expect("pop"); + match stock { + simlin_engine::datamodel::Variable::Stock(s) => match &s.equation { + simlin_engine::datamodel::Equation::Arrayed(_, elements, _, _) => { + elements.iter().map(|(sub, _, _, _)| sub.clone()).collect() + } + other => panic!("pop is arrayed: {other:?}"), + }, + other => panic!("pop is a stock: {other:?}"), + } + }; + let expected = [ + "nyc,young", + "nyc,old", + "boston,young", + "boston,old", + "la,young", + "la,old", + ]; + for project in [&canonical, &spaced, &capitalized] { + assert_eq!(stored_subscripts(project), expected); + } + + let (results, canonical_series) = simulate(&canonical); + assert!(results.step_count > 1); + let final_pop: f64 = canonical_series["pop[nyc,young]"][results.step_count - 1]; + assert!(final_pop > 100.0, "the fixture grows: {final_pop}"); + for (label, project) in [("spaced", &spaced), ("capitalized", &capitalized)] { + let (_, series) = simulate(project); + assert_eq!(series.len(), canonical_series.len(), "{label}"); + for (name, values) in &canonical_series { + assert_eq!(&series[name], values, "{label}: series {name} differs"); + } + } +} + +/// A subscript naming no element combination of the variable's dimensions is +/// reported on that variable, naming the subscript, and is never a silent +/// default: the diagnostic is the `UnknownElementSubscript` advisory, which +/// matches by the same key the compiler expands with. +#[test] +fn a_subscript_naming_no_element_is_reported_on_the_variable() { + let project = load(&sliced_sum_json(|r, a| { + if r == "nyc" && a == "young" { + "nyc, yung".to_string() + } else { + format!("{r},{a}") + } + })); + let mut db = SimlinDb::default(); + let sync = sync_from_datamodel_incremental(&mut db, &project, None); + let diagnostics = collect_all_diagnostics(&db, sync.project, LtmOverlay::Off); + let unknown: Vec<(String, String)> = diagnostics + .iter() + .filter_map(|d| match &d.error { + DiagnosticError::Model(e) if e.code == ErrorCode::UnknownElementSubscript => Some(( + d.variable.clone().unwrap_or_default(), + e.get_details().unwrap_or_default().to_string(), + )), + _ => None, + }) + .collect(); + assert_eq!( + unknown.len(), + 1, + "one unknown-subscript diagnostic: {diagnostics:?}" + ); + let (variable, message) = &unknown[0]; + assert_eq!(variable, "pop"); + assert!( + message.contains("nyc,yung"), + "the diagnostic names the subscript: {message}" + ); + // The typo also leaves the declared element `nyc,young` with no arm: + // that is the zero consequence, and the sibling advisory names it. + let missing = missing_element_findings(&diagnostics); + assert_eq!( + missing, + vec![("pop".to_string(), "'nyc,young'".to_string())], + "{diagnostics:?}" + ); +} + +/// Every `MissingElementEquation` advisory, as `(variable, the quoted +/// element list the message names)`. +fn missing_element_findings( + diagnostics: &[simlin_engine::db::Diagnostic], +) -> Vec<(String, String)> { + diagnostics + .iter() + .filter_map(|d| match &d.error { + DiagnosticError::Model(e) if e.code == ErrorCode::MissingElementEquation => { + assert_eq!(d.severity, simlin_engine::db::DiagnosticSeverity::Warning); + let details = e.get_details().unwrap_or_default().to_string(); + let elements = details + .split(" has no equation for ") + .nth(1) + .and_then(|rest| rest.split(": no element entry").next()) + .unwrap_or_else(|| panic!("the warning names the elements: {details}")) + .to_string(); + Some((d.variable.clone().unwrap_or_default(), elements)) + } + _ => None, + }) + .collect() +} + +/// `sliced_sum` with `pop`'s `la,old` arm dropped and no default: the +/// declared element has no equation, so `pop` carries a Warning naming it, +/// and the element simulates as the fabricated 0 the warning announces -- +/// a loud zero, never a silent one. +#[test] +fn a_declared_element_with_no_arm_is_a_warning_naming_it() { + let text = sliced_sum_json(|r, a| format!("{r},{a}")) + .replace(r#", {"subscript": "la,old", "equation": "40"}"#, ""); + assert!(!text.contains("la,old"), "the arm is dropped"); + let project = load(&text); + let mut db = SimlinDb::default(); + let sync = sync_from_datamodel_incremental(&mut db, &project, None); + let diagnostics = collect_all_diagnostics(&db, sync.project, LtmOverlay::Off); + assert_eq!( + missing_element_findings(&diagnostics), + vec![("pop".to_string(), "'la,old'".to_string())], + "{diagnostics:?}" + ); + let (_, series) = simulate(&project); + assert!( + series["pop[la,old]"].iter().all(|v| *v == 0.0), + "the armless element is the fabricated 0 the warning names" + ); + assert!( + series["pop[la,young]"][0] == 60.0, + "its siblings are untouched" + ); +} + +/// The same dropped arm under an EXCEPT default: the default applies to +/// the missing element, so there is nothing to report and it simulates. +#[test] +fn an_except_default_covers_a_missing_element() { + let text = sliced_sum_json(|r, a| format!("{r},{a}")) + .replace(r#", {"subscript": "la,old", "equation": "40"}"#, "") + .replace( + r#""arrayedEquation": {"dimensions": ["Region", "Age"], "elements": ["#, + r#""arrayedEquation": {"dimensions": ["Region", "Age"], "equation": "40", "hasExceptDefault": true, "elements": ["#, + ); + assert!( + text.contains("hasExceptDefault"), + "the default is spliced in" + ); + let project = load(&text); + let mut db = SimlinDb::default(); + let sync = sync_from_datamodel_incremental(&mut db, &project, None); + let diagnostics = collect_all_diagnostics(&db, sync.project, LtmOverlay::Off); + assert!( + missing_element_findings(&diagnostics).is_empty(), + "{diagnostics:?}" + ); + let (results, series) = simulate(&project); + assert_eq!(series["pop[la,old]"][0], 40.0); + assert!(series["pop[la,old]"][results.step_count - 1] > 40.0); +} + +/// An entry whose equation is empty is not an arm: the compiler drops it and +/// the element takes the fabricated 0, so the warning names it exactly as it +/// names an absent entry. +#[test] +fn an_entry_with_an_empty_equation_is_not_an_arm() { + let text = sliced_sum_json(|r, a| format!("{r},{a}")).replace( + r#"{"subscript": "young", "equation": "0.08"}"#, + r#"{"subscript": "young", "equation": ""}"#, + ); + assert!(text.contains(r#""equation": """#), "the arm is emptied"); + let project = load(&text); + let mut db = SimlinDb::default(); + let sync = sync_from_datamodel_incremental(&mut db, &project, None); + let diagnostics = collect_all_diagnostics(&db, sync.project, LtmOverlay::Off); + assert_eq!( + missing_element_findings(&diagnostics), + vec![("rate".to_string(), "'young'".to_string())], + "{diagnostics:?}" + ); + let (_, series) = simulate(&project); + assert!(series["rate[young]"].iter().all(|v| *v == 0.0)); + assert!(series["rate[old]"].iter().all(|v| *v == 0.02)); +} + +/// An EXCEPT default that is the empty string parses to nothing and covers +/// nothing: the dropped `la,old` arm is reported as if there were no default. +#[test] +fn an_empty_except_default_covers_nothing() { + let text = sliced_sum_json(|r, a| format!("{r},{a}")) + .replace(r#", {"subscript": "la,old", "equation": "40"}"#, "") + .replace( + r#""arrayedEquation": {"dimensions": ["Region", "Age"], "elements": ["#, + r#""arrayedEquation": {"dimensions": ["Region", "Age"], "equation": "", "hasExceptDefault": true, "elements": ["#, + ); + assert!(text.contains(r#""equation": "", "hasExceptDefault": true"#)); + let project = load(&text); + let mut db = SimlinDb::default(); + let sync = sync_from_datamodel_incremental(&mut db, &project, None); + let diagnostics = collect_all_diagnostics(&db, sync.project, LtmOverlay::Off); + assert_eq!( + missing_element_findings(&diagnostics), + vec![("pop".to_string(), "'la,old'".to_string())], + "{diagnostics:?}" + ); + let (_, series) = simulate(&project); + assert!(series["pop[la,old]"].iter().all(|v| *v == 0.0)); +} + +/// The XMILE spelling of the same shape: `` with no +/// `` and no `` is an entry that names `b` and gives it nothing, so +/// `v[b]` is the fabricated 0, its reader reads 0, and the warning names `b`. +#[test] +fn an_xmile_element_with_neither_equation_nor_gf_is_reported() { + let xmile = r#" + +
ttt
+ 02
1
+ + + + 1 + + + + v[b] + +
"#; + let project = simlin_engine::compat::open_xmile(&mut xmile.as_bytes()).expect("XMILE parses"); + let mut db = SimlinDb::default(); + let sync = sync_from_datamodel_incremental(&mut db, &project, None); + let diagnostics = collect_all_diagnostics(&db, sync.project, LtmOverlay::Off); + assert_eq!( + missing_element_findings(&diagnostics), + vec![("v".to_string(), "'b'".to_string())], + "{diagnostics:?}" + ); + let (_, series) = simulate(&project); + assert!(series["v[a]"].iter().all(|v| *v == 1.0)); + assert!(series["v[b]"].iter().all(|v| *v == 0.0)); + assert!(series["reads_v"].iter().all(|v| *v == 0.0)); +} + +/// An arm naming an element that exists in the project but not in the +/// variable's own dimensions (`rate` is over `Age`; the arm names the +/// `Region` element `nyc`) is an unused arm: a Warning naming it, no +/// missing-element finding, and the model simulates -- every element of +/// `Age` has its equation. +#[test] +fn an_arm_for_an_element_of_another_dimension_is_only_a_warning() { + let text = sliced_sum_json(|r, a| format!("{r},{a}")).replace( + r#"{"subscript": "old", "equation": "0.02"}"#, + r#"{"subscript": "old", "equation": "0.02"}, {"subscript": "nyc", "equation": "0.5"}"#, + ); + assert!( + text.contains(r#""subscript": "nyc""#), + "the extra arm is spliced in" + ); + let project = load(&text); + let mut db = SimlinDb::default(); + let sync = sync_from_datamodel_incremental(&mut db, &project, None); + let diagnostics = collect_all_diagnostics(&db, sync.project, LtmOverlay::Off); + assert!( + missing_element_findings(&diagnostics).is_empty(), + "{diagnostics:?}" + ); + let unused: Vec<(String, String)> = diagnostics + .iter() + .filter_map(|d| match &d.error { + DiagnosticError::Model(e) if e.code == ErrorCode::UnknownElementSubscript => { + assert_eq!(d.severity, simlin_engine::db::DiagnosticSeverity::Warning); + Some(( + d.variable.clone().unwrap_or_default(), + e.get_details().unwrap_or_default().to_string(), + )) + } + _ => None, + }) + .collect(); + assert_eq!(unused.len(), 1, "{diagnostics:?}"); + assert_eq!(unused[0].0, "rate"); + assert!( + unused[0].1.contains("'nyc'"), + "names the arm: {}", + unused[0].1 + ); + let (_, series) = simulate(&project); + assert!(series["rate[young]"][0] == 0.08 && series["rate[old]"][0] == 0.02); +} diff --git a/src/simlin-engine/tests/integration/main.rs b/src/simlin-engine/tests/integration/main.rs index f9ef6d2de..756ec4752 100644 --- a/src/simlin-engine/tests/integration/main.rs +++ b/src/simlin-engine/tests/integration/main.rs @@ -28,6 +28,7 @@ mod test_helpers; mod clearn_unit_errors; mod compiler_vector; +mod element_subscripts; mod json_roundtrip; mod layout; mod ltm_array_agg; diff --git a/src/simlin-engine/tests/integration/simulate.rs b/src/simlin-engine/tests/integration/simulate.rs index 205fcfa22..0d4f63d53 100644 --- a/src/simlin-engine/tests/integration/simulate.rs +++ b/src/simlin-engine/tests/integration/simulate.rs @@ -2547,15 +2547,45 @@ TIME STEP = 1 ~~| assert!((get("g[a2]") - 7.0).abs() < 1e-10, "g[A2] should be 7"); assert!((get("g[a3]") - 7.0).abs() < 1e-10, "g[A3] should be 7"); - // h[DimA] :EXCEPT: [SubA] = 8 (no overrides for A2, A3) + // h[DimA] :EXCEPT: [SubA] = 8 (no overrides for A2, A3). Vensim defines + // no h[A2] or h[A3] at all -- its output for the sdeverywhere `except` + // model lists only h[A1] -- so the 0 here is Simlin's fabricated value + // for an element with no equation, named by the MissingElementEquation + // warning, not a Vensim result. assert!((get("h[a1]") - 8.0).abs() < 1e-10, "h[A1] should be 8"); assert!( (get("h[a2]") - 0.0).abs() < 1e-10, - "h[A2] should be 0 (undefined)" + "h[A2] is Simlin's fabricated 0 (Vensim has no such element)" ); assert!( (get("h[a3]") - 0.0).abs() < 1e-10, - "h[A3] should be 0 (undefined)" + "h[A3] is Simlin's fabricated 0 (Vensim has no such element)" + ); + // ... and the warning names those two elements on h. + let mut db = SimlinDb::default(); + let sync = sync_from_datamodel_incremental(&mut db, &datamodel_project, None); + let diagnostics = simlin_engine::db::collect_all_diagnostics( + &db, + sync.project, + simlin_engine::db::LtmOverlay::Off, + ); + let h_missing: Vec = diagnostics + .iter() + .filter_map(|d| match &d.error { + simlin_engine::db::DiagnosticError::Model(e) + if e.code == simlin_engine::common::ErrorCode::MissingElementEquation + && d.variable.as_deref() == Some("h") => + { + e.get_details() + } + _ => None, + }) + .collect(); + assert_eq!(h_missing.len(), 1, "{diagnostics:?}"); + assert!( + h_missing[0].starts_with("array variable 'h' has no equation for 'a2', 'a3'"), + "{}", + h_missing[0] ); // p[DimA] :EXCEPT: [A1] = 2, p[A1] = 5 From e85914f59096886b21fccc70c809d9b82ae52ac6 Mon Sep 17 00:00:00 2001 From: Bobby Powers Date: Wed, 9 Sep 2026 20:28:08 -0700 Subject: [PATCH 07/10] engine: report a loop's node sequence from its element circuit A per-element growth loop closed through a variable-backed reducer (growth[r] = pop*0.02*(1 - total/1000), total = SUM(pop)) was reported as growth -> pop[boston] -> total -> growth[boston]: four nodes for a three-link loop, growth twice, once bare. The reported sequence was read from links[0].from plus every link's to, and Link.from doubles as the link-score name-resolution flag: build_element_level_loops strips it to the variable level on a same-element A2A hop (a bracketed from selects the FixedIndex / cross-dimensional score name, an unbracketed one the Bare or per-target-element form), while to keeps the element subscript whenever the target is arrayed. Raw scores were right; identity was not, so the de-subscripting oracle could not match these loops to the scalar model's circuits. loop_node_sequence in db/analysis.rs is now the one owner of the reported sequence: the enumerated loops of model_detected_loops, a pin's detected_loop_from_loop, and discovery's FoundLoop summaries (analysis.rs loop_variables, whose from-led reading equals the to-led one on discovery's element-level links) all read through it. It reads the circuit off each link's to, starting from the last link's (the first node), never from from: every node of an elementary circuit is the to of exactly one link, and a stitched cross-agg loop visits only its agg twice, which the existing synthetic-agg trim absorbs (the stitcher takes pairwise node-disjoint petals, db/ltm/loops.rs, so no other node can repeat; a debug assertion states that). Only the mixed branch of build_element_level_loops strips from; the cross-element builder keeps it subscripted. Links, ids and the pin-dedup rotation are untouched, so the loop-score equations are the same text as before: the LTM dumper's every synthetic-variable equation and every saved series on bare_reducer, variable_backed and scalar_cofactor are identical before and after. Pinned: the three mixed loops of bare_reducer (SUM(pop)) and variable_backed (SUM(pop[*])) report growth[e] -> pop[e] -> total, each node once, beside the A2A loop pop -> growth, through Simlin JSON in the engine and through pysimlin Model.loops and Run.loops; scalar_cofactor's reducer loops report growth[e] -> pop[e] -> weighted with every sequence visiting each node once; a pin naming {growth, pop, total} in discovery mode reports its three element-level instances (pin1:1..3) as the same circuits through the pin path; and the owner's unit tests cover the same-element A2A hop, bare and cross-element links, a slotted agg hop and an empty loop. On the de-subscripting oracle every loop of the three models now matches the scalar model's circuit (6, 6 and 7 matched, none unmatched on either side; before, the three mixed loops of each were unmatched on both sides). --- src/pysimlin/tests/test_loop_node_sequence.py | 83 +++++++ src/simlin-engine/src/analysis.rs | 19 +- src/simlin-engine/src/db.rs | 1 + src/simlin-engine/src/db/analysis.rs | 180 +++++++++++---- .../tests/integration/ltm_loop_nodes.rs | 209 ++++++++++++++++++ src/simlin-engine/tests/integration/main.rs | 1 + 6 files changed, 438 insertions(+), 55 deletions(-) create mode 100644 src/pysimlin/tests/test_loop_node_sequence.py create mode 100644 src/simlin-engine/tests/integration/ltm_loop_nodes.rs diff --git a/src/pysimlin/tests/test_loop_node_sequence.py b/src/pysimlin/tests/test_loop_node_sequence.py new file mode 100644 index 000000000..5314a8f9d --- /dev/null +++ b/src/pysimlin/tests/test_loop_node_sequence.py @@ -0,0 +1,83 @@ +"""A detected loop's ``variables`` is its element circuit, each node once. + +The fixture is the arrays investigation's ``bare_reducer`` model: a +per-element growth loop closed through ``total = SUM(pop)``. Its three mixed +loops were reported as ``growth -> pop[boston] -> total -> growth[boston]`` +(four nodes, ``growth`` twice, once bare); the engine's +``mixed_loops_through_a_variable_backed_reducer_report_the_element_circuit`` +pins the same sequences on the Rust side. +""" + +from __future__ import annotations + +import json +from typing import TYPE_CHECKING, Any + +import simlin + +if TYPE_CHECKING: + from pathlib import Path + + +def _bare_reducer_json() -> dict[str, Any]: + return { + "name": "bare_reducer", + "simSpecs": {"startTime": 0.0, "endTime": 20.0, "dt": "1", "method": "euler"}, + "models": [ + { + "name": "main", + "stocks": [ + { + "name": "pop", + "inflows": ["growth"], + "outflows": [], + "arrayedEquation": { + "dimensions": ["Region"], + "elements": [ + {"subscript": "nyc", "equation": "100"}, + {"subscript": "boston", "equation": "200"}, + {"subscript": "la", "equation": "300"}, + ], + }, + } + ], + "flows": [ + { + "name": "growth", + "arrayedEquation": { + "dimensions": ["Region"], + "equation": "pop * 0.02 * (1 - total / 1000)", + }, + } + ], + "auxiliaries": [{"name": "total", "equation": "SUM(pop)"}], + "views": [], + } + ], + "dimensions": [{"name": "Region", "elements": ["nyc", "boston", "la"]}], + "units": [], + } + + +def _rotation(nodes: tuple[str, ...]) -> tuple[str, ...]: + start = min(range(len(nodes)), key=lambda i: nodes[i]) + return nodes[start:] + nodes[:start] + + +def test_mixed_loops_through_a_reducer_report_the_element_circuit(tmp_path: Path) -> None: + path = tmp_path / "bare_reducer.sd.json" + path.write_text(json.dumps(_bare_reducer_json())) + model = simlin.load(path) + cycles = set() + for loop in model.loops: + assert len(set(loop.variables)) == len(loop.variables), (loop.id, loop.variables) + cycles.add(_rotation(tuple(loop.variables))) + assert cycles == { + _rotation(("pop", "growth")), + _rotation(("growth[nyc]", "pop[nyc]", "total")), + _rotation(("growth[boston]", "pop[boston]", "total")), + _rotation(("growth[la]", "pop[la]", "total")), + } + # The runtime surface reports the same circuits. + run = model.run(analyze_loops=True) + assert {_rotation(tuple(lp.variables)) for lp in run.loops} == cycles diff --git a/src/simlin-engine/src/analysis.rs b/src/simlin-engine/src/analysis.rs index c8bd91159..84be08db4 100644 --- a/src/simlin-engine/src/analysis.rs +++ b/src/simlin-engine/src/analysis.rs @@ -592,19 +592,14 @@ fn signed_relative_importance(fl: &crate::ltm_finding::FoundLoop) -> Vec { .collect() } -/// Extract the ordered variable names from a `FoundLoop`. +/// The ordered variable names of a `FoundLoop`: the node sequence around the +/// cycle WITHOUT a trailing repeat of the first node, read by the one owner +/// the structural surface uses (`db::loop_node_sequence`), so both kinds of +/// loop populate the same `Loop` type consistently: consumers that render the +/// cycle (e.g. pysimlin's `Loop.__str__`) close it themselves by appending +/// the first variable, and a stored repeat would double that closing node. fn loop_variables(fl: &crate::ltm_finding::FoundLoop) -> Vec { - // The bare node sequence around the cycle (each link's `from`), WITHOUT a - // trailing repeat of the first node. This matches the structural-loop - // convention (`db::analysis` `model_detected_loops`) so both kinds of loop - // populate the same `Loop` type consistently: consumers that render the - // cycle (e.g. pysimlin's `Loop.__str__`) close it themselves by appending - // the first variable, and a stored repeat would double that closing node. - fl.loop_info - .links - .iter() - .map(|l| l.from.to_string()) - .collect() + crate::db::loop_node_sequence(&fl.loop_info) } /// Convert a `FoundLoop` to a `LoopSummary`, resolving a human-readable diff --git a/src/simlin-engine/src/db.rs b/src/simlin-engine/src/db.rs index 41d93c49f..44d858f07 100644 --- a/src/simlin-engine/src/db.rs +++ b/src/simlin-engine/src/db.rs @@ -174,6 +174,7 @@ mod analysis; pub use analysis::RefShape; pub use analysis::causal_graph_from_element_edges; pub use analysis::causal_graph_from_element_edges_with_modules; +pub(crate) use analysis::loop_node_sequence; pub(crate) use analysis::model_lowered_variables; pub(crate) use analysis::unique_module_output; pub use analysis::{ModuleOutputsRead, causal_graph_from_edges}; diff --git a/src/simlin-engine/src/db/analysis.rs b/src/simlin-engine/src/db/analysis.rs index 0ddc0ce26..d11b5848c 100644 --- a/src/simlin-engine/src/db/analysis.rs +++ b/src/simlin-engine/src/db/analysis.rs @@ -2698,31 +2698,11 @@ pub fn model_detected_loops( resolve_loop_partitions(&final_loops, &partitions, dims.as_slice()); let enumerated_loops = loops.into_iter().map(|l| { - // Extract variable names from the loop's links. Synthetic - // `$⁚ltm⁚agg⁚{n}` hops are an LTM scoring implementation detail and - // are trimmed from the user-facing list (mirroring - // `detected_loop_from_loop`); they stay in the links so the id sort - // key and the pin-dedup rotation see them. A cross-element loop's - // links carry element subscripts, so its variables are - // element-subscripted (`pool[a]`) -- the same convention a - // cross-element pin's `detected_loop_from_loop` output uses. - let is_agg = - |n: &str| crate::ltm_agg::is_synthetic_agg_name(crate::ltm::strip_subscript(n)); - let mut vars = Vec::new(); - let mut seen = std::collections::HashSet::new(); + // The reported node sequence (`loop_node_sequence`): the element + // circuit with synthetic agg hops trimmed. The hops stay in the + // links so the id sort key and the pin-dedup rotation see them. + let vars = loop_node_sequence(&l); let loop_key = loop_rotation(&l); - if !l.links.is_empty() { - let first = l.links[0].from.to_string(); - if !is_agg(&first) && seen.insert(first.clone()) { - vars.push(first); - } - for link in &l.links { - let to = link.to.to_string(); - if !is_agg(&to) && seen.insert(to.clone()) { - vars.push(to); - } - } - } // Structural classification has no runtime score data, so // confidence is binary: 1.0 when every link in the loop has // a determined polarity (R/B), 0.0 when any link is unknown @@ -2769,6 +2749,52 @@ pub fn model_detected_loops( } } +/// The node sequence a loop is reported with (`DetectedLoop::variables`): +/// the element-level circuit `n_0 -> n_1 -> ... -> n_{k-1}`, each node +/// once, the cycle implicitly closed, with synthetic `$⁚ltm⁚agg⁚{n}` hops +/// trimmed -- they are an LTM scoring implementation detail, like +/// macro/module internals, and a pin expanded through a hoisted reducer +/// carries them in its links (GH #737). A cross-element loop's nodes are +/// element-subscripted (`pool[a]`); an A2A loop's are bare variables. +/// +/// Read off each link's `to`, starting from the last link's (the first +/// node), never from `Link.from`. Every link builder keeps an element +/// subscript on `to` whenever the target is arrayed, but `from` doubles as +/// the link-score name-resolution flag: the mixed branch of +/// `db/ltm/loops.rs` `build_element_level_loops` strips it to the variable +/// level on a same-element A2A hop (`build_element_subscripted_links`, the +/// cross-element builder, keeps a subscripted `from`), so the circuit +/// `growth[boston] -> pop[boston] -> total` arrives as the links +/// `growth -> pop[boston]`, `pop[boston] -> total`, `total -> growth[boston]` +/// and a `from`-led reading reports four nodes with `growth` twice. Every +/// node of an elementary circuit is the `to` of exactly one link; a +/// stitched cross-agg loop visits its agg twice, which the trim absorbs. +/// Discovery's `FoundLoop` links (`crate::analysis`) read through here too, +/// so the two surfaces report one convention. +pub(crate) fn loop_node_sequence(l: &crate::ltm::Loop) -> Vec { + let is_agg = |n: &str| crate::ltm_agg::is_synthetic_agg_name(crate::ltm::strip_subscript(n)); + let Some(last) = l.links.last() else { + return Vec::new(); + }; + let mut vars = Vec::with_capacity(l.links.len()); + let mut seen = std::collections::HashSet::new(); + for link in std::iter::once(last).chain(l.links.iter().take(l.links.len() - 1)) { + let node = link.to.to_string(); + if is_agg(&node) { + continue; + } + let fresh = seen.insert(node.clone()); + debug_assert!( + fresh, + "a loop's circuit visits each node once: {node} repeats" + ); + if fresh { + vars.push(node); + } + } + vars +} + /// Build a `DetectedLoop` (the FFI loop surface) from one of a pin's scored /// loops, preserving its pin-derived id (`pin{n}` / `pin{n}⁚{j}`). Mirrors the /// per-loop body of `model_detected_loops`: the variable list is the cycle's @@ -2776,25 +2802,7 @@ pub fn model_detected_loops( /// structural polarity confidence is binary (1.0 for a fully-known polarity /// loop, 0.0 when any link is Unknown). fn detected_loop_from_loop(l: &crate::ltm::Loop, pin_name: &str) -> DetectedLoop { - let mut vars = Vec::new(); - let mut seen = std::collections::HashSet::new(); - // Synthetic `$⁚ltm⁚agg⁚{n}` hops are an LTM scoring implementation - // detail; like macro/module internals they are trimmed from the reported - // node sequence (a pin expanded through a hoisted reducer carries them - // in its links -- GH #737). - let is_agg = |n: &str| crate::ltm_agg::is_synthetic_agg_name(crate::ltm::strip_subscript(n)); - if !l.links.is_empty() { - let first = l.links[0].from.to_string(); - if !is_agg(&first) && seen.insert(first.clone()) { - vars.push(first); - } - for link in &l.links { - let to = link.to.to_string(); - if !is_agg(&to) && seen.insert(to.clone()) { - vars.push(to); - } - } - } + let vars = loop_node_sequence(l); let polarity = detected_polarity_from_ltm(&l.polarity); let polarity_confidence = match polarity { DetectedLoopPolarity::Undetermined => 0.0, @@ -4390,6 +4398,92 @@ mod polarity_confidence_tests { /// key for scalar and arrayed models alike -- pinned by /// `exhaustive_and_discovery_partitions_agree_..._scalar` and /// `exhaustive_and_discovery_partitions_agree_on_stock_sets_arrayed`. +/// `loop_node_sequence` reads the circuit off each link's `to`. The link +/// shapes here are the documented output of `build_element_level_loops` +/// for the hops named, not free inventions: a same-element A2A hop keeps +/// `to[e]` and strips `from`, a cross-dimensional hop keeps `from[e]` and +/// a bare `to`, an A2A loop's links are bare, and a hoisted reducer's agg +/// hop names the agg. `ltm_loop_nodes` (integration) pins the same +/// sequences through the real pipeline. +#[cfg(test)] +mod loop_node_sequence_tests { + use super::loop_node_sequence; + use crate::common::Ident; + use crate::ltm::{Link, LinkPolarity, Loop, LoopPolarity}; + + fn link(from: &str, to: &str) -> Link { + Link { + from: Ident::new(from), + to: Ident::new(to), + polarity: LinkPolarity::Positive, + } + } + + fn loop_of(links: Vec) -> Loop { + Loop { + id: String::new(), + links, + stocks: vec![], + polarity: LoopPolarity::Reinforcing, + dimensions: vec![], + slot_links: vec![], + } + } + + /// The mixed circuit `growth[boston] -> pop[boston] -> total`: the first + /// link's `from` is the stripped `growth`, so a `from`-led reading would + /// report `growth, pop[boston], total, growth[boston]`. + #[test] + fn a_mixed_loop_through_a_reducer_is_its_element_circuit() { + let l = loop_of(vec![ + link("growth", "pop[boston]"), + link("pop[boston]", "total"), + link("total", "growth[boston]"), + ]); + assert_eq!( + loop_node_sequence(&l), + vec!["growth[boston]", "pop[boston]", "total"] + ); + } + + /// Bare A2A links and cross-element links read the same either way; the + /// sequence starts at the first link's source node. + #[test] + fn bare_and_cross_element_loops_start_at_the_first_source() { + let a2a = loop_of(vec![link("pop", "growth"), link("growth", "pop")]); + assert_eq!(loop_node_sequence(&a2a), vec!["pop", "growth"]); + let cross = loop_of(vec![ + link("pop[nyc]", "migration[boston]"), + link("migration[boston]", "pop[boston]"), + link("pop[boston]", "migration[nyc]"), + link("migration[nyc]", "pop[nyc]"), + ]); + assert_eq!( + loop_node_sequence(&cross), + vec![ + "pop[nyc]", + "migration[boston]", + "pop[boston]", + "migration[nyc]" + ] + ); + } + + /// A synthetic agg hop is trimmed, including its slot suffix, and an + /// empty loop reports no nodes. + #[test] + fn agg_hops_are_trimmed_and_an_empty_loop_is_empty() { + let agg = format!("{}[nyc]", crate::ltm_agg::synthetic_agg_name(3)); + let l = loop_of(vec![ + link("pop[nyc]", &agg), + link(&agg, "growth[nyc]"), + link("growth", "pop[nyc]"), + ]); + assert_eq!(loop_node_sequence(&l), vec!["pop[nyc]", "growth[nyc]"]); + assert!(loop_node_sequence(&loop_of(vec![])).is_empty()); + } +} + #[cfg(test)] mod detected_loop_partition_tests { use super::*; diff --git a/src/simlin-engine/tests/integration/ltm_loop_nodes.rs b/src/simlin-engine/tests/integration/ltm_loop_nodes.rs new file mode 100644 index 000000000..889ea23a6 --- /dev/null +++ b/src/simlin-engine/tests/integration/ltm_loop_nodes.rs @@ -0,0 +1,209 @@ +// Copyright 2026 The Simlin Authors. All rights reserved. +// Use of this source code is governed by the Apache License, +// Version 2.0, that can be found in the LICENSE file. + +//! The node sequence a detected loop reports is its element circuit, each +//! node once. The fixtures are the arrays investigation's `bare_reducer`, +//! `variable_backed` and `scalar_cofactor` models: a per-element growth loop +//! closed through a variable-backed reducer (`total = SUM(pop)`), whose +//! mixed loops were reported as `growth -> pop[boston] -> total -> +//! growth[boston]` -- four nodes for a three-link loop, `growth` twice, once +//! bare -- because the same-element hop's `from` is stripped for link-score +//! name resolution. + +use std::collections::BTreeSet; + +use simlin_engine::db::{ + SimlinDb, model_detected_loops, set_project_ltm_discovery_mode, sync_from_datamodel_incremental, +}; +use simlin_engine::json; + +/// The investigation's corpus shape as Simlin JSON: `pop[Region]` (100, 200, +/// 300) with inflow `growth[Region] = pop * 0.02 * (1 - / +/// )` and the reducer auxiliaries `auxes` (JSON objects). +fn reducer_feedback( + name: &str, + reducer_var: &str, + cap: &str, + auxes: &str, +) -> simlin_engine::datamodel::Project { + reducer_feedback_with_pins(name, reducer_var, cap, auxes, "") +} + +/// [`reducer_feedback`] with `loopMetadata` entries (JSON objects). The +/// variables carry uids -- 1 for `pop`, 2 for `growth`, and whatever the +/// `auxes` objects declare -- so a pin can name them. +fn reducer_feedback_with_pins( + name: &str, + reducer_var: &str, + cap: &str, + auxes: &str, + loop_metadata: &str, +) -> simlin_engine::datamodel::Project { + let text = format!( + r#"{{ + "name": "{name}", + "simSpecs": {{"startTime": 0.0, "endTime": 20.0, "dt": "1", "method": "euler"}}, + "models": [{{ + "name": "main", + "stocks": [{{"name": "pop", "uid": 1, "inflows": ["growth"], "outflows": [], + "arrayedEquation": {{"dimensions": ["Region"], "elements": [ + {{"subscript": "nyc", "equation": "100"}}, {{"subscript": "boston", "equation": "200"}}, + {{"subscript": "la", "equation": "300"}}]}}}}], + "flows": [{{"name": "growth", "uid": 2, "arrayedEquation": {{"dimensions": ["Region"], + "equation": "pop * 0.02 * (1 - {reducer_var} / {cap})"}}}}], + "auxiliaries": [{auxes}], + "loopMetadata": [{loop_metadata}], + "views": [] + }}], + "dimensions": [{{"name": "Region", "elements": ["nyc", "boston", "la"]}}], + "units": [] +}}"# + ); + let project: json::Project = serde_json::from_str(&text).expect("valid Simlin JSON"); + project.into() +} + +/// Rotate a node sequence to start at its smallest node, so two spellings +/// of one cycle compare equal. +fn rotation(nodes: &[String]) -> Vec { + let Some(start) = (0..nodes.len()).min_by_key(|&i| &nodes[i]) else { + return Vec::new(); + }; + nodes[start..] + .iter() + .chain(&nodes[..start]) + .cloned() + .collect() +} + +/// Every reported loop's node sequence, rotation-normalized, for `project`. +fn reported_cycles(project: &simlin_engine::datamodel::Project) -> BTreeSet> { + let mut db = SimlinDb::default(); + let sync = sync_from_datamodel_incremental(&mut db, project, None); + let source_model = sync.models["main"].source_model; + let detected = model_detected_loops(&db, source_model, sync.project); + detected + .loops + .iter() + .map(|l| { + let distinct: BTreeSet<&String> = l.variables.iter().collect(); + assert_eq!( + distinct.len(), + l.variables.len(), + "{}: a loop visits each node once: {:?}", + l.id, + l.variables + ); + rotation(&l.variables) + }) + .collect() +} + +fn cycle(nodes: &[&str]) -> Vec { + rotation(&nodes.iter().map(|n| n.to_string()).collect::>()) +} + +/// The three per-element loops through the reducer are `growth[e] -> +/// pop[e] -> total`, three nodes each once, beside the A2A loop `pop -> +/// growth`, for a bare argument (`SUM(pop)`) and a starred one +/// (`SUM(pop[*])`). +#[test] +fn mixed_loops_through_a_variable_backed_reducer_report_the_element_circuit() { + let expected: BTreeSet> = [ + cycle(&["pop", "growth"]), + cycle(&["growth[nyc]", "pop[nyc]", "total"]), + cycle(&["growth[boston]", "pop[boston]", "total"]), + cycle(&["growth[la]", "pop[la]", "total"]), + ] + .into_iter() + .collect(); + for (name, reducer) in [ + ("bare_reducer", "SUM(pop)"), + ("variable_backed", "SUM(pop[*])"), + ] { + let auxes = format!(r#"{{"name": "total", "uid": 3, "equation": "{reducer}"}}"#); + let cycles = reported_cycles(&reducer_feedback(name, "total", "1000", &auxes)); + assert_eq!(cycles, expected, "{name}"); + } +} + +/// The scalar-cofactor model reads `pop[nyc]` inside the reducer's scale +/// (`weighted = SUM(pop[*] * scale)`, `scale = 1 + pop[nyc] / 10000`), so +/// `pop[nyc]` also closes loops through `scale`; every reported sequence +/// still visits each node once and the reducer loops are the three-node +/// circuits. +#[test] +fn scalar_cofactor_reducer_loops_report_each_node_once() { + let auxes = r#"{"name": "weighted", "uid": 3, "equation": "SUM(pop[*] * scale)"}, + {"name": "scale", "uid": 4, "equation": "1 + pop[nyc] / 10000"}"#; + let cycles = reported_cycles(&reducer_feedback( + "scalar_cofactor", + "weighted", + "5000", + auxes, + )); + for region in ["nyc", "boston", "la"] { + let through_reducer = cycle(&[ + &format!("growth[{region}]"), + &format!("pop[{region}]"), + "weighted", + ]); + assert!( + cycles.contains(&through_reducer), + "{region}: the reducer loop is its element circuit; got {cycles:?}" + ); + } + assert!(cycles.contains(&cycle(&["pop", "growth"])), "{cycles:?}"); +} + +/// The pin path, in discovery mode: a pin naming `{growth, pop, total}` is +/// expanded on the element graph into the three mixed circuits, each +/// reported through `detected_loop_from_loop` -- the same owner -- as +/// `growth[e] -> pop[e] -> total`, and in discovery mode those are the only +/// loops the structural surface reports. +#[test] +fn a_pin_through_a_reducer_reports_the_element_circuits_in_discovery_mode() { + let project = reducer_feedback_with_pins( + "bare_reducer_pinned", + "total", + "1000", + r#"{"name": "total", "uid": 3, "equation": "SUM(pop)"}"#, + r#"{"uids": [1, 2, 3], "name": "Reducer feedback"}"#, + ); + let mut db = SimlinDb::default(); + let sync = sync_from_datamodel_incremental(&mut db, &project, None); + set_project_ltm_discovery_mode(&mut db, sync.project, true); + let source_model = sync.models["main"].source_model; + let detected = model_detected_loops(&db, source_model, sync.project); + let mut ids: Vec<&str> = detected.loops.iter().map(|l| l.id.as_str()).collect(); + ids.sort_unstable(); + assert_eq!( + ids, + ["pin1\u{205A}1", "pin1\u{205A}2", "pin1\u{205A}3"], + "the pin's three element-level instances are the reported loops" + ); + let cycles: BTreeSet> = detected + .loops + .iter() + .map(|l| { + let distinct: BTreeSet<&String> = l.variables.iter().collect(); + assert_eq!( + distinct.len(), + l.variables.len(), + "{}: {:?}", + l.id, + l.variables + ); + rotation(&l.variables) + }) + .collect(); + let expected: BTreeSet> = [ + cycle(&["growth[nyc]", "pop[nyc]", "total"]), + cycle(&["growth[boston]", "pop[boston]", "total"]), + cycle(&["growth[la]", "pop[la]", "total"]), + ] + .into_iter() + .collect(); + assert_eq!(cycles, expected); +} diff --git a/src/simlin-engine/tests/integration/main.rs b/src/simlin-engine/tests/integration/main.rs index 756ec4752..bd34ca1bc 100644 --- a/src/simlin-engine/tests/integration/main.rs +++ b/src/simlin-engine/tests/integration/main.rs @@ -34,6 +34,7 @@ mod layout; mod ltm_array_agg; mod ltm_discovery_large_models; mod ltm_dt_invariance; +mod ltm_loop_nodes; mod ltm_relative_scores; // Compares xmutil-based MDL parsing against the native Rust parser, so it // needs the optional xmutil C++ converter compiled in. From f5e032575bba55bdefd20b3a769ed9152c4bf8c0 Mon Sep 17 00:00:00 2001 From: Bobby Powers Date: Wed, 9 Sep 2026 20:45:09 -0700 Subject: [PATCH 08/10] engine: freeze the clock inside the ceteris-paribus partial MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The changed-first partial froze model dependencies only, so TIME, STEP, RAMP and PULSE advanced inside it: a source with no influence on its target scored +/-1 whenever an exogenous forcing moved the target, and the forcing was credited to every link into it (GH #1016; 41% of C-LEARN's link-score arms carried a live time term). The 2020 paper's partial f(x_t, y_{t-1}) - z_{t-1} reads every input but the isolated one at the previous step, and the clock is an input of f like any other; the identity "a source f does not read scores 0" only closes with the clock among the "all other inputs". This is the reading of the formula; what Vensim and Stella do here is unverified and the decision does not rest on it. One rule, at both changed-first freeze sites (the wrap walker, whose index pass and table-index freeze route through it, and the reducer-body freeze): a call of a builtin the compiler itself classifies Invariance::TimeDependent (the one owner of "reads the clock") is read at the previous step -- a bare TIME through the per-model helper $⁚ltm⁚freeze⁚time = PREVIOUS(TIME), any other call lagged whole with its arguments verbatim (a capture lags the call once; lagging a frozen dep inside it as well would read it two steps back) -- unless the call holds an occurrence of the isolated input's live shape, the occurrence IR's verdict, in which case the call stays live, clock included: the isolated input and the clock are then inseparable in one call, and its whole change is attributed to the input. That residual is documented and pinned, not diagnosed; a read of another element of an arrayed input is not the live shape and freezes with the call. Inside a frozen dependency's subscript index the call is left verbatim, because the enclosing freeze already lags the whole read once (PREVIOUS(arr[TIME]) is arr_{t-1}[TIME_{t-1}]); a model dependency in that position is lagged twice today, the pre-existing residual the zero-slot predicate documents, and the clock does not join it. The run constants DT, INITIAL_TIME and FINAL_TIME are Pure and untouched. The changed-last duals leave the clock live deliberately: their numerator reads every other input at t on both sides, so the clock's motion cancels there. The guard form's own TIME = INITIAL_TIME arm is outside the partial. The helper exists because a frozen clock is one value for the whole model: spelled inline, PREVIOUS(TIME) is a capture per arm and per occurrence, which on C-LEARN v77 cost 9,632 result slots (a third of the row). With the helper, minted once by model_ltm_variables when any arm reads it and registered as a freeze helper so it sorts ahead of every score, the width moves 28,725 -> 28,980 (+255: 36 helper instances, their 36 captures, and 183 captures of time-dependent calls frozen whole) and the emitted count 6,224 -> 6,227 (the helper in main, the ramp_from_to macro model and the stdlib npv template). The zero-slot predicate recognizes the helper as the one-step lag it is, so an arm whose only varying content was the clock is now provably PREVIOUS(target) and omitted to an exact zero. Pinned (tests/integration/ltm_frozen_clock.rs): x = 5 + STEP(2, 2) + 0*s scores LS(s -> x) = 0 at every step (was 1 at t = 2); the isolated loop f = 0.1*s + RAMP(1, 0) scores LS(s -> f) = 0.1*Δs/|Δf| = 0.500, 0.545, 0.587, 0.624, 0.658 at t = 1..5 with the loop score equal to it (was 1); z = x * TIME scores the changed-first numerator x_t * TIME_{t-1} - z_{t-1}, distinguishable from the clock-live form (a constant 1) at every step; the GH #763 repro grow = 1 + MIN(pop[*] * TIME) keeps the PREVIOUS(agg) anchor, so the row that is never the argmin scores 0; z = x + arr[TIME] over arr = [1, 2, 4, 8, 16] scores 0.1667, 0.1379, 0.1213, 0.1118 at t = 2..5, the once-lagged form, where a twice-lagged clock flips the sign from t = 3; and f = 1 + STEP(0.1*s, 2) pins the residual. The text-level rows enumerate every Invariance class a partial can meet, both changed-last duals keep their clock live, and the helper is minted once per model and only when read. Measured on the LTM corpus and C-LEARN v77 with a simlin simulate --ltm run on the previous and the new CLI: the five corpus models (logistic_growth, hero_culture, arrayed_population, cross_agg, cross_element) change no link-score or loop-score series. C-LEARN changes 146 of 5,998 link-score series, 86 of them to an identical zero; every one lies on an edge no retained loop uses, and analyze() reports the same 153 loops, identical relative-score series and identical 565 dominant periods before and after, so no dominant loop changes at any step. The remaining changes replace the clock-credited +/-1 with the paper's value: rs_global_pfc -> proportion_of_global_to_cop_pfc[..] (ZIDZ of a lookup of Time over the global sum) goes from an alternating +/-1 to the large negative ratio of a target that barely moves while its inputs cancel, the "very large values" case section 3.1 of the reference describes. The C-LEARN slot-maxima digest is re-pinned with its derivation split across the two commits that moved it: the net-flow aux commit (05e4dae4) moved slots 20,221 -> 20,337 and nonzero slots 3,106 -> 3,198 without re-pinning this #[ignore]d release-only gate, and the frozen clock moves them 20,337 -> 20,338 and 3,198 -> 2,876 (the omitted clock-only arms). The lookup-pin and value-gate expectations are re-derived under the frozen clock; the reference doc's section 3.1 note, the design doc, the engine CLAUDE.md and the performance doc state the rule. --- docs/design/engine-performance.md | 19 +- docs/design/ltm--loops-that-matter.md | 13 +- docs/reference/ltm--loops-that-matter.md | 21 ++ src/simlin-engine/CLAUDE.md | 2 +- src/simlin-engine/src/db/ltm/equation.rs | 17 + src/simlin-engine/src/db/ltm/link_scores.rs | 6 +- src/simlin-engine/src/db/ltm/mod.rs | 20 ++ src/simlin-engine/src/db/ltm_tests.rs | 108 +++++- .../src/db/ltm_value_gate_tests.rs | 47 +-- .../src/db/ltm_value_golden/value_gate.txt | 13 +- src/simlin-engine/src/ltm_augment.rs | 165 ++++++++- .../src/ltm_augment_array_freeze.rs | 38 +++ .../src/ltm_augment_pin_tests.rs | 12 +- src/simlin-engine/src/ltm_augment_tests.rs | 185 ++++++++++ .../src/ltm_augment_zero_slot.rs | 40 ++- .../tests/integration/ltm_frozen_clock.rs | 319 ++++++++++++++++++ src/simlin-engine/tests/integration/main.rs | 1 + .../tests/integration/simulate.rs | 18 +- .../tests/integration/simulate_ltm.rs | 32 +- 19 files changed, 989 insertions(+), 87 deletions(-) create mode 100644 src/simlin-engine/tests/integration/ltm_frozen_clock.rs diff --git a/docs/design/engine-performance.md b/docs/design/engine-performance.md index c13e5e119..c825ad066 100644 --- a/docs/design/engine-performance.md +++ b/docs/design/engine-performance.md @@ -982,14 +982,17 @@ short of it rather than taking a ~2-3%. 10. **LTM link-score arms** — the dominant cost of an LTM-enabled run on an arrayed model, and mostly a generation question rather than a VM one. An arm whose ceteris-paribus partial is *provably* `PREVIOUS(target)` is - omitted and lowers to a single zero-store; on C-LEARN that is 4,335 arms - and −19.2% of the flow program. The residual is gated on a semantics - question, not on engineering: ~5,000 further arms are blocked solely by a - live `time()`, because TIME is excluded from the freeze (GH #1016), and - resolving that would roughly double the win. Do **not** substitute the - cheaper negative test ("the link's source stayed frozen") — it asks a - different question and silently rewrites 187 result slots. GH #977 carries - the decomposition and the standing constraints. + omitted and lowers to a single zero-store. The clock is a frozen input of + the partial (GH #1016: `TIME` and the time-dependent builtins read at the + previous step), so an arm whose only varying content was the clock is + omitted too. A frozen bare `TIME` reads one per-model helper + (`$⁚ltm⁚freeze⁚time = PREVIOUS(TIME)`; spelled inline it was a capture per + arm and per occurrence, 9,632 slots on C-LEARN), while a frozen + time-dependent CALL is still a capture per arm (183 slots there), which + is the cost side of the same rule. Do + **not** substitute the cheaper negative test ("the link's source stayed + frozen") — it asks a different question and silently rewrites real + scores. GH #977 carries the decomposition and the standing constraints. "Provably" carries a LAG-ALIGNMENT requirement that a walk stopping at the first `PREVIOUS` will miss: the partial equals `PREVIOUS(target)` only if diff --git a/docs/design/ltm--loops-that-matter.md b/docs/design/ltm--loops-that-matter.md index 1b747bb9a..9edc6d815 100644 --- a/docs/design/ltm--loops-that-matter.md +++ b/docs/design/ltm--loops-that-matter.md @@ -2128,7 +2128,18 @@ cases remain deliberate carve-outs: recursively transforming it to wrap non-excluded dependencies in `PREVIOUS()`, and printing the result back to equation text. This is done once at augmentation time (not per-timestep), producing a static equation that the - simulation engine evaluates normally. + simulation engine evaluates normally. The clock is one of the "all others" + (GH #1016): inside a changed-first partial `TIME` reads the per-model helper + `$⁚ltm⁚freeze⁚time = PREVIOUS(TIME)` and a time-dependent call (`STEP`, + `RAMP`, `PULSE`, by the builtin's own `Invariance::TimeDependent`) is lagged + whole, arguments verbatim, unless the call holds an occurrence of the + isolated input's live shape, in which case it stays live, clock included; + inside a frozen dependency's subscript index the enclosing freeze already + lags the clock once and it is left alone. The changed-last fallback leaves + the clock live, as it leaves every other input live. A source with no + influence on its target therefore scores 0 under an exogenous forcing, and + a forcing on a loop takes its own share of the target's change rather than + the loop's (`tests/integration/ltm_frozen_clock.rs`). ## Test Coverage diff --git a/docs/reference/ltm--loops-that-matter.md b/docs/reference/ltm--loops-that-matter.md index 55260d5ec..119ab0316 100644 --- a/docs/reference/ltm--loops-that-matter.md +++ b/docs/reference/ltm--loops-that-matter.md @@ -312,6 +312,27 @@ re-evaluation per input per variable). > loses nothing. The two renderings differ only in which step's co-factor weights the > isolated input's change. +> **Simlin implementation note: the clock is a frozen input.** +> "All other inputs" includes the clock. In the changed-first partial `TIME` is +> read at the previous step, and a call of a time-dependent builtin (`STEP`, +> `RAMP`, `PULSE`) is evaluated whole at the previous step, so +> `f(x_current, y_previous)` is the target's equation with the isolated input +> alone advanced. Without that, a source with no influence on its target scores +> +/-1 whenever an exogenous forcing moves the target, and the forcing is +> credited to every link into it; with it, the identity `Delta_x(z) = 0` for an +> input `x` that `f` does not read holds exactly, and an exogenous forcing on a +> loop takes its own share of `Delta(z)` rather than the loop's. The run +> constants `DT`, `INITIAL TIME` and `FINAL TIME` are not clock reads. One +> residual: a time-dependent call that reads the isolated input itself +> (`STEP(x, 2)`; a read of another element of an arrayed input is not one) +> cannot be evaluated at the previous step without lagging `x` too, so that +> call is left live, clock included, and its whole change is attributed to +> `x`. The changed-last fallback leaves the clock live by +> construction: `z(x_current, w_current) - z(x_previous, w_current)` reads +> every other input at the current step on both sides, so the clock's own +> motion cancels. What Vensim and Stella do here is unverified; this is the +> reading of the formula above (GH #1016). + ### 3.2 Flow-to-Stock Link Score For a stock S with inflow i and outflow o, the link scores from the flows to the stock diff --git a/src/simlin-engine/CLAUDE.md b/src/simlin-engine/CLAUDE.md index 0b5c85fb0..f6c524917 100644 --- a/src/simlin-engine/CLAUDE.md +++ b/src/simlin-engine/CLAUDE.md @@ -135,7 +135,7 @@ LTM instruments a model with synthetic variables (link scores, loop scores, aggr - `equation.rs`: `LtmEquation`, the typed per-arm `Expr0` a synthetic variable carries (a text generator's arm is parsed once, `LtmArm::new`; a typed generator's arm is its tree, `LtmArm::from_typed`). `link_scores.rs`: the emitters, each parsing its arm once and reading the completeness guard off that arm; `shaped_link_score` is the per-shape entry. - `loops.rs` (including cross-aggregate loop recovery), `pinned.rs` (`LOOPSCORE` pins), `compile.rs` (`lower_ltm_variable`, the one lowering of a generated equation or a generated capture under the shapes of its reads; fragment compilation; the compile-failure diagnostic pass), `parse.rs`. - `LtmImplicitVarMeta` is captures only: a generated equation is built from an expanded tree and mints no module instance, so the module universe under LTM is the source models' own (`db::assemble::enumerate_module_instances`). -- `ltm_augment.rs` and its `ltm_augment_*.rs` children -- the ceteris-paribus equation generators: wrap every non-source reference of the target's equation in `PREVIOUS`, pin arrayed dependencies to the element the target reads (`subscript_idents_in_expr0`, on the tree), and freeze subscript indices. A flow-to-stock score is the closed-form partial of the stock's synthetic net-flow aux `$⁚ltm⁚net⁚{stock} = inflows - outflows` (`generate_net_flow_equation`, minted beside the score by `shaped_link_score` and deduplicated by name), `sign * |Δflow / Δnet|` over the same `[t - dt, t]` window as every other link; the aux is not a causal node. Production never parses a target equation; it lowers the target's `Expr2` back to `Expr0` (`patch::expr2_to_expr0`) and consults the IR by structural path, and a target with no lowered body is declined (`PartialEquationErrorKind::MissingTypedTarget`), never scored around user text or a `0` body. An aggregate node's reducer is the classified `BuiltinFn` (`AggNode::reducer`, projected by `reducer_expr0`); `AggNode::reducer_key` is its printed identity for the dedup, the IR's routing and the wrap's matching, never an equation to parse. +- `ltm_augment.rs` and its `ltm_augment_*.rs` children -- the ceteris-paribus equation generators: wrap every non-source reference of the target's equation in `PREVIOUS`, pin arrayed dependencies to the element the target reads (`subscript_idents_in_expr0`, on the tree), and freeze subscript indices. The clock is a frozen input too (GH #1016): a changed-first partial reads `TIME` and a time-dependent call (`STEP`, `RAMP`, `PULSE`; the builtin's own `Invariance::TimeDependent`, read through `is_time_dependent_builtin`) at the previous step -- a bare `TIME` through the per-model helper `$⁚ltm⁚freeze⁚time = PREVIOUS(TIME)` (`FROZEN_CLOCK_HELPER`, minted once by `model_ltm_variables` when any arm reads it), a call lagged whole with its arguments verbatim, or left to the enclosing freeze inside a frozen dependency's subscript index -- unless the call holds an occurrence of the isolated input's live shape (the occurrence IR's verdict), in which case it stays live, clock included; the changed-last duals (`wrap_live_shaped_in_previous`, `wrap_matching_in_previous`) leave the clock live because their numerator reads every other input at `t` on both sides. A flow-to-stock score is the closed-form partial of the stock's synthetic net-flow aux `$⁚ltm⁚net⁚{stock} = inflows - outflows` (`generate_net_flow_equation`, minted beside the score by `shaped_link_score` and deduplicated by name), `sign * |Δflow / Δnet|` over the same `[t - dt, t]` window as every other link; the aux is not a causal node. Production never parses a target equation; it lowers the target's `Expr2` back to `Expr0` (`patch::expr2_to_expr0`) and consults the IR by structural path, and a target with no lowered body is declined (`PartialEquationErrorKind::MissingTypedTarget`), never scored around user text or a `0` body. An aggregate node's reducer is the classified `BuiltinFn` (`AggNode::reducer`, projected by `reducer_expr0`); `AggNode::reducer_key` is its printed identity for the dedup, the IR's routing and the wrap's matching, never an equation to parse. - `ltm/` -- the vocabulary (`Link`, `Loop`, polarity in `types.rs`), `CausalGraph` (`graph.rs`), Johnson's circuit enumerator and Tarjan SCC (`indexed.rs`), cycle partitions (`partitions.rs`), static polarity analysis (`polarity.rs`). - `ltm_finding.rs` with `ltm_finding_enum.rs` and `ltm_finding_fallback.rs` -- post-simulation discovery: exact union-graph enumeration, a shortest-path sampling fallback under a wall-clock and memory budget, retention against the full loop universe, and the coverage-aware cap. - `ltm_post.rs`, `ltm_dominance.rs` -- relative loop and link scores (normalized within a cycle partition, in emission order so the IEEE sum is stable) and dominant-period selection. diff --git a/src/simlin-engine/src/db/ltm/equation.rs b/src/simlin-engine/src/db/ltm/equation.rs index bd160019d..fe4b00c9c 100644 --- a/src/simlin-engine/src/db/ltm/equation.rs +++ b/src/simlin-engine/src/db/ltm/equation.rs @@ -196,6 +196,23 @@ impl LtmEquation { } } + /// Every arm of the equation, in slot order (the default last). + pub(crate) fn arms(&self) -> impl Iterator { + let (single, elements, default): (Option<&LtmArm>, &[(String, LtmArm)], Option<&LtmArm>) = + match self { + LtmEquation::Scalar(arm) | LtmEquation::ApplyToAll(_, arm) => { + (Some(arm), &[], None) + } + LtmEquation::Arrayed { + elements, default, .. + } => (None, elements.as_slice(), default.as_ref()), + }; + single + .into_iter() + .chain(elements.iter().map(|(_, arm)| arm)) + .chain(default) + } + /// The equation's diagnostic source text concatenated into a single string /// (the generator's exact spelling): the scalar / apply-to-all formula /// verbatim, or -- for the per-element (`Arrayed`) variant -- every element diff --git a/src/simlin-engine/src/db/ltm/link_scores.rs b/src/simlin-engine/src/db/ltm/link_scores.rs index d048c9637..b9682ee1d 100644 --- a/src/simlin-engine/src/db/ltm/link_scores.rs +++ b/src/simlin-engine/src/db/ltm/link_scores.rs @@ -3144,8 +3144,10 @@ fn emit_per_element_link_scores( /// the freeze set (`model_deps`) and -- when arrayed -- its declared /// dimension count (`arrayed_dep_dims`, the row-pinning gate). Identifiers /// in the body that are not model variables (dimension/element names in -/// subscripts, TIME) are excluded, so the body partial leaves them live, -/// matching `build_partial_equation_shaped`'s deps-only freezing. +/// subscripts) are excluded, so the body partial leaves them live; the clock +/// is not an identifier at all but a builtin, frozen by +/// `ltm_augment::is_time_dependent_builtin`'s rule (GH #1016), matching +/// `build_partial_equation_shaped`'s convention. fn reducer_body_ctx_parts( db: &dyn Db, source_vars: &HashMap, diff --git a/src/simlin-engine/src/db/ltm/mod.rs b/src/simlin-engine/src/db/ltm/mod.rs index bce79439d..cf630c9ae 100644 --- a/src/simlin-engine/src/db/ltm/mod.rs +++ b/src/simlin-engine/src/db/ltm/mod.rs @@ -1899,6 +1899,26 @@ pub fn model_ltm_variables( }); } + // The frozen clock's helper (GH #1016) is one value for the whole model, + // so it is minted here, once, when any arm reads it -- the wrap spells a + // frozen `TIME` as a reference to it rather than threading a helper + // channel through every generator (the reducer-body and per-element + // generators have none). Registered as the freeze helper it is, so the + // dedup and the category sort below treat it like every other. + let reads_frozen_clock = vars.iter().any(|v| { + v.equation.arms().any(|arm| { + arm.expr.as_deref().is_some_and(|expr| { + crate::ltm_augment::expr_reference_idents(expr) + .contains(crate::ltm_augment::FROZEN_CLOCK_HELPER) + }) + }) + }); + if reads_frozen_clock { + vars.push(compile::freeze_helper_var( + crate::ltm_augment::frozen_clock_helper(), + )); + } + // A score's companion variables are minted per score with content-derived // names -- a freeze helper (GH #995) once per partial that references the // frozen slice, a stock's net-flow aux once per flow of the stock -- so diff --git a/src/simlin-engine/src/db/ltm_tests.rs b/src/simlin-engine/src/db/ltm_tests.rs index d05bc6048..e7fe2c234 100644 --- a/src/simlin-engine/src/db/ltm_tests.rs +++ b/src/simlin-engine/src/db/ltm_tests.rs @@ -2323,7 +2323,7 @@ fn a_lookup_table_index_is_element_pinned_in_a_per_element_partial() { let score_name = offsets .keys() .map(|k| k.as_str().to_string()) - .find(|k| k.contains("link_score") && k.contains("factor\u{2192}out[b]")) + .find(|k| k == "$\u{205A}ltm\u{205A}link_score\u{205A}factor\u{2192}out[b]") .expect("the factor->out[b] per-element link score must be emitted"); let mut vm = crate::vm::Vm::new(compiled).expect("vm"); vm.run_to_end().expect("run"); @@ -2333,13 +2333,28 @@ fn a_lookup_table_index_is_element_pinned_in_a_per_element_partial() { .map(|s| results.data[s * results.step_size + base]) .collect(); - // The EXACT series, both elements. `b`'s table has slope 3 and `a`'s has 2, - // and `drift` supplies a second varying contribution so the score does not - // saturate at 1 -- which is what makes the pinned element observable in the - // VALUE. Pinning is the whole subject of this test, so the assertion has to - // be the number, not a property the number happens to satisfy. - let expected_a = [0.0, 0.0, 0.805_022_617_376_384_4, 0.809_579_376_075_400_5]; - let expected_b = [0.0, 0.0, 0.860_979_814_269_031_9, 0.864_449_016_428_508_1]; + // The EXACT series, both elements. The partial reads the clock frozen + // (`LOOKUP(tbl[e], "$⁚ltm⁚freeze⁚time")`, the per-model `PREVIOUS(TIME)` + // helper, GH #1016), so the numerator is + // `L_e(TIME_{t-1}) * Δfactor`: the table's own slope over the step and the + // drift are the rest of `Δout[e]` and are not credited to `factor`, which + // is why the scores are small. `b`'s table has slope 3 and `a`'s has 2, so + // `L_e(TIME_{t-1})` -- and with it the score -- differs per element, which + // is what makes the pinned element observable in the VALUE. Pinning is the + // whole subject of this test, so the assertion has to be the number, not a + // property the number happens to satisfy. + let expected_a = [ + 0.0, + 0.0, + 0.004_757_448_136_016_217_4, + 0.018_677_978_159_516_23, + ]; + let expected_b = [ + 0.0, + 0.0, + 0.005_088_138_797_753_429, + 0.019_943_887_314_840_866, + ]; assert_eq!( series.len(), expected_b.len(), @@ -2358,7 +2373,7 @@ fn a_lookup_table_index_is_element_pinned_in_a_per_element_partial() { let a_off = offsets .keys() .map(|k| k.as_str().to_string()) - .find(|k| k.contains("link_score") && k.contains("factor\u{2192}out[a]")) + .find(|k| k == "$\u{205A}ltm\u{205A}link_score\u{205A}factor\u{2192}out[a]") .map(|n| offsets[&crate::common::Ident::new(&n)]) .expect("the factor->out[a] per-element link score must be emitted"); let series_a: Vec = (0..results.step_count) @@ -2764,3 +2779,78 @@ fn delay3_input_to_stock_link_score_reads_the_bound_port() { ); } } + +/// The frozen clock's helper (GH #1016) is one variable per model: minted +/// once when any arm reads it, named as a freeze helper so it sorts ahead of +/// every score, defined as `PREVIOUS(TIME)`, and absent from a model whose +/// partials read no clock. +#[test] +fn the_frozen_clock_helper_is_minted_once_per_model_that_reads_the_clock() { + use crate::ltm_augment::FROZEN_CLOCK_HELPER; + + // Two targets read the clock, on two loops through one stock. + let project = TestProject::new("frozen_clock_helper") + .with_sim_time(0.0, 3.0, 1.0) + .stock("s", "10", &["a", "b"], &[], None) + .flow("a", "0.1 * s + TIME", None) + .flow("b", "0.2 * s + STEP(1, 2) * TIME", None) + .build_datamodel(); + let db = SimlinDb::default(); + let sync = sync_from_datamodel(&db, &project); + let ltm = crate::db::model_ltm_variables(&db, sync.models["main"].source, sync.project); + + let helpers: Vec<&crate::db::LtmSyntheticVar> = ltm + .vars + .iter() + .filter(|v| v.name == FROZEN_CLOCK_HELPER) + .collect(); + assert_eq!( + helpers.len(), + 1, + "one clock helper for the model; have {:?}", + ltm.vars.iter().map(|v| &v.name).collect::>() + ); + let helper = helpers[0]; + assert_eq!(helper.equation.source_text(), "PREVIOUS(TIME)"); + assert!(helper.dimensions.is_empty()); + assert!(helper.compile_directly); + + // Both scores read it, and it is evaluated before either. + let helper_pos = ltm + .vars + .iter() + .position(|v| v.name == FROZEN_CLOCK_HELPER) + .unwrap(); + let mut readers = 0; + for (pos, var) in ltm.vars.iter().enumerate() { + let reads = var.equation.arms().any(|arm| { + arm.expr.as_deref().is_some_and(|e| { + crate::ltm_augment::expr_reference_idents(e).contains(FROZEN_CLOCK_HELPER) + }) + }); + if reads { + readers += 1; + assert!( + pos > helper_pos, + "{} reads the clock helper but is ordered before it", + var.name + ); + } + } + assert_eq!(readers, 2, "the s -> a and s -> b scores read the helper"); + + // The other arm of the decision: a model whose partials read no clock + // mints no helper. + let project = TestProject::new("no_clock") + .with_sim_time(0.0, 3.0, 1.0) + .stock("s", "10", &["a"], &[], None) + .flow("a", "0.1 * s", None) + .build_datamodel(); + let db = SimlinDb::default(); + let sync = sync_from_datamodel(&db, &project); + let ltm = crate::db::model_ltm_variables(&db, sync.models["main"].source, sync.project); + assert!( + ltm.vars.iter().all(|v| v.name != FROZEN_CLOCK_HELPER), + "no clock read, no helper" + ); +} diff --git a/src/simlin-engine/src/db/ltm_value_gate_tests.rs b/src/simlin-engine/src/db/ltm_value_gate_tests.rs index 115a4c4fc..e367fbf13 100644 --- a/src/simlin-engine/src/db/ltm_value_gate_tests.rs +++ b/src/simlin-engine/src/db/ltm_value_gate_tests.rs @@ -39,13 +39,11 @@ use crate::test_common::TestProject; /// each arm's fate is decided independently: /// /// * `nyc` reads the link source `pop[nyc]` AND carries `TIME`. For the -/// `pop[nyc]` link this arm is live on both counts; for the OTHER links it is -/// the load-bearing row -- every occurrence of their source is frozen, and -/// the arm must still be materialized because `TIME` advances. This is the -/// mechanism that makes the naive "the source stayed frozen" collapse unsound -/// (5,035 of C-LEARN's 9,514 no-live-source arms are blocked solely by a live -/// `time()`; GH #1016). If a future relaxation drops it, this arm goes to zero -/// and the assertion below reds. +/// `pop[nyc]` link this arm is live and scored; for the OTHER links every +/// occurrence of their source is frozen and so is the clock (GH #1016: the +/// partial reads `PREVIOUS(TIME)`), so the arm is provably `PREVIOUS(growth)` +/// and is omitted to an exact `+0.0` -- the row that shows the clock is a +/// frozen input rather than the thing that keeps a source-free arm alive. /// * `boston` reads `alt[a1]` -- a source in a dimension DISJOINT from the /// target's, subscripted by a bare element name. This is the ACCESS SHAPE in /// which GH #977's 322 unwrapped-bare-variable arms arise (raw @@ -221,27 +219,36 @@ fn ltm_slot_values_are_pinned_on_the_value_gate_fixture() { assert_value_golden("value_gate", &render_slab(&series)); } -/// Mechanism 1: an arm with NO live source reference but a live `TIME` must be -/// materialized and must carry a non-zero value. +/// Mechanism 1: an arm with NO live source reference whose only varying +/// content is the clock is a structural zero, because the clock is a frozen +/// input of the partial (GH #1016). /// /// The `alt[a1] -> growth` link's `nyc` slot is that arm: `pop[nyc]` and `base` /// are frozen for this link, `alt[a1]` does not appear in the `nyc` equation at -/// all, and what remains live is `TIME * 0.002`. Under the negative "the -/// source stayed frozen" criterion this slot would be dropped to zero; under -/// the positive predicate `TIME` is `BuiltinReach::Varying`, so the arm stays. +/// all, and `TIME * 0.002` is read as `PREVIOUS(TIME) * 0.002` -- so the +/// partial is `PREVIOUS(growth[nyc])` term for term, the positive predicate +/// establishes it, and the slot is omitted to an exact `+0.0`. The same arm for +/// the `pop[nyc]` link reads its source live and is scored, which is what +/// shows the omission is the arm's, not the fixture's. /// -/// This assertion does not depend on the golden, which is the point: a careless -/// `UPDATE_LTM_VALUE_GOLDEN=1` re-capture would bless the zeroed slot, and this -/// would still red. +/// Neither assertion depends on the golden, which is the point: a careless +/// `UPDATE_LTM_VALUE_GOLDEN=1` re-capture would bless either a materialized +/// clock-driven score or a zeroed live arm, and this would still red. #[test] -fn a_time_bearing_arm_with_no_live_source_is_not_zeroed() { +fn a_time_bearing_arm_with_no_live_source_is_a_structural_zero() { let series = ltm_slot_series(<m_value_gate_project()); // Region declaration order: nyc=0, boston=1, la=2. - let nyc = slot(&series, "link_score\u{205A}alt[a1]\u{2192}growth", 0); + let nyc_for_alt = slot(&series, "link_score\u{205A}alt[a1]\u{2192}growth", 0); assert!( - nyc.iter().any(|v| v.abs() > 1e-12 && v.is_finite()), - "the TIME-bearing `nyc` arm was zeroed: an arm whose only live content \ - is a time-dependent builtin is NOT a structural zero; got {nyc:?}" + nyc_for_alt.iter().all(|v| v.to_bits() == 0.0f64.to_bits()), + "the `nyc` arm of the alt[a1] link reads no live source and the clock \ + is frozen, so it is omitted to an exact +0.0; got {nyc_for_alt:?}" + ); + let nyc_for_pop = slot(&series, "link_score\u{205A}pop[nyc]\u{2192}growth", 0); + assert!( + nyc_for_pop.iter().any(|v| v.abs() > 1e-12 && v.is_finite()), + "the same arm reads pop[nyc] live for its own link and is scored; \ + got {nyc_for_pop:?}" ); } diff --git a/src/simlin-engine/src/db/ltm_value_golden/value_gate.txt b/src/simlin-engine/src/db/ltm_value_golden/value_gate.txt index bf7a11b0a..042520c94 100644 --- a/src/simlin-engine/src/db/ltm_value_golden/value_gate.txt +++ b/src/simlin-engine/src/db/ltm_value_golden/value_gate.txt @@ -1,4 +1,5 @@ -$⁚ltm⁚link_score⁚alt[a1]→growth[0] 0.000000000000e0 6.666666666667e-1 6.600660066007e-1 6.535306996046e-1 6.470600986184e-1 6.406535629885e-1 +$⁚ltm⁚freeze⁚time[0] 0.000000000000e0 0.000000000000e0 1.000000000000e0 2.000000000000e0 3.000000000000e0 4.000000000000e0 +$⁚ltm⁚link_score⁚alt[a1]→growth[0] 0.000000000000e0 0.000000000000e0 0.000000000000e0 0.000000000000e0 0.000000000000e0 0.000000000000e0 $⁚ltm⁚link_score⁚alt[a1]→growth[1] 0.000000000000e0 1.000000000000e0 1.000000000000e0 1.000000000000e0 1.000000000000e0 1.000000000000e0 $⁚ltm⁚link_score⁚alt[a1]→growth[2] 0.000000000000e0 0.000000000000e0 0.000000000000e0 0.000000000000e0 0.000000000000e0 0.000000000000e0 $⁚ltm⁚link_score⁚growth→pop[0] 0.000000000000e0 1.000000000000e0 1.000000000000e0 1.000000000000e0 1.000000000000e0 1.000000000000e0 @@ -6,14 +7,14 @@ $⁚ltm⁚link_score⁚growth→pop[1] 0.000000000000e0 1.000000000000e0 1.00000 $⁚ltm⁚link_score⁚growth→pop[2] 0.000000000000e0 0.000000000000e0 0.000000000000e0 0.000000000000e0 0.000000000000e0 0.000000000000e0 $⁚ltm⁚link_score⁚pop[boston]→pop_total[0] 0.000000000000e0 1.578947368421e-1 1.649162354629e-1 1.715697797158e-1 1.778798792994e-1 1.838689119179e-1 $⁚ltm⁚link_score⁚pop[la]→pop_total[0] 0.000000000000e0 3.157894736842e-1 3.073927967621e-1 2.993785051921e-1 2.917210936525e-1 2.843972770067e-1 -$⁚ltm⁚link_score⁚pop[nyc]→growth[0] 0.000000000000e0 1.000000000000e0 1.000000000000e0 1.000000000000e0 1.000000000000e0 1.000000000000e0 +$⁚ltm⁚link_score⁚pop[nyc]→growth[0] 0.000000000000e0 3.333333333333e-1 3.399339933993e-1 3.464693003954e-1 3.529399013816e-1 3.593464370115e-1 $⁚ltm⁚link_score⁚pop[nyc]→growth[1] 0.000000000000e0 0.000000000000e0 0.000000000000e0 0.000000000000e0 0.000000000000e0 0.000000000000e0 $⁚ltm⁚link_score⁚pop[nyc]→growth[2] 0.000000000000e0 0.000000000000e0 0.000000000000e0 0.000000000000e0 0.000000000000e0 0.000000000000e0 $⁚ltm⁚link_score⁚pop[nyc]→pop_total[0] 0.000000000000e0 5.263157894737e-1 5.276909677750e-1 5.290517150920e-1 5.303990270480e-1 5.317338110755e-1 -$⁚ltm⁚link_score⁚pop_total→alt[a1][0] 0.000000000000e0 1.000000000000e0 1.000000000000e0 1.000000000000e0 1.000000000000e0 1.000000000000e0 -$⁚ltm⁚link_score⁚pop_total→alt[a2][0] 0.000000000000e0 1.000000000000e0 1.000000000000e0 1.000000000000e0 1.000000000000e0 1.000000000000e0 -$⁚ltm⁚loop_score⁚r1[0] 0.000000000000e0 1.578947368421e-1 1.649162354629e-1 1.715697797158e-1 1.778798792994e-1 1.838689119179e-1 -$⁚ltm⁚loop_score⁚r2[0] 0.000000000000e0 1.000000000000e0 1.000000000000e0 1.000000000000e0 1.000000000000e0 1.000000000000e0 +$⁚ltm⁚link_score⁚pop_total→alt[a1][0] 0.000000000000e0 8.675799086758e-2 8.891713245778e-2 9.108062465248e-2 9.324847077762e-2 9.542067375910e-2 +$⁚ltm⁚link_score⁚pop_total→alt[a2][0] 0.000000000000e0 8.675799086758e-2 8.891713245778e-2 9.108062465248e-2 9.324847077762e-2 9.542067375910e-2 +$⁚ltm⁚loop_score⁚r1[0] 0.000000000000e0 1.369863013699e-2 1.466387875309e-2 1.562668270800e-2 1.658702672678e-2 1.754489545855e-2 +$⁚ltm⁚loop_score⁚r2[0] 0.000000000000e0 3.333333333333e-1 3.399339933993e-1 3.464693003954e-1 3.529399013816e-1 3.593464370115e-1 $⁚ltm⁚net⁚pop[0] 1.000000000000e-1 1.030000000000e-1 1.060300000000e-1 1.090903000000e-1 1.121812030000e-1 1.153030150300e-1 $⁚ltm⁚net⁚pop[1] 3.000000000000e-2 3.219000000000e-2 3.438519000000e-2 3.658560519000e-2 3.879128109519e-2 4.100225357929e-2 $⁚ltm⁚net⁚pop[2] 6.000000000000e-2 6.000000000000e-2 6.000000000000e-2 6.000000000000e-2 6.000000000000e-2 6.000000000000e-2 diff --git a/src/simlin-engine/src/ltm_augment.rs b/src/simlin-engine/src/ltm_augment.rs index c5f2f30f7..a1629ffd7 100644 --- a/src/simlin-engine/src/ltm_augment.rs +++ b/src/simlin-engine/src/ltm_augment.rs @@ -11,7 +11,7 @@ //! value is at the step after the start. use crate::ast::{Expr0, IndexExpr0, print_eqn}; -use crate::builtins::UntypedBuiltinFn; +use crate::builtins::{BuiltinSig, Invariance, UntypedBuiltinFn}; use crate::canonicalize; use crate::common::{Canonical, Ident, RawIdent}; use crate::datamodel::{self, Equation}; @@ -106,6 +106,36 @@ fn is_array_reducer_name(name: &str, arity: usize) -> bool { crate::ltm_agg::reducer_kind_from_name(&name.to_ascii_lowercase(), arity).is_some() } +/// Does `name` call a builtin that reads the clock -- `TIME`, `STEP`, `RAMP`, +/// `PULSE`? Decided by the signature's `Invariance`, the compiler's own +/// statement of which builtins vary with time (`compiler::invariance` reads +/// the same field to refuse hoisting them), so the ceteris-paribus freeze and +/// the hoisting classifier cannot disagree about what a clock read is. The run +/// constants `DT`, `INITIAL_TIME` and `FINAL_TIME` are `Pure`: they do not +/// change across a run, so freezing them would change nothing. +fn is_time_dependent_builtin(name: &str) -> bool { + BuiltinSig::by_name(&name.to_ascii_lowercase()) + .is_some_and(|sig| sig.invariance == Invariance::TimeDependent) +} + +/// The frozen read of a time-dependent `call` (GH #1016): a bare `TIME` in a +/// value position is the per-model helper [`FROZEN_CLOCK_HELPER`] +/// (`PREVIOUS(TIME)`, one value for the whole model, so one variable rather +/// than a capture per arm); every other form -- `STEP`/`RAMP`/`PULSE`, whose +/// arguments make each call its own, and a `TIME` in a subscript index, whose +/// first-DT value must be the un-lagged index -- is the call lagged whole by +/// [`freeze_at_previous`], arguments verbatim. +fn freeze_clock_read(call: Expr0, loc: crate::builtins::Loc, in_subscript_index: bool) -> Expr0 { + if !in_subscript_index + && let Expr0::App(UntypedBuiltinFn(name, args), _) = &call + && args.is_empty() + && name.eq_ignore_ascii_case("time") + { + return Expr0::Var(RawIdent::new_from_str(FROZEN_CLOCK_HELPER), loc); + } + freeze_at_previous(call, loc, in_subscript_index) +} + /// Whether any subexpression of `expr` prints exactly as `reducer_text` -- the /// [`WrapCtx::live_reducer_text`] containment test (Track A stage 1, finding 2). /// @@ -267,7 +297,10 @@ use freeze::freeze_at_previous; #[path = "ltm_augment_array_freeze.rs"] mod array_freeze; -pub(crate) use array_freeze::{ArrayFreezeHelper, FREEZE_HELPER_PREFIX, materialize_array_freezes}; +pub(crate) use array_freeze::{ + ArrayFreezeHelper, FREEZE_HELPER_PREFIX, FROZEN_CLOCK_HELPER, frozen_clock_helper, + materialize_array_freezes, +}; /// Deciding when a per-element link-score arm is provably `PREVIOUS(target)` /// and may therefore be OMITTED rather than materialized (GH #977), in its own @@ -361,6 +394,36 @@ fn other_dep_verdict( /// even when the outer subscript matches the live shape, so nested /// references like `arr[other_var]` still get wrapped. /// +/// The clock is a frozen input too (GH #1016). The paper's partial +/// `f(x_t, y_{t-1}) - z_{t-1}` reads every input but the isolated one at the +/// previous step, and `TIME` is an input of `f` like any other: with it live, +/// a source with no influence on its target scores +/-1 whenever an exogenous +/// forcing moves the target, and the forcing is credited to every link into +/// it. So a call of a time-dependent builtin ([`is_time_dependent_builtin`]: +/// `TIME`, `STEP`, `RAMP`, `PULSE`) is read at the previous step +/// ([`freeze_clock_read`]: a bare `TIME` through the per-model helper +/// [`FROZEN_CLOCK_HELPER`], any other call wrapped WHOLE in `PREVIOUS` with +/// its arguments verbatim -- the capture lags the whole call by one step, and +/// lagging a frozen dep inside it as well would read that dep two steps back) +/// -- unless the call holds an occurrence of the isolated input's LIVE SHAPE +/// (the occurrence IR's verdict, `subtree_has_live_shape`, the same one the +/// reducer arm takes; a read of another element of an arrayed input, or an +/// index-nested read, is not one), in which case the call stays live, clock +/// included: the isolated input and the clock are then inseparable inside one +/// call, and the score attributes that call's whole change to the input (the +/// stated residual, pinned by `tests/integration/ltm_frozen_clock.rs`). Inside +/// a subtree the wrap is about to lag (`frozen`: a frozen dep's subscript +/// index) the call is left verbatim, because the enclosing synthesized freeze +/// already lags it once -- a capture evaluates the whole `arr[TIME]` at `t` +/// and `PREVIOUS` reads it a step back, index included -- and lagging the +/// clock again would read `arr_{t-1}[TIME_{t-2}]`. (A model dependency in +/// that position IS lagged twice, `PREVIOUS(q[PREVIOUS(ctr, ctr)])`, the +/// residual `ltm_augment_zero_slot` documents; the clock does not join it.) +/// In a LIVE reference's index the frozen clock takes the un-lagged read as +/// its first-DT value like every other index freeze ([`freeze_at_previous`]). +/// The guard form's own `TIME = INITIAL_TIME` arm is outside the partial and +/// reads the clock live. +/// /// `ctx.iter_ctx` carries the GH #511 iterated-dimension context (the live /// source's declared dimension names + the target equation's iterated /// dimensions + a `DimensionsContext` for the mapped case). An @@ -723,6 +786,43 @@ fn wrap_non_matching_in_previous( None => call, }; } + // Whether the reducer held LIVE by `live_reducer_text` sits inside + // this call's arguments (Track A stage 1, finding 2; see the reducer + // arm below, which shares the answer). + let holds_live_reducer = live_reducer_text + .is_some_and(|text| args.iter().any(|a| expr0_contains_reducer_text(a, text))); + // The clock is a frozen input (GH #1016; see the rustdoc): a + // time-dependent call holding no live-shape occurrence of the + // isolated input is read at the previous step, arguments verbatim + // -- unless the wrap is about to lag the subtree it sits in, where + // that enclosing freeze already lags it once. Frozen whole or left + // to the enclosing freeze, the wrap never descends into it, so the + // `PerElement` pin-only descent reaches the source references it + // may still hold (another element's, an index-nested one), exactly + // as for a frozen-whole reducer. + if is_time_dependent_builtin(&name) + && !holds_live_reducer + && !ctx + .occ + .subtree_has_live_shape(path, live_source, live_shape) + { + let call = Expr0::App(UntypedBuiltinFn(name, args), loc); + let call = match ctx.pin { + Some(pin_ctx) => post_transform::pin_only_source_refs( + call, + pin_ctx, + ctx.occ, + path, + &mut out.missing_occurrence, + ), + None => call, + }; + return if frozen { + call + } else { + freeze_clock_read(call, loc, in_subscript_index) + }; + } // A LOOKUP call's first argument names a graphical-function table // (a lookup-only variable, or the WITH-LOOKUP self-reference); the // table HEAD is static data the compiler resolves to a table id, not @@ -806,8 +906,6 @@ fn wrap_non_matching_in_previous( // slice partial). The top-of-function guard has already declined to // hold THIS reducer live (its own text does not equal // `live_reducer_text`), so this only affects an enclosing reducer. - let holds_live_reducer = live_reducer_text - .is_some_and(|text| args.iter().any(|a| expr0_contains_reducer_text(a, text))); if is_array_reducer_name(&name, args.len()) && !holds_live_reducer && !ctx @@ -1355,7 +1453,11 @@ fn contains_unfreezable_previous(expr: &Expr0) -> bool { /// untouched (already lagged/frozen; double-wrapping would read two steps /// back). Non-matching occurrences of `live_source` -- and all other /// references -- stay current: their influence is attributed by their own -/// link-score variables. +/// link-score variables. The clock stays current too, deliberately: the +/// changed-last partial `z(x_t, w_t) - z(x_{t-1}, w_t)` reads every other +/// input, `TIME` included, at `t` on both sides, so the clock's own motion +/// cancels and only the isolated input's change is attributed (the +/// changed-first dual freezes it instead; GH #1016). /// /// Boundary: unlike its changed-first dual, this walker never recurses /// into subscript INDEX expressions -- so every wrap it emits is the @@ -1724,7 +1826,9 @@ fn shaped_guard_form_text( /// "feeder frozen" evaluation of a hoisted reducer's equation. References /// already inside a `PREVIOUS(...)`/`INIT(...)` call are left untouched /// (their contents are already lagged/frozen; double-wrapping would read -/// two steps back). Subscript index expressions are recursed into so a +/// two steps back), and the clock stays live: this is the changed-last +/// convention, which reads every input but the feeder at `t` on both sides +/// of its numerator (see [`wrap_live_shaped_in_previous`]). Subscript index expressions are recursed into so a /// `arr[target + 1]` style index reference is frozen too; the outer /// subscripted variable itself is wrapped only when it names `target` /// (defensive -- the feeder this is used for is scalar and so is always a @@ -4664,9 +4768,10 @@ pub(crate) struct ReducerBodyCtx<'a> { /// against the row's axis; an unprovable correspondence bails. pub arrayed_dep_dims: &'a HashMap, /// Every model-variable ident the body may reference -- the freeze set. - /// References to idents NOT in this set (TIME, function names resolved - /// as `App`s, dimension/element names) stay live, matching - /// `build_partial_equation_shaped`'s deps-only freezing convention. + /// References to idents NOT in this set (function names resolved as + /// `App`s, dimension/element names) stay live; the clock is frozen by the + /// builtin rule ([`is_time_dependent_builtin`]), not by membership here, + /// matching `build_partial_equation_shaped`'s convention. pub model_deps: &'a HashSet, /// Canonical dimension names of the live source's axes, in declared /// order -- parallel to the row tuple. @@ -4971,7 +5076,14 @@ fn pin_body_to_row(expr: Expr0, ctx: &ReducerBodyCtx<'_>, row_parts: &[String]) /// Wrap every model-variable reference of a row-pinned body in /// `PREVIOUS()` (the value-position form of [`freeze_at_previous`]), except -/// occurrences of `keep_live` (when given). Subscript +/// occurrences of `keep_live` (when given), and every clock read the same +/// way -- a time-dependent call ([`is_time_dependent_builtin`]) whose +/// arguments do not read `keep_live` is read at the previous step +/// ([`freeze_clock_read`]), the rule [`wrap_non_matching_in_previous`] +/// states (GH #1016). Freezing the +/// clock alongside the model references is what keeps the all-frozen row +/// terms equal to the ones that produced `PREVIOUS(agg)`, the anchor identity +/// the nonlinear body partial rests on (GH #763). Subscript /// indices are never recursed into: on an arrayed MODEL dep's subscript, /// pinning has already replaced them with literal qualified elements (not /// causal references); on a non-model head (whose expression indices @@ -4984,6 +5096,14 @@ fn freeze_pinned_body(expr: Expr0, freeze: &HashSet, keep_live: Option<& let c = canonicalize(ident); freeze.contains(c.as_ref()) && Some(c.as_ref()) != keep_live }; + // By ident, which in a ROW-PINNED body is the occurrence-level answer the + // wrap walker takes from the IR: every surviving reference to the source + // is the row's own element (`pin_body_to_row` bails on a fixed-literal + // self-reference), so "the call names the source" is "the call reads the + // live row". + let reads_live = |arg: &Expr0| -> bool { + keep_live.is_some_and(|live| expr_reference_idents(arg).contains(live)) + }; match expr { Expr0::Const(..) => expr, Expr0::Var(ref ident, loc) => { @@ -5004,6 +5124,13 @@ fn freeze_pinned_body(expr: Expr0, freeze: &HashSet, keep_live: Option<& if name.eq_ignore_ascii_case("previous") || name.eq_ignore_ascii_case("init") { return Expr0::App(UntypedBuiltinFn(name, args), loc); } + if is_time_dependent_builtin(&name) && !args.iter().any(reads_live) { + return freeze_clock_read( + Expr0::App(UntypedBuiltinFn(name, args), loc), + loc, + false, + ); + } let args = args .into_iter() .map(|a| freeze_pinned_body(a, freeze, keep_live)) @@ -5294,14 +5421,16 @@ fn generate_linear_body_partial( /// keeps the GH #483 unrolled population-variance form (divisor `N`, /// inlined mean) over the body terms. /// -/// Anchor caveat (GH #763): "frozen" freezes MODEL references only, so a -/// body referencing TIME, a time builtin (PULSE/STEP/RAMP), or a nested -/// `PREVIOUS(x)` keeps that factor live in every term, and then -/// `R(all-frozen terms) != PREVIOUS(agg)` -- the anchor subtraction -/// attributes the time-drift to every row, including rows whose true -/// partial is 0 (destroying the frozen-argmin-scores-0 property of -/// MIN/MAX). For pure-model-ref bodies the anchor identity holds exactly -/// because per-variable `PREVIOUS` sampling commutes with arithmetic. +/// Anchor identity: "frozen" freezes the model references AND the clock +/// ([`freeze_pinned_body`], GH #1016), so an all-frozen term is the +/// previous step's evaluation of that row and `R(all-frozen terms) = +/// PREVIOUS(agg)` exactly, because per-variable `PREVIOUS` sampling commutes +/// with arithmetic; a row that is never the argmin therefore scores 0 +/// (`tests/integration/ltm_frozen_clock.rs`, the GH #763 repro). The one +/// residue is an ORIGINAL `PREVIOUS(x)` in the body: it is left untouched, +/// so the frozen term reads `x(t-1)` where the anchor read `x(t-2)`, and the +/// misalignment is attributed to every row -- the same lag-misalignment +/// class `ltm_augment_zero_slot` documents. /// /// When the pinned body is the bare source reference the legacy /// [`generate_nonlinear_partial`] is returned byte-identically; RANK is diff --git a/src/simlin-engine/src/ltm_augment_array_freeze.rs b/src/simlin-engine/src/ltm_augment_array_freeze.rs index 56cdb3216..27d89b19d 100644 --- a/src/simlin-engine/src/ltm_augment_array_freeze.rs +++ b/src/simlin-engine/src/ltm_augment_array_freeze.rs @@ -101,6 +101,30 @@ pub(crate) struct ArrayFreezeHelper { /// The reserved name prefix for materialized freeze helpers. pub(crate) const FREEZE_HELPER_PREFIX: &str = "$\u{205A}ltm\u{205A}freeze\u{205A}"; +/// The one freeze helper every model shares: `$⁚ltm⁚freeze⁚time = +/// PREVIOUS(TIME)`, the clock read at the previous step (GH #1016). +/// +/// A frozen clock is one value for the whole model, so a partial reads this +/// helper rather than spelling `PREVIOUS(TIME)` inline: the inline spelling is +/// a capture per arm and per occurrence (`snapshot_arg` routes a builtin +/// argument through a capture), which on C-LEARN v77 cost 9,632 result slots +/// -- a third of the row -- for one number. Minted once per model by +/// `db::ltm::model_ltm_variables` when any arm references it, and sorted with +/// the other freeze helpers ahead of every score. Its name is under +/// [`FREEZE_HELPER_PREFIX`] so the companion dedup and the evaluation-order +/// category treat it as the freeze helper it is. +pub(crate) const FROZEN_CLOCK_HELPER: &str = "$\u{205A}ltm\u{205A}freeze\u{205A}time"; + +/// The [`FROZEN_CLOCK_HELPER`] as the scalar one-arm helper the emission loop +/// registers through `freeze_helper_var`. +pub(crate) fn frozen_clock_helper() -> ArrayFreezeHelper { + ArrayFreezeHelper { + name: FROZEN_CLOCK_HELPER.to_string(), + dims: Vec::new(), + arms: vec![(String::new(), "PREVIOUS(TIME)".to_string())], + } +} + /// One axis of a materializable slice, post-classification. enum FreezeAxis { /// A statically-kept index: the verbatim text baked into every arm. @@ -850,6 +874,20 @@ mod tests { /// Two identical freezes in one expression share one helper; the helper /// name survives canonicalization unchanged (assembly canonicalizes every /// LTM var name, so a name that mutates there would orphan its fragment). + /// The clock helper must sort and dedup as a freeze helper, which is a + /// property of its NAME. + #[test] + fn the_frozen_clock_helper_is_named_under_the_freeze_prefix() { + assert!(FROZEN_CLOCK_HELPER.starts_with(FREEZE_HELPER_PREFIX)); + let helper = frozen_clock_helper(); + assert_eq!(helper.name, FROZEN_CLOCK_HELPER); + assert!(helper.dims.is_empty()); + assert_eq!( + helper.arms, + vec![(String::new(), "PREVIOUS(TIME)".to_string())] + ); + } + #[test] fn identical_freezes_dedup_and_names_are_canonical() { let (dep_dims, ctx) = dims_fixture(); diff --git a/src/simlin-engine/src/ltm_augment_pin_tests.rs b/src/simlin-engine/src/ltm_augment_pin_tests.rs index 00958ef35..193d821d6 100644 --- a/src/simlin-engine/src/ltm_augment_pin_tests.rs +++ b/src/simlin-engine/src/ltm_augment_pin_tests.rs @@ -1436,12 +1436,16 @@ fn per_element_pin_index_verdict_enumeration() { // not) inside an `Expr0::App` without a fourth copy of the builtin // classification `builtins`/`compiler::invariance` own -- so it // declined conservatively. It no longer has to decide: the rule keeps - // the index either way, and the wrap's index pass leaves a 0-arity - // builtin live exactly as it does everywhere else in a partial (TIME - // is not a dep being isolated, and the guard form reads it live too). + // the index either way, and what the wrap's index pass does with it + // is the builtin's own classification (`Invariance::TimeDependent`, + // read through `ltm_augment::is_time_dependent_builtin`): the clock is + // a frozen input of the partial (GH #1016), so the bare column's + // runtime read is lagged with the un-lagged index as its first-DT + // value, exactly like `idx` two rows down, while inside the + // pre-existing freeze it is already lagged and left alone. "a 0-arity builtin index", "pop[Region, TIME]", - Some("pop[region\u{B7}boston, time()]"), + Some("pop[region\u{B7}boston, PREVIOUS(time(), time())]"), Some("pop[region\u{B7}boston, time()]"), ), // --- unspellable: a COMPILABILITY verdict, so loud in BOTH contexts ---- diff --git a/src/simlin-engine/src/ltm_augment_tests.rs b/src/simlin-engine/src/ltm_augment_tests.rs index f4f463960..5fa4b9a79 100644 --- a/src/simlin-engine/src/ltm_augment_tests.rs +++ b/src/simlin-engine/src/ltm_augment_tests.rs @@ -2268,6 +2268,151 @@ fn partial_equation_unknown_ident_unchanged() { assert!(!partial.contains("PREVIOUS(unknown)"), "partial: {partial}"); } +// -- GH #1016: the clock is a frozen input of a changed-first partial -- +// +// The paper's partial reads every input but the isolated one at the previous +// step, and the clock is an input like any other. The rule is one owner: a +// builtin whose signature is `Invariance::TimeDependent` (`TIME`, `STEP`, +// `RAMP`, `PULSE`) is read at the previous step -- a bare `TIME` as the +// per-model helper `$⁚ltm⁚freeze⁚time = PREVIOUS(TIME)`, a call as +// `PREVIOUS()` with its arguments verbatim (a capture lags the whole +// call once; freezing inside as well would read two steps back) -- unless the +// call's arguments read the live source, in +// which case the call stays live, clock included (the stated residual: the +// isolated input and the clock are then inseparable in one call). The run +// constants `DT`, `INITIAL_TIME`, `FINAL_TIME` are `Invariance::Pure` and are +// left alone: freezing them would not change them. + +/// The rows enumerate `Invariance`'s two classes a partial can meet -- the +/// `TimeDependent` forms (bare `TIME`; a call of constants; a call of a frozen +/// dep; a call reading the live source; a call reading ANOTHER element of the +/// source, which is not the live shape; the other two time-dependent +/// builtins; a `TIME` and a call inside a FROZEN dep's index, left to that +/// enclosing freeze) and the `Pure` run constants. The `Lagged` and +/// `Snapshot` classes are the pre-existing `PREVIOUS`/`INIT` passthrough, +/// pinned by `partial_equation_does_not_rewrap_inside_previous`; the clock +/// inside a LIVE reference's index (`PREVIOUS(time(), time())`) is the +/// "a 0-arity builtin index" row of +/// `pin_tests::per_element_pin_index_verdict_enumeration`. +#[test] +fn partial_equation_freezes_the_clock() { + let deps = deps_set(&["pop", "helper", "arr"]); + let live = Ident::::new("pop"); + let shape = RefShape::Bare; + let dims = region_dim_elements(); + let rows: [(&str, &str, &str); 8] = [ + ( + "a bare TIME reads the per-model frozen-clock helper", + "pop + TIME", + "pop + \"$\u{205A}ltm\u{205A}freeze\u{205A}time\"", + ), + ( + "a time-dependent call of constants, wrapped whole", + "pop + STEP(2, 2)", + "pop + PREVIOUS(step(2, 2))", + ), + ( + "a time-dependent call of a frozen dep: the call is lagged once, its \ + argument is not lagged again", + "pop + STEP(helper, 2)", + "pop + PREVIOUS(step(helper, 2))", + ), + ( + "a time-dependent call reading the live source keeps its clock live", + "STEP(pop, 2) + helper", + "step(pop, 2) + PREVIOUS(helper)", + ), + ( + "a time-dependent call reading ANOTHER element of the source is not \ + the live shape and is frozen whole", + "pop + STEP(pop[nyc], 2)", + "pop + PREVIOUS(step(pop[nyc], 2))", + ), + ( + "RAMP and PULSE are time-dependent too", + "pop + RAMP(1, 0) + PULSE(1, 2, 3)", + "pop + PREVIOUS(ramp(1, 0)) + PREVIOUS(pulse(1, 2, 3))", + ), + ( + "a TIME index of a FROZEN dep is left to the enclosing freeze, which \ + lags the whole read once (lagging it again would read \ + arr_{t-1}[TIME_{t-2}])", + "pop + arr[TIME]", + "pop + PREVIOUS(arr[time()])", + ), + ( + "a time-dependent call in a FROZEN dep's index, likewise", + "pop + arr[STEP(2, 2)]", + "pop + PREVIOUS(arr[step(2, 2)])", + ), + ]; + for (label, eqn, want) in rows { + let partial = + build_partial_equation_shaped(eqn, &deps, &live, &shape, &dims, None, None).unwrap(); + assert_eq!(partial, want, "{label}"); + } + // The run constants are `Invariance::Pure`: nothing to freeze. + let partial = build_partial_equation_shaped( + "pop * DT + INITIAL_TIME + FINAL_TIME", + &deps, + &live, + &shape, + &dims, + None, + None, + ) + .unwrap(); + assert!( + !partial.contains("PREVIOUS("), + "the run constants are not clock reads; got: {partial}" + ); +} + +/// The changed-LAST dual leaves the clock live, as it leaves every other +/// input live: `Delta_x z = z(x_t, w_t) - z(x_{t-1}, w_t)` reads the clock at +/// `t` on both sides, so only the feeder is lagged in the frozen evaluation. +#[test] +fn scalar_feeder_changed_last_keeps_the_clock_live() { + let eq = generate_scalar_feeder_to_agg_equation( + "scale", + "$\u{205A}ltm\u{205A}agg\u{205A}0", + &expr("sum(pop[*] * scale * TIME)"), + None, + ); + assert!( + eq.contains("sum(pop[*] * PREVIOUS(scale) * time())"), + "the feeder-frozen evaluation keeps the clock live; got: {eq}" + ); + assert!(!eq.contains("PREVIOUS(time"), "got: {eq}"); +} + +/// A reducer body reading the clock (GH #763): the row-pinned terms freeze +/// the clock with the model references, so the all-frozen terms reproduce +/// `PREVIOUS(agg)` and the anchor identity holds. +#[test] +fn nonlinear_body_partial_freezes_the_clock_in_every_term() { + let elements = vec!["region·nyc".to_string(), "region·boston".to_string()]; + let fixture = BodyCtxFixture::new("pop[*] * TIME", "pop", &[("pop", 1)], &[], &["region"]); + let eq = generate_element_to_scalar_equation( + "pop", + "total", + "region·nyc", + &elements, + &ReducerKind::Nonlinear, + "MIN", + true, + Some(&fixture.ctx()), + None, + ); + let clock = "\"$\u{205A}ltm\u{205A}freeze\u{205A}time\""; + assert!( + eq.contains(&format!( + "MIN((pop[region·nyc] * {clock}), (PREVIOUS(pop[region·boston]) * {clock}))" + )), + "got: {eq}" + ); +} + // -- GH #311: parse failure must be a loud error, never a silent // semantics-changing fallback -- // @@ -5281,6 +5426,46 @@ fn shaped_guard_form_falls_back_to_changed_last_for_unfreezable_co_source() { ); } +/// The changed-last leg leaves the clock LIVE (GH #1016): its numerator +/// `z(x_t, w_t, T_t) - z(x_{t-1}, w_t, T_t)` reads the clock at `t` on both +/// sides, so only the isolated input is lagged in the frozen evaluation -- +/// the same GH #743 shape with `+ TIME` on the target, where the frozen +/// evaluation keeps `time()` and neither the clock helper nor a +/// `PREVIOUS(time())` appears. +#[test] +fn shaped_guard_form_changed_last_keeps_the_clock_live() { + let deps = deps_set(&["matrix", "frac"]); + let live = Ident::::new("frac"); + let source_dims = vec![vec!["r1".to_string(), "r2".to_string()]]; + let source_dim_names = vec!["d1".to_string()]; + let target_iterated = vec!["d1".to_string()]; + let iter_ctx = IteratedDimCtx { + source_dim_names: &source_dim_names, + target_iterated_dims: &target_iterated, + dep_dims: None, + }; + let text = sgft( + "SUM(matrix[D1, *] * frac[D1]) + TIME", + &deps, + &live, + &RefShape::Bare, + &source_dims, + &source_dim_names, + Some(&iter_ctx), + None, + "growth", + None, + ) + .unwrap(); + assert_eq!( + text, + "if (TIME = INITIAL_TIME) then 0 \ + else if ((growth - PREVIOUS(growth)) = 0) OR ((frac - PREVIOUS(frac)) = 0) then 0 \ + else SAFEDIV((growth - (sum(matrix[d1, *] * PREVIOUS(frac)) + time())), \ + ABS((growth - PREVIOUS(growth))), 0) * SIGN((frac - PREVIOUS(frac)))" + ); +} + /// A shape where the changed-first partial is freezable stays byte-identical /// to the historical output: the chooser is a pure pass-through to the /// changed-first guard form whenever that form compiles. diff --git a/src/simlin-engine/src/ltm_augment_zero_slot.rs b/src/simlin-engine/src/ltm_augment_zero_slot.rs index 199860ddc..f6d5549ea 100644 --- a/src/simlin-engine/src/ltm_augment_zero_slot.rs +++ b/src/simlin-engine/src/ltm_augment_zero_slot.rs @@ -9,6 +9,9 @@ use crate::ast::{Expr0, IndexExpr0}; use crate::builtins::UntypedBuiltinFn; +use crate::canonicalize; + +use super::FROZEN_CLOCK_HELPER; /// Whether the caller's result is a whole VARIABLE's equation or one slot of an /// `Ast::Arrayed` one -- which is the only thing that decides whether a @@ -70,13 +73,14 @@ pub(crate) enum ZeroSlotPolicy { /// /// The tempting negative test -- "the link's source /// stayed frozen" -- says nothing about what else the arm reads, and - /// collapsing on it changes 187 C-LEARN result slots across 35 link-score - /// variables (151 by >= 1.0, worst 8,086.97 -> 0), because the wrap does not - /// freeze everything that varies: a live `time()` remains, and a + /// collapsing on it rewrote real scores to zero when it was tried on + /// C-LEARN, because the wrap does not freeze everything that varies: a /// raw-vs-canonical element-spelling mismatch can leave the source itself - /// unwrapped. Those are tracked as #1016 and the wrap defects in #977; this - /// predicate is correct whether or not they are fixed, because it asks about - /// the emitted tree rather than about the wrap's bookkeeping. + /// unwrapped (the wrap defects in #977). This predicate is correct whether + /// or not that is fixed, because it asks about the emitted tree rather + /// than about the wrap's bookkeeping: a frozen clock reads the clock helper + /// or `PREVIOUS()`, which the walk sees as the one-step lag it is, + /// and a live read is a live read. OmitStructuralZero, } @@ -139,8 +143,11 @@ enum BuiltinReach { /// `lookup` is deliberately `Varying` even though a graphical function is a /// compile-time constant: it would only matter for an arm whose lookup index is /// itself invariant, and GH #977 measured that relaxation as buying **exactly -/// zero** additional arms on C-LEARN (those arms hit a live `time()` inside the -/// lookup's own index immediately afterwards). Conservative and free. +/// zero** additional arms on C-LEARN, measured with the clock live in every +/// partial and not redone under the frozen clock, where a lookup's argument +/// reads the clock helper and a table-head index reads +/// `PREVIOUS(time(), time())`; a relaxation would also have to treat the table +/// HEAD, a `Var`, as static. Conservative and free. fn classify_builtin_reach(name: &str) -> BuiltinReach { // Lowercased at parse time, but classify case-insensitively so a future // caller with raw source spelling cannot silently fall into `Varying`. @@ -186,7 +193,13 @@ fn classify_builtin_reach(name: &str) -> BuiltinReach { /// A `Var` or `Subscript` reached outside a frozen subtree is a live read and /// ends the walk, which is why subscript INDICES are never descended into: the /// whole reference is already `NotEstablished`, so `IndexExpr0` needs no arm -/// here and a new index variant cannot change any verdict. +/// here and a new index variant cannot change any verdict. The one `Var` that +/// is not a live read is the frozen clock's helper ([`FROZEN_CLOCK_HELPER`]), +/// whose definition IS `PREVIOUS(TIME)`: the wrap spells a frozen `TIME` as a +/// reference to it rather than inline, and it is the one-step lag the anchor +/// expects, so it is `Established` exactly as the inline `PREVIOUS(TIME)` +/// would be. No other freeze helper is recognized here: their arms are +/// one-step lags too, but that relaxation is unmeasured and not taken. pub(super) fn partial_is_provably_previous_target(original: &Expr0, partial: &Expr0) -> bool { !contains_previous_call(original) && reach_of(partial) == Reach::Established } @@ -226,8 +239,13 @@ fn contains_previous_call(expr: &Expr0) -> bool { fn reach_of(expr: &Expr0) -> Reach { match expr { Expr0::Const(..) => Reach::Established, - // A live read of model state: the value it yields this step is exactly - // what the wrap was supposed to freeze and did not. + // The frozen clock's helper is `PREVIOUS(TIME)` by definition (see the + // rustdoc); every other bare reference is a live read of model state, + // the value it yields this step being exactly what the wrap was + // supposed to freeze and did not. + Expr0::Var(name, _) if canonicalize(name.as_str()).as_ref() == FROZEN_CLOCK_HELPER => { + Reach::Established + } Expr0::Var(..) => Reach::NotEstablished, Expr0::Subscript(..) => Reach::NotEstablished, Expr0::Op1(_, inner, _) => reach_of(inner), diff --git a/src/simlin-engine/tests/integration/ltm_frozen_clock.rs b/src/simlin-engine/tests/integration/ltm_frozen_clock.rs new file mode 100644 index 000000000..a16b9af62 --- /dev/null +++ b/src/simlin-engine/tests/integration/ltm_frozen_clock.rs @@ -0,0 +1,319 @@ +// Copyright 2026 The Simlin Authors. All rights reserved. +// Use of this source code is governed by the Apache License, +// Version 2.0, that can be found in the LICENSE file. + +//! The clock is a frozen input of a ceteris-paribus partial (GH #1016). +//! +//! The 2020 paper's partial `Delta_x z = f(x_t, y_{t-1}) - z_{t-1}` reads +//! every input but the isolated one at the previous step, and the clock is an +//! input like any other: `TIME` and the time-dependent builtins (`STEP`, +//! `RAMP`, `PULSE`) are read at the previous step inside the partial. What +//! that buys, in order: +//! +//! * a source with no influence on its target scores 0 even while an +//! exogenous forcing moves the target +//! (`a_source_with_no_influence_scores_zero_under_a_step_forcing`); +//! * an exogenous forcing on an isolated loop is NOT credited to the loop +//! (`an_exogenous_ramp_takes_its_share_of_an_isolated_loop`); +//! * a target that reads the clock directly scores the paper's partial, +//! `x_t * TIME_{t-1} - z_{t-1}` +//! (`a_target_reading_time_scores_the_changed_first_partial`); +//! * a reducer body reading the clock keeps the `PREVIOUS(agg)` anchor, so a +//! row that is never the argmin scores 0 (GH #763, +//! `a_time_bearing_reducer_body_keeps_the_frozen_argmin_at_zero`). +//! +//! The stated residual: a time-dependent call whose arguments read the live +//! source keeps its clock live for that call +//! (`a_time_call_reading_the_live_source_keeps_its_clock_live`). And the +//! boundary of the freeze: a clock read inside a frozen dependency's subscript +//! index is lagged exactly once, by the enclosing freeze +//! (`a_clock_read_in_a_frozen_deps_index_is_lagged_once`). + +use simlin_engine::datamodel::Project; +use simlin_engine::test_common::TestProject; + +use crate::test_helpers::{ltm_run, ltm_series}; + +/// `$⁚ltm⁚link_score⁚{from}→{to}`. +fn link_score(from: &str, to: &str) -> String { + format!("$\u{205A}ltm\u{205A}link_score\u{205A}{from}\u{2192}{to}") +} + +/// `$⁚ltm⁚loop_score⁚{id}`. +fn loop_score(id: &str) -> String { + format!("$\u{205A}ltm\u{205A}loop_score\u{205A}{id}") +} + +fn assert_series_eq(got: &[f64], want: &[f64], tol: f64, what: &str) { + assert_eq!(got.len(), want.len(), "{what}: length"); + for (t, (g, w)) in got.iter().zip(want).enumerate() { + assert!( + (g - w).abs() < tol, + "{what} at step {t}: got {g}, want {w} (got {got:?}, want {want:?})" + ); + } +} + +/// `x = 5 + STEP(2, 2) + 0 * s` on the structural loop `s -> x -> s`: `s` +/// has no influence on `x`, and the step forcing moves `x` at t = 2. +fn inert_source_under_step() -> Project { + TestProject::new("inert_step") + .with_sim_time(0.0, 4.0, 1.0) + .stock("s", "0", &["x"], &[], None) + .flow("x", "5 + STEP(2, 2) + 0 * s", None) + .build_datamodel() +} + +#[test] +fn a_source_with_no_influence_scores_zero_under_a_step_forcing() { + let run = ltm_run(&inert_source_under_step(), false); + let x = ltm_series(&run.results, "x", 0); + // The forcing is real: x jumps from 5 to 7 at t = 2. + assert_series_eq(&x, &[5.0, 5.0, 7.0, 7.0, 7.0], 1e-12, "x"); + // s changes at every step (its inflow is never 0), so the zero below is + // the partial's, not the guard's Δsource = 0 arm. + let s = ltm_series(&run.results, "s", 0); + assert!(s.windows(2).all(|w| w[1] > w[0]), "s must move: {s:?}"); + let score = ltm_series(&run.results, &link_score("s", "x"), 0); + assert_series_eq(&score, &[0.0; 5], 1e-12, "LS(s -> x)"); +} + +/// The isolated loop `s -> f -> s` with `f = 0.1 * s + RAMP(1, 0)`. +fn ramped_isolated_loop() -> Project { + TestProject::new("ramped_loop") + .with_sim_time(0.0, 5.0, 1.0) + .stock("s", "100", &["f"], &[], None) + .flow("f", "0.1 * s + RAMP(1, 0)", None) + .build_datamodel() +} + +#[test] +fn an_exogenous_ramp_takes_its_share_of_an_isolated_loop() { + let run = ltm_run(&ramped_isolated_loop(), false); + let s = ltm_series(&run.results, "s", 0); + let f = ltm_series(&run.results, "f", 0); + let s_to_f = ltm_series(&run.results, &link_score("s", "f"), 0); + let f_to_s = ltm_series(&run.results, &link_score("f", "s"), 0); + let loop_id = run.loop_through("f").to_string(); + let loop_series = ltm_series(&run.results, &loop_score(&loop_id), 0); + + // Hand calculation: the partial holds the ramp at its previous value, so + // the loop's share of Δf is 0.1 * Δs and the ramp's slope (1 per time + // unit, dt = 1) is the rest. + let mut want = vec![0.0]; + for t in 1..s.len() { + let ds = s[t] - s[t - 1]; + let df = f[t] - f[t - 1]; + want.push(0.1 * ds / df.abs()); + } + assert_series_eq(&s_to_f, &want, 1e-12, "LS(s -> f)"); + // The pinned values: 0.500, 0.545, 0.587, 0.624, 0.658 at t = 1..5. + assert_series_eq( + &s_to_f, + &[0.0, 0.500, 0.545, 0.587, 0.624, 0.658], + 5e-4, + "LS(s -> f) rounded", + ); + // The single inflow scores 1 into its stock, so the loop score IS the + // stock-to-flow share. + assert_series_eq( + &f_to_s, + &[0.0, 1.0, 1.0, 1.0, 1.0, 1.0], + 1e-12, + "LS(f -> s)", + ); + assert_series_eq(&loop_series, &want, 1e-12, "loop score"); +} + +/// `z = x * TIME` on the loop `s -> x -> z -> s`. +fn target_reading_time() -> Project { + TestProject::new("time_target") + .with_sim_time(0.0, 5.0, 1.0) + .stock("s", "10", &["z"], &[], None) + .aux("x", "2 + 0.01 * s", None) + .flow("z", "x * TIME", None) + .build_datamodel() +} + +#[test] +fn a_target_reading_time_scores_the_changed_first_partial() { + let run = ltm_run(&target_reading_time(), false); + let x = ltm_series(&run.results, "x", 0); + let z = ltm_series(&run.results, "z", 0); + let time = ltm_series(&run.results, "time", 0); + let score = ltm_series(&run.results, &link_score("x", "z"), 0); + + // The paper's changed-first numerator with the clock frozen: + // N = x_t * TIME_{t-1} - z_{t-1}, + // scored as SAFEDIV(N, |Δz|, 0) * SIGN(Δx). With the clock live the + // numerator would be x_t * TIME_t - z_{t-1} = Δz, a score of 1. + let mut want = vec![0.0]; + let mut live_clock_would_give = vec![0.0]; + for t in 1..z.len() { + let dz = z[t] - z[t - 1]; + let dx = x[t] - x[t - 1]; + let n = x[t] * time[t - 1] - z[t - 1]; + want.push(n / dz.abs() * dx.signum()); + live_clock_would_give.push((x[t] * time[t] - z[t - 1]) / dz.abs() * dx.signum()); + } + assert_series_eq(&score, &want, 1e-12, "LS(x -> z)"); + // The two conventions are distinguishable on this model at every step + // after the first: the clock-live numerator is the whole of Δz. + for t in 2..z.len() { + assert!( + (live_clock_would_give[t] - 1.0).abs() < 1e-12 && (want[t] - 1.0).abs() > 0.1, + "step {t}: frozen-clock {} vs live-clock {}", + want[t], + live_clock_would_give[t] + ); + } +} + +/// `grow = 1 + MIN(pop[*] * TIME)` (GH #763's repro) on a loop through both +/// rows: `south` starts larger and grows at the same rate, so it is never +/// the argmin. +fn time_bearing_min_body() -> Project { + TestProject::new("min_time_body") + .with_sim_time(0.0, 4.0, 1.0) + .named_dimension("Region", &["north", "south"]) + .array_stock("pop[Region]", "10", &["growth"], &[], None) + .array_flow("growth[Region]", "pop * 0.1 * grow", None) + .aux("grow", "1 + MIN(pop[*] * TIME)", None) + .build_datamodel() +} + +#[test] +fn a_time_bearing_reducer_body_keeps_the_frozen_argmin_at_zero() { + let mut project = time_bearing_min_body(); + // south starts at 20: the same equations, a larger start, never the argmin. + for var in &mut project.models[0].variables { + if let simlin_engine::datamodel::Variable::Stock(stock) = var + && stock.ident == "pop" + { + stock.equation = simlin_engine::datamodel::Equation::Arrayed( + vec!["Region".to_string()], + vec![ + ("north".to_string(), "10".to_string(), None, None), + ("south".to_string(), "20".to_string(), None, None), + ], + None, + false, + ); + } + } + let run = ltm_run(&project, false); + // The hoisted MIN is `$⁚ltm⁚agg⁚0`, scored per source element. + let agg = "$\u{205A}ltm\u{205A}agg\u{205A}0"; + let north = ltm_series(&run.results, &link_score("pop[north]", agg), 0); + let south = ltm_series(&run.results, &link_score("pop[south]", agg), 0); + // At t = 1 the frozen clock reads TIME = 0, so every term is 0 and the + // whole change of MIN is the clock's: north scores from t = 2. + assert!( + north[2..].iter().all(|v| v.abs() > 1e-9), + "north is the argmin at every step and carries the score: {north:?}" + ); + assert_series_eq(&south, &[0.0; 5], 1e-12, "LS(pop[south] -> MIN)"); +} + +/// `z = x + arr[TIME]` with `arr = [1, 2, 4, 8, 16]` on the loop +/// `s -> x -> z -> s`, run from t = 1 so `TIME` indexes `arr` directly. +fn frozen_dep_indexed_by_time() -> Project { + let mut project = TestProject::new("time_index") + .with_sim_time(1.0, 5.0, 1.0) + .named_dimension("Slot", &["a1", "a2", "a3", "a4", "a5"]) + .array_aux("arr[Slot]", "1") + .stock("s", "10", &["z"], &[], None) + .aux("x", "0.1 * s", None) + .flow("z", "x + arr[TIME]", None) + .build_datamodel(); + for var in &mut project.models[0].variables { + if let simlin_engine::datamodel::Variable::Aux(aux) = var + && aux.ident == "arr" + { + aux.equation = simlin_engine::datamodel::Equation::Arrayed( + vec!["Slot".to_string()], + ["1", "2", "4", "8", "16"] + .iter() + .enumerate() + .map(|(i, v)| (format!("a{}", i + 1), v.to_string(), None, None)) + .collect(), + None, + false, + ); + } + } + project +} + +/// `arr` is a frozen dependency of `z` for the `x -> z` link, so the partial +/// reads `PREVIOUS(arr[TIME])`: one capture evaluates `arr[TIME]` at `t` and +/// `PREVIOUS` reads it a step back, index included, which is the paper's +/// `arr_{t-1}[TIME_{t-1}]`. Lagging the clock inside as well would read +/// `arr_{t-1}[TIME_{t-2}]` and flip the score's sign from the second scored +/// step. +#[test] +fn a_clock_read_in_a_frozen_deps_index_is_lagged_once() { + let run = ltm_run(&frozen_dep_indexed_by_time(), false); + let x = ltm_series(&run.results, "x", 0); + let z = ltm_series(&run.results, "z", 0); + let score = ltm_series(&run.results, &link_score("x", "z"), 0); + let arr = [1.0, 2.0, 4.0, 8.0, 16.0]; + // The partial is x_t + arr[TIME_{t-1}] and the anchor z_{t-1} is + // x_{t-1} + arr[TIME_{t-1}], so the numerator is Δx. + let mut want = vec![0.0]; + let mut double_lag_would_give = vec![0.0]; + for t in 1..z.len() { + let dx = x[t] - x[t - 1]; + let dz = z[t] - z[t - 1]; + want.push(dx / dz.abs() * dx.signum()); + // TIME at step t is t + 1 (the run starts at 1); arr[TIME_{t-2}] is + // arr[t - 1] 1-based, i.e. index t - 2, undefined at the first step. + if t >= 2 { + let n = dx + arr[t - 2] - arr[t - 1]; + double_lag_would_give.push(n / dz.abs() * dx.signum()); + } else { + double_lag_would_give.push(want[t]); + } + } + assert_series_eq(&score, &want, 1e-12, "LS(x -> z)"); + assert_series_eq( + &score, + &[0.0, 0.1667, 0.1379, 0.1213, 0.1118], + 5e-5, + "LS(x -> z) rounded", + ); + // The double-lag form is a different, sign-flipped number from t = 3. + for t in 2..z.len() { + assert!( + double_lag_would_give[t] < 0.0 && want[t] > 0.0, + "step {t}: once-lagged {} vs twice-lagged {}", + want[t], + double_lag_would_give[t] + ); + } +} + +/// `f = STEP(0.1 * s, 2)` on the isolated loop: the call reads the live +/// source, so its clock stays live (the stated residual), and the step of the +/// forcing is credited to the loop. +fn time_call_reading_the_source() -> Project { + TestProject::new("step_of_source") + .with_sim_time(0.0, 4.0, 1.0) + .stock("s", "100", &["f"], &[], None) + .flow("f", "1 + STEP(0.1 * s, 2)", None) + .build_datamodel() +} + +#[test] +fn a_time_call_reading_the_live_source_keeps_its_clock_live() { + let run = ltm_run(&time_call_reading_the_source(), false); + let f = ltm_series(&run.results, "f", 0); + let s_to_f = ltm_series(&run.results, &link_score("s", "f"), 0); + // f is 1 until the step fires at t = 2, then 1 + 0.1 * s. + assert!((f[1] - 1.0).abs() < 1e-12 && f[2] > 10.0, "f: {f:?}"); + // At t = 2 the partial is 1 + STEP(0.1 * s_t, 2) with the clock live, so + // the whole jump is attributed to s: the score is 1 (the residual). From + // t = 3 the step is on in both the partial and the anchor and the score + // is the ordinary 0.1 * Δs / |Δf| = 1. + assert_series_eq(&s_to_f, &[0.0, 0.0, 1.0, 1.0, 1.0], 1e-12, "LS(s -> f)"); +} diff --git a/src/simlin-engine/tests/integration/main.rs b/src/simlin-engine/tests/integration/main.rs index 162babd16..52c4762c5 100644 --- a/src/simlin-engine/tests/integration/main.rs +++ b/src/simlin-engine/tests/integration/main.rs @@ -34,6 +34,7 @@ mod ltm_array_agg; mod ltm_discovery_large_models; mod ltm_dt_invariance; mod ltm_flow_to_stock; +mod ltm_frozen_clock; mod ltm_integration_method; // Compares xmutil-based MDL parsing against the native Rust parser, so it // needs the optional xmutil C++ converter compiled in. diff --git a/src/simlin-engine/tests/integration/simulate.rs b/src/simlin-engine/tests/integration/simulate.rs index 11c40ef5d..f49fc49e4 100644 --- a/src/simlin-engine/tests/integration/simulate.rs +++ b/src/simlin-engine/tests/integration/simulate.rs @@ -6915,6 +6915,22 @@ fn corpus_clearn_macros_import() { /// or one of those six scalars; the margin is 36,811 free against the /// 65,536-slot ceiling. /// +/// The frozen clock (GH #1016) moved the count UP, 6,224 -> 6,227 (+3), and +/// the width UP, 28,725 -> 28,980 (+255). The count is the per-model clock +/// helper `$⁚ltm⁚freeze⁚time` (`examples/ltm_var_dump.rs`: one each in +/// `main`, the `ramp_from_to` macro model and the stdlib `npv` template, +/// the three models whose partials read `TIME`). The width is the +/// result-column diff of a C-LEARN `simlin simulate --ltm` run on the +/// previous and the new CLI, every added column one of three kinds and +/// nothing removed: 36 clock-helper instances (`main`'s plus one per +/// `ramp_from_to` call site), their 36 `PREVIOUS(TIME)` captures, and 183 +/// captures of time-dependent CALLS frozen whole (`PREVIOUS(STEP(..))` and +/// the like), one per arm and per occurrence; a call inside a frozen +/// dependency's subscript index is left to that enclosing freeze and mints +/// none. Without the shared helper the bare `TIME` reads alone cost 9,632 +/// capture slots, which is why the helper exists. The margin is 36,556 free +/// against the 65,536-slot ceiling. +/// /// The pin below catches emission changes in EITHER direction, and re-deriving /// it means re-measuring BOTH numbers, not just the count. #[test] @@ -6939,7 +6955,7 @@ fn clearn_ltm_var_count_guardrail() { }) .sum(); assert_eq!( - total, 6224, + total, 6227, "C-LEARN's emitted LTM var count moved; if this is an intentional \ emission change, re-derive the layout-slot impact (the #654 \ ceiling) and update this pin with the new numbers" diff --git a/src/simlin-engine/tests/integration/simulate_ltm.rs b/src/simlin-engine/tests/integration/simulate_ltm.rs index 4aea4b9b8..a9814d78e 100644 --- a/src/simlin-engine/tests/integration/simulate_ltm.rs +++ b/src/simlin-engine/tests/integration/simulate_ltm.rs @@ -11258,6 +11258,26 @@ fn test_whole_rhs_mapped_reducer_routes_through_synthetic_agg() { /// when the omission is disabled" would be equally consistent with a digest that /// cannot see the omission at all. /// +/// Two emission changes moved the pin, measured one commit apart so each +/// carries its own numbers (this gate is `#[ignore]`d and release-only, so +/// neither the hook nor `cargo test` runs it; the first change landed without +/// re-pinning it and the second re-derived both): +/// +/// * the flow-to-stock score's net-flow aux: slots 20,221 -> 20,337 (+116: +/// the 24 `$⁚ltm⁚net⁚{stock}` auxes of `main`, one slot per stock element, +/// while the two arrayed scores it added and the six per-element scalars it +/// removed are a wash), `nonzero_slots` 3,106 -> 3,198 (+92, the net-aux +/// slots whose stock has a moving flow), every slot still finite. +/// * the frozen clock (GH #1016): slots 20,337 -> 20,338 (+1, the +/// `$⁚ltm⁚freeze⁚time` helper), `nonzero_slots` 3,198 -> 2,876 (-322: the +/// arms whose only varying content was the clock, which now read +/// `PREVIOUS(TIME)` and are provably `PREVIOUS(target)`, so they are +/// omitted to an exact zero -- the DROP this gate exists to catch, here +/// derived rather than absorbed), every slot still finite. The 146 +/// link-score series that change on a `simlin simulate --ltm` run lie on +/// no retained loop: the 153 loops, their relative-score series and the 565 +/// dominant periods `analyze()` reports are identical before and after. +/// /// Run with: /// cargo test -p simlin-engine --release --test integration -- --ignored \ /// clearn_ltm_slot_maxima_digest @@ -11572,10 +11592,10 @@ fn the_digest_sees_both_a_value_swap_and_a_rebinding() { } /// Pinned by `clearn_ltm_slot_maxima_digest`; see its rustdoc before changing. -const CLEARN_LTM_SLOTS: usize = 20_221; +const CLEARN_LTM_SLOTS: usize = 20_338; const CLEARN_LTM_UNKNOWN_EXTENT: usize = 0; -const CLEARN_LTM_NONZERO_SLOTS: usize = 3_106; -const CLEARN_LTM_FINITE_SLOTS: usize = 20_221; -const CLEARN_LTM_MANTISSA_DIGEST: i64 = 790_401_758_590; -const CLEARN_LTM_EXPONENT_DIGEST: i64 = 2_212; -const CLEARN_LTM_IDENTITY_DIGEST: u64 = 16_953_901_100_024_641_861; +const CLEARN_LTM_NONZERO_SLOTS: usize = 2_876; +const CLEARN_LTM_FINITE_SLOTS: usize = 20_338; +const CLEARN_LTM_MANTISSA_DIGEST: i64 = 763_629_105_850; +const CLEARN_LTM_EXPONENT_DIGEST: i64 = 2_449; +const CLEARN_LTM_IDENTITY_DIGEST: u64 = 10_121_477_288_851_905_900; From 765b7a3b6b3db764333f881d87ff79122900a019 Mon Sep 17 00:00:00 2001 From: Bobby Powers Date: Wed, 9 Sep 2026 20:48:09 -0700 Subject: [PATCH 09/10] doc: single-mode LTM design plan The 2026-09-07 LTM audit established that both hero models auto-flip to discovery, so the exhaustive path's compile-time loop list is empty on every model that matters and the TS engine has no discovery binding; that the post-simulation product of a loop's recorded link scores reproduces the in-simulation loop_score column to one ulp on all seven fixtures; and that all-edge instrumentation costs no more than loop-edge instrumentation plus per-loop score variables. This plan makes the post-simulation pipeline the only loop-scoring path, keeps Johnson only for the structural preview and ids, injects pins into the pipeline, treats an inline reducer as a real loop node, deletes static polarity and the shortest-path fallback, and gates sub-model pathway emission on SCC membership. The plan embeds five product decisions (ids without a polarity letter, element-level loops as the reported unit, named-aggregate reducer semantics, static polarity deleted, synthetic CLI columns hidden) and its acceptance criteria are drafted, not yet validated with the owner; both are called out in the document. --- docs/README.md | 1 + .../2026-09-07-ltm-single-mode.md | 350 ++++++++++++++++++ 2 files changed, 351 insertions(+) create mode 100644 docs/design-plans/2026-09-07-ltm-single-mode.md diff --git a/docs/README.md b/docs/README.md index 25d052dad..2d2e6fd9f 100644 --- a/docs/README.md +++ b/docs/README.md @@ -31,6 +31,7 @@ - [design-plans/](design-plans/) -- Design plans (architecture and phasing for major efforts) - [design-plans/2026-08-25-compiler-unification.md](design-plans/2026-08-25-compiler-unification.md) -- Engine compiler de-duplication, landed: one fragment compiler (`lower_fragment`/`DepShape`), one `BuiltinFn::signature` table, one temp allocator and materialization pass, one axis matcher (`match_axes`), structured `DepRef`s, one `Variable`, one `Diagnostic`, a `(variable, project)`-keyed parse with AST-carried captures and implicit modules, per-variable lowered memos, Loops That Matter as a consumer of each; every pinned semantic divergence and a per-commit measured ledger - [design-plans/2026-09-04-link-scores-from-fragments.md](design-plans/2026-09-04-link-scores-from-fragments.md) -- LTM link scores synthesized from the target's compiled fragment instead of generated equation text: the rewrite, its per-link keying, the formula families as typed builders, one `LinkScore` opcode, a native post-pass; the roofline for "LTM always on" + - [design-plans/2026-09-07-ltm-single-mode.md](design-plans/2026-09-07-ltm-single-mode.md) -- Single-mode LTM: score every causal edge, compute every loop score after the run from the recorded series (one pipeline for every surface), Johnson only for the structural preview and ids, pins injected into the pipeline, a reducer as a real loop node, static polarity and the shortest-path fallback deleted, sub-model pathway scores emitted only for instances that can lie on a loop; 8 phases - [design-plans/2026-08-26-compiler-unification-phase7-investigation.md](design-plans/2026-08-26-compiler-unification-phase7-investigation.md) -- Companion to the compiler unification plan: every parse-time decision that reads model state, where each moves under a `(variable, project)` parse key, every text/name identity site of synthesized helpers, the salsa keys, the capture runlist contract - [design-plans/2026-08-17-pysimlin-widget.md](design-plans/2026-08-17-pysimlin-widget.md) -- pysimlin file-backed models + anywidget in-notebook editor (file on disk as sync authority) - [design-plans/2026-04-05-server-rewrite.md](design-plans/2026-04-05-server-rewrite.md) -- Local-first `simlin-serve` binary: filesystem-backed editor + in-process MCP diff --git a/docs/design-plans/2026-09-07-ltm-single-mode.md b/docs/design-plans/2026-09-07-ltm-single-mode.md new file mode 100644 index 000000000..090441474 --- /dev/null +++ b/docs/design-plans/2026-09-07-ltm-single-mode.md @@ -0,0 +1,350 @@ +# Single-mode LTM Design + +## Summary + +Simlin implements Loops That Matter (LTM), a method that measures how much each feedback loop in a system dynamics model accounts for the model's behavior at each point in time. The engine does this two different ways today. The exhaustive path enumerates loops at compile time and synthesizes one extra simulation variable per loop, which computes that loop's score as the simulation runs. The discovery path instead instruments every causal link and reconstructs the loops afterward from the recorded link-score series. A model-size gate picks between them automatically, and both of the large real models the project measures itself against (World3 and C-LEARN) trip it. So the exhaustive path only ever runs on small test fixtures, every consumer that reads its compile-time loop list is empty on the models that matter, and the TypeScript engine has no binding to the other path at all. + +This plan deletes the exhaustive path and makes the post-simulation pipeline the only one. An audit established that the post-simulation product of a loop's link scores reproduces the in-simulation loop-score column to one ulp on all seven LTM fixtures, so the in-simulation form carries nothing worth keeping; those columns are captured as goldens before the code producing them is removed, and the replacement is asserted against them. Collapsing to one mode also removes the reason six decisions each had two implementations (cycle partitions, loop deduplication, score normalization, polarity, the per-exit-port module override, and cross-aggregate loop recovery), so most phases are consolidations onto a single owner rather than new machinery. Along the way the plan drops a sampling fallback whose measured recall was 0.20 on the only model dense enough to need it, replaces compile-time polarity analysis with classification from the run, makes an inline array reducer a real reported node instead of stitching loops back together around it, and stops emitting module instrumentation for instances that cannot lie on a loop. The bar for landing is that every surface reports loops on every model, the scores match the retained goldens, and compile time, run time and slot count stay within five percent of today's figures. + +## Definition of Done + +Each item states what is true of the tree when this plan has landed, with the command or test that shows it. + +1. **One instrumentation.** An LTM compile scores every causal edge of every model; there is no loop-edge-only instrumentation, no `$⁚ltm⁚loop_score⁚` synthetic variable, no `ltm_discovery_mode` input, no `LtmMode` enum, no auto-flip gate and no auto-flip warning. `rg "loop_score⁚|ltm_discovery_mode|LtmMode|MAX_LTM_SCC_NODES|auto-switched" src/` is empty outside `docs/`. +2. **One loop-scoring pipeline.** Every consumer that reports loop scores (libsimlin, pysimlin, MCP, `@simlin/engine`, the CLI `--ltm` report, the layout) reads them from one post-simulation pipeline over the recorded link-score series (`ltm_finding::discover_loops_with_graph` and its successors). Per-loop, per-slot raw and partition-relative series are bit-identical to the retained goldens of the seven LTM fixtures captured from the exhaustive `loop_score` columns before this plan lands (`tests/integration/ltm_single_mode_parity.rs`). +3. **Every surface gets loops on every model.** World3 and C-LEARN report loops through libsimlin `simlin_analyze_get_loops_runtime`, pysimlin `Run.loops`, MCP `read_model`, `@simlin/engine` `Run.loops`, and the CLI `--ltm` report, with the same ids and scores on each (`tests/integration/ltm_discovery_large_models.rs`, `src/engine/tests/wasm-ltm.test.ts`, `src/pysimlin/tests/test_ltm.py`). +4. **Pins are reported regardless of retention or the cap**, scored by the same pipeline, with the user's name and a stable id (`tests/integration/simulate_ltm_pinned.rs`). +5. **The shortest-path fallback is deleted.** On budget exhaustion the partial universe is ranked and reported with `enumeration_complete == false`; `src/simlin-engine/src/ltm_finding_fallback.rs`, its tests, `examples/ltm_fallback_eval.rs`, `FallbackConfig`, `ENUM_BUDGET_FRACTION` and `fallback_candidates` do not exist. +6. **One owner per decision.** Cycle partitions come from `db::analysis::model_element_cycle_partitions`; loop identity from one element-level canonical key in `ltm/mod.rs`; relative normalization from one function in `ltm_post.rs`; runtime polarity from `LoopPolarity::from_runtime_scores` over the relative series; the per-exit-port module override from `ltm_finding.rs`; each verified by `rg` in the phase that consolidates it. +7. **An inline reducer is a node.** A loop through `SUM(pop[*])` is an elementary circuit of the element graph in which the reducer's `$⁚ltm⁚agg⁚{n}` node is a real, reported node (displayed as the reducer's spelling); there is no petal stitching, no `MAX_AGG_PETALS`, no `MAX_CROSS_AGG_LOOPS`, no `agg_recovery_truncated`. The three-element `test/cross_agg_ltm` fixture reports three loops with relative scores summing to 1 (`tests/integration/ltm_array_agg.rs`). +8. **Static polarity is gone.** `ltm/polarity.rs` and its tests do not exist; a loop's polarity is a field derived from its runtime relative series, `None` before a run; loop ids carry no polarity letter. +9. **Sub-model pathway and composite scores are emitted only for module instances that can lie on a feedback loop** (an instance whose node is in a nontrivial SCC of its parent's causal graph). `Theil_2011.mdl` (zero loops, many SMOOTH instances) compiles under the overlay to at most 1.5x its plain slot count (`examples/ltm_full_bench.rs`). +10. **Cost and memory on the corpus are not worse than today.** For every model in `docs/design/engine-performance.md`'s ledger the LTM compile time, run time and slot count are within 5 percent of the pre-plan figures, and C-LEARN's discovery pass stays under 100 ms (ledger rows in this document). + +## Acceptance Criteria + +Drafted from the Definition of Done; to be validated with the owner before an implementation plan is written from them. + +### ltm-single-mode.AC1: One instrumentation +- **ltm-single-mode.AC1.1 Success:** `model_ltm_variables` on `test/logistic_growth_ltm` emits exactly the seven link-score variables of today's discovery instrumentation and no loop-score variable. +- **ltm-single-mode.AC1.2 Success:** Compiling World3 under the overlay emits no diagnostic (today: the auto-flip warning). +- **ltm-single-mode.AC1.3 Success:** `SourceProject` has no `ltm_discovery_mode` field; `analyze_model` sets no input and opens no revision (`db::exec_probe::ProbedDb` shows zero re-executed `model_ltm_variables` bodies on a second `analyze_model` call on the same db). +- **ltm-single-mode.AC1.4 Failure:** A model whose LTM compile fails (an unfreezable partial in every link of a loop) still simulates plainly; the failure is one `Warning` naming the link, never an empty loop list without a reason. + +### ltm-single-mode.AC2: One loop-scoring pipeline +- **ltm-single-mode.AC2.1 Success:** For each of the seven fixtures (`logistic_growth_ltm`, `arms_race_3party`, `decoupled_stocks`, `arrayed_population_ltm`, `cross_element_ltm`, `cross_agg_ltm`, `hero_culture_ltm`), the post-simulation raw loop score of every loop at every saved step equals the retained exhaustive golden to within one ulp, and the partition-relative score equals the golden recomputed under per-partition normalization. +- **ltm-single-mode.AC2.2 Success:** libsimlin, pysimlin, MCP and `@simlin/engine` report identical loop ids, polarities and relative series for `hero_culture` (cross-surface test). +- **ltm-single-mode.AC2.3 Success:** The layout's loop importance (`layout::detect_ltm_loops`) is computed from the same results the display run produces, not from a second simulation (`ProbedDb` shows one `Vm::run_to_end` per layout on a loop-bearing model). +- **ltm-single-mode.AC2.4 Edge:** A stateless model (no stocks, no lagged deps) emits no LTM variables and reports an empty loop list with `enumeration_complete == true`. + +### ltm-single-mode.AC3: Every surface, every model +- **ltm-single-mode.AC3.1 Success:** World3 reports 200 loops (the coverage-aware cap) on every surface; C-LEARN reports 153. +- **ltm-single-mode.AC3.2 Success:** `@simlin/engine` `Model.run({analyzeLtm: true})` on World3 under the `'vm'` and `'wasm'` engines returns the same loops as libsimlin (the wasm engine reaches the pipeline through `simlin_analyze_discover_loops_from_wasm_results`). +- **ltm-single-mode.AC3.3 Success:** pysimlin `Model.run()` on World3 returns a non-empty `Run.loops` and emits no `RuntimeWarning`. +- **ltm-single-mode.AC3.4 Success:** `Run.links` on the TS surface excludes `$⁚` internal nodes by default (`includeInternal` defaults to false, matching pysimlin). + +### ltm-single-mode.AC4: Pins +- **ltm-single-mode.AC4.1 Success:** A pinned loop below `MIN_CONTRIBUTION` at every step is still reported, with its relative series (near 0) and its user name. +- **ltm-single-mode.AC4.2 Success:** A pinned loop that the enumerator also finds is reported once, under the pin's name, with the enumerator's id. +- **ltm-single-mode.AC4.3 Failure:** A pin naming a variable set with no closed cycle is reported in `PinnedLoopsResult::invalid` with the pin's identity, as today. +- **ltm-single-mode.AC4.4 Success:** An arrayed pin reports one loop per element instance, each with its own series. + +### ltm-single-mode.AC5: Fallback deleted +- **ltm-single-mode.AC5.1 Success:** With `MAX_DISCOVERY_ENUM_CIRCUITS` lowered by the test-only `EnumBudgetGuard` to a value below World3's universe, discovery reports `enumeration_complete == false`, a non-empty ranked list drawn from the partial universe, and `universe_loops == None`. +- **ltm-single-mode.AC5.2 Success:** The wall-clock budget is spent entirely on the enumeration; an expired deadline yields the partial result of AC5.1. + +### ltm-single-mode.AC6: One owner +- **ltm-single-mode.AC6.1 Success:** `rg -n "compute_cycle_partitions|model_cycle_partitions\b" src/simlin-engine/src` finds only `model_element_cycle_partitions` and its readers. +- **ltm-single-mode.AC6.2 Success:** Two circuits that are rotations of each other, in exhaustive Johnson output, in a pin expansion and in discovery, all resolve to one key from `ltm::canonical_cycle_key`. +- **ltm-single-mode.AC6.3 Success:** The relative series of any loop, read through libsimlin's per-element accessor, pysimlin, and the layout, are the same bytes (one normalization function). + +### ltm-single-mode.AC7: Reducer is a node +- **ltm-single-mode.AC7.1 Success:** `test/cross_agg_ltm` reports exactly three loops, `pop[x] -> SUM(pop) -> growth[x] -> pop[x]` for x in {a, b, c}, each with raw score 1/3 and relative score 1/3. +- **ltm-single-mode.AC7.2 Success:** The de-subscripted twin of that model with the SUM named (`total = SUM(pop[*])`) reports the same three loops with the same scores (the oracle in the audit's scratchpad, promoted to `tests/integration/ltm_desubscript_oracle.rs`). +- **ltm-single-mode.AC7.3 Success:** A `migration_matrix`-class model (3x3 flow matrix through one reducer) completes with no truncation flag and a loop count linear in the element count. + +### ltm-single-mode.AC8: Static polarity gone +- **ltm-single-mode.AC8.1 Success:** Before a run, `Model.get_loops()` (pysimlin) and `simlin_analyze_get_loops` report loops with `polarity == None` and ids without a letter prefix; after a run, every loop in the corpus carries `Reinforcing` or `Balancing` with confidence 1.0 except the yeast model's R loop (`Undetermined`). +- **ltm-single-mode.AC8.2 Success:** A loop's id is the same before and after the run and across two runs with different parameters. + +### ltm-single-mode.AC9: Module emission gated on loop membership +- **ltm-single-mode.AC9.1 Success:** `Theil_2011.mdl` under the overlay has at most 1.5x its plain slot count and emits no pathway or composite variables. +- **ltm-single-mode.AC9.2 Success:** `smooth3.mdl` (a SMOOTH inside a loop) still emits the pathway and composite variables of the instances on the loop, and their loop scores match the retained goldens. + +### ltm-single-mode.AC10: Cost +- **ltm-single-mode.AC10.1 Success:** The ledger rows for every corpus model are within 5 percent of the pre-plan compile, run and slot figures; C-LEARN discovery under 100 ms. + +## Glossary + +- **LTM (Loops That Matter)**: The feedback-loop dominance method Simlin implements (Schoenberg, Eberlein et al.). It quantifies, at each timestep, how much of a model's observed behavior each feedback loop accounts for. Full write-up in `docs/reference/ltm--loops-that-matter.md`. +- **Link score**: A dimensionless per-timestep measure of one causal link's contribution to the change in its target variable, carrying both a magnitude and a sign. Computed by re-evaluating the target's equation with every input except the one under test held at its previous value. +- **Loop score**: The product of the link scores around a closed loop. Its sign is the loop's polarity; its magnitude is the "force" the loop exerts on the stocks it touches. A loop alone on its stocks always scores exactly plus or minus 1. +- **Relative loop score**: A loop score normalized by the sum of absolute loop scores of all loops in its cycle partition. Lands in [-1, 1], and the absolute values within a partition sum to 1. This is what consumers plot, because raw scores diverge toward infinity near dominance shifts where competing loops cancel. +- **Polarity (reinforcing / balancing / undetermined)**: An even number of negative links makes a loop reinforcing; an odd number makes it balancing; a link of unknown sign makes it undetermined. Static polarity is derived from the equation AST at compile time; runtime polarity is classified from the sign of the recorded score series, with a confidence value. This plan deletes static polarity. +- **Cycle partition**: A group of stocks connected to each other through feedback, computed as a strongly connected component of the stock-to-stock reachability graph. Relative scores are only meaningful within a partition. +- **SCC (strongly connected component)**: A maximal set of graph nodes where every node reaches every other. Every feedback loop lies entirely inside one SCC, so a node in no nontrivial SCC cannot be on a loop. +- **Element graph**: The causal graph after arrayed variables are expanded to one node per array element, so `pop[nyc]` and `pop[boston]` are distinct nodes and reported loops are element-specific. +- **Slot**: One element position of an arrayed variable, and equivalently one output column of a run. "Slot count" is the proxy for how much the LTM overlay inflates a compiled model. +- **A2A (apply-to-all), and the A2A collapse**: Apply-to-all is the ordinary arrayed equation form, one equation covering every element diagonally. Today's exhaustive path collapses such a family into a single `Loop` carrying per-slot link lists (`slot_links`); this plan reports N separate element-level loops that share a variable-level shape instead. +- **Aggregate node (`$⁚ltm⁚agg⁚{n}`)**: An analysis-only stand-in for an inline array reducer such as the `SUM(pop[*])` inside `share[r] = pop[r] / SUM(pop[*])`. Causality is routed through it rather than scored as one lumped link. Model equations are never rewritten. +- **Reducer**: A builtin that collapses an array dimension: `SUM`, `MEAN`, `MIN`, `MAX`, `STDDEV`, `RANK`, `SIZE`. +- **Petal stitching**: The current mechanism for loops that visit one aggregate node more than once. It enumerates `agg -> ... -> agg` segments ("petals") and emits one loop per pairwise-disjoint subset, which is exponential in the petal count, needs budgets and a truncation flag, and matches neither the de-subscripted model's loop set nor the named-aggregate model's. Deleted by this plan. +- **Module instance / sub-model**: A nested model used as a variable, either a stdlib macro (`SMOOTH`, `DELAY`, `TREND`) or a user-defined sub-model. Its inputs and outputs are entry and exit ports. +- **Pathway score**: The product of link scores along one internal path through a module instance, from an entry port to an exit port. +- **Composite score**: The pathway score with the largest absolute magnitude at each timestep. It serves as the parent model's link score for the edge into the instance, hiding the instance's internals the way the papers hide a macro's. +- **Per-exit-port override**: A correction that rescores a loop's module link against the pathway ending at the port the loop actually leaves through, rather than the composite's max-magnitude pick across all ports. +- **Pin (pinned loop)**: A loop a modeler names explicitly by its variable set. The reference calls this `LOOPSCORE`: the named loop is reported and scored regardless of what discovery found. Pins keep `pin{n}` ids and are exempt from retention and from the report cap. +- **Elementary circuit**: A directed cycle that visits no node twice; what "a loop" means algorithmically. +- **Johnson's algorithm**: The standard output-sensitive enumerator for all elementary circuits of a directed graph. Used at compile time to produce the structural loop list. +- **Structural loops**: The pre-run loop list, produced by a budgeted Johnson run over the element graph. After this plan it supplies loop identity and ids only, never scores. +- **Canonical cycle key**: A rotation-invariant identity for a cycle, preserving direction. `A -> B -> C -> A` and `B -> C -> A -> B` collapse to one key, while `A -> C -> B -> A` stays a distinct loop (GH #308). +- **Union graph**: The post-simulation graph built from just those causal edges whose recorded link score was ever active. +- **Activity bitset**: One bit per saved step per edge, recording whether that edge was active there. ANDing the bitsets along a path gives exactly the steps it can score, and an empty AND proves no extension of that path can score either. +- **The universe**: The complete set of cycles that could ever have a nonzero score, which exact enumeration produces within its budgets. Relative scores are fractions, so a correct denominator requires the whole population. +- **Retention / `MIN_CONTRIBUTION`**: A candidate loop is kept only if at some single timestep its absolute score reaches 0.1% of its partition's total score mass at that step. +- **Coverage-aware cap**: The rule deciding which loops make the 200-loop report under pressure: any loop that is the top loop at some timestep within a competing partition keeps its slot unconditionally, and remaining slots are filled in ranking order. +- **`enumeration_complete` / `universe_loops`**: Result fields saying whether the exact enumerator finished, and how many circuits the universe held. +- **Shortest-path fallback**: The second candidate generator that samples cycles when exact enumeration exhausts its budget or deadline. Deleted by this plan. +- **De-subscripting oracle**: A test harness that mechanically expands an arrayed model into the equivalent scalar model, simulates both, and compares values, link scores, loops and relative series. Reference section 15.4 makes the de-subscripted scalar model the correctness standard for arrayed LTM. +- **salsa**: The Rust incremental-computation framework the compiler is built on: inputs, memoized queries, query arguments that key separate memos, and revisions. +- **Overlay (`db::LtmOverlay`)**: Whether a compile assembles the LTM instrumentation, passed as an argument to every compile query rather than stored as an input on the project, so both variants stay memoized side by side. +- **Synthetic variable**: An auxiliary the LTM pass adds to the model to compute a score, named with a `$` prefix and the U+205A separator (`$⁚ltm⁚link_score⁚x→y`). +- **`Results`**: A run's output: a flat slab of every saved variable at every saved step, plus per-variable offsets naming the columns. +- **`ProbedDb`**: A test-only wrapper that records which tracked query bodies salsa actually executed over a measured region; used to prove absences (no re-execution, no second simulation). +- **Golden**: A retained expected output, captured from today's code before that code is deleted, against which the replacement is asserted. +- **World3**: The Limits to Growth world model; large and densely connected (a 166-node SCC), the stress case for loop enumeration. +- **C-LEARN**: Climate Interactive's climate policy model and the largest model in the repository; the compile and run performance benchmark of `docs/design/engine-performance.md`. + +## Architecture + +### The finding this plan acts on + +The audit of 2026-09-07 established three facts that make the present two-mode design indefensible: + +1. Both hero models (World3, C-LEARN) exceed `MAX_LTM_SCC_NODES = 50` and auto-flip to discovery, so exhaustive mode runs only on toy fixtures, and every consumer of the structural loop surface (`model_detected_loops`, libsimlin `get_loops`, pysimlin `Run.loops`, the CLI, the layout, TS `Model.loops()`) is empty on the models the project cares about. The TS engine has no discovery binding at all. +2. The post-simulation product of a loop's recorded link-score series reproduces the in-simulation `$⁚ltm⁚loop_score⁚` column to one ulp on all seven fixtures (38 loops, arrayed slots and cross-aggregate loops included). In-simulation loop scores carry no information the post-simulation product lacks. +3. All-edge instrumentation costs no more than loop-edge instrumentation plus loop-score variables on every corpus model (C-LEARN compile 0.80 vs 0.74 s, run 1.31 vs 1.33 s; toy models emit fewer variables under all-edge instrumentation because per-loop variables dominate). + +Two further findings shape the plan: the exhaustive surface's relative normalization (per `(partition, slot)` bucket, `ltm_post::compute_rel_loop_scores_per_element`) is wrong on every coupled arrayed model while discovery's per-partition normalization is right; and static polarity is `Unknown` on 195 of World3's 200 and 150 of C-LEARN's 153 reported loops while runtime classification resolves every loop in the corpus at confidence 1.0. + +### The pipeline + +``` +compile (LtmOverlay::On) + model_causal_edges / model_element_causal_edges which edges exist (unchanged) + model_ltm_reference_sites, enumerate_agg_nodes how each edge reads (unchanged) + model_ltm_variables one link score per element edge, + sub-model pathway/composite vars for + instances in a nontrivial SCC, + NO loop scores, NO mode + model_structural_loops (Johnson, budgeted) the pre-run loop list: element cycles, + canonical keys, ids; pins validated here +simulate + the VM writes every link-score series into Results (unchanged) +post-simulation (one function, every consumer) + ltm_finding::score_loops(results, graph, structural, pins, budget) + ActivityGraph::build -> enumerate_active_circuits the universe (exact within budget) + retain_circuits, materialize, pins injected raw series per loop and slot + ltm_post::relative_scores(per partition) the one normalization + LoopPolarity::from_runtime_scores(relative series) polarity + confidence + rank, coverage-aware cap, ids from structural keys the report +``` + +`LtmOverlay` stays the compile key (`db::LtmOverlay`); there is no second key. What used to distinguish the modes (which edges get scores; whether loop scores are variables; whether Johnson's list or the post-simulation list is "the" loop list) collapses to: scores for all edges, loops from the post-simulation pipeline, Johnson for the structural preview and for ids. + +### Loop identity and ids + +A loop's identity is its element-level cycle: the canonical rotation of its element-subscripted node sequence (`ltm::canonical_cycle_key`, the one implementation that replaces `canonical_rotation`, `canonical_cycle_rotation`, `strip_element_subscript`, `dedup_trimmed_twins` and the rotation matching in `build_element_level_loops`). Two directed cycles over one node set are distinct keys (GH #308). + +Ids are assigned from the structural list when Johnson completes within `MAX_LTM_CIRCUITS`: loops sorted by canonical key get `l1, l2, ...`, and a post-simulation loop whose key appears in the structural list takes that id. A post-simulation loop absent from the structural list (Johnson did not complete, or the loop runs through a reducer node the structural list also carries, so this is rare) gets an id from its key's position in the sorted reported set, after the structural ids. Pins keep `pin{n}` ids and carry the user's name. Ids therefore carry no polarity letter; the display label a UI shows (`R1`, `B2` by importance rank, the papers' convention) is a presentation concern computed from the run, not an identity. + +This retires the exhaustive surface's A2A collapse (one `Loop` with `slot_links` and N slots): a loop is one element cycle, and an arrayed family is N loops sharing a variable-level shape. The FFI's subscripted access (`simlin_analyze_get_relative_loop_score("r1[nyc]")`, `simlin_analyze_get_loop_element_count`, pysimlin's `element=` accessors) goes with it; consumers that want the family group by the variable-level shape, which every reported loop exposes. + +### Pins + +`db/ltm/pinned.rs::model_pinned_loops` keeps its validation (it reads the causal graph, not a run) and its `pin{n}` ids, but emits no `loop_score` variable. The post-simulation pipeline injects each valid pin's element cycle(s) into the candidate set before retention, marks them exempt from retention and from the coverage-aware cap, and scores them with everyone else. This is the reference's `LOOPSCORE` semantics: the loop is reported regardless of what discovery found (reference section 10.2). + +### Reducers + +An inline reducer subexpression is already a node in the element graph (`$⁚ltm⁚agg⁚{n}`, `ltm_agg.rs`, `db/analysis.rs::emit_agg_routed_edges`). Today it is trimmed from reported loops and the loops that visit it more than once are reconstructed by petal stitching (`db/ltm/loops.rs::stitch_cross_agg_petals`), which emits one loop per petal subset. The audit showed this is neither the de-subscripted model's loop universe (which has `sum C(N,k) (k-1)!` circuits) nor the named-aggregate model's (N loops), and that its relative scores are off by construction. This plan adopts the named-aggregate semantics, which is also the papers' treatment of hidden structure (a macro's internal pathways collapse to the macro; reference sections 6.3 and 6.4): the reducer node is a real node, loops are the elementary circuits of the element graph, and the node is reported in the loop's sequence with the reducer's spelling (`SUM(pop)`) so a reader sees the loop close through the aggregate rather than reading `growth[a] -> pop[a]` as a self-loop. The stitcher, its budgets and its truncation flag are deleted. + +### Sub-model instances + +Pathway and composite variables exist so a loop through a module instance can be scored through the instance's strongest internal pathway (reference section 6.3) and, in the post-simulation pipeline, re-selected per exit port (`ltm_finding::recompute_module_input_edge_series`). An instance whose node is not in a nontrivial SCC of its parent's causal graph cannot lie on a loop, so no consumer ever reads its pathway scores; `model_ltm_variables` emits them only for instances in a nontrivial SCC (`causal_graph_from_edges(..).scc_of(node)`). The exhaustive-side override machinery (`db/ltm/mod.rs::compute_module_link_overrides`, `max_abs_alias_selection`, the `⁚via⁚` / `⁚viaacc⁚` alias variables) goes with the loop-score variables; the post-simulation recompute is the one owner. + +### Polarity + +`LoopPolarity::from_runtime_scores` classifies from a loop's partition-relative series (bounded, dominance-weighted; the audit's synthetic case shows raw sums let three inflection steps outweigh two hundred balancing steps). Before a run a loop has `polarity: None`. `ltm/polarity.rs` (static AST polarity), `CausalGraph::get_link_polarity` / `all_link_polarities` / `calculate_polarity`, `db::analysis::compute_link_polarities` and the agg-hop recovery in `db/ltm/loops.rs` are deleted. libsimlin `get_links` and the MCP `read_model` relationships report the runtime sign of each link's recorded series (same classifier, link series) when a run exists and `None` otherwise. + +### Enumeration budgets + +`enumerate_active_circuits` keeps its three budgets and the caller's deadline. When any trips, the partial circuit list is retained, ranked and reported with `enumeration_complete == false` and `universe_loops == None`. The audit measured the shortest-path fallback at recall 0.20 at the cap on World3 (the only corpus model dense enough to need it, where the exact enumerator finishes in 0.56 s) and 0.98 on C-LEARN (where the exact enumerator finishes in 47 ms); it never fires within default budgets on any model in the repository, so it is deleted rather than kept as a second generator. + +### Surfaces + +- libsimlin: `simlin_analyze_get_loops` (structural, pre-run), `simlin_analyze_get_loops_runtime` (post-run, the pipeline), `simlin_analyze_discover_loops` (renamed `simlin_analyze_loops`, the same pipeline over a sim's results), `simlin_analyze_loops_from_wasm_results` (the pipeline over a wasm slab, twin of `simlin_analyze_rel_loop_score_from_wasm_results`, which it replaces). `simlin_sim_get_ltm_mode`, `simlin_analyze_get_relative_loop_score`, `simlin_analyze_get_rel_loop_score`, `simlin_analyze_get_loop_score`, `simlin_analyze_get_loop_element_count` and the `SimState` partition/denominator caches are deleted; per-loop series come from the loops result. +- `@simlin/engine`: `Run.loops` populated on both engines through the new binding; `includeInternal` defaults to false; `Model.loops()` returns the structural list with `polarity: null`. +- pysimlin: `Run.loops` from the pipeline on every model; `LtmMode`, `Sim.get_loops_runtime`, the `element=` accessors and `Run._populate_loop_behavior`'s column reader are deleted; `Model.get_loops()` structural with `polarity None`. +- MCP: `analyze_model` calls the pipeline directly; `enumerationComplete` stays on the wire. +- CLI `--ltm`: prints the post-simulation loop report; `$⁚` columns are excluded from the TSV unless `--ltm-columns` is given. +- Layout: `layout::detect_ltm_loops` takes a `Results` (the display run's, when the caller has one) and calls the pipeline; it no longer runs its own simulation when results are supplied. + +## Existing Patterns + +- The post-simulation pipeline is discovery's existing `ltm_finding` module (`docs/design-plans/2026-08-17-ltm-discovery-exact.md`): union graph, activity bitsets, exact enumeration, retention against the universe, coverage-aware cap. This plan makes it the only pipeline; it changes nothing in its numerics. +- `db::LtmOverlay` as a query argument (`docs/design/engine-performance.md`, "The LTM overlay is an argument, not an input") is the pattern this plan extends by deleting the remaining mode input rather than converting it (GH #1056). +- One owner per compiler decision (`src/simlin-engine/CLAUDE.md`, "One owner") is the rule the consolidation phases enforce; the audit listed six LTM violations (partitions, dedup, normalization, runtime polarity readers, per-exit-port override, cross-agg stitching), each traceable to the mode split. +- The parity-against-a-retained-oracle pattern (`tests/integration/simulate_ltm_wasm.rs`) is reused: the seven fixtures' exhaustive `loop_score` series are captured as goldens before the exhaustive path is deleted, and the pipeline is asserted against them. +- The de-subscripting oracle built in the audit (`scratchpad/B/desub.py`, `oracle.py`) becomes an integration harness over `datamodel::Project`, following the CLAUDE.md rule that the only correctness standard for arrayed LTM is the de-subscripted scalar model (reference section 15.4). + +Divergence from the design doc `docs/design/ltm--loops-that-matter.md`: its "Two Modes of Operation", "Cross-agg loop recovery", "Per-Slot Loop Score Equations", "Static Polarity" and "Passthrough composites and per-exit-port loop scoring" sections describe machinery this plan deletes; Phase 8 rewrites the document as if the code had always been single-mode. + +## Implementation Phases + + +### Phase 1: Goldens and the one normalization +**Goal:** Capture the exhaustive path's loop-score series as goldens, and make per-partition normalization the single owner every reader uses. + +**Components:** +- `tests/integration/ltm_single_mode_parity.rs` -- for each of the seven fixtures, the exhaustive `loop_score` series per loop and slot, captured once into `test//loop_scores.golden.tsv`, and the assertion that the post-simulation pipeline reproduces them to one ulp. +- `src/simlin-engine/src/ltm_post.rs` -- `relative_scores(per-partition)` as the one normalization; `compute_rel_loop_scores`, the per-element bucket grid, `LoopElementIndex` and the streaming readers deleted; `ltm_finding::signed_relative_scores` calls it. +- `src/libsimlin/src/analysis.rs`, `src/simlin-engine/src/layout/detect_ltm_loops.rs` -- read the owner. + +**Dependencies:** Work package 1 item B1 (the per-partition rule) merged. + +**Done when:** `ltm_single_mode_parity.rs` passes against the goldens through the discovery pipeline; `AC2.1`, `AC6.3`. + + + +### Phase 2: One instrumentation, loops from the pipeline +**Goal:** `model_ltm_variables` emits link scores for every edge and no loop scores; every consumer reads loops from the post-simulation pipeline; the mode is gone. + +**Components:** +- `src/simlin-engine/src/db/ltm/mod.rs` -- the exhaustive branch, loop-score emission, `compute_module_link_overrides`, `max_abs_alias_selection`, `model_ltm_mode`, the auto-flip warnings and `LtmMode` deleted; pathway/composite emission kept (gated in Phase 6). +- `src/simlin-engine/src/ltm_augment.rs` -- `generate_loop_score_variables`, `generate_dimensioned_loop_score_equation`, `generate_link_product` and `resolve_link_score_name_for_loop` deleted. +- `src/simlin-engine/src/db/input.rs`, `db.rs`, `analysis.rs`, `wasmgen/module.rs` -- `ltm_discovery_mode` and its setters/threading deleted; `analyze_model` calls the pipeline. +- `src/simlin-engine/src/ltm_finding.rs` -- `score_loops(results, graph, structural, pins, budget)` as the one entry point; pin injection before retention (exempt from retention and cap). +- `src/simlin-engine/src/db/ltm/pinned.rs` -- validation kept; `expand_pin_on_element_graph` produces element cycles for injection; no `Loop` emission. +- `src/libsimlin/src/analysis.rs`, `simulation.rs`, `model.rs` -- `get_loops_runtime` and `discover_loops` on the pipeline; `simlin_sim_get_ltm_mode` and the per-loop score accessors deleted; `compile_to_wasm` loses its discovery flag. +- `src/pysimlin/simlin/{model,run,sim,analysis}.py` -- `Run.loops` from the pipeline; `LtmMode` and the column reader deleted. +- `src/simlin-mcp-core/src/tools/*` -- unchanged call shape, no flag. +- `src/simlin-cli/src/main.rs` -- the `--ltm` report from the pipeline. +- Tests: `simulate_ltm.rs`, `simulate_ltm_pinned.rs`, `db/ltm_unified_tests.rs`, `db/ltm_module_tests.rs`, `db/ltm_tests.rs` rewritten against `score_loops`' output; mode tests deleted. + +**Dependencies:** Phase 1. + +**Done when:** every LTM test is green against the pipeline; `AC1.1`-`AC1.4`, `AC2.2`, `AC2.4`, `AC4.1`-`AC4.4`. + + + +### Phase 3: Loop identity, ids and the structural list +**Goal:** One canonical cycle key; ids from the structural list; the A2A collapse retired. + +**Components:** +- `src/simlin-engine/src/ltm/mod.rs` -- `canonical_cycle_key` (element-level, direction-preserving) replacing `canonical_rotation`, `db::analysis::canonical_cycle_rotation` / `strip_element_subscript` / `strip_subscript`, `dedup_trimmed_twins`, `by_reported_cycle`, and `assign_pin_ids`' rotation match. +- `src/simlin-engine/src/db/analysis.rs` -- `model_structural_loops` (budgeted Johnson over the element graph, ids `l{n}` by key order) replacing `model_loop_circuits_tiered`, `classify_cycle`, `build_loops_from_tiered` and `model_detected_loops`' mode branch; `model_element_cycle_partitions` the one partition owner (`model_cycle_partitions` and `CausalGraph::compute_cycle_partitions` deleted). +- `src/simlin-engine/src/db/ltm/loops.rs` -- the A2A half of `build_element_level_loops` and `Loop::slot_links` deleted. +- `src/simlin-engine/src/ltm/types.rs` -- `Loop { key, id, nodes, links, stocks, partition, dimensions_shape }` without `slot_links`. +- Consumers of subscripted loop access (`src/libsimlin/src/analysis.rs` `parse_subscripted_loop_id`, `get_loop_element_count`; pysimlin `element=`) deleted; `DiscoveredLoop`/`LoopSummary` expose the variable-level shape for grouping. + +**Dependencies:** Phase 2. + +**Done when:** `AC6.1`, `AC6.2`, `AC8.2`; the seven fixtures report the element-level loop set the parity goldens name. + + + +### Phase 4: Delete the fallback +**Goal:** The exact enumerator is the only candidate generator; budget exhaustion reports a partial universe. + +**Components:** +- `src/simlin-engine/src/ltm_finding_fallback.rs`, `ltm_finding_fallback_tests.rs`, `examples/ltm_fallback_eval.rs`, the fallback rows of `examples/ltm_discovery_bench.rs` -- deleted. +- `src/simlin-engine/src/ltm_finding.rs`, `ltm_finding_enum.rs` -- `enumerate_active_circuits` returns the partial list on a budget trip; `retain_circuits` and ranking run over it; `ENUM_BUDGET_FRACTION`, `FallbackConfig` and `fallback_candidates` deleted; `DiscoveryResult::truncated` folded into `enumeration_complete`. +- `docs/design/ltm--loops-that-matter.md` -- the fallback section replaced by the partial-universe rule. + +**Dependencies:** Phase 2. + +**Done when:** `AC5.1`, `AC5.2`; `rg -n "fallback" src/simlin-engine/src/ltm_finding*.rs` is empty. + + + +### Phase 5: A reducer is a node +**Goal:** Loops through an inline reducer are the element graph's elementary circuits through the reducer node; the stitcher is gone. + +**Components:** +- `src/simlin-engine/src/db/ltm/loops.rs` -- `stitch_cross_agg_petals`, `recover_cross_agg_loops`, `collect_agg_petals`, `MAX_AGG_PETALS`, `MAX_CROSS_AGG_LOOPS`, `AggLoopBudgetGuard`, `recover_agg_hop_polarities` deleted. +- `src/simlin-engine/src/ltm_finding.rs` -- `stitch_cross_agg_node_paths`, `trim_synthetic_aggs_from_loop_links` and `agg_recovery_truncated` deleted; the reported node sequence keeps the agg node, displayed by `AggNode::display_spelling` (the reducer's spelled text). +- `src/simlin-engine/src/ltm_agg.rs` -- `display_spelling` on `AggNode`. +- Surfaces: `agg_recovery_truncated` removed from `ModelAnalysis`, the FFI, pysimlin and MCP; loop node lists may contain an agg display node, flagged `synthetic: true` so a UI can style it. +- `tests/integration/ltm_array_agg.rs` -- cross-agg expectations rewritten to the named-aggregate semantics; `tests/integration/ltm_desubscript_oracle.rs` -- the audit's oracle over `datamodel::Project` (de-subscript, simulate both, compare values bit-exact, link scores, loops and relative series) on the arrayed fixtures and a generated corpus. + +**Dependencies:** Phase 3. + +**Done when:** `AC7.1`-`AC7.3`. + + + +### Phase 6: Static polarity deleted; module emission gated +**Goal:** Polarity is a runtime field; sub-model pathway/composite variables exist only for instances that can be on a loop. + +**Components:** +- `src/simlin-engine/src/ltm/polarity.rs`, `polarity_tests.rs`, `with_lookup_tests.rs` -- deleted; `CausalGraph::get_link_polarity`, `all_link_polarities`, `calculate_polarity`, `db::analysis::compute_link_polarities` deleted; `Link` loses its static `polarity`. +- `src/simlin-engine/src/ltm/types.rs` -- `Loop::polarity: Option`; `from_runtime_scores` over the relative series (the one classifier), applied to link series for `get_links` and the MCP relationships. +- `src/simlin-engine/src/db/ltm/mod.rs` -- pathway/composite emission gated on `causal_graph_from_edges(edges).scc_of(instance).len() > 1`. +- `src/simlin-engine/src/ltm/graph.rs` -- `all_links` and `find_loops_with_limit` deleted (no callers). +- pysimlin `test_ltm_polarity.py`, `engine/ltm/tests.rs` id assertions rewritten. + +**Dependencies:** Phase 3. + +**Done when:** `AC8.1`, `AC9.1`, `AC9.2`. + + + +### Phase 7: Every surface, and the layout on the display run +**Goal:** The TS engine, both wasm paths, pysimlin, the CLI and the layout all read the pipeline; no surface is empty on any model. + +**Components:** +- `src/libsimlin/src/analysis.rs` -- `simlin_analyze_loops_from_wasm_results` replacing `simlin_analyze_rel_loop_score_from_wasm_results`; `simlin_analyze_links_from_wasm_results` keeps its role. +- `src/engine/src/{model,sim,run}.ts`, `src/engine/src/internal/{analysis,wasmgen}.ts`, `direct-backend.ts`, `worker-server.ts`, `worker-protocol.ts` -- `Run.loops` on both engines; `includeInternal` default false; `Model.loops()` structural. +- `src/simlin-engine/src/layout/detect_ltm_loops.rs`, `layout/mod.rs` -- accept a `Results`; `src/libsimlin/src/layout.rs`, `src/simlin-mcp-core/src/tools/edit_model.rs` -- pass the display run's results where one exists. +- `src/simlin-cli/src/main.rs` -- `$⁚` columns excluded unless `--ltm-columns`. +- `src/engine/tests/wasm-ltm.test.ts`, `src/pysimlin/tests/test_ltm.py`, `tests/integration/ltm_discovery_large_models.rs` -- the cross-surface assertions of `AC3`. + +**Dependencies:** Phase 2 (pipeline), Phase 6 (polarity field shape). + +**Done when:** `AC2.3`, `AC3.1`-`AC3.4`. + + + +### Phase 8: Docs, ledger, epic +**Goal:** The documentation describes single-mode LTM as the only design; the cost ledger is recorded; the epic reflects what closed. + +**Components:** +- `docs/design/ltm--loops-that-matter.md` -- rewritten: pipeline, identity and ids, pins, reducers, module gating, polarity, budgets, surfaces; mode, stitching, per-slot equations, static polarity and fallback sections removed. +- `docs/reference/ltm--loops-that-matter.md` -- sections 15.2, 16.10 and the Simlin implementation notes updated to the named-aggregate semantics. +- `src/simlin-engine/CLAUDE.md`, `src/libsimlin/CLAUDE.md`, `src/pysimlin/CLAUDE.md`, `src/engine/CLAUDE.md` -- surface changes. +- `docs/design-plans/2026-09-04-link-scores-from-fragments.md` -- its "What this plan does not change" list and its Phase 4 (no loop/pathway/composite text generators remain to convert) updated. +- This document's ledger section: pre- and post-plan compile, run, slots, discovery time for the corpus (`examples/ltm_full_bench.rs`). +- GH epic #488: clusters B, E, G updated; #1056, #760, #309 (resolution unchanged, noted), #677, #674 (pins carry direction now), #658, #755, #672, #665, #685, #701 closed or rescoped in the PR body. + +**Dependencies:** Phases 1-7. + +**Done when:** `AC10.1`; the docs contain no changelog sentences about the modes. + + +## Additional Considerations + +**Owner decisions embedded in this plan** (each defaulted to the audit's recommendation; a different answer changes a phase, not the architecture): + +- D1. Ids without a polarity letter (`l{n}`), polarity as a runtime field; display labels `R{n}`/`B{n}` are a UI concern. Alternative: keep `r/b/u` prefixes computed after the run, which makes ids unstable across parameter changes. +- D2. Element-level loops as the reported unit (an arrayed family is N loops sharing a shape). Alternative: keep the A2A collapse on the reported surface, which keeps `slot_links`, the subscripted FFI access and ~430 lines of the tiered enumerator. +- D3. Named-aggregate semantics for inline reducers (a reducer is a node, shown in the loop). Alternative: enumerate the de-subscripted model's `(k-1)!` orderings, which is what reference 15.4 literally promises and is super-exponential. +- D4. Static polarity deleted (pre-run loops carry no polarity). Alternative: keep a ten-line flow-to-stock sign rule for an unrun diagram; the audit found nothing else `polarity.rs` computes survives contact with a real model. +- D5. The CLI hides `$⁚` columns by default. + +**Discovery cost on dense models.** World3's exact enumeration is 0.56 s per run; under always-on that is paid on every simulation. The wall-clock budget parameter of the pipeline exists; a product default for it, and whether a run should reuse the previous run's loop set when the model is unchanged, are decisions for the always-on product plan, not this one. + +**Sampling resolution.** The pipeline scores loops at saved-step resolution (GH #309). The fragments plan's Phase 6 retains every-dt states natively, which is the input a per-dt pipeline needs; this plan does not change the resolution. + +**Interaction with the fragments plan.** Independent and preferably first: the fragments plan's Phase 4 converts the loop-score, pathway and composite text generators to typed builders, and after this plan the loop-score generators and the exhaustive override aliases do not exist. Its "What this plan does not change" list (discovery mode, pinned loops, `ltm_discovery_mode`) needs rewording either way. + +**What is not in scope.** The loop UI in `src/diagram`, the RK and conveyor policies for always-on, the equilibrium diagnostic (#504), and the per-dt sampling of #309 belong to the always-on product plan that follows this one. From b27550ae7eec8452dcdc5f3fd7b481465c9988b2 Mon Sep 17 00:00:00 2001 From: Bobby Powers Date: Wed, 9 Sep 2026 20:55:31 -0700 Subject: [PATCH 10/10] engine: name a pin by its own uids when it cannot be scored Every one of the hero_culture fixture's fourteen loopMetadata entries was rejected at sync ("a pinned loop must name at least two variables"), so the fixture never exercised a pin and the rejection named only the loop's name, five of which read "Technical Debt Spiral". The fixture was the wrong side. LoopMetadata.uids are VARIABLE uids everywhere else: the TypeScript types say so, SetLoopName mints a uid on each named variable, and the MCP open path assigns every variable a uid so that SetLoopName can resolve it. The fixture's entries were written with the uids of the variables' diagram elements while no variable carried a uid at all, so the projection resolved nothing. Variable and view-element uids share one number space (patch::next_available_uid takes the maximum over both), so the regenerated fixture gives each of the 27 variables the uid of its own element and leaves the entries untouched; all fourteen now resolve, validate against the causal graph, and score. A pin that cannot be scored is now reported with the entry's own identity: PinnedLoopSpec carries the entry's uids as written, the subset that matched no variable, and whether any variable of the model carried a uid at sync (uids live only on the datamodel, so the fact travels on the spec); every rejection arm goes through one reject site that appends the unresolved-uid clause, so the two-variable gate reads "'Ghost' (uids [7, 4]) names 0; uids [4, 7] match no variable's uid; no variable in the model carries a uid (a pin names variables by their uid, which SetLoopName assigns)", the not-a-cycle arm names the entry's uids the same way, and the no-stock and failed-expansion arms name a dropped uid too. A pin that drops a stale uid but whose survivors still form a cycle scores that cycle and says so, naming the dropped uids and the cycle scored instead ("... dropped uids [99], which match no variable's uid, and scored the loop its remaining variables form instead: births -> population"): a renamed or re-created variable must never silently turn a pin into a different loop under the same name. The name alone was never an identity. Pinned: the fixture's fourteen pins produce no warning and, in discovery mode, fourteen non-trivial pin{n} loop scores whose relative scores are finite, bounded and non-zero; a pin on a model whose variables carry no uid is reported with its name, its uids, the unresolved uids and the no-uid cause; a pin with one stale uid beside a resolving one names the stale uid and not the no-uid cause; a pin whose survivors are no cycle names its uids and the dropped one; a pin with a stale uid beside two resolving ones scores their cycle under pin1 and warns naming it, and that warning has its row in the once-per-project warning-family test. --- .../src/db/diagnostic_payload_tests.rs | 16 ++ src/simlin-engine/src/db/diagnostic_tests.rs | 9 + src/simlin-engine/src/db/input.rs | 13 + src/simlin-engine/src/db/ltm/mod.rs | 5 + src/simlin-engine/src/db/ltm/pinned.rs | 90 +++++-- src/simlin-engine/src/db/sync.rs | 23 +- .../tests/integration/simulate_ltm_pinned.rs | 237 ++++++++++++++++++ test/hero_culture_ltm/hero_culture.sd.json | 27 ++ 8 files changed, 397 insertions(+), 23 deletions(-) diff --git a/src/simlin-engine/src/db/diagnostic_payload_tests.rs b/src/simlin-engine/src/db/diagnostic_payload_tests.rs index 38e53d8df..38935d6bb 100644 --- a/src/simlin-engine/src/db/diagnostic_payload_tests.rs +++ b/src/simlin-engine/src/db/diagnostic_payload_tests.rs @@ -775,6 +775,22 @@ fn every_warning_family_is_emitted_once_across_revisions() { guards: no_guards, matches: |d| assembly_reason_contains(d, "pinned loop 'bogus'"), }, + Family { + name: "ltm: a pin that scored the loop its surviving variables form", + child: || { + let mut project = loop_child(); + pin_loop(&mut project.models[0], "growth", &["s", "in_f"]); + // A uid no variable carries: dropped, the survivors still + // form the loop, and the pin scores it with a warning. + project.models[0].loop_metadata[0].uids.push(99); + project + }, + probe: "s", + wiring: &[], + discovery: false, + guards: no_guards, + matches: |d| assembly_reason_contains(d, "dropped uids [99]"), + }, Family { name: "ltm: an arrayed edge whose dimensions do not correspond", child: || { diff --git a/src/simlin-engine/src/db/diagnostic_tests.rs b/src/simlin-engine/src/db/diagnostic_tests.rs index b4f19520d..24ecb50ac 100644 --- a/src/simlin-engine/src/db/diagnostic_tests.rs +++ b/src/simlin-engine/src/db/diagnostic_tests.rs @@ -1077,6 +1077,9 @@ fn test_diagnostics_stable_across_unrelated_input_change() { .to(vec![PinnedLoopSpec { name: "dummy_loop".to_string(), variables: vec![], + uids: vec![], + unresolved_uids: vec![], + model_variables_carry_uids: true, description: String::new(), }]); @@ -2820,6 +2823,9 @@ fn macro_registry_build_error_survives_an_unrelated_input_change() { .to(vec![PinnedLoopSpec { name: "dummy_loop".to_string(), variables: vec![], + uids: vec![], + unresolved_uids: vec![], + model_variables_carry_uids: true, description: String::new(), }]); @@ -2949,6 +2955,9 @@ fn unit_definition_errors_survive_an_unrelated_input_change() { .to(vec![PinnedLoopSpec { name: "dummy_loop".to_string(), variables: vec![], + uids: vec![], + unresolved_uids: vec![], + model_variables_carry_uids: true, description: String::new(), }]); diff --git a/src/simlin-engine/src/db/input.rs b/src/simlin-engine/src/db/input.rs index dda293ef5..aa6063ccb 100644 --- a/src/simlin-engine/src/db/input.rs +++ b/src/simlin-engine/src/db/input.rs @@ -126,6 +126,19 @@ pub struct PinnedLoopSpec { /// deduplicated so the spec is order-independent (a loop's identity is /// its node set; the cycle order is recovered from the causal graph). pub variables: Vec, + /// The entry's `uids` exactly as written in the `LoopMetadata`: the + /// identity a user sees in the file, reported when the pin cannot be + /// scored so the entry can be found (the name alone is not unique). + pub uids: Vec, + /// The subset of `uids` that matched no variable's uid, ascending: a + /// stale reference after a delete, or a file whose pins were written + /// against something other than variable uids (view element uids, say). + pub unresolved_uids: Vec, + /// Whether any variable of the model carried a uid at sync. UIDs live + /// only on the datamodel, so this is recorded here for the failure + /// message: when no variable has one, the pin's uids cannot name + /// anything whatever they are, and saying so points at the cause. + pub model_variables_carry_uids: bool, /// The user-supplied description (empty when none was given). pub description: String, } diff --git a/src/simlin-engine/src/db/ltm/mod.rs b/src/simlin-engine/src/db/ltm/mod.rs index b3cff01dd..9d0d81bcd 100644 --- a/src/simlin-engine/src/db/ltm/mod.rs +++ b/src/simlin-engine/src/db/ltm/mod.rs @@ -1723,6 +1723,11 @@ pub fn model_ltm_variables( let _ = name; warnings.warn(None, reason.clone()); } + for (_name, message) in &pinned.warnings { + // A pin that scored a different loop than written (a stale uid was + // dropped) is the other silent wrong number: say which loop it scored. + warnings.warn(None, message.clone()); + } if !pinned.loops.is_empty() { // The variable-level node set of each already-emitted enumerated loop, // keyed by canonical rotation, so a pin that duplicates one is skipped. diff --git a/src/simlin-engine/src/db/ltm/pinned.rs b/src/simlin-engine/src/db/ltm/pinned.rs index ea8ad56b4..fd67c8792 100644 --- a/src/simlin-engine/src/db/ltm/pinned.rs +++ b/src/simlin-engine/src/db/ltm/pinned.rs @@ -69,6 +69,26 @@ pub struct PinnedLoopsResult { /// scorable feedback loop. The reason is a human-readable explanation for /// the surfaced diagnostic. pub invalid: Vec<(String, String)>, + /// `(name, message)` for each pinned loop that DID score, but not the + /// loop as written: a uid that matched no variable was dropped and the + /// survivors still formed a cycle, so the score is that cycle's. Surfaced + /// as a Warning so a renamed or re-created variable never silently turns + /// a pin into a different loop under the same name. + pub warnings: Vec<(String, String)>, +} + +impl PinnedLoopsResult { + /// Record a pin that scores nothing. The one site that appends the + /// unresolved-uid clause, so every rejection arm -- too few variables, + /// no cycle, no stock, a failed element expansion -- names a dropped + /// uid: a pin whose dropped uid was its only stock, say, must still say + /// which uid went missing. + fn reject(&mut self, spec: &crate::db::input::PinnedLoopSpec, reason: String) { + self.invalid.push(( + spec.name.clone(), + format!("{reason}{}", unresolved_uids_clause(spec)), + )); + } } /// Resolve and validate a model's pinned loops against its causal graph. @@ -134,15 +154,16 @@ pub(crate) fn model_pinned_loops( let id = format!("pin{}", idx + 1); if spec.variables.len() < 2 { - result.invalid.push(( - spec.name.clone(), + result.reject( + spec, format!( "a pinned loop must name at least two variables that form a feedback loop; \ - '{}' names {}", + '{}' (uids {:?}) names {}", spec.name, - spec.variables.len() + spec.uids, + spec.variables.len(), ), - )); + ); continue; } @@ -150,15 +171,16 @@ pub(crate) fn model_pinned_loops( spec.variables.iter().map(|v| Ident::new(v)).collect(); let Some(cycle) = graph.order_variable_cycle(&vars) else { - result.invalid.push(( - spec.name.clone(), + result.reject( + spec, format!( - "the variables named by pinned loop '{}' do not form a closed feedback loop \ - in the model's causal graph: [{}]", + "the variables named by pinned loop '{}' (uids {:?}) do not form a closed \ + feedback loop in the model's causal graph: [{}]", spec.name, - spec.variables.join(", ") + spec.uids, + spec.variables.join(", "), ), - )); + ); continue; }; @@ -195,14 +217,14 @@ pub(crate) fn model_pinned_loops( .enrich_with_module_stocks(&cycle, parent_stocks) .is_empty(); if !has_stock && !cycle_has_lagged_edge(db, project, source_vars, &cycle) { - result.invalid.push(( - spec.name.clone(), + result.reject( + spec, format!( "pinned loop '{}' contains no stock; a feedback loop must pass through at \ least one stock (or other state, such as a PREVIOUS-lagged reference)", spec.name ), - )); + ); continue; } @@ -236,15 +258,25 @@ pub(crate) fn model_pinned_loops( loops } Err(reason) => { - result.invalid.push(( - spec.name.clone(), - format!("pinned loop '{}' {reason}", spec.name), - )); + result.reject(spec, format!("pinned loop '{}' {reason}", spec.name)); continue; } } } }; + if !spec.unresolved_uids.is_empty() { + result.warnings.push(( + spec.name.clone(), + format!( + "pinned loop '{}' (uids {:?}) dropped uids {:?}, which match no variable's \ + uid, and scored the loop its remaining variables form instead: {}", + spec.name, + spec.uids, + spec.unresolved_uids, + cycle_strs.join(" -> ") + ), + )); + } result.loops.push(PinnedLoop { loops, name: spec.name.clone(), @@ -254,6 +286,28 @@ pub(crate) fn model_pinned_loops( result } +/// The clause a pin's failure message carries when some of its uids matched +/// no variable, so the entry can be found in the file and the cause seen: a +/// pin names variables by their uid, and a file whose variables carry none +/// (`PinnedLoopSpec::model_variables_carry_uids`) says so outright, because that is the shape a +/// pin written against view-element uids takes. Empty when every uid +/// resolved. +fn unresolved_uids_clause(spec: &crate::db::input::PinnedLoopSpec) -> String { + if spec.unresolved_uids.is_empty() { + return String::new(); + } + let cause = if !spec.model_variables_carry_uids { + "; no variable in the model carries a uid (a pin names variables by their uid, \ + which SetLoopName assigns)" + } else { + "" + }; + format!( + "; uids {:?} match no variable's uid{cause}", + spec.unresolved_uids + ) +} + /// Whether any edge `from -> to` of the ordered cycle is a PREVIOUS-lagged /// reference: `to` reads `from` ONLY through `PREVIOUS(...)` in its dt /// equation (`DepRefs::dt_previous_only`). Such an edge is the one-DT memory diff --git a/src/simlin-engine/src/db/sync.rs b/src/simlin-engine/src/db/sync.rs index cadc6e5e2..f537fdac2 100644 --- a/src/simlin-engine/src/db/sync.rs +++ b/src/simlin-engine/src/db/sync.rs @@ -168,11 +168,13 @@ fn macro_declarations_from_datamodel( /// /// UIDs live only on the datamodel `Variable`s and are never synced into the /// db, so we must resolve them here at sync time. A UID with no matching -/// variable (a stale reference after a delete) is dropped from that loop's -/// set; the LTM pin-resolution query then validates whatever survives against -/// the causal graph (an incomplete set fails the cycle check and surfaces a -/// diagnostic rather than scoring a partial loop). Deleted entries are -/// excluded entirely -- a deleted pin contributes no `loop_score`. +/// variable (a stale reference after a delete, or a pin written against +/// view-element uids) is dropped from that loop's set and recorded on the +/// spec as unresolved; the LTM pin-resolution query then validates whatever +/// survives against the causal graph (an incomplete set fails the cycle +/// check and surfaces a diagnostic that names the entry's uids rather than +/// scoring a partial loop). Deleted entries are excluded entirely -- a +/// deleted pin contributes no `loop_score`. fn pinned_loops_from_datamodel(model: &datamodel::Model) -> Vec { use std::collections::BTreeSet; @@ -203,9 +205,20 @@ fn pinned_loops_from_datamodel(model: &datamodel::Model) -> Vec .collect::>() .into_iter() .collect(); + let unresolved_uids: Vec = lm + .uids + .iter() + .filter(|uid| !uid_to_name.contains_key(uid)) + .copied() + .collect::>() + .into_iter() + .collect(); PinnedLoopSpec { name: lm.name.clone(), variables, + uids: lm.uids.clone(), + unresolved_uids, + model_variables_carry_uids: !uid_to_name.is_empty(), description: lm.description.clone(), } }) diff --git a/src/simlin-engine/tests/integration/simulate_ltm_pinned.rs b/src/simlin-engine/tests/integration/simulate_ltm_pinned.rs index d04fd73ac..2d88f5068 100644 --- a/src/simlin-engine/tests/integration/simulate_ltm_pinned.rs +++ b/src/simlin-engine/tests/integration/simulate_ltm_pinned.rs @@ -2095,3 +2095,240 @@ fn pinned_loop_through_unscoreable_edge_skipped_with_single_warnings() { vm.run_to_end() .expect("simulation should run to completion"); } + +/// The `hero_culture_ltm` fixture's fourteen `loopMetadata` entries name +/// their variables by uid (each variable carries the uid of its diagram +/// element, the one number space `patch::next_available_uid` spans), so +/// every pin resolves, validates against the causal graph, and -- in +/// discovery mode, where a pin is the only way a specific loop is scored -- +/// emits a non-trivial `pin{n}` loop score. +#[test] +fn hero_culture_fixture_pins_all_resolve_and_score() { + let f = std::fs::File::open("../../test/hero_culture_ltm/hero_culture.sd.json").unwrap(); + let json_project = + simlin_engine::json::Project::from_reader(std::io::BufReader::new(f)).unwrap(); + let project: datamodel::Project = json_project.into(); + let pins = project.models[0].loop_metadata.len(); + assert_eq!(pins, 14, "the fixture declares fourteen pins"); + + let mut db = SimlinDb::default(); + let sync = sync_from_datamodel_incremental(&mut db, &project, None); + set_project_ltm_discovery_mode(&mut db, sync.project, true); + let diagnostics = collect_all_diagnostics(&db, sync.project, simlin_engine::db::LtmOverlay::On); + let pin_warnings: Vec = diagnostics + .iter() + .map(|d| format!("{:?}", d.error)) + .filter(|m| m.contains("pinned")) + .collect(); + assert!( + pin_warnings.is_empty(), + "every fixture pin resolves and validates; got {pin_warnings:#?}" + ); + + let source_model = sync.models["main"].source_model; + let ltm = model_ltm_variables(&db, source_model, sync.project); + let pin_ids: Vec = (1..=pins).map(|n| format!("pin{n}")).collect(); + for id in &pin_ids { + assert!( + ltm.vars + .iter() + .any(|v| v.name == format!("$\u{205A}ltm\u{205A}loop_score\u{205A}{id}")), + "{id} is scored; have {:?}", + ltm.vars + .iter() + .filter(|v| v.name.contains("loop_score")) + .map(|v| v.name.as_str()) + .collect::>() + ); + } + + let compiled = + compile_project_incremental(&db, sync.project, "main", simlin_engine::db::LtmOverlay::On) + .expect("the fixture compiles with its pins"); + let mut vm = Vm::new(compiled).unwrap(); + vm.run_to_end().unwrap(); + let results = vm.into_results(); + for id in &pin_ids { + assert_loop_score_is_link_product(&results, id); + } + // Relative-score sanity through the one normalization owner: every pin's + // share of its partition is finite and bounded, and is non-zero at some + // step (a pin that never carries any share would be scoring nothing). + let relative = simlin_engine::ltm_post::compute_rel_loop_scores(&results, <m.loop_partitions); + for id in &pin_ids { + let series = relative + .get(id) + .unwrap_or_else(|| panic!("{id} has a relative series; have {:?}", relative.keys())); + assert!( + series + .iter() + .all(|v| v.is_finite() && v.abs() <= 1.0 + 1e-12), + "{id}: relative scores are finite and bounded; got {series:?}" + ); + assert!( + series.iter().any(|v| *v != 0.0), + "{id}: the pin carries a share of its partition at some step" + ); + } +} + +/// A pin whose uids match no variable is reported with the entry's own +/// identity -- its name AND its uids as written -- and with the cause when +/// the model's variables carry no uid at all, which is the shape a pin +/// written against diagram-element uids takes (the fixture above before it +/// was regenerated). The name alone is not an identity: the fixture has four +/// entries named "Technical Debt Spiral". +#[test] +fn a_pin_whose_uids_match_no_variable_is_reported_with_its_identity() { + // No `assign_uids`: the variables carry none, so nothing can match. + let mut project = TestProject::new("ghost_pin") + .with_sim_time(0.0, 20.0, 0.25) + .stock("population", "100", &["births"], &[], None) + .flow("births", "population * 0.08", None) + .build_datamodel(); + project.models[0] + .loop_metadata + .push(datamodel::LoopMetadata { + uids: vec![7, 4], + deleted: false, + name: "Ghost".to_string(), + description: String::new(), + }); + + let mut db = SimlinDb::default(); + let sync = sync_from_datamodel_incremental(&mut db, &project, None); + let diagnostics = collect_all_diagnostics(&db, sync.project, simlin_engine::db::LtmOverlay::On); + let message = diagnostics + .iter() + .map(|d| format!("{:?}", d.error)) + .find(|m| m.contains("'Ghost'")) + .unwrap_or_else(|| panic!("the pin is reported; got {:?}", diagnostics)); + for needle in [ + "(uids [7, 4])", + "names 0", + "uids [4, 7] match no variable's uid", + "no variable in the model carries a uid", + ] { + assert!(message.contains(needle), "{needle:?} in {message}"); + } +} + +/// A pin with one stale uid beside one that resolves fails the two-variable +/// gate and names the stale uid; the model's other variables DO carry uids, +/// so the message does not claim none do. +#[test] +fn a_pin_with_a_stale_uid_names_it() { + let mut project = two_loop_population(); + // `population` is uid 1 (`assign_uids` numbers in declaration order); + // 99 was never assigned. + project.models[0] + .loop_metadata + .push(datamodel::LoopMetadata { + uids: vec![1, 99], + deleted: false, + name: "Half".to_string(), + description: String::new(), + }); + + let mut db = SimlinDb::default(); + let sync = sync_from_datamodel_incremental(&mut db, &project, None); + let diagnostics = collect_all_diagnostics(&db, sync.project, simlin_engine::db::LtmOverlay::On); + let message = diagnostics + .iter() + .map(|d| format!("{:?}", d.error)) + .find(|m| m.contains("'Half'")) + .unwrap_or_else(|| panic!("the pin is reported; got {:?}", diagnostics)); + assert!(message.contains("(uids [1, 99]) names 1"), "{message}"); + assert!( + message.contains("uids [99] match no variable's uid"), + "{message}" + ); + assert!( + !message.contains("no variable in the model carries a uid"), + "{message}" + ); +} + +/// A pin that drops a stale uid but whose surviving variables still form a +/// cycle scores that cycle -- and says so: a user who deleted or re-created a +/// variable would otherwise get a different loop under the pin's name with +/// no signal. The warning names the pin, its uids as written, the dropped +/// uids and the cycle scored instead; the pin still scores. +#[test] +fn a_valid_pin_that_dropped_a_stale_uid_warns_and_names_the_cycle_it_scored() { + let mut project = two_loop_population(); + // `assign_uids` numbers in declaration order: population 1, births 2, + // crowding 3, deaths 4; 99 was never assigned. + project.models[0] + .loop_metadata + .push(datamodel::LoopMetadata { + uids: vec![1, 2, 99], + deleted: false, + name: "Growth".to_string(), + description: String::new(), + }); + + let mut db = SimlinDb::default(); + let sync = sync_from_datamodel_incremental(&mut db, &project, None); + set_project_ltm_discovery_mode(&mut db, sync.project, true); + let diagnostics = collect_all_diagnostics(&db, sync.project, simlin_engine::db::LtmOverlay::On); + let message = diagnostics + .iter() + .map(|d| format!("{:?}", d.error)) + .find(|m| m.contains("'Growth'")) + .unwrap_or_else(|| panic!("the pin is reported; got {:?}", diagnostics)); + for needle in [ + "(uids [1, 2, 99])", + "dropped uids [99], which match no variable's uid", + "scored the loop its remaining variables form instead: births -> population", + ] { + assert!(message.contains(needle), "{needle:?} in {message}"); + } + + // The pin still scores: in discovery mode it is the only scored loop. + let source_model = sync.models["main"].source_model; + let ltm = model_ltm_variables(&db, source_model, sync.project); + assert!( + ltm.vars + .iter() + .any(|v| v.name == "$\u{205A}ltm\u{205A}loop_score\u{205A}pin1"), + "the surviving cycle is scored under pin1" + ); +} + +/// The not-a-cycle arm names the entry's uids as written and the dropped +/// one: `population` and `crowding` do not close a loop on their own, and +/// 99 matched nothing. +#[test] +fn a_pin_that_is_no_cycle_names_its_uids_and_the_dropped_one() { + let mut project = two_loop_population(); + // population 1, crowding 3 (`assign_uids` declaration order); 99 stale. + project.models[0] + .loop_metadata + .push(datamodel::LoopMetadata { + uids: vec![1, 3, 99], + deleted: false, + name: "Split".to_string(), + description: String::new(), + }); + + let mut db = SimlinDb::default(); + let sync = sync_from_datamodel_incremental(&mut db, &project, None); + let diagnostics = collect_all_diagnostics(&db, sync.project, simlin_engine::db::LtmOverlay::On); + let message = diagnostics + .iter() + .map(|d| format!("{:?}", d.error)) + .find(|m| m.contains("'Split'")) + .unwrap_or_else(|| panic!("the pin is reported; got {:?}", diagnostics)); + for needle in [ + "pinned loop 'Split' (uids [1, 3, 99]) do not form a closed feedback loop", + "[crowding, population]", + "; uids [99] match no variable's uid", + ] { + assert!(message.contains(needle), "{needle:?} in {message}"); + } + assert!( + !message.contains("no variable in the model carries a uid"), + "{message}" + ); +} diff --git a/test/hero_culture_ltm/hero_culture.sd.json b/test/hero_culture_ltm/hero_culture.sd.json index d189f11f3..b12760172 100644 --- a/test/hero_culture_ltm/hero_culture.sd.json +++ b/test/hero_culture_ltm/hero_culture.sd.json @@ -12,6 +12,7 @@ "stocks": [ { "name": "hero_culture", + "uid": 1, "initialEquation": "10", "units": "dimensionless", "inflows": [ @@ -24,6 +25,7 @@ }, { "name": "incidents", + "uid": 2, "initialEquation": "5", "units": "incidents", "inflows": [ @@ -36,6 +38,7 @@ }, { "name": "system_reliability", + "uid": 3, "initialEquation": "80", "units": "dimensionless", "inflows": [ @@ -48,6 +51,7 @@ }, { "name": "technical_debt", + "uid": 4, "initialEquation": "20", "units": "dimensionless", "inflows": [ @@ -62,48 +66,56 @@ "flows": [ { "name": "culture_decay", + "uid": 5, "equation": "Hero_Culture * culture_decay_fraction", "units": "dimensionless/Month", "documentation": "Natural decay of hero culture without reinforcement" }, { "name": "culture_reinforcement", + "uid": 6, "equation": "Incidents * heroic_response_fraction * hero_recognition_factor / response_time", "units": "dimensionless/Month", "documentation": "Recognition and rewards for firefighting reinforce hero culture" }, { "name": "debt_accumulation", + "uid": 7, "equation": "Incidents * heroic_response_fraction * debt_per_heroic_fix / response_time", "units": "dimensionless/Month", "documentation": "Technical debt from heroic quick fixes that skip proper solutions" }, { "name": "debt_paydown", + "uid": 8, "equation": "investment_in_prevention * debt_paydown_effectiveness", "units": "dimensionless/Month", "documentation": "Technical debt reduction through proper engineering investment" }, { "name": "incident_occurrence", + "uid": 9, "equation": "base_incident_rate * (100 - System_Reliability) / 50 + Technical_Debt * debt_incident_factor", "units": "incidents/Month", "documentation": "New incidents arising from system fragility and technical debt" }, { "name": "incident_resolution", + "uid": 10, "equation": "Incidents * (heroic_response_fraction + fundamental_fix_fraction) / response_time", "units": "incidents/Month", "documentation": "Incidents resolved through heroic intervention or fundamental fixes" }, { "name": "reliability_erosion", + "uid": 11, "equation": "Incidents * heroic_response_fraction * quick_fix_erosion / response_time + Technical_Debt * debt_erosion_factor", "units": "dimensionless/Month", "documentation": "Reliability loss from quick fixes and mounting tech debt" }, { "name": "reliability_improvement", + "uid": 12, "equation": "investment_in_prevention * prevention_effectiveness", "units": "dimensionless/Month", "documentation": "Improvements from proactive reliability investments" @@ -112,90 +124,105 @@ "auxiliaries": [ { "name": "base_incident_rate", + "uid": 13, "equation": "2", "units": "incidents/Month", "documentation": "Baseline rate of incidents in a normal system" }, { "name": "culture_decay_fraction", + "uid": 14, "equation": "0.02", "units": "1/Month", "documentation": "Natural decay rate of hero culture" }, { "name": "debt_erosion_factor", + "uid": 15, "equation": "0.02", "units": "1/Month", "documentation": "Rate at which technical debt erodes reliability" }, { "name": "debt_incident_factor", + "uid": 16, "equation": "0.05", "units": "incidents/Month", "documentation": "Incidents generated per unit of technical debt" }, { "name": "debt_paydown_effectiveness", + "uid": 17, "equation": "0.01", "units": "dimensionless", "documentation": "Rate at which prevention investment pays down debt" }, { "name": "debt_per_heroic_fix", + "uid": 18, "equation": "0.3", "units": "dimensionless/incidents", "documentation": "Technical debt created per incident fixed heroically" }, { "name": "fundamental_fix_fraction", + "uid": 19, "equation": "0.2 * (1 - Hero_Culture / 100)", "units": "dimensionless", "documentation": "Fraction of incidents fixed properly, decreases as hero culture dominates" }, { "name": "hero_recognition_factor", + "uid": 20, "equation": "0.05", "units": "dimensionless/incidents", "documentation": "How much hero culture grows per heroic incident resolution" }, { "name": "heroic_response_fraction", + "uid": 21, "equation": "0.3 * (1 + Hero_Culture / 50)", "units": "dimensionless", "documentation": "Fraction of incidents addressed heroically, increases with hero culture" }, { "name": "investment_in_prevention", + "uid": 22, "equation": "total_engineering_capacity * (1 - Hero_Culture / 100) * prevention_priority", "units": "dimensionless/Month", "documentation": "Resources devoted to proactive reliability work" }, { "name": "prevention_effectiveness", + "uid": 23, "equation": "0.02", "units": "dimensionless", "documentation": "How much reliability improves per unit of prevention investment" }, { "name": "prevention_priority", + "uid": 24, "equation": "0.003", "units": "1/Month", "documentation": "Management priority given to prevention vs. feature work" }, { "name": "quick_fix_erosion", + "uid": 25, "equation": "0.1", "units": "dimensionless/incidents", "documentation": "Reliability lost per incident resolved via quick fix" }, { "name": "response_time", + "uid": 26, "equation": "2", "units": "Month", "documentation": "Average time to resolve an incident" }, { "name": "total_engineering_capacity", + "uid": 27, "equation": "100", "units": "dimensionless", "documentation": "Total available engineering capacity (normalized)"