You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
The panel's is_real vote is a hard binary ("true issue or false positive — purely epistemic", includes/panel-concern-brief.md:8-10). That is the right shape for its original job — killing hallucinations (a finding describing code that isn't there, or misreading what is). It is the wrong shape for a conditional / stochastic hazard: a defect whose mechanism is 100% present
in the code but whose manifestation depends on a future or external state. The binary gives such a
finding nowhere honest to land, so a conscientious panellist routes it to not_real, and the true
signal is destroyed at the vote.
This is the per-vote end of the same pipe as #61 / #62 (which concern the aggregation /
synthesiser end). Track together.
Concrete evidence (n=1, retrospective)
Scoped quality A/B: 5 code-review-suite specialists were run standalone against commit cf9bc9d of HavenEngineering/finance-erp-apps PR #158 (the exact commit an external reviewer, "Marlon", reviewed
with his own tool), and scored against his 5 findings.
Relevant finding: ZB61 IndexOfOptional silent-failure in MarginReportReader.cs. The A&L
sub-department column is read optionally (missing → "") rather than required (throw). If ZB61 is
ever absent from A&L output (report-layout drift, path edited in MarginReportFilter but not here —
the path is duplicated as a private const in both), every A&L row's sub-department silently blanks
to "", which reads as the legitimate value 000 = None — wrong data shown to finance during the
parallel-run sense-check, with no error.
The mechanism is unconditionally present in the code.
Whether it bites is conditional (depends on future report/layout state).
Our correctness specialist stood on the exact line and did not raise this. It raised a different, speculative concern ("verify the IsDescription guard is present") that is provably false — the guard is at MarginReportReader.cs:82. Its own prose hedged: "I cannot see the full
body… this may already be handled." — i.e. its true internal state was uncertain, but with only
a binary to land on it laundered that into a confident-sounding Important that was actually a
coin-flip.
So even before the panel: a genuinely uncertain internal state was quantised into a false certainty.
The two-stage quantisation problem
Per-vote: each panellist rounds an internal probability to {real, not_real}.
Aggregation: three binaries collapse to a consensus fraction that can only be {0, 33, 67, 100}.
A hazard all three independently read as "plausible, ~40%" rounds unanimously to not_real → 0/3 → dropped. The "live but uncertain" signal is destroyed at the vote, not merely diluted at
aggregation.
Why the existing severity machinery does not save it
The brief clearly tried to handle stochasticity — but on the severity axis, which fires too
late:
panel-concern-brief.md:20-30 — "Severity rates the impact if the issue manifested… decided before any mitigation… a coarse control reduces likelihood, not impact, and does not lower
severity."
That rule only engages after a panellist has already decided to call the finding is_real: true.
The lossy step is upstream of that, at is_real itself, where "true issue vs false positive"
invites routing "doesn't manifest in any state I can see today" straight to not_real. The
likelihood-doesn't-lower-severity rule never gets to fire, because the finding was already killed at
the epistemic gate.
The genuine counter-tension (why the binary exists)
The whole disposition block (panel-concern-brief.md:34-48: "be pedantic, default to raise, never
hedge to a note") is engineered to stop reviewers wishy-washing everything into a mushy middle. A
naive graded-confidence slider invites exactly that drift — everything settles at 50 and nothing
clears a bar. So this is not "binary bad, slider good."
The real diagnosis: the binary conflates two genuinely different questions:
Does the described mechanism exist in the code? — a true epistemic gate; keep it binary and
decisive (this is the hallucination filter, and it works).
How likely is it to bite? — a probability that currently has no home and gets smuggled into
the is_real vote, corrupting it.
Candidate direction (for brainstorming, not decided)
Split the gate. Make is_real strictly "the described mechanism exists in the code" (pure
hallucination filter, stays binary). Give conditional hazards a separate likelihood / trigger
dimension that feeds the rubric alongside impact — so a latent-but-unlikely defect survives as
"real mechanism · conditional trigger · high impact" instead of vanishing at 0/3. This preserves the
pedantic "default to raise" spine while giving stochastic findings an honest home.
Open questions for brainstorming:
Does a likelihood axis reintroduce the exact mushy-middle drift the binary was defending against?
How does the rubric (includes/verdict-rubric.md) combine a likelihood dimension with severity +
tractability without exploding the tier logic?
Is this better solved upstream (specialist prompts learning to raise the conditional finding
crisply) so the panel never has to adjudicate a mis-stated one? Note the comment-truth findings in
the same A/B were an origination failure, not an adjudication one — the panel only ever grades
what it's handed (plus what it independently re-derives).
Both are the aggregation-end manifestations of "the pipeline is lossy about low-consensus /
uncertain findings"; this issue is the per-vote-end manifestation.
Provenance
Retrospective analysis only; nothing posted to PR #158 (merged). n=1 — one PR, one commit. The split
is interpretable and consistent, not statistically established. Needs a proper brainstorming pass
before any prompt/rubric change.
Summary (needs brainstorming — not a settled fix)
The panel's
is_realvote is a hard binary ("true issue or false positive — purely epistemic",includes/panel-concern-brief.md:8-10). That is the right shape for its original job — killinghallucinations (a finding describing code that isn't there, or misreading what is). It is the
wrong shape for a conditional / stochastic hazard: a defect whose mechanism is 100% present
in the code but whose manifestation depends on a future or external state. The binary gives such a
finding nowhere honest to land, so a conscientious panellist routes it to
not_real, and the truesignal is destroyed at the vote.
This is the per-vote end of the same pipe as #61 / #62 (which concern the aggregation /
synthesiser end). Track together.
Concrete evidence (n=1, retrospective)
Scoped quality A/B: 5
code-review-suitespecialists were run standalone against commitcf9bc9dofHavenEngineering/finance-erp-appsPR #158 (the exact commit an external reviewer, "Marlon", reviewedwith his own tool), and scored against his 5 findings.
Relevant finding: ZB61
IndexOfOptionalsilent-failure inMarginReportReader.cs. The A&Lsub-department column is read optionally (missing →
"") rather than required (throw). If ZB61 isever absent from A&L output (report-layout drift, path edited in
MarginReportFilterbut not here —the path is duplicated as a private const in both), every A&L row's sub-department silently blanks
to
"", which reads as the legitimate value000 = None— wrong data shown to finance during theparallel-run sense-check, with no error.
correctnessspecialist stood on the exact line and did not raise this. It raised adifferent, speculative concern ("verify the
IsDescriptionguard is present") that is provablyfalse — the guard is at
MarginReportReader.cs:82. Its own prose hedged: "I cannot see the fullbody… this may already be handled." — i.e. its true internal state was uncertain, but with only
a binary to land on it laundered that into a confident-sounding
Importantthat was actually acoin-flip.
So even before the panel: a genuinely uncertain internal state was quantised into a false certainty.
The two-stage quantisation problem
{real, not_real}.{0, 33, 67, 100}.A hazard all three independently read as "plausible, ~40%" rounds unanimously to
not_real→ 0/3 →dropped. The "live but uncertain" signal is destroyed at the vote, not merely diluted at
aggregation.
Why the existing severity machinery does not save it
The brief clearly tried to handle stochasticity — but on the severity axis, which fires too
late:
panel-concern-brief.md:20-30— "Severity rates the impact if the issue manifested… decidedbefore any mitigation… a coarse control reduces likelihood, not impact, and does not lower
severity."
That rule only engages after a panellist has already decided to call the finding
is_real: true.The lossy step is upstream of that, at
is_realitself, where "true issue vs false positive"invites routing "doesn't manifest in any state I can see today" straight to
not_real. Thelikelihood-doesn't-lower-severity rule never gets to fire, because the finding was already killed at
the epistemic gate.
The genuine counter-tension (why the binary exists)
The whole disposition block (
panel-concern-brief.md:34-48: "be pedantic, default to raise, neverhedge to a note") is engineered to stop reviewers wishy-washing everything into a mushy middle. A
naive graded-confidence slider invites exactly that drift — everything settles at 50 and nothing
clears a bar. So this is not "binary bad, slider good."
The real diagnosis: the binary conflates two genuinely different questions:
decisive (this is the hallucination filter, and it works).
the
is_realvote, corrupting it.Candidate direction (for brainstorming, not decided)
Split the gate. Make
is_realstrictly "the described mechanism exists in the code" (purehallucination filter, stays binary). Give conditional hazards a separate likelihood / trigger
dimension that feeds the rubric alongside impact — so a latent-but-unlikely defect survives as
"real mechanism · conditional trigger · high impact" instead of vanishing at 0/3. This preserves the
pedantic "default to raise" spine while giving stochastic findings an honest home.
Open questions for brainstorming:
includes/verdict-rubric.md) combine a likelihood dimension with severity +tractability without exploding the tier logic?
crisply) so the panel never has to adjudicate a mis-stated one? Note the comment-truth findings in
the same A/B were an origination failure, not an adjudication one — the panel only ever grades
what it's handed (plus what it independently re-derives).
Related
uncertain findings"; this issue is the per-vote-end manifestation.
Provenance
Retrospective analysis only; nothing posted to PR #158 (merged). n=1 — one PR, one commit. The split
is interpretable and consistent, not statistically established. Needs a proper brainstorming pass
before any prompt/rubric change.