Skip to content

Count each node's conducting transistors, and stop walking a group of one - #21

Open
Roxxik wants to merge 2 commits into
mist64:masterfrom
Roxxik:on-degree
Open

Roxxik wants to merge 2 commits into
mist64:masterfrom
Roxxik:on-degree

Conversation

@Roxxik

@Roxxik Roxxik commented Jul 30, 2026

Copy link
Copy Markdown

Two thirds of the groups recalcNode builds hold a single node. Every one of
them is discovered the same way: clear the group, add the node, test the gate of
every transistor that touches it, find nothing, and settle. The information
needed to skip all of that is cheap to maintain, and this PR maintains it.

The counter

A transistor is conducting exactly when its gate node is high. So the number of
conducting transistors on a node only changes where a node value changes, and
that is a place the solver already walks a list of the nodes behind that node's
transistors. Counting there is nearly free.

The existing list does not suit a counter. It is deduplicated, and it drops vss
and vcc, whereas a counter has to move once per transistor endpoint and must
see a node held to ground as connected. So this adds an undeduplicated endpoint
list that keeps the power nodes. That list turns out to hold the same nodes as
nodes_dependant, which is the deduplicated version of it, so the two merge: the
low path now walks the endpoints once and both counts them down and queues them,
and listout_add filters the repeats out as it already did. nodes_dependant
is gone.

vss and vcc are deliberately absent from the endpoint list and never have a
value assigned, so their counters would sit at zero forever and wrongly qualify
them. They are parked on a nonzero value at setup.

The fast path, twice

First cut: a node with no conducting transistor at all takes its value from
its own pullup or pulldown, and with neither it keeps the value it has. That is a
third of all recalcNode calls, settled without touching a group.

Second cut: the rest of the groups of one are nodes whose only conducting
transistors lead to vss or vcc. The group walk stops at the rails anyway (it
has to, or it would run on into every other group deriving its value from the
same rail), so those nodes were paying for a walk that could not have found
anyone.

Counting rail connections separately from ordinary ones catches them. A node is a
group of one exactly when the ordinary count is zero, whatever it is tied to on
the rails, and what remains decides the value in the order addNodeToGroup would
have applied it: vss beats vcc, and vcc beats the node's own pulldown or
pullup.

The three counters live in one word, ten bits each, so the test stays a single
load and mask. Ten bits is ample: the busiest node on this netlist reaches 12, 9
and 1. What an endpoint contributes is fixed by what sits on the far side of its
transistor, so it is computed once at setup and stored next to the endpoint,
which keeps the update a branchless +=.

Fast-path coverage goes from a third of recalcNode calls to two thirds.

Verification

This PR adds a -DDEBUG_ON_DEGREE build. A counter that drifts out of step with
the netlist does not crash: it silently skips work that needed doing, and what
surfaces is a wrong value on some node far away from the cause. So the build
recomputes every counter from scratch and asserts it against the maintained one
on entry to recalcNodeList, catching a drift where it happens rather than where
it eventually shows. It is clean over a whole run.

measure is byte-identical over all 256 opcodes, with and without the checker.
cbmbasic --benchmark prints the same output as master, down to the final CPU
state line.

Neither of those looks past the registers, so I checked the netlist itself as
well. A patch I keep locally, not part of this PR, folds the values of all 1725
nodes into a running hash every sixteenth cycle and prints it at the end of a
run. Master and this branch end on the same hash on every workload I have, so no
node takes a different value at any point. Both commits were checked separately.

Performance

Against master, cbmbasic to READY., minimum of twelve interleaved runs:

master this PR
wall 0.8675 s 0.5972 s -31.2%
cycles 2,951,788,903 2,041,778,467 -30.8%
instructions 7,129,394,678 4,803,623,873 -32.6%
branch misses 51,796,352 32,376,654 -37.5%

Split between the two commits, each measured on cbmbasic cycles against its own
parent: skipping the nodes no transistor connects to anything is -5.13%, and the
group-of-one fast path on top of it is -26.72%. A much longer
instruction-level workload I run locally (Harte) shows the same figure as cbmbasic for
the pair, so this is not specific to what BASIC happens to exercise on startup.

The counter costs memory. The endpoint list and the per-endpoint contribution it
carries take the allocations from roughly 89 KB to 111 KB.

Note for review

The fast path is a dozen lines and reads straight off the counter. The counter is
the fiddly part: it has to stay exactly in step with every node value change for
a whole run, and outside the assertion build nothing in the solver would notice
if it drifted.

I used an LLM to help me out in this work, but I manually reviewed all changes made.

Roxxik added 2 commits July 30, 2026 02:59
A third of recalcNode's calls are for a node that no turned-on
transistor connects to anything, which can only take its own pullup or
pulldown value, or keep the one it has. Count the turned-on transistors
per node and those calls become a load and a branch instead of a group
walk. The counters only move where a node value changes, which already
walks the nodes behind that node's transistors, so maintaining them is
nearly free.

nodes_dependant cannot serve as that walk: it is deduplicated and drops
vss and vcc, while a counter has to move once per endpoint and must see
a node held to ground as connected. An undeduplicated endpoint list
replaces it, with listout_add filtering the repeats as it already did.

-DDEBUG_ON_DEGREE asserts every counter against a from-scratch recount.
measure is byte-identical over all 256 opcodes with and without it.
Two thirds of the groups recalcNode builds hold a single node, but the
on-degree fast path only catches the half of them nothing is connected
to. The rest are alone because their only turned-on transistors lead to
vss or vcc, and the walk stops at the rails anyway.

Count rail connections apart from ordinary ones, three counters packed
into one word so the test stays a single load and mask. A node is alone
exactly when the ordinary count is zero, and what remains decides the
value in addNodeToGroup's order: vss beats vcc, vcc beats the node's own
pulldown or pullup. Fast path coverage goes from a third of recalcNode
calls to two thirds.

cbmbasic to READY. is 26.72% fewer cycles than the commit before it,
minimum of twelve interleaved runs. measure is byte-identical over all
256 opcodes, with and without the on-degree checker.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant