Conversation
A third of recalcNode's calls are for a node that no turned-on transistor connects to anything, which can only take its own pullup or pulldown value, or keep the one it has. Count the turned-on transistors per node and those calls become a load and a branch instead of a group walk. The counters only move where a node value changes, which already walks the nodes behind that node's transistors, so maintaining them is nearly free. nodes_dependant cannot serve as that walk: it is deduplicated and drops vss and vcc, while a counter has to move once per endpoint and must see a node held to ground as connected. An undeduplicated endpoint list replaces it, with listout_add filtering the repeats as it already did. -DDEBUG_ON_DEGREE asserts every counter against a from-scratch recount. measure is byte-identical over all 256 opcodes with and without it.
Two thirds of the groups recalcNode builds hold a single node, but the on-degree fast path only catches the half of them nothing is connected to. The rest are alone because their only turned-on transistors lead to vss or vcc, and the walk stops at the rails anyway. Count rail connections apart from ordinary ones, three counters packed into one word so the test stays a single load and mask. A node is alone exactly when the ordinary count is zero, and what remains decides the value in addNodeToGroup's order: vss beats vcc, vcc beats the node's own pulldown or pullup. Fast path coverage goes from a third of recalcNode calls to two thirds. cbmbasic to READY. is 26.72% fewer cycles than the commit before it, minimum of twelve interleaved runs. measure is byte-identical over all 256 opcodes, with and without the on-degree checker.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Two thirds of the groups
recalcNodebuilds hold a single node. Every one ofthem is discovered the same way: clear the group, add the node, test the gate of
every transistor that touches it, find nothing, and settle. The information
needed to skip all of that is cheap to maintain, and this PR maintains it.
The counter
A transistor is conducting exactly when its gate node is high. So the number of
conducting transistors on a node only changes where a node value changes, and
that is a place the solver already walks a list of the nodes behind that node's
transistors. Counting there is nearly free.
The existing list does not suit a counter. It is deduplicated, and it drops
vssand
vcc, whereas a counter has to move once per transistor endpoint and mustsee a node held to ground as connected. So this adds an undeduplicated endpoint
list that keeps the power nodes. That list turns out to hold the same nodes as
nodes_dependant, which is the deduplicated version of it, so the two merge: thelow path now walks the endpoints once and both counts them down and queues them,
and
listout_addfilters the repeats out as it already did.nodes_dependantis gone.
vssandvccare deliberately absent from the endpoint list and never have avalue assigned, so their counters would sit at zero forever and wrongly qualify
them. They are parked on a nonzero value at setup.
The fast path, twice
First cut: a node with no conducting transistor at all takes its value from
its own pullup or pulldown, and with neither it keeps the value it has. That is a
third of all
recalcNodecalls, settled without touching a group.Second cut: the rest of the groups of one are nodes whose only conducting
transistors lead to
vssorvcc. The group walk stops at the rails anyway (ithas to, or it would run on into every other group deriving its value from the
same rail), so those nodes were paying for a walk that could not have found
anyone.
Counting rail connections separately from ordinary ones catches them. A node is a
group of one exactly when the ordinary count is zero, whatever it is tied to on
the rails, and what remains decides the value in the order
addNodeToGroupwouldhave applied it:
vssbeatsvcc, andvccbeats the node's own pulldown orpullup.
The three counters live in one word, ten bits each, so the test stays a single
load and mask. Ten bits is ample: the busiest node on this netlist reaches 12, 9
and 1. What an endpoint contributes is fixed by what sits on the far side of its
transistor, so it is computed once at setup and stored next to the endpoint,
which keeps the update a branchless
+=.Fast-path coverage goes from a third of
recalcNodecalls to two thirds.Verification
This PR adds a
-DDEBUG_ON_DEGREEbuild. A counter that drifts out of step withthe netlist does not crash: it silently skips work that needed doing, and what
surfaces is a wrong value on some node far away from the cause. So the build
recomputes every counter from scratch and asserts it against the maintained one
on entry to
recalcNodeList, catching a drift where it happens rather than whereit eventually shows. It is clean over a whole run.
measureis byte-identical over all 256 opcodes, with and without the checker.cbmbasic --benchmarkprints the same output as master, down to the final CPUstate line.
Neither of those looks past the registers, so I checked the netlist itself as
well. A patch I keep locally, not part of this PR, folds the values of all 1725
nodes into a running hash every sixteenth cycle and prints it at the end of a
run. Master and this branch end on the same hash on every workload I have, so no
node takes a different value at any point. Both commits were checked separately.
Performance
Against master, cbmbasic to
READY., minimum of twelve interleaved runs:Split between the two commits, each measured on cbmbasic cycles against its own
parent: skipping the nodes no transistor connects to anything is -5.13%, and the
group-of-one fast path on top of it is -26.72%. A much longer
instruction-level workload I run locally (Harte) shows the same figure as cbmbasic for
the pair, so this is not specific to what BASIC happens to exercise on startup.
The counter costs memory. The endpoint list and the per-endpoint contribution it
carries take the allocations from roughly 89 KB to 111 KB.
Note for review
The fast path is a dozen lines and reads straight off the counter. The counter is
the fiddly part: it has to stay exactly in step with every node value change for
a whole run, and outside the assertion build nothing in the solver would notice
if it drifted.
I used an LLM to help me out in this work, but I manually reviewed all changes made.