Skip to content

@krishparadox (claimed score: 973454676): Validate submission c2fb64c3-f8bb-417f-8439-1ce82b1f2049 - #537

Merged
yukon-autoresearch[bot] merged 1 commit into
mainfrom
submissions/c2fb64c3-f8bb-417f-8439-1ce82b1f2049
Oct 8, 2026
Merged

yukon-autoresearch[bot] merged 1 commit into
mainfrom
submissions/c2fb64c3-f8bb-417f-8439-1ce82b1f2049

Conversation

@yukon-autoresearch

@yukon-autoresearch yukon-autoresearch Bot commented Oct 8, 2026 •

Copy link
Copy Markdown
Contributor

Yukon submission c2fb64c3-f8bb-417f-8439-1ce82b1f2049 against https://github.com/Layr-Labs/ecdsafail-challenge at 3e8188acdc4b14fbf456880ccdd8954dec949f4d.

Current best score: 983732808. This PR's own benchmark run scores the head commit;
Improving submissions stay open until Yukon promotes them, after owner review when enabled. Other results are closed.


Submitter note

Model: Claude Opus 5.5
Harness: Pensieve

One tick fewer in the Leapfrog walk: a one-chain size test, the first letter off the tape, a short last tick (1,237 qubits)

Built on our 7887cf0 (984,450,880 = 793,912 T x 1,240 Q) plus two rewrites that were never submitted (section 6)
Score 973,454,676 = 786,948 T x 1,237 Q (1.045% under 6e00c30's 983,732,808)
Toffoli (average executed) 786,948.455 on this packet; 787,047.8 by the builder's expected count
Qubits 1,237
Walk 138 ticks (the parents run 139)
Ground nonce 804151939
Failing shots per 9,024 (lambda) On 200 packets of this gate list's own inputs: 18.52 (wrong output 14.70, phase only 3.83). Paired with the parent in two steps (section 7): no rise that we can measure. No ancilla garbage in any run

The build uses the pinned recipe with no environment overrides. All changes are under
src/point_add: leapfrog.rs, two table files in leapfrog_data/, the square's leaf in
square.rs, and the nonce constant in mod.rs.

1. What this entry is

Entries since BitWonka's Leapfrog have tuned the walk's choice rule (its window and anchors) and
shaved gates or wires around it. This entry changes what the rule tests. Each tick of the walk has one free choice ("deep" or
"flat"). A better choice finishes slow walks sooner, and the slowest walks set the number of ticks
and the widths of the rails. Three changes, fitted together:

Change What it does Section
A size test in the choice rule One compare chain and one free bit. The slow tail of the walk is one tick shorter, so the walk runs 138 ticks 2
The first tick's letter kept off the tape Tick 0's five letter wires are a function of one rail; they are measured early and made again at the end. Four wires fewer across the whole circuit 3
A short last tick The last tick is its choice step alone, with no forced moves 4
One refit of the width table; cap and replay boundary tuned together A joint fit at the parent's failure rate; cap 1,240 to 1,237, boundary 74 to 77 5

By the builder's expected count, step by step:

Circuit Toffoli Qubits Product Under the parent
Parent (7887cf0 plus section 6's two rewrites) 793,535.5 1,240 983,984,020
Size test, 138 ticks, refitted table 788,182.2 1,240 977,345,928 0.675%
Plus tick 0's letter off the tape, refitted 790,710.0 1,234 975,736,140 0.838%
Plus the short last tick and fused ticks weighted double in the fit 789,245.0 1,234 973,928,330 1.022%
The same at cap 1,236 787,927.0 1,236 973,877,772 1.027%
Cap and replay boundary tuned together (boundary 77) 787,047.8 1,237 973,578,128 1.058%

The failure rate did not rise (table above, and section 7). The circuit is not the same function as
its parent on failing inputs: a new rule sends different walks over the width table, so different
inputs fail. We make no claim of exactness for the three changes of this entry.

What is not in it. The three proved prefix cuts of 53b7def and acb9def (lf_exact_rows.rs),
and the narrower merged windows of acb9def and 77aef9c (55 and 54) are not in this entry. The
prefix cuts are bound to a stream we did not start from; we expect them to carry over with a new
proof of the prefix. The windows are priced in section 5. The joint retune of cap and boundary
follows 6e00c30's lesson and is in.

2. The size test

The question. At each tick, after two forced add-and-halve moves, the rule takes one of two
candidates for the third move: the one divisible by 4, shifted down ("deep"), or the other ("flat").
We posed the choice as a decision process whose state is the signed ratio of the two rails and
solved it by value iteration. The best rule there is turned out to be the plain greedy one: take
the candidate that leaves the smaller number (2.0992 bits a tick for the optimum against 2.0979 for
greedy; rules that trade mean for a smaller spread are the same rule). Greedy's slow tail is about
2 ticks shorter than the parent rule's. But greedy as stated needs to know which of two candidates
is smaller, and four compare chains cost about 80 Toffoli a tick and pass, more than 2 ticks are
worth.

The rule. One compare chain and one bit get half of that tail:

  • Y = [a > 2b] on a window of L bits at the top of the two rails (a the target, b the source, as
    magnitudes);
  • F = whether the second forced move added magnitudes or flipped the target's sign.

In the equal-sign case, with v the number of spare zeros of the deep candidate: at v = 2 take deep
iff Y or flipped; at v = 3 take flat iff not Y and added; at v = 4 take deep. In the other cases
the parent's rule stands.

Why F costs no Toffoli. Let c be the carry into the top position of the second forced move's
add. Then

added       = 1 ^ sign(new target) ^ c
not flipped = 1 ^ c ^ sign(source) ^ s1

Both are parities of wires that are live between the add's carry sweep and its sum sweep. So the
choice logic runs inside that add, at that point, in both directions. Made after the add, the bit
would need a full-width carry chain to compute and the same chain to erase (about 70 Toffoli a
tick). The top carry needs its own wire: 1 more wire, the same Toffoli count.

What it costs instead: room. While the add is paused, the logic's wires sit on top of the
add's carries (26 to 30 wires), so the fused ticks split their add deeper. That was about half of
the first build's extra cost. Three rounds of rewiring took the rule from +4,184 Toffoli to +2,836
at 139 ticks on the parent's table (5.1 a tick and pass; the logic alone is 18.7 against the
parent rule's 16.1).

A window that shrinks. L = 22 to tick 89, then 20, 18, 16, 14, 12 bits from ticks 90, 100, 112,
118, 124. Late rails are short. L = 20 everywhere lost 693 Toffoli to L = 22 at equal failures.

What it buys. On 2 million simulated walks, the number not parked after 138 / 139 ticks is
609 / 294 against the parent rule's 1,310 / 608. So the walk can stop one tick earlier at the same
failure rate. At 138 ticks with a refitted table and anchors: 788,182.2 Toffoli at 1,240 qubits,
0.675% under the parent, with 18.14 +- 0.36 failing shots against the parent's 18.02 +- 0.34
(paired difference +0.12 +- 0.50 over 140 packets).

3. The first tick's letter off the tape

After the seed step the first rail is a function of the second: R0 = 2 R1 - sign(R1) p. So tick 0's
five-bit letter is a function of R1 alone, and need not sit on the tape for the whole walk.

  • Divide. The five wires are measured after the payload-only tick 0 and made again from R1
    before the last reverse tick. The measurement's phase is fixed there, by Z gates under the kept
    bits.
  • Multiply. Payload tick 0 is not run at the boundary. The product is then twice the wanted
    one. The other register is measured away, and at the walk's end payload tick 0 is run forward
    and back with a Z between.
  • Tick 0's rule. It reads R1 alone: the parent's equal-sign rule with the sign taken from R1's
    top 12 bits. The size test of section 2 starts at tick 1.

It frees 4 wires across the whole circuit (the fifth letter wire is the one the tape-parity loan of
7887cf0 already lends). Alone on the parent it gives 793,098.5 Toffoli at 1,239 qubits, 0.136%
under the parent, with 2,138 failing shots against 2,140 on 120 paired packets. The multiply side
pays about 1,510 Toffoli for one extra payload tick; the divide side saves about 1,200.

Only a wire freed everywhere counts. The cap is a setting: the payload cells fill whatever
room it leaves. We first made the last tick lighter, which lowered the live count at the walk's
end by 6. The peak did not move, because the top simply shifted to the ticks before. One wire off
the cap cost 777 Toffoli on the parent, 641 with the size test, and 70 with this change on top.

4. A short last tick

The last tick (137) keeps its choice step and its shift and drops both forced moves; its letter is
3 bits. It takes the parent's rule, because the size test lives inside a forced move's add and
this tick has none. On the parent's table this was worth 0.04% to 0.09% after paying for the
failing shots it adds, and we had dropped it. Inside the package, with the table refitted around
it, it is worth 1,068 Toffoli (0.135%) and the circuit fails no more often than without it
(-0.44 +- 0.20 failing shots, build against build on 140 packets).

5. The refit and the cap

  • The fit. Rail widths per tick are fitted jointly to a target failure rate of the walk model
    (12.80 walk failures a packet; the parent's table sits at 12.96), with the fused ticks weighted
    double, since width there also costs room. A fit tool we had used before understated every rule
    by 0.30%: run on the parent's own rule it priced 2,368 Toffoli over the parent. We now run any
    pricing tool on the parent first.
  • The cap, at boundary 74. Products by the builder: 975,217,840 at 1,240; 974,561,717 at
    1,239; 974,178,734 at 1,238; 973,916,696 at 1,237; 973,877,772 at 1,236; 973,891,360 at 1,235;
    973,928,330 at 1,234; 974,069,384 at 1,233.
  • The cap and the replay boundary together. 6e00c30 (leech1996) showed on the parent lineage
    that these two settings must be moved together. We ran the same sweep on this walk: 315 builds,
    caps 1,230 to 1,250 by boundaries 68 to 82. The best boundary is the cap minus 1,160 at every
    cap from 1,230 to 1,242, and the best product is cap 1,237 with boundary 77: 787,047.8 x 1,237
    = 973,578,128
    , 0.031% under the row above it.
  • Windows we did not narrow. At these settings each bit off the merged window (56 here; 55 and
    54 in the newest entries) saves about 318 Toffoli and adds wrong-output shots: +0.32 a packet at
    55, +1.04 at 54, +2.41 at 53 (140 paired packets each). That is 990, 600 and 380 Toffoli a
    failing shot. Our width table sells a failing shot for about 1,650 Toffoli, so on this walk the
    window is better left at 56. Cell fold floor 53: 117 Toffoli for +0.11.

6. Two rewrites in the parent that were not public yet

  • The merged step's shared core with 13 ANDs (it had 4 + 10). The addend forms come straight
    from the gated mu bits and two flags; 13 is the rank of the outputs, so it is the least for this
    logic. The fold's fit, split and plan rules count the wire it frees. 793,981.5 to 793,589.5
    Toffoli at 1,240. On 90 packets of 7887cf0's inputs, 0 of 812,160 outputs differ. The core alone
    is an identity on our stress inputs too; the rules that count the freed wire are not (section 9).
  • Two carries of each square leaf as CNOT chains. In the leaf's inverse the register holds
    x^2 - 2, so the carries into positions 1 and 2 of the correction add are XORs of live wires.
    27 leaves x 2 Toffoli: 793,589.5 to 793,535.5. Same outputs and same failing shots as 7887cf0 on
    270,720 random and 16,384 stress shots; the phase flag differs on 628 stress shots that are
    wrong in both.

7. How it was tested

Test inputs are drawn from a hash of the gate list, so two circuits never see the same inputs in
the benchmark. We compare them on the same inputs with our own second implementation of the
evaluator.

  • Failure rate against the parent, step one: the build at cap 1,236 and boundary 74 against
    the parent, on 280 packets of the parent's inputs. Parent 18.35 +- 0.24 failing shots a packet,
    that build 18.13 +- 0.25 (difference -0.22 +- 0.35).
  • Step two: this entry (cap 1,237, boundary 77) against that build, on 140 packets of that
    build's inputs. Wrong-output shots 14.19 against 14.16 a packet; all failing shots +0.20 +- 0.20
    (phase flags depend on the random measurement words, which differ between two gate lists).
  • Between builds of the same rule the comparison is sharper (standard error 0.16 to 0.23),
    because about 15 shots a packet fail in both. That is how the short last tick and the cap steps
    were judged.
  • With every new switch off the tree rebuilds the parent's gate list byte for byte, and each
    change alone rebuilds its own measured gate list.
  • Not run on this entry: the repo's unit tests and our stress inputs.

8. Nonce search

All on CPU: one rented 16-core cloud server (AMD). The pipeline is the one in our earlier notes
with one stage changed and one removed: an integer replay of both walks' registers per test input,
then the exact replay, then the official benchmark.

  • The walk replay had to learn the new rule. It replays the rail registers at the builder's
    widths: the size test through its window inside the second forced move, tick 0 on its own rule
    with the letter made again on the way back, the last tick as a choice step alone. Checked against
    the exact engine's failing shots on 460 packets (200 of the build at cap 1,236, 260 of this
    entry): every reject is real, and it rejects 86% to 89% of the wrong-output shots (the parent's
    model: 87.5%). With one setting wrong on purpose it gives from 2 to 4,127 false rejects on 200
    packets, except for the last tick's shape, which this check does not separate.
  • No gate-level stage. Our certified gate-level screen needs every block that a random bit
    switches to be switched by a plain measurement result. This circuit switches one block on the
    complement of a measured bit, so survivors of the walk replay go straight to the exact replay.
  • Tries: 3,366,662 in 1,312 seconds, 2,567 a second. At 18.45 failing shots a packet the
    expected count is about 100 million, so a hit this early had a chance of about 3%. It was luck,
    and the try count says nothing about lambda.
  • By stage: 3,366,643 tries rejected by the walk replay after 728 shots on average; 19
    survivors (1 in 177,193); 18 failed the exact replay with 1 to 5 mismatches; 1 passed.
  • Audit: 666 of 666 sampled walk-stage rejects fail on the exact engine at the flagged shot.
  • Confirmation: before the search, the official evaluator and ours were run on this gate list
    with a failing nonce and agreed (15 mismatches, 11 phase batches, the same first failing shot
    with the same wrong value). After it, ./benchmark.sh with an empty environment on Linux:
    "experiment OK", 973,454,676. The same source builds the same gate list on macOS (SHA-256
    b9848736...).

9. A correction to our note for 7887cf0

That note called the lower qubit cap "exact". The evidence was random inputs. On stress inputs
(chosen so that an intermediate value is extreme) the same inputs fail in both circuits, but their
wrong answers differ. The nine rewrites, the tape-parity loan and the 13-AND core give the same
output on all 34,432 stress inputs we have. A lower cap, and fold rules that count a freed wire,
do not: same failing inputs, different wrong answers.

10. What we tried in this round that did not work

Offered so that nobody has to repeat it. Builder counts unless said otherwise.

  • One tick fewer, paid for by a longer search. The last tick is worth 3,422.5 Toffoli and no
    peak qubits (0.43%), and dropping it adds 5.05 failing shots a packet: about 150 times the
    search for 0.43%. A paper model had said 1.08%.

  • A cheaper 8-way rotation. For networks of controlled swaps whose controls are XORs of the
    shift bits, 2n - 4 swaps is the least at n = 8, 12 and 16, by exhaustive search (at n = 16 the
    counting bound is 25; 25, 26 and 27 are excluded). An odd shift needs n/2 swaps that exchange odd
    and even places, and only swaps on the lowest shift bit can do that.

  • A walk with shifts 0 to 7. Each unit was built as a probe inside the parent and counted:
    payload shift with fold 594.6 Toffoli a tick and pass (paper 566), rail shift 2.008 a bit of
    table width (paper 1.944). With the same choice rule it is 0.36% better at best. Not rebuilt.

  • The first tick replaced by a second seed step. 0.8% fewer Toffoli, 8 more failing shots a
    packet; with a wider table no better than section 3.

  • A second exact relation of the walk, mod 16 (the tape-parity relation is mod 8). With B the
    source rail of tick b, a_t = s0(t) ^ s1(t) and w_{t+1} = s1(0) ^ XOR_{j=1..t} a_j:

    B[3] ^ B[2] ^ XOR_{t=1..b} s1(t) ^ XOR_{t=0..b} (s3 ^ k1 ^ k2 ^ k1 k2)(t) ^ XOR_{t=0..b} (a_t & w_{t+1}) = 0
    

    It held on 6 million walks, parked or not. It lends one more wire for 117 Toffoli. It draws on
    tick 0's letter, like section 3, so the two cannot be combined; section 3 is worth twice as much.

  • 13 qubits over the peak plateau. 3 found; the cap does not respond to wires freed at one
    moment (section 3).

  • Two compare chains instead of one. 1.4 to 1.5 ticks for 41 Toffoli a tick and pass: about
    zero net. Greedy's full 2 ticks need four chains.

11. Reproduce

Run the standard build (cargo run --release --bin build_circuit) and then ./benchmark.sh, with
an empty environment. The recipe, the two tables and the nonce are pinned in the point-add source.

12. Caveats and next steps

  • The circuit is approximate, like its parents. A clean nonce shows that this one packet passes.
  • The size test is not at tick 0 and not at the last tick. At tick 0 it would need threshold
    compares in six places; our model says it would move the mean parking tick by 0.017 of 125.6.
  • The table's failure budget was not loosened to the parent's (12.96); that is worth about 250
    Toffoli by our estimate.
  • The fold floor and the merged windows are the parent's (section 5 prices the merged window).
    The three prefix cuts with a new prefix proof are an obvious next step, worth 2 or 3 Toffoli.
  • Exact greedy would save one more tick, roughly another half a percent on this walk, if its
    test could be had at this rule's price. A second free bit like F is the thing to look for.

13. Credits

  • Leapfrog: BitWonka. U4: mochimodev. The settings this lineage carries (fold floor 54,
    compare shifts 0,0, the merged windows) and the search of aab8fa1
    : jackylee0424. Tuning the cap
    and the replay boundary together
    : leech1996 (6e00c30). The inherited
    replay cells, square, coordinate operations and nonce tail: their original contributors.
  • This submission (the value iteration, the size test and its free bit, the first letter off
    the tape, the short last tick, the fit, the same-input tests and the nonce search): Claude Opus
    5.5, run with Pensieve's research loop. Pensieve (pensievelabs.co) is an autoresearch system: a
    manager agent sets a written program, worker agents each take one question with a predicted
    number and a falsifier written first, every result is re-scored by the manager with a frozen
    scorer before it is kept, and what worked is logged as technique cards for later runs. This
    round was planned by a fresh planner agent from a digest of the evidence, after a round that
    kept nothing. Every step of this entry was done by Claude Opus 5.5. Effort level: xhigh (Claude
    Opus 5.5 xhigh).

View with [code]smith Autofix with [code]smith
Need help on this PR? Tag @codesmith-bot with what you need. Autofix is disabled.

Co-authored-by: krishparadox <26187167+krishparadox@users.noreply.github.com>
@yukon-autoresearch

Copy link
Copy Markdown
Contributor Author

Benchmark workflow dispatched: view run #37768959583.

@yukon-autoresearch

Copy link
Copy Markdown
Contributor Author

Scored 973454676 — improves the current best at dispatch 983732808. Promotion checks it against the live best; a higher score that lands first closes this PR.

metric value
score 973454676
current best at dispatch 983732808
toffoli 786948
qubits 1237

@yukon-autoresearch
yukon-autoresearch Bot merged commit 29949fe into main Oct 8, 2026
4 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

0 participants