Skip to content

fix: close gaps in agreement cancels, chain reads and pacing - #744

Draft
MoonBoi9001 wants to merge 16 commits into
mb9/reviewfrom
mb9/fix-second-review-findings-on-agreement-cancels
Draft

MoonBoi9001 wants to merge 16 commits into
mb9/reviewfrom
mb9/fix-second-review-findings-on-agreement-cancels

Conversation

@MoonBoi9001

@MoonBoi9001 MoonBoi9001 commented Oct 9, 2026 •

Copy link
Copy Markdown
Member

This fixes 10 problems a second review of #715 found in how dipper ends agreements, reads the chain and paces offers, with 1 commit and, where one fits, 1 test per fix. The 2 changes worth a closer look are that no reassessment now gives an abandoned agreement's slot to a new indexer while it may still be paid, and that the chain listener retries a page of changes that failed to apply for up to 10 minutes.

The registry trait gave 9 event and audit methods a do-nothing default, so a registry
missing one would drop accepts, ends and their announcements without a word.
Only the test stub keeps those defaults now.
The example config's webhook placeholder wasn't a URL, so a copy that kept it stopped dipper
at startup. A test now reads the example's alerts section to keep it valid.
Marking a stale agreement for cancelling had no time limit, so a hung database held up every
other agreement's check. The mark now gives up after the usual database timeout.
A cancel can queue behind another transaction for up to the 7-minute worker job limit, but the
retry waited only 2 minutes, so it could send a second cancel. It now waits 8 minutes.
When an agreement the indexer rejected shows up accepted on-chain but has ended by the time
dipper checks, it stays rejected. Recording its accept announced it with no end to follow.
A failed read of the chain's time skipped every cancel in the sweep, though only offers never
accepted need it, to tell whether their deadline has passed. Now only those wait.
Offer pacing counted only agreements waiting for acceptance, so an offer dipper was withdrawing
didn't count, though the indexer can still accept it until the withdrawal lands or its deadline.
When the database failed while applying a page of on-chain changes, the listener moved past
the page anyway, so a change like reopening a live cancelled agreement was lost for good.
Once confirmed, the newest block dipper had seen took any block up to a week ahead from 1
endpoint, so a faulty one could have every correct endpoint refused as behind for hours.
A comment in the test stub still said its do-nothing event and audit methods copied defaults
from the real registry trait, which no longer has any.
Only the liveness check waited for a stale agreement's cancel before replacing it; any other
reassessment saw the request short and filled the slot, paying 2 indexers for it until then.
The function that did both, with the new database time limit, grew past the project's limit
on how hard a function may be to follow.
Formatting only, so the format check in CI passes.
Shrinking the selection by the held slots made the kept agreements look surplus, so a target
of 1 with 1 held slot cancelled the healthy one. Held slots now only limit new offers.
Noting that a stale agreement had ended on-chain had no time limit, so a hung database could
hold up the liveness check. It now gives up after the usual database timeout.
A change that could never apply, such as a row that can't be read, held the listener's place
for good and stalled every later accept and cancel. It now moves on after 10 minutes.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant