Skip to content

Release 4.5.0 - #325

Merged
admin-sbneto merged 27 commits into
masterfrom
develop
Sep 15, 2026
Merged

admin-sbneto merged 27 commits into
masterfrom
develop

Conversation

@vhmartinezm

Copy link
Copy Markdown
Contributor

Release 4.5.0

ruleset_list(sort=) and ruleset_list(exclude_favorites=) — the SDK half of the hunt page's active-first ruleset order. Both are opt-in: omit them and every call is byte-for-byte what it was.

Written on the async method, which is canonical, and regenerated into the sync mirror.

Why this release is on the critical path

polyswarm-cli raises its floor to polyswarm_api>=4.5.0 for these keywords. Until 4.5.0 is on PyPI that floor is unsatisfiable for anyone installing from the index — the CLI's own CI resolves it from the branch archive, real users cannot. This release has to land before the CLI's.

What callers need to know

Two properties of the new order, both server-side facts rather than SDK behaviour, documented on the method and in specs/03-endpoints.md:

  • The rank is the stored hunt link, which is wider than what livescan_id renders from. A legacy row whose hunt was stopped without clearing the link leads the list while serializing a null id. Read the field, never the position.
  • The key is mutable, unlike the id-desc default: a ruleset whose hunt stops mid-walk is yielded twice, one that starts mid-walk is skipped for the rest of that walk — in any walk, including a fresh one. The generator streams pages and does not dedupe; callers consuming more than one page dedupe by id. That id is unique but unordered (the server orders on its own insertion key), so it is safe to dedupe with and useless to resume a walk with.

Deploy steps

  1. Merge, tag and publish 4.5.0 to PyPI.
  2. Then the CLI release.

kyle-buchmiller and others added 27 commits August 21, 2026 14:22
A hunt result says which rule matched and its tags, never why. The server
now returns the yara strings behind a hit; parse them onto LiveHuntResult
and HistoricalHuntResult, and so onto their list subclasses.

Read with .get() rather than a subscript. The key is additive, so a server
predating it omits it entirely and a subscript would raise on every result.

Three states reach callers and they are not interchangeable: None (not
reported -- an older server, removed evidence, or a list endpoint, which
omits it rather than fetch a blob per row), [] (the rule matched with no
byte evidence to show), and a populated list (the evidence, as a lower
bound rather than a match count).
Pure-unit: the resources parse a dict, so no HTTP and no stack is involved.

Covers all four affected classes -- the list subclasses inherit __init__ and
must behave identically -- and asserts the states stay distinguishable: an
absent key and an explicit null both read as None, while an empty list stays
an empty list rather than collapsing into it. Entries pass through verbatim,
since only yara knows whether `data` is a hex dump or escaped text.
Records what each state means, the per-entry dict shape, and the two
properties consumers get wrong: a populated list is a lower bound rather
than a match count, and `truncated` means "there was more than this" rather
than an exact size.

Also states that the evidence rides on the detail routes only -- the
paginated feed methods will always yield None -- so a caller looking for
strings knows to fetch a single result.
Comment text only -- the parsed AST is identical before and after. Both copies
edited identically, since the two resource classes carry the same note.

Keeps what the comment exists for: why .get() and not a subscript, and that
the three states are distinct.
The suite was pure-unit only, so nothing asserted that the detail route
actually emits the key or that the list route omits it. Those tests exercise
dict.get and would pass identically if the server never grew the field and
the attribute were a permanent None -- which is the gap the e2e-first
invariant exists to close: a fabricated response asserts what we think the
server returns, a cassette asserts what it actually returned.

test_live and test_async_live already poll the detail route on a rule that
matched their own artifact, so the assertions cost two lines each and pin the
list-vs-detail split that was previously stated only in prose.

Cassettes re-recorded delete-driven against a live stack running the matching
server and analyzer branches. Both now carry real evidence -- the per-test
rule keys on the test's uid, so the recorded hit is `$u` at offset 69 with
the uid's own length -- and null on the list rows.

Verified the assertions are load-bearing: against the previous cassettes they
fail with "detail route should carry the yara evidence / assert None".
The resources spec is what to read before changing a resource's parsing, and
it carries the two closest precedents -- known_good / known_good_sources and
state are both documented there as additive, .get()-parsed fields in exactly
this situation. matched_strings had no note, so a reader following the spec
index would not find it.

Points at the three-state table in the downstream-contract spec rather than
restating it, and records the part a parser needs up front: which value means
"not reported" versus "matched with no evidence", and that the list endpoints
always yield the former by design.
`matched_strings_dropped` on both hunt-result resources, inherited by their list
subclasses. The server bounds how many matched-string bytes one result may carry,
and a truncated list is otherwise indistinguishable from a complete one: a caller
reading 75 entries concludes the rule hit 75 times when it hit 400.

A sibling attribute rather than a key inside `matched_strings`, which stays a
plain list -- so this is additive to the shape already documented, not a change
to it. Parsed with .get() like the field beside it: None means nothing was
withheld, which is also what a server predating the bound reports, so callers
need no special case for the older shape. It can never accompany an empty list,
since a match's first string is never withheld.

Also corrects the downstream contract, which told consumers fast-scan reports
only the first offset per string and that this was one reason the list is a
lower bound. Measured against yara 4.5.2 that is false -- fast mode collapses
repeats of a single string only when the rule's condition does not need them,
and never limits how many distinct strings a rule reports. The remaining reasons
are sound and the byte bound is now a fourth.

195 tests pass.
…e wire shape

The cassettes were recorded before this attribute existed -- two commits before
-- so neither carried the key. The pure-unit tests exercise dict.get and would
have passed identically if the server never emitted it, which is the exact gap
the earlier commit here argued the e2e-first invariant exists to close. Applied
unevenly is worse than not applied.

Re-recorded delete-driven against a stack running the matching server branch.
The assertion reads `'matched_strings_dropped' in result.json` rather than
checking the attribute for None, because `is None` cannot distinguish a served
null from an absent key -- and what needs pinning is that the server SENDS the
field. Its value is null there: the per-test rule is small and withholds
nothing, so the null arm is what a passing hunt actually looks like.

Also corrects both specs, which said list endpoints OMIT the key. They send an
explicit null; omitted is the older-server case. That wording survived from a
design that was reverted, and the recorded cassettes disagreed with it. And in
the new section, "None means nothing was withheld" was stated flatly while the
section above it is careful that None on matched_strings means "we don't know" --
the same conflation, now reading the same way in both places.

195 tests pass.
… their claims

The previous commit said the withheld-count conflation now "reads the same way
in both places". It did not: the flat claim -- None means nothing was withheld --
was corrected in one spec and left standing in specs/02-resources.md and in both
copies of the resources.py comment. specs/02 contradicted itself four lines
apart, telling a reader None means nothing was withheld and then that the list
endpoints always send None by design, so for a list row it asserted both
"nothing withheld" and "we did not look". All three now carry the same reading.

Two tests did not test what they were named for:

- test_dropped_is_independent_of_the_strings_list passed a count and never read
  it -- removing the kwarg left it passing identically, duplicating the test
  above it. Renamed for the property it actually pins and now asserts the count,
  so the populated-list-plus-count pairing is covered.
- test_raw_json_still_carries_the_key compared a dict to itself: __init__ does
  self.json = content, so the assertion could not fail. Replaced with the
  property that could -- parsing must not drop keys from .json, and the parsed
  attributes must agree with the raw payload.

Also records what is and is not pinned against a real server. The live pair is
verified end to end; the historical pair follows by SYMMETRY, because the e2e
stack does not reliably populate historical results inside a test window. Both
specs stated the contract for "all four classes" as established fact, which is
the same "asserts what we think the server returns" problem the e2e-first
invariant exists to prevent. Softened, with the gap recorded in
99-open-questions.md alongside what would close it.

specs/04-testing.md's module inventory gains this module, and known_good_test.py
which was already missing.

195 tests pass.
…oute framing

The per-entry dict shape is contract -- specs/05 documents offset, identifier,
length, data and truncated -- but the only assertion on it compared against a
hand-written fixture, which is what we THINK the server sends rather than what it
does. A server-side key rename would have passed the whole suite VCR-off. Both
cassettes already carry a real entry, so the live tests now assert the key set
against the recorded response, which costs nothing and closes the gap.

Separately, the spec framed the `…List` subclasses as the list-route parsers.
That is not true of the live pair: live_feed builds its request with
LiveHuntResult.list(...), which hits /hunt/live/list but parses rows as
LiveHuntResult. LiveHuntResultList is only ever a delete builder and is never
instantiated from a response; only historical_results yields …List instances.

The framing matters rather than being pedantry: the section exists to say that
None is ambiguous, and a reader who takes the class as the route concludes a
LiveHuntResult must have come from the detail route and reads its None as
"nothing to show". Corrected in the spec and in the test module's comment.

195 tests pass.
The previous commit corrected the class-to-route framing and got the
replacement wrong. It said LiveHuntResultList "is never instantiated from a
response; only historical_results() yields ...List instances". Both halves are
false: _build_request sets result_parser=cls, so the paginated body returned by
DELETE /hunt/live/list is parsed through LiveHuntResultList. The cassette
re-recorded in this PR contains exactly that exchange.

That makes delete responses a FOURTH source of None -- alongside an older
server, a list route and removed evidence -- and none of the three places that
enumerate the causes listed it. A caller iterating live_feed_delete() holds
...List objects with both fields None while the spec told them that class only
comes from historical_results(), which is the worst combination: an ambiguous
value plus a mental model that resolves it wrongly.

Corrected in specs/02, the three-state table in specs/05, and both copies of
the resources.py comment.

195 tests pass.
The previous commit replaced a tautological assertion and moved the tautology
rather than removing it: `assert result.json is content` followed by
`set(content) <= set(result.json)` reduces to `set(content) <= set(content)`.
Nothing checked the property the test is named for either, since self.json keeps
a live reference to the dict that was passed in.

Now compared against an independent deepcopy, so "parsing did not mutate the
payload" and "parsing did not drop keys" can both actually fail. It also survives
a future change that defensively copies content, which the identity assertion
would have failed for the wrong reason.

Also records that the non-null matched_strings_dropped path is not pinned against
a live server. The live tests assert only the None arm -- correctly, since the
per-test rule withholds nothing -- so the two strongest claims the contract makes
about the field rest on hand-written dicts. Producing an over-budget match on the
e2e stack is disproportionate to what it would pin, so this is recorded in
99-open-questions.md rather than tested, alongside the historical-pair gap.

195 tests pass.
…icated comments

Two assertions could not fail:

- The list-row check read `my_results[0].matched_strings is None`, two lines above
  a comment arguing that `is None` cannot distinguish a served null from an absent
  key -- which is why the count below it reads .json. The same argument applies
  here and specs/05 makes the stronger claim (an explicit null), so it now reads
  .json too.
- test_populated_list_is_passed_through_verbatim passed the module-level _STRINGS
  into the content dict, which stores it BY REFERENCE, so the assertions compared
  the object with itself and no in-place reshape could fail them. Deep-copied, the
  same fix test_parsing_does_not_mutate_the_raw_json already needed.

Also trims the two resources.py comment blocks, which restated the contract spec
nearly verbatim at both call sites. Three copies of one argument is three places
to drift, and several commits in this PR were spent reconciling exactly that. The
decision and a pointer are enough; specs/05 owns the reasoning.

195 tests pass.
02-resources.md states that None reaches a caller from four places, not three
-- and then two summaries in that same file enumerate three, as does the
matched_strings_dropped paragraph in 05. The file contradicted itself within
one section, which is the defect an earlier commit was written to fix.

Rather than restate the list a fourth time, the summaries now defer to the
one table that owns it. A pointer cannot drift out of step with what it
points at.
The live assertions are the only thing distinguishing served evidence from a
served null -- the pure-unit tier exercises dict.get and passes identically
whether or not the server ever grew the field. So the assertion stays exactly
as strict.

Only the message widens. The value is null unless the producer that emits the
strings is deployed, so the likeliest cause of this failing is an image that
predates it -- while everything else about the stack looks healthy, which
makes the failure read as a server bug. The message now names the real cause.
Expose matched strings on hunt-result resources
…r-side

`ruleset_list(sort='active_first')` asks the server for rulesets with a running
live hunt first — as recorded by the server's live-hunt link, the same one
`livescan_id` renders from — newest first within each block
(`GET /v3/hunt/rule/list?sort=active_first`). Unset sends no `sort`, so the
request stays byte-compatible with the pre-sort contract and the list keeps its
newest-first default. The SDK never re-orders rows: the list is
keyset-paginated, so a client-side sort would reorder one page and misrepresent
the rest; a page's `offset` is only valid under the same `sort`, and the server
refuses a cursor minted under the other order.

Canonical change in aio/api.py; api.py is the regenerated unasync mirror
(scripts/regenerate_sync.py, ruff on PATH). YaraRuleset.list already forwards
arbitrary keywords through core._params, so the resource needs no change.

Tests on two tiers. Pure-unit pins the wire shape on both transports, the
omitted default, composition with the filters, and that `_next_page` carries
`sort` onto page 2. Live-e2e (sync + async, cassettes recorded against a stack
running the server branch) pins what no builder test can: the server actually
applies the order — this test's running ruleset precedes its newer idle one
under `sort='active_first'` and follows it under the default — and refuses an
unknown sort rather than ignoring it.
The CLI client adopts `ruleset_list(sort=)` in the paired change set and
expresses that as `polyswarm_api>=4.5.0` rather than probing the installed SDK
(the workspace's cross-repo dependency standard; AGENTS.md's standing
exception). A floor cannot name a version this repo has not declared, so the
bump lands here, in the feature PR, not at the release step.

Minor, not major: one new optional keyword with a default that preserves
today's behaviour. Bumped with bump-my-version; the emitted string is a clean
`4.5.0` (a `.devN` form would sort below the floor and send the CLI's CI to
PyPI for a version that does not exist). Order is forced as before: this repo
releases before the CLI can.
…pe by id

The server assigns that obligation to clients and pins it with boundary
tests; it appeared nowhere on this side. Documented on the async method
(the canonical source), regenerated into the sync mirror, and recorded in
the endpoints table. No behaviour change: dedupe does not belong in the
shared streaming generator, which every list endpoint uses and which must
not grow an unbounded id set for one of them.
…livescan_id renders

Three separate things the sort docstring got wrong or over-promised:

The rank is not "the same link livescan_id renders from". The server orders
on the stored link and serializes the id under a stricter predicate, so a
legacy row whose hunt was stopped without clearing the link leads the list
while rendering a null id. A caller reading the leading block as "running"
— or taking rows until the first null id — reads it backwards. The docstring
now says to read the field and never the position.

"A fresh walk from the first page is always self-consistent" was wrong in
the same breath as the duplicate it describes: the mutable key repeats and
skips rows WITHIN a walk, so starting fresh does not avoid it. Retracted.

And the live ordering test's poll called list.index() on a row a lagging
replica may not have returned yet. poll_equals absorbs NotFound/NoResults,
not ValueError, so the lag the poll exists for would have errored the test
on its first attempt instead of retrying. Both twins now read as "not yet".
The sort paragraph called the unsorted list "the id-desc default", which
reads as a promise about the id callers can see. It is not one: the server
orders on its own insertion key and renders a random 17-digit number as id,
so a caller who recorded the smallest id yielded and resumed below it would
silently skip or repeat rows. The dedupe advice in the same paragraph stands
— that only needs uniqueness — and now says so explicitly.
…uilder test

Two leftovers from the ordering correction, both caught in review.

The default-order assertion was the one unpolled call of a helper that now
returns None while either row is missing — deliberate, so the poll can retry.
Unpolled, a lagging read replica turns that into 'None is False' instead of a
retry, against the invariant that these pass on the live stack with VCR off.
It is polled now, with want=False, which the helper's own guard accepts.

And the builder test's class docstring still called the unsorted list id-desc,
the exact claim the rest of this branch exists to correct — the cassette shows
the default page returning the lower visible id first.
…es separately

The server gained the inverse of favorites_only: a paginated list with the
favorites taken out. It exists because the favorites are a separate, unpaginated
fetch bounded by the account's budget, so leaving them in the page too makes a
client either render a row twice or render a short page.

Appended to the signature rather than placed beside favorites_only, so a caller
passing the later filters positionally keeps working.
…till untested

The only coverage the parameter shipped with called the generic resource
builder, which this change does not touch — dropping the keyword from both
transports left the whole suite green. The new cases drive the sync and async
client methods, and deleting the pass-through now fails exactly one test.

The e2e arm specs/04 asks for is still missing, and the comment at the live
sort test says so plainly, along with what compensates and what does not:
a rename made in lockstep with the server's spelling is the case only a live
request catches. Recording it needs a stack whose key-management service
carries the fixture account.
feat: ruleset_list(sort=) for the server's active-first order (4.5.0)
@claude

claude Bot commented Sep 15, 2026

Copy link
Copy Markdown

Review — release 4.5.0

Gitflow and the bump are correct: develop → master, 4.4.0 → 4.5.0 consistent across pyproject.toml version, [tool.bumpversion] current_version, and __init__.py __version__, and the emitted string is a clean X.Y.0 (AGENTS.md §Gitflow's standing-exception check — the CLI floor >=4.5.0 is satisfiable).

Correctness looks clean. Things I specifically checked and found good:

  • sort / exclude_favorites ride the query via the generic core._params path; exclude_favorites coerces to the int-bool the server parses, same as favorites_only. Unset → omitted, so the no-filter request stays byte-compatible (test_list_omits_every_unset_filter).
  • sort survives onto page 2 — _next_page clones request.params wholesale, and test_sort_survives_onto_the_next_page drives that directly rather than through the stubbed _paginate. That's the one place a rewrite could silently drop it, so good that it's pinned.
  • Both new cassettes are genuinely recorded, not copied between the sync/async twins (distinct ids, timestamps, rule bodies, content-lengths), and both carry all three list variants — sort=active_first, unsorted, and sort=bogus → 400 → RequestException via _raise_for_status's else arm.
  • poll_equals(_running_precedes_idle, False) is False is the right idiom for a falsy want, and the membership-tolerant return None keeps replica lag out of .index().
  • Re-recorded test_live / test_async_live cassettes back every new assertion: detail route carries matched_strings with exactly {offset, identifier, length, data, truncated}, list rows carry an explicit null, and matched_strings_dropped is present-and-null.

Spec drift

1. specs/05-downstream-contract.md overclaims what is pinned for the live pair.

The live pair (LiveHuntResult / LiveHuntResultList) is verified end to end — test_live / test_async_live assert the detail route carries evidence, that list rows do not, and that the server serves matched_strings_dropped.

All three enumerated assertions read a LiveHuntResult. Nothing asserts the …List half: client_scan_test.py:535 and async_client_test.py:933 call live_feed_delete([result_id]) without binding the return value, so the delete response's matched_strings is None — which specs/02 calls out as the fourth, most surprising cause of None — is asserted only by hand-written dicts in hunt_matched_strings_test.py. Given this PR files two open-questions entries for exactly this class of gap, either narrow the sentence to LiveHuntResult or assert on the live_feed_delete return.

2. The exclude_favorites e2e gap lives only in a code comment. The long comment above test_rules_sort_active_first (client_scan_test.py) is a careful statement of an invariant-1 gap — the server ignores unknown query args, so a lockstep rename stays green — but it's the one gap in this PR that didn't get a specs/99-open-questions.md entry. Move it there alongside the two that did; a comment on an unrelated test is not where the next contributor will look.

3. specs/03-endpoints.md versions only half the new surface. sort='active_first' is labelled (4.5.0); exclude_favorites lands in the same release and carries no version marker, even though the CLI floor pins both keywords to it.

PR hygiene

4. The description omits the larger half of the release. The body covers only ruleset_list(sort=) / (exclude_favorites=), but the diff also adds matched_strings and matched_strings_dropped as public attributes on four resources classes (LiveHuntResult, LiveHuntResultList, HistoricalHuntResult, HistoricalHuntResultList) plus three new spec sections. That is additive-and-minor, so it doesn't change the bump — but this is the release PR, and its description is the closest thing PyPI consumers get to release notes.

5. Ticket IDs will land on master. Two merge commits in this PR carry branch names in their subjects:

  • Merge pull request #322 from polyswarm/DN-8378-yara-matched-strings
  • Merge pull request #324 from polyswarm/DN-8445-ruleset-list-active-first

AGENTS.md §Commit + PR hygiene: "Don't reference ticket IDs or internal project codes in commit messages… This repo is public; published artefacts shouldn't leak internal references." Squash this merge so the codes don't reach published history, and drop the DN- prefix from feature branch names going forward.

@admin-sbneto
admin-sbneto merged commit 38cd189 into master Sep 15, 2026
3 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Development

Successfully merging this pull request may close these issues.

3 participants