Skip to content

Release 4.4.0 - #323

Merged
admin-sbneto merged 38 commits into
masterfrom
develop
Sep 8, 2026
Merged

admin-sbneto merged 38 commits into
masterfrom
develop

Conversation

@vhmartinezm

Copy link
Copy Markdown
Contributor

TL;DR — changelog for 4.4.0

  • Ruleset resources carry the hunt-page tracking fields: favorite, favorited_at, rule_count, historical_hunt_count, new_results_count, new_results_counted_at; hunt resources carry the frozen-rule provenance (source_rule, source_rule_changed).
  • New ruleset_favorite(id, favorite) on both clients; the server's FAVORITE_LIMIT refusal surfaces as a typed error with favorites_used / favorites_limit.
  • ruleset_list accepts name, status, favorites_only, has_new_results (conjunctive) and asserts them universally.
  • live_feed gains livescan_id scoping and an optional max_results total bound (0 = unbounded); since stays in seconds.
  • Internal result-bound helper made private; e2e stack hostnames renamed to the canonical short names (ai, sandbox) with cassettes re-recorded (test: e2e stack hostname is now 'ai' (was artifact-index-e2e) #320).
  • Version is already 4.4.0 on develop (bumped in Hunt-page ruleset tracking: favorites, rule counts, hunt provenance #321 so the CLI could pin the floor) — this PR publishes it. PyPI upload fires on the version change on master.

Requires

  • artifact-index#1963 deployed to prod first — the read surface these fields come from. Until then, published 4.4.0 renders null for the new fields against prod (additive, no failure).
  • polyswarm-cli Release 4.4.0 merges after this is on PyPI (it pins polyswarm_api>=4.4.0).

sbneto and others added 30 commits August 6, 2026 18:10
The e2e docker-compose stack renamed its services to canonical short names;
the artifact-index API service there is now reachable as 'ai'. The suite's
hardcoded live-stack URIs and all VCR cassettes follow (VCR matches on host,
so the cassettes must carry the same hostname the tests dial; this is a
mechanical hostname substitution — method/path/query/bodies untouched).
…service-e2e)

Companion to the 'ai' rename: SANDBOX_SERVICE_URI in both suite files plus the
sandbox-bearing cassettes (byte-safe replace — one cassette carries binary
recording bodies). Suite verified locally VCR-on: 163 passed.
…nstants

The previous commit's replace loop aborted on a binary-bodied cassette midway
through the file list, leaving the two SANDBOX_SERVICE_URI constants on the old
hostname while the cassettes had moved — suite verified VCR-on after this fix:
163 passed.
test: e2e stack hostname is now 'ai' (was artifact-index-e2e)
…nance

New fields parsed on existing resources (all additive; an older server
leaves them None):

- YaraRuleset: favorite, favorited_at, rule_count (None means the server
  had no answer, distinct from 0), historical_hunt_count, and
  new_results_count (only when the list is asked to include counts).
- HistoricalHunt: rule_id (the source ruleset), rule_modified (freeze-time
  audit value), and source_rule_changed — a tri-state answering "has the
  source ruleset's body changed since the hunt froze it?" (None = unknown,
  not 'unchanged').

New endpoints and filters:

- ruleset_favorite(id, favorite): idempotent star/unstar; the response
  carries favorites_used/favorites_limit; over-budget refusals surface a
  machine-readable FAVORITE_LIMIT error.
- ruleset_list(name=, status=, favorites_only=, has_new_results=, since=,
  include_counts=): the hunt-page filters, conjunctive and optional.
- live_results_count(since=): per-live-hunt result counts in a window,
  one aggregate for every 'new results' badge.
- live_feed(livescan_id=): scope the feed to one live hunt.

Sync and asyncio clients both. The rules live-suite tests now create a
uid-namespaced single-rule ruleset (deterministic rule_count, no name
collisions on the shared stack) and exercise the favorite round-trip,
provenance, counter increment, and the changed-since-freeze flip; their
cassettes are removed to re-record against a stack that serves the new
fields.
The sync api.py was hand-edited; scripts/regenerate_sync.py places
live_results_count in aio's order and applies ruff's formatting, which
is what the unasync-mirror CI gate diffs against.

The rules live-tests' three read-after-write assertions (the counter,
and both sides of the changed-since-freeze flip) now poll: those GETs
read the replica, and on a real-replica stack the stale read of the
flip is a silent False. Sleeps are free on VCR replay.
Ids exceed JavaScript's safe-integer range; the counts entries carry the
same digit string YaraRuleset.livescan_id does.
Recorded against the branch server image (both tests green live first);
the offline suite replays them — 163 passed with no stack.
Review findings, all four:

- specs updated in the same PR as required: 03-endpoints gains
  ruleset_favorite / live_results_count rows, the real ruleset_list
  signature and live_feed's livescan_id; 02-resources catalogues both
  new resources — including why YaraRulesetFavorite empties
  RESOURCE_ID_KEYS (the server reads the toggle from the PUT body; the
  empty key list is the only thing routing id there) and the bool→int
  body serialisation; 05's commonly-imported list carries both.
- the rules live-tests now exercise every previously-uncovered surface
  against the real stack (and the cassettes record it): live_start →
  status=active filter → include_counts observed as a computed 0
  (distinct from null) → live_results_count (our zero-result hunt
  ABSENT from counts, keyed by the same digit strings ruleset_get
  renders) → the livescan_id-scoped feed → live_stop, with the stop in
  a finally because a running hunt blocks ruleset deletion.
- pure-unit builder tests (hunt_tracking_builder_test.py, the
  known_good_test pattern) pin the request shapes: the favorite PUT's
  body routing incl. 1/0 bools, counts query routing + None omission,
  the list filters' int bools and byte-compatible no-filter request,
  and livescan_id stringification.
- the unstar stays a contract assertion with slot hygiene documented:
  ruleset_delete soft-deletes and the budget counts only deleted=false
  rows, so a failed run's star frees itself with the rule. The limit
  pin vs used bound is now commented as deliberate.

Also: test/eicar.yara deleted (no test references it since the uid_yara
move; the helper docstring no longer names the file).
…ew 2)

All five follow-ups:

- live_results_count moved to the _single Live-hunts table in
  specs/03 — it returns one resource, and _single-vs-_paginate is that
  document's organizing invariant.
- specs/04's fixture inventory drops the retired test/eicar.yara and
  names the new pure-unit module.
- Parse-side pins for the counts resource: the cassettes only carry
  EMPTY counts (fresh zero-result hunt), so the {livescan_id, count}
  entry shape, the digit-string join key and the null-counts coalesce
  now have canned-payload tests.
- The livescan_id feed assertions no longer read as if they verify the
  scoping: with a zero-result hunt they pin the wire shape and the
  empty pass-through only, and the comments now say so (the scoping
  semantics are pinned by the server's own HTTP suite).
- FAVORITE_LIMIT's machine-readable contract is now documented
  (specs/05: no typed exception by design; the path is
  exc.request.errors with the code plus the same counters a successful
  toggle returns) and pinned by a respx refusal test — mocked because a
  genuinely full budget on the shared stack would race every other run.
The docstring said SECONDS and claimed the server was tightening the parameter
to `is not None` so that since=0 would mean an empty window. No such change
was ever made, and the opposite is the contract: the server applies the filter
on a truthiness test, so absent-or-0 means no time filter at all and the feed
pages over everything. The unit is now minutes on the wire, which is a break
worth naming — a caller passing an explicit since in seconds gets a 60x wider
window, silently — so specs/05 records it as a behaviour change rather than a
documentation correction, for the develop -> master bump decision.

max_results is additive and defaults to None, the historical behaviour: every
page, up to the client's page cap. It bounds the yielded count and sizes the
page request via core.page_size_for, so a small ask does not fetch a full
default page. The bound is client-side by nature — the server has never served
an unbounded query, it caps every page; this client is what follows cursors
until has_more clears.

page_size_for lives in core.py: pure, shared by both transports, and not
unasync-processed, so there is exactly one copy. Sync mirror regenerated.
It said live_start never gets a livescan_id on the local e2e stack. The rules
tests now assert the opposite as a hard contract and test/vcr/test_rules.vcr
records a real one. What is still true is that no microengines process
submissions, so the feed stays empty — which is what the zero-result feed check
in those tests rests on.
`rule.id in by_name` and `any(... and r.favorite ...)` both hold if the
server ignored the query param and returned the unfiltered list — which is the
regression an SDK-side filter test exists to catch, and the one most likely
when the paired server change lands. Both arms now assert over every row. No
re-record: the recorded responses already satisfy the stronger form.
Truncation and page-size selection are both decisions the code now makes and
neither had a test. Pinned in the pure-unit tier: producing more feed rows than
a page holds on the shared e2e stack would mean generating real live-hunt
volume, and _paginate is the seam that would otherwise keep following cursors.
The truncation tested `is not None` while page_size_for treats 0 as no bound,
so max_results=0 asked for a server-default page and then stopped after the
first result. 0 now means no bound on both halves, which is also what `since`
means by 0 on the same call.
max_results read 0 two contradictory ways: page_size_for treated falsy as
unbounded while the generator tested `is not None`, so max_results=0 asked for
a full server-default page and then stopped after one row. Negatives were worse
— page_size_for returned -1, which would have put limit=-1 on the wire and the
server answers that with nothing.

as_result_bound is now the single definition (None/0/negative -> no bound) and
both halves read it, so they cannot disagree again. The bound and the page it
implies stay separate numbers on purpose: a caller asking for 5000 gets 1000-row
pages and still stops at 5000.

Also pins the wiring itself. The truncation tests replace _paginate wholesale,
so the descriptor was never built and dropping `limit=page_size_for(...)`
entirely kept every test green; TestLiveFeedLimitOnTheWire fails without it
(verified), and asserts the unbounded case sends no limit at all, which is what
keeps the default request byte-compatible with the recorded cassettes.

specs/01, specs/03 and specs/05 pick up the new core symbols, the MINUTES unit
and the absent-or-0 contract — specs/03 still documented SECONDS, contradicting
the docstrings and specs/05 inside the same change.
…r arm

'rules' is a prefix of 'ruleset', so the long-pole fragment also matched the
instant ruleset_favorite respx suite and scheduled it ahead of real long poles.
'test_rules' matches only the two intended nodeids.

Nothing covered an explicit favorites_only=False / has_new_results=False: they
serialise to query 0 (core._params coerces bools before routing), and an
inverted filter is the one failure mode that silently returns the wrong rows.

Also records where MAX_PAGE_SIZE comes from — the server's AI_MAX_QUERY_RESULTS
— and why it sits in core.py rather than settings.py.
The long-pole fragment went from 'rules' to 'test_rules' to fix a false positive
and introduced a false negative: 'test_async_rules' is test_ + async_rules, so
the heaviest new test in the suite fell to the backfill tail. '_rules' matches
both nodeids and still misses the respx suite, where the character before
'rules' is a slash. The replaced line also left its old trailing comment behind.

Every max_results test drove the GENERATED sync mirror. The async for +
yielded/return shape is the part unasync rewrites rather than copies, so the
canonical loop had no coverage at all — verified the new tests fail when the
canonical bound check alone is broken.
The wire is not moving to minutes (prod traffic makes it a silent 60x widening
for clients we don't control), so this is a documentation correction again
rather than a behaviour change — which also removes it from the develop -> master
bump decision. max_results and the absent-or-0 contract are unaffected.

Also shortens the core helper comments: the argument belongs in specs/05, not
beside every line that touches it.
MAX_PAGE_SIZE mirrored the server's AI_MAX_QUERY_RESULTS code default of 1000,
but the chart sets 300 in every environment — so live_feed(max_results=500)
would have sent limit=500 and got a 400. The cap is an env var the deployment
chooses, so the SDK cannot know it: max_results now bounds only how many
results the generator yields, and the request is unchanged (which also keeps
the default call byte-compatible with every cassette).

Also polls the favorite read-after-write — the star is written on the line
above, and it was the one such assertion here that neither polled nor guarded
204. One read per attempt, so the recorded interactions are unchanged.
The two are independent: the page stays the server's to choose (50 for web,
capped by AI_MAX_QUERY_RESULTS) and _next_page echoes it, so a bounded read
keeps paginating in those same small chunks and stops once it has enough.
Sending limit=max_results conflated them — the 400 above the deployment's cap
was a symptom of that, not the reason.
The two-line read-after-write note was pasted twice in a row, in both the
sync and async live tests.
The Versioning table had no row for adding an optional keyword to an
existing public method, so the nearest match was `Signature change on a
public method | major` — which scores this change major even though every
existing call site keeps working untouched.

The table already carves out the additive exception cases on exactly that
reasoning ("no consumer has to change"); this applies the same rule to
keyword arguments, and narrows the signature row to the changes a caller
must actually react to.
Three accuracy fixes from an audit of the revert commits:

- The `since` docstring said the parameter "said minutes before 4.4".
  The version is chosen at the `develop -> master` step, not here, so if
  that release cuts as anything else the shipped docstring is wrong and
  nothing would catch it. "in earlier releases" carries the same meaning
  with no forward reference.
- `specs/05` described the CLI's `1440` default in the present tense
  inside a paragraph that is otherwise entirely historical. That default
  was retired in this same change set.
- `core.py`'s module docstring lists the pure helpers; `as_result_bound`
  was added to `specs/01` but never to the list beside the code.
The CLI expresses its need for this change set's surfaces as a version
requirement — `polyswarm_api>=4.4.0` — rather than probing the installed SDK
at runtime. A pin is checkable by pip at install time, before any code runs,
and it cannot name a version this repo has not declared. So the bump lands
here, in the feature PR, instead of at the release step.

Minor, not major: every addition is additive — new methods, a new resource,
new optional keywords, new parsed fields — so no existing call site changes.

Bumped with bump-my-version and verified the emitted string is a clean
`4.4.0`: the serialize config can produce `4.4.0.devN+sha`, and PEP 440 orders
that BELOW 4.4.0, which would silently fail the CLI's floor and send its CI to
PyPI for a version that does not exist yet.

AGENTS.md and specs/05 record the exception and the ordering it forces: this
repo must release before the CLI can, since the CLI's floor is unsatisfiable
from PyPI until then. Consumer CI is unaffected — it installs this repo from
git by branch name.
sbneto and others added 8 commits August 31, 2026 13:35
Both AGENTS.md and the specs/02 inline comment framed the attribute as "set
this when the identifier isn't `id`". `YaraRulesetFavorite` is the first
resource where the key is not the identifier at all: it names `community` so
that `id` stays in the PUT body, because the default would move `id` to the
query string and the server 400s. specs/02's prose and its per-resource entry
already describe it correctly; the orientation doc and the code comment were
the two places still under-describing the mechanism.
`favorites_limit == 5` mirrors a code default the deployment can override —
`AI_FAVORITE_RULESETS_LIMIT` is read from the environment — which is the same
mistake as the MAX_PAGE_SIZE constant this branch already removed: the client
asserting a server-owned number it cannot see. Against a live stack with VCR
off and a different cap configured, the exact pin fails on a correct server.

Assert the relation instead; the used-vs-limit check beside it still carries
the meaning.
Six create-response assertions sat above the `try:` whose `finally:` deletes
the ruleset, so any of them failing leaks one on the shared e2e stack. Two of
them predate this branch; the four tracking-field assertions widened it. The
cost is not abstract: a leaked STARRED ruleset holds one of the team's five
favorite slots, and the favorite tests need a free one. `try` now opens
immediately after `ruleset_create`, in both transports.

specs/04 gains the `poll_equals` / `poll_equals_async` entry it was missing
while documenting its sibling `run_concurrently`. It carries a non-obvious
invariant worth writing down: `want` must never be None, because not-found
during the lag window also reads as None and would turn a vanished resource
into a passing assertion — poll a boolean instead. Also records the
mirror-image limit, that a False poll cannot ride out a stale False.

AGENTS.md now names where the bump exception comes from: it is this repo's
instance of a workspace-level standard, not a rule the repo grants itself.
Two comments carried an "(F9)" review token. AGENTS.md keeps internal
references out of published artefacts; the rule names commits and PR text, but
a committed source comment outlives both. The sentences read the same without
it.
The rules fragment has now been wrong three times — "rules" caught "ruleset",
"test_rules" missed "test_async_rules", and "_rules" is itself a prefix of
"_ruleset" so it front-loaded the instant unit tests. Every substring of those
two test names is a prefix of test_ruleset_*, so no fourth fragment fixes it.

Give the matcher a way to say "exactly this test" instead: a fragment starting
with "::" is matched as a nodeid SUFFIX. The two long poles now name themselves
exactly; every other fragment keeps substring matching, which is right for the
ones that legitimately span several tests.

Verified: test_rules and test_async_rules rank as long poles, the unit test
backfills to the tail.
…nternals

`as_result_bound` was listed on the documented public surface, which put a
four-line internal helper under "rename or removal of a public symbol = major".
Nothing outside this package calls it — not the paired CLI — and every sibling
helper in core.py is already underscore-private. Renamed to `_as_result_bound`
and dropped from the specs/05 surface listing; specs/01 still records it among
the helpers, where it belongs.

This repo is public, so two other things should not have been in it:

- Production telemetry. specs/05 published request volumes, distinct API-key
  and user-agent counts, and the observed value range for a parameter. The
  decision it supports — the wire stays seconds, because re-reading live
  traffic as minutes widens every window 60x silently — needs none of it.
- A deployment's configured page cap. The mechanism (a per-deployment env var,
  set well below a large bound) is what explains the 400; the number is not.

Also stops the spec arguing with its own drafts. "An earlier revision of this
document claimed…" and "the MAX_PAGE_SIZE mistake again" describe review
history, which git already holds; a spec should read as the current state so a
reader does not have to work out which claim is live.
…om source

The bullet read as if any sibling needing a new surface justifies bumping the
version in a feature PR. It does not. The exception exists because a floor
cannot name a version this repo has not declared — which is only a problem for
a consumer whose CI installs this repo from source by branch and can therefore
adopt before publication.

A consumer that installs published artifacts has no such need: it adopts after
the release, and this repo bumps at its own release step as usual. Stating the
condition keeps the exception from being applied where it does not hold.
Hunt-page ruleset tracking: favorites, rule counts, hunt provenance
@claude

claude Bot commented Sep 8, 2026

Copy link
Copy Markdown

Review — release 4.4.0 (develop → master)

Checked against AGENTS.md and specs/01–05. The SDK surface is clean: YaraRulesetFavorite narrowing RESOURCE_ID_KEYS to [community] matches both the recorded wire (PUT /hunt/rule/favorite?community=gamma with body {"id":"…","favorite":1}) and the specs/02 note; max_results is a pure client-side bound that never reaches the query (pinned on both the canonical async loop and the generated mirror); every new resource field is an additive .get() parse; and 4.3.0 → 4.4.0 as a minor is correct under the newly added bump-table rows, with the bump having landed on develop under the standing downstream-floor exception.

Three things to fix.

1. specs/04-testing.md:293 still tells you to dial the old hostname (spec drift)

The local-pinned-stack loop reads:

-e publishes host ports — required for running the suite from the host (the tests target artifact-index-e2e:9696, which resolves to 127.0.0.1 via /etc/hosts …)

After this PR the suite targets ai:9696 (and sandbox:54110). Anyone following the documented loop adds the wrong /etc/hosts entry and gets connection failures. This is the spec for the very workflow the rename changed, so it belongs in the same PR — per AGENTS.md, "Update the spec in the same PR as the code change."

2. docker/Dockerfile:19-20 — the rename is not finished

The comment there still names artifact-index-e2e:9696 and sandbox-service-e2e:54110. Comment-only, but this is the test image the whole e2e fleet resolves as polyswarm-api:latest, and one of the commits here is titled "finish the sandbox hostname rename". Grepping the repo for the two old names returns exactly this file and specs/04 above — nothing else.

3. PR body names fields that do not exist

The changelog says hunt resources carry source_rule, source_rule_changed. The actual attributes are rule_id, rule_modified, source_rule_changed (src/polyswarm_api/resources.py:839-845) — which is what the tests, the cassettes, and specs/02/03 all use. Since this description is the 4.4.0 release changelog, worth correcting so the two do not disagree.

Non-blocking: specs/05-downstream-contract.md:137 gains a stray blank line inside the core symbol-list fence, and poll_equals sleeps once more after its final failed attempt (~1s x 9 poll sites, live run only).

@admin-sbneto
admin-sbneto merged commit 018d737 into master Sep 8, 2026
3 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Development

Successfully merging this pull request may close these issues.

3 participants