Release 4.4.0 - #323
Release 4.4.0#323
Conversation
The e2e docker-compose stack renamed its services to canonical short names; the artifact-index API service there is now reachable as 'ai'. The suite's hardcoded live-stack URIs and all VCR cassettes follow (VCR matches on host, so the cassettes must carry the same hostname the tests dial; this is a mechanical hostname substitution — method/path/query/bodies untouched).
…service-e2e) Companion to the 'ai' rename: SANDBOX_SERVICE_URI in both suite files plus the sandbox-bearing cassettes (byte-safe replace — one cassette carries binary recording bodies). Suite verified locally VCR-on: 163 passed.
…nstants The previous commit's replace loop aborted on a binary-bodied cassette midway through the file list, leaving the two SANDBOX_SERVICE_URI constants on the old hostname while the cassettes had moved — suite verified VCR-on after this fix: 163 passed.
test: e2e stack hostname is now 'ai' (was artifact-index-e2e)
…nance New fields parsed on existing resources (all additive; an older server leaves them None): - YaraRuleset: favorite, favorited_at, rule_count (None means the server had no answer, distinct from 0), historical_hunt_count, and new_results_count (only when the list is asked to include counts). - HistoricalHunt: rule_id (the source ruleset), rule_modified (freeze-time audit value), and source_rule_changed — a tri-state answering "has the source ruleset's body changed since the hunt froze it?" (None = unknown, not 'unchanged'). New endpoints and filters: - ruleset_favorite(id, favorite): idempotent star/unstar; the response carries favorites_used/favorites_limit; over-budget refusals surface a machine-readable FAVORITE_LIMIT error. - ruleset_list(name=, status=, favorites_only=, has_new_results=, since=, include_counts=): the hunt-page filters, conjunctive and optional. - live_results_count(since=): per-live-hunt result counts in a window, one aggregate for every 'new results' badge. - live_feed(livescan_id=): scope the feed to one live hunt. Sync and asyncio clients both. The rules live-suite tests now create a uid-namespaced single-rule ruleset (deterministic rule_count, no name collisions on the shared stack) and exercise the favorite round-trip, provenance, counter increment, and the changed-since-freeze flip; their cassettes are removed to re-record against a stack that serves the new fields.
The sync api.py was hand-edited; scripts/regenerate_sync.py places live_results_count in aio's order and applies ruff's formatting, which is what the unasync-mirror CI gate diffs against. The rules live-tests' three read-after-write assertions (the counter, and both sides of the changed-since-freeze flip) now poll: those GETs read the replica, and on a real-replica stack the stale read of the flip is a silent False. Sleeps are free on VCR replay.
Ids exceed JavaScript's safe-integer range; the counts entries carry the same digit string YaraRuleset.livescan_id does.
Recorded against the branch server image (both tests green live first); the offline suite replays them — 163 passed with no stack.
Review findings, all four: - specs updated in the same PR as required: 03-endpoints gains ruleset_favorite / live_results_count rows, the real ruleset_list signature and live_feed's livescan_id; 02-resources catalogues both new resources — including why YaraRulesetFavorite empties RESOURCE_ID_KEYS (the server reads the toggle from the PUT body; the empty key list is the only thing routing id there) and the bool→int body serialisation; 05's commonly-imported list carries both. - the rules live-tests now exercise every previously-uncovered surface against the real stack (and the cassettes record it): live_start → status=active filter → include_counts observed as a computed 0 (distinct from null) → live_results_count (our zero-result hunt ABSENT from counts, keyed by the same digit strings ruleset_get renders) → the livescan_id-scoped feed → live_stop, with the stop in a finally because a running hunt blocks ruleset deletion. - pure-unit builder tests (hunt_tracking_builder_test.py, the known_good_test pattern) pin the request shapes: the favorite PUT's body routing incl. 1/0 bools, counts query routing + None omission, the list filters' int bools and byte-compatible no-filter request, and livescan_id stringification. - the unstar stays a contract assertion with slot hygiene documented: ruleset_delete soft-deletes and the budget counts only deleted=false rows, so a failed run's star frees itself with the rule. The limit pin vs used bound is now commented as deliberate. Also: test/eicar.yara deleted (no test references it since the uid_yara move; the helper docstring no longer names the file).
…ew 2)
All five follow-ups:
- live_results_count moved to the _single Live-hunts table in
specs/03 — it returns one resource, and _single-vs-_paginate is that
document's organizing invariant.
- specs/04's fixture inventory drops the retired test/eicar.yara and
names the new pure-unit module.
- Parse-side pins for the counts resource: the cassettes only carry
EMPTY counts (fresh zero-result hunt), so the {livescan_id, count}
entry shape, the digit-string join key and the null-counts coalesce
now have canned-payload tests.
- The livescan_id feed assertions no longer read as if they verify the
scoping: with a zero-result hunt they pin the wire shape and the
empty pass-through only, and the comments now say so (the scoping
semantics are pinned by the server's own HTTP suite).
- FAVORITE_LIMIT's machine-readable contract is now documented
(specs/05: no typed exception by design; the path is
exc.request.errors with the code plus the same counters a successful
toggle returns) and pinned by a respx refusal test — mocked because a
genuinely full budget on the shared stack would race every other run.
…ts died at get_sources)
The docstring said SECONDS and claimed the server was tightening the parameter to `is not None` so that since=0 would mean an empty window. No such change was ever made, and the opposite is the contract: the server applies the filter on a truthiness test, so absent-or-0 means no time filter at all and the feed pages over everything. The unit is now minutes on the wire, which is a break worth naming — a caller passing an explicit since in seconds gets a 60x wider window, silently — so specs/05 records it as a behaviour change rather than a documentation correction, for the develop -> master bump decision. max_results is additive and defaults to None, the historical behaviour: every page, up to the client's page cap. It bounds the yielded count and sizes the page request via core.page_size_for, so a small ask does not fetch a full default page. The bound is client-side by nature — the server has never served an unbounded query, it caps every page; this client is what follows cursors until has_more clears. page_size_for lives in core.py: pure, shared by both transports, and not unasync-processed, so there is exactly one copy. Sync mirror regenerated.
It said live_start never gets a livescan_id on the local e2e stack. The rules tests now assert the opposite as a hard contract and test/vcr/test_rules.vcr records a real one. What is still true is that no microengines process submissions, so the feed stays empty — which is what the zero-result feed check in those tests rests on.
`rule.id in by_name` and `any(... and r.favorite ...)` both hold if the server ignored the query param and returned the unfiltered list — which is the regression an SDK-side filter test exists to catch, and the one most likely when the paired server change lands. Both arms now assert over every row. No re-record: the recorded responses already satisfy the stronger form.
Truncation and page-size selection are both decisions the code now makes and neither had a test. Pinned in the pure-unit tier: producing more feed rows than a page holds on the shared e2e stack would mean generating real live-hunt volume, and _paginate is the seam that would otherwise keep following cursors.
The truncation tested `is not None` while page_size_for treats 0 as no bound, so max_results=0 asked for a server-default page and then stopped after the first result. 0 now means no bound on both halves, which is also what `since` means by 0 on the same call.
max_results read 0 two contradictory ways: page_size_for treated falsy as unbounded while the generator tested `is not None`, so max_results=0 asked for a full server-default page and then stopped after one row. Negatives were worse — page_size_for returned -1, which would have put limit=-1 on the wire and the server answers that with nothing. as_result_bound is now the single definition (None/0/negative -> no bound) and both halves read it, so they cannot disagree again. The bound and the page it implies stay separate numbers on purpose: a caller asking for 5000 gets 1000-row pages and still stops at 5000. Also pins the wiring itself. The truncation tests replace _paginate wholesale, so the descriptor was never built and dropping `limit=page_size_for(...)` entirely kept every test green; TestLiveFeedLimitOnTheWire fails without it (verified), and asserts the unbounded case sends no limit at all, which is what keeps the default request byte-compatible with the recorded cassettes. specs/01, specs/03 and specs/05 pick up the new core symbols, the MINUTES unit and the absent-or-0 contract — specs/03 still documented SECONDS, contradicting the docstrings and specs/05 inside the same change.
…r arm 'rules' is a prefix of 'ruleset', so the long-pole fragment also matched the instant ruleset_favorite respx suite and scheduled it ahead of real long poles. 'test_rules' matches only the two intended nodeids. Nothing covered an explicit favorites_only=False / has_new_results=False: they serialise to query 0 (core._params coerces bools before routing), and an inverted filter is the one failure mode that silently returns the wrong rows. Also records where MAX_PAGE_SIZE comes from — the server's AI_MAX_QUERY_RESULTS — and why it sits in core.py rather than settings.py.
The long-pole fragment went from 'rules' to 'test_rules' to fix a false positive and introduced a false negative: 'test_async_rules' is test_ + async_rules, so the heaviest new test in the suite fell to the backfill tail. '_rules' matches both nodeids and still misses the respx suite, where the character before 'rules' is a slash. The replaced line also left its old trailing comment behind. Every max_results test drove the GENERATED sync mirror. The async for + yielded/return shape is the part unasync rewrites rather than copies, so the canonical loop had no coverage at all — verified the new tests fail when the canonical bound check alone is broken.
The wire is not moving to minutes (prod traffic makes it a silent 60x widening for clients we don't control), so this is a documentation correction again rather than a behaviour change — which also removes it from the develop -> master bump decision. max_results and the absent-or-0 contract are unaffected. Also shortens the core helper comments: the argument belongs in specs/05, not beside every line that touches it.
MAX_PAGE_SIZE mirrored the server's AI_MAX_QUERY_RESULTS code default of 1000, but the chart sets 300 in every environment — so live_feed(max_results=500) would have sent limit=500 and got a 400. The cap is an env var the deployment chooses, so the SDK cannot know it: max_results now bounds only how many results the generator yields, and the request is unchanged (which also keeps the default call byte-compatible with every cassette). Also polls the favorite read-after-write — the star is written on the line above, and it was the one such assertion here that neither polled nor guarded 204. One read per attempt, so the recorded interactions are unchanged.
The two are independent: the page stays the server's to choose (50 for web, capped by AI_MAX_QUERY_RESULTS) and _next_page echoes it, so a bounded read keeps paginating in those same small chunks and stops once it has enough. Sending limit=max_results conflated them — the 400 above the deployment's cap was a symptom of that, not the reason.
The two-line read-after-write note was pasted twice in a row, in both the sync and async live tests.
The Versioning table had no row for adding an optional keyword to an
existing public method, so the nearest match was `Signature change on a
public method | major` — which scores this change major even though every
existing call site keeps working untouched.
The table already carves out the additive exception cases on exactly that
reasoning ("no consumer has to change"); this applies the same rule to
keyword arguments, and narrows the signature row to the changes a caller
must actually react to.
Three accuracy fixes from an audit of the revert commits: - The `since` docstring said the parameter "said minutes before 4.4". The version is chosen at the `develop -> master` step, not here, so if that release cuts as anything else the shipped docstring is wrong and nothing would catch it. "in earlier releases" carries the same meaning with no forward reference. - `specs/05` described the CLI's `1440` default in the present tense inside a paragraph that is otherwise entirely historical. That default was retired in this same change set. - `core.py`'s module docstring lists the pure helpers; `as_result_bound` was added to `specs/01` but never to the list beside the code.
The CLI expresses its need for this change set's surfaces as a version requirement — `polyswarm_api>=4.4.0` — rather than probing the installed SDK at runtime. A pin is checkable by pip at install time, before any code runs, and it cannot name a version this repo has not declared. So the bump lands here, in the feature PR, instead of at the release step. Minor, not major: every addition is additive — new methods, a new resource, new optional keywords, new parsed fields — so no existing call site changes. Bumped with bump-my-version and verified the emitted string is a clean `4.4.0`: the serialize config can produce `4.4.0.devN+sha`, and PEP 440 orders that BELOW 4.4.0, which would silently fail the CLI's floor and send its CI to PyPI for a version that does not exist yet. AGENTS.md and specs/05 record the exception and the ordering it forces: this repo must release before the CLI can, since the CLI's floor is unsatisfiable from PyPI until then. Consumer CI is unaffected — it installs this repo from git by branch name.
Both AGENTS.md and the specs/02 inline comment framed the attribute as "set this when the identifier isn't `id`". `YaraRulesetFavorite` is the first resource where the key is not the identifier at all: it names `community` so that `id` stays in the PUT body, because the default would move `id` to the query string and the server 400s. specs/02's prose and its per-resource entry already describe it correctly; the orientation doc and the code comment were the two places still under-describing the mechanism.
`favorites_limit == 5` mirrors a code default the deployment can override — `AI_FAVORITE_RULESETS_LIMIT` is read from the environment — which is the same mistake as the MAX_PAGE_SIZE constant this branch already removed: the client asserting a server-owned number it cannot see. Against a live stack with VCR off and a different cap configured, the exact pin fails on a correct server. Assert the relation instead; the used-vs-limit check beside it still carries the meaning.
Six create-response assertions sat above the `try:` whose `finally:` deletes the ruleset, so any of them failing leaks one on the shared e2e stack. Two of them predate this branch; the four tracking-field assertions widened it. The cost is not abstract: a leaked STARRED ruleset holds one of the team's five favorite slots, and the favorite tests need a free one. `try` now opens immediately after `ruleset_create`, in both transports. specs/04 gains the `poll_equals` / `poll_equals_async` entry it was missing while documenting its sibling `run_concurrently`. It carries a non-obvious invariant worth writing down: `want` must never be None, because not-found during the lag window also reads as None and would turn a vanished resource into a passing assertion — poll a boolean instead. Also records the mirror-image limit, that a False poll cannot ride out a stale False. AGENTS.md now names where the bump exception comes from: it is this repo's instance of a workspace-level standard, not a rule the repo grants itself.
Two comments carried an "(F9)" review token. AGENTS.md keeps internal references out of published artefacts; the rule names commits and PR text, but a committed source comment outlives both. The sentences read the same without it.
The rules fragment has now been wrong three times — "rules" caught "ruleset", "test_rules" missed "test_async_rules", and "_rules" is itself a prefix of "_ruleset" so it front-loaded the instant unit tests. Every substring of those two test names is a prefix of test_ruleset_*, so no fourth fragment fixes it. Give the matcher a way to say "exactly this test" instead: a fragment starting with "::" is matched as a nodeid SUFFIX. The two long poles now name themselves exactly; every other fragment keeps substring matching, which is right for the ones that legitimately span several tests. Verified: test_rules and test_async_rules rank as long poles, the unit test backfills to the tail.
…nternals `as_result_bound` was listed on the documented public surface, which put a four-line internal helper under "rename or removal of a public symbol = major". Nothing outside this package calls it — not the paired CLI — and every sibling helper in core.py is already underscore-private. Renamed to `_as_result_bound` and dropped from the specs/05 surface listing; specs/01 still records it among the helpers, where it belongs. This repo is public, so two other things should not have been in it: - Production telemetry. specs/05 published request volumes, distinct API-key and user-agent counts, and the observed value range for a parameter. The decision it supports — the wire stays seconds, because re-reading live traffic as minutes widens every window 60x silently — needs none of it. - A deployment's configured page cap. The mechanism (a per-deployment env var, set well below a large bound) is what explains the 400; the number is not. Also stops the spec arguing with its own drafts. "An earlier revision of this document claimed…" and "the MAX_PAGE_SIZE mistake again" describe review history, which git already holds; a spec should read as the current state so a reader does not have to work out which claim is live.
…om source The bullet read as if any sibling needing a new surface justifies bumping the version in a feature PR. It does not. The exception exists because a floor cannot name a version this repo has not declared — which is only a problem for a consumer whose CI installs this repo from source by branch and can therefore adopt before publication. A consumer that installs published artifacts has no such need: it adopts after the release, and this repo bumps at its own release step as usual. Stating the condition keeps the exception from being applied where it does not hold.
Hunt-page ruleset tracking: favorites, rule counts, hunt provenance
|
Review — release 4.4.0 ( Checked against Three things to fix. 1. The local-pinned-stack loop reads:
After this PR the suite targets 2. The comment there still names 3. PR body names fields that do not exist The changelog says hunt resources carry Non-blocking: |
TL;DR — changelog for 4.4.0
favorite,favorited_at,rule_count,historical_hunt_count,new_results_count,new_results_counted_at; hunt resources carry the frozen-rule provenance (source_rule,source_rule_changed).ruleset_favorite(id, favorite)on both clients; the server'sFAVORITE_LIMITrefusal surfaces as a typed error withfavorites_used/favorites_limit.ruleset_listacceptsname,status,favorites_only,has_new_results(conjunctive) and asserts them universally.live_feedgainslivescan_idscoping and an optionalmax_resultstotal bound (0= unbounded);sincestays in seconds.ai,sandbox) with cassettes re-recorded (test: e2e stack hostname is now 'ai' (was artifact-index-e2e) #320).develop(bumped in Hunt-page ruleset tracking: favorites, rule counts, hunt provenance #321 so the CLI could pin the floor) — this PR publishes it. PyPI upload fires on theversionchange onmaster.Requires
nullfor the new fields against prod (additive, no failure).polyswarm_api>=4.4.0).