Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
16 commits
Select commit Hold shift + click to select a range
cdb2a77
feat(hunts): expose matched strings on hunt-result resources
kyle-buchmiller Aug 21, 2026
08a31f4
test(hunts): pin the three matched-strings states
kyle-buchmiller Aug 21, 2026
021d356
docs(specs): document matched_strings as a three-state attribute
kyle-buchmiller Aug 21, 2026
b29bc3d
docs(hunts): tighten the matched_strings comment
kyle-buchmiller Aug 25, 2026
74d0da9
test: pin matched_strings against the live server, not just dict.get
kyle-buchmiller Aug 25, 2026
84c8a0b
docs(specs): note matched_strings on the hunt-result resources
kyle-buchmiller Aug 25, 2026
4d57f66
feat(hunts): expose the withheld-string count on hunt results
kyle-buchmiller Aug 28, 2026
49a64db
test: pin matched_strings_dropped against the live server, and fix th…
kyle-buchmiller Aug 29, 2026
8299433
docs: finish the None-ambiguity correction, and make two tests assert…
kyle-buchmiller Aug 29, 2026
b482498
test: pin the per-entry shape against the server, and fix the class-r…
kyle-buchmiller Aug 30, 2026
024449d
docs: the ...List classes DO parse responses -- delete responses
kyle-buchmiller Aug 30, 2026
28cabee
test: make both raw-json assertions falsifiable
kyle-buchmiller Aug 30, 2026
a4934b2
test: assert the list null on .json, deep-copy the fixture, trim dupl…
kyle-buchmiller Aug 31, 2026
09e9343
Merge origin/develop into the matched-strings branch
kyle-buchmiller Sep 3, 2026
13c191c
docs: let the four-cause table own the None enumeration
kyle-buchmiller Sep 3, 2026
1d2dab9
test: say what a null matched_strings means on a green stack
kyle-buchmiller Sep 3, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
44 changes: 44 additions & 0 deletions specs/02-resources.md
Original file line number Diff line number Diff line change
Expand Up @@ -327,6 +327,50 @@ Classmethod builders (each returns a `PolyswarmRequest` descriptor):

**No instance methods** that issue HTTP. Uploading to the pre-signed S3 URL is done via the session: `await api.session.upload_file(instance.upload_url, artifact)` (or `api.session.upload_file(...)` for sync).

### `LiveHuntResult` / `HistoricalHuntResult`

**`matched_strings`.** The yara strings behind a hunt hit, so a consumer can see *why*
a rule fired rather than only which one did. Additive and optional, parsed with `.get()`
like `known_good` / `state` above — a server too old to emit it parses to `None` with no
behaviour change, and a subscript would raise on every result instead.

It is **three-state** and the states are not interchangeable; the table, the per-entry
dict shape and the lower-bound caveat live in
[`05-downstream-contract.md`](./05-downstream-contract.md)
§"`matched_strings` on hunt results" — read it there rather than inferring from the
attribute. The short version a parser needs: `None` means *not reported* — **four**
distinct causes, enumerated in that table and revisited under "four places, not three"
below — `[]` means *matched with no byte evidence*, and a populated list is evidence.

**`matched_strings_dropped`.** A sibling `int`/`None`, parsed the same additive way:
how many matched strings the server's byte budget withheld from this result. `None` is
**ambiguous in the same way as `matched_strings`** and must be read the same way: on a
**detail** route it means nothing was withheld; under any of the other three causes it
means nothing looked. It is never a claim that the
evidence is complete — which matters because the list endpoints always send `None` (below),
so on a list row the two readings are not interchangeable.

A non-null count always means "the list you have is short by this much". It never
accompanies an empty `matched_strings` — a match's first string is never withheld.

Both `…List` subclasses inherit these from their parent's `__init__`, so all four
hunt-result classes carry them — but on the list endpoints the values are always `None`
by design.

**Do not read the class as telling you the route.** For the live pair it does not:
`live_feed()` builds its request with `LiveHuntResult.list(...)`, which hits
`/hunt/live/list` but parses rows as **`LiveHuntResult`** (see
[`03-endpoints.md`](./03-endpoints.md)). `LiveHuntResultList` is only ever a *delete*
**builder** — but the delete response is parsed **through** it (`_build_request` sets
`result_parser=cls`), so `live_feed_delete()` and `historical_results_delete()` both yield
`…List` instances, with both fields `None`. From a *read*, only `historical_results()`
yields them.

So `None` reaches a caller from four places, not three: an older server, deleted evidence,
a list route, and a **delete response**. A `LiveHuntResult` carrying
`matched_strings is None` may well have come from the list route — which is exactly why
that `None` is ambiguous and must not be read as "nothing to show".

### `LocalArtifact`

A file-system or in-memory artifact prepared for upload. Constructed via:
Expand Down
2 changes: 2 additions & 0 deletions specs/04-testing.md
Original file line number Diff line number Diff line change
Expand Up @@ -24,6 +24,8 @@ How the test suite is organised. Three layers: pure unit tests (no HTTP at all
- `test/metadata_field_properties_test.py` — the canonical example of the parametrised `ClientTestCase` harness with `respx`-backed mocking.
- `test/client_scan_test.py` — sync, VCR-backed integration tests (not yet on the parametrised harness — follow-up work).
- `test/async_client_test.py` — async, VCR-backed integration tests (not yet on the parametrised harness — follow-up work).
- `test/hunt_matched_strings_test.py` — pure-unit tests for the three-state `matched_strings` contract and its `matched_strings_dropped` sibling, across all four hunt-result classes. The endpoint behaviour they cannot see (that the detail route actually emits the keys and the list route sends `null`) is pinned live in `client_scan_test.py::test_live` / `async_client_test.py::test_async_live` — the e2e-first + pure-unit pairing invariant 1 asks for.
- `test/known_good_test.py` — pure-unit tests for the known-good resource fields.
- `test/jmespath_test.py` — unit tests for `BaseJsonResource.jmespath`.
- `test/vcr/*.vcr` — recorded cassettes.
- `test/malicious` — fixture file for upload tests (`test/eicar.yara` was retired when the rules tests moved to per-test `uid_yara` bodies).
Expand Down
55 changes: 55 additions & 0 deletions specs/05-downstream-contract.md
Original file line number Diff line number Diff line change
Expand Up @@ -273,6 +273,61 @@ What is **not** part of the contract:
- The exact server JSON shape — that lives in the artifact-index repo's contract.
- The order of fields in the JSON.

### `matched_strings` on hunt results — a three-state attribute

`LiveHuntResult.matched_strings` / `HistoricalHuntResult.matched_strings` (and therefore
their `…List` subclasses) carry the yara strings behind a hunt hit. It is read with
`.get()` rather than a subscript, deliberately: the key is **additive**, so a server
older than it omits the key entirely and a subscript would raise on every result.

Three values are possible and consumers **must not** collapse them:

| Value | Meaning |
|---|---|
| `None` | Not reported. **Four** causes, which `.get()` collapses: the **list** endpoints send an explicit `null` rather than fetch a blob per row; **delete** responses (`live_feed_delete` / `historical_results_delete`) are parsed through the `…List` classes and carry `null` the same way; a server predating the field omits it entirely; and stored evidence may have been deleted. "We don't know", *not* "there was nothing". |
| `[]` | The rule matched and there is no byte evidence to show — a rule with no strings section, one whose matching strings are all `private`, or one that matched on absence (`not $a`, `none of them`). |
| `[…]` | The evidence. A **lower bound**, not a match count: `any of them` prints only the strings that hit, `private` strings never appear, and the server may withhold some past a size limit (see `matched_strings_dropped`). |

Each entry is a dict:

```python
{'offset': 78, 'identifier': '$stub', 'length': 14, 'data': '54 68 69 …', 'truncated': False}
```

- `data` is kept **exactly as yara rendered it** — a hex string comes back as byte pairs, a text string as ASCII with `\xNN` escapes. Only yara knows which applies, so it is not decoded back to bytes.
- `length` is the **stored** length, capped server-side. Past the cap the true length is unrecoverable.
- `truncated` means "there was more than this". It over-reports at exactly the cap, because nothing in the output distinguishes a match that ended there from one that was cut.

### `matched_strings_dropped` — the count that keeps a short list honest

A sibling attribute on the same four classes, `int` or `None`. It is how many matched
strings the server's per-result byte budget withheld, and it exists because a truncated
list is otherwise indistinguishable from a complete one: a consumer reading twelve
entries would conclude the rule hit twelve times when it hit thirty-one.

`None` carries the same ambiguity as `matched_strings` itself and should be read the same
way: on a **detail** route it means nothing was withheld; under any of the other three
causes in the table above, it means nothing looked. It is not
a claim that the evidence is complete. It is deliberately a
**sibling** rather than a key inside `matched_strings`, which stays a plain list.

A populated `matched_strings` with a non-null count is the normal shape for a verbose
ruleset. The first string of a match is never withheld, so this can never accompany an
empty list.

**How much of this is pinned against a real server.** The live pair
(`LiveHuntResult` / `LiveHuntResultList`) is verified end to end — `test_live` /
`test_async_live` assert the detail route carries evidence, that list rows do not, and
that the server serves `matched_strings_dropped`. The **historical** pair follows by
symmetry, not by measurement: the e2e stack does not reliably populate historical results
inside a test window, so nothing pins that those routes emit either key. The server
renders both pairs through the same helpers, which is why symmetry is a reasonable
assumption — but it is an assumption. See `specs/99-open-questions.md`.

**Evidence lives on the detail routes only.** `live_feed()` and `historical_results()`
page over list endpoints and will always yield `None` here; fetch a single result
(`live_result(id)` / `historical_result(id)`) to get the strings.

## Pagination

Generator endpoints return an iterable:
Expand Down
40 changes: 40 additions & 0 deletions specs/99-open-questions.md
Original file line number Diff line number Diff line change
Expand Up @@ -170,3 +170,43 @@ def test_rescans(self):
```

Cleanup with `try/finally` + `except NotFoundException: pass` tolerates the ioc-cache divergence (GET-by-host can serve a cached id that DELETE-by-id no longer finds). When that artifact-index bug is fixed the `except` becomes redundant.

## Historical hunt-result fields are not pinned against a live server

**Status:** gap, blocked on the e2e stack.

`matched_strings` / `matched_strings_dropped` are asserted end to end for the **live**
hunt pair only. `test_historical_results` tolerates an empty result set by design — the
stack does not reliably populate historical results inside a test window — so nothing
verifies that `/hunt/historical/results` emits either key, or that
`/hunt/historical/results/list` sends the explicit `null`.

Both specs previously stated the contract for "all four classes" as established fact;
they now say the historical half follows by symmetry. The server renders both pairs
through the same serializer helpers, so the assumption is reasonable — but a fabricated
response asserts what we *think* the server returns (invariant 1), and that is the state
the historical half is in.

**Action:** if the stack gains a way to produce a historical result deterministically,
add the same assertions to a historical live test and delete this entry.

## The non-null `matched_strings_dropped` path is not pinned against a live server

**Status:** gap, probably not worth closing with a test.

`test_live` / `test_async_live` assert only the `is None` arm — correctly, since the
per-test rule is small and the server withholds nothing from it. So the two strongest
claims `05-downstream-contract.md` makes about this field rest entirely on hand-written
pure-unit dicts:

- a non-null count means "the list you have is short by this much", and
- it can never accompany an empty `matched_strings`, because a match's first string is
never withheld.

That is the same "asserts what we *think* the server returns" gap invariant 1 exists to
close, and it sits alongside the historical-pair entry above.

Producing an over-budget match on the e2e stack means a rule whose matches exceed the
server's per-hunt byte budget across a single artifact — engineering a fixture for that is
disproportionate to what it would pin. **Recorded rather than tested, deliberately.** If a
stack fixture ever produces one cheaply, assert both claims there and delete this entry.
10 changes: 10 additions & 0 deletions src/polyswarm_api/resources.py
Original file line number Diff line number Diff line change
Expand Up @@ -805,6 +805,11 @@ def __init__(self, content, api=None):
self.sha1 = content.get('sha1')
self.rule_name = content['rule_name']
self.tags = content['tags']
# `.get()`, not a subscript -- both keys are additive, so an older server omits
# them. None is AMBIGUOUS on both (four causes) and is never a claim that the
# evidence is complete. Contract: specs/05-downstream-contract.md.
self.matched_strings = content.get('matched_strings')
self.matched_strings_dropped = content.get('matched_strings_dropped')
self.polyscore = content['polyscore']
self.malware_family = content['malware_family']
self.detections = content['detections']
Expand Down Expand Up @@ -863,6 +868,11 @@ def __init__(self, content, api=None):
self.created = core.parse_isoformat(content['created'])
self.rule_name = content['rule_name']
self.tags = content['tags']
# `.get()`, not a subscript -- both keys are additive, so an older server omits
# them. None is AMBIGUOUS on both (four causes) and is never a claim that the
# evidence is complete. Contract: specs/05-downstream-contract.md.
self.matched_strings = content.get('matched_strings')
self.matched_strings_dropped = content.get('matched_strings_dropped')
self.polyscore = content['polyscore']
self.malware_family = content['malware_family']
self.detections = content['detections']
Expand Down
24 changes: 24 additions & 0 deletions test/async_client_test.py
Original file line number Diff line number Diff line change
Expand Up @@ -865,6 +865,30 @@ async def test_async_live(self, uid):
result = await api.live_result(result_id)
assert result.download_url

# The list/detail split, pinned against the real server rather than prose.
# The pure-unit tests exercise dict.get and would pass identically if the
# server never grew the field; only a cassette shows what it actually sent.
assert result.matched_strings, (
'detail route should carry the yara evidence. A null here against an\n'
'otherwise-green stack means the analyzer image predates the change\n'
'that emits `strings` -- check the analyzer, not this repo.')
# On .json for the same reason as the count below: the attribute cannot
# distinguish a served null from an absent key, and specs/05 claims the
# list route sends an explicit null.
assert my_results[0].json['matched_strings'] is None, \
'list rows carry the key as null, not the evidence'
# On .json, not the attribute: `is None` cannot tell a served null from an
# absent key, and what needs pinning is that the server SENDS this field.
assert 'matched_strings_dropped' in result.json, \
'server must serve the withheld-count field'
assert result.matched_strings_dropped is None, \
'nothing withheld for a match this small'
# The per-entry shape is contract (specs/05) but was pinned only by a hand-written
# fixture -- i.e. what we THINK the server sends. This asserts it against what the
# server actually sent, so a key rename cannot pass the suite VCR-off.
assert set(result.matched_strings[0]) == {
'offset', 'identifier', 'length', 'data', 'truncated'}, result.matched_strings[0]

await api.live_feed_delete([result_id])
with pytest.raises(exceptions.NotFoundException):
await api.live_result(result_id)
Expand Down
24 changes: 24 additions & 0 deletions test/client_scan_test.py
Original file line number Diff line number Diff line change
Expand Up @@ -508,6 +508,30 @@ def test_live(self):
result = api.live_result(result_id)
assert result.download_url

# The list/detail split, pinned against the real server rather than prose.
# The pure-unit tests exercise dict.get and would pass identically if the
# server never grew the field; only a cassette shows what it actually sent.
assert result.matched_strings, (
'detail route should carry the yara evidence. A null here against an\n'
'otherwise-green stack means the analyzer image predates the change\n'
'that emits `strings` -- check the analyzer, not this repo.')
# On .json for the same reason as the count below: the attribute cannot
# distinguish a served null from an absent key, and specs/05 claims the
# list route sends an explicit null.
assert my_results[0].json['matched_strings'] is None, \
'list rows carry the key as null, not the evidence'
# On .json, not the attribute: `is None` cannot tell a served null from an
# absent key, and what needs pinning is that the server SENDS this field.
assert 'matched_strings_dropped' in result.json, \
'server must serve the withheld-count field'
assert result.matched_strings_dropped is None, \
'nothing withheld for a match this small'
# The per-entry shape is contract (specs/05) but was pinned only by a hand-written
# fixture -- i.e. what we THINK the server sends. This asserts it against what the
# server actually sent, so a key rename cannot pass the suite VCR-off.
assert set(result.matched_strings[0]) == {
'offset', 'identifier', 'length', 'data', 'truncated'}, result.matched_strings[0]

api.live_feed_delete([result_id])
with pytest.raises(exceptions.NotFoundException):
api.live_result(result_id)
Expand Down
Loading
Loading