How the CLI is tested: the CliRunner harness, the mocking styles (SDK-boundary mocks, SDK-transport mocks, VCR cassettes), the formatter-unit style for pure rendering, the cassette layout and record workflow, and how to run the suite. Files: tests/, tests/vcr/, src/conftest.py, pyproject.toml ([project.optional-dependencies].tests, [tool.pytest.ini_options]).
-
Anything that is command behaviour is driven through
click.testing.CliRunner— argument parsing, the SDK call, the wiring, the exit code: exercise the real command tree, never an internal function standing in for it. No live PolySwarm stack is required. The one sanctioned exception is pure rendering logic — see Style 3. -
Mock at exactly one point per path — the SDK method, the SDK transport, or VCR. A test patches
polyswarm_api.api.PolyswarmAPI.<method>(unit-style, the default), patches the SDK transportPolyswarmSession.executewhen the behaviour under test runs inside the SDK method (Style 4), or lets VCR replay recorded HTTP (end-to-end). Never two of them for the same path. The CLI's own code is exercised either way. -
VCR is an efficiency cache, not a load-bearing requirement. The suite must pass against a live e2e stack with VCR off. Don't hardcode
record_mode='none'; if a test only works against its recorded cassette, that's a bug in the test. Note this is about a test's logic, not its fixtures: a.clicksnapshot pins server-generated ids and timestamps, so re-recording needs a stack in a particular state — see Re-recording a cassette. -
Never
cpa cassette from a sibling test, never hand-edit cassette bytes. Re-record against a live stack.
pip install .[tests] # pytest, pytest-cov, PyYAML, vcrpy
pytest # testpaths = ["tests"]; capture to a file for analysissrc/conftest.py puts src/ on sys.path so imports resolve without a src. prefix. The SDK (polyswarm_api) must be installed too — in CI it's pulled from the SDK repo's branch archive (see 05-sdk-contract.md); locally, install the SDK you're testing against (an editable checkout of the SDK works).
For commands whose behaviour is "call this SDK method, render the result," patch the method on the SDK client and assert the call + the rendered output. Example: tests/field_property_test.py patches polyswarm_api.api.PolyswarmAPI.metadata_field_properties_{write,get,delete,list} and drives CliRunner. This insulates the test from SDK transport internals — it survives SDK refactors as long as the method name/signature is stable.
Use this style when the point of the test is the command wiring + formatting, not the wire shape.
For end-to-end coverage (the CLI calling the SDK calling a recorded server), tests/cli_test.py uses vcrpy. Each test has two fixtures in tests/vcr/:
tests/vcr/<name>.vcr— the recorded HTTP interaction(s).tests/vcr/<name>.click— the expected rendered CLI output (compared againstresult.output).
The VCR object is built with the default request matcher (method, scheme, host, port, path, query) — it does not match on headers or body, so cassettes survive User-Agent changes and request-body relocation, and the query matcher compares params as a set (order-independent). Default record_mode='once'.
Helpers in cli_test.py: _run_cli(args) invokes the command tree under a cassette; _assert_text_result / _assert_json_result compare result.output and the exit code against the .click fixture.
Some cassettes need a stack state, not just a live stack. The ruleset-favorite ones do,
because _assert_text_result compares the whole rendered block verbatim and the budget
counter is part of it: test_ruleset_favorite_text pins Favorites used: 2 of 5 and
test_ruleset_unfavorite_text pins 1 of 5. Their counters only make sense as a
sequence (star, star, unstar), so re-recording one alone produces a set that contradicts
itself.
Re-recording them is harder than it looks, and the shipped set is not what a single
pass produces — test_ruleset_favorite_text acts on a ruleset that does not appear in
test_ruleset_list_json's inventory at all, so the two were recorded against different
stack states. Two couplings to plan around before starting:
unittestruns methods in sorted-name order, andtest_ruleset_list_jsonsorts betweenfavorite_textandunfavorite_text. Its snapshot pins"favorite": falseon every ruleset, so a full-suite recording captures it while a ruleset is starred and that snapshot changes too. Re-record it with the others, or not at all.- Renaming or reordering any of these methods changes the sequence silently: replay keeps passing, and only the next re-record surfaces the contradiction.
If keeping them consistent stops being worth it, break the coupling rather than documenting
a wider one — _assert_text_result already takes a replace= hook, and normalising the
budget counter there makes each cassette independent of what ran before it.
A refusal at the favorite cap is deliberately not recorded here. Saturating the budget
takes all five team slots, which is exactly the state the tests above must not be in, so
the cassette and its siblings cannot both be satisfiable in one VCR-off run. That refusal is
pinned without a stack instead: the CLI's message and exit code at the SDK boundary, the
envelope spelling against a real PolyswarmRequest, and the wire shape by the SDK's own
stubbed-transport suite, which declined a recording for the same reason.
rm tests/vcr/<name>.vcr # (and regenerate <name>.click from the new run)
pytest tests/cli_test.py::<Class>::<test> # records against whatever stack your env points atPoint your environment at a live e2e stack, run the test, and commit the freshly recorded .vcr (+ updated .click). The deletion is what makes VCR record.
For rendering logic with no command-tree behaviour — which labelled line a given field set produces — construct the formatter directly (TextOutput(color=False)) and call the resource method with an SDK resource built from a literal dict, asserting on the returned lines (write=False, no stream, no cassette). Example: tests/known_good_field_test.py renders ArtifactInstances that differ only in state / known_good and asserts which Detections/Status line comes out. This is the right choice when the branch matrix is wide and every branch is a function of the resource's fields — a cassette per branch would mean recording a server state that only the formatter cares about.
Use it only for that. Argument parsing, SDK calls, generator consumption, ctx.obj wiring and exit codes are command behaviour: a formatter unit test can't observe them, so those need Style 1 or Style 2. A command whose rendering is covered by Style 3 still needs at least one CliRunner test proving the command reaches the formatter at all.
For behaviour that happens inside an SDK endpoint method, a Style 1 mock proves nothing: patching PolyswarmAPI.search_url replaces the very code under test. IoC refanging is the case today — the SDK refangs search_url's argument (and the other URL / domain / IP inputs) before it builds the request, so the observable effect is the request that reaches the transport.
Patch polyswarm_api.session.PolyswarmSession.execute — the SDK's documented session customization point — and assert on the PolyswarmRequest it receives: request.params (query string) and request.input_json (JSON body), read-only. Example: tests/refang_test.py.
Use this style only when Style 1 would bypass the behaviour under test. Moving these tests "up" to endpoint-method mocks would keep them green while dropping their coverage. The dependency it adds (the execute(request) seam and those two request fields) is recorded in 05-sdk-contract.md; an SDK rename there breaks these tests, which is the intended signal.
- The command parses its arguments and calls the expected SDK method with the expected arguments.
- The result is rendered through the right formatter method, for both
--output-format textandjsonwhere both matter. - Error paths map to the right exit code (no-results → 1, partial → 3, internal/other → 2) — see
01-architecture.md§exit-code mapping.
This spec describes the harness as it stands. Not yet documented (add as the suite grows): a per-command coverage matrix, a documented "VCR-off against live e2e" CI job, and conventions for fixture/.click generation. See 99-open-questions.md.
A test never asks the installed SDK whether it has a feature. The pin guarantees
it. (For the one pre-existing render guard this does not cover, see
03-formatters.md §Known-good artifact instances.) When this repo needs a surface the SDK does not yet publish, the SDK bumps its
version and pyproject.toml raises polyswarm_api>= to it. pip then refuses the
combination that would fail, at install time, before a single test runs — so a test
can simply use the surface.
Do not reintroduce per-test skip guards — hasattr on a method, a built resource's
attribute, a parameter in the installed signature. They are the shape this rule exists to
exclude, and the reasons are worth keeping written down:
- The fact lived twice. The pin said one thing; each guard re-derived the same thing at runtime. Nothing kept them in sync, so every guard was one edit away from disagreeing with the tests it gated.
- Both ways of disagreeing are defects. A guard that checks less than its test uses lets the test run and fail where it should have skipped. One that checks more skips a test that would have passed — silently dropping coverage while CI stays green. The second is the dangerous one, because nothing reports it.
- It never verified the thing it claimed to protect. A green run against the paired SDK said nothing about the floor install the guards existed for.
The version contract has none of that: the claim is checked once, by a tool, against the artifact that will actually be installed.
When you need a new SDK surface, in order: add it in polyswarm-api → bump that
repo's version (minor, for an additive surface) in the same PR, because this repo's
floor cannot name a version the SDK has not declared → raise the floor here → use the
surface in code and tests with no guard. CI installs the SDK from git by branch name,
so an unreleased version is not an obstacle; see
05-sdk-contract.md §Current floor for the ordering that
forces at release time.
The failure mode to expect, and it is a good one: if the paired SDK branch is
missing, CI falls back to the SDK's develop, whose version does not satisfy the new
floor, and pip install .[tests] fails loudly. Before, that fallback silently tested
against the wrong SDK.