Skip to content

docs(benchmarks): measure the storage paths, and state the limits they imply - #23

Closed
glatinone wants to merge 1 commit into
test/python-sdk-livefrom
docs/measured-performance
Closed

glatinone wants to merge 1 commit into
test/python-sdk-livefrom
docs/measured-performance

Conversation

@glatinone

Copy link
Copy Markdown
Owner

docs(benchmarks): measure the storage paths, and state the limits they imply

Both backends rank the whole candidate set before cutting a page: Chroma reads the
collection, Postgres selects every matching row. That is a deliberate trade - a
decay-weighted re-rank can only lift a fresher cell above a stale one if the cell was
actually fetched - and it means search cost grows with how much is stored. The
documentation did not say so anywhere, and a reader had to infer it from the code.

benchmarks/run.py measures ingest, search, a by-id read and a listing page at three
collection sizes; benchmarks/attribution.py splits a search into the collection read
and the Python work on top of it (deserialise, check access, blend the decay score), so
the shape of the curve is explained rather than assumed. docs/performance.md carries
the numbers with the machine that produced them and says plainly what is not
measured: the embedding model, concurrency, and selective filters.

Three decisions in the harness came from getting it wrong first, and both the scripts
and the docs record them rather than hiding them:

  • One operation, repeated. The first version rotated the search term between
    samples, which measures six different operations and reports the spread between them
    as if it were the spread of one. It produced a table where a search over 5000 cells
    looked faster than the same search over 1000. Every sample now uses the same query.
  • min / median / p95, never an average. The first table was taken while the same
    machine was doing other work and showed no growth at all between 1000 and 5000
    cells. Re-measuring on a quiet machine, and reporting the minimum beside the
    distribution, made the curve visible - and the docs say so, because a reader who
    runs this on a busy machine will otherwise get a different and wrong shape.
  • No timing assertions in CI. A shared runner cannot produce a number worth
    publishing, so CI runs the harness in a smoke mode (--sizes 20 --queries 2) to
    prove it works and asserts nothing about how long anything took. Without that, a
    broken harness rots quietly until somebody needs it.

The headline numbers, measured on this machine (Windows 11, 12 CPUs, Python 3.12):
a by-id read stays flat (5.6 -> 6.9 ms from 100 to 5000 cells); search grows with the
store (4.0 -> 29.1 -> 163.8 ms minimum); the listing page grows fastest of the three
(3.6 -> 388 ms) because it deserialises every cell in the collection to keep twenty;
and ingest cost per cell rises with the store (244 -> 47 cells/s). Postgres numbers
are welcome by pull request against a service container - this machine has no
PostgreSQL, which is why the tables are Chroma-only and say so.

…y imply

Both backends rank the whole candidate set before cutting a page: Chroma reads the
collection, Postgres selects every matching row. That is a deliberate trade - a
decay-weighted re-rank can only lift a fresher cell above a stale one if the cell was
actually fetched - and it means search cost grows with how much is stored. The
documentation did not say so anywhere, and a reader had to infer it from the code.

`benchmarks/run.py` measures ingest, search, a by-id read and a listing page at three
collection sizes; `benchmarks/attribution.py` splits a search into the collection read
and the Python work on top of it (deserialise, check access, blend the decay score), so
the shape of the curve is explained rather than assumed. `docs/performance.md` carries
the numbers with the machine that produced them and says plainly what is *not*
measured: the embedding model, concurrency, and selective filters.

Three decisions in the harness came from getting it wrong first, and both the scripts
and the docs record them rather than hiding them:

- **One operation, repeated.** The first version rotated the search term between
  samples, which measures six different operations and reports the spread between them
  as if it were the spread of one. It produced a table where a search over 5000 cells
  looked faster than the same search over 1000. Every sample now uses the same query.
- **min / median / p95, never an average.** The first table was taken while the same
  machine was doing other work and showed no growth at all between 1000 and 5000
  cells. Re-measuring on a quiet machine, and reporting the minimum beside the
  distribution, made the curve visible - and the docs say so, because a reader who
  runs this on a busy machine will otherwise get a different and wrong shape.
- **No timing assertions in CI.** A shared runner cannot produce a number worth
  publishing, so CI runs the harness in a smoke mode (`--sizes 20 --queries 2`) to
  prove it works and asserts nothing about how long anything took. Without that, a
  broken harness rots quietly until somebody needs it.

The headline numbers, measured on this machine (Windows 11, 12 CPUs, Python 3.12):
a by-id read stays flat (5.6 -> 6.9 ms from 100 to 5000 cells); search grows with the
store (4.0 -> 29.1 -> 163.8 ms minimum); the listing page grows fastest of the three
(3.6 -> 388 ms) because it deserialises every cell in the collection to keep twenty;
and ingest cost per cell rises with the store (244 -> 47 cells/s). Postgres numbers
are welcome by pull request against a service container - this machine has no
PostgreSQL, which is why the tables are Chroma-only and say so.
@glatinone

Copy link
Copy Markdown
Owner Author

Landed on master in the v0.1.0 chain: the branch was fast-forward merged as part of b940905..92b88ee and released as v0.1.0. Closing so the open list matches reality - the commits are in master, and the tag points at them.

@glatinone glatinone closed this Oct 4, 2026

This branch was successfully deployed

1 active deployment
github-pages — 92b88eed Deployed Oct 4, 2026 by glatinone via deploy #9
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant