Repository navigation
Conversation
…y imply Both backends rank the whole candidate set before cutting a page: Chroma reads the collection, Postgres selects every matching row. That is a deliberate trade - a decay-weighted re-rank can only lift a fresher cell above a stale one if the cell was actually fetched - and it means search cost grows with how much is stored. The documentation did not say so anywhere, and a reader had to infer it from the code. `benchmarks/run.py` measures ingest, search, a by-id read and a listing page at three collection sizes; `benchmarks/attribution.py` splits a search into the collection read and the Python work on top of it (deserialise, check access, blend the decay score), so the shape of the curve is explained rather than assumed. `docs/performance.md` carries the numbers with the machine that produced them and says plainly what is *not* measured: the embedding model, concurrency, and selective filters. Three decisions in the harness came from getting it wrong first, and both the scripts and the docs record them rather than hiding them: - **One operation, repeated.** The first version rotated the search term between samples, which measures six different operations and reports the spread between them as if it were the spread of one. It produced a table where a search over 5000 cells looked faster than the same search over 1000. Every sample now uses the same query. - **min / median / p95, never an average.** The first table was taken while the same machine was doing other work and showed no growth at all between 1000 and 5000 cells. Re-measuring on a quiet machine, and reporting the minimum beside the distribution, made the curve visible - and the docs say so, because a reader who runs this on a busy machine will otherwise get a different and wrong shape. - **No timing assertions in CI.** A shared runner cannot produce a number worth publishing, so CI runs the harness in a smoke mode (`--sizes 20 --queries 2`) to prove it works and asserts nothing about how long anything took. Without that, a broken harness rots quietly until somebody needs it. The headline numbers, measured on this machine (Windows 11, 12 CPUs, Python 3.12): a by-id read stays flat (5.6 -> 6.9 ms from 100 to 5000 cells); search grows with the store (4.0 -> 29.1 -> 163.8 ms minimum); the listing page grows fastest of the three (3.6 -> 388 ms) because it deserialises every cell in the collection to keep twenty; and ingest cost per cell rises with the store (244 -> 47 cells/s). Postgres numbers are welcome by pull request against a service container - this machine has no PostgreSQL, which is why the tables are Chroma-only and say so.
Owner
Author
|
Landed on master in the v0.1.0 chain: the branch was fast-forward merged as part of |
This branch was successfully deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
docs(benchmarks): measure the storage paths, and state the limits they imply
Both backends rank the whole candidate set before cutting a page: Chroma reads the
collection, Postgres selects every matching row. That is a deliberate trade - a
decay-weighted re-rank can only lift a fresher cell above a stale one if the cell was
actually fetched - and it means search cost grows with how much is stored. The
documentation did not say so anywhere, and a reader had to infer it from the code.
benchmarks/run.pymeasures ingest, search, a by-id read and a listing page at threecollection sizes;
benchmarks/attribution.pysplits a search into the collection readand the Python work on top of it (deserialise, check access, blend the decay score), so
the shape of the curve is explained rather than assumed.
docs/performance.mdcarriesthe numbers with the machine that produced them and says plainly what is not
measured: the embedding model, concurrency, and selective filters.
Three decisions in the harness came from getting it wrong first, and both the scripts
and the docs record them rather than hiding them:
samples, which measures six different operations and reports the spread between them
as if it were the spread of one. It produced a table where a search over 5000 cells
looked faster than the same search over 1000. Every sample now uses the same query.
machine was doing other work and showed no growth at all between 1000 and 5000
cells. Re-measuring on a quiet machine, and reporting the minimum beside the
distribution, made the curve visible - and the docs say so, because a reader who
runs this on a busy machine will otherwise get a different and wrong shape.
publishing, so CI runs the harness in a smoke mode (
--sizes 20 --queries 2) toprove it works and asserts nothing about how long anything took. Without that, a
broken harness rots quietly until somebody needs it.
The headline numbers, measured on this machine (Windows 11, 12 CPUs, Python 3.12):
a by-id read stays flat (5.6 -> 6.9 ms from 100 to 5000 cells); search grows with the
store (4.0 -> 29.1 -> 163.8 ms minimum); the listing page grows fastest of the three
(3.6 -> 388 ms) because it deserialises every cell in the collection to keep twenty;
and ingest cost per cell rises with the store (244 -> 47 cells/s). Postgres numbers
are welcome by pull request against a service container - this machine has no
PostgreSQL, which is why the tables are Chroma-only and say so.