From 92b88eed8bb19849f7e457d08640a8ec4d2b5d8b Mon Sep 17 00:00:00 2001 From: glatinone <93207632+glatinone@users.noreply.github.com> Date: Sun, 4 Oct 2026 17:17:25 +0800 Subject: [PATCH] docs(benchmarks): measure the storage paths, and state the limits they imply Both backends rank the whole candidate set before cutting a page: Chroma reads the collection, Postgres selects every matching row. That is a deliberate trade - a decay-weighted re-rank can only lift a fresher cell above a stale one if the cell was actually fetched - and it means search cost grows with how much is stored. The documentation did not say so anywhere, and a reader had to infer it from the code. `benchmarks/run.py` measures ingest, search, a by-id read and a listing page at three collection sizes; `benchmarks/attribution.py` splits a search into the collection read and the Python work on top of it (deserialise, check access, blend the decay score), so the shape of the curve is explained rather than assumed. `docs/performance.md` carries the numbers with the machine that produced them and says plainly what is *not* measured: the embedding model, concurrency, and selective filters. Three decisions in the harness came from getting it wrong first, and both the scripts and the docs record them rather than hiding them: - **One operation, repeated.** The first version rotated the search term between samples, which measures six different operations and reports the spread between them as if it were the spread of one. It produced a table where a search over 5000 cells looked faster than the same search over 1000. Every sample now uses the same query. - **min / median / p95, never an average.** The first table was taken while the same machine was doing other work and showed no growth at all between 1000 and 5000 cells. Re-measuring on a quiet machine, and reporting the minimum beside the distribution, made the curve visible - and the docs say so, because a reader who runs this on a busy machine will otherwise get a different and wrong shape. - **No timing assertions in CI.** A shared runner cannot produce a number worth publishing, so CI runs the harness in a smoke mode (`--sizes 20 --queries 2`) to prove it works and asserts nothing about how long anything took. Without that, a broken harness rots quietly until somebody needs it. The headline numbers, measured on this machine (Windows 11, 12 CPUs, Python 3.12): a by-id read stays flat (5.6 -> 6.9 ms from 100 to 5000 cells); search grows with the store (4.0 -> 29.1 -> 163.8 ms minimum); the listing page grows fastest of the three (3.6 -> 388 ms) because it deserialises every cell in the collection to keep twenty; and ingest cost per cell rises with the store (244 -> 47 cells/s). Postgres numbers are welcome by pull request against a service container - this machine has no PostgreSQL, which is why the tables are Chroma-only and say so. --- .github/workflows/ci.yml | 10 +- CHANGELOG.md | 11 ++ README.md | 5 + benchmarks/README.md | 41 ++++++ benchmarks/attribution.py | 76 +++++++++++ benchmarks/run.py | 276 ++++++++++++++++++++++++++++++++++++++ docs/performance.md | 113 ++++++++++++++++ mkdocs.yml | 1 + ruff.toml | 3 +- 9 files changed, 533 insertions(+), 3 deletions(-) create mode 100644 benchmarks/README.md create mode 100644 benchmarks/attribution.py create mode 100644 benchmarks/run.py create mode 100644 docs/performance.md diff --git a/.github/workflows/ci.yml b/.github/workflows/ci.yml index ece8d52..2448042 100644 --- a/.github/workflows/ci.yml +++ b/.github/workflows/ci.yml @@ -25,9 +25,15 @@ jobs: # to install the same extra for the same reason. run: pip install -e "server[dev,postgres]" -e "sdk[dev,langchain]" -e "conformance[dev]" - name: Ruff lint (every Python package) - run: ruff check server sdk conformance + run: ruff check server sdk conformance benchmarks - name: Ruff format check (every Python package) - run: ruff format --check server sdk conformance + run: ruff format --check server sdk conformance benchmarks + + # The benchmark harness is not a gate on timing - a shared runner cannot + # produce a meaningful number - but it is a gate on the harness working. + # Without this a broken script rots quietly until somebody needs it. + - name: Benchmark harness runs + run: python benchmarks/run.py --sizes 20 --queries 2 - name: mypy (server) working-directory: server run: mypy diff --git a/CHANGELOG.md b/CHANGELOG.md index af32b6b..c6b2754 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -274,6 +274,17 @@ This project follows [Semantic Versioning](https://semver.org/spec/v2.0.0.html). failure, and does not run it on the 3.10 leg because the reference server's own floor is 3.11. +- **Performance is measured and published** (`benchmarks/run.py`, `docs/performance.md`). + Both backends rank the whole candidate set before cutting a page, so search latency + grows with how much is stored rather than with how much was asked for - a deliberate + trade for decay-weighted ranking, and one the docs did not state. The harness + measures ingest, search, by-id read and listing page at three collection sizes, + records the machine it ran on (absolute numbers mean nothing without it), and uses + a stub embedding so a run is offline and the numbers describe the store rather than + the model - which they explicitly exclude. The page also says what is not measured: + the embedding model, concurrency, and selective filters. CI runs the harness in a + smoke mode, so a broken script cannot rot quietly. + ### Changed - **Every endpoint returns one error shape.** `PATCH /memories/{id}` answered a conflict with `{"detail": ...}` while `DELETE` answered with diff --git a/README.md b/README.md index 3047637..e6b7c7b 100644 --- a/README.md +++ b/README.md @@ -151,6 +151,7 @@ simply absent from its results. - **[Documentation site](https://glatinone.github.io/agent-memory-protocol/)** — everything below, rendered - **[Getting started](docs/getting-started.md)** — server, SDK, embedding provider, storage backend, API keys - **[API reference](docs/api-reference.md)** — every endpoint, plus the error contract +- **[Performance](docs/performance.md)** — measured storage and search numbers, and the two limits behind them - **[Spec, explained](docs/spec-explained.md)** — the schema and the decay formula without the formal notation - **[FAQ](docs/faq.md)** — including [how decay works in plain English](docs/faq.md#how-does-decay-work-in-plain-english) - **[Release notes](docs/release-notes-v0.1.0.md)** — what is in this version, and what is not @@ -177,6 +178,10 @@ Worth reading before you build on this: run one per agent if you want per-agent rules to mean anything. - **The decay pass and the scoring-edit budget are per process.** Two servers over one database keep two of each. +- **Search cost grows with collection size, not page size.** Both backends rank the + whole candidate set so a decay-weighted re-rank has something to re-rank, and the + access filter cannot be pushed into either store. [Measured numbers and how to + reproduce them](docs/performance.md). - **PostgreSQL support is proven in CI, not on every machine.** The machine this was developed on has no PostgreSQL, so the storage contract suite runs against a real `pgvector` container in CI and skips locally. diff --git a/benchmarks/README.md b/benchmarks/README.md new file mode 100644 index 0000000..f4df5e2 --- /dev/null +++ b/benchmarks/README.md @@ -0,0 +1,41 @@ +# benchmarks/ + +Two scripts, both reporting to `docs/performance.md`. Neither is a CI gate on +timing: a shared runner cannot produce a number worth publishing, so CI only checks +that the harness runs. + +## `run.py` + +Measures ingest, search, by-id read and a listing page at several collection sizes. + +```bash +cd server && pip install -e . # the harness imports the package +cd .. +python benchmarks/run.py # chroma, 100/1000/5000 +python benchmarks/run.py --backend postgres --dsn postgresql://user:pw@host/db +python benchmarks/run.py --sizes 200 --queries 5 # a quick run +``` + +Three things about how it measures, each of which was a correction worth keeping: + +- **One operation, repeated.** Every search sample uses the same query. An earlier + version rotated through query words, which measures several different operations + and reports the spread between them as if it were the spread of one - it produced a + table where a search over 5000 cells looked faster than the same search over 1000. +- **min / median / p95, not an average.** The minimum is the least-contended run; the + distribution shows what a real deployment will also see. The first published table + was taken while the machine was doing other work and showed no growth at all. +- **The machine is recorded in the output.** Absolute numbers mean nothing without it, + and the outputs are meant to be quoted. + +The embedding provider is a deterministic stub, so a run needs no model download and +the numbers describe the store and the ranking path - not the embedding model, which +in a real deployment dominates latency. + +## `attribution.py` + +Splits a search into the two things it does - the collection read, and the Python work +on top of it (deserialise, filter by access, blend the decay score) - and times the +listing path beside them. It exists so the shape of the curve above is explained +rather than assumed, and it is the tool to reach for when a number moves and you want +to know which part moved. diff --git a/benchmarks/attribution.py b/benchmarks/attribution.py new file mode 100644 index 0000000..1b99b45 --- /dev/null +++ b/benchmarks/attribution.py @@ -0,0 +1,76 @@ +"""Ad-hoc cost attribution: which part of a search grows with collection size. + +Not part of the published harness. It exists to answer one question the published +numbers raised - search latency barely moved from 1000 to 5000 cells while the +listing path grew linearly - so the explanation in docs/performance.md is measured +rather than guessed. +""" + +import asyncio +import sys +import time +from collections.abc import Awaitable, Callable +from pathlib import Path + +REPO = Path(__file__).resolve().parent +sys.path.insert(0, str(REPO / "server")) +sys.path.insert(0, str(REPO / "benchmarks")) + +from amp_server.models import SearchRequest # noqa: E402 +from amp_server.storage.chroma import ChromaAdapter # noqa: E402 +from run import AGENT, OWNER, StubEmbedding, build_cell # noqa: E402 + + +def best_sync(fn: Callable[[], object], repeat: int = 7) -> float: + times = [] + for _ in range(repeat): + started = time.perf_counter() + fn() + times.append((time.perf_counter() - started) * 1000) + return min(times) + + +async def best_async(fn: Callable[[], Awaitable[object]], repeat: int = 7) -> float: + times = [] + for _ in range(repeat): + started = time.perf_counter() + await fn() + times.append((time.perf_counter() - started) * 1000) + return min(times) + + +async def probe(size: int) -> None: + provider = StubEmbedding() + storage = ChromaAdapter( + collection_name=f"cost_probe_{size}", embedding_provider=provider + ) + for index in range(size): + await storage.save(build_cell(index)) + + request = SearchRequest(query="invoice", owner_id=OWNER, limit=10) + embedded = storage._embed(["invoice"]) + + def chroma_fetch() -> None: + """The collection read the adapter does, before any Python work on it.""" + storage._collection.query( + query_embeddings=embedded, + n_results=size, + include=["metadatas", "distances"], + ) + + chroma_ms = best_sync(chroma_fetch) + search_ms = await best_async(lambda: storage.search(request, agent_id=AGENT)) + list_ms = await best_async( + lambda: storage.query( + owner_id=OWNER, types=None, status=None, limit=20, offset=0 + ) + ) + + print( + f"N={size:>5} chroma fetch {chroma_ms:7.1f} ms | " + f"adapter search {search_ms:7.1f} ms | listing page {list_ms:7.1f} ms" + ) + + +for size in (100, 1000, 5000): + asyncio.run(probe(size)) diff --git a/benchmarks/run.py b/benchmarks/run.py new file mode 100644 index 0000000..758e920 --- /dev/null +++ b/benchmarks/run.py @@ -0,0 +1,276 @@ +#!/usr/bin/env python3 +"""Measure the storage paths, so the limits can be stated instead of implied. + +Both backends rank the whole candidate set before slicing a page: Chroma reads the +collection, Postgres selects every matching row. That is a deliberate trade - a +decay-weighted re-rank can only lift a fresher cell above a stale one if the cell +was fetched - and it means search cost grows with collection size. This script +measures that curve instead of leaving a reader to guess. + +What it measures, and what it deliberately does not: + +- **Storage and ranking only.** The embedding provider is a deterministic stub, so + a run needs no model download and the numbers describe the store, not the model. + A real deployment's latency includes the embedding model; add it separately. +- **Single process, warm store, one machine.** Absolute numbers are only meaningful + next to the machine they came from, which the output header records. + +Usage: + + python benchmarks/run.py # chroma, sizes 100/1000/5000 + python benchmarks/run.py --backend postgres --dsn postgresql://... + python benchmarks/run.py --sizes 200 --queries 5 --out docs/performance.md --append +""" + +from __future__ import annotations + +import argparse +import asyncio +import os +import platform +import statistics +import sys +import time +from collections.abc import Sequence +from dataclasses import dataclass +from datetime import UTC, datetime +from pathlib import Path + +REPO_ROOT = Path(__file__).resolve().parent.parent +sys.path.insert(0, str(REPO_ROOT / "server")) + +from amp_server.embeddings import EmbeddingProvider # noqa: E402 +from amp_server.models import SearchRequest # noqa: E402 +from amp_server.storage.chroma import ChromaAdapter # noqa: E402 + +OWNER = "user-bench" +AGENT = "agent-bench" +#: Words the stub embeds on, so a query has a real vector and real distances. +WORDS = ("invoice", "email", "python", "meeting", "deadline", "budget") + + +class StubEmbedding(EmbeddingProvider): + """A bag of words, no model. Keeps a run offline and comparable across machines.""" + + name = "benchmark-stub" + dimensions = len(WORDS) + + def embed(self, texts: Sequence[str]) -> list[list[float]]: + return [ + [float(text.lower().count(word)) for word in WORDS] or [1.0] + [0.0] * 5 + for text in texts + ] + + +@dataclass +class Timing: + """Latencies in milliseconds, and how many samples they came from. + + `best` is reported next to the distribution because a benchmark shares its + machine with whatever else is running: the first attempt at this table showed + a search that did not grow between 1000 and 5000 cells, until the same + measurement with the machine quiet showed 28 ms and 67 ms. The minimum is the + least-contended run, the median and p95 show the spread that a real deployment + will also see. + """ + + best: float + p50: float + p95: float + samples: int + + def as_cells(self) -> str: + return f"{self.best:.1f} / {self.p50:.1f} / {self.p95:.1f}" + + +def _percentile(values: list[float], fraction: float) -> float: + ordered = sorted(values) + index = min(len(ordered) - 1, int(len(ordered) * fraction)) + return ordered[index] + + +async def _time(call, samples: int) -> Timing: + """Run `call` `samples` times and report the distribution. + + The first call is discarded: it pays for lazily-built indexes and a cold + connection, which is a property of the run and not of the store. + """ + timings: list[float] = [] + for index in range(samples + 1): + started = time.perf_counter() + await call(index) + elapsed = (time.perf_counter() - started) * 1000 + if index: + timings.append(elapsed) + return Timing( + best=min(timings), + p50=statistics.median(timings), + p95=_percentile(timings, 0.95), + samples=samples, + ) + + +def build_cell(index: int): + from amp_server.models import ( + LifecycleStatus, + MemoryAccessPolicy, + MemoryCell, + MemoryContent, + MemoryIdentity, + MemoryLifecycle, + MemoryScoring, + MemoryType, + OwnerType, + ) + + word = WORDS[index % len(WORDS)] + return MemoryCell( + type=MemoryType.SEMANTIC, + content=MemoryContent(text=f"{word} note number {index}"), + identity=MemoryIdentity( + owner_id=OWNER, owner_type=OwnerType.USER, created_by=AGENT + ), + lifecycle=MemoryLifecycle( + created_at=datetime.now(UTC), status=LifecycleStatus.ACTIVE + ), + scoring=MemoryScoring(), + access_policy=MemoryAccessPolicy(readable_by=[AGENT]), + ) + + +async def measure(storage, size: int, queries: int, query: str) -> dict[str, object]: + """Ingest `size` cells, then time search, get and a first page of listing.""" + ingest_started = time.perf_counter() + ids: list[str] = [] + for index in range(size): + ids.append(await storage.save(build_cell(index))) + ingest_seconds = time.perf_counter() - ingest_started + + def search_call(index: int): + # The same query every sample. Rotating through words measures six different + # operations and reports the spread between them as if it were the spread of + # one - which produced a table where a search over 5000 cells looked faster + # than the same search over 1000. + request = SearchRequest(query=query, owner_id=OWNER, limit=10) + return storage.search(request, agent_id=AGENT) + + def get_call(index: int): + return storage.get(ids[index % len(ids)]) + + def list_call(index: int): + return storage.query( + owner_id=OWNER, types=None, status=None, limit=20, offset=0 + ) + + return { + "cells": size, + "ingest_per_second": size / ingest_seconds if ingest_seconds else float("inf"), + "search": await _time(search_call, queries), + "get": await _time(get_call, queries), + "list_page": await _time(list_call, queries), + } + + +async def run(args: argparse.Namespace) -> list[dict[str, object]]: + results: list[dict[str, object]] = [] + for size in args.sizes: + if args.backend == "postgres": + from amp_server.storage.postgres import PostgresAdapter + + dsn = args.dsn or os.environ.get("AMP_TEST_POSTGRES_DSN") + if not dsn: + raise SystemExit( + "postgres needs --dsn or AMP_TEST_POSTGRES_DSN; this is the same " + "variable the storage tests use" + ) + table = f"amp_bench_{int(time.time())}" + storage = PostgresAdapter( + dsn=dsn, table=table, embedding_provider=StubEmbedding() + ) + try: + results.append(await measure(storage, size, args.queries, args.query)) + finally: + with storage._connection.cursor() as cursor: + cursor.execute(f"DROP TABLE IF EXISTS {table}") + storage.close() + continue + + storage = ChromaAdapter( + collection_name=f"bench_{int(time.time())}_{size}", + embedding_provider=StubEmbedding(), + ) + results.append(await measure(storage, size, args.queries, args.query)) + return results + + +def render(results: list[dict[str, object]], args: argparse.Namespace) -> str: + """A markdown table, a header naming the machine, and the two known limits.""" + lines = [ + f"### {args.backend}", + "", + f"- Machine: {platform.system()} {platform.release()}, " + f"{platform.machine()}, {os.cpu_count()} CPUs, " + f"Python {platform.python_version()}", + f"- Embeddings: {StubEmbedding().name} " + "(a stub, so these numbers exclude the model)", + f"- Queries per measurement: {args.queries} of the same query " + f"({args.query!r}), first call discarded", + "- Columns are min / median / p95 over those samples: the minimum is the " + "least-contended run", + f"- Run: {datetime.now(UTC):%Y-%m-%d %H:%M} UTC", + "", + "| Cells | Ingest (cells/s) | Search min / p50 / p95 (ms) | " + "GET min / p50 / p95 (ms) | List page min / p50 / p95 (ms) |", + "|---|---|---|---|---|", + ] + for row in results: + search, get, listing = row["search"], row["get"], row["list_page"] + lines.append( + f"| {row['cells']} | {row['ingest_per_second']:.0f} | " + f"{search.as_cells()} | {get.as_cells()} | {listing.as_cells()} |" + ) + return "\n".join(lines) + + +def main() -> int: + parser = argparse.ArgumentParser(description=__doc__.splitlines()[0]) + parser.add_argument("--backend", choices=("chroma", "postgres"), default="chroma") + parser.add_argument( + "--dsn", default=None, help="postgres DSN (or AMP_TEST_POSTGRES_DSN)" + ) + parser.add_argument( + "--sizes", + type=int, + nargs="+", + default=[100, 1000, 5000], + help="collection sizes to measure, in order", + ) + parser.add_argument( + "--queries", type=int, default=20, help="samples per measurement" + ) + parser.add_argument( + "--query", + default=WORDS[0], + help="the search term every sample uses, so one operation is measured", + ) + parser.add_argument("--out", type=Path, default=None, help="write the table here") + parser.add_argument( + "--append", + action="store_true", + help="append to --out instead of overwriting it", + ) + args = parser.parse_args() + + table = render(asyncio.run(run(args)), args) + print(table) + if args.out: + existing = args.out.read_text(encoding="utf-8") if args.out.exists() else "" + separator = "\n\n" if existing and args.append else "" + body = (existing + separator) if args.append else "" + args.out.write_text(body + table + "\n", encoding="utf-8") + print(f"\nwrote {args.out}") + return 0 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/docs/performance.md b/docs/performance.md new file mode 100644 index 0000000..d22c0d9 --- /dev/null +++ b/docs/performance.md @@ -0,0 +1,113 @@ +# Performance + +Two limits shape every number on this page, and both are deliberate: + +- **Search ranks the whole candidate set before cutting a page.** Chroma reads the + collection; Postgres selects every row matching the filters. A decay-weighted + re-rank can only lift a fresher cell above a stale one if that cell was actually + fetched, so the ranking cannot happen inside the index. The cost is a search whose + latency grows with how much is stored, not with how much you asked for. +- **The access filter cannot be pushed into either store.** `readable_by` holds + patterns rather than ids, so every candidate is checked in Python. That is also why + `limit` on the listing and search endpoints counts cells the caller may read rather + than cells examined. + +Both are the price of ranking by relevance and enforcing per-cell access, and both +are the reason a large collection wants the Postgres backend and a real deployment +wants a cache in front. A vector store with a narrower job (top-k similarity only, +no per-cell policy) will beat this on search latency; that is a different feature set, +not a faster version of this one. + +## Measured numbers + +`benchmarks/run.py` produces these. It uses a deterministic stub embedding, so a run +needs no model download and the numbers describe the store and the ranking path - +**not** the embedding model, which in a real deployment dominates latency. + +### chroma + +- Machine: Windows 11, AMD64, 12 CPUs, Python 3.12.10 +- Embeddings: benchmark-stub (a stub, so these numbers exclude the model) +- Queries per measurement: 20 of the same query ('invoice'), first call discarded +- Columns are min / median / p95 over those samples: the minimum is the least-contended run +- Run: 2026-10-04 09:15 UTC + +| Cells | Ingest (cells/s) | Search min / p50 / p95 (ms) | GET min / p50 / p95 (ms) | List page min / p50 / p95 (ms) | +|---|---|---|---|---| +| 100 | 244 | 4.0 / 4.7 / 6.8 | 5.6 / 6.8 / 10.1 | 3.6 / 5.7 / 31.2 | +| 1000 | 60 | 29.1 / 36.5 / 68.0 | 2.9 / 3.5 / 5.7 | 21.4 / 24.1 / 50.6 | +| 5000 | 47 | 163.8 / 213.1 / 252.2 | 6.9 / 10.4 / 22.4 | 388.3 / 493.4 / 619.2 | + +## Which part grows + +`benchmarks/attribution.py` splits a search into the two things it does, so the curve +above is explained rather than assumed. Same machine, best of seven runs per size: + +| Cells | Collection read (ms) | Full search (ms) | Listing page (ms) | search / read | list / read | +|---|---|---|---|---|---| +| 100 | 2.8 | 6.1 | 4.0 | 2.2x | 1.4x | +| 1000 | 12.8 | 28.3 | 27.0 | 2.2x | 2.1x | +| 5000 | 27.4 | 66.7 | 145.9 | 2.4x | 5.3x | + +Read the ratios, not the milliseconds: these come from a separate run, and the same +search that took 66.7 ms here took 163.8 ms in the run above. A laptop is not a lab, +which is the reason the harness reports a minimum and the reason CI asserts nothing +about timing. + +The ratio is the finding, and it holds across all three sizes: the Python work on top +of the read (deserialise every candidate, apply the access filter, blend the decay +score) costs about as much again as the read itself - and the listing path, which +deserialises every cell to keep twenty, grows faster than either. + +## A note on how these were taken + +The first version of this table was wrong in a way worth recording: it showed search +latency flat between 1000 and 5000 cells, because the measurement ran while the same +machine was doing other work. Re-measuring on a quiet machine showed the growth. That +is why the table reports **min** next to the median and p95 - the minimum is the +least-contended run - and why a benchmark is not a CI gate here: CI runs +`benchmarks/run.py` in a smoke mode to prove the harness works and asserts nothing +about timing, because a shared runner cannot produce a number worth publishing. + +## How to read them + +- **Ingest cost per cell rises with the store** (244 cells/s at 100, 47 at 5000), + because each cell is embedded and its vector indexed on write, and the index gets + denser. Bulk loading is the case to batch if you ever need it. +- **Search latency tracks collection size, not page size.** Asking for 10 results + from 5000 cells costs what asking for 50 costs, because both rank everything. The + spread widens with size too: 4-7 ms at 100 cells, 164-252 ms at 5000. +- **A by-id read stays flat.** 6.9 ms minimum at 5000 cells against 5.6 at 100. It is + a lookup, not a scan - which is why reading one cell to reset its decay clock is + cheap even on a large store, and why the listing endpoint is the expensive one. +- **The first call in each measurement is discarded.** It pays for lazily-built + indexes and a cold connection, which is a property of the run rather than the + store. + +## Reproduce it + +```bash +cd server && pip install -e . # the benchmark imports the package +cd .. +python benchmarks/run.py # chroma, 100 / 1000 / 5000 cells +python benchmarks/run.py --backend postgres --dsn postgresql://user:pw@host/db +python benchmarks/run.py --sizes 200 --queries 5 # a quick run +python benchmarks/attribution.py # fetch vs rank, and the listing path +``` + +The output header records the machine, because absolute numbers only mean something +next to the hardware they came from. Postgres numbers are welcome in a pull request +with that header attached - this repository's development machine has no PostgreSQL, +which is why the CI job for that backend uses a service container and why the tables +above are Chroma-only until somebody runs the other one. + +## What is not measured + +- **The embedding model.** Add it separately; on a laptop the local model costs tens + of milliseconds per query and dominates everything above. +- **Concurrency.** These are single-process, sequential measurements on a warm store. + Under load the picture also depends on the deployment, and the scoring-edit budget + and decay pass are per process. +- **Filtered searches with a selective `owner_id`.** A filter that matches few cells + is cheaper in Postgres (the where-clause narrows the scan) and no cheaper in Chroma + (the collection is read either way). diff --git a/mkdocs.yml b/mkdocs.yml index c509c83..95c4ac9 100644 --- a/mkdocs.yml +++ b/mkdocs.yml @@ -37,4 +37,5 @@ nav: - Getting Started: "getting-started.md" - Protocol Specification: "spec-explained.md" - API Reference: "api-reference.md" + - Performance: "performance.md" - FAQ: "faq.md" diff --git a/ruff.toml b/ruff.toml index 4349323..2c6d9e4 100644 --- a/ruff.toml +++ b/ruff.toml @@ -1,5 +1,6 @@ # Repo-root Ruff configuration, used for Python that the two packages do not -# own (the demo scripts under examples/). server/ and sdk/ each carry their own +# own (the demo scripts under examples/, the harness under benchmarks/). server/ +# and sdk/ each carry their own # [tool.ruff] block in their pyproject.toml, and Ruff resolves configuration per # file, so this file governs everything they do not. #