diff --git a/.github/workflows/ci.yml b/.github/workflows/ci.yml index ece8d52..2448042 100644 --- a/.github/workflows/ci.yml +++ b/.github/workflows/ci.yml @@ -25,9 +25,15 @@ jobs: # to install the same extra for the same reason. run: pip install -e "server[dev,postgres]" -e "sdk[dev,langchain]" -e "conformance[dev]" - name: Ruff lint (every Python package) - run: ruff check server sdk conformance + run: ruff check server sdk conformance benchmarks - name: Ruff format check (every Python package) - run: ruff format --check server sdk conformance + run: ruff format --check server sdk conformance benchmarks + + # The benchmark harness is not a gate on timing - a shared runner cannot + # produce a meaningful number - but it is a gate on the harness working. + # Without this a broken script rots quietly until somebody needs it. + - name: Benchmark harness runs + run: python benchmarks/run.py --sizes 20 --queries 2 - name: mypy (server) working-directory: server run: mypy diff --git a/CHANGELOG.md b/CHANGELOG.md index af32b6b..c6b2754 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -274,6 +274,17 @@ This project follows [Semantic Versioning](https://semver.org/spec/v2.0.0.html). failure, and does not run it on the 3.10 leg because the reference server's own floor is 3.11. +- **Performance is measured and published** (`benchmarks/run.py`, `docs/performance.md`). + Both backends rank the whole candidate set before cutting a page, so search latency + grows with how much is stored rather than with how much was asked for - a deliberate + trade for decay-weighted ranking, and one the docs did not state. The harness + measures ingest, search, by-id read and listing page at three collection sizes, + records the machine it ran on (absolute numbers mean nothing without it), and uses + a stub embedding so a run is offline and the numbers describe the store rather than + the model - which they explicitly exclude. The page also says what is not measured: + the embedding model, concurrency, and selective filters. CI runs the harness in a + smoke mode, so a broken script cannot rot quietly. + ### Changed - **Every endpoint returns one error shape.** `PATCH /memories/{id}` answered a conflict with `{"detail": ...}` while `DELETE` answered with diff --git a/README.md b/README.md index 3047637..e6b7c7b 100644 --- a/README.md +++ b/README.md @@ -151,6 +151,7 @@ simply absent from its results. - **[Documentation site](https://glatinone.github.io/agent-memory-protocol/)** — everything below, rendered - **[Getting started](docs/getting-started.md)** — server, SDK, embedding provider, storage backend, API keys - **[API reference](docs/api-reference.md)** — every endpoint, plus the error contract +- **[Performance](docs/performance.md)** — measured storage and search numbers, and the two limits behind them - **[Spec, explained](docs/spec-explained.md)** — the schema and the decay formula without the formal notation - **[FAQ](docs/faq.md)** — including [how decay works in plain English](docs/faq.md#how-does-decay-work-in-plain-english) - **[Release notes](docs/release-notes-v0.1.0.md)** — what is in this version, and what is not @@ -177,6 +178,10 @@ Worth reading before you build on this: run one per agent if you want per-agent rules to mean anything. - **The decay pass and the scoring-edit budget are per process.** Two servers over one database keep two of each. +- **Search cost grows with collection size, not page size.** Both backends rank the + whole candidate set so a decay-weighted re-rank has something to re-rank, and the + access filter cannot be pushed into either store. [Measured numbers and how to + reproduce them](docs/performance.md). - **PostgreSQL support is proven in CI, not on every machine.** The machine this was developed on has no PostgreSQL, so the storage contract suite runs against a real `pgvector` container in CI and skips locally. diff --git a/benchmarks/README.md b/benchmarks/README.md new file mode 100644 index 0000000..f4df5e2 --- /dev/null +++ b/benchmarks/README.md @@ -0,0 +1,41 @@ +# benchmarks/ + +Two scripts, both reporting to `docs/performance.md`. Neither is a CI gate on +timing: a shared runner cannot produce a number worth publishing, so CI only checks +that the harness runs. + +## `run.py` + +Measures ingest, search, by-id read and a listing page at several collection sizes. + +```bash +cd server && pip install -e . # the harness imports the package +cd .. +python benchmarks/run.py # chroma, 100/1000/5000 +python benchmarks/run.py --backend postgres --dsn postgresql://user:pw@host/db +python benchmarks/run.py --sizes 200 --queries 5 # a quick run +``` + +Three things about how it measures, each of which was a correction worth keeping: + +- **One operation, repeated.** Every search sample uses the same query. An earlier + version rotated through query words, which measures several different operations + and reports the spread between them as if it were the spread of one - it produced a + table where a search over 5000 cells looked faster than the same search over 1000. +- **min / median / p95, not an average.** The minimum is the least-contended run; the + distribution shows what a real deployment will also see. The first published table + was taken while the machine was doing other work and showed no growth at all. +- **The machine is recorded in the output.** Absolute numbers mean nothing without it, + and the outputs are meant to be quoted. + +The embedding provider is a deterministic stub, so a run needs no model download and +the numbers describe the store and the ranking path - not the embedding model, which +in a real deployment dominates latency. + +## `attribution.py` + +Splits a search into the two things it does - the collection read, and the Python work +on top of it (deserialise, filter by access, blend the decay score) - and times the +listing path beside them. It exists so the shape of the curve above is explained +rather than assumed, and it is the tool to reach for when a number moves and you want +to know which part moved. diff --git a/benchmarks/attribution.py b/benchmarks/attribution.py new file mode 100644 index 0000000..1b99b45 --- /dev/null +++ b/benchmarks/attribution.py @@ -0,0 +1,76 @@ +"""Ad-hoc cost attribution: which part of a search grows with collection size. + +Not part of the published harness. It exists to answer one question the published +numbers raised - search latency barely moved from 1000 to 5000 cells while the +listing path grew linearly - so the explanation in docs/performance.md is measured +rather than guessed. +""" + +import asyncio +import sys +import time +from collections.abc import Awaitable, Callable +from pathlib import Path + +REPO = Path(__file__).resolve().parent +sys.path.insert(0, str(REPO / "server")) +sys.path.insert(0, str(REPO / "benchmarks")) + +from amp_server.models import SearchRequest # noqa: E402 +from amp_server.storage.chroma import ChromaAdapter # noqa: E402 +from run import AGENT, OWNER, StubEmbedding, build_cell # noqa: E402 + + +def best_sync(fn: Callable[[], object], repeat: int = 7) -> float: + times = [] + for _ in range(repeat): + started = time.perf_counter() + fn() + times.append((time.perf_counter() - started) * 1000) + return min(times) + + +async def best_async(fn: Callable[[], Awaitable[object]], repeat: int = 7) -> float: + times = [] + for _ in range(repeat): + started = time.perf_counter() + await fn() + times.append((time.perf_counter() - started) * 1000) + return min(times) + + +async def probe(size: int) -> None: + provider = StubEmbedding() + storage = ChromaAdapter( + collection_name=f"cost_probe_{size}", embedding_provider=provider + ) + for index in range(size): + await storage.save(build_cell(index)) + + request = SearchRequest(query="invoice", owner_id=OWNER, limit=10) + embedded = storage._embed(["invoice"]) + + def chroma_fetch() -> None: + """The collection read the adapter does, before any Python work on it.""" + storage._collection.query( + query_embeddings=embedded, + n_results=size, + include=["metadatas", "distances"], + ) + + chroma_ms = best_sync(chroma_fetch) + search_ms = await best_async(lambda: storage.search(request, agent_id=AGENT)) + list_ms = await best_async( + lambda: storage.query( + owner_id=OWNER, types=None, status=None, limit=20, offset=0 + ) + ) + + print( + f"N={size:>5} chroma fetch {chroma_ms:7.1f} ms | " + f"adapter search {search_ms:7.1f} ms | listing page {list_ms:7.1f} ms" + ) + + +for size in (100, 1000, 5000): + asyncio.run(probe(size)) diff --git a/benchmarks/run.py b/benchmarks/run.py new file mode 100644 index 0000000..758e920 --- /dev/null +++ b/benchmarks/run.py @@ -0,0 +1,276 @@ +#!/usr/bin/env python3 +"""Measure the storage paths, so the limits can be stated instead of implied. + +Both backends rank the whole candidate set before slicing a page: Chroma reads the +collection, Postgres selects every matching row. That is a deliberate trade - a +decay-weighted re-rank can only lift a fresher cell above a stale one if the cell +was fetched - and it means search cost grows with collection size. This script +measures that curve instead of leaving a reader to guess. + +What it measures, and what it deliberately does not: + +- **Storage and ranking only.** The embedding provider is a deterministic stub, so + a run needs no model download and the numbers describe the store, not the model. + A real deployment's latency includes the embedding model; add it separately. +- **Single process, warm store, one machine.** Absolute numbers are only meaningful + next to the machine they came from, which the output header records. + +Usage: + + python benchmarks/run.py # chroma, sizes 100/1000/5000 + python benchmarks/run.py --backend postgres --dsn postgresql://... + python benchmarks/run.py --sizes 200 --queries 5 --out docs/performance.md --append +""" + +from __future__ import annotations + +import argparse +import asyncio +import os +import platform +import statistics +import sys +import time +from collections.abc import Sequence +from dataclasses import dataclass +from datetime import UTC, datetime +from pathlib import Path + +REPO_ROOT = Path(__file__).resolve().parent.parent +sys.path.insert(0, str(REPO_ROOT / "server")) + +from amp_server.embeddings import EmbeddingProvider # noqa: E402 +from amp_server.models import SearchRequest # noqa: E402 +from amp_server.storage.chroma import ChromaAdapter # noqa: E402 + +OWNER = "user-bench" +AGENT = "agent-bench" +#: Words the stub embeds on, so a query has a real vector and real distances. +WORDS = ("invoice", "email", "python", "meeting", "deadline", "budget") + + +class StubEmbedding(EmbeddingProvider): + """A bag of words, no model. Keeps a run offline and comparable across machines.""" + + name = "benchmark-stub" + dimensions = len(WORDS) + + def embed(self, texts: Sequence[str]) -> list[list[float]]: + return [ + [float(text.lower().count(word)) for word in WORDS] or [1.0] + [0.0] * 5 + for text in texts + ] + + +@dataclass +class Timing: + """Latencies in milliseconds, and how many samples they came from. + + `best` is reported next to the distribution because a benchmark shares its + machine with whatever else is running: the first attempt at this table showed + a search that did not grow between 1000 and 5000 cells, until the same + measurement with the machine quiet showed 28 ms and 67 ms. The minimum is the + least-contended run, the median and p95 show the spread that a real deployment + will also see. + """ + + best: float + p50: float + p95: float + samples: int + + def as_cells(self) -> str: + return f"{self.best:.1f} / {self.p50:.1f} / {self.p95:.1f}" + + +def _percentile(values: list[float], fraction: float) -> float: + ordered = sorted(values) + index = min(len(ordered) - 1, int(len(ordered) * fraction)) + return ordered[index] + + +async def _time(call, samples: int) -> Timing: + """Run `call` `samples` times and report the distribution. + + The first call is discarded: it pays for lazily-built indexes and a cold + connection, which is a property of the run and not of the store. + """ + timings: list[float] = [] + for index in range(samples + 1): + started = time.perf_counter() + await call(index) + elapsed = (time.perf_counter() - started) * 1000 + if index: + timings.append(elapsed) + return Timing( + best=min(timings), + p50=statistics.median(timings), + p95=_percentile(timings, 0.95), + samples=samples, + ) + + +def build_cell(index: int): + from amp_server.models import ( + LifecycleStatus, + MemoryAccessPolicy, + MemoryCell, + MemoryContent, + MemoryIdentity, + MemoryLifecycle, + MemoryScoring, + MemoryType, + OwnerType, + ) + + word = WORDS[index % len(WORDS)] + return MemoryCell( + type=MemoryType.SEMANTIC, + content=MemoryContent(text=f"{word} note number {index}"), + identity=MemoryIdentity( + owner_id=OWNER, owner_type=OwnerType.USER, created_by=AGENT + ), + lifecycle=MemoryLifecycle( + created_at=datetime.now(UTC), status=LifecycleStatus.ACTIVE + ), + scoring=MemoryScoring(), + access_policy=MemoryAccessPolicy(readable_by=[AGENT]), + ) + + +async def measure(storage, size: int, queries: int, query: str) -> dict[str, object]: + """Ingest `size` cells, then time search, get and a first page of listing.""" + ingest_started = time.perf_counter() + ids: list[str] = [] + for index in range(size): + ids.append(await storage.save(build_cell(index))) + ingest_seconds = time.perf_counter() - ingest_started + + def search_call(index: int): + # The same query every sample. Rotating through words measures six different + # operations and reports the spread between them as if it were the spread of + # one - which produced a table where a search over 5000 cells looked faster + # than the same search over 1000. + request = SearchRequest(query=query, owner_id=OWNER, limit=10) + return storage.search(request, agent_id=AGENT) + + def get_call(index: int): + return storage.get(ids[index % len(ids)]) + + def list_call(index: int): + return storage.query( + owner_id=OWNER, types=None, status=None, limit=20, offset=0 + ) + + return { + "cells": size, + "ingest_per_second": size / ingest_seconds if ingest_seconds else float("inf"), + "search": await _time(search_call, queries), + "get": await _time(get_call, queries), + "list_page": await _time(list_call, queries), + } + + +async def run(args: argparse.Namespace) -> list[dict[str, object]]: + results: list[dict[str, object]] = [] + for size in args.sizes: + if args.backend == "postgres": + from amp_server.storage.postgres import PostgresAdapter + + dsn = args.dsn or os.environ.get("AMP_TEST_POSTGRES_DSN") + if not dsn: + raise SystemExit( + "postgres needs --dsn or AMP_TEST_POSTGRES_DSN; this is the same " + "variable the storage tests use" + ) + table = f"amp_bench_{int(time.time())}" + storage = PostgresAdapter( + dsn=dsn, table=table, embedding_provider=StubEmbedding() + ) + try: + results.append(await measure(storage, size, args.queries, args.query)) + finally: + with storage._connection.cursor() as cursor: + cursor.execute(f"DROP TABLE IF EXISTS {table}") + storage.close() + continue + + storage = ChromaAdapter( + collection_name=f"bench_{int(time.time())}_{size}", + embedding_provider=StubEmbedding(), + ) + results.append(await measure(storage, size, args.queries, args.query)) + return results + + +def render(results: list[dict[str, object]], args: argparse.Namespace) -> str: + """A markdown table, a header naming the machine, and the two known limits.""" + lines = [ + f"### {args.backend}", + "", + f"- Machine: {platform.system()} {platform.release()}, " + f"{platform.machine()}, {os.cpu_count()} CPUs, " + f"Python {platform.python_version()}", + f"- Embeddings: {StubEmbedding().name} " + "(a stub, so these numbers exclude the model)", + f"- Queries per measurement: {args.queries} of the same query " + f"({args.query!r}), first call discarded", + "- Columns are min / median / p95 over those samples: the minimum is the " + "least-contended run", + f"- Run: {datetime.now(UTC):%Y-%m-%d %H:%M} UTC", + "", + "| Cells | Ingest (cells/s) | Search min / p50 / p95 (ms) | " + "GET min / p50 / p95 (ms) | List page min / p50 / p95 (ms) |", + "|---|---|---|---|---|", + ] + for row in results: + search, get, listing = row["search"], row["get"], row["list_page"] + lines.append( + f"| {row['cells']} | {row['ingest_per_second']:.0f} | " + f"{search.as_cells()} | {get.as_cells()} | {listing.as_cells()} |" + ) + return "\n".join(lines) + + +def main() -> int: + parser = argparse.ArgumentParser(description=__doc__.splitlines()[0]) + parser.add_argument("--backend", choices=("chroma", "postgres"), default="chroma") + parser.add_argument( + "--dsn", default=None, help="postgres DSN (or AMP_TEST_POSTGRES_DSN)" + ) + parser.add_argument( + "--sizes", + type=int, + nargs="+", + default=[100, 1000, 5000], + help="collection sizes to measure, in order", + ) + parser.add_argument( + "--queries", type=int, default=20, help="samples per measurement" + ) + parser.add_argument( + "--query", + default=WORDS[0], + help="the search term every sample uses, so one operation is measured", + ) + parser.add_argument("--out", type=Path, default=None, help="write the table here") + parser.add_argument( + "--append", + action="store_true", + help="append to --out instead of overwriting it", + ) + args = parser.parse_args() + + table = render(asyncio.run(run(args)), args) + print(table) + if args.out: + existing = args.out.read_text(encoding="utf-8") if args.out.exists() else "" + separator = "\n\n" if existing and args.append else "" + body = (existing + separator) if args.append else "" + args.out.write_text(body + table + "\n", encoding="utf-8") + print(f"\nwrote {args.out}") + return 0 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/docs/performance.md b/docs/performance.md new file mode 100644 index 0000000..d22c0d9 --- /dev/null +++ b/docs/performance.md @@ -0,0 +1,113 @@ +# Performance + +Two limits shape every number on this page, and both are deliberate: + +- **Search ranks the whole candidate set before cutting a page.** Chroma reads the + collection; Postgres selects every row matching the filters. A decay-weighted + re-rank can only lift a fresher cell above a stale one if that cell was actually + fetched, so the ranking cannot happen inside the index. The cost is a search whose + latency grows with how much is stored, not with how much you asked for. +- **The access filter cannot be pushed into either store.** `readable_by` holds + patterns rather than ids, so every candidate is checked in Python. That is also why + `limit` on the listing and search endpoints counts cells the caller may read rather + than cells examined. + +Both are the price of ranking by relevance and enforcing per-cell access, and both +are the reason a large collection wants the Postgres backend and a real deployment +wants a cache in front. A vector store with a narrower job (top-k similarity only, +no per-cell policy) will beat this on search latency; that is a different feature set, +not a faster version of this one. + +## Measured numbers + +`benchmarks/run.py` produces these. It uses a deterministic stub embedding, so a run +needs no model download and the numbers describe the store and the ranking path - +**not** the embedding model, which in a real deployment dominates latency. + +### chroma + +- Machine: Windows 11, AMD64, 12 CPUs, Python 3.12.10 +- Embeddings: benchmark-stub (a stub, so these numbers exclude the model) +- Queries per measurement: 20 of the same query ('invoice'), first call discarded +- Columns are min / median / p95 over those samples: the minimum is the least-contended run +- Run: 2026-10-04 09:15 UTC + +| Cells | Ingest (cells/s) | Search min / p50 / p95 (ms) | GET min / p50 / p95 (ms) | List page min / p50 / p95 (ms) | +|---|---|---|---|---| +| 100 | 244 | 4.0 / 4.7 / 6.8 | 5.6 / 6.8 / 10.1 | 3.6 / 5.7 / 31.2 | +| 1000 | 60 | 29.1 / 36.5 / 68.0 | 2.9 / 3.5 / 5.7 | 21.4 / 24.1 / 50.6 | +| 5000 | 47 | 163.8 / 213.1 / 252.2 | 6.9 / 10.4 / 22.4 | 388.3 / 493.4 / 619.2 | + +## Which part grows + +`benchmarks/attribution.py` splits a search into the two things it does, so the curve +above is explained rather than assumed. Same machine, best of seven runs per size: + +| Cells | Collection read (ms) | Full search (ms) | Listing page (ms) | search / read | list / read | +|---|---|---|---|---|---| +| 100 | 2.8 | 6.1 | 4.0 | 2.2x | 1.4x | +| 1000 | 12.8 | 28.3 | 27.0 | 2.2x | 2.1x | +| 5000 | 27.4 | 66.7 | 145.9 | 2.4x | 5.3x | + +Read the ratios, not the milliseconds: these come from a separate run, and the same +search that took 66.7 ms here took 163.8 ms in the run above. A laptop is not a lab, +which is the reason the harness reports a minimum and the reason CI asserts nothing +about timing. + +The ratio is the finding, and it holds across all three sizes: the Python work on top +of the read (deserialise every candidate, apply the access filter, blend the decay +score) costs about as much again as the read itself - and the listing path, which +deserialises every cell to keep twenty, grows faster than either. + +## A note on how these were taken + +The first version of this table was wrong in a way worth recording: it showed search +latency flat between 1000 and 5000 cells, because the measurement ran while the same +machine was doing other work. Re-measuring on a quiet machine showed the growth. That +is why the table reports **min** next to the median and p95 - the minimum is the +least-contended run - and why a benchmark is not a CI gate here: CI runs +`benchmarks/run.py` in a smoke mode to prove the harness works and asserts nothing +about timing, because a shared runner cannot produce a number worth publishing. + +## How to read them + +- **Ingest cost per cell rises with the store** (244 cells/s at 100, 47 at 5000), + because each cell is embedded and its vector indexed on write, and the index gets + denser. Bulk loading is the case to batch if you ever need it. +- **Search latency tracks collection size, not page size.** Asking for 10 results + from 5000 cells costs what asking for 50 costs, because both rank everything. The + spread widens with size too: 4-7 ms at 100 cells, 164-252 ms at 5000. +- **A by-id read stays flat.** 6.9 ms minimum at 5000 cells against 5.6 at 100. It is + a lookup, not a scan - which is why reading one cell to reset its decay clock is + cheap even on a large store, and why the listing endpoint is the expensive one. +- **The first call in each measurement is discarded.** It pays for lazily-built + indexes and a cold connection, which is a property of the run rather than the + store. + +## Reproduce it + +```bash +cd server && pip install -e . # the benchmark imports the package +cd .. +python benchmarks/run.py # chroma, 100 / 1000 / 5000 cells +python benchmarks/run.py --backend postgres --dsn postgresql://user:pw@host/db +python benchmarks/run.py --sizes 200 --queries 5 # a quick run +python benchmarks/attribution.py # fetch vs rank, and the listing path +``` + +The output header records the machine, because absolute numbers only mean something +next to the hardware they came from. Postgres numbers are welcome in a pull request +with that header attached - this repository's development machine has no PostgreSQL, +which is why the CI job for that backend uses a service container and why the tables +above are Chroma-only until somebody runs the other one. + +## What is not measured + +- **The embedding model.** Add it separately; on a laptop the local model costs tens + of milliseconds per query and dominates everything above. +- **Concurrency.** These are single-process, sequential measurements on a warm store. + Under load the picture also depends on the deployment, and the scoring-edit budget + and decay pass are per process. +- **Filtered searches with a selective `owner_id`.** A filter that matches few cells + is cheaper in Postgres (the where-clause narrows the scan) and no cheaper in Chroma + (the collection is read either way). diff --git a/mkdocs.yml b/mkdocs.yml index c509c83..95c4ac9 100644 --- a/mkdocs.yml +++ b/mkdocs.yml @@ -37,4 +37,5 @@ nav: - Getting Started: "getting-started.md" - Protocol Specification: "spec-explained.md" - API Reference: "api-reference.md" + - Performance: "performance.md" - FAQ: "faq.md" diff --git a/ruff.toml b/ruff.toml index 4349323..2c6d9e4 100644 --- a/ruff.toml +++ b/ruff.toml @@ -1,5 +1,6 @@ # Repo-root Ruff configuration, used for Python that the two packages do not -# own (the demo scripts under examples/). server/ and sdk/ each carry their own +# own (the demo scripts under examples/, the harness under benchmarks/). server/ +# and sdk/ each carry their own # [tool.ruff] block in their pyproject.toml, and Ruff resolves configuration per # file, so this file governs everything they do not. #