Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
10 changes: 8 additions & 2 deletions .github/workflows/ci.yml
Original file line number Diff line number Diff line change
Expand Up @@ -25,9 +25,15 @@ jobs:
# to install the same extra for the same reason.
run: pip install -e "server[dev,postgres]" -e "sdk[dev,langchain]" -e "conformance[dev]"
- name: Ruff lint (every Python package)
run: ruff check server sdk conformance
run: ruff check server sdk conformance benchmarks
- name: Ruff format check (every Python package)
run: ruff format --check server sdk conformance
run: ruff format --check server sdk conformance benchmarks

# The benchmark harness is not a gate on timing - a shared runner cannot
# produce a meaningful number - but it is a gate on the harness working.
# Without this a broken script rots quietly until somebody needs it.
- name: Benchmark harness runs
run: python benchmarks/run.py --sizes 20 --queries 2
- name: mypy (server)
working-directory: server
run: mypy
Expand Down
11 changes: 11 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -274,6 +274,17 @@ This project follows [Semantic Versioning](https://semver.org/spec/v2.0.0.html).
failure, and does not run it on the 3.10 leg because the reference server's own
floor is 3.11.

- **Performance is measured and published** (`benchmarks/run.py`, `docs/performance.md`).
Both backends rank the whole candidate set before cutting a page, so search latency
grows with how much is stored rather than with how much was asked for - a deliberate
trade for decay-weighted ranking, and one the docs did not state. The harness
measures ingest, search, by-id read and listing page at three collection sizes,
records the machine it ran on (absolute numbers mean nothing without it), and uses
a stub embedding so a run is offline and the numbers describe the store rather than
the model - which they explicitly exclude. The page also says what is not measured:
the embedding model, concurrency, and selective filters. CI runs the harness in a
smoke mode, so a broken script cannot rot quietly.

### Changed
- **Every endpoint returns one error shape.** `PATCH /memories/{id}` answered a
conflict with `{"detail": ...}` while `DELETE` answered with
Expand Down
5 changes: 5 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -151,6 +151,7 @@ simply absent from its results.
- **[Documentation site](https://glatinone.github.io/agent-memory-protocol/)** — everything below, rendered
- **[Getting started](docs/getting-started.md)** — server, SDK, embedding provider, storage backend, API keys
- **[API reference](docs/api-reference.md)** — every endpoint, plus the error contract
- **[Performance](docs/performance.md)** — measured storage and search numbers, and the two limits behind them
- **[Spec, explained](docs/spec-explained.md)** — the schema and the decay formula without the formal notation
- **[FAQ](docs/faq.md)** — including [how decay works in plain English](docs/faq.md#how-does-decay-work-in-plain-english)
- **[Release notes](docs/release-notes-v0.1.0.md)** — what is in this version, and what is not
Expand All @@ -177,6 +178,10 @@ Worth reading before you build on this:
run one per agent if you want per-agent rules to mean anything.
- **The decay pass and the scoring-edit budget are per process.** Two servers over
one database keep two of each.
- **Search cost grows with collection size, not page size.** Both backends rank the
whole candidate set so a decay-weighted re-rank has something to re-rank, and the
access filter cannot be pushed into either store. [Measured numbers and how to
reproduce them](docs/performance.md).
- **PostgreSQL support is proven in CI, not on every machine.** The machine this was
developed on has no PostgreSQL, so the storage contract suite runs against a real
`pgvector` container in CI and skips locally.
Expand Down
41 changes: 41 additions & 0 deletions benchmarks/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,41 @@
# benchmarks/

Two scripts, both reporting to `docs/performance.md`. Neither is a CI gate on
timing: a shared runner cannot produce a number worth publishing, so CI only checks
that the harness runs.

## `run.py`

Measures ingest, search, by-id read and a listing page at several collection sizes.

```bash
cd server && pip install -e . # the harness imports the package
cd ..
python benchmarks/run.py # chroma, 100/1000/5000
python benchmarks/run.py --backend postgres --dsn postgresql://user:pw@host/db
python benchmarks/run.py --sizes 200 --queries 5 # a quick run
```

Three things about how it measures, each of which was a correction worth keeping:

- **One operation, repeated.** Every search sample uses the same query. An earlier
version rotated through query words, which measures several different operations
and reports the spread between them as if it were the spread of one - it produced a
table where a search over 5000 cells looked faster than the same search over 1000.
- **min / median / p95, not an average.** The minimum is the least-contended run; the
distribution shows what a real deployment will also see. The first published table
was taken while the machine was doing other work and showed no growth at all.
- **The machine is recorded in the output.** Absolute numbers mean nothing without it,
and the outputs are meant to be quoted.

The embedding provider is a deterministic stub, so a run needs no model download and
the numbers describe the store and the ranking path - not the embedding model, which
in a real deployment dominates latency.

## `attribution.py`

Splits a search into the two things it does - the collection read, and the Python work
on top of it (deserialise, filter by access, blend the decay score) - and times the
listing path beside them. It exists so the shape of the curve above is explained
rather than assumed, and it is the tool to reach for when a number moves and you want
to know which part moved.
76 changes: 76 additions & 0 deletions benchmarks/attribution.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,76 @@
"""Ad-hoc cost attribution: which part of a search grows with collection size.

Not part of the published harness. It exists to answer one question the published
numbers raised - search latency barely moved from 1000 to 5000 cells while the
listing path grew linearly - so the explanation in docs/performance.md is measured
rather than guessed.
"""

import asyncio
import sys
import time
from collections.abc import Awaitable, Callable
from pathlib import Path

REPO = Path(__file__).resolve().parent
sys.path.insert(0, str(REPO / "server"))
sys.path.insert(0, str(REPO / "benchmarks"))

from amp_server.models import SearchRequest # noqa: E402
from amp_server.storage.chroma import ChromaAdapter # noqa: E402
from run import AGENT, OWNER, StubEmbedding, build_cell # noqa: E402


def best_sync(fn: Callable[[], object], repeat: int = 7) -> float:
times = []
for _ in range(repeat):
started = time.perf_counter()
fn()
times.append((time.perf_counter() - started) * 1000)
return min(times)


async def best_async(fn: Callable[[], Awaitable[object]], repeat: int = 7) -> float:
times = []
for _ in range(repeat):
started = time.perf_counter()
await fn()
times.append((time.perf_counter() - started) * 1000)
return min(times)


async def probe(size: int) -> None:
provider = StubEmbedding()
storage = ChromaAdapter(
collection_name=f"cost_probe_{size}", embedding_provider=provider
)
for index in range(size):
await storage.save(build_cell(index))

request = SearchRequest(query="invoice", owner_id=OWNER, limit=10)
embedded = storage._embed(["invoice"])

def chroma_fetch() -> None:
"""The collection read the adapter does, before any Python work on it."""
storage._collection.query(
query_embeddings=embedded,
n_results=size,
include=["metadatas", "distances"],
)

chroma_ms = best_sync(chroma_fetch)
search_ms = await best_async(lambda: storage.search(request, agent_id=AGENT))
list_ms = await best_async(
lambda: storage.query(
owner_id=OWNER, types=None, status=None, limit=20, offset=0
)
)

print(
f"N={size:>5} chroma fetch {chroma_ms:7.1f} ms | "
f"adapter search {search_ms:7.1f} ms | listing page {list_ms:7.1f} ms"
)


for size in (100, 1000, 5000):
asyncio.run(probe(size))
Loading
Loading