Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions llp/0000-fluiddb.explainer.md
Original file line number Diff line number Diff line change
Expand Up @@ -68,6 +68,7 @@ live here as numbered, living documents, and code points at them with `@ref` com
| 0024 / 0024.000 | RFC / Research | TypeScript replay evaluation protocol and engineering verification |
| 0024.001 / 0024.001.000 | RFC / Research | Swappable SDK components, portable evaluator, and real-provider comparison |
| 0024.002 / 0024.002.000 | RFC / Research | Longer product continuity, complete legacy import and native loading recovery |
| 0024.003 / 0024.003.000 | RFC / Research | Product save latency: controlled adapter comparison, guard phases and real embedding timings |

## Where things live

Expand Down
10 changes: 10 additions & 0 deletions llp/0002-findings.research.md
Original file line number Diff line number Diff line change
Expand Up @@ -365,6 +365,16 @@ and ~0.3 s per call.

## Evaluation

- **Fast retrieval benchmarks do not establish fast product saves.**
- [observed] LLP 0024.003.000: managed Live delegates to a Responses model before a real function call;
the explicit SDK save itself only embeds and atomically writes. Under controlled database delays,
BetterMind's full hydration/repeated checks made host saves grow from 2.3s at 30 facts to 4.3s at 1,000.
Scoped guard preparation and selective detail hydration reduced these to 1.7s and 2.8s (26–36% across
four fixture sizes), with no SDK version change. Statement/metadata scans remain linear.
- Ten short fictional production-adapter embedding requests had median 161ms and maximum 2,602ms, all
on the first network attempt. This local sample does not explain the reported 31–35s incident.
- [high for controlled operation costs; insufficient for production tail latency or proactivity]

- **First-save erasure needs a generation even before any memory exists.**
- [observed] LLP 0023.003: independently reproduced a save reading revision=null, waiting for embeddings,
then committing after another instance erased the empty person. Every adapter now retains a fresh
Expand Down
61 changes: 61 additions & 0 deletions llp/0024.003-save-latency.rfc.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,61 @@
# LLP 0024.003: Measure the product save path

**Type:** RFC
**Status:** Draft
**Systems:** Research, Evaluation, Writes, SDK
**Author:** Codex
**Date:** 2026-09-29
**Related:** LLP 0001, LLP 0023.001, LLP 0024.002

## Question

[confirmed] (Adam Zvada, 2026-09-29): the reported 31–35-second memory saving experience is unacceptable;
investigate the actual XO/tool boundary and experiment with faster saves.

[observed] This document preserves the protocol written before the runs in BetterMind's ignored
`.context/memory-save-benchmark/protocol.md`. That file was written before baseline measurements; its provider
probe section was added before live requests. This tracked copy was added after running, without changing
hypotheses or thresholds. The results belong in LLP 0024.003.000.

## Hypotheses and protocol

- H1: full fact hydration and repeated guard reads add multiple serial database round trips to the host's save,
compared with a direct explicit SDK save through the same product adapter.
- H2: one verified local preparation phase and selective detail hydration reduce controlled host latency by
at least 25%, retaining the same persisted fact, one embedding invocation and one atomic transaction.
- H3: ten short two-input requests to the configured production embedding adapter have median latency below
1.5 seconds and maximum below 5 seconds from the development machine.

H1/H2 use fictional histories of 0, 30, 300 and 1,000 independent untagged SDK saves. Compare the real
BetterMind host helper and direct SDK method over the product's Firestore adapter and local JSON storage.
Inject 100ms per read RPC, batch-get chunks of 100 IDs, four transaction stages at 100ms each, and 500ms per
embedding invocation. Each case starts with its own store. Verify exactly one new fact after the timed save.
Count read RPCs, requested documents/query results, transactions and embedding calls separately. Transaction
reads are excluded from the read counters. No provider credentials or network are used; this part costs zero.

H3 sends ten distinct fictional short notes with their agent-save source labels to `text-embedding-3-small`,
1,536 dimensions, two attempts and 15 seconds per attempt: the existing product configuration. Cache every
request and response through `evals/transport.ts`. Record success, complete call duration and actual network
attempt count. Cached replays cannot supply live latency. Never use account data.

## Budget and deviations

The shared previously approved synthetic-test cap remains $2; this is not a new allowance. Before H3,
reserve $0.01 in its existing ledger, using a deliberately conservative $0.50/M embedding-token rate and the
transport's enforced request reservations. Cache-only is the default; live mode is explicit.

[observed] The first four requests succeeded. The transport's conservative allowance, including 4,096 tokens
of request overhead, stopped sample five before network use. Continue the same samples with a $0.04 reservation
within the same ledger, reusing the four caches and retaining their original live durations. This deviation
was recorded before the continuation. It is a budget stop, not a provider failure. Both run records remain.

## Constraints and limits

The implementation must preserve LLP 0023.001#sdk-boundary and #database-adapters: actual persistence before a
saved receipt, immutable retry identity, truthful source provenance, revision-guarded atomic commits and
erasure/forget fences. Model calls cannot hold a database transaction. The product's guard read cache must
end and revalidate before embedding or a no-write receipt, and cannot survive a retry attempt.

These experiments do not measure Live delegation latency, physical app playback, model proactivity or
production latency percentiles. Ten successful embedding calls cannot diagnose one reported slow incident.
Selective hydration is not an indexed constant-time lookup; statements and metadata are still scanned.
88 changes: 88 additions & 0 deletions llp/0024.003.000-save-latency-results.research.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,88 @@
# LLP 0024.003.000: Host save overhead and embedding latency

**Type:** Research
**Status:** Draft
**Systems:** Research, Evaluation, Writes, SDK
**Author:** Codex
**Date:** 2026-09-29
**Related:** LLP 0024.003, LLP 0023.001, LLP 0002

## What ran

[observed] Protocol: LLP 0024.003. BetterMind baseline `1be2de8f` consumes the unchanged public packages
`@fluiddb/fluiddb@1.0.0-next.6` and `xo-harness@0.3.1`. The candidate changes only the product adapter and
save helper, plus the extraction scheduling bug described below. Reproducer in that repository:
`npm run bench:memory-save -- .context/save-benchmark.json`.

The baseline/candidate results, protocol and cached provider responses live in BetterMind's ignored
`.context/memory-save-benchmark/`. No account data was changed or used in these experiments.

## The configured call path

[observed] BetterMind `src/domain/voice/providers.ts` leaves the installed XO Live adapter's defaults in
place: `gpt-live-1` delegates through managed Responses to `gpt-5.6-luna`, which chooses a function and
arguments. The harness invokes the host's `remember_memory`. FluidDB's explicit save then runs code,
one embedding invocation for the fact/source label, and an atomic write. It does not extract the already
supplied fact with another language model. MCP is not part of this SDK integration.

Historical retrieval/answer-quality experiments did not measure this voice-to-tool-to-Firestore path.
The current embedder's two 15-second attempts plus a one-second timeout retry wait can account for about
31 seconds, but this is a configured failure path, not the measured cause of the user's incident.

## Controlled storage results

All values below include the protocol's artificial 100ms read round trip, 400ms transaction and 500ms
embedding invocation. They are not production percentiles. The direct SDK path uses the same product adapter.

| Existing facts | SDK before / after | Host before / after | Host read RPCs before / after |
| --- | --- | --- | --- |
| 0 | 1,717 / 1,413ms | 2,321 / 1,711ms | 22 / 14 |
| 30 | 1,710 / 1,409ms | 2,318 / 1,713ms | 24 / 14 |
| 300 | 1,709 / 1,406ms | 2,835 / 2,024ms | 31 / 17 |
| 1,000 | 1,710 / 1,408ms | 4,311 / 2,774ms | 52 / 24 |

Every case persisted exactly one new fact with one embedding invocation and one transaction. H1 is supported
by this controlled comparison. H2 meets the predeclared 25% threshold: host reduction is 26–36%.

The product now shares guard reads during local save preparation, revalidates before embedding/no-write
receipts, and releases the cache before the provider and commit. Independent record and guard requests
overlap. Full group/window hydration is limited to retry IDs, fingerprint matches, replacement aliases,
pin promotion and records supplying known names. Fresh original-generation checks still fence retries and
all writes. A test rejects unrelated source/group hydration in a 120-fact store and bounds guard reads;
other regressions preserve normalized duplicates, imported aliases, people tags, pins, cancellation,
concurrent saves, forgetting and erasure.

The helper still scans statements and metadata. Tagged names or corrections may hydrate more, and a
correction invalidating the dossier can read other windows. These results do not establish an indexed
large-store implementation or faster model decisions.

## Real embedding sample

Ten distinct fictional two-input requests succeeded on their first network attempt. Durations in sample
order were 2,602, 156, 185, 1,462, 288, 146, 145, 148, 162 and 159ms. Median: 160.5ms; maximum: 2,602ms.
H3 passes its predeclared thresholds. Original live timings were retained for the first four cached requests
in the continuation; no cache timing was counted as a live request.

At the deliberately conservative $0.50/M embedding-token rate, reported token usage estimates $0.000285.
The ledger retains $0.023232 of conservative request reservations across the two runs, bringing the existing
$2 campaign's conservative upper bound to $0.888851. There were no unknown-usage or unpriced responses in
this probe. Earlier unknown voice usage remains reserved. The local sample is small and does not reproduce
the reported long save, diagnose it or establish production tail latency.

## Adjacent regression and remaining work

A review reproduction found that updating a tool in an existing text-bearing chat message reset pending
extraction's generation and retry count despite unchanged speech. Comparing eligible conversation text with
the committed version preserves the pending work for tool-only changes and still queues changed/removed
evidence. Tests reproduced the failure in Realtime and aggregated text-turn projections before the fix.

Review also reproduced an incomplete retry boundary: the product REST client emitted generic errors for
transactional HTTP 409 `ABORTED`, while the save helper retried only SDK revision conflicts. The fix preserves
a typed abort at transaction read/commit boundaries and checks the original source/erasure fences before a
bounded retry, including when the competing revision is not yet visible. Other transport/provider failures
are not retried. This is a product transport fix; the SDK contract and release stay unchanged.

[inferred] The remaining slow-tail/proactivity work needs measured delegation and function durations,
indexed candidate lookup, and potentially a separate durable acceptance/indexing contract. Returning early
with a fake saved receipt or merely reducing provider timeouts does not implement that contract. Native
Live's revisable fragments and delayed-receipt speech also remain separate provider/product concerns.
Loading