Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
32 changes: 32 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,6 +2,38 @@

## [Unreleased]

### Direct SQLite upgrade guidance for PR #179 users

- **Stop or upgrade every original #179 (`a54a8ec`) client, including readers,
before resuming traffic with #180 (`e39357b`) or later.** One indexed BSON/date
read by a stale client can recreate an incompatible `_bson_v1` index, breaking
older writers and plain SQLite `UPDATE`, `REINDEX`, and `VACUUM`. Fixed clients
clean it up again on database open, so a mixed deployment can repeatedly
break and repair access.
- After stopping the affected clients, reopen each database with the fixed
version and perform a storage operation, such as
`client[database_name].list_collection_names()`, to remove the owned legacy
indexes. Constructing a client alone is insufficient. Pre-#179 writers remain
compatible with the repaired store; explicit unique/partial indexes keep
their separate SQLite function requirements. See the
[README upgrade steps](README.md#upgrading-sqlite-stores-from-pr-179).
- Budget for initial key materialization before serving traffic. Michael
Kennedy's direct SQLite retest measured a 200 MiB synthetic collection's cold
read at 199.83 ms before the fix versus 1,386.07 ms after, and a real 75,617-row
date range at 380.29 ms versus 760.54 ms. Warm reads stayed similar (59.69 to
63.81 ms and 4.38 to 4.49 ms). These external measurements describe his
workloads, not a latency guarantee. See the
[benchmark comparison](docs/BENCHMARKS.md#tm-053-external-direct-sqlite-retest).
- Add a tested [startup warm-up example](examples/sqlite_warmup.py): consume a
supported query for each relevant declared BSON/date index after migrations
and bulk loading, before traffic. Limiting results does not limit key-build
work; later writes still require key refresh on the next relevant read.
- Record Michael's acceptance of TM-053 and TM-054 at `fac214c` (merged as
`e39357b`): 903 application tests passed, with no candidate-narrowing or
MongoDB differential regressions. His unrelated application type-checker
gate failed on both comparison pins. See the
[acceptance record](docs/TALKPYTHON_ACCEPTANCE.md).

### TM-053 SQLite portability and partial-index validation
- Replace read-created BSON/date expression indexes with stored canonical key
columns and native indexes. Native SQL triggers invalidate changed rows;
Expand Down
82 changes: 78 additions & 4 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -32,6 +32,10 @@ the current stable package exactly; `master` may contain later changes. See
[`CHANGELOG.md`](https://github.com/schapman1974/tinymongo/blob/master/CHANGELOG.md)
for the complete release history and any unreleased work.

If you deployed the SQLite changes from PR #179 (`a54a8ec`), follow the
[SQLite upgrade steps](#upgrading-sqlite-stores-from-pr-179) before resuming
traffic with PR #180 (`e39357b`) or later.

The default JSON backend has a small dependency set. Optional database backends
may install native binary wheels supplied by DuckDB, PyArrow, or SQL drivers.

Expand Down Expand Up @@ -341,6 +345,8 @@ run or application restart. The command-line tool intentionally omits the
memory backend because every CLI invocation exits immediately; use it through
the Python API instead.

## SQLite read planning

SQLite, DuckDB, and Parquet compile supported Mongo-style filters into SQL over
the `_id` column and JSON document payload. Unsupported filter shapes fall back
to Python document matching so existing TinyMongo behavior remains available.
Expand All @@ -364,10 +370,9 @@ avoid this setup and refresh work.
These read-created indexes require no application-defined SQLite function, so
they permit plain SQLite writes, `VACUUM`, `REINDEX`, backups, and dump restores.
Opening an older store removes TinyMongo's legacy `_bson_v1` query indexes;
schema changes noticed by later queries also trigger cleanup. Upgrade clients
running the original #179 implementation, because their reads can recreate
those incompatible indexes. Writers from before #179 can use the new key
columns and invalidation triggers without registering a new function. Explicit
schema changes noticed by later queries also trigger cleanup. Writers from
before #179 can use the new key columns and invalidation triggers without
registering a new function. Explicit
unique and partial indexes retain their separate function requirements; this
change does not make every TinyMongo SQLite schema independent of the library.

Expand All @@ -384,6 +389,75 @@ post-image validation.
Older blob-format SQLite and DuckDB files are migrated to collection tables when
opened.

### Upgrading SQLite stores from PR #179

These steps apply to direct SQLite stores (`backend="sqlite"`).

**Stop or upgrade every process running the original PR #179 implementation
(`a54a8ec`), including readers, before resuming traffic.** A single indexed BSON
or date read by a stale #179 client can recreate a legacy `_bson_v1` index and
break older writers and plain SQLite operations such as `UPDATE`, `REINDEX`,
and `VACUUM`. A fixed client removes that index on its next database open, so
mixed versions can repeatedly break and repair access to the same store.

1. Stop the #179 readers and writers, including background workers and scripts
that share the SQLite files. Upgrade them to PR #180 (`e39357b`) or later.
2. With the fixed version, reopen each affected SQLite database and perform a
storage operation, for example
`client[database_name].list_collection_names()`. This removes the owned
legacy indexes automatically; constructing a client alone does not open
the database files.
3. After migrations or bulk loading, warm the indexes used by critical reads
as described below, then resume traffic with the upgraded processes.

Pre-#179 writers can continue using the repaired store; the requirement above
targets clients that can recreate the #179 query indexes. The portability
change applies to the read-created BSON/date indexes. Explicit unique and
partial indexes still have their separate SQLite function requirements.

### Warm SQLite reads before serving traffic

The portability fix trades a more expensive first read for native SQLite
maintenance and stable warm reads. In Michael Kennedy's direct SQLite retest,
the first indexed read of a 200 MiB synthetic collection rose from 199.83 ms
to 1,386.07 ms; a date range over 75,617 real `opt_ins` documents rose from
380.29 ms to 760.54 ms. Warm reads were similar: 59.69 to 63.81 ms and 4.38 to
4.49 ms respectively. These are measurements of his datasets, not timing
guarantees. See the [external retest](docs/BENCHMARKS.md#tm-053-external-direct-sqlite-retest)
for the full comparison and [his report](https://github.com/schapman1974/tinymongo/issues/136#issuecomment-5736577143)
for the environment.

Run representative indexed reads after migrations and bulk loading, before
marking the application ready for traffic. Warm each relevant declared
top-level BSON/date index with a separate supported predicate; a query using
several indexed fields may select only one candidate source. Partial indexes,
dotted fields, and unsupported predicates do not gain acceleration from this
step. Ordinary numeric ranges use a different path and do not need these
materialized BSON keys.

Consume the cursor: `list(collection.find(query).limit(1))` executes the read,
whereas constructing the cursor alone does not. The limit bounds the returned
documents, **not the first key build**, which scans the collection and holds a
write lock. Keys persist across client restarts, but later inserts and updates
leave uncomputed keys for the next relevant read to refresh. Warm-up moves the
initial cost into startup; it does not eliminate refresh work after writes.

The runnable [startup warm-up example](examples/sqlite_warmup.py) uses public
APIs and synthetic data in a temporary direct SQLite store. It demonstrates
declaring indexes, loading data, and consuming representative date and BSON
scalar reads before serving requests. From a checkout, install the optional
BSON types used by the example and run it:

```bash
python -m pip install -e '.[bson]'
python -m examples.sqlite_warmup
```

Adapt its warm-up function to your existing collections and query values;
keep index creation and data migration in your application's setup phase.

## Backend guides and benchmarks

Local load-test results for these backends are documented in
[backend benchmarks](https://github.com/schapman1974/tinymongo/blob/master/docs/BENCHMARKS.md).

Expand Down
49 changes: 46 additions & 3 deletions docs/BENCHMARKS.md
Original file line number Diff line number Diff line change
Expand Up @@ -269,8 +269,51 @@ large-document penalty, but is not a private application rerun.
The new query indexes contain no application-defined function. Local
portability regressions cover plain SQLite updates and inserts, `VACUUM`,
`REINDEX`, backup and dump restore, and cleanup of legacy `_bson_v1` indexes.
Clients running the original #179 read planner still need upgrading because
they can recreate those legacy indexes. Existing explicit unique and partial
indexes have separate function requirements, outside this query-cache change.
Stop or upgrade every original #179 process sharing the file, including readers
and writers: one of its reads can recreate those legacy indexes. Pre-#179
writers can remain. Existing explicit unique and partial indexes have separate
function requirements, outside this query-cache change.
Dropping a declared index removes its derived index and trigger and clears its
keys; the empty column remains for SQLite versions without `DROP COLUMN`.

## TM-053 external direct SQLite retest

Michael's [external #180 acceptance report](https://github.com/schapman1974/tinymongo/issues/136#issuecomment-5736577143)
compared direct SQLite at `a54a8ec` with the stored-key implementation at
`fac214c`. The latter merged as `e39357b` while he measured; he checked that
the three source files were byte-identical. These are his measurements from
serialized runs in the same session, separate from the local synthetic results
above. His environment was Python 3.14.6, PyMongo 4.17, and macOS 26.2 arm64.

| Cold first read | `a54a8ec` | `fac214c` | Ratio |
| --- | ---: | ---: | ---: |
| Synthetic, 0.6 MiB | 3.92 ms | 11.61 ms | 3.0x |
| Synthetic, 3.5 MiB | 6.41 ms | 26.77 ms | 4.2x |
| Synthetic, 25 MiB | 30.40 ms | 187.00 ms | 6.2x |
| Synthetic, 200 MiB | 199.83 ms | 1,386.07 ms | 6.9x |
| Real `opt_ins` date range, 75,617 documents | 380.29 ms | 760.54 ms | 2.0x |
| Real `episodes` date range, 563 documents | 24.62 ms | 81.28 ms | 3.3x |

Warm reads stayed approximately flat: the 200 MiB synthetic case took 59.69 ms
before and 63.81 ms after, while the real `opt_ins` range took 4.38 and 4.49 ms.
Stored keys preserve the TM-042 steady-state improvement but make their initial
materialization more expensive. Budget this work during deployment warm-up:
the first relevant query took about 1.4 seconds for his 200 MiB collection and
0.76 seconds for the real `opt_ins` collection. These measurements do not imply
that every new filter shape rebuilds the keys; queries can reuse materialized
keys for the same declared field index.

His synthetic write checks found a 1.0x before/after-index ratio at every payload
on both pins. The real `opt_ins` insert stayed near one second; migration took
11.1 versus 10.9 seconds, and application startup took 0.8 versus 1.1 seconds.
The seeded application suite passed 903 tests on both pins, and the 45-shape
query differential found no regressions. The full acceptance and the remaining
application-side type-checker gate are recorded in
[the external acceptance results](TALKPYTHON_ACCEPTANCE.md#michael-kennedys-180-acceptance-external-evidence).

Before warming or serving a shared file, stop or upgrade all original #179
readers and writers. Michael verified that one stale #179 read can recreate the
legacy function-dependent index and block writers and plain SQLite maintenance
again, even after #180 repaired the store. Pre-#179 writers do not recreate that
index and can remain. Warm-up cannot make an active mix of #179 and #180 clients
safe from this recurring incompatibility.
78 changes: 70 additions & 8 deletions docs/TALKPYTHON_ACCEPTANCE.md
Original file line number Diff line number Diff line change
Expand Up @@ -322,14 +322,16 @@ JSON/memory application tests remained blocked by the startup cost tracked in

## TM-053 and remaining partial-filter validation follow-up (local evidence)

The combined follow-up stores canonical BSON/date keys in ordinary SQLite
columns, with native indexes and SQL triggers that invalidate keys when a
document changes. Relevant reads refresh only uncomputed keys transactionally
and also include rows invalidated concurrently as conservative candidates.
The combined follow-up, merged as #180 (`e39357b`), stores canonical BSON/date
keys in ordinary SQLite columns, with native indexes and SQL triggers that
invalidate keys when a document changes. Relevant reads refresh only uncomputed
keys transactionally and also include rows invalidated concurrently as
conservative candidates.
Opening a store removes owned legacy `_bson_v1` indexes; queries repeat cleanup
when they discover schema changes. The new read-created schema requires no
query-key Python function. Older pre-#179 writers remain compatible, while
original #179 readers must upgrade so they cannot recreate the legacy indexes.
query-key Python function. Older pre-#179 writers remain compatible. Stop or
upgrade every original #179 process sharing the file, including both readers
and writers, so its reads cannot recreate the legacy indexes.
Explicit unique and partial indexes retain separate function requirements.

`tests/test_sqlite_query_portability.py` covers plain SQLite maintenance,
Expand All @@ -344,5 +346,65 @@ valid predicates may appear in an index. Unknown operators inside field
predicates and malformed nested operands return code `2`; valid prohibited
predicates retain code `67`. The shared index-validation contracts cover the
embedded backends, sync/async APIs, and the MongoDB reference backend. These
local checks are not a rerun of Michael's private application; his next SQLite
pass should confirm portability and query behavior against the new pin.
local checks are separate from Michael's subsequent private application rerun,
recorded below.

## Michael Kennedy's #180 acceptance (external evidence)

Michael's [retest of `fac214c`](https://github.com/schapman1974/tinymongo/issues/136#issuecomment-5736577143)
compared the #180 branch with `a54a8ec` using direct `sqlite` storage. The branch
merged as `e39357b` during his measurements; he verified the three source files
were byte-identical between those pins. He confirmed both TM-053 and the
remaining partial-filter error-code mismatch, TM-054, were fixed, with no
regressions. These are his external results, not a local rerun or new evidence
for other storage backends.

- **903 application tests passed on both pins** against the seeded SQLite
store. The separate type-checker gate failed on both because of pre-existing
application diagnostics, unrelated to TinyMongo.
- All 45 differential query shapes retained their results: 27 narrowed, with
zero narrowing bugs and zero divergences from MongoDB 8.2. Eight selective
shapes still examined 0–1% of a 2,000-document collection.
- A date read created a native `_bson_v2` index and invalidation trigger, with
no `_bson_v1` dependency. A pre-#179 client's insert and plain SQLite
`UPDATE`, `REINDEX`, and `VACUUM` succeeded. Opening a store containing a
legacy query index removed it and restored writer and maintenance access.
His application test run itself also left the store free of legacy indexes.
- The 13-case partial-filter matrix had zero error-code divergences, down
from one. Unknown operators returned `2`; prohibited predicates retained
`67`, and non-mapping filters retained `14`. The pre-existing distinction
between `OperationFailure` and PyMongo's `WriteError` remained; `WriteError`
subclasses `OperationFailure`, so existing exception handlers still match.

His mixed-version check makes the upgrade order significant: a single read by
an original #179 (`a54a8ec`) client recreated `_bson_v1` alongside the native
index and again prevented other writers and plain SQLite maintenance from
working. Opening the store with #180 repaired it again. **Stop or upgrade all
original #179 readers and writers before returning the shared file to service.**
Pre-#179 writers can remain; they do not create the incompatible query index.
Repeated cleanup cannot make a live mix of #179 and #180 clients stable.

He also measured the cost of first-use key materialization separately from
warm reads:

| Direct SQLite workload | `a54a8ec` cold | `fac214c` cold | `a54a8ec` warm | `fac214c` warm |
| --- | ---: | ---: | ---: | ---: |
| Synthetic, 200 MiB | 199.83 ms | 1,386.07 ms | 59.69 ms | 63.81 ms |
| Real `opt_ins` date range, 75,617 documents | 380.29 ms | 760.54 ms | 4.38 ms | 4.49 ms |

Warm performance stayed approximately flat, while cold key construction was
6.9x slower in the large synthetic case and 2.0x slower on `opt_ins`. These
are first-use costs to budget before serving traffic, not steady-state query
latencies. The [full external cold-read table](BENCHMARKS.md#tm-053-external-direct-sqlite-retest)
includes the smaller synthetic collections and `episodes`.

His write checks found a 1.0x before/after-index ratio at every synthetic payload
on both pins; a real `opt_ins` insert stayed near one second. Migration took
11.1 versus 10.9 seconds, and application startup took 0.8 versus 1.1 seconds.
He separately corrected the NUL documentation request: the PostgreSQL
limitation was already documented at the tested pin in `REMOTE_SQL.md`.

The external run used Python 3.14.6, PyMongo 4.17, MongoDB 8.2.3 as the query
oracle, and macOS 26.2 arm64, with serialized runs against an 81,579-document
real dataset. It does not replace the historical results for JSON, memory, or
sharded SQLite.
Loading
Loading