diff --git a/CHANGELOG.md b/CHANGELOG.md index 361ca37..7d116f6 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -2,6 +2,38 @@ ## [Unreleased] +### Direct SQLite upgrade guidance for PR #179 users + +- **Stop or upgrade every original #179 (`a54a8ec`) client, including readers, + before resuming traffic with #180 (`e39357b`) or later.** One indexed BSON/date + read by a stale client can recreate an incompatible `_bson_v1` index, breaking + older writers and plain SQLite `UPDATE`, `REINDEX`, and `VACUUM`. Fixed clients + clean it up again on database open, so a mixed deployment can repeatedly + break and repair access. +- After stopping the affected clients, reopen each database with the fixed + version and perform a storage operation, such as + `client[database_name].list_collection_names()`, to remove the owned legacy + indexes. Constructing a client alone is insufficient. Pre-#179 writers remain + compatible with the repaired store; explicit unique/partial indexes keep + their separate SQLite function requirements. See the + [README upgrade steps](README.md#upgrading-sqlite-stores-from-pr-179). +- Budget for initial key materialization before serving traffic. Michael + Kennedy's direct SQLite retest measured a 200 MiB synthetic collection's cold + read at 199.83 ms before the fix versus 1,386.07 ms after, and a real 75,617-row + date range at 380.29 ms versus 760.54 ms. Warm reads stayed similar (59.69 to + 63.81 ms and 4.38 to 4.49 ms). These external measurements describe his + workloads, not a latency guarantee. See the + [benchmark comparison](docs/BENCHMARKS.md#tm-053-external-direct-sqlite-retest). +- Add a tested [startup warm-up example](examples/sqlite_warmup.py): consume a + supported query for each relevant declared BSON/date index after migrations + and bulk loading, before traffic. Limiting results does not limit key-build + work; later writes still require key refresh on the next relevant read. +- Record Michael's acceptance of TM-053 and TM-054 at `fac214c` (merged as + `e39357b`): 903 application tests passed, with no candidate-narrowing or + MongoDB differential regressions. His unrelated application type-checker + gate failed on both comparison pins. See the + [acceptance record](docs/TALKPYTHON_ACCEPTANCE.md). + ### TM-053 SQLite portability and partial-index validation - Replace read-created BSON/date expression indexes with stored canonical key columns and native indexes. Native SQL triggers invalidate changed rows; diff --git a/README.md b/README.md index 4a54e72..c2a878e 100644 --- a/README.md +++ b/README.md @@ -32,6 +32,10 @@ the current stable package exactly; `master` may contain later changes. See [`CHANGELOG.md`](https://github.com/schapman1974/tinymongo/blob/master/CHANGELOG.md) for the complete release history and any unreleased work. +If you deployed the SQLite changes from PR #179 (`a54a8ec`), follow the +[SQLite upgrade steps](#upgrading-sqlite-stores-from-pr-179) before resuming +traffic with PR #180 (`e39357b`) or later. + The default JSON backend has a small dependency set. Optional database backends may install native binary wheels supplied by DuckDB, PyArrow, or SQL drivers. @@ -341,6 +345,8 @@ run or application restart. The command-line tool intentionally omits the memory backend because every CLI invocation exits immediately; use it through the Python API instead. +## SQLite read planning + SQLite, DuckDB, and Parquet compile supported Mongo-style filters into SQL over the `_id` column and JSON document payload. Unsupported filter shapes fall back to Python document matching so existing TinyMongo behavior remains available. @@ -364,10 +370,9 @@ avoid this setup and refresh work. These read-created indexes require no application-defined SQLite function, so they permit plain SQLite writes, `VACUUM`, `REINDEX`, backups, and dump restores. Opening an older store removes TinyMongo's legacy `_bson_v1` query indexes; -schema changes noticed by later queries also trigger cleanup. Upgrade clients -running the original #179 implementation, because their reads can recreate -those incompatible indexes. Writers from before #179 can use the new key -columns and invalidation triggers without registering a new function. Explicit +schema changes noticed by later queries also trigger cleanup. Writers from +before #179 can use the new key columns and invalidation triggers without +registering a new function. Explicit unique and partial indexes retain their separate function requirements; this change does not make every TinyMongo SQLite schema independent of the library. @@ -384,6 +389,75 @@ post-image validation. Older blob-format SQLite and DuckDB files are migrated to collection tables when opened. +### Upgrading SQLite stores from PR #179 + +These steps apply to direct SQLite stores (`backend="sqlite"`). + +**Stop or upgrade every process running the original PR #179 implementation +(`a54a8ec`), including readers, before resuming traffic.** A single indexed BSON +or date read by a stale #179 client can recreate a legacy `_bson_v1` index and +break older writers and plain SQLite operations such as `UPDATE`, `REINDEX`, +and `VACUUM`. A fixed client removes that index on its next database open, so +mixed versions can repeatedly break and repair access to the same store. + +1. Stop the #179 readers and writers, including background workers and scripts + that share the SQLite files. Upgrade them to PR #180 (`e39357b`) or later. +2. With the fixed version, reopen each affected SQLite database and perform a + storage operation, for example + `client[database_name].list_collection_names()`. This removes the owned + legacy indexes automatically; constructing a client alone does not open + the database files. +3. After migrations or bulk loading, warm the indexes used by critical reads + as described below, then resume traffic with the upgraded processes. + +Pre-#179 writers can continue using the repaired store; the requirement above +targets clients that can recreate the #179 query indexes. The portability +change applies to the read-created BSON/date indexes. Explicit unique and +partial indexes still have their separate SQLite function requirements. + +### Warm SQLite reads before serving traffic + +The portability fix trades a more expensive first read for native SQLite +maintenance and stable warm reads. In Michael Kennedy's direct SQLite retest, +the first indexed read of a 200 MiB synthetic collection rose from 199.83 ms +to 1,386.07 ms; a date range over 75,617 real `opt_ins` documents rose from +380.29 ms to 760.54 ms. Warm reads were similar: 59.69 to 63.81 ms and 4.38 to +4.49 ms respectively. These are measurements of his datasets, not timing +guarantees. See the [external retest](docs/BENCHMARKS.md#tm-053-external-direct-sqlite-retest) +for the full comparison and [his report](https://github.com/schapman1974/tinymongo/issues/136#issuecomment-5736577143) +for the environment. + +Run representative indexed reads after migrations and bulk loading, before +marking the application ready for traffic. Warm each relevant declared +top-level BSON/date index with a separate supported predicate; a query using +several indexed fields may select only one candidate source. Partial indexes, +dotted fields, and unsupported predicates do not gain acceleration from this +step. Ordinary numeric ranges use a different path and do not need these +materialized BSON keys. + +Consume the cursor: `list(collection.find(query).limit(1))` executes the read, +whereas constructing the cursor alone does not. The limit bounds the returned +documents, **not the first key build**, which scans the collection and holds a +write lock. Keys persist across client restarts, but later inserts and updates +leave uncomputed keys for the next relevant read to refresh. Warm-up moves the +initial cost into startup; it does not eliminate refresh work after writes. + +The runnable [startup warm-up example](examples/sqlite_warmup.py) uses public +APIs and synthetic data in a temporary direct SQLite store. It demonstrates +declaring indexes, loading data, and consuming representative date and BSON +scalar reads before serving requests. From a checkout, install the optional +BSON types used by the example and run it: + +```bash +python -m pip install -e '.[bson]' +python -m examples.sqlite_warmup +``` + +Adapt its warm-up function to your existing collections and query values; +keep index creation and data migration in your application's setup phase. + +## Backend guides and benchmarks + Local load-test results for these backends are documented in [backend benchmarks](https://github.com/schapman1974/tinymongo/blob/master/docs/BENCHMARKS.md). diff --git a/docs/BENCHMARKS.md b/docs/BENCHMARKS.md index 242ca87..fbc9b85 100644 --- a/docs/BENCHMARKS.md +++ b/docs/BENCHMARKS.md @@ -269,8 +269,51 @@ large-document penalty, but is not a private application rerun. The new query indexes contain no application-defined function. Local portability regressions cover plain SQLite updates and inserts, `VACUUM`, `REINDEX`, backup and dump restore, and cleanup of legacy `_bson_v1` indexes. -Clients running the original #179 read planner still need upgrading because -they can recreate those legacy indexes. Existing explicit unique and partial -indexes have separate function requirements, outside this query-cache change. +Stop or upgrade every original #179 process sharing the file, including readers +and writers: one of its reads can recreate those legacy indexes. Pre-#179 +writers can remain. Existing explicit unique and partial indexes have separate +function requirements, outside this query-cache change. Dropping a declared index removes its derived index and trigger and clears its keys; the empty column remains for SQLite versions without `DROP COLUMN`. + +## TM-053 external direct SQLite retest + +Michael's [external #180 acceptance report](https://github.com/schapman1974/tinymongo/issues/136#issuecomment-5736577143) +compared direct SQLite at `a54a8ec` with the stored-key implementation at +`fac214c`. The latter merged as `e39357b` while he measured; he checked that +the three source files were byte-identical. These are his measurements from +serialized runs in the same session, separate from the local synthetic results +above. His environment was Python 3.14.6, PyMongo 4.17, and macOS 26.2 arm64. + +| Cold first read | `a54a8ec` | `fac214c` | Ratio | +| --- | ---: | ---: | ---: | +| Synthetic, 0.6 MiB | 3.92 ms | 11.61 ms | 3.0x | +| Synthetic, 3.5 MiB | 6.41 ms | 26.77 ms | 4.2x | +| Synthetic, 25 MiB | 30.40 ms | 187.00 ms | 6.2x | +| Synthetic, 200 MiB | 199.83 ms | 1,386.07 ms | 6.9x | +| Real `opt_ins` date range, 75,617 documents | 380.29 ms | 760.54 ms | 2.0x | +| Real `episodes` date range, 563 documents | 24.62 ms | 81.28 ms | 3.3x | + +Warm reads stayed approximately flat: the 200 MiB synthetic case took 59.69 ms +before and 63.81 ms after, while the real `opt_ins` range took 4.38 and 4.49 ms. +Stored keys preserve the TM-042 steady-state improvement but make their initial +materialization more expensive. Budget this work during deployment warm-up: +the first relevant query took about 1.4 seconds for his 200 MiB collection and +0.76 seconds for the real `opt_ins` collection. These measurements do not imply +that every new filter shape rebuilds the keys; queries can reuse materialized +keys for the same declared field index. + +His synthetic write checks found a 1.0x before/after-index ratio at every payload +on both pins. The real `opt_ins` insert stayed near one second; migration took +11.1 versus 10.9 seconds, and application startup took 0.8 versus 1.1 seconds. +The seeded application suite passed 903 tests on both pins, and the 45-shape +query differential found no regressions. The full acceptance and the remaining +application-side type-checker gate are recorded in +[the external acceptance results](TALKPYTHON_ACCEPTANCE.md#michael-kennedys-180-acceptance-external-evidence). + +Before warming or serving a shared file, stop or upgrade all original #179 +readers and writers. Michael verified that one stale #179 read can recreate the +legacy function-dependent index and block writers and plain SQLite maintenance +again, even after #180 repaired the store. Pre-#179 writers do not recreate that +index and can remain. Warm-up cannot make an active mix of #179 and #180 clients +safe from this recurring incompatibility. diff --git a/docs/TALKPYTHON_ACCEPTANCE.md b/docs/TALKPYTHON_ACCEPTANCE.md index 684003a..cd21756 100644 --- a/docs/TALKPYTHON_ACCEPTANCE.md +++ b/docs/TALKPYTHON_ACCEPTANCE.md @@ -322,14 +322,16 @@ JSON/memory application tests remained blocked by the startup cost tracked in ## TM-053 and remaining partial-filter validation follow-up (local evidence) -The combined follow-up stores canonical BSON/date keys in ordinary SQLite -columns, with native indexes and SQL triggers that invalidate keys when a -document changes. Relevant reads refresh only uncomputed keys transactionally -and also include rows invalidated concurrently as conservative candidates. +The combined follow-up, merged as #180 (`e39357b`), stores canonical BSON/date +keys in ordinary SQLite columns, with native indexes and SQL triggers that +invalidate keys when a document changes. Relevant reads refresh only uncomputed +keys transactionally and also include rows invalidated concurrently as +conservative candidates. Opening a store removes owned legacy `_bson_v1` indexes; queries repeat cleanup when they discover schema changes. The new read-created schema requires no -query-key Python function. Older pre-#179 writers remain compatible, while -original #179 readers must upgrade so they cannot recreate the legacy indexes. +query-key Python function. Older pre-#179 writers remain compatible. Stop or +upgrade every original #179 process sharing the file, including both readers +and writers, so its reads cannot recreate the legacy indexes. Explicit unique and partial indexes retain separate function requirements. `tests/test_sqlite_query_portability.py` covers plain SQLite maintenance, @@ -344,5 +346,65 @@ valid predicates may appear in an index. Unknown operators inside field predicates and malformed nested operands return code `2`; valid prohibited predicates retain code `67`. The shared index-validation contracts cover the embedded backends, sync/async APIs, and the MongoDB reference backend. These -local checks are not a rerun of Michael's private application; his next SQLite -pass should confirm portability and query behavior against the new pin. +local checks are separate from Michael's subsequent private application rerun, +recorded below. + +## Michael Kennedy's #180 acceptance (external evidence) + +Michael's [retest of `fac214c`](https://github.com/schapman1974/tinymongo/issues/136#issuecomment-5736577143) +compared the #180 branch with `a54a8ec` using direct `sqlite` storage. The branch +merged as `e39357b` during his measurements; he verified the three source files +were byte-identical between those pins. He confirmed both TM-053 and the +remaining partial-filter error-code mismatch, TM-054, were fixed, with no +regressions. These are his external results, not a local rerun or new evidence +for other storage backends. + +- **903 application tests passed on both pins** against the seeded SQLite + store. The separate type-checker gate failed on both because of pre-existing + application diagnostics, unrelated to TinyMongo. +- All 45 differential query shapes retained their results: 27 narrowed, with + zero narrowing bugs and zero divergences from MongoDB 8.2. Eight selective + shapes still examined 0–1% of a 2,000-document collection. +- A date read created a native `_bson_v2` index and invalidation trigger, with + no `_bson_v1` dependency. A pre-#179 client's insert and plain SQLite + `UPDATE`, `REINDEX`, and `VACUUM` succeeded. Opening a store containing a + legacy query index removed it and restored writer and maintenance access. + His application test run itself also left the store free of legacy indexes. +- The 13-case partial-filter matrix had zero error-code divergences, down + from one. Unknown operators returned `2`; prohibited predicates retained + `67`, and non-mapping filters retained `14`. The pre-existing distinction + between `OperationFailure` and PyMongo's `WriteError` remained; `WriteError` + subclasses `OperationFailure`, so existing exception handlers still match. + +His mixed-version check makes the upgrade order significant: a single read by +an original #179 (`a54a8ec`) client recreated `_bson_v1` alongside the native +index and again prevented other writers and plain SQLite maintenance from +working. Opening the store with #180 repaired it again. **Stop or upgrade all +original #179 readers and writers before returning the shared file to service.** +Pre-#179 writers can remain; they do not create the incompatible query index. +Repeated cleanup cannot make a live mix of #179 and #180 clients stable. + +He also measured the cost of first-use key materialization separately from +warm reads: + +| Direct SQLite workload | `a54a8ec` cold | `fac214c` cold | `a54a8ec` warm | `fac214c` warm | +| --- | ---: | ---: | ---: | ---: | +| Synthetic, 200 MiB | 199.83 ms | 1,386.07 ms | 59.69 ms | 63.81 ms | +| Real `opt_ins` date range, 75,617 documents | 380.29 ms | 760.54 ms | 4.38 ms | 4.49 ms | + +Warm performance stayed approximately flat, while cold key construction was +6.9x slower in the large synthetic case and 2.0x slower on `opt_ins`. These +are first-use costs to budget before serving traffic, not steady-state query +latencies. The [full external cold-read table](BENCHMARKS.md#tm-053-external-direct-sqlite-retest) +includes the smaller synthetic collections and `episodes`. + +His write checks found a 1.0x before/after-index ratio at every synthetic payload +on both pins; a real `opt_ins` insert stayed near one second. Migration took +11.1 versus 10.9 seconds, and application startup took 0.8 versus 1.1 seconds. +He separately corrected the NUL documentation request: the PostgreSQL +limitation was already documented at the tested pin in `REMOTE_SQL.md`. + +The external run used Python 3.14.6, PyMongo 4.17, MongoDB 8.2.3 as the query +oracle, and macOS 26.2 arm64, with serialized runs against an 81,579-document +real dataset. It does not replace the historical results for JSON, memory, or +sharded SQLite. diff --git a/examples/sqlite_warmup.py b/examples/sqlite_warmup.py new file mode 100644 index 0000000..5cf7770 --- /dev/null +++ b/examples/sqlite_warmup.py @@ -0,0 +1,79 @@ +"""Warm SQLite query keys during application startup using the public API. + +Run this temporary-store demonstration from the repository root: + + python -m pip install -e '.[bson]' + python -m examples.sqlite_warmup + +For your application, finish schema/data migrations and declare the indexes, +then call ``warm_sqlite_queries`` before accepting traffic. First use may scan +the whole collection, write keys, and hold the SQLite write lock. ``limit(1)`` +bounds returned documents; it does not limit that setup work. +""" + +import tempfile +from datetime import datetime, timedelta + +from bson import ObjectId + +from tinymongo import TinyMongoClient + + +def warm_sqlite_queries(collection, queries): + """Execute one representative indexed read per field to warm query keys. + + Use separate filters for declared, non-partial indexes whose leading fields + are top-level BSON scalars or dates. Combining fields in one filter may warm + only one index. A probe need not match a document to warm a populated + collection. Empty collections have no keys to compute; later writes leave + keys for the next relevant read to refresh. + + Creating a cursor alone is insufficient: consume it to execute the read. + Repeating these probes after a restart reuses persisted keys, and repeating + them after writes refreshes invalidated keys. + """ + for query in queries: + list(collection.find(query, {"_id": 1}).limit(1)) + + +def run_example(): + """Seed and warm an isolated SQLite store, then simulate serving a read.""" + published_at = datetime(2026, 1, 1) + author_id = ObjectId("000000000000000000000001") + with tempfile.TemporaryDirectory(prefix="tinymongo-sqlite-warmup-") as path: + with TinyMongoClient(path, backend="sqlite") as client: + database = client.app + # Opening a handle is lazy; this operation performs upgrade cleanup. + database.list_collection_names() + articles = database.articles + + # Application schema/data migrations happen before warm-up. + articles.insert_many( + [ + { + "_id": number, + "published_at": published_at + timedelta(days=number), + "author_id": author_id, + "title": "Example article {0}".format(number), + } + for number in range(3) + ] + ) + articles.create_index("published_at") + articles.create_index("author_id") + + warm_sqlite_queries( + articles, + [ + {"published_at": {"$gte": published_at}}, + {"author_id": author_id}, + ], + ) + + # Start accepting application traffic only after warm-up returns. + return articles.count_documents({"author_id": author_id}) + + +if __name__ == "__main__": + count = run_example() + print("Startup warm-up finished for {0} example articles.".format(count)) diff --git a/tests/test_sqlite_warmup_example.py b/tests/test_sqlite_warmup_example.py new file mode 100644 index 0000000..f73a22e --- /dev/null +++ b/tests/test_sqlite_warmup_example.py @@ -0,0 +1,152 @@ +"""The documented startup probes actually prepare and refresh SQLite keys.""" + +import sqlite3 +import subprocess +import sys +from datetime import datetime, timedelta +from pathlib import Path + +import pytest +from bson import ObjectId + +from examples.sqlite_warmup import warm_sqlite_queries +from tinymongo import TinyMongoClient, table_backends + + +DATE = datetime(2026, 1, 1) +AUTHOR = ObjectId("000000000000000000000001") +OTHER_AUTHOR = ObjectId("000000000000000000000002") + + +def _uncomputed_key_counts(collection): + # Inspect durable state rather than assuming a consumed cursor warmed it. + with sqlite3.connect(collection.parent.engine.path) as connection: + columns = [ + row[1] + for row in connection.execute('PRAGMA table_info("articles")') + if row[1].endswith("_bson_key_v2") + ] + return [ + connection.execute( + 'SELECT COUNT(*) FROM "articles" WHERE "{0}" IS NULL'.format(column) + ).fetchone()[0] + for column in columns + ] + + +@pytest.mark.parametrize("matching_probes", [True, False]) +def test_startup_warmup_persists_keys_and_refreshes_later_writes( + tmp_path, monkeypatch, matching_probes +): + probes = [ + { + "published_at": { + "$gte": DATE if matching_probes else DATE + timedelta(days=100) + } + }, + {"author_id": AUTHOR if matching_probes else OTHER_AUTHOR}, + ] + computed = [] + original = table_backends._sqlite_bson_query_key_from_row + + def record_computation(data, field): + computed.append(field) + return original(data, field) + + monkeypatch.setattr( + table_backends, "_sqlite_bson_query_key_from_row", record_computation + ) + with TinyMongoClient(str(tmp_path), backend="sqlite") as client: + articles = client.app.articles + articles.insert_many( + [ + { + "_id": number, + "published_at": DATE + timedelta(days=number), + "author_id": AUTHOR, + } + for number in range(3) + ] + ) + articles.create_index("published_at") + articles.create_index("author_id") + assert _uncomputed_key_counts(articles) == [] + + warm_sqlite_queries(articles, probes) + assert _uncomputed_key_counts(articles) == [0, 0] + assert computed.count("published_at") == 3 + assert computed.count("author_id") == 3 + + computed.clear() + with TinyMongoClient(str(tmp_path), backend="sqlite") as client: + articles = client.app.articles + assert _uncomputed_key_counts(articles) == [0, 0] + warm_sqlite_queries(articles, probes) + assert computed == [] # Restarting does not rebuild existing keys. + + changed = {"published_at": DATE + timedelta(days=20), "author_id": OTHER_AUTHOR} + articles.update_one({"_id": 0}, {"$set": changed}) + articles.insert_one(dict(changed, _id=3)) + assert _uncomputed_key_counts(articles) == [2, 2] + + warm_sqlite_queries(articles, probes) + assert _uncomputed_key_counts(articles) == [0, 0] + assert computed.count("published_at") == 2 + assert computed.count("author_id") == 2 + assert [row["_id"] for row in articles.find({"author_id": OTHER_AUTHOR})] == [ + 0, + 3, + ] + assert articles.count_documents({"published_at": changed["published_at"]}) == 2 + + +def test_sqlite_warmup_example_runs_as_documented(): + result = subprocess.run( + [sys.executable, "-m", "examples.sqlite_warmup"], + cwd=str(Path(__file__).resolve().parents[1]), + capture_output=True, + text=True, + check=True, + timeout=30, + ) + assert result.stdout.strip() == "Startup warm-up finished for 3 example articles." + + +def test_documented_collection_listing_performs_legacy_index_cleanup(tmp_path): + with TinyMongoClient(str(tmp_path), backend="sqlite") as client: + articles = client.app.articles + articles.insert_one({"_id": 1, "author_id": AUTHOR}) + articles.create_index("author_id") + engine = articles.parent.engine + path = engine.path + spec = engine.get_index_specs("articles")[0] + legacy_name = engine._physical_index_name("articles", spec) + "_bson_v1" + with sqlite3.connect(path) as connection: + connection.create_function( + "tinymongo_bson_query_key_v1", + 2, + table_backends._sqlite_bson_query_key_from_row, + deterministic=True, + ) + connection.execute( + 'CREATE INDEX "{0}" ON articles ' + "(tinymongo_bson_query_key_v1(data, 'author_id'))".format(legacy_name) + ) + + with TinyMongoClient(str(tmp_path), backend="sqlite") as client: + database = client["app"] + with sqlite3.connect(path) as connection: + assert connection.execute( + "SELECT name FROM sqlite_master WHERE name = ?", (legacy_name,) + ).fetchone() == (legacy_name,) + assert database.list_collection_names() == ["articles"] + with sqlite3.connect(path) as connection: + assert ( + connection.execute( + "SELECT name FROM sqlite_master WHERE name = ?", (legacy_name,) + ).fetchone() + is None + ) + connection.execute("VACUUM") + connection.execute("REINDEX") + connection.execute("UPDATE articles SET data = data")