Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
17 commits
Select commit Hold shift + click to select a range
ff77a4d
Verify a backup is readable before restore_backup replaces the database
cb1kenobi Oct 8, 2026
b1b1544
Keep the marker after any publication, and the displaced mode on reco…
cb1kenobi Oct 8, 2026
183c53a
Probe the offline restore's lock before staging too
cb1kenobi Oct 8, 2026
4430600
Refuse to restore into a database directory that is a mount point
cb1kenobi Oct 8, 2026
6665dd6
Refuse a restore whose staging copy would not fit beside the database
cb1kenobi Oct 8, 2026
624d78a
Drop a failed publication's candidate before the rerun measures space
cb1kenobi Oct 8, 2026
9944799
Retire .replaced atomically before the marker clears
cb1kenobi Oct 8, 2026
0979503
Keep a refusal's status code when a restore needs a rerun, and cover …
cb1kenobi Oct 8, 2026
df23f32
Merge remote-tracking branch 'origin/main' into fix/restore-verify-be…
cb1kenobi Oct 8, 2026
3da1cac
Frame-review the new restore staging module, which imports the storag…
cb1kenobi Oct 8, 2026
569d858
Read a caught value's message and code defensively in restore staging
cb1kenobi Oct 8, 2026
3d0e562
Verify closure before an online rerun drops a published candidate
cb1kenobi Oct 8, 2026
be52c8d
Keep a database loaded through a rescan during online staging
cb1kenobi Oct 9, 2026
0f1ac65
Name a stray .replaced instead of blaming an open database for it
cb1kenobi Oct 9, 2026
bf1e140
Say what a restore marker proves, and stop logging an outage while se…
cb1kenobi Oct 9, 2026
5583a42
Merge branch 'main' into fix/restore-verify-before-purge-2965
cb1kenobi Oct 9, 2026
146e1fd
Merge branch 'main' into fix/restore-verify-before-purge-2965
cb1kenobi Oct 9, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions .github/workflows/review-coverage.yml
Original file line number Diff line number Diff line change
Expand Up @@ -123,6 +123,7 @@ jobs:
dataLayer/harperBridge/ResourceBridge.ts
dataLayer/harperBridge/lmdbBridge/lmdbMethods/lmdbGetBackup.js
dataLayer/restoreMarker.ts
dataLayer/restoreStaging.ts
dataLayer/rocksdbBackup.ts
dataLayer/schemaDescribe.ts
resources/DatabaseTransaction.ts
Expand Down
69 changes: 53 additions & 16 deletions dataLayer/DESIGN.md
Original file line number Diff line number Diff line change
Expand Up @@ -56,9 +56,44 @@ The certificate verification tables (`hdb_certificate_cache`, `hdb_crl_cache`, `

## RocksDB backup/restore: the restore lock + marker protocol (`dataLayer/restoreMarker.ts`, `dataLayer/rocksdbBackup.ts`)

The `restore_backup` operation restores a user database on a live server by closing it across all
worker threads, purging its directory (`backups.restore` with `purgeAllFiles`), and reloading it.
Three non-obvious mechanics keep that safe:
The `restore_backup` operation restores a user database on a live server by restoring the backup
into a staging directory, closing the database across all worker threads, swapping the staged copy
in, and reloading it. Several non-obvious mechanics keep that safe:

- **Nothing is destroyed until this build has opened the backup (harper#2965,
`dataLayer/restoreStaging.ts`).** `backups.restore`'s `purgeAllFiles` clears its target _before_
copying, and an unreadable backup only fails inside the copy (checksum) or at the next open
(an unsupported table `format_version`, or a transaction log with an unsupported header, which
`open` silently skips). So the restore lands in `` `restore`/<key>.staging ``, every staged
transaction-log store must pass `validateTransactionLogStore`, and the engine is opened and stamped
there; only then is the database published by two renames — itself to `<key>.replaced`, staging to
its name. The first rename is the one-way door the marker protocol guards, so a staging failure
(corrupt, unreadable, `ENOSPC`) is a pre-destruction failure. Readability, not completeness, is the
bar: validation is non-strict, since a torn log tail is something open-time recovery truncates and
the operator has no override to restore past a refusal. The open's proof is bounded: with the
default table cache RocksDB loads only part of a large database's tables at open, so an
unsupported table outside that set still surfaces later, on a cold read. Opening every table
(`maxOpenFiles: -1`) would close that gap at the cost of a descriptor per table, which a large
database can exhaust. Cost: disk for one extra engine copy while staging, written online while
every database on that filesystem keeps serving, so a copy that would not leave headroom (the
larger of 256 MiB and a tenth of the copy) is refused with a 507 before it starts, where the
purge it replaced freed the space first. Two limits: an in-place restore needs room for two
copies even offline, where nothing else is serving, and the check is per restore, so concurrent
restores of different databases on one filesystem are not reserved against each other. Blob
roots are not staged (they span filesystems and the archive capabilities already gate their
encodings), so a blob-copy failure after publication still requires a rerun. A database
directory that is a symlink is refused, since the swap would replace the link with a directory,
and so is one that is a mount point of another filesystem, since staging (beside it) would land
elsewhere and the rename could only fail after a full copy. The check compares devices, so a
same-device bind mount passes it and fails at the first rename, with the database untouched.
**`.replaced` outlives every attempt under a preexisting marker**: a crash between the renames leaves it as the only copy of the database, so
a rerun keeps it until its own replacement publishes (and drops whatever is at the database path
then, a disposable candidate, before the space check; online, only once nothing holds it open). That inference needs `.replaced` never to
outlive its restore, so a finished restore renames it to `<key>.discarded` while its marker still
stands, and only the removal of that may fail quietly; a `.replaced` that no marker accounts for
is refused (409), never trusted. A failed second rename moves it back, and
only a rollback whose directories were fsynced counts as "nothing destroyed". Once staging is
published the marker stays on any later failure, even where nothing was displaced.

- **Two files in an isolated `` `restore` `` directory beside (never inside) the database directory**,
each keyed by `sha256(basename(dbPath)).slice(0,32)`: `<key>.lock`, an OS-level exclusive flock
Expand Down Expand Up @@ -89,15 +124,18 @@ Three non-obvious mechanics keep that safe:
it and broadcasting a reload would surface the earlier attempt's partial/corrupt directory as
healthy. Only a _fresh_ marker on a _previously healthy_ database that failed before destruction is
safe to clear.
- **The ITC close broadcast is normally best-effort, so closure is verified before the purge.** A
- **The ITC close broadcast is normally best-effort, so closure is verified before publication.** A
SCHEMA broadcast (`signalSchemaChange`) usually resolves after remote handlers complete but times
out at 30s "best-effort" and swallows errors. The restore `close` phase is stricter: it waits until
every eligible recipient acknowledges or its port closes, and aborts the restore with a retryable
409 if that does not happen within 30s, because proceeding past an unconfirmed blob-save barrier
would re-open the race the barrier exists to close. A destructive purge still
would re-open the race the barrier exists to close. Publication still
verifies closure independently: `restoreBackup` polls rocksdb-js `registryStatus()` (process-global
across worker threads) until the database path has no open instance, and aborts with a 409 —
_cleaning up the marker, since nothing was destroyed_ — if handles remain.
_cleaning up the marker, since nothing was destroyed_ — if handles remain. The close can only reach
a loaded database, and online staging holds the marker for the whole copy, so a rescan during it
keeps a marked root this thread already has open (`restoreBlocksLoad`) instead of unloading it and
orphaning the handle.
- **The close acknowledgement fences blob saves, deferred reclamation and orphan cleanup, not just
database handles.** A store
handle can close while a `saveBlob` file pipeline it started is still pending, because blob roots
Expand Down Expand Up @@ -128,7 +166,7 @@ Three non-obvious mechanics keep that safe:
correct.** rocksdb-js's registry is process-global but records only a per-path refCount, with no
attribution to a thread or component; Harper keeps no component→database ownership map. So when a
loaded component holds its own handle on the target database, `registryStatus()` stays non-zero,
Harper can neither identify nor force-close that handle, and an in-place purge would corrupt a live
Harper can neither identify nor force-close that handle, and swapping its directory would corrupt a live
instance.
`verifyDatabaseClosed` therefore waits only a short grace period (`DATABASE_CLOSE_WAIT_MS`, for a
just-finished job worker's own close to drain) and then fails fast with a 409 that points at
Expand All @@ -150,13 +188,12 @@ Three non-obvious mechanics keep that safe:
calls `closeLoadedDatabases()` (`resources/databases.ts`) in its `finally`, closing every loaded
user database on that thread (the non-enumerable `system` DB is intentionally skipped), so an
exited job worker leaves no residual handle to be mistaken for a live holder.
- **A restore stamps a new database generation before `completeRestore`.** The restored files carry
the backup's generation, so both paths open the restored directory privately, stamp it and flush
(`stampDatabaseDirectory`, [database generation](../resources/DESIGN.md#database-generation-and-resumable-positions))
inside the destructive section: a failed stamp leaves the marker, and the rerun re-purges and
re-stamps. The stamp is as durable as the marker protocol it runs inside.
- **A restore stamps a new database generation before publishing.** The restored files carry
the backup's generation, so both paths stamp the staged copy and flush
(`stampDatabaseDirectory`, [database generation](../resources/DESIGN.md#database-generation-and-resumable-positions));
that open is also the readability proof above, and a failed stamp destroys nothing.
- **`dropDatabase` and `restore_backup` serialize on the same lock, not a check-then-act probe.**
A drop's `destroy()` interleaving with a restore's purge-and-copy on the same directory would gut
A drop's `destroy()` interleaving with a restore's publication on the same directory would gut
a "successful" restore (or vice versa). `dropDatabase` takes the restore lock for every RocksDB or
LMDB root and publishes a positional `.dropping` marker beside each root before deleting any of
them. A restore in progress makes the acquire fail with 409; `beginRestore` likewise refuses a
Expand All @@ -180,13 +217,13 @@ Three non-obvious mechanics keep that safe:
cheap pre-check that answers a plainly blocked caller without taking a lock.
- **The offline restore probes RocksDB's own `LOCK` file, and fails closed.** The offline path runs
only when the CLI sees no server (a PID heuristic; the PID file is briefly absent
mid-`harper restart`), and `backups.restore`'s `purgeAllFiles` never takes RocksDB's lock — so
before purging, `restoreBackupOffline` opens the database to probe. It now takes the restore
mid-`harper restart`), and publishing by rename never takes RocksDB's lock — so before
staging and again before publishing, `restoreBackupOffline` opens the database to probe. It now takes the restore
lock+marker _before_ probing (so a server that starts afterward sees the marker and refuses to
load), and recognizes the rocksdb-js lock error by message (`isRocksDbLockError`; at 2.5.0, when
this was written, a plain `Error` with no `code`) —
`IO error: While lock file: <db>/LOCK: Resource temporarily unavailable` — aborting with a 409
rather than purging a database another process holds open. Any _other_ open failure
rather than replacing a database another process holds open. Any _other_ open failure
(corrupt/half-restored) is exactly what restore recovers, so only a lock conflict aborts.

Known limitation: the flock is process-owned; if the restore job's worker _thread_ dies without
Expand Down
19 changes: 18 additions & 1 deletion dataLayer/restoreMarker.ts
Original file line number Diff line number Diff line change
Expand Up @@ -13,7 +13,7 @@ import logger from '../utility/logging/harper_logger.ts';
* mutate the same directory concurrently.
*
* Restore metadata lives in an isolated `` `restore` `` directory *beside* the database directory
* (never inside it, since a restore purges the destination). Each database's two files are keyed by
* (never inside it, since a restore replaces the destination). Each database's entries are keyed by
* a hash of the database directory name rather than being suffixed onto the name itself. That keeps
* them out of the database-name namespace — a legal database literally named `orders.restoring`
* would otherwise be mistaken for the restore marker of `orders`, and a 250-character name plus a
Expand Down Expand Up @@ -49,6 +49,8 @@ import logger from '../utility/logging/harper_logger.ts';
* roots, and retrying the same drop removes remaining roots and the blob directories recorded for
* their physical store identities before clearing the markers. The recorded identity is covered by
* a digest, so damaged marker content fails closed instead of redirecting blob deletion.
* - `<meta-dir>/<key>.staging/`, `<key>.replaced/` and `<key>.discarded/` — a restore's proven
* replacement, the database it displaces, and that database once the restore no longer needs it; see `restoreStaging.ts` for when each may be removed.
*/

// The backtick makes this an illegal database name (schemaRegex rejects `/` and backtick only), so
Expand All @@ -57,6 +59,9 @@ export const RESTORE_META_DIR = '`restore`';
export const RESTORE_LOCK_SUFFIX = '.lock';
export const RESTORING_MARKER_SUFFIX = '.restoring';
export const DROPPING_MARKER_SUFFIX = '.dropping';
const RESTORE_STAGING_SUFFIX = '.staging';
const RESTORE_REPLACED_SUFFIX = '.replaced';
const RESTORE_DISCARDED_SUFFIX = '.discarded';
// Deliberately not a `.restoring` suffix: `scanBlockedRestores` selects markers by that suffix, and
// a half-written temp must never be mistaken for one.
const MARKER_TEMP_SUFFIX = '.tmp';
Expand Down Expand Up @@ -90,6 +95,18 @@ export function restoringMarkerPath(dbPath: string): string {
return join(restoreMetaDir(dbPath), restoreMetaKey(dbPath) + RESTORING_MARKER_SUFFIX);
}

export function restoreStagingPath(dbPath: string): string {
return join(restoreMetaDir(dbPath), restoreMetaKey(dbPath) + RESTORE_STAGING_SUFFIX);
}

export function restoreReplacedPath(dbPath: string): string {
return join(restoreMetaDir(dbPath), restoreMetaKey(dbPath) + RESTORE_REPLACED_SUFFIX);
}

export function restoreDiscardedPath(dbPath: string): string {
return join(restoreMetaDir(dbPath), restoreMetaKey(dbPath) + RESTORE_DISCARDED_SUFFIX);
}

export function droppingMarkerPath(dbPath: string): string {
return join(restoreMetaDir(dbPath), restoreMetaKey(dbPath) + DROPPING_MARKER_SUFFIX);
}
Expand Down
Loading
Loading