Skip to content
Draft
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
14 changes: 10 additions & 4 deletions reference/backups/operations.md
Original file line number Diff line number Diff line change
Expand Up @@ -108,17 +108,23 @@ harper purge_backups database=data keep_count=3

## `restore_backup`

<VersionBadge version="v5.2.0" /> <EngineBadge engines="RocksDB" />
<VersionBadge version="v5.2.0" /> <VersionBadge type="changed" version="v5.3.2" /> <EngineBadge engines="RocksDB" />

Restores a database in place from a managed backup. `backup_id` defaults to the latest backup. The audit/transaction log is restored alongside the data, and — for a database with file-backed blobs — the blob roots are purged and rewritten from the backup's blob snapshot.

Through a running server this runs as a background [job](../operations-api/operations.md#jobs): Harper closes the database across all worker threads, restores it, and reloads it. This works only when no loaded component is holding the database open — if one does, the job ends in `ERROR` (surfaced by [`get_job`](../operations-api/operations.md#get_job)) telling you to restore offline. Restoring the `system` database is rejected up front, before a job is created, as is `target_database` while the server is running. These cases require running the command from the CLI with the server stopped. See [when can a database be restored?](./overview.md#when-can-a-database-be-restored)
Through a running server this runs as a background [job](../operations-api/operations.md#jobs): Harper stages the backup (see below), closes the database across all worker threads, swaps the staged copy in, and reloads it. This works only when no loaded component is holding the database open — if one does, the job ends in `ERROR` (surfaced by [`get_job`](../operations-api/operations.md#get_job)) telling you to restore offline. Restoring the `system` database is rejected up front, before a job is created, as is `target_database` while the server is running. These cases require running the command from the CLI with the server stopped. See [when can a database be restored?](./overview.md#when-can-a-database-be-restored)
Comment thread
cb1kenobi marked this conversation as resolved.

As of v5.3.2, a restore checks that this version of Harper can read the backup's database files before it replaces them. Harper first restores the backup into a staging directory beside the database, checks its transaction logs, and opens it. Only then does it swap the staged copy in for the database. A backup whose database files are corrupt, or written in a format this version cannot read, is refused: the restore fails with an error saying the database `was not modified`, followed by the cause, and the database keeps its data and its blobs. A rerun over an earlier restore that did not finish is the exception: the database is still incomplete from that attempt, so it stays blocked, and the error says to rerun the restore instead. The check covers the database files and transaction logs only. File-backed blobs are not staged or checked; as before, they are rewritten from the backup's blob snapshot after the swap, so a failure there leaves the database marked incompletely restored until a rerun succeeds.

Through a running server, the database stays open and keeps serving reads and writes while the backup is staged, which can take minutes for a large database. Writes made during that time succeed, and are then replaced by the restore along with everything else written after the backup was taken.

Staging needs free space beside the database directory for the backup's restored database files, including its transaction logs, which `list_backups`'s `size` leaves out. A restore checks this before staging, and if the copy would not leave headroom (the larger of 256 MiB and a tenth of the copy) for the databases still serving on that filesystem, it fails with status code 507 and the database is not modified. The database directory must be a real directory on the same filesystem as its parent: a database whose directory is a symbolic link or a mount point is refused before anything is staged. Point the configured database path at the link's target, or mount the volume at the parent directory instead. If an earlier restore of that database did not finish, keep Harper stopped, make the change, and rerun the restore offline before starting Harper again: moving the path also moves the restore's tracking files, so a restart in between would load the half-restored directory.

As of v5.3.1, subscriptions opened before a restore do not carry over to the restored data. When Harper reloads the database, each one ends, and its last message is a `DatabaseGenerationChangedError` (status code 409, code `DATABASE_GENERATION_CHANGED`). Resubscribe to resynchronize against the restored state. A restore that fails before it changes anything reloads the database as it was; those subscriptions end instead with a retryable `DatabaseClosingError` (status code 503, code `DATABASE_CLOSING`); resubscribe to continue.
As of v5.3.1, subscriptions opened before a restore do not carry over to the restored data. When Harper reloads the database, each one ends, and its last message is a `DatabaseGenerationChangedError` (status code 409, code `DATABASE_GENERATION_CHANGED`). Resubscribe to resynchronize against the restored state. A restore that fails after closing the database but before changing it reloads the database as it was; those subscriptions end instead with a retryable `DatabaseClosingError` (status code 503, code `DATABASE_CLOSING`); resubscribe to continue. As of v5.3.2, a restore refused before the database is closed (an unreadable backup, too little space, or a symbolic-link or mount-point directory) leaves the database and its subscriptions running.

From the CLI with the server stopped, `target_database=<name>` restores into a separate database instead of overwriting the source. The target must not already exist, or must be an empty directory; Harper picks the new database up on the next start.

If a restore is interrupted before it completes (crash, power loss), Harper marks the database as incompletely restored and refuses to load it on the next start, logging an incomplete-restore error. Recover by rerunning `restore_backup` for the same database and `backup_id` — do not try to load or hand-repair the directory.
If a restore is interrupted before it completes (crash, power loss), Harper marks the database as incompletely restored and refuses to load it on the next start, logging an incomplete-restore error. Recover by rerunning `restore_backup` for the same database and `backup_id` — do not try to load or hand-repair the directory. An interruption while the backup is still being staged leaves the database's files untouched, but the database stays blocked until the rerun. An interruption after the swap has begun leaves the previous database's files (not its blobs) set aside, and Harper keeps them until a rerun completes. A restore that is not interrupted removes them when it finishes.

```json
{ "operation": "restore_backup", "database": "data", "backup_id": 1 }
Expand Down
5 changes: 3 additions & 2 deletions reference/backups/overview.md
Original file line number Diff line number Diff line change
Expand Up @@ -38,11 +38,12 @@ The offline path matters for restore. RocksDB is single-writer, so an in-place r
- **Backups live on the node that created them.** The backup repository is a local directory. RocksDB shares files across backup IDs (a backup ID is not a self-contained folder), so disaster-recovery copies must take the entire per-database repository — `<backupPath>/<database>` — not an individual backup, and must do so while no backup operation is running, or from an atomic filesystem snapshot. A live recursive copy can race `create_backup`/`delete_backup`/`purge_backups` and produce an unrestorable copy. Alternatively, use `get_backup` to pull a snapshot from a running server.
- **`get_backup` always streams the current state.** It cannot download a historical managed backup; to move a retained backup off-host, copy its whole per-database repository as above.
- **A restore is a point-in-time rollback.** In a replicated cluster, coordinate a restore with replication before bringing the node back.
- **A restore stages a copy first.** <VersionBadge type="changed" version="v5.3.2" /> A restore stages the backup's database files beside the database directory and checks them before replacing the database, so the parent directory's filesystem needs room for them, plus headroom, or the restore is refused with status code 507. An online restore keeps serving the database while it stages, and writes made during staging are replaced by the restore. Blobs are not staged or checked. A database directory that is a symbolic link or a mount point cannot be restored in place. See [`restore_backup`](./operations.md#restore_backup).
- **An interrupted restore leaves the database unloadable.** If `restore_backup` is interrupted before completing (crash, power loss), Harper marks the database as incompletely restored and skips loading it on the next start, logging an incomplete-restore error. Rerun `restore_backup` for the same database and `backup_id` to recover; do not load or hand-repair the directory.

### When can a database be restored?

An in-place restore purges and rewrites the database's files, which requires the database to be fully closed first. Whether a restore can run online depends on what is holding the database open:
An in-place restore replaces the database's files, which requires the database to be fully closed first. Whether a restore can run online depends on what is holding the database open:

| Database | Online `restore_backup` (server running) | Offline `harper restore_backup` (server stopped) |
| -------------------------------------------------- | ---------------------------------------- | ------------------------------------------------ |
Expand Down Expand Up @@ -81,7 +82,7 @@ Restore the latest backup, or pass `backup_id=<id>` for an earlier one:
harper restore_backup database=data
```

With the server running, this restores the database in place — Harper closes the database across its worker threads, restores it, and reloads it — as long as nothing is holding the database open (see [when can a database be restored?](#when-can-a-database-be-restored)). With the server stopped, the same command restores the files directly and works for any database.
With the server running, this restores the database in place — Harper stages and checks the backup while the database keeps serving, then closes the database across its worker threads, swaps the staged copy in, and reloads it — as long as nothing is holding the database open (see [when can a database be restored?](#when-can-a-database-be-restored)). With the server stopped, the same command restores the files directly and works for any database.

## Example: download a snapshot and restore it manually

Expand Down
4 changes: 4 additions & 0 deletions release-notes/v5-lincoln/5.3.md
Original file line number Diff line number Diff line change
Expand Up @@ -132,6 +132,10 @@ Each SSH key's block in `<rootPath>/ssh/config` now runs from its `#<name>` line

An online `restore_backup` now ends every subscription to the database that was opened before the restore, once Harper reloads the database. Each one's last message is a `DatabaseGenerationChangedError` (status code 409, code `DATABASE_GENERATION_CHANGED`). Resubscribe to resynchronize against the restored state. A restore that fails before it changes anything reloads the database as it was, and its subscriptions end with a retryable `DatabaseClosingError` (status code 503, code `DATABASE_CLOSING`) instead. Previously such a subscription could stall silently, or start receiving the restored database's writes once another subscriber attached. See [`restore_backup`](/reference/v5/backups/operations#restore_backup).

### Restore Checks a Backup Before Replacing the Database (5.3.2)

`restore_backup` no longer deletes a database before checking that the backup can be read. It restores the backup into a staging directory beside the database, checks its transaction logs and opens it, and only then swaps it in. A backup whose database files are corrupt, or written in a format this version cannot read, is now refused with the database left as it was; previously the restore deleted the database first and then failed. Staging needs free space beside the database directory for the backup's database files and transaction logs; a restore that would not leave headroom for the databases still serving there (the larger of 256 MiB and a tenth of the copy) is refused with status code 507 before it starts. An online restore keeps serving the database while it stages, and writes made during staging are replaced by the restore. A database whose directory is a symbolic link or a mount point is now refused before anything is staged, so a restore that used to run on such a layout fails until the configured path points at a real directory on its parent's filesystem; if an earlier restore of that database did not finish, rerun it offline right after the change, before starting Harper. A restore refused before the database is closed leaves its subscriptions running. File-backed blobs are still rewritten after the swap and are not checked. See [`restore_backup`](/reference/v5/backups/operations#restore_backup).

## Security

### Route-Owned Authentication
Expand Down
Loading