Add database backups and automatic restore

Daemon job "backup" (online copy, quick_check, rotation), restore of the
newest valid copy when the database cannot be opened, change journal reset
after restore, last_restore in /health, docs and tests.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
This commit is contained in:
ayurishchevandClaude Sonnet 5 committed 2026-09-21 08:55:42 +03:00
1 parent c99665542a
commit 02f7b49e12
10 files changed
+294 -21

No files matched your search

+16 -6
View File
@@ -45,6 +45,7 @@ The system consists of **two independent processes**: the collector daemon (`col
}
```
`ttl_days` - how long an address is kept after it was last seen (default `90`, `0` = keep forever).
Optional `backup_keep` - how many database backups to keep (default `7`); the backup cron is `schedule.backup` (default `30 4 * * *`), see section 9.
Optional `changes_retention_days` - how long the change journal for `/addresses/diff` is kept (default `30`, `0` = forever).
5. **Tests** (optional) run in a container, not in the local `venv`:
@@ -358,7 +359,7 @@ Body:
"cron": "*/15 * * * *"
}
```
*Note: `type` can be `asn` or `fqdn`. The change is written to `config.json`; the collector daemon applies it within 30 seconds (`applied_within_seconds` in the response). An invalid cron expression is rejected by the API (400) and ignored by the daemon.*
*Note: `type` can be `asn`, `fqdn` or `backup` (database backup, see section 9). The change is written to `config.json`; the collector daemon applies it within 30 seconds (`applied_within_seconds` in the response). An invalid cron expression is rejected by the API (400) and ignored by the daemon.*
### Endpoints: Manage ASNs and FQDNs
The lists of monitored sources are stored in `config.json` and can be managed through the API. Reading is open; changing requires the `X-API-Key` header (same token as `POST /schedule`).
@@ -394,7 +395,7 @@ Body is optional: `type` is `asn`, `fqdn` or `all` (default). The API does not c
### Endpoint: Health
**GET** `/health`
Reports the state of the collector daemon (read from `status.json`): `collector_alive`, the cron / `running` / last run / `last_finished` / last error / next run of each job, and the number of stored addresses. The daemon writes a heartbeat every 30 seconds; `collector_alive` is `false` if it is older than 120 seconds or the daemon never ran. `status` is `ok` only if the daemon is alive and no job failed in its last run; otherwise `degraded` (HTTP code is still 200).
Reports the state of the collector daemon (read from `status.json`): `collector_alive`, the cron / `running` / last run / `last_finished` / last error / next run of each job, the number of stored addresses and `last_restore` (`null`, or the record of the last automatic restore of the database from a backup, see section 9). The daemon writes a heartbeat every 30 seconds; `collector_alive` is `false` if it is older than 120 seconds or the daemon never ran. `status` is `ok` only if the daemon is alive and no job failed in its last run; otherwise `degraded` (HTTP code is still 200).
---
@@ -433,7 +434,7 @@ When the collector runs (whether manually or via schedule):
### Scheduler Logic
`collector_daemon.py` uses `APScheduler` (`BlockingScheduler`) in its own process; the API has no scheduler.
1. **Startup**: the daemon takes an exclusive lock (`collector.daemon.lock`, single instance), loads the `schedule` block from `config.json` and creates two independent jobs (`asn_job`, `fqdn_job`; defaults `0 2 * * *` and `0 3 * * *`). Overlapping runs of the same job are not allowed.
1. **Startup**: the daemon takes an exclusive lock (`collector.daemon.lock`, single instance), loads the `schedule` block from `config.json` and creates three independent jobs (`asn_job`, `fqdn_job`, `backup_job`; defaults `0 2 * * *`, `0 3 * * *` and `30 4 * * *`). Overlapping runs of the same job are not allowed.
2. **Runtime updates (POST /schedule)**: the API validates the cron expression and writes `config.json` atomically. Every 30 seconds the daemon compares the `schedule` block with the active one and reschedules the changed job (`reschedule_job`). An invalid cron expression is logged and ignored, the previous schedule stays.
3. **Status**: at the start and end of every run and every 30 seconds the daemon atomically writes `status.json` (jobs state + `updated_at` heartbeat); `GET /health` reads it.
**Manual runs**: every 5 seconds the daemon checks `collect_request.json` (written by `POST /collect`) and starts the requested collections as one-off jobs; one run per type at a time.
@@ -443,7 +444,7 @@ When the collector runs (whether manually or via schedule):
## 8. Application Setup: Docker Compose
One image (`Dockerfile`), two services: `api` (uvicorn) and `collector` (`collector_daemon.py`). They share the named volume `ripe_data` mounted at `/data` (`RIPE_DATA_DIR`), which holds `ripe.db` (with `-wal`/`-shm`), `config.json`, `status.json` and lock files. Containers run as non-root (uid 10001) with a read-only root filesystem, dropped capabilities and rotated logs. Both have healthchecks (`healthcheck.py`).
One image (`Dockerfile`), two services: `api` (uvicorn) and `collector` (`collector_daemon.py`). They share the named volume `ripe_data` mounted at `/data` (`RIPE_DATA_DIR`), which holds `ripe.db` (with `-wal`/`-shm`), `backups/` (database copies, see section 9), `config.json`, `status.json` and lock files. Containers run as non-root (uid 10001) with a read-only root filesystem, dropped capabilities and rotated logs. Both have healthchecks (`healthcheck.py`).
### Start
```bash
@@ -517,5 +518,14 @@ Automatic: on the first start of any process (API, collector or CLI) `data.json`
### Rollback to the JSON version
Stop both services, rename the `*.migrated-*` files back to `data.json` / `fqdn_data.json`, remove `ripe.db*`, start the previous version of the code. Addresses collected after the migration are lost in that case.
### Backup
Copy the database consistently with `sqlite3 ripe.db ".backup ripe.db.bak"` (do not copy `ripe.db` alone while the services are running - part of the data may still be in the `-wal` file).
### Backup and automatic restore
**Backup job.** The collector daemon runs the job `backup` (default `30 4 * * *`, change it with `POST /schedule` and `"type": "backup"` or in `schedule.backup`). It makes an online copy of `ripe.db` (safe while the services run), checks the live database and the copy with `PRAGMA quick_check`, and keeps the last `backup_keep` copies (default 7; older ones are deleted only after a new copy succeeded). If the live database fails the check or the copy is bad, no copy is written, the older copies stay, and `GET /health` shows `degraded` (`jobs.backup.last_error`).
Copies are named `ripe-<UTC time>.db` and stored in `RIPE_BACKUP_DIR` (default `<RIPE_DATA_DIR>/backups`, i.e. `/data/backups` in Docker). **By default they are on the same volume as the database**: this protects against a corrupted file, not against losing the volume. For that, mount a separate volume/host directory and set `RIPE_BACKUP_DIR` to it, or copy the directory elsewhere regularly (e.g. `docker compose cp collector:/data/backups ./backups`).
**Automatic restore.** If `ripe.db` cannot be opened as a database (any process: API, daemon, CLI), the file is moved to `ripe.db.corrupt-<timestamp>` and the newest copy that passes the integrity check is put in its place; concurrent processes are serialized with a lock. The event is logged (ERROR), written to `last_restore.json` and shown in `GET /health` as `last_restore`. Data collected after that copy is lost; the collector gathers it again on the next runs. If there is no valid copy, the behaviour is as before: the API answers `503`.
- The change journal goes back with the copy, so after a restore all cursors and times issued earlier answer `410` on `/addresses/diff` (the client makes a full download). The journal counter is shifted by 1,000,000 for that (a heuristic: it assumes fewer changes than that between two copies).
- Only corruption detected **when the database is opened** is restored automatically. Damage inside the file shows up as `503` on reads and is caught by the `backup` job (`quick_check`); restore it by hand: stop both services, keep the damaged `ripe.db*`, copy the chosen `backups/ripe-*.db` to `ripe.db`, start the services.
Manual consistent copy: `sqlite3 ripe.db ".backup ripe.db.bak"` (do not copy `ripe.db` alone while the services are running - part of the data may still be in the `-wal` file).