Collectors record the result of every source (schema v3, table source_status); /health reports failing sources (failures in a row >= source_failure_threshold) and turns degraded; GET /sources shows the full state; runs end with a summary log line instead of "data saved". Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
31 KiB
RIPE AS CIDR & FQDN IP Collector
This project collects CIDR prefixes for specified Autonomous Systems (AS) from the RIPE NCC API and resolves IP addresses for specified FQDNs. It accumulates these addresses over time, maintaining a history of discovered prefixes. It also provides a FastAPI-based HTTP interface to retrieve the collected data.
The system consists of two independent processes: the collector daemon (collector_daemon.py, runs the schedule) and the API server (api_server.py, serves data and edits the config). They communicate only through files: the collected addresses live in the SQLite database ripe.db, the rest in JSON (config.json, status.json). See section 9.
Quick start: the whole system (API + collector) runs with
docker compose up -d, see section 8. Sections 1-4 describe the manual installation (venv + systemd/OpenRC), which remains supported.
Upgrading from a version where the scheduler lived in the API: running only
ripe-apiis no longer enough - without theripe-collectorservice nothing is collected.GET /healthreportscollector_alive: false/degradedin that case.
1. Preparation and Installation
Prerequisites
- Python 3.8+
pipandvenv
Installation Steps
-
Clone the repository (or copy the files) to your desired location, e.g.,
/opt/ripe_collector.mkdir -p /opt/ripe_collector cd /opt/ripe_collector # Copy files: *.py (api_server, cidr_collector, collector_daemon, db, formatters, storage, healthcheck), # requirements.txt, constraints.txt, config.json -
Create a Virtual Environment:
python3 -m venv venv -
Install Dependencies:
source venv/bin/activate pip install -c constraints.txt -r requirements.txt # exact versions, see "Dependency versions" below deactivate -
Initial Configuration: Edit
config.jsonto set your initial ASNs and FQDNs.{ "asns": [62041], "fqdns": ["google.com"], "ttl_days": 90 }ttl_days- how long an address is kept after it was last seen (default90,0= keep forever). Optionalsource_failure_threshold(default3) - how many runs in a row a source may fail beforeGET /healthreportsdegraded(see the sources status below). Optionalripestat_sourceapp(defaultripe-cidr-collector) - the application name sent to RIPEstat assourceapp; RIPEstat asks clients to identify themselves, you can add a contact (my-collector admin@example.org). Optionalallow_non_global_ips(defaultfalse) - by default only public addresses resolved from FQDNs are stored; loopback, private, link-local and unspecified addresses (127.0.0.1,10.x,0.0.0.0, ...) are ignored and logged. Settruefor internal names. Optionalbackup_keep- how many database backups to keep (default7); the backup cron isschedule.backup(default30 4 * * *), see section 9. Optionalchanges_retention_days- how long the change journal for/addresses/diffis kept (default30,0= forever). -
Tests (optional) run in a container, not in the local
venv:docker build -f Dockerfile.test -t ripe-collector-test . docker run --rm ripe-collector-testThe image (
Dockerfile.test,python:3.11-slim) contains only the code and dev dependencies - production data (ripe.db,data.json,fqdn_data.json) is excluded via.dockerignore. -
Knowledge graph (optional):
/graphify .(Claude Code skill) buildsgraphify-out/-graph.html(interactive),GRAPH_REPORT.md,graph.json- from code anddocs/; refresh with/graphify . --updateafter changes. The directory is excluded from git and from Docker images. -
Repository:
mainbranch, remoteorigin=https://artstore.rxmsk.ru/ayurishchev/ripe-cidr-collector.git. Committed: code, tests, Docker files,config.json(initial sources),.env.example,docs/. Not committed (.gitignore):venv/,.env, databases*.db*, collected datadata.json/fqdn_data.json,graphify-out/, service files. Every change follows the plan -> implementation -> summary flow indocs/and gets its own commit.
Dependency versions
requirements.txt / requirements-dev.txt list direct dependencies with exact versions (==); constraints.txt is the full pip freeze (including transitive packages) of the image the tests pass on. Both Dockerfiles install with -c constraints.txt. To update: change the versions in a container, run the tests, regenerate constraints.txt (docker run --rm --entrypoint pip <test-image> freeze > constraints.txt) and commit all three files together.
2. Running the Collector
The collector normally runs as the daemon collector_daemon.py (see sections 3-4): it keeps the schedule from config.json (schedule.asn, schedule.fqdn) and picks up changes made via POST /schedule within 30 seconds, without a restart. Only one daemon instance can run at a time (a second one exits with an error).
/opt/ripe_collector/venv/bin/python3 /opt/ripe_collector/collector_daemon.py
The script cidr_collector.py is used for manual runs and for managing the lists (see section 6).
Manual Run
/opt/ripe_collector/venv/bin/python3 /opt/ripe_collector/cidr_collector.py run
Alternative: cron instead of the daemon
If you prefer cron, run the collector once per day (do not combine with the daemon - it is redundant). To run daily at 02:00 AM:
Run the job as the same user as the API (
crontab -u ripe -e), otherwise data files may end up with the wrong owner.
- Open crontab:
crontab -e - Add the line:
0 2 * * * /opt/ripe_collector/venv/bin/python3 /opt/ripe_collector/cidr_collector.py run >> /var/log/ripe_collector.log 2>&1
3. Application Setup: Systemd (Ubuntu, Debian)
This section describes how to run two system services: the API Server (api_server.py) and the collector daemon (collector_daemon.py).
Create Service File
Create a dedicated user and the token file (the token protects POST /schedule):
useradd --system --home /opt/ripe_collector --shell /usr/sbin/nologin ripe
chown -R ripe: /opt/ripe_collector
echo "RIPE_API_TOKEN=$(openssl rand -hex 32)" > /etc/ripe-api.env
chmod 600 /etc/ripe-api.env
Create /etc/systemd/system/ripe-api.service:
[Unit]
Description=RIPE CIDR Collector API
After=network.target
[Service]
User=ripe
EnvironmentFile=/etc/ripe-api.env
WorkingDirectory=/opt/ripe_collector
NoNewPrivileges=true
ProtectSystem=strict
ReadWritePaths=/opt/ripe_collector
ExecStart=/opt/ripe_collector/venv/bin/uvicorn api_server:app --host 0.0.0.0 --port 8000
Restart=always
[Install]
WantedBy=multi-user.target
Create /etc/systemd/system/ripe-collector.service (the daemon does not need the token):
[Unit]
Description=RIPE CIDR Collector daemon
After=network.target
[Service]
User=ripe
WorkingDirectory=/opt/ripe_collector
NoNewPrivileges=true
ProtectSystem=strict
ReadWritePaths=/opt/ripe_collector
ExecStart=/opt/ripe_collector/venv/bin/python collector_daemon.py
Restart=always
[Install]
WantedBy=multi-user.target
Enable and Start
# Reload systemd
sudo systemctl daemon-reload
# Enable services to start on boot and start them immediately
sudo systemctl enable --now ripe-collector ripe-api
# Check status
sudo systemctl status ripe-collector ripe-api
4. Application Setup: RC-Script (Alpine Linux)
For Alpine Linux using OpenRC.
Create Init Script
Create /etc/init.d/ripe-api:
#!/sbin/openrc-run
name="ripe-api"
description="RIPE CIDR Collector API"
command="/opt/ripe_collector/venv/bin/uvicorn"
# --host and --port and module:app passed as arguments
command_args="api_server:app --host 0.0.0.0 --port 8000"
command_background="yes"
pidfile="/run/${RC_SVCNAME}.pid"
directory="/opt/ripe_collector"
depend() {
need net
}
Add the token and a non-root user (/etc/conf.d/ripe-api is read by OpenRC):
adduser -S -h /opt/ripe_collector ripe
chown -R ripe /opt/ripe_collector
echo 'export RIPE_API_TOKEN="<random-token>"' > /etc/conf.d/ripe-api
chmod 600 /etc/conf.d/ripe-api
and add command_user="ripe" to the init script.
Create the collector daemon script /etc/init.d/ripe-collector:
#!/sbin/openrc-run
name="ripe-collector"
description="RIPE CIDR Collector daemon"
command="/opt/ripe_collector/venv/bin/python"
command_args="collector_daemon.py"
command_background="yes"
command_user="ripe"
pidfile="/run/${RC_SVCNAME}.pid"
directory="/opt/ripe_collector"
depend() {
need net
}
Make Executable
chmod +x /etc/init.d/ripe-api /etc/init.d/ripe-collector
Enable and Start
# Add to default runlevel
rc-update add ripe-collector default
rc-update add ripe-api default
# Start services
service ripe-collector start
service ripe-api start
# Check status
service ripe-collector status
service ripe-api status
5. API Usage Documentation
The API runs by default on port 8000. It allows retrieving the collected data in a flat JSON list.
Security model: read endpoints (/addresses, /schedule GET, /health) are open, since consumers (routers, firewalls) usually can only do a plain GET - restrict access to port 8000 with a firewall. POST /schedule requires the header X-API-Key: <RIPE_API_TOKEN>; if the RIPE_API_TOKEN environment variable is not set, POST is disabled (503). If the data files are unreadable, the API answers 503.
Base URL
http://<server-ip>:8000
Endpoint: Get Addresses
GET /addresses
Retrieves the list of collected IP addresses/CIDRs.
| Parameter | Type | Required | Default | Description |
|---|---|---|---|---|
type |
string | No | all |
Filter by source type. Options: cidr (ASNs only), fqdn (Domains only), all (Both). |
format |
string | No | json |
Output format: json, nftables, mikrotik, bird, frr (see below). |
ip_version |
string | No | all |
4, 6 or all. |
aggregate |
bool | No | false |
Collapse overlapping/adjacent prefixes (/23 + /24 -> /23); a host IP inside a wider prefix is dropped. Applied after cidr and fqdn sources are merged. |
name |
string | No | ripe |
Name of the list/set in generated configs ([A-Za-z][A-Za-z0-9_]{0,31}). |
With the defaults the response is the same flat JSON list as before. Other formats return text/plain scripts that are safe to re-apply (they replace the previous list).
Example 1: Get All Addresses (Default)
Request:
curl -X GET "http://localhost:8000/addresses"
Response (JSON):
[
"142.250.1.1",
"149.154.160.0/22",
"149.154.160.0/23",
"2001:4860:4860::8888",
"91.108.4.0/22"
]
Example 2: Get Only CIDRs (from ASNs)
Request:
curl -X GET "http://localhost:8000/addresses?type=cidr"
Response (JSON):
[
"149.154.160.0/22",
"149.154.160.0/23",
"91.108.4.0/22"
]
Example 3: Get Only Resolved IPs (from FQDNs)
Request:
curl -X GET "http://localhost:8000/addresses?type=fqdn"
Response (JSON):
[
"142.250.1.1",
"2001:4860:4860::8888"
]
Ready-to-use configuration formats
format |
Result | How to apply |
|---|---|---|
nftables |
table inet <name> with interval sets <name>_v4 / <name>_v6 (auto-merge); other objects of the table are untouched |
curl -s "$URL/addresses?format=nftables" | nft -f - |
mikrotik |
/ip firewall address-list and /ipv6 firewall address-list named <name> (old entries are removed first, so there is a short window with an empty list) |
/tool fetch url="$URL/addresses?format=mikrotik" dst-path=ripe.rsc then /import ripe.rsc |
bird |
BIRD 2 prefix sets define <NAME>_V4 = [...], <NAME>_V6 for use in filters (net ~ RIPE_V4) |
save to a file, include it, birdc configure |
frr |
ip prefix-list <name>_v4 / ipv6 prefix-list <name>_v6 with permit entries |
curl -s "$URL/addresses?format=frr" > ripe.conf && vtysh -f ripe.conf |
Examples:
# IPv4 only, aggregated, as an nftables set named "tg"
curl -s "http://localhost:8000/addresses?format=nftables&ip_version=4&aggregate=true&name=tg"
A version without addresses is omitted from the script. The bird and nftables outputs were syntax-checked with bird -p and nft -c.
Endpoint: Changes since the last sync (diff)
GET /addresses/diff?since=<cursor | time>[&type=cidr|fqdn|all][&ip_version=all|4|6] - only what was added and removed, instead of the full list. Read access is open, like /addresses.
curl "http://localhost:8000/addresses/diff?since=0"
# {"since":"0","now":"2026-09-21T04:09:44Z","cursor":42,"added":["3.0.0.0/24"],"removed":["2.0.0.0/24"]}
curl "http://localhost:8000/addresses/diff?since=42" # next sync: use the returned cursor
sinceis a cursor from the previous answer (recommended: exact, independent of clocks) or an ISO 8601 time (no time zone = UTC; write+in a URL as%2B). Time is compared with millisecond precision, borders are inclusive, so an entry may be reported twice - repeating it is harmless. A cursor is at most 30 ASCII digits; anything else that is not an ISO 8601 time (including out-of-range dates) is400.- The result is the net effect: an address added and removed (or the other way) within the interval is not reported; an address that another source still holds is not reported as removed. Both CIDRs and IPs of FQDNs count. Output is JSON only, no aggregation.
- The first sync: take the full
/addresseslist and the cursor from its response headerX-Changes-Cursor(taken in the same database snapshot as the data, so the two match exactly), then call/addresses/diff?since=<that cursor>regularly. 400- invalidsince;410 Gone-sinceis older than the journal (seechanges_retention_days) or the cursor does not belong to this database: fetch the full/addresseslist and continue from the newcursor.
Endpoint: Manage Schedule
GET /schedule
Returns the current cron schedules.
POST /schedule (requires X-API-Key)
Updates the schedule for a specific collector type.
curl -X POST http://localhost:8000/schedule -H "X-API-Key: $RIPE_API_TOKEN" \
-H "Content-Type: application/json" -d '{"type": "asn", "cron": "*/15 * * * *"}'
Body:
{
"type": "asn",
"cron": "*/15 * * * *"
}
Note: type can be asn, fqdn or backup (database backup, see section 9). The change is written to config.json; the collector daemon applies it within 30 seconds (applied_within_seconds in the response). An invalid cron expression is rejected by the API (400) and ignored by the daemon.
Endpoints: Manage ASNs and FQDNs
The lists of monitored sources are stored in config.json and can be managed through the API. Reading is open; changing requires the X-API-Key header (same token as POST /schedule).
| Method and path | Description |
|---|---|
GET /asns, GET /fqdns |
Current lists: {"asns": [...]}, {"fqdns": [...]} |
POST /asns {"asn": 62041} |
Add an ASN (1..4294967295). 201 if added, 200 if it was already there. |
POST /fqdns {"fqdn": "example.com"} |
Add a domain (lower-cased, trailing dot removed; IP addresses and invalid labels are rejected with 422). |
DELETE /asns/{asn}?purge=false, DELETE /fqdns/{fqdn}?purge=false |
Remove a source (404 if unknown). |
curl -X POST http://localhost:8000/asns -H "X-API-Key: $RIPE_API_TOKEN" \
-H "Content-Type: application/json" -d '{"asn": 62041}'
curl -X DELETE "http://localhost:8000/fqdns/example.com?purge=true" -H "X-API-Key: $RIPE_API_TOKEN"
- A new source is collected on the next scheduled run of the daemon; to collect it right away call
POST /collect(see below). - What happens to the data of a removed source: by default it is kept and its addresses expire by
ttl_days(the collector keeps ageing entries of sources that are no longer configured and deletes the entry when it is empty). Withpurge=truethe collected addresses are deleted immediately and disappear from/addresses. Withttl_days = 0nothing expires - usepurge=true. - The CLI commands (
add,remove,add-fqdn,remove-fqdn) use the same locked, atomic config update, so they are safe to use alongside the API.
Endpoint: Collect now
POST /collect (requires X-API-Key) - starts a collection immediately instead of waiting for the schedule.
curl -X POST http://localhost:8000/collect -H "X-API-Key: $RIPE_API_TOKEN" \
-H "Content-Type: application/json" -d '{"type": "asn"}'
Body is optional: type is asn, fqdn or all (default). The API does not collect itself: it passes the request to the collector daemon through the file collect_request.json, which the daemon checks every 5 seconds, so the collection starts within ~5 s. Answer 202 lists the requested types; follow the progress in GET /health (running, last_run, last_finished, last_error of each job).
- Repeated requests are merged into one; if a collection of that type is already running (scheduled or manual) the new start is skipped.
- If the collector daemon is not running (no fresh heartbeat), the answer is
503and nothing is queued.
Endpoint: Health
GET /health
Reports the state of the collector daemon (read from status.json): collector_alive, the cron / running / last run / last_finished / last error / next run of each job, the number of stored addresses and last_restore (null, or {at, backup} - time and file name of the last automatic restore from a backup; directories and the error text stay in last_restore.json) and db_recreated (null, or {at, pending} after the database was recreated without a backup; pending: true makes status degraded), see section 9. The daemon writes a heartbeat every 30 seconds; collector_alive is false if it is older than 120 seconds or the daemon never ran. status is ok only if the daemon is alive and no job failed in its last run; otherwise degraded (HTTP code is still 200).
sources shows the state of the data sources: for asn and fqdn the number of configured sources (total) and the list of failing ones (failing: source, failures, last_success, error_kind). A source is failing when its last source_failure_threshold runs in a row failed (default 3, so a single RIPEstat hiccup does not turn the status yellow); any failing source makes status degraded. error_kind is a short category without details (timeout, connection, http_429, http_4xx, http_5xx, invalid_response, dns, no_global_addresses, unknown); the full error text stays in the collector log.
Endpoint: Sources status
GET /sources (open) - the state of every configured source: kind, source, addresses (how many values it gave), last_attempt, last_success, failures (runs failed in a row), error_kind, plus failure_threshold. A source that was never polled has empty state; a source that "succeeds" but gives addresses: 0 is easy to spot here. The state is kept in the database (table source_status), survives restarts and is included in backups; it is removed with the source (purge or the next collection after removal).
6. Advanced CLI Usage
The collector script supports running modes independently:
# Run both (Default)
python3 cidr_collector.py run
# Run only ASN collection
python3 cidr_collector.py run --mode asn
# Run only FQDN collection
python3 cidr_collector.py run --mode fqdn
7. Internal Logic & Architecture
Collector Logic
When the collector runs (whether manually or via schedule):
- Instantiation: Creates a new instance of
CIDRCollectororFQDNCollector. This forces a fresh read ofconfig.json, ensuring any added ASNs/FQDNs are immediately processed. - Fetching:
- ASN: Queries RIPE NCC API (
stat.ripe.net) with thesourceappparameter. Connection failures and codes 429/500/502/503/504 are retried up to 3 times with growing pauses (0, 2, 4 s; aRetry-Afterpause is capped at 30 s), 10 s per attempt - at worst about 50 s per ASN. If all attempts fail the ASN is skipped for this run and nothing is deleted; the failure is recorded forGET /healthandGET /sources. Every run ends with one summary line in the log (ASN collection finished: 6 sources, 5 ok, 1 failed, +12/-3 prefixes). - FQDN: Uses Python's
socket.getaddrinfoto resolve A and AAAA records. Non-public addresses are dropped (seeallow_non_global_ips); if nothing is left the run is treated like a DNS failure (nothing is deleted).
- ASN: Queries RIPE NCC API (
- Merge in one transaction: the fetched addresses are merged into the SQLite table
addresses(db.py) in a single write transaction, so readers (the API) never see a half-updated state. - Accumulation with TTL: Each address has
first_seen/last_seen.- New addresses are inserted; already seen ones get
last_seenrefreshed. - Addresses not seen for longer than
ttl_daysare removed - but only after a successful fetch. A RIPE/DNS failure never deletes anything. - Sources that are no longer configured are not fetched any more; their addresses simply expire by TTL.
- New addresses are inserted; already seen ones get
- Persistence: SQLite in WAL mode - the API reads while the collector writes. A corrupted database file is renamed to
ripe.db.corrupt-<timestamp>and the API answers503instead of serving wrong data.
Scheduler Logic
collector_daemon.py uses APScheduler (BlockingScheduler) in its own process; the API has no scheduler.
- Startup: the daemon takes an exclusive lock (
collector.daemon.lock, single instance), loads thescheduleblock fromconfig.jsonand creates three independent jobs (asn_job,fqdn_job,backup_job; defaults0 2 * * *,0 3 * * *and30 4 * * *). Overlapping runs of the same job are not allowed. - Runtime updates (POST /schedule): the API validates the cron expression and writes
config.jsonatomically. Every 30 seconds the daemon compares thescheduleblock with the active one and reschedules the changed job (reschedule_job). An invalid cron expression is logged and ignored, the previous schedule stays. - Status: at the start and end of every run and every 30 seconds the daemon atomically writes
status.json(jobs state +updated_atheartbeat);GET /healthreads it. Manual runs: every 5 seconds the daemon checkscollect_request.json(written byPOST /collect) and starts the requested collections as one-off jobs; one run per type at a time. - Concurrency: a running job completes normally when the schedule changes; the new schedule applies to the next calculated run time. On
SIGTERMthe daemon exits after the running collection finishes. Restarting the API does not affect collection.
8. Application Setup: Docker Compose
One image (Dockerfile), two services: api (uvicorn) and collector (collector_daemon.py). They share the named volume ripe_data mounted at /data (RIPE_DATA_DIR), which holds ripe.db (with -wal/-shm), backups/ (database copies, see section 9), config.json, status.json and lock files. Containers run as non-root (uid 10001) with a read-only root filesystem, dropped capabilities and rotated logs. Both have healthchecks (healthcheck.py).
Start
cp .env.example .env
# set RIPE_API_TOKEN (openssl rand -hex 32) and, if needed, TZ / API_PORT
docker compose up -d
docker compose ps # both services become "healthy" within about a minute
docker compose logs -f collector
Then add sources through the API (section 5), e.g. POST /asns with the token. With an empty volume the lists are empty.
Settings (.env)
| Variable | Default | Description |
|---|---|---|
RIPE_API_TOKEN |
- | Token for changing requests (X-API-Key). Without it they are disabled (503). |
TZ |
UTC |
Time zone in which the cron schedules are evaluated (e.g. Europe/Moscow). |
API_PORT |
8000 |
Host port of the API. |
Migrating existing data into the volume
Run once, from the directory with your current config.json, data.json, fqdn_data.json, before the first up (the JSON data files are imported into ripe.db automatically on the first start, see section 9):
docker compose create
docker volume ls | grep ripe_data # the volume is named <project>_ripe_data, e.g. ripe_cidr_collector_ripe_data
docker run --rm -v <project>_ripe_data:/data -v "$PWD":/src:ro alpine \
sh -c 'cp /src/config.json /src/data.json /src/fqdn_data.json /data/ && chown 10001 /data/*.json'
docker compose up -d
Operations
docker compose build && docker compose up -d # update after code changes
docker compose stop collector # graceful: a running collection finishes (up to 60s)
docker compose down # stop; data stays in the volume (add -v to delete it)
Notes and risks
- Port 8000 is published without TLS, so
X-API-Keytravels in clear text. Restrict access with a firewall or put a TLS reverse proxy in front (bind the port to127.0.0.1by changingportsindocker-compose.yml). - The image copies all
*.pymodules from the project root and imports the main ones during the build, so a missing or broken module fails the build instead of the start. - Dependencies are pinned (see "Dependency versions" below), so rebuilds give the same libraries. The base image
python:3.11-slimis not pinned by digest: Python patch releases arrive on rebuild. - Files written by the services in the volume (
config.json,status.json) have mode600; both services run as the same user. docker composeuses the image nameripe-cidr-collector;Dockerfile.testis used only for running the tests (section 1, step 5).
9. Data Storage (SQLite)
Collected addresses are stored in ripe.db (in the project directory, or in RIPE_DATA_DIR - the /data volume in Docker). config.json (sources, schedule, ttl_days) and status.json (collector heartbeat) stay JSON.
CREATE TABLE addresses (
kind TEXT NOT NULL, -- 'asn' or 'fqdn'
source TEXT NOT NULL, -- '62041' or 'example.com'
value TEXT NOT NULL, -- prefix or IP
first_seen TEXT NOT NULL, -- ISO time
last_seen TEXT NOT NULL,
PRIMARY KEY (kind, source, value)
);
The schema version is stored in PRAGMA user_version. The database runs in WAL mode (files ripe.db-wal, ripe.db-shm next to it); API and collector (also two containers sharing one local volume) can work with it concurrently. Do not place it on a network file system.
Change journal
For /addresses/diff the database keeps the tables changes(id, ts, kind, value, action add|del) and meta (the journal "horizon"). SQL triggers on addresses write to changes, so every path (collection, TTL, purge, removal of a source) is covered: add - the value appeared and no other source had it, del - the last entry of the value is gone. ts is UTC. The import from JSON is not written to the journal. Old records are deleted by the collector (changes_retention_days); the schema version is 3 (an older database is upgraded automatically on the first start, data is kept). Version 3 adds the table source_status(kind, source, last_attempt, last_success, failures, error_kind).
Inspect the data:
sqlite3 ripe.db "SELECT kind, source, count(*), max(last_seen) FROM addresses GROUP BY 1, 2"
Migration from the JSON storage
Automatic: on the first start of any process (API, collector or CLI) data.json and fqdn_data.json are imported into ripe.db in one transaction (concurrent starts are safe). Existing first_seen/last_seen are preserved; data in the oldest format without them gets last_seen = migration time (TTL starts counting from then). The originals are not deleted but renamed to data.json.migrated-<timestamp> and fqdn_data.json.migrated-<timestamp>.
Rollback to the JSON version
Stop both services, rename the *.migrated-* files back to data.json / fqdn_data.json, remove ripe.db*, start the previous version of the code. Addresses collected after the migration are lost in that case.
Backup and automatic restore
Backup job. The collector daemon runs the job backup (default 30 4 * * *, change it with POST /schedule and "type": "backup" or in schedule.backup). It makes an online copy of ripe.db (safe while the services run), checks the live database and the copy with PRAGMA quick_check, and keeps the last backup_keep copies (default 7; older ones are deleted only after a new copy succeeded). An empty database is not copied while sources are configured (a warning is logged, the job stays successful, older copies are kept): an empty copy would push good ones out of the rotation and restoring it would give an empty list. If the live database fails the check or the copy is bad, no copy is written, the older copies stay, and GET /health shows degraded (jobs.backup.last_error).
Copies are named ripe-<UTC time>.db and stored in RIPE_BACKUP_DIR (default <RIPE_DATA_DIR>/backups, i.e. /data/backups in Docker). By default they are on the same volume as the database: this protects against a corrupted file, not against losing the volume. For that, mount a separate volume/host directory and set RIPE_BACKUP_DIR to it, or copy the directory elsewhere regularly (e.g. docker compose cp collector:/data/backups ./backups).
Automatic restore. If ripe.db cannot be opened as a database (any process: API, daemon, CLI), the file is moved to ripe.db.corrupt-<timestamp> and the newest copy that passes the integrity check is put in its place; concurrent processes are serialized with a lock. The event is logged (ERROR), written to last_restore.json and shown in GET /health as last_restore. Data collected after that copy is lost; the collector gathers it again on the next runs. If the newest valid copy has no data although sources are configured, it is restored and the same marker is written. If there is no valid copy, a new empty database is created and the marker file db_recreated.json is written. While the data is not gathered again, GET /addresses and GET /addresses/diff answer 503 with Retry-After: 300 instead of an empty list (a consumer could take it for the truth and wipe its rules); a type without configured sources (e.g. no FQDNs) is not blocked. The collector removes the marker after a run when every configured type has data again; to accept an empty result earlier, delete db_recreated.json by hand. GET /health shows degraded and db_recreated (at, pending) meanwhile. A fresh installation (no database yet) and a manually deleted database file are not treated as a loss and return an empty list until the first collection.
- The change journal goes back with the copy, so after a restore all cursors and times issued earlier answer
410on/addresses/diff(the client makes a full download). The journal counter is shifted by 1,000,000 for that (a heuristic: it assumes fewer changes than that between two copies). - Only corruption detected when the database is opened is restored automatically. Damage inside the file shows up as
503on reads and is caught by thebackupjob (quick_check); restore it by hand: stop both services, keep the damagedripe.db*, copy the chosenbackups/ripe-*.dbtoripe.db, start the services.
Manual consistent copy: sqlite3 ripe.db ".backup ripe.db.bak" (do not copy ripe.db alone while the services are running - part of the data may still be in the -wal file).