# RIPE AS CIDR & FQDN IP Collector This project collects CIDR prefixes for specified Autonomous Systems (AS) from the RIPE NCC API and resolves IP addresses for specified FQDNs. It accumulates these addresses over time, maintaining a history of discovered prefixes. It also provides a FastAPI-based HTTP interface to retrieve the collected data. The system consists of **two independent processes**: the collector daemon (`collector_daemon.py`, runs the schedule) and the API server (`api_server.py`, serves data and edits the config). They communicate only through files: the collected addresses live in the SQLite database `ripe.db`, the rest in JSON (`config.json`, `status.json`). See section 9. > **Quick start:** the whole system (API + collector) runs with `docker compose up -d`, see section 8. Sections 1-4 describe the manual installation (venv + systemd/OpenRC), which remains supported. > **Upgrading from a version where the scheduler lived in the API:** running only `ripe-api` is no longer enough - without the `ripe-collector` service nothing is collected. `GET /health` reports `collector_alive: false` / `degraded` in that case. ## 1. Preparation and Installation ### Prerequisites - Python 3.8+ - `pip` and `venv` ### Installation Steps 1. **Clone the repository** (or copy the files) to your desired location, e.g., `/opt/ripe_collector`. ```bash mkdir -p /opt/ripe_collector cd /opt/ripe_collector # Copy files: cidr_collector.py, api_server.py, storage.py, requirements.txt, config.json ``` 2. **Create a Virtual Environment**: ```bash python3 -m venv venv ``` 3. **Install Dependencies**: ```bash source venv/bin/activate pip install -r requirements.txt deactivate ``` 4. **Initial Configuration**: Edit `config.json` to set your initial ASNs and FQDNs. ```json { "asns": [62041], "fqdns": ["google.com"], "ttl_days": 90 } ``` `ttl_days` - how long an address is kept after it was last seen (default `90`, `0` = keep forever). Optional `changes_retention_days` - how long the change journal for `/addresses/diff` is kept (default `30`, `0` = forever). 5. **Tests** (optional) run in a container, not in the local `venv`: ```bash docker build -f Dockerfile.test -t ripe-collector-test . docker run --rm ripe-collector-test ``` The image (`Dockerfile.test`, `python:3.11-slim`) contains only the code and dev dependencies - production data (`ripe.db`, `data.json`, `fqdn_data.json`) is excluded via `.dockerignore`. 6. **Knowledge graph** (optional): `/graphify .` (Claude Code skill) builds `graphify-out/` - `graph.html` (interactive), `GRAPH_REPORT.md`, `graph.json` - from code and `docs/`; refresh with `/graphify . --update` after changes. The directory is excluded from git and from Docker images. 7. **Repository:** `main` branch, remote `origin` = `https://artstore.rxmsk.ru/ayurishchev/ripe-cidr-collector.git`. Committed: code, tests, Docker files, `config.json` (initial sources), `.env.example`, `docs/`. Not committed (`.gitignore`): `venv/`, `.env`, databases `*.db*`, collected data `data.json` / `fqdn_data.json`, `graphify-out/`, service files. Every change follows the plan -> implementation -> summary flow in `docs/` and gets its own commit. --- ## 2. Running the Collector The collector normally runs as the **daemon** `collector_daemon.py` (see sections 3-4): it keeps the schedule from `config.json` (`schedule.asn`, `schedule.fqdn`) and picks up changes made via `POST /schedule` within 30 seconds, without a restart. Only one daemon instance can run at a time (a second one exits with an error). ```bash /opt/ripe_collector/venv/bin/python3 /opt/ripe_collector/collector_daemon.py ``` The script `cidr_collector.py` is used for manual runs and for managing the lists (see section 6). ### Manual Run ```bash /opt/ripe_collector/venv/bin/python3 /opt/ripe_collector/cidr_collector.py run ``` ### Alternative: cron instead of the daemon If you prefer cron, run the collector once per day (do not combine with the daemon - it is redundant). To run daily at 02:00 AM: > Run the job as the same user as the API (`crontab -u ripe -e`), otherwise data files may end up with the wrong owner. 1. Open crontab: ```bash crontab -e ``` 2. Add the line: ```cron 0 2 * * * /opt/ripe_collector/venv/bin/python3 /opt/ripe_collector/cidr_collector.py run >> /var/log/ripe_collector.log 2>&1 ``` --- ## 3. Application Setup: Systemd (Ubuntu, Debian) This section describes how to run two system services: the **API Server** (`api_server.py`) and the **collector daemon** (`collector_daemon.py`). ### Create Service File Create a dedicated user and the token file (the token protects `POST /schedule`): ```bash useradd --system --home /opt/ripe_collector --shell /usr/sbin/nologin ripe chown -R ripe: /opt/ripe_collector echo "RIPE_API_TOKEN=$(openssl rand -hex 32)" > /etc/ripe-api.env chmod 600 /etc/ripe-api.env ``` Create `/etc/systemd/system/ripe-api.service`: ```ini [Unit] Description=RIPE CIDR Collector API After=network.target [Service] User=ripe EnvironmentFile=/etc/ripe-api.env WorkingDirectory=/opt/ripe_collector NoNewPrivileges=true ProtectSystem=strict ReadWritePaths=/opt/ripe_collector ExecStart=/opt/ripe_collector/venv/bin/uvicorn api_server:app --host 0.0.0.0 --port 8000 Restart=always [Install] WantedBy=multi-user.target ``` Create `/etc/systemd/system/ripe-collector.service` (the daemon does not need the token): ```ini [Unit] Description=RIPE CIDR Collector daemon After=network.target [Service] User=ripe WorkingDirectory=/opt/ripe_collector NoNewPrivileges=true ProtectSystem=strict ReadWritePaths=/opt/ripe_collector ExecStart=/opt/ripe_collector/venv/bin/python collector_daemon.py Restart=always [Install] WantedBy=multi-user.target ``` ### Enable and Start ```bash # Reload systemd sudo systemctl daemon-reload # Enable services to start on boot and start them immediately sudo systemctl enable --now ripe-collector ripe-api # Check status sudo systemctl status ripe-collector ripe-api ``` --- ## 4. Application Setup: RC-Script (Alpine Linux) For Alpine Linux using OpenRC. ### Create Init Script Create `/etc/init.d/ripe-api`: ```sh #!/sbin/openrc-run name="ripe-api" description="RIPE CIDR Collector API" command="/opt/ripe_collector/venv/bin/uvicorn" # --host and --port and module:app passed as arguments command_args="api_server:app --host 0.0.0.0 --port 8000" command_background="yes" pidfile="/run/${RC_SVCNAME}.pid" directory="/opt/ripe_collector" depend() { need net } ``` Add the token and a non-root user (`/etc/conf.d/ripe-api` is read by OpenRC): ```sh adduser -S -h /opt/ripe_collector ripe chown -R ripe /opt/ripe_collector echo 'export RIPE_API_TOKEN=""' > /etc/conf.d/ripe-api chmod 600 /etc/conf.d/ripe-api ``` and add `command_user="ripe"` to the init script. Create the collector daemon script `/etc/init.d/ripe-collector`: ```sh #!/sbin/openrc-run name="ripe-collector" description="RIPE CIDR Collector daemon" command="/opt/ripe_collector/venv/bin/python" command_args="collector_daemon.py" command_background="yes" command_user="ripe" pidfile="/run/${RC_SVCNAME}.pid" directory="/opt/ripe_collector" depend() { need net } ``` ### Make Executable ```bash chmod +x /etc/init.d/ripe-api /etc/init.d/ripe-collector ``` ### Enable and Start ```bash # Add to default runlevel rc-update add ripe-collector default rc-update add ripe-api default # Start services service ripe-collector start service ripe-api start # Check status service ripe-collector status service ripe-api status ``` --- ## 5. API Usage Documentation The API runs by default on port `8000`. It allows retrieving the collected data in a flat JSON list. **Security model:** read endpoints (`/addresses`, `/schedule` GET, `/health`) are open, since consumers (routers, firewalls) usually can only do a plain GET - restrict access to port 8000 with a firewall. `POST /schedule` requires the header `X-API-Key: `; if the `RIPE_API_TOKEN` environment variable is not set, `POST` is disabled (503). If the data files are unreadable, the API answers `503`. ### Base URL `http://:8000` ### Endpoint: Get Addresses **GET** `/addresses` Retrieves the list of collected IP addresses/CIDRs. | Parameter | Type | Required | Default | Description | | :--- | :--- | :--- | :--- | :--- | | `type` | string | No | `all` | Filter by source type. Options: `cidr` (ASNs only), `fqdn` (Domains only), `all` (Both). | | `format` | string | No | `json` | Output format: `json`, `nftables`, `mikrotik`, `bird`, `frr` (see below). | | `ip_version` | string | No | `all` | `4`, `6` or `all`. | | `aggregate` | bool | No | `false` | Collapse overlapping/adjacent prefixes (`/23` + `/24` -> `/23`); a host IP inside a wider prefix is dropped. Applied after `cidr` and `fqdn` sources are merged. | | `name` | string | No | `ripe` | Name of the list/set in generated configs (`[A-Za-z][A-Za-z0-9_]{0,31}`). | With the defaults the response is the same flat JSON list as before. Other formats return `text/plain` scripts that are safe to re-apply (they replace the previous list). #### Example 1: Get All Addresses (Default) **Request:** ```bash curl -X GET "http://localhost:8000/addresses" ``` **Response (JSON):** ```json [ "142.250.1.1", "149.154.160.0/22", "149.154.160.0/23", "2001:4860:4860::8888", "91.108.4.0/22" ] ``` #### Example 2: Get Only CIDRs (from ASNs) **Request:** ```bash curl -X GET "http://localhost:8000/addresses?type=cidr" ``` **Response (JSON):** ```json [ "149.154.160.0/22", "149.154.160.0/23", "91.108.4.0/22" ] ``` #### Example 3: Get Only Resolved IPs (from FQDNs) **Request:** ```bash curl -X GET "http://localhost:8000/addresses?type=fqdn" ``` **Response (JSON):** ```json [ "142.250.1.1", "2001:4860:4860::8888" ] ``` #### Ready-to-use configuration formats | `format` | Result | How to apply | | :--- | :--- | :--- | | `nftables` | table `inet ` with interval sets `_v4` / `_v6` (`auto-merge`); other objects of the table are untouched | `curl -s "$URL/addresses?format=nftables" \| nft -f -` | | `mikrotik` | `/ip firewall address-list` and `/ipv6 firewall address-list` named `` (old entries are removed first, so there is a short window with an empty list) | `/tool fetch url="$URL/addresses?format=mikrotik" dst-path=ripe.rsc` then `/import ripe.rsc` | | `bird` | BIRD 2 prefix sets `define _V4 = [...]`, `_V6` for use in filters (`net ~ RIPE_V4`) | save to a file, `include` it, `birdc configure` | | `frr` | `ip prefix-list _v4` / `ipv6 prefix-list _v6` with `permit` entries | `curl -s "$URL/addresses?format=frr" > ripe.conf && vtysh -f ripe.conf` | Examples: ```bash # IPv4 only, aggregated, as an nftables set named "tg" curl -s "http://localhost:8000/addresses?format=nftables&ip_version=4&aggregate=true&name=tg" ``` A version without addresses is omitted from the script. The `bird` and `nftables` outputs were syntax-checked with `bird -p` and `nft -c`. ### Endpoint: Changes since the last sync (diff) **GET** `/addresses/diff?since=[&type=cidr|fqdn|all][&ip_version=all|4|6]` - only what was added and removed, instead of the full list. Read access is open, like `/addresses`. ```bash curl "http://localhost:8000/addresses/diff?since=0" # {"since":"0","now":"2026-09-21T04:09:44Z","cursor":42,"added":["3.0.0.0/24"],"removed":["2.0.0.0/24"]} curl "http://localhost:8000/addresses/diff?since=42" # next sync: use the returned cursor ``` - `since` is a **cursor** from the previous answer (recommended: exact, independent of clocks) or an ISO 8601 time (no time zone = UTC; write `+` in a URL as `%2B`). Time is compared with millisecond precision, borders are inclusive, so an entry may be reported twice - repeating it is harmless. - The result is the net effect: an address added and removed (or the other way) within the interval is not reported; an address that another source still holds is not reported as removed. Both CIDRs and IPs of FQDNs count. Output is JSON only, no aggregation. - The first sync: take the full `/addresses` list and the cursor from its response header `X-Changes-Cursor` (read before the data), then call `/addresses/diff?since=` regularly. - `400` - invalid `since`; `410 Gone` - `since` is older than the journal (see `changes_retention_days`) or the cursor does not belong to this database: fetch the full `/addresses` list and continue from the new `cursor`. ### Endpoint: Manage Schedule **GET** `/schedule` Returns the current cron schedules. **POST** `/schedule` (requires `X-API-Key`) Updates the schedule for a specific collector type. ```bash curl -X POST http://localhost:8000/schedule -H "X-API-Key: $RIPE_API_TOKEN" \ -H "Content-Type: application/json" -d '{"type": "asn", "cron": "*/15 * * * *"}' ``` Body: ```json { "type": "asn", "cron": "*/15 * * * *" } ``` *Note: `type` can be `asn` or `fqdn`. The change is written to `config.json`; the collector daemon applies it within 30 seconds (`applied_within_seconds` in the response). An invalid cron expression is rejected by the API (400) and ignored by the daemon.* ### Endpoints: Manage ASNs and FQDNs The lists of monitored sources are stored in `config.json` and can be managed through the API. Reading is open; changing requires the `X-API-Key` header (same token as `POST /schedule`). | Method and path | Description | | :--- | :--- | | `GET /asns`, `GET /fqdns` | Current lists: `{"asns": [...]}`, `{"fqdns": [...]}` | | `POST /asns` `{"asn": 62041}` | Add an ASN (`1..4294967295`). `201` if added, `200` if it was already there. | | `POST /fqdns` `{"fqdn": "example.com"}` | Add a domain (lower-cased, trailing dot removed; IP addresses and invalid labels are rejected with `422`). | | `DELETE /asns/{asn}?purge=false`, `DELETE /fqdns/{fqdn}?purge=false` | Remove a source (`404` if unknown). | ```bash curl -X POST http://localhost:8000/asns -H "X-API-Key: $RIPE_API_TOKEN" \ -H "Content-Type: application/json" -d '{"asn": 62041}' curl -X DELETE "http://localhost:8000/fqdns/example.com?purge=true" -H "X-API-Key: $RIPE_API_TOKEN" ``` - A new source is collected on the next scheduled run of the daemon; to collect it right away call `POST /collect` (see below). - **What happens to the data of a removed source:** by default it is kept and its addresses expire by `ttl_days` (the collector keeps ageing entries of sources that are no longer configured and deletes the entry when it is empty). With `purge=true` the collected addresses are deleted immediately and disappear from `/addresses`. With `ttl_days = 0` nothing expires - use `purge=true`. - The CLI commands (`add`, `remove`, `add-fqdn`, `remove-fqdn`) use the same locked, atomic config update, so they are safe to use alongside the API. ### Endpoint: Collect now **POST** `/collect` (requires `X-API-Key`) - starts a collection immediately instead of waiting for the schedule. ```bash curl -X POST http://localhost:8000/collect -H "X-API-Key: $RIPE_API_TOKEN" \ -H "Content-Type: application/json" -d '{"type": "asn"}' ``` Body is optional: `type` is `asn`, `fqdn` or `all` (default). The API does not collect itself: it passes the request to the collector daemon through the file `collect_request.json`, which the daemon checks every 5 seconds, so the collection starts within ~5 s. Answer `202` lists the requested types; follow the progress in `GET /health` (`running`, `last_run`, `last_finished`, `last_error` of each job). - Repeated requests are merged into one; if a collection of that type is already running (scheduled or manual) the new start is skipped. - If the collector daemon is not running (no fresh heartbeat), the answer is `503` and nothing is queued. ### Endpoint: Health **GET** `/health` Reports the state of the collector daemon (read from `status.json`): `collector_alive`, the cron / `running` / last run / `last_finished` / last error / next run of each job, and the number of stored addresses. The daemon writes a heartbeat every 30 seconds; `collector_alive` is `false` if it is older than 120 seconds or the daemon never ran. `status` is `ok` only if the daemon is alive and no job failed in its last run; otherwise `degraded` (HTTP code is still 200). --- ## 6. Advanced CLI Usage The collector script supports running modes independently: ```bash # Run both (Default) python3 cidr_collector.py run # Run only ASN collection python3 cidr_collector.py run --mode asn # Run only FQDN collection python3 cidr_collector.py run --mode fqdn ``` --- ## 7. Internal Logic & Architecture ### Collector Logic When the collector runs (whether manually or via schedule): 1. **Instantiation**: Creates a new instance of `CIDRCollector` or `FQDNCollector`. This forces a fresh read of `config.json`, ensuring any added ASNs/FQDNs are immediately processed. 2. **Fetching**: * **ASN**: Queries RIPE NCC API (`stat.ripe.net`). * **FQDN**: Uses Python's `socket.getaddrinfo` to resolve A and AAAA records. 3. **Merge in one transaction**: the fetched addresses are merged into the SQLite table `addresses` (`db.py`) in a single write transaction, so readers (the API) never see a half-updated state. 4. **Accumulation with TTL**: Each address has `first_seen`/`last_seen`. * New addresses are inserted; already seen ones get `last_seen` refreshed. * Addresses not seen for longer than `ttl_days` are removed - but only after a *successful* fetch. A RIPE/DNS failure never deletes anything. * Sources that are no longer configured are not fetched any more; their addresses simply expire by TTL. 5. **Persistence**: SQLite in WAL mode - the API reads while the collector writes. A corrupted database file is renamed to `ripe.db.corrupt-` and the API answers `503` instead of serving wrong data. ### Scheduler Logic `collector_daemon.py` uses `APScheduler` (`BlockingScheduler`) in its own process; the API has no scheduler. 1. **Startup**: the daemon takes an exclusive lock (`collector.daemon.lock`, single instance), loads the `schedule` block from `config.json` and creates two independent jobs (`asn_job`, `fqdn_job`; defaults `0 2 * * *` and `0 3 * * *`). Overlapping runs of the same job are not allowed. 2. **Runtime updates (POST /schedule)**: the API validates the cron expression and writes `config.json` atomically. Every 30 seconds the daemon compares the `schedule` block with the active one and reschedules the changed job (`reschedule_job`). An invalid cron expression is logged and ignored, the previous schedule stays. 3. **Status**: at the start and end of every run and every 30 seconds the daemon atomically writes `status.json` (jobs state + `updated_at` heartbeat); `GET /health` reads it. **Manual runs**: every 5 seconds the daemon checks `collect_request.json` (written by `POST /collect`) and starts the requested collections as one-off jobs; one run per type at a time. 4. **Concurrency**: a running job completes normally when the schedule changes; the new schedule applies to the next calculated run time. On `SIGTERM` the daemon exits after the running collection finishes. Restarting the API does not affect collection. --- ## 8. Application Setup: Docker Compose One image (`Dockerfile`), two services: `api` (uvicorn) and `collector` (`collector_daemon.py`). They share the named volume `ripe_data` mounted at `/data` (`RIPE_DATA_DIR`), which holds `ripe.db` (with `-wal`/`-shm`), `config.json`, `status.json` and lock files. Containers run as non-root (uid 10001) with a read-only root filesystem, dropped capabilities and rotated logs. Both have healthchecks (`healthcheck.py`). ### Start ```bash cp .env.example .env # set RIPE_API_TOKEN (openssl rand -hex 32) and, if needed, TZ / API_PORT docker compose up -d docker compose ps # both services become "healthy" within about a minute docker compose logs -f collector ``` Then add sources through the API (section 5), e.g. `POST /asns` with the token. With an empty volume the lists are empty. ### Settings (`.env`) | Variable | Default | Description | | :--- | :--- | :--- | | `RIPE_API_TOKEN` | - | Token for changing requests (`X-API-Key`). Without it they are disabled (503). | | `TZ` | `UTC` | Time zone in which the cron schedules are evaluated (e.g. `Europe/Moscow`). | | `API_PORT` | `8000` | Host port of the API. | ### Migrating existing data into the volume Run once, from the directory with your current `config.json`, `data.json`, `fqdn_data.json`, before the first `up` (the JSON data files are imported into `ripe.db` automatically on the first start, see section 9): ```bash docker compose create docker volume ls | grep ripe_data # the volume is named _ripe_data, e.g. ripe_cidr_collector_ripe_data docker run --rm -v _ripe_data:/data -v "$PWD":/src:ro alpine \ sh -c 'cp /src/config.json /src/data.json /src/fqdn_data.json /data/ && chown 10001 /data/*.json' docker compose up -d ``` ### Operations ```bash docker compose build && docker compose up -d # update after code changes docker compose stop collector # graceful: a running collection finishes (up to 60s) docker compose down # stop; data stays in the volume (add -v to delete it) ``` ### Notes and risks - **Port 8000 is published without TLS**, so `X-API-Key` travels in clear text. Restrict access with a firewall or put a TLS reverse proxy in front (bind the port to `127.0.0.1` by changing `ports` in `docker-compose.yml`). - The image installs unpinned dependencies from `requirements.txt`; rebuilds may pick up newer versions. - Files written by the services in the volume (`config.json`, `status.json`) have mode `600`; both services run as the same user. - `docker compose` uses the image name `ripe-cidr-collector`; `Dockerfile.test` is used only for running the tests (section 1, step 5). --- ## 9. Data Storage (SQLite) Collected addresses are stored in `ripe.db` (in the project directory, or in `RIPE_DATA_DIR` - the `/data` volume in Docker). `config.json` (sources, schedule, `ttl_days`) and `status.json` (collector heartbeat) stay JSON. ```sql CREATE TABLE addresses ( kind TEXT NOT NULL, -- 'asn' or 'fqdn' source TEXT NOT NULL, -- '62041' or 'example.com' value TEXT NOT NULL, -- prefix or IP first_seen TEXT NOT NULL, -- ISO time last_seen TEXT NOT NULL, PRIMARY KEY (kind, source, value) ); ``` The schema version is stored in `PRAGMA user_version`. The database runs in WAL mode (files `ripe.db-wal`, `ripe.db-shm` next to it); API and collector (also two containers sharing one local volume) can work with it concurrently. Do not place it on a network file system. ### Change journal For `/addresses/diff` the database keeps the tables `changes(id, ts, kind, value, action add|del)` and `meta` (the journal "horizon"). SQL triggers on `addresses` write to `changes`, so every path (collection, TTL, `purge`, removal of a source) is covered: `add` - the value appeared and no other source had it, `del` - the last entry of the value is gone. `ts` is UTC. The import from JSON is not written to the journal. Old records are deleted by the collector (`changes_retention_days`); the schema version is 2 (a version 1 database is upgraded automatically on the first start, data is kept). Inspect the data: ```bash sqlite3 ripe.db "SELECT kind, source, count(*), max(last_seen) FROM addresses GROUP BY 1, 2" ``` ### Migration from the JSON storage Automatic: on the first start of any process (API, collector or CLI) `data.json` and `fqdn_data.json` are imported into `ripe.db` in one transaction (concurrent starts are safe). Existing `first_seen`/`last_seen` are preserved; data in the oldest format without them gets `last_seen` = migration time (TTL starts counting from then). The originals are **not deleted** but renamed to `data.json.migrated-` and `fqdn_data.json.migrated-`. ### Rollback to the JSON version Stop both services, rename the `*.migrated-*` files back to `data.json` / `fqdn_data.json`, remove `ripe.db*`, start the previous version of the code. Addresses collected after the migration are lost in that case. ### Backup Copy the database consistently with `sqlite3 ripe.db ".backup ripe.db.bak"` (do not copy `ripe.db` alone while the services are running - part of the data may still be in the `-wal` file).