Add per-source status tracking, /health sources block and /sources
Collectors record the result of every source (schema v3, table source_status); /health reports failing sources (failures in a row >= source_failure_threshold) and turns degraded; GET /sources shows the full state; runs end with a summary log line instead of "data saved". Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
This commit is contained in:
1 parent
ab89c87321
commit
cb935fef0d
8 files changed
+351
-13
No files matched your search
@@ -45,6 +45,7 @@ The system consists of **two independent processes**: the collector daemon (`col
|
||||
}
|
||||
```
|
||||
`ttl_days` - how long an address is kept after it was last seen (default `90`, `0` = keep forever).
|
||||
Optional `source_failure_threshold` (default `3`) - how many runs in a row a source may fail before `GET /health` reports `degraded` (see the sources status below).
|
||||
Optional `ripestat_sourceapp` (default `ripe-cidr-collector`) - the application name sent to RIPEstat as `sourceapp`; RIPEstat asks clients to identify themselves, you can add a contact (`my-collector admin@example.org`).
|
||||
Optional `allow_non_global_ips` (default `false`) - by default only public addresses resolved from FQDNs are stored; loopback, private, link-local and unspecified addresses (`127.0.0.1`, `10.x`, `0.0.0.0`, ...) are ignored and logged. Set `true` for internal names.
|
||||
Optional `backup_keep` - how many database backups to keep (default `7`); the backup cron is `schedule.backup` (default `30 4 * * *`), see section 9.
|
||||
@@ -399,6 +400,11 @@ Body is optional: `type` is `asn`, `fqdn` or `all` (default). The API does not c
|
||||
**GET** `/health`
|
||||
Reports the state of the collector daemon (read from `status.json`): `collector_alive`, the cron / `running` / last run / `last_finished` / last error / next run of each job, the number of stored addresses and `last_restore` (`null`, or `{at, backup}` - time and file name of the last automatic restore from a backup; directories and the error text stay in `last_restore.json`) and `db_recreated` (`null`, or `{at, pending}` after the database was recreated without a backup; `pending: true` makes `status` `degraded`), see section 9. The daemon writes a heartbeat every 30 seconds; `collector_alive` is `false` if it is older than 120 seconds or the daemon never ran. `status` is `ok` only if the daemon is alive and no job failed in its last run; otherwise `degraded` (HTTP code is still 200).
|
||||
|
||||
`sources` shows the state of the data sources: for `asn` and `fqdn` the number of configured sources (`total`) and the list of **failing** ones (`failing`: `source`, `failures`, `last_success`, `error_kind`). A source is failing when its last `source_failure_threshold` runs in a row failed (default 3, so a single RIPEstat hiccup does not turn the status yellow); any failing source makes `status` `degraded`. `error_kind` is a short category without details (`timeout`, `connection`, `http_429`, `http_4xx`, `http_5xx`, `invalid_response`, `dns`, `no_global_addresses`, `unknown`); the full error text stays in the collector log.
|
||||
|
||||
### Endpoint: Sources status
|
||||
**GET** `/sources` (open) - the state of every configured source: `kind`, `source`, `addresses` (how many values it gave), `last_attempt`, `last_success`, `failures` (runs failed in a row), `error_kind`, plus `failure_threshold`. A source that was never polled has empty state; a source that "succeeds" but gives `addresses: 0` is easy to spot here. The state is kept in the database (table `source_status`), survives restarts and is included in backups; it is removed with the source (`purge` or the next collection after removal).
|
||||
|
||||
---
|
||||
|
||||
## 6. Advanced CLI Usage
|
||||
@@ -424,7 +430,7 @@ python3 cidr_collector.py run --mode fqdn
|
||||
When the collector runs (whether manually or via schedule):
|
||||
1. **Instantiation**: Creates a new instance of `CIDRCollector` or `FQDNCollector`. This forces a fresh read of `config.json`, ensuring any added ASNs/FQDNs are immediately processed.
|
||||
2. **Fetching**:
|
||||
* **ASN**: Queries RIPE NCC API (`stat.ripe.net`) with the `sourceapp` parameter. Connection failures and codes 429/500/502/503/504 are retried up to 3 times with growing pauses (0, 2, 4 s; a `Retry-After` pause is capped at 30 s), 10 s per attempt - at worst about 50 s per ASN. If all attempts fail the ASN is skipped for this run and nothing is deleted.
|
||||
* **ASN**: Queries RIPE NCC API (`stat.ripe.net`) with the `sourceapp` parameter. Connection failures and codes 429/500/502/503/504 are retried up to 3 times with growing pauses (0, 2, 4 s; a `Retry-After` pause is capped at 30 s), 10 s per attempt - at worst about 50 s per ASN. If all attempts fail the ASN is skipped for this run and nothing is deleted; the failure is recorded for `GET /health` and `GET /sources`. Every run ends with one summary line in the log (`ASN collection finished: 6 sources, 5 ok, 1 failed, +12/-3 prefixes`).
|
||||
* **FQDN**: Uses Python's `socket.getaddrinfo` to resolve A and AAAA records. Non-public addresses are dropped (see `allow_non_global_ips`); if nothing is left the run is treated like a DNS failure (nothing is deleted).
|
||||
3. **Merge in one transaction**: the fetched addresses are merged into the SQLite table `addresses` (`db.py`) in a single write transaction, so readers (the API) never see a half-updated state.
|
||||
4. **Accumulation with TTL**: Each address has `first_seen`/`last_seen`.
|
||||
@@ -508,7 +514,7 @@ CREATE TABLE addresses (
|
||||
The schema version is stored in `PRAGMA user_version`. The database runs in WAL mode (files `ripe.db-wal`, `ripe.db-shm` next to it); API and collector (also two containers sharing one local volume) can work with it concurrently. Do not place it on a network file system.
|
||||
|
||||
### Change journal
|
||||
For `/addresses/diff` the database keeps the tables `changes(id, ts, kind, value, action add|del)` and `meta` (the journal "horizon"). SQL triggers on `addresses` write to `changes`, so every path (collection, TTL, `purge`, removal of a source) is covered: `add` - the value appeared and no other source had it, `del` - the last entry of the value is gone. `ts` is UTC. The import from JSON is not written to the journal. Old records are deleted by the collector (`changes_retention_days`); the schema version is 2 (a version 1 database is upgraded automatically on the first start, data is kept).
|
||||
For `/addresses/diff` the database keeps the tables `changes(id, ts, kind, value, action add|del)` and `meta` (the journal "horizon"). SQL triggers on `addresses` write to `changes`, so every path (collection, TTL, `purge`, removal of a source) is covered: `add` - the value appeared and no other source had it, `del` - the last entry of the value is gone. `ts` is UTC. The import from JSON is not written to the journal. Old records are deleted by the collector (`changes_retention_days`); the schema version is 3 (an older database is upgraded automatically on the first start, data is kept). Version 3 adds the table `source_status(kind, source, last_attempt, last_success, failures, error_kind)`.
|
||||
|
||||
Inspect the data:
|
||||
```bash
|
||||
|
||||
Reference in new issue
Block a user