Compare commits

...
8 Commits
Author SHA1 Message Date
ayurishchevandClaude Sonnet 5.5 1366ecbdea Add the admin guide for manual database cleanup (SQL)
docs/ADMIN_CLEANUP.md: what is in the control-api database and what must
not be touched, preparation (stop, backup, checks), ready SQL for a full
reset before a new run, the event log, the check registry, single
addresses and compaction, verification after the cleanup, restore from a
backup, and what to do through the API instead. Every SQL block was run
on a copy of the production backup. Linked from the README.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-10-02 16:45:59 +03:00
ayurishchevandClaude Sonnet 5.5 0532baff09 Keep one address per validator; fix heartbeat handling and queue clear
A mass check on 2026-10-02 stalled 7 of 20 validators and sent 42
addresses to fail without a single check. A validator busy with slow
checks went silent, was marked unreachable, and its next heartbeat put it
back to idle while it still held the address; it was handed a second one,
whose association never ran (the in-flight guard was keyed by validator),
and both waited for their leases to expire.

- Heartbeat/re-register return an unreachable validator to assigned when
  it still holds an address, else idle.
- A validator is released only from the address it currently holds
  (ReleaseFIP, RequeueOrFail, MarkFIPOccupied, FreeValidator); an
  unreachable validator stays unreachable until its next heartbeat, so a
  dead validator is no longer handed a new address every lease period.
- ClaimNextQueued refuses a validator that still has an address; a
  ReconcileValidators pass on every tick repairs rows that disagree with
  the queue.
- Association guard is keyed by address, not validator.
- The agent sends heartbeats from their own goroutine.
- Clear queue / delete: detach only floating IPs of unfinished rows (done,
  failed and occupied rows kept their fip_id and made a clear issue >1000
  sequential cloud calls: 256 s), at most 8 in parallel; the operation no
  longer dies with the client connection (10 minute limit).

Includes the incident analysis and the plan under analysis/ and
docs/changes/, and rebuilt bin/control-api and bin/validator-agent.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-10-02 14:42:12 +03:00
ayurishchevandClaude Sonnet 5.5 cf4a883363 ansible: fix container name and git user, refuse to start a second agent
The real container on the validators is named validator-agent (the image
is cloud-ip-validator-validator-agent); the playbook used the image name
as the container name, so it would have started a second agent next to
the old one with the same validator_id. The clone on the validators is
owned by root, so git must run as root (git_user), otherwise fetch fails
with "cannot open .git/FETCH_HEAD: Permission denied".

Preflight now stops when the host has another container of this agent
(by name or image) besides container_name.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-10-02 10:01:05 +03:00
ayurishchevandClaude Sonnet 5.5 49890ff5de Add Ansible playbook to deliver validator-agent to the validators
Run from the jump host: on each validator it updates the git clone in
/opt/cloud-ip-validator, builds the image there, stops and removes the
current container and starts a new one from the new image. Run
parameters live in an env file (deploy/ansible/env/validator-agent.env,
git-ignored, template committed).

The image is built before the running container is touched, so a failed
build leaves the old container running. Hosts are updated in waves
(1, 4, rest) and any failure stops the run. validator_id comes from the
inventory and is checked against the running container before it is
replaced. Only ansible.builtin modules are used, so the validators need
no extra packages.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-10-02 09:23:39 +03:00
ayurishchevandClaude Sonnet 5.5 abbee9a08a Add self-check via control-api (self_check.methods)
control-api is hosted outside the cloud and validators reach it directly,
so it sees the floating IP as the connection's source address. New open
route GET /api/v1/agents/{id}/observed-ip returns that address (taken only
from the TCP peer; forwarding headers are ignored so a validator cannot
forge it).

The agent gets self_check.methods, a priority-ordered list of ip_echo
(unchanged) and control_api; the default stays [ip_echo]. The self-check
passes when any method confirms the address; the next method is tried on
no answer and on a mismatch. Each method has its own timeout so a hung
first method cannot starve the fallback, and control_api uses a new TCP
connection per call (a connection opened before the floating IP was
attached would keep reporting the old address).

Also: docker agent template/env, example config, docs, plan in
docs/changes, e2e script switch E2E_SELF_CHECK_METHODS, rebuilt
bin/control-api and bin/validator-agent.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-10-02 03:24:20 +03:00
ayurishchevandClaude Sonnet 5.5 146259cabb Detach the floating IP when a self-check fails
A failed self-check returned the address to the queue and freed the
validator in the database, but left the floating IP attached to the
validator's port. Every later association on that port then failed with
409 ("fixed IP already has a floating IP"), so one failed self-check
poisoned a validator for good; on 2026-10-01 all 20 validators were
poisoned within 23 minutes after ifconfig.me timeouts.

SelfCheckResult now disassociates the floating IP before requeueing, and
ignores a late failed report for an address the validator no longer
holds (it could belong to another validator by then).

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-10-02 03:24:19 +03:00
ayurishchevandClaude Sonnet 5.5 1ee5757004 Free validator ports in the cloud on clear/cancel/delete
An address still being associated (assigning_fip) has no fip_id in the
database, so "clear queue" did not detach it, and the association then
finished after the row was gone, leaving the floating IP on the validator
port for good.

- After clear/cancel/delete, ask the cloud which floating IPs sit on the
  affected validator ports (new ListFloatingIPsByPort) and detach those
  that this system queued (known in ip_registry); foreign ones are left.
- SetFIPAssociated applies only to a row still in assigning_fip; if the
  address was removed meanwhile, associateFIP detaches the floating IP.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-10-01 21:00:31 +03:00
ayurishchevandClaude Sonnet 5.5 c0e300f71c Run per-validator OpenStack work in parallel so all validators start at once
The orchestrator claimed idle validators one by one and associated each
floating IP synchronously (~30 s per address), so 20 validators started
about 30 s apart. Aggregation/disassociation and lease reclaim were
sequential in the same way.

The slow OpenStack calls now run in one goroutine per address, guarded by
an in-flight set against duplicates. control-api runs Tick in Async mode
(Tick does not wait); tests keep the waiting mode.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-10-01 20:37:32 +03:00
54 changed files with 3516 additions and 112 deletions

No files matched your search

+2
View File
@@ -5,6 +5,8 @@
!/deploy/docker/.env.example
!/deploy/docker/.env.prod.example
/deploy/docker/control-api/control-api.docker.yaml
# Ansible: рабочий env-файл запуска validator-agent (шаблон .example остаётся в git).
/deploy/ansible/env/*.env
/rxprod-compose/.env
# Live runtime database for the rxprod-compose control-api container.
+7 -4
View File
@@ -48,7 +48,7 @@ docker compose up -d --build # весь стенд на одной
**Остальные компоненты**
| Компонент | Ключи |
|---|---|
| `validator-agent` | `validator_id`, `control_api_url`, `control_api_token_env` (`CONTROL_API_AGENT_TOKEN`), `poll_interval_seconds`, `self_check.*` (таймаут, `ip_echo_urls`), `checks.*` (таймауты HTTPS/ICMP, число ICMP-пакетов, `ssh.*`) |
| `validator-agent` | `validator_id`, `control_api_url`, `control_api_token_env` (`CONTROL_API_AGENT_TOKEN`), `poll_interval_seconds`, `self_check.*` (таймаут, `methods` — `ip_echo`/`control_api`, `ip_echo_urls`), `checks.*` (таймауты HTTPS/ICMP, число ICMP-пакетов, `ssh.*`) |
| `prober` | `site_id`, `control_api_url`, `control_api_token_env` (`CONTROL_API_AGENT_TOKEN`), `poll_interval_seconds`, `checks.*` (таймауты TCP/ICMP, число ICMP-пакетов) |
| `admin-dashboard` | `server.listen_addr` (`:8090`), `control_api.base_url`, `control_api.timeout_seconds`, `control_api.token_env` (`ADMIN_DASHBOARD_CONTROL_API_TOKEN`), `auth.username_env` / `password_env` / `session_secret_env` (`ADMIN_DASHBOARD_USERNAME` / `_PASSWORD` / `_SESSION_SECRET`), `auth.session_ttl_minutes` (480), `overview.last_completed_count` (20), `overview.poll_interval_seconds` (5) |
@@ -118,7 +118,7 @@ docs/ документация и планы доработок
## API (`/api/v1`)
| Область | Эндпоинты |
|---|---|
| Валидатор | `POST /agents/register`, `POST /agents/{id}/heartbeat`, `GET /agents/{id}/assignment`, `POST /agents/{id}/self-check\|events\|results\|complete` |
| Валидатор | `POST /agents/register`, `POST /agents/{id}/heartbeat`, `GET /agents/{id}/assignment`, `GET /agents/{id}/observed-ip`, `POST /agents/{id}/self-check\|events\|results\|complete` |
| Пробер | `POST /probers/register`, `POST /probers/{site_id}/heartbeat`, `GET /probers/{site_id}/assignments`, `POST /probers/{site_id}/results` |
| Очередь | `GET /admin/status`, `GET\|POST /admin/ips` (`limit/offset/state/q/result/order` — постранично), `GET /admin/ips/{ip}`, `POST /admin/ips/{ip}/cancel`, `DELETE /admin/ips/{ip}`, `POST /admin/ips/delete\|clear`, `POST\|GET /admin/ips/scan` (фоновый скан: `202`, `dry_run`, `wait`; статус и прогресс) |
| Реестр | `GET /admin/registry` (`limit/offset/q/last_result` — постранично), `GET /admin/registry/{ip}` |
@@ -138,7 +138,7 @@ docs/ документация и планы доработок
|---|---|---|
| admin | `CONTROL_API_ADMIN_TOKEN` | все `/api/v1/admin/*` |
| agent | `CONTROL_API_AGENT_TOKEN` | запись результатов валидатора и пробера: `self-check`, `events`, `results`, `complete` |
| открыто | — | `GET /healthz`, `register`, `heartbeat` и получение задания (`GET assignment` / `assignments`) |
| открыто | — | `GET /healthz`, `register`, `heartbeat`, получение задания (`GET assignment` / `assignments`) и `GET /agents/{id}/observed-ip` |
Токены разные: административный не открывает методы агентов, и наоборот. Валидатор и пробер получают настройку и задание без токена, но не могут отправить результат без токена агентов.
- **Пустой токен — уровень открыт** (обратная совместимость): `control-api` стартует с предупреждением в логе. На реальном стенде задайте оба токена и ограничьте доступ на уровне сети ([docs/SETUP.md](docs/SETUP.md#сетевые-доступы)); токены идут открытым текстом без TLS — публикуйте через reverse-proxy с TLS. Включать токен агентов нужно **после** его раздачи валидаторам и проберам ([порядок](docs/SETUP.md#5-аутентификация-токены-и-пароль-дашборда)).
- **Дашборд закрыт логином и паролем** (один администратор; пароль и ключ сессии — из env). Сессия — подписанная cookie (`HttpOnly`, `SameSite=Strict`, без состояния на сервере), CSRF-защита по `Origin`, 5 неудачных входов за 10 минут с одного IP → `429`. Без заданных логина/пароля дашборд открыт (с предупреждением в логе). Дашборд ходит в API с токеном администратора. Подробности — [docs/DASHBOARD.md](docs/DASHBOARD.md#вход-и-сессия).
@@ -149,7 +149,7 @@ docs/ документация и планы доработок
## Публикация и эксплуатация
- **Бинарники.** Готовые linux/amd64 лежат в `bin/` и **не обновляются автоматически**: после правок кода пересоберите их и обновите `SHA256SUMS` (команды — [docs/SETUP.md](docs/SETUP.md#вариант-b-сборка-из-исходников)); Dockerfile копируют именно `bin/*`.
- **systemd.** Юниты в `deploy/systemd/`; у `control-api` — `EnvironmentFile` с учётными данными OpenStack.
- **Docker.** `deploy/docker/docker-compose.yml` + `docker-compose.override.yml` (dev, mock) или `docker-compose.prod.yml` (без публикации портов, `restart: unless-stopped`). Какие сервисы поднимаются на хосте, задаёт `COMPOSE_PROFILES`: `control-plane`, `dashboard`, `prober`, `validator`. БД — volume `cloud-ip-validator-db`.
- **Docker.** `deploy/docker/docker-compose.yml` + `docker-compose.override.yml` (dev, mock) или `docker-compose.prod.yml` (без публикации портов, `restart: unless-stopped`). Какие сервисы поднимаются на хосте, задаёт `COMPOSE_PROFILES`: `control-plane`, `dashboard`, `prober`, `validator`. БД — volume `cloud-ip-validator-db`. Массовая доставка `validator-agent` на валидаторы (сборка образа на каждом хосте, замена контейнера) — Ansible-сценарий [`deploy/ansible/`](deploy/ansible/README.md).
- **Реальный стенд.** `rxprod-compose/` — compose с готовыми образами, собственным `control-api.yaml` и каталогом БД `capi-db/`; `.env` с учётными данными в репозиторий не входит.
- **Миграции** применяются при старте `control-api`; версия схемы — `PRAGMA user_version`. Начальная загрузка (`validators`, `sites`, `targets`, `check_types`, `inbound_checks`) выполняется только в пустые таблицы.
- Остановка (`SIGTERM`) корректно завершает HTTP-сервер и фоновые циклы. Состояние автоцикла и очереди сохраняется в БД.
@@ -182,6 +182,7 @@ scripts/run-local-e2e.sh # сквозной прог
| [docs/USAGE.md](docs/USAGE.md) | Повседневная работа: очередь, сканирование Floating IP, автоматический цикл, статус, реестр и история, разбор результатов |
| [docs/API.md](docs/API.md) | Спецификация HTTP API `control-api` и примеры запросов (curl) |
| [docs/DASHBOARD.md](docs/DASHBOARD.md) | Устройство `admin-dashboard`: страницы, поиск и фильтр, обработка ошибок |
| [docs/ADMIN_CLEANUP.md](docs/ADMIN_CLEANUP.md) | Ручная очистка БД администратором (SQL): сброс данных перед новым прогоном, журнал событий, реестр проверок, отдельные адреса, резервная копия и восстановление |
| [docs/DIAGRAMS.md](docs/DIAGRAMS.md) | Диаграммы потоков данных: control plane, egress-проверка, телеметрия |
| [docs/LOCAL_E2E.md](docs/LOCAL_E2E.md) | Полностью офлайн-прогон всей системы одним скриптом |
| [docs/changes/](docs/changes/) | Планы доработок и отчёты ревью с отметкой времени в имени файла (последняя: [скан при тысячах адресов](docs/changes/2026-10-01_18-59_fip-scan-at-scale-review.md)) |
@@ -200,6 +201,8 @@ scripts/run-local-e2e.sh # сквозной прог
| Дата | Веха | Документ |
|---|---|---|
| 2026-10-02 | Валидатор не получает второй адрес при потере heartbeat (иначе адреса уходили в `fail` без проверок); heartbeat агента в отдельном потоке; быстрая очистка очереди, не зависящая от соединения клиента | [план](docs/changes/2026-10-02_09-02_orchestrator-validator-state-and-clear-plan.md) · [анализ инцидента](analysis/2026-10-02_08-56_1026-addresses_mass-check-analysis.md) · [USAGE](docs/USAGE.md#управление-валидаторами) |
| 2026-10-02 | Самопроверка через ручку control-api (`self_check.methods`, способ `control_api` рядом с IP-echo); отвязка Floating IP при провале self-check | [план](docs/changes/2026-10-02_03-06_self-check-control-api-plan.md) · [API](docs/API.md#get-apiv1agentsidobserved-ip) |
| 2026-10-01 | Скан Floating IP при тысячах адресов: фоновый постраничный скан с прогрессом, фаза `scanning` в автоцикле, постраничные `/ips` и `/registry`, «Обзор» на счётчиках | [план](docs/changes/2026-10-01_18-19_fip-scan-at-scale-plan.md) · [ревью и тесты](docs/changes/2026-10-01_18-59_fip-scan-at-scale-review.md) · [USAGE](docs/USAGE.md#сканирование-floating-ip-из-openstack) · [API](docs/API.md#post-apiv1adminipsscan) |
| 2026-10-01 | Аутентификация: токены администратора и агентов для API, логин и пароль для дашборда | [план](docs/changes/2026-10-01_11-12_authentication-plan.md) · [ревью и тесты](docs/changes/2026-10-01_11-31_authentication-review.md) · [API](docs/API.md#аутентификация) |
| 2026-10-01 | Автоматический цикл проверок по сценарию: очистка → скан FIP → проверка → пауза | [USAGE](docs/USAGE.md#автоматический-цикл-проверок) · [API](docs/API.md#автоматический-цикл-проверок) |
@@ -0,0 +1,167 @@
# Аналитический разбор массовой проверки: 1026 адресов
> Время отчёта: 2026-10-02 08:56 UTC · Данные: снимок БД control-api на 08:46:53 UTC и лог control-api за 4 часа до остановки
> Проверено адресов: **1026** (984 завершены `done` + 42 завершены `failed`) из 6440 в очереди
> Окно прогона: 07:14:58 – 08:46:53 UTC (1 ч 32 мин). Остановлен вручную в 08:53:46 UTC.
## 1. Остановка проверок
- Проверки остановлены в 08:53:46 UTC операцией «Очистить всё» (`POST /api/v1/admin/ips/clear`).
- После остановки: очередь пуста, все 20 валидаторов `idle`, последняя выдача адреса в 08:53:43, новых выдач нет.
Реестр с историей адресов сохранён (6445 записей).
- Отвязка Floating IP с портов валидаторов при очистке: 13 зависших привязок снято прямым опросом портов
(`detached floating ip from validator port`), ещё 8 привязок, завершавшихся во время очистки, система сняла сама
(`address removed during association`). Состояние портов в самом OpenStack на момент отчёта не проверялось.
- **Первая попытка очистки не сработала.** Очистка шла дольше таймаута клиента (120 с) и оборвалась: 600 отвязок завершились
с ошибкой `context canceled`, ничего не удалилось, проверки продолжались. Повторная очистка без таймаута выполнялась 256 с.
- **Причина долгой очистки (дефект кода, не исправлен):** у завершённых адресов (`done`) в БД остаётся `fip_id`, и очистка
последовательно отвязывает все такие FIP, хотя они уже свободны (более 1000 вызовов OpenStack по ~0,2 с).
## 2. Итоги прогона
| Показатель | Значение |
|---|---|
| Адресов в очереди | 6440 |
| Завершено | 1026 (984 `done` + 42 `failed`) |
| `pass` | 227 (23% от завершённых) |
| `partial` | 757 (77%) |
| `fail` | 42 (все `failed`, вердикта по существу нет, см. раздел 4) |
| Остались в очереди / в работе на момент снимка | 5392 в очереди, 22 в работе |
Пропускная способность по 10-минутным окнам (завершено адресов за минуту): 07:10 — 5,8; 07:20 — 13,4; 07:30 — 11,6;
07:40 — 10,1; 07:50 — 10,6; 08:00 — 9,4; 08:10 — 9,7; 08:20 — 10,5; 08:30 — 10,6; 08:40 — 6,7 (окно неполное).
Цикл одного адреса (от выдачи валидатору до итога): минимум 65 с, медиана 85 с, p90 100 с, p99 115 с, максимум 125 с,
среднее 84 с (по 984 адресам `done`). Для 20 валидаторов это теоретически ~14 адресов в минуту; фактически ~10,4,
потеря около 27% из-за залипших валидаторов (раздел 4).
## 3. Фактура по накопившимся ошибкам
### 3.1. Классы ошибок (с 07:10)
В логе control-api за 4 часа на уровне `ERROR` только один вид сообщений: 41 ошибка привязки Floating IP.
| Ошибка / событие | Число | Что это |
|---|---|---|
| `lease expired` (возврат адреса по истечении лизинга) | 198 | Валидатор не подхватил задание за время лизинга. Затронуто минимум 77 адресов (часть событий без привязки к адресу) |
| Привязка FIP: `409 Cannot associate floating IP … fixed IP already has a floating IP` | 41 | На порту валидатора уже висит другой FIP. Порты: v1 — 13, v7 — 6, v13 — 6, v3 — 5, v12 — 5, v16 — 4, v17 — 2 |
| `validator_unreachable` | 7 | По одному разу: v1 и v12 (07:18:03), v13 (07:33:53), v3 (07:34:33), v7 (07:40:03), v16 (07:42:33), v17 (08:17:53) |
| `site_unreachable` | 5 | rxmsk (08:09, 08:45) и misha-v (08:09, 08:20, 08:45) |
| «Очистить всё»: `context canceled` | 600 | Последствие обрыва первой очистки (08:49), см. раздел 1 |
Чего не было: **провалов self-check — 0 из 1015** результатов; все 1015 прошли способом `control_api` (запасной `ip_echo`
не понадобился). Событий `fip_occupied` — 0, адресов `occupied` — 0.
Журнал событий за прогон: `self_check_result` 2030 (по две записи на адрес), `fip_associated` 1120, `config_received` 1015,
`aggregated` 1004, `retry_or_fail` 239 (198 лизинг + 41 привязка), `lease_expired` 198, `validator_unreachable` 7,
`site_unreachable` 5.
### 3.2. По валидаторам
| Валидатор | Завершено (`done`) | `failed` | Сбросов лизинга | `unreachable` |
|---|---|---|---|---|
| vkiplab-v1 | 2 | 17 | 49 | 1 |
| vkiplab-v12 | 2 | 13 | 51 | 1 |
| vkiplab-v13 | 13 | 10 | 41 | 1 |
| vkiplab-v16 | 38 | 2 | 19 | 1 |
| vkiplab-v7 | 36 | 0 | 21 | 1 |
| vkiplab-v17 | 50 | 0 | 10 | 1 |
| vkiplab-v3 | 50 | 0 | 7 | 1 |
| остальные 13 (v2, v4–v6, v8–v11, v14, v15, v18–v20) | 60–63 | 0 | 0 | 0 |
(Сбросы лизинга и `unreachable` в таблице — только за прогон, с 07:10. У v14, v18 и v2 в истории БД есть сбросы лизинга
за 1 октября, к этому прогону они не относятся.)
- **v1 и v12** не подхватили ни одного задания после 07:17 (последний подхват 07:17:29 и 07:17:26).
- **v13** — после 07:33:16.
- **v7, v16, v17, v3** залипали временно и затем восстановились; механизм восстановления не выяснен.
### 3.3. 42 адреса `fail`
- По валидаторам: v1 — 17, v12 — 13, v13 — 10, v16 — 2. Все 42 исчерпали повторы: `retry_count` = 4 и `attempt_number` = 4 у каждого.
- Причины неудачных попыток по этим адресам: 143 сброса лизинга и 25 ошибок привязки `409`.
- **Ни одной проверки по ним не выполнено** (в таблице проверок у этих адресов 0 записей): адреса не получили вердикта,
и `fail` здесь не характеризует сами адреса.
- Появлялись равномерно с 07:31 до 08:46 (4–10 за 10 минут).
- Подсети: 37.139.x, 79.137.x, 83.166.x и другие, без концентрации.
## 4. Причины
Подтверждены по БД и логам control-api. Логи агентов на самих валидаторах не изучались.
**Причина 1. Валидатор получает два адреса сразу и застревает.** Два дефекта вместе:
- Heartbeat (`queries_validators.go`) возвращает валидатор из `unreachable` в `idle`, не проверяя, что за ним числится адрес.
Адрес с долгими внешними проверками (3–4 таймаута по 10 с) блокирует агента больше 30 с (порог
`heartbeat_timeout_seconds`), control-api помечает валидатор недоступным, затем возвращает в `idle` занятым.
- При завершении старого адреса `ReleaseFIP` (`queries_ipqueue.go`) освобождает валидатор по его имени, а не по адресу.
- Подтверждение: **35 двойных выдач** (два адреса одному валидатору с интервалом ~5 с) в логе: v1 — 10, v12 — 9, v7 — 6,
v13 — 4, v17 — 3, v16 — 2, v3 — 1. Из 1265 выдач в логе.
**Причина 2. Регрессия моей правки с параллельной привязкой.** Защита от дублей в `orchestrator.go:149` ключуется по
валидатору (`assign:<validator>`). Вторая выдача того же валидатора пропускает привязку, и адрес стоит в `assigning_fip`
до истечения лизинга. В БД у валидатора одно поле `current_ip_id`, оно указывает на последний выданный адрес, агент по нему
получает пустое задание, лизинг истекает, валидатор берёт новый адрес, и круг повторяется. Ошибки `409` — следствие: на
порту остаётся FIP первого адреса, привязать второй нельзя.
**Причина 3. Агент молчит во время долгих проверок.** Heartbeat отправляется только между заданиями; адреса с несколькими
таймаутами ведут к `unreachable` (пусковой механизм причины 1). Пример: v1 проверял `37.139.32.1`, v12 — `37.139.32.4`;
у обоих проваливались все четыре внешних HTTPS-цели (~40 с таймаутов), в 07:18:03 оба помечены `unreachable`.
**Дефект очистки.** См. раздел 1: отвязка всех `done`-адресов последовательно.
## 5. Результаты проверок самих адресов
**Почему 77% `partial`:** 735 адресов проваливают только исходящие HTTPS, ещё 22 — исходящие и входящие.
| Цель (egress HTTPS) | Адресов с провалом (из 984) |
|---|---|
| `packages.ubuntu.com` | 698 (71%) |
| `repo.almalinux.org/almalinux/` | 304 (31%) |
| `github.com` | 251 (26%) |
| `hub.docker.com` | 17 (2%) |
Число провалов на один `partial`-адрес: 1 — 392 адреса, 2 — 202, 3 — 136, 4 и больше — 27.
По подсетям (доля адресов с провалом цели):
| Подсеть | Адресов | `packages.ubuntu.com` | `repo.almalinux.org` | `github.com` | `hub.docker.com` | Доля `partial` |
|---|---|---|---|---|---|---|
| 83.166.x | 435 | 80% | 53% | 42% | 0% | 92% |
| 37.139.x | 418 | 62% | 11% | 10% | 4% | 63% |
| 79.137.x | 109 | 66% | 19% | 18% | 0% | 68% |
| 5.188.x | 22 | 59% | 0% | 0% | 0% | 59% |
- `packages.ubuntu.com` проваливается у всех подсетей и во все окна. Доля проваленных строк egress для этой цели по 10-минутным
окнам выросла с 26% до 40% (строки включают HTTPS и ICMP, поэтому реальная доля HTTPS вдвое выше).
- Доля `partial` почти одинакова у всех валидаторов (72–84%; v1 — 100% по 2 адресам, v12 — 50% по 2): причина в самом адресе
или во внешнем ресурсе, а не в валидаторе.
- Провалы `repo.almalinux.org` и `github.com` сильно зависят от подсети (83.166.x — 53% и 42%, 37.139.x — 11% и 10%).
- Успешные HTTPS-проверки: медиана 200 мс, p90 6707 мс (много ответов близко к таймауту 10 с).
**Входящие проверки.** Провалы только у `inbound-site-1` (33 пробы из 2934) и `inbound-site-3` (50 из 2931); `inbound-site-2` — 0 из 2952,
`inbound-site-4` — 1 из 2952. Провалы всплесками в 08:00, 08:20 и 08:40 по ICMP, SSH и TCP 22; у 20 адресов упали все пробы
одной площадки. Пробер rxmsk и misha-v в эти же минуты отмечены `unreachable`. Причина недоступности проберов не выяснена.
## 6. Рекомендации
1. **Исправить оркестратор** (до повторного запуска): защита от дублей по адресу, `unreachable` → `assigned` вместо `idle`,
освобождение валидатора только по текущему адресу, heartbeat агента в отдельном потоке.
2. **Исправить очистку:** отвязывать только действительно привязанные FIP (без `fip_released_at`) и параллельно.
3. **Перепроверить 42 адреса** после исправления: вердикта по существу они не получили. Все `fail` с причиной «lease expired» считать недействительными.
4. **Решить по `packages.ubuntu.com`:** 71% `partial` и 10 с таймаута на каждый такой адрес; если ресурс нестабилен, убрать или заменить.
5. **Проверить причины `site_unreachable`** у проберов (возможна перегрузка при большой очереди).
6. Проверить в OpenStack, что порты валидаторов свободны от Floating IP.
## 7. Что не проверено
- Состояние портов и Floating IP в самом OpenStack.
- Логи агентов `validator-agent` на валидаторах и логи проберов (причины блокировок heartbeat подтверждены по времени и
событиям control-api, не по логам агентов).
- Почему залипшие v7, v16, v17, v3 восстановились.
- Причины провалов внешних HTTPS-целей (ресурс, сеть облака или фильтрация) и недоступности проберов.
## 8. Источники данных
- Снимок БД control-api: `.backup` на 08:46:53 UTC (таблицы `ip_queue`, `checks`, `events`, `validators`, `sites`).
- Лог контейнера control-api за 4 часа до остановки (1504 строки) и за период очистки.
- Статус и результат очистки: HTTP 200 за 256,7 с.
@@ -0,0 +1,42 @@
37.139.32.70
37.139.32.71
37.139.32.72
37.139.32.151
37.139.33.219
37.139.33.227
37.139.33.230
37.139.34.75
37.139.34.94
37.139.34.128
37.139.40.122
37.139.41.103
37.139.41.129
37.139.41.157
37.139.41.202
37.139.42.126
37.139.43.9
79.137.174.150
79.137.174.172
79.137.174.205
79.137.174.210
79.137.175.14
79.137.175.51
79.137.175.71
79.137.175.162
83.166.232.116
83.166.232.117
83.166.232.159
83.166.233.37
83.166.233.75
83.166.233.78
83.166.234.136
83.166.235.75
83.166.235.132
83.166.235.247
83.166.237.59
83.166.237.228
83.166.238.19
83.166.248.57
83.166.248.60
83.166.248.92
83.166.248.188
+2 -2
View File
@@ -1,4 +1,4 @@
26ff297b66a0c974b7143179e44a58f476265482cd1929bc01a95ad29e50ca7a control-api
091ad94b5706b1778b181542de7a4b251cd16c421144269d0955769603d51e06 validator-agent
9447699d9eed4b9fb5d417769fc868986a10357ec633ec3742883b3002aaafc3 control-api
9fb6608b84143f7c4f318f3cc92dcd9f95c7831d627b23d67cce5a5908ced704 validator-agent
3e9e14dbb361ee76aaad7c1da6864b3ea111e0ed151403f904b12485631bbf75 prober
d72234688eb1954dbff420ca1c8e83b83ceab561b18a0336af0f73fe4d9ac8be admin-dashboard
BIN
View File
Binary file not shown.
Binary file not shown.
+1
View File
@@ -59,6 +59,7 @@ func run(configPath string, log *slog.Logger) error {
}
orch := orchestrator.New(database, osClient, cfg, log)
orch.Async = true // slow OpenStack calls run per address, Tick never waits for them
// Background jobs (the floating-IP scan) live as long as the process, not
// as long as the HTTP request or loop iteration that started them.
orch.SetContext(ctx)
+13 -1
View File
@@ -13,7 +13,19 @@ poll_interval_seconds: 5
self_check:
timeout_seconds: 10
# Must be a resource genuinely outside the cloud project — OpenStack only
# Способы самопроверки в порядке приоритета (допустимо: ip_echo,
# control_api); по умолчанию [ip_echo]. Самопроверка успешна, если адрес
# подтвердил любой способ: пробуются по порядку, остановка на первом
# успешном, к следующему переходим и при отсутствии ответа, и при
# несовпадении адреса. Таймаут timeout_seconds действует на каждый способ
# отдельно (зависший первый способ не лишает второй времени). control_api спрашивает у control-api, с какого адреса он видит
# это соединение; рекомендуется [control_api, ip_echo], когда control-api
# стоит вне облака и валидатор ходит к нему напрямую (через внешнюю сеть).
# Ограничение: если control-api достижим по внутренней сети облака, он
# увидит частный адрес валидатора и control_api всегда даст несовпадение —
# тогда оставьте только ip_echo (или он сработает вторым в списке).
methods: [control_api, ip_echo]
# Used by the ip_echo method. Must be a resource genuinely outside the cloud project — OpenStack only
# applies floating-IP SNAT to traffic leaving via the external network,
# so anything reachable over the project's internal network (including
# control-api itself, if it's on the same internal network) would report
+110
View File
@@ -0,0 +1,110 @@
# Доставка validator-agent на валидаторы (Ansible)
Сценарий запускается с jump-хоста и на каждой ВМ-валидаторе: обновляет git-клон в `/opt/cloud-ip-validator`,
**собирает образ на самом хосте**, останавливает и удаляет старый контейнер `validator-agent` (образ — `cloud-ip-validator-validator-agent`)
и поднимает на его месте новый. Параметры запуска агента вынесены в env-файл.
Порядок безопасен: сначала проверки, обновление кода и сборка образа, и только потом замена контейнера. Если что-то
упало до замены, старый контейнер продолжает работать. Простой валидатора — секунды (`stop` + `rm` + `run`).
## Требования
| Где | Что |
|---|---|
| jump-хост (Debian 13) | `apt install ansible-core`; SSH-доступ по ключу ко всем валидаторам |
| валидаторы | Docker, git, клон репозитория в `/opt/cloud-ip-validator`, `sudo` без пароля для SSH-пользователя, доступ к Gitea и Docker Hub (`alpine:3.20`), архитектура x86_64 |
Только модули `ansible.builtin`: Python Docker SDK и дополнительные коллекции на валидаторах не нужны.
Go на валидаторах не нужен: образ копирует закоммиченный `bin/validator-agent` (его сумма проверяется по `bin/SHA256SUMS`).
## Подготовка (один раз)
```bash
cd deploy/ansible
cp env/validator-agent.env.example env/validator-agent.env
chmod 600 env/validator-agent.env
$EDITOR env/validator-agent.env # адрес control-api, токен, способы самопроверки, таймауты
ansible validators -m ping # проверка связи (20 хостов: validator-1 ... validator-20)
```
- В `inventory/hosts.yml` 20 хостов `validator-1 … validator-20` с адресами (`ansible_host`) и переменной `validator_id`
(`vkiplab-v1 … vkiplab-v20`, как в control-api). Соответствие `validator-N` → `vkiplab-vN` задано по порядку номеров;
**сценарий сверяет `validator_id` с тем, что записано в работающем контейнере на хосте, и при расхождении останавливается до замены
контейнера**. `VALIDATOR_AGENT_VALIDATOR_ID` в env-файл писать не нужно: сценарий подставляет его из inventory.
- Отпечатки хостов запоминаются при первом подключении (`StrictHostKeyChecking=accept-new` в `ansible.cfg`); сменившийся отпечаток известного хоста — ошибка.
- SSH: пользователь `debian` и ключ `~/.ssh/vk_cloud_priv.key` на jump-хосте (права 0600) — одинаковые на всех хостах; меняются в
`inventory/group_vars/validators.yml` (`ansible_user`, `ansible_ssh_private_key_file`). Для docker и записи env-файла сценарий
повышает права через `sudo`.
- Токен в env-файле можно зашифровать: `ansible-vault encrypt env/validator-agent.env`, запускать с `--ask-vault-pass`.
Рабочий `env/validator-agent.env` в git не попадает (`.gitignore`).
## Запуск
```bash
cd deploy/ansible
ansible-playbook playbooks/deploy-validator-agent.yml --limit vkiplab-v1 # канарейка: один валидатор
ansible-playbook playbooks/deploy-validator-agent.yml # все валидаторы волнами
ansible-playbook playbooks/deploy-validator-agent.yml --check # только проверки (preflight), без изменений
```
Волны по умолчанию: 1 хост, затем 4, затем все остальные (`deploy_serial: [1, 4, "100%"]`). Любой сбой в волне останавливает
прогон: следующая волна не начнётся. Все сразу: `-e '{"deploy_serial": ["100%"]}'`.
Что выкатывается — `deploy_ref` (по умолчанию `main`): ветка, тег или коммит. Выкатить нужно **запушенный** коммит:
валидаторы берут код из репозитория, а не с jump-хоста. После правок кода сначала пересоберите `bin/validator-agent`
(см. [SETUP.md](../../docs/SETUP.md#обновление-образов-после-изменения-кода)) и закоммитьте его вместе с `bin/SHA256SUMS`.
### Откат
```bash
ansible-playbook playbooks/deploy-validator-agent.yml -e deploy_ref=<предыдущий коммит или тег>
```
Для тега или коммита клон переходит в detached HEAD; следующий запуск с `deploy_ref=main` возвращает его на ветку.
## Что делает сценарий на каждом хосте
1. **preflight** — env-файл на jump-хосте существует и в нём задан `VALIDATOR_AGENT_CONTROL_API_URL` (адрес-пример
`example.com` не принимается); на валидаторе отвечает Docker, есть git и клон, архитектура x86_64; `validator_id` из inventory совпадает
с `validator_id` работающего контейнера; на хосте нет другого контейнера этого агента (по имени или образу) — иначе рядом со старым
запустился бы второй с тем же `validator_id`. Показывает текущий контейнер.
2. **git** — от имени `git_user` (на валидаторах `root`: клон принадлежит ему) `fetch`, затем клон сбрасывается на `deploy_ref` (`checkout --force`). Через `git`, а не модуль `git`:
учётные данные, уже настроенные в клоне, не трогаются. Локальные правки отслеживаемых файлов в клоне будут сброшены;
неотслеживаемые и игнорируемые (`.env.*`) — нет.
3. **build** — проверка `bin/validator-agent` по `bin/SHA256SUMS`; `docker build --platform linux/amd64` с контекстом в корне репозитория,
образ получает метку ревизии (`cloud-ip-validator-validator-agent:<хеш>`) и `latest`.
4. **replace** — env-файл копируется в `/opt/cloud-ip-validator/deploy/docker/.env.validator` (0600; путь закрыт `.gitignore`),
затем `docker stop` → `docker rm -f` → `docker run -d --restart unless-stopped --cap-add NET_RAW --env-file …`.
5. **verify** — ждёт строку `registered` в логе агента (регистрация в control-api), проверяет, что контейнер запущен, без перезапусков
и на только что собранном образе. При неудаче хост падает с последними строками лога, следующие волны не стартуют.
Затем остаются 3 последних образа с метками ревизий (`keep_images`), остальные и «висячие» удаляются.
Все шаги можно запускать по тегам: `--tags preflight|git|build|replace|verify`.
## Параметры
**`env/validator-agent.env`** — запуск агента (читает `deploy/docker/validator-agent/docker-entrypoint.sh`):
`VALIDATOR_AGENT_CONTROL_API_URL`, `CONTROL_API_AGENT_TOKEN`, `VALIDATOR_AGENT_SELF_CHECK_METHODS`,
`VALIDATOR_AGENT_SELF_CHECK_TIMEOUT_SECONDS`, `VALIDATOR_AGENT_POLL_INTERVAL_SECONDS`, `VALIDATOR_AGENT_HTTPS_TIMEOUT_SECONDS`,
`VALIDATOR_AGENT_ICMP_TIMEOUT_SECONDS`, `VALIDATOR_AGENT_ICMP_COUNT`, `VALIDATOR_AGENT_SSH_ENABLED`, `VALIDATOR_AGENT_SSH_TIMEOUT_SECONDS`.
Смена только env-файла не требует изменения кода: достаточно запустить сценарий (контейнер пересоздаётся с новыми значениями).
**`inventory/group_vars/validators.yml`** — доставка: `repo_dir`, `repo_remote`, `deploy_ref`, `git_user` (пользователь, у которого в клоне
настроен доступ к репозиторию; пусто — SSH-пользователь), `image_name`, `container_name`, `platform`, `keep_images`, `restart_policy`,
`capabilities`, `log_max_size`, `log_max_file`, `stop_timeout`, `verify_retries`, `verify_delay`, `local_env_file`. Любой параметр
переопределяется ключом `-e`.
## Разбор сбоев
- **«Нет env-файла» / «не задан VALIDATOR_AGENT_CONTROL_API_URL»** — см. «Подготовка».
- **«в запущенном контейнере validator_id=…, а в inventory …»** — соответствие `validator-N` и `validator_id` в `inventory/hosts.yml`
неверно для этого хоста: исправьте inventory (контейнер при этом не тронут).
- **Контейнер не прошёл проверку** — в сообщении последние строки лога. `registration failed` — control-api недоступен с валидатора
или `validator_id` не заведён в control-api. Контейнер остаётся на хосте для разбора (`docker logs`).
- **`Permission denied` на `.git/FETCH_HEAD`, `dubious ownership`, `could not read Username`** — git запущен не от владельца клона или
учётные данные есть у другого пользователя: задайте `git_user` (на валидаторах — `root`).
- **«найден другой контейнер агента»** — на хосте есть контейнер с похожим именем или образом, не совпадающий с `container_name`:
проверьте имя (`docker ps -a`) и при необходимости удалите лишний контейнер вручную.
- **`bin/validator-agent` не совпадает с суммой** — в репозитории устарел `bin/SHA256SUMS`: пересоберите бинарник и обновите сумму.
- **Сборка падает на `FROM alpine:3.20` / `apk add`** — с валидатора нет доступа к Docker Hub / репозиториям Alpine.
- **Старый образ другого имени остаётся** — сценарий чистит только образы `image_name`; образ с прежним именем удалите вручную (`docker rmi`).
+15
View File
@@ -0,0 +1,15 @@
# Запускать из каталога deploy/ansible (там же лежит этот файл).
[defaults]
inventory = inventory/hosts.yml
roles_path = roles
forks = 20
retry_files_enabled = False
interpreter_python = auto_silent
callback_result_format = yaml
[ssh_connection]
pipelining = True
# accept-new: отпечаток нового хоста запоминается при первом подключении (без
# интерактивного вопроса); изменившийся отпечаток известного хоста по-прежнему
# приводит к ошибке.
ssh_args = -o ControlMaster=auto -o ControlPersist=60s -o StrictHostKeyChecking=accept-new
+29
View File
@@ -0,0 +1,29 @@
# Параметры запуска validator-agent. Скопируйте в validator-agent.env и заполните:
# cp validator-agent.env.example validator-agent.env && chmod 600 validator-agent.env
# Файл передаётся контейнеру как `docker run --env-file` (формат KEY=VALUE, без
# кавычек и пробелов вокруг "="). Читает их deploy/docker/validator-agent/docker-entrypoint.sh.
# VALIDATOR_AGENT_VALIDATOR_ID сюда НЕ пишется: сценарий подставляет имя хоста.
# Адрес control-api (обязательно). Для внешнего размещения — адрес, доступный
# с валидаторов напрямую (через него же работает способ самопроверки control_api).
VALIDATOR_AGENT_CONTROL_API_URL=https://control-api.example.com
# Токен агентов (CONTROL_API_AGENT_TOKEN на стороне control-api). Пусто — если
# токен на control-api ещё не включён.
CONTROL_API_AGENT_TOKEN=
# Способы самопроверки в порядке приоритета: ip_echo, control_api.
# Самопроверка проходит, если адрес подтвердил любой способ. Без пробелов.
VALIDATOR_AGENT_SELF_CHECK_METHODS=[control_api,ip_echo]
# Таймаут одного способа, секунд.
VALIDATOR_AGENT_SELF_CHECK_TIMEOUT_SECONDS=10
# Период опроса control-api, секунд.
VALIDATOR_AGENT_POLL_INTERVAL_SECONDS=5
# Исходящие проверки.
VALIDATOR_AGENT_HTTPS_TIMEOUT_SECONDS=10
VALIDATOR_AGENT_ICMP_TIMEOUT_SECONDS=5
VALIDATOR_AGENT_ICMP_COUNT=3
VALIDATOR_AGENT_SSH_ENABLED=false
VALIDATOR_AGENT_SSH_TIMEOUT_SECONDS=5
@@ -0,0 +1,50 @@
---
# Параметры доставки validator-agent (не секреты). Любой из них можно
# переопределить в командной строке: -e deploy_ref=<коммит|тег>.
# Параметры запуска самого агента (адрес control-api, токен, способы
# самопроверки, таймауты) лежат в env-файле, см. local_env_file.
# SSH: пользователь и ключ одинаковы на jump-хосте и на валидаторах. Путь к
# ключу — на jump-хосте (ключ с правами 0600). Для docker и записи env-файла
# сценарий повышает права через sudo (become).
ansible_user: debian
ansible_ssh_private_key_file: ~/.ssh/vk_cloud_priv.key
# --- git-клон на валидаторе ---------------------------------------------
repo_dir: /opt/cloud-ip-validator
repo_remote: origin
# Ветка, тег или коммит, который нужно выкатить (откат: -e deploy_ref=<коммит>).
deploy_ref: main
# Пользователь, от имени которого выполняется git в клоне (через sudo): владелец
# клона. На валидаторах клон принадлежит root, репозиторий читается без учётных
# данных. Пусто — git работает от SSH-пользователя без sudo.
git_user: root
# --- образ и контейнер --------------------------------------------------
image_name: cloud-ip-validator-validator-agent
# Имя контейнера на валидаторах (имя образа — другое, см. image_name).
container_name: validator-agent
platform: linux/amd64
dockerfile: deploy/docker/validator-agent/Dockerfile
# Сколько образов с метками ревизий хранить (кроме latest); старые удаляются.
keep_images: 3
# --- запуск контейнера (как в deploy/docker/RUN.txt) --------------------
restart_policy: unless-stopped
capabilities: [NET_RAW]
log_max_size: 10m
log_max_file: "3"
stop_timeout: 10
# --- проверка после запуска ---------------------------------------------
verify_retries: 10
verify_delay: 3
# --- env-файл на jump-хосте ---------------------------------------------
# Параметры запуска агента. Рабочий файл создаётся из validator-agent.env.example
# и в git не попадает. Файл можно зашифровать: ansible-vault encrypt <файл>
# (тогда запускайте с --ask-vault-pass или --vault-password-file).
local_env_file: "{{ playbook_dir }}/../env/validator-agent.env"
# Куда файл копируется на валидатор (путь закрыт .gitignore репозитория,
# переживает git reset).
remote_env_file: "{{ repo_dir }}/deploy/docker/.env.validator"
+69
View File
@@ -0,0 +1,69 @@
# Валидаторы. Имя хоста (validator-N) — это имя ВМ; validator_id в control-api
# другой (vkiplab-vN) и задаётся переменной validator_id. Соответствие
# validator-N -> vkiplab-vN предполагается по порядку номеров; сценарий
# сверяет validator_id с тем, что записано в работающем контейнере на хосте,
# и при расхождении останавливается ДО замены контейнера.
all:
children:
validators:
hosts:
validator-1:
ansible_host: 10.11.12.161
validator_id: vkiplab-v1
validator-2:
ansible_host: 10.11.12.177
validator_id: vkiplab-v2
validator-3:
ansible_host: 10.11.12.33
validator_id: vkiplab-v3
validator-4:
ansible_host: 10.11.12.41
validator_id: vkiplab-v4
validator-5:
ansible_host: 10.11.12.193
validator_id: vkiplab-v5
validator-6:
ansible_host: 10.11.12.197
validator_id: vkiplab-v6
validator-7:
ansible_host: 10.11.12.198
validator_id: vkiplab-v7
validator-8:
ansible_host: 10.11.12.169
validator_id: vkiplab-v8
validator-9:
ansible_host: 10.11.12.185
validator_id: vkiplab-v9
validator-10:
ansible_host: 10.11.12.186
validator_id: vkiplab-v10
validator-11:
ansible_host: 10.11.12.199
validator_id: vkiplab-v11
validator-12:
ansible_host: 10.11.12.196
validator_id: vkiplab-v12
validator-13:
ansible_host: 10.11.12.194
validator_id: vkiplab-v13
validator-14:
ansible_host: 10.11.12.195
validator_id: vkiplab-v14
validator-15:
ansible_host: 10.11.12.65
validator_id: vkiplab-v15
validator-16:
ansible_host: 10.11.12.189
validator_id: vkiplab-v16
validator-17:
ansible_host: 10.11.12.190
validator_id: vkiplab-v17
validator-18:
ansible_host: 10.11.12.168
validator_id: vkiplab-v18
validator-19:
ansible_host: 10.11.12.191
validator_id: vkiplab-v19
validator-20:
ansible_host: 10.11.12.73
validator_id: vkiplab-v20
@@ -0,0 +1,17 @@
---
# Доставка validator-agent на все валидаторы: обновить git-клон, собрать образ
# на каждом хосте, остановить и удалить старый контейнер, поднять новый.
# Запуск (из deploy/ansible): ansible-playbook playbooks/deploy-validator-agent.yml
- name: Deliver validator-agent to the validators
hosts: validators
become: true
gather_facts: false
# Волны: один хост (канарейка), затем четыре, затем все остальные. Любой
# сбой останавливает прогон — следующая волна не начнётся.
# Все сразу: -e '{"deploy_serial": ["100%"]}'.
serial: "{{ deploy_serial }}"
max_fail_percentage: 0
vars:
deploy_serial: [1, 4, "100%"]
roles:
- validator_agent
@@ -0,0 +1,40 @@
---
# Dockerfile ничего не компилирует: в образ копируется закоммиченный
# bin/validator-agent. Не даём выкатить бинарник, не совпадающий с суммой.
- name: Check bin/validator-agent against SHA256SUMS
ansible.builtin.shell: |
set -o pipefail
grep -E '[[:space:]]validator-agent$' SHA256SUMS | sha256sum -c -
args:
chdir: "{{ repo_dir }}/bin"
executable: /bin/bash
changed_when: false
# Контекст сборки — корень репозитория (так и в SETUP.md). Слои кэшируются,
# при смене bin/ образ пересобирается сам.
- name: Build the image
ansible.builtin.command:
argv:
- docker
- build
- --platform
- "{{ platform }}"
- --label
- "git.rev={{ rev_after.stdout }}"
- --label
- deployed.by=ansible
- -t
- "{{ image_ref }}"
- -f
- "{{ dockerfile }}"
- .
chdir: "{{ repo_dir }}"
- name: Tag the image as latest
ansible.builtin.command: "docker tag {{ image_ref }} {{ image_name }}:latest"
- name: Read the image id
ansible.builtin.command:
argv: [docker, image, inspect, --format, "{% raw %}{{.Id}}{% endraw %}", "{{ image_ref }}"]
changed_when: false
register: built_image
@@ -0,0 +1,80 @@
---
# Команды git вместо модуля git: модуль переписывает URL remote и может
# затереть учётные данные, уже настроенные в клоне. Работаем от git_user
# (или от SSH-пользователя, если он не задан).
- name: Remember the current revision
ansible.builtin.command: "{{ git_cmd }} rev-parse HEAD"
become: "{{ git_user | length > 0 }}"
become_user: "{{ git_user }}"
changed_when: false
register: rev_before
- name: Fetch the remote
ansible.builtin.command: "{{ git_cmd }} fetch --prune --tags {{ repo_remote }}"
become: "{{ git_user | length > 0 }}"
become_user: "{{ git_user }}"
changed_when: false
# deploy_ref — ветка, тег или коммит. Ветка берётся из remote (свежая),
# тег и коммит — как есть.
- name: Resolve deploy_ref as a remote branch
ansible.builtin.command: >-
{{ git_cmd }} rev-parse --verify --quiet
refs/remotes/{{ repo_remote }}/{{ deploy_ref }}^{commit}
become: "{{ git_user | length > 0 }}"
become_user: "{{ git_user }}"
changed_when: false
failed_when: false
register: ref_branch
- name: Resolve deploy_ref as a tag or commit
ansible.builtin.command: "{{ git_cmd }} rev-parse --verify --quiet {{ deploy_ref }}^{commit}"
become: "{{ git_user | length > 0 }}"
become_user: "{{ git_user }}"
changed_when: false
failed_when: false
register: ref_other
when: ref_branch.rc != 0
- name: Fail if deploy_ref does not exist
ansible.builtin.assert:
that: ref_branch.rc == 0 or (ref_other.rc | default(1)) == 0
fail_msg: "deploy_ref={{ deploy_ref }} не найден в {{ repo_dir }} ({{ repo_remote }})."
quiet: true
- name: Fix the target revision
ansible.builtin.set_fact:
target_rev: "{{ ref_branch.stdout if ref_branch.rc == 0 else ref_other.stdout }}"
# Ветка: остаёмся на локальной ветке (клон не уходит в detached HEAD),
# сброс на remote. Тег или коммит: detached HEAD.
- name: Check out the branch
ansible.builtin.command: "{{ git_cmd }} checkout --force -B {{ deploy_ref }} {{ target_rev }}"
become: "{{ git_user | length > 0 }}"
become_user: "{{ git_user }}"
when: ref_branch.rc == 0
changed_when: rev_before.stdout != target_rev
- name: Check out the tag or commit
ansible.builtin.command: "{{ git_cmd }} checkout --force --detach {{ target_rev }}"
become: "{{ git_user | length > 0 }}"
become_user: "{{ git_user }}"
when: ref_branch.rc != 0
changed_when: rev_before.stdout != target_rev
- name: Read the deployed revision
ansible.builtin.command: "{{ git_cmd }} rev-parse HEAD"
become: "{{ git_user | length > 0 }}"
become_user: "{{ git_user }}"
changed_when: false
register: rev_after
- name: Check that the clone is at the target revision
ansible.builtin.assert:
that: rev_after.stdout == target_rev
fail_msg: "Клон на {{ rev_after.stdout }}, ожидалось {{ target_rev }}."
quiet: true
- name: Remember the short revision
ansible.builtin.set_fact:
deploy_rev: "{{ rev_after.stdout[:12] }}"
@@ -0,0 +1,33 @@
---
# Порядок важен: сначала всё, что не трогает работающий контейнер (проверки,
# обновление кода, сборка образа), и только потом замена контейнера. Если
# что-то упало до replace, старый контейнер продолжает работать.
- name: Preflight checks
ansible.builtin.import_tasks: preflight.yml
tags: [preflight]
- name: Dry run stops after preflight
ansible.builtin.debug:
msg: "check mode: git, build, replace and verify are skipped"
when: ansible_check_mode
tags: [always]
- name: Update the git clone
ansible.builtin.import_tasks: git.yml
when: not ansible_check_mode
tags: [git]
- name: Build the image
ansible.builtin.import_tasks: build.yml
when: not ansible_check_mode
tags: [build]
- name: Replace the container
ansible.builtin.import_tasks: replace.yml
when: not ansible_check_mode
tags: [replace]
- name: Verify the new container
ansible.builtin.import_tasks: verify.yml
when: not ansible_check_mode
tags: [verify]
@@ -0,0 +1,150 @@
---
# --- на jump-хосте (один раз) -------------------------------------------
- name: Check that the env file exists on the jump host
ansible.builtin.stat:
path: "{{ local_env_file }}"
delegate_to: localhost
become: false
run_once: true
check_mode: false
register: env_file_stat
- name: Fail early without an env file
ansible.builtin.assert:
that: env_file_stat.stat.exists
fail_msg: >-
Нет env-файла {{ local_env_file }}. Создайте его:
cp env/validator-agent.env.example env/validator-agent.env и заполните.
quiet: true
run_once: true
# Содержимое файла (в нём токен) не выводится: разбор идёт в задаче с no_log,
# а проверка и её сообщение — по готовым булевым значениям.
- name: Inspect the env file without printing it
ansible.builtin.set_fact:
env_url_set: "{{ env_file_text is regex('(?m)^VALIDATOR_AGENT_CONTROL_API_URL=\\S+') }}"
env_url_is_example: "{{ env_file_text is regex('(?m)^VALIDATOR_AGENT_CONTROL_API_URL=\\S*example\\.com') }}"
vars:
env_file_text: "{{ lookup('ansible.builtin.file', local_env_file) }}"
run_once: true
no_log: true
- name: Check that the env file sets the control-api address
ansible.builtin.assert:
that:
- env_url_set | bool
- not (env_url_is_example | bool)
fail_msg: >-
В {{ local_env_file }} не задан VALIDATOR_AGENT_CONTROL_API_URL
(или остался адрес-пример example.com).
quiet: true
run_once: true
# --- на каждом валидаторе -----------------------------------------------
- name: Check that Docker answers
ansible.builtin.command: docker version --format {% raw %}'{{.Server.Version}}'{% endraw %}
changed_when: false
check_mode: false
- name: Check that git is installed
ansible.builtin.command: git --version
changed_when: false
check_mode: false
- name: Check that the git clone exists
ansible.builtin.stat:
path: "{{ repo_dir }}/.git"
check_mode: false
register: clone_stat
- name: Fail without a clone
ansible.builtin.assert:
that: clone_stat.stat.exists
fail_msg: "Нет git-клона {{ repo_dir }} на {{ inventory_hostname }}."
quiet: true
- name: Read the CPU architecture
ansible.builtin.command: uname -m
changed_when: false
check_mode: false
register: arch
- name: The image is linux/amd64 only
ansible.builtin.assert:
that: arch.stdout in ['x86_64', 'amd64']
fail_msg: "Архитектура {{ arch.stdout }}: образ {{ platform }} здесь не запустится (exec format error)."
quiet: true
- name: Look at the current container
ansible.builtin.command: >-
docker container inspect --format
{% raw %}'{{.Config.Image}} {{.State.Status}}'{% endraw %}
{{ container_name }}
register: current_container
changed_when: false
failed_when: false
check_mode: false
# Защита от второго агента: если на хосте уже есть другой контейнер этого
# агента (по имени или по образу), сценарий остановится, а не запустит
# рядом ещё один с тем же validator_id.
- name: List containers on the validator
ansible.builtin.command: docker ps -a --format {% raw %}'{{.Names}}|{{.Image}}'{% endraw %}
register: all_containers
changed_when: false
check_mode: false
- name: Check that there is no other agent container
ansible.builtin.assert:
that: (other_agents | from_json) | length == 0
fail_msg: >-
{{ inventory_hostname }}: найден другой контейнер агента: {{ (other_agents | from_json) | join(', ') }}.
Сценарий заменяет только контейнер {{ container_name }}. Проверьте container_name в
inventory/group_vars/validators.yml или удалите лишний контейнер вручную.
quiet: true
vars:
other_agents: >-
{%- set found = [] -%}
{%- for line in all_containers.stdout_lines -%}
{%- set row = line.split('|') -%}
{%- if row[0] != container_name and ('validator-agent' in row[0] or row[1] == image_name or row[1].startswith(image_name ~ ':')) -%}
{%- set _ = found.append(row[0] ~ ' (' ~ row[1] ~ ')') -%}
{%- endif -%}
{%- endfor -%}
{{- found | to_json -}}
# validator_id работающего контейнера — эталон: если он отличается от
# inventory, заменять контейнер нельзя (агент зарегистрировался бы под чужим
# именем, адреса привязывались бы к порту другой ВМ). Выводится только он,
# а не все переменные окружения (там токен).
- name: Read validator_id of the running container
ansible.builtin.shell: |
set -o pipefail
docker container inspect --format '{% raw %}{{range .Config.Env}}{{println .}}{{end}}{% endraw %}' {{ container_name }} \
| sed -n 's/^VALIDATOR_AGENT_VALIDATOR_ID=//p'
args:
executable: /bin/bash
register: running_validator_id
changed_when: false
failed_when: false
check_mode: false
when: current_container.rc == 0
- name: Check validator_id against the running container
ansible.builtin.assert:
that: >-
current_container.rc != 0
or (running_validator_id.stdout | trim) == ''
or (running_validator_id.stdout | trim) == effective_validator_id
fail_msg: >-
{{ inventory_hostname }}: в запущенном контейнере validator_id={{ running_validator_id.stdout | default('') | trim }},
а в inventory {{ effective_validator_id }}. Проверьте соответствие имени ВМ и validator_id
в inventory/hosts.yml; контейнер не тронут.
quiet: true
- name: Report the current container
ansible.builtin.debug:
msg: >-
{{ container_name }}:
{{ current_container.stdout if current_container.rc == 0 else 'контейнера нет (будет создан)' }};
validator_id для запуска: {{ effective_validator_id }}
@@ -0,0 +1,45 @@
---
# Образ уже собран: простой валидатора — только stop + rm + run.
- name: Copy the env file to the validator
ansible.builtin.copy:
src: "{{ local_env_file }}"
dest: "{{ remote_env_file }}"
owner: root
group: root
mode: "0600"
no_log: true
- name: Check whether the container exists
ansible.builtin.command: "docker container inspect {{ container_name }}"
register: container_exists
changed_when: false
failed_when: false
- name: Stop the current container
ansible.builtin.command: "docker stop -t {{ stop_timeout }} {{ container_name }}"
when: container_exists.rc == 0
- name: Remove the current container
ansible.builtin.command: "docker rm -f {{ container_name }}"
when: container_exists.rc == 0
# --restart нужен: агент завершается, если регистрация в control-api не
# удалась, и должен подняться снова. validator_id берётся из inventory
# (validator_id) и перекрывает env-файл.
- name: Start the new container
ansible.builtin.command:
argv: >-
{{ ['docker', 'run', '-d',
'--name', container_name,
'--restart', restart_policy,
'--platform', platform,
'--env-file', remote_env_file,
'-e', 'VALIDATOR_AGENT_VALIDATOR_ID=' ~ effective_validator_id,
'--log-driver', 'json-file',
'--log-opt', 'max-size=' ~ log_max_size,
'--log-opt', 'max-file=' ~ log_max_file,
'--label', 'git.rev=' ~ rev_after.stdout,
'--label', 'deployed.by=ansible']
+ (capabilities | map('regex_replace', '^(.*)$', '--cap-add=\1') | list)
+ [image_ref] }}
register: started
@@ -0,0 +1,65 @@
---
- name: Verify the new container
block:
# Агент пишет "registered" после успешной регистрации в control-api.
- name: Wait for the agent to register in control-api
ansible.builtin.command: "docker logs --tail 200 {{ container_name }}"
register: agent_logs
changed_when: false
until: agent_logs.stdout is search('msg=registered validator_id=' ~ effective_validator_id ~ '(\s|$)') or agent_logs.stderr is search('msg=registered validator_id=' ~ effective_validator_id ~ '(\s|$)')
retries: "{{ verify_retries | int }}"
delay: "{{ verify_delay | int }}"
- name: Inspect the container
ansible.builtin.command:
argv: [docker, inspect, --format, "{% raw %}{{.State.Running}} {{.RestartCount}} {{.Image}}{% endraw %}", "{{ container_name }}"]
register: container_state
changed_when: false
- name: Check the container state
ansible.builtin.assert:
that:
- container_state.stdout.split()[0] == 'true'
- container_state.stdout.split()[1] == '0'
- container_state.stdout.split()[2] == built_image.stdout
fail_msg: >-
Контейнер {{ container_name }} в состоянии «{{ container_state.stdout }}»
(ожидалось: запущен, 0 перезапусков, образ {{ built_image.stdout }}).
quiet: true
rescue:
- name: Collect the container log
ansible.builtin.command: "docker logs --tail 30 {{ container_name }}"
register: failed_logs
changed_when: false
failed_when: false
- name: Fail the host and stop the next waves
ansible.builtin.fail:
msg: |-
{{ inventory_hostname }}: контейнер не прошёл проверку после запуска.
Последние строки лога:
{{ failed_logs.stdout }}{{ failed_logs.stderr }}
# Старые образы с метками ревизий: оставляем keep_images последних (docker
# выводит от новых к старым), образ работающего контейнера docker не удалит.
- name: List image tags
ansible.builtin.command:
argv: [docker, images, "{{ image_name }}", --format, "{% raw %}{{.Tag}}{% endraw %}"]
register: image_tags
changed_when: false
- name: Remove old revision images
ansible.builtin.command: "docker rmi {{ image_name }}:{{ item }}"
loop: "{{ (image_tags.stdout_lines | reject('equalto', 'latest') | list)[keep_images | int:] }}"
changed_when: true
failed_when: false
- name: Remove dangling images
ansible.builtin.command: docker image prune -f
changed_when: false
- name: Summary
ansible.builtin.debug:
msg: >-
{{ inventory_hostname }} ({{ effective_validator_id }}): {{ deploy_rev }} ({{ deploy_ref }}),
образ {{ built_image.stdout[:19] }}, контейнер {{ container_name }} запущен
@@ -0,0 +1,7 @@
---
# git с явным safe.directory: клон может принадлежать другому пользователю.
git_cmd: "git -c safe.directory={{ repo_dir }} -C {{ repo_dir }}"
# validator_id в control-api: из inventory, иначе имя хоста.
effective_validator_id: "{{ validator_id | default(inventory_hostname) }}"
# Образ с меткой ревизии, собранный в этом прогоне.
image_ref: "{{ image_name }}:{{ deploy_rev | default('unknown') }}"
+1
View File
@@ -100,6 +100,7 @@ services:
VALIDATOR_AGENT_CONTROL_API_URL: "${VALIDATOR_AGENT_CONTROL_API_URL:-http://control-api:8080}"
VALIDATOR_AGENT_POLL_INTERVAL_SECONDS: "${VALIDATOR_AGENT_POLL_INTERVAL_SECONDS:-5}"
VALIDATOR_AGENT_SELF_CHECK_TIMEOUT_SECONDS: "${VALIDATOR_AGENT_SELF_CHECK_TIMEOUT_SECONDS:-10}"
VALIDATOR_AGENT_SELF_CHECK_METHODS: "${VALIDATOR_AGENT_SELF_CHECK_METHODS:-[ip_echo]}"
VALIDATOR_AGENT_HTTPS_TIMEOUT_SECONDS: "${VALIDATOR_AGENT_HTTPS_TIMEOUT_SECONDS:-10}"
VALIDATOR_AGENT_ICMP_TIMEOUT_SECONDS: "${VALIDATOR_AGENT_ICMP_TIMEOUT_SECONDS:-5}"
VALIDATOR_AGENT_ICMP_COUNT: "${VALIDATOR_AGENT_ICMP_COUNT:-3}"
@@ -6,6 +6,7 @@ set -eu
export VALIDATOR_AGENT_POLL_INTERVAL_SECONDS="${VALIDATOR_AGENT_POLL_INTERVAL_SECONDS:-5}"
export VALIDATOR_AGENT_SELF_CHECK_TIMEOUT_SECONDS="${VALIDATOR_AGENT_SELF_CHECK_TIMEOUT_SECONDS:-10}"
export VALIDATOR_AGENT_SELF_CHECK_METHODS="${VALIDATOR_AGENT_SELF_CHECK_METHODS:-[ip_echo]}"
export VALIDATOR_AGENT_HTTPS_TIMEOUT_SECONDS="${VALIDATOR_AGENT_HTTPS_TIMEOUT_SECONDS:-10}"
export VALIDATOR_AGENT_ICMP_TIMEOUT_SECONDS="${VALIDATOR_AGENT_ICMP_TIMEOUT_SECONDS:-5}"
export VALIDATOR_AGENT_ICMP_COUNT="${VALIDATOR_AGENT_ICMP_COUNT:-3}"
@@ -16,7 +17,7 @@ export VALIDATOR_AGENT_SSH_TIMEOUT_SECONDS="${VALIDATOR_AGENT_SSH_TIMEOUT_SECOND
# NOT templated into the YAML: the binary reads them straight from this
# container's environment (names default in the config loader), so secrets
# never land in a file inside the container.
envsubst '${VALIDATOR_AGENT_VALIDATOR_ID} ${VALIDATOR_AGENT_CONTROL_API_URL} ${VALIDATOR_AGENT_POLL_INTERVAL_SECONDS} ${VALIDATOR_AGENT_SELF_CHECK_TIMEOUT_SECONDS} ${VALIDATOR_AGENT_HTTPS_TIMEOUT_SECONDS} ${VALIDATOR_AGENT_ICMP_TIMEOUT_SECONDS} ${VALIDATOR_AGENT_ICMP_COUNT} ${VALIDATOR_AGENT_SSH_ENABLED} ${VALIDATOR_AGENT_SSH_TIMEOUT_SECONDS}' \
envsubst '${VALIDATOR_AGENT_VALIDATOR_ID} ${VALIDATOR_AGENT_CONTROL_API_URL} ${VALIDATOR_AGENT_POLL_INTERVAL_SECONDS} ${VALIDATOR_AGENT_SELF_CHECK_TIMEOUT_SECONDS} ${VALIDATOR_AGENT_SELF_CHECK_METHODS} ${VALIDATOR_AGENT_HTTPS_TIMEOUT_SECONDS} ${VALIDATOR_AGENT_ICMP_TIMEOUT_SECONDS} ${VALIDATOR_AGENT_ICMP_COUNT} ${VALIDATOR_AGENT_SSH_ENABLED} ${VALIDATOR_AGENT_SSH_TIMEOUT_SECONDS}' \
< /etc/validator-agent/validator-agent.yaml.tmpl > /etc/validator-agent/validator-agent.yaml
exec /usr/local/bin/validator-agent -config /etc/validator-agent/validator-agent.yaml
@@ -4,6 +4,7 @@ poll_interval_seconds: ${VALIDATOR_AGENT_POLL_INTERVAL_SECONDS}
self_check:
timeout_seconds: ${VALIDATOR_AGENT_SELF_CHECK_TIMEOUT_SECONDS}
methods: ${VALIDATOR_AGENT_SELF_CHECK_METHODS}
checks:
https_timeout_seconds: ${VALIDATOR_AGENT_HTTPS_TIMEOUT_SECONDS}
+248
View File
@@ -0,0 +1,248 @@
# Ручная очистка базы данных control-api (SQL)
Когда нужна: подготовка к новому полному прогону, разбор после инцидента, освобождение места, удаление отдельных адресов.
Все примеры проверены 2026-10-02 на копии боевой БД (20 валидаторов, 4 площадки, 6445 адресов в реестре, 29 889 проверок,
19 897 событий): без ошибок, `integrity_check` = `ok`, нарушений внешних ключей нет.
> **Сначала API.** Если в очереди есть адреса в работе, очищайте очередь штатно: кнопка «Очистить всё» на `/ips` или
> `POST /api/v1/admin/ips/clear` (см. [USAGE.md](USAGE.md#удаление-адресов-из-очереди)). Только так control-api отвяжет Floating IP от портов
> валидаторов. SQL ниже работает с базой «в покое»: он не обращается к OpenStack.
## 1. Что в базе и что нельзя трогать
| Группа | Таблицы | Можно чистить |
|---|---|---|
| Данные прогона | `ip_queue` (очередь), `ip_registry` (реестр адресов), `checks` (реестр проверок), `ip_site_checks` (признаки площадок по адресам в работе), `events` (журнал событий) | да |
| Настройки (не трогать) | `validators`, `sites`, `target_groups` (цели), `check_types`, `inbound_checks_settings`, `settings`, `auto_cycle` | **нет** |
| Служебное | `sqlite_sequence` (нумерация записей), `PRAGMA user_version` (версия схемы) | нумерацию можно сбросить, версию не менять |
Связи (внешние ключи): `checks`, `events`, `ip_site_checks` ссылаются на `ip_queue`; `checks`, `events`, `ip_queue` — на `ip_registry`;
`validators.current_ip_id` — на `ip_queue`. Поэтому порядок удаления всегда такой: сначала `validators.current_ip_id` в `NULL`, затем
`ip_site_checks`, `checks`, `events`, `ip_queue` и в конце `ip_registry`. Времена в БД хранятся строками `2026-10-02T07:14:58.857Z` (UTC).
## 2. Подготовка (всегда, перед любой очисткой)
**2.1. Где лежит база и чем её открывать.** Нужен клиент `sqlite3` на хосте (`apt install sqlite3`).
| Развёртывание | Файл базы |
|---|---|
| `rxprod-compose/` (боевой стенд) | `rxprod-compose/capi-db/control-api.db` |
| Docker (`deploy/docker`) | в томе `cloud-ip-validator-db`: `docker volume inspect cloud-ip-validator-db --format '{{.Mountpoint}}'`, файл `control-api.db` в этом каталоге (нужен root) |
| systemd | `/var/lib/cloud-ip-validator/control-api.db` (`database.path` в `control-api.yaml`) |
Дальше в примерах `DB=путь/к/control-api.db`.
**2.2. Проверить, что в очереди ничего не в работе:**
```bash
sqlite3 -readonly "$DB" "
SELECT state, COUNT(*) FROM ip_queue
WHERE state NOT IN ('done', 'failed', 'occupied') GROUP BY state;"
```
Пустой вывод — всё завершено. Адреса в работе есть — сначала «Очистить всё» через API (см. выше). Затем проверьте в OpenStack, что на портах
валидаторов нет лишних Floating IP; по базе видно только то, что control-api считает привязанным:
```bash
sqlite3 -readonly "$DB" "
SELECT ip_address, state, fip_id FROM ip_queue
WHERE fip_id <> '' AND state NOT IN ('done', 'failed', 'occupied');"
```
**2.3. Остановить control-api.** Он держит базу открытой и пишет в неё на каждом такте; ручные правки поверх работающего процесса
ненадёжны.
```bash
cd rxprod-compose && docker compose stop control-api # compose-развёртывание
sudo systemctl stop control-api # systemd
```
**2.4. Сделать резервную копию** (штатной командой SQLite, не `cp`: у базы есть журнал `-wal`):
```bash
TS=$(date -u +%Y-%m-%d_%H-%M)
sqlite3 "$DB" ".backup '$(dirname "$DB")/backup-$TS-before-cleanup.db'"
sqlite3 -readonly "$(dirname "$DB")/backup-$TS-before-cleanup.db" "PRAGMA integrity_check;" # должно быть: ok
```
Лог контейнера до чистки при необходимости сохраните отдельно: `docker logs <контейнер> > backup-$TS.container.log 2>&1`.
**2.5. Посмотреть, что и сколько лежит** (до и после чистки):
```bash
sqlite3 -readonly "$DB" "
SELECT 'ip_queue' AS tbl, COUNT(*) AS n FROM ip_queue
UNION ALL SELECT 'ip_registry', COUNT(*) FROM ip_registry
UNION ALL SELECT 'checks', COUNT(*) FROM checks
UNION ALL SELECT 'ip_site_checks', COUNT(*) FROM ip_site_checks
UNION ALL SELECT 'events', COUNT(*) FROM events
UNION ALL SELECT 'validators', COUNT(*) FROM validators
UNION ALL SELECT 'sites', COUNT(*) FROM sites;"
```
## 3. Сценарии
Команды выполняются так: `sqlite3 "$DB"` и вставить блок, либо сохранить блок в файл и выполнить `sqlite3 "$DB" < файл.sql`.
### 3.1. Полный сброс данных прогона (перед новым полным прогоном)
Очищает очередь, реестр адресов, реестр проверок, журнал событий. Валидаторы, площадки, цели, типы проверок и все настройки остаются.
```sql
PRAGMA foreign_keys = ON;
BEGIN;
-- валидаторы больше не ссылаются на адреса очереди
UPDATE validators SET current_ip_id = NULL WHERE current_ip_id IS NOT NULL;
UPDATE validators SET state = 'idle' WHERE state = 'assigned';
-- порядок важен: сначала зависимые таблицы
DELETE FROM ip_site_checks;
DELETE FROM checks;
DELETE FROM events;
DELETE FROM ip_queue;
DELETE FROM ip_registry;
-- нумерация снова с 1 (необязательно)
DELETE FROM sqlite_sequence WHERE name IN ('ip_registry', 'checks', 'ip_queue', 'events');
COMMIT;
```
Затем освободите место (отдельной командой, не внутри транзакции):
```sql
PRAGMA wal_checkpoint(TRUNCATE);
VACUUM;
```
Файл сжимается до сотен килобайт (на проверочной копии: 20 МБ → 128 КБ).
### 3.2. Только журнал событий
Очередь, реестр и проверки не затрагиваются.
Весь журнал:
```sql
DELETE FROM events;
DELETE FROM sqlite_sequence WHERE name = 'events';
```
Только старше 7 дней (число дней меняйте в `'-7 days'`):
```sql
DELETE FROM events
WHERE occurred_at < strftime('%Y-%m-%dT%H:%M:%fZ', 'now', '-7 days');
```
### 3.3. Только реестр проверок (история проверок)
Адреса и их итоги в очереди остаются; пропадает подробная история проверок (на странице «Реестр» исчезнут результаты).
Вся история:
```sql
DELETE FROM checks;
DELETE FROM sqlite_sequence WHERE name = 'checks';
```
Оставить последние 3 цикла каждого адреса (число `3` меняйте):
```sql
DELETE FROM checks
WHERE cycle_id <= (SELECT MAX(c2.cycle_id) FROM checks c2 WHERE c2.registry_id = checks.registry_id) - 3;
```
Постоянное ограничение глубины истории лучше задать настройкой `history_retention_cycles` (страница `/settings` или
`PUT /api/v1/admin/config/orchestrator`): control-api сам подрезает историю при завершении каждого адреса. SQL выше нужен для разовой чистки.
### 3.4. Удалить конкретные адреса целиком
Удаляет адрес из очереди и реестра вместе со всей его историей (проверки и события). Список адресов подставьте в первую команду
`CREATE TEMP TABLE doomed_reg`. Адрес, который сейчас проверяется, удалять этим способом нельзя: используйте API (раздел 5).
```sql
PRAGMA foreign_keys = ON;
BEGIN;
CREATE TEMP TABLE doomed_reg AS
SELECT id FROM ip_registry WHERE ip_address IN ('5.188.140.6', '5.188.140.62'); -- ваши адреса
CREATE TEMP TABLE doomed_ip AS
SELECT id FROM ip_queue WHERE registry_id IN (SELECT id FROM doomed_reg);
UPDATE validators SET current_ip_id = NULL WHERE current_ip_id IN (SELECT id FROM doomed_ip);
DELETE FROM ip_site_checks WHERE ip_id IN (SELECT id FROM doomed_ip);
DELETE FROM checks WHERE registry_id IN (SELECT id FROM doomed_reg);
DELETE FROM events WHERE registry_id IN (SELECT id FROM doomed_reg) OR ip_id IN (SELECT id FROM doomed_ip);
DELETE FROM ip_queue WHERE id IN (SELECT id FROM doomed_ip);
DELETE FROM ip_registry WHERE id IN (SELECT id FROM doomed_reg);
DROP TABLE doomed_ip;
DROP TABLE doomed_reg;
COMMIT;
```
### 3.5. Только освободить место
Если удалили много, а файл не уменьшился (SQLite не отдаёт место ОС до `VACUUM`):
```sql
PRAGMA wal_checkpoint(TRUNCATE);
VACUUM;
```
Размер и свободные страницы:
```sql
SELECT page_count * page_size / 1024 AS size_kb, freelist_count * page_size / 1024 AS free_kb
FROM pragma_page_count(), pragma_page_size(), pragma_freelist_count();
```
## 4. После очистки: запуск и проверка
```bash
cd rxprod-compose && docker compose up -d --no-deps control-api # или: sudo systemctl start control-api
```
Если нужно, чтобы и лог контейнера начался с нуля, пересоздайте контейнер: `docker compose up -d --force-recreate --no-deps control-api`
(старый лог сохраните заранее, п. 2.4).
Проверка базы (до запуска или на копии):
```bash
sqlite3 -readonly "$DB" "
PRAGMA integrity_check;
PRAGMA foreign_key_check;
SELECT state, COUNT(*) FROM validators GROUP BY state;
SELECT COUNT(*) AS validators_with_address FROM validators WHERE current_ip_id IS NOT NULL;
SELECT COUNT(*) AS queue_rows FROM ip_queue;"
```
Ожидается: `ok`; пустой результат `foreign_key_check`; у валидаторов состояние `idle`; `validators_with_address` = 0; для полного сброса `queue_rows` = 0.
Через API: `GET /api/v1/admin/status` (`total_ips` = 0, `total_validators` = число валидаторов) и `GET /api/v1/admin/validators`
(через 10–15 секунд после запуска у всех свежий `last_heartbeat_at`).
## 5. Что делать через API, а не через SQL
| Задача | Как |
|---|---|
| Остановить проверку адреса, удалить адрес в работе | `POST /api/v1/admin/ips/{ip}/cancel`, `DELETE /api/v1/admin/ips/{ip}` |
| Очистить всю очередь с отвязкой Floating IP | `POST /api/v1/admin/ips/clear` |
| Перепроверить завершённые адреса | `POST /api/v1/admin/ips` со списком адресов |
| Валидатор «завис» с адресом | ничего не править: лизинг истечёт, адрес вернётся в очередь, валидатор освободится сам |
| Добавить/убрать валидатор, площадку, цель | `/api/v1/admin/config/*` или страницы дашборда |
## 6. Восстановление из копии
```bash
cd rxprod-compose && docker compose stop control-api
cp capi-db/backup-<метка>-before-cleanup.db capi-db/control-api.db
rm -f capi-db/control-api.db-wal capi-db/control-api.db-shm # старый журнал к новой копии не относится
docker compose up -d --no-deps control-api
```
## 7. Ловушки
- Не выполняйте `DELETE` при работающем control-api: он пишет в ту же базу.
- Не удаляйте строки настроек (`validators`, `sites`, `target_groups`, `check_types`, `inbound_checks_settings`, `settings`, `auto_cycle`):
после этого control-api либо не стартует, либо работает без площадок и целей.
- Не нарушайте порядок удаления (раздел 1) и не отключайте `PRAGMA foreign_keys = ON` в блоках выше: база сама остановит ошибочное удаление.
- `VACUUM` нельзя вызывать внутри транзакции и пока control-api запущен.
- Не копируйте файл базы командой `cp` при работающем процессе: журнал `-wal` останется в неконсистентном состоянии. Используйте `.backup`.
- Не меняйте `PRAGMA user_version`: по нему control-api применяет миграции схемы.
- Не правьте состояние адресов и валидаторов вручную (`state`, `owner_validator_id`, `lease_expires_at`) вместо API: control-api сверяет их на каждом такте и исправит расхождение,
но до этого результат непредсказуем.
+32 -4
View File
@@ -41,10 +41,10 @@ JSON, базовый префикс прикладных методов — `/ap
|---|---|---|
| **admin** | `CONTROL_API_ADMIN_TOKEN` | все `/api/v1/admin/*` (очередь, реестр, автоцикл, `config/*`) |
| **agent** | `CONTROL_API_AGENT_TOKEN` | запись результатов: `POST /agents/{id}/self-check`, `/events`, `/results`, `/complete` и `POST /probers/{site_id}/results` |
| **открыто** | — | `GET /healthz`; `POST /agents/register`, `POST /agents/{id}/heartbeat`, `GET /agents/{id}/assignment`; `POST /probers/register`, `POST /probers/{site_id}/heartbeat`, `GET /probers/{site_id}/assignments` |
| **открыто** | — | `GET /healthz`; `POST /agents/register`, `POST /agents/{id}/heartbeat`, `GET /agents/{id}/assignment`, `GET /agents/{id}/observed-ip`; `POST /probers/register`, `POST /probers/{site_id}/heartbeat`, `GET /probers/{site_id}/assignments` |
- Токены разные: токен администратора **не** подходит для методов агентов, и наоборот.
- Валидатор и пробер могут без токена зарегистрироваться, слать heartbeat и забирать задание (настройку); отправка результатов без токена агентов — `401`.
- Валидатор и пробер могут без токена зарегистрироваться, слать heartbeat и забирать задание (настройку); валидатор также может спросить, с какого адреса его видит control-api (`observed-ip`); отправка результатов без токена агентов — `401`.
- Имена переменных меняются в секции `auth` конфига control-api (`admin_token_env`, `agent_token_env`); сами значения в YAML не хранятся.
- Ответ при отказе: `401 {"error": "unauthorized"}` с заголовком `WWW-Authenticate: Bearer`.
- Токен не задан (пустая переменная) — уровень открыт; в логе control-api при старте предупреждение. Токены нужно генерировать случайными: `openssl rand -hex 32`.
@@ -145,6 +145,30 @@ curl -s -H "Authorization: Bearer $ADMIN_TOKEN" http://<control-api>:8080/api/v1
`check_config` — уже развёрнутая конфигурация проверок (тип + список
целей), агенту не нужно самому сопоставлять группы целей.
### `GET /api/v1/agents/{id}/observed-ip`
С какого адреса control-api видит соединение валидатора. Используется
self-check способом `control_api` (`self_check.methods` в
`validator-agent.yaml`) как альтернатива внешнему IP-echo сервису. Уровень
доступа — открыто: отдаётся только адрес самого вызывающего.
Ответ `200`:
```json
{"ip": "203.0.113.10", "source": "remote_addr"}
```
- Адрес берётся только из адреса TCP-соединения (`RemoteAddr`), приведённого
к каноничному виду (`::ffff:1.2.3.4` → `1.2.3.4`). Заголовки
`X-Forwarded-For` / `X-Real-IP` **не учитываются**: иначе валидатор мог бы
подделать адрес и пройти проверку. Метод рассчитан на прямое подключение
без обратного прокси.
- `404`, если `validator_id` не зарегистрирован.
- Способ корректен, только если соединение выходит через внешнюю сеть
(SNAT Floating IP). Если control-api достижим из облака по внутренней
сети, он увидит приватный адрес валидатора. Если порт control-api
опубликован через Docker, проверьте, что ручка показывает внешний адрес
клиента, а не адрес шлюза Docker.
### `POST /api/v1/agents/{id}/self-check`
Отчёт о результате self-check — подтверждение, что исходящий трафик
@@ -152,7 +176,11 @@ curl -s -H "Authorization: Bearer $ADMIN_TOKEN" http://<control-api>:8080/api/v1
определяет это **сам**, обращаясь к внешнему (снаружи облака) IP-echo
сервису (`self_check.ip_echo_urls` в `validator-agent.yaml`, например
`api.ipify.org`) и сравнивая ответ с `ip_address` из задания — control-api
в этом определении не участвует. Важно, что ресурс должен быть именно
в этом определении не участвует. Дополнительно можно включить способ
`control_api` (`self_check.methods`): агент спрашивает у control-api через
`GET /agents/{id}/observed-ip`, с какого адреса тот его видит. Способы
пробуются по приоритету, достаточно подтверждения любым из них; без
настройки работает только IP-echo. Важно, что ресурс должен быть именно
внешним: OpenStack применяет SNAT через Floating IP только к трафику,
уходящему через внешнюю сеть, поэтому обращение к чему-либо внутри
проекта (в том числе к самому control-api, если он в той же внутренней
@@ -165,7 +193,7 @@ curl -s -H "Authorization: Bearer $ADMIN_TOKEN" http://<control-api>:8080/api/v1
"ip_id": 42,
"detected_egress_ip": "203.0.113.10",
"success": true,
"detail": "matched"
"detail": "matched (control_api)"
}
```
+18 -5
View File
@@ -701,10 +701,17 @@ docker run -d --platform linux/amd64 --cap-add NET_RAW --name validator-agent \
| `VALIDATOR_AGENT_SSH_TIMEOUT_SECONDS` | нет | `5` |
`--cap-add NET_RAW` обязателен для ICMP-проверок, как и у `prober`.
`self_check.ip_echo_urls` в переменные не вынесен — при отсутствии в
конфиге агент сам подставляет дефолт (`api.ipify.org`, `ifconfig.me`);
свой список задавайте через смонтированный конфиг вместо шаблона, если
нужно переопределить.
`self_check.ip_echo_urls` и `self_check.methods` в переменные не вынесены —
при отсутствии в конфиге агент сам подставляет дефолты (`api.ipify.org`,
`ifconfig.me` и `methods: [ip_echo]`); свои значения задавайте через
смонтированный конфиг вместо шаблона, если нужно переопределить.
`methods` — способы самопроверки в порядке приоритета (`ip_echo`,
`control_api`), достаточно подтверждения любым. Способ `control_api`
спрашивает у control-api, с какого адреса он видит валидатора
(`GET /agents/{id}/observed-ip`); при внешнем размещении control-api
рекомендуется `[control_api, ip_echo]`. Ограничение: если control-api
достижим из облака по внутренней сети, он увидит приватный адрес
валидатора и этот способ всегда даст несовпадение — используйте `ip_echo`.
### Обновление образов после изменения кода
@@ -732,6 +739,10 @@ docker compose up -d --build
сборке) и `docker rm -f <имя> && docker run ... ` (или `docker restart`,
если менялись только переменные окружения, а не сам бинарник/образ).
Для массового обновления валидаторов (git-клон, сборка образа на хосте,
замена контейнера, проверка регистрации) есть Ansible-сценарий:
[`deploy/ansible/`](../deploy/ansible/README.md).
### Диагностика Docker-развёртывания
- **Контейнер сразу падает, в логах `exec format error`** — образ собран
@@ -804,7 +815,9 @@ curl -s http://<control-api>:8080/api/v1/admin/validators | python3 -m json.tool
через floating IP, который в данный момент привязан к валидатору.
- `validator-agent` → внешние IP-echo сервисы из `self_check.ip_echo_urls`
(по умолчанию `api.ipify.org`, `ifconfig.me`) — **обязательно вне
облака**: это и есть механизм self-check (см.
облака**: это и есть механизм self-check способом `ip_echo` (при
`self_check.methods` с `control_api` достаточно ещё и доступа к control-api
по внешней сети; см.
[DIAGRAMS.md](DIAGRAMS.md#2-поток-данных-от-валидатора-к-целевому-серверу-egress-проверка)).
Если валидатор не может достучаться ни до одного из этих адресов,
self-check никогда не пройдёт и IP будет бесконечно возвращаться в
+26 -2
View File
@@ -408,6 +408,17 @@ curl -s http://<control-api>:8080/api/v1/admin/validators | python3 -m json.tool
(занят), `unreachable` (пропустил heartbeat дольше
`orchestrator.heartbeat_timeout_seconds`).
Правила, которые держат состояние валидатора согласованным:
- валидатор держит **не более одного адреса**; адрес освобождает валидатор
только пока он остаётся его текущим (запоздалое завершение старого адреса
чужого валидатора не освобождает);
- после пропущенного heartbeat валидатор, у которого есть адрес, возвращается
в `assigned`, а не в `idle`, и не получает второй адрес; без адреса — в `idle`;
- `unreachable`-валидатор не получает адресов, пока не пришлёт heartbeat
(даже если лизинг его адреса истёк и адрес вернулся в очередь);
- на каждом такте оркестратор сверяет валидаторы с очередью и исправляет
расхождения (в логе `repaired validators that disagreed with the queue`).
**Добавление нового валидатора (без перезапуска control-api):**
1. Поднимите новую ВМ в сервисном проекте облака, узнайте её Neutron
`port_id`.
@@ -690,6 +701,12 @@ curl -s -X POST http://<control-api>:8080/api/v1/admin/ips/delete \
curl -s -X POST http://<control-api>:8080/api/v1/admin/ips/clear
```
Очистка отвязывает Floating IP только у адресов, которые ещё в работе
(завершённые уже свободны), затем сама опрашивает порты валидаторов в облаке и
снимает оставшиеся привязки адресов из реестра. Занимает секунды. Она не
прерывается разрывом соединения (таймаутом клиента или дашборда): операция
доводится до конца на стороне control-api, предел — 10 минут.
В `admin-dashboard` то же самое доступно на странице `/ips`: чекбоксы у
каждой строки + кнопка «Удалить выбранные» для точечного/массового
удаления, кнопка «Удалить» в каждой строке, и отдельная кнопка «Очистить
@@ -742,7 +759,9 @@ curl -s -X PUT http://<control-api>:8080/api/v1/admin/config/orchestrator \
дело обычно в self-check: он запрашивает внешние (вне облака) сервисы из
`self_check.ip_echo_urls` в конфиге валидатора (по умолчанию
`api.ipify.org`, `ifconfig.me`) — если у ВМ-валидатора нет исходящего
доступа в интернет к этим адресам, запрос не проходит вообще, и агент
доступа в интернет к этим адресам, запрос не проходит вообще (при
`self_check.methods: [control_api, ip_echo]` агент сперва спросит адрес у
control-api, и проверка может пройти и без IP-echo), и агент
даже не может *сообщить* результат control-api (ни успешный, ни
неуспешный) — тогда статус реально зависает до истечения
`orchestrator.lease_ttl_seconds`, после чего адрес возвращается в
@@ -763,7 +782,12 @@ https://api.ipify.org`) и логи `journalctl -u validator-agent` на пре
облака (см. `self_check.ip_echo_urls`) — запрос к чему-либо внутри
проекта (в том числе к самому control-api, если он в той же внутренней
сети) покажет приватный адрес валидатора независимо от того, правильно
ли привязан FIP, и всегда будет давать ложный провал.
ли привязан FIP, и всегда будет давать ложный провал. Это относится и к
способу `control_api` (`self_check.methods`): он корректен только когда
валидатор ходит к control-api через внешнюю сеть; при внутреннем доступе в
`detail` будет подсказка про приватный адрес — оставьте `ip_echo`. В
`detail` события `self_check_result` указан сработавший способ
(`matched (control_api)`) либо причина по каждому способу.
**Площадка (`site-N`) никогда не отчитывается по конкретному IP.**
Сперва проверьте статус самой площадки — `GET
@@ -0,0 +1,110 @@
# План: самопроверка через control-api (дополнительный способ сверки публичного IP)
> Дата: 2026-10-02 03:06 · Статус: **реализовано и проверено** (юнит-тесты, локальный e2e с `ip_echo` и с `control_api`) (решения пользователя — в конце)
## Зачем
Самопроверка агента подтверждает, что исходящий трафик валидатора идёт через выданный Floating IP: агент спрашивает свой
публичный адрес у внешнего IP-echo сервиса и сверяет с назначенным адресом. Сейчас это единственный способ, и он хрупкий:
1 октября таймауты `https://ifconfig.me/ip` дали волну провалов self-check (33 случая за вечер).
`control-api` в текущем развёртывании стоит **во внешнем окружении** (вне облака), валидаторы подключаются к нему **напрямую**.
Значит, соединение валидатора с ним выходит наружу через Floating IP, и `control-api` сам видит публичный адрес источника.
Это даёт второй способ сверки без сторонних сервисов: агент спрашивает у `control-api`, с какого адреса тот его видит.
Требование: **существующий способ (IP-echo) сохраняется**, новый добавляется как опция агента.
## Решение в двух строках
1. `control-api` получает ручку «с какого адреса ты меня видишь».
2. Агент получает настройку `self_check.methods` — список способов в порядке приоритета; по умолчанию `[ip_echo]` (всё как сейчас).
## Конфигурация агента
```yaml
self_check:
timeout_seconds: 10
methods: [control_api, ip_echo] # по умолчанию [ip_echo]
ip_echo_urls: [...] # без изменений
```
- Допустимые значения: `ip_echo` (текущий способ), `control_api` (новый). Неизвестное значение — ошибка при старте агента.
- **Самопроверка успешна, если её подтвердил любой из способов.** Способы пробуются по порядку приоритета (первый —
главный); остановка на первом успешном. К следующему способу переходим и при отсутствии ответа (ошибка, таймаут,
404/5xx), и при несовпадении адреса. Провал — только если не подтвердил ни один способ; в `detail` попадает причина по
каждому способу.
- Внутри способа `ip_echo` поведение прежнее: URL перебираются по порядку, переход к следующему URL только при ошибке.
- Пустой список или отсутствие ключа → `[ip_echo]`. Агент без новой настройки ведёт себя ровно как раньше.
- Для внешнего размещения `control-api` в примерах и рекомендациях стоит `[control_api, ip_echo]`: способ через API в приоритете.
- Итог в `detail`: `detected_egress_ip=<ip> matched (control_api)`; видно в событиях адреса и в дашборде.
## Control API
**Новая ручка:** `GET /api/v1/agents/{id}/observed-ip` → `200 {"ip":"90.156.213.5","source":"remote_addr"}`.
- Уровень доступа — **открыто**, как `heartbeat` и `assignment`: ручка отдаёт только адрес самого вызывающего, секретов нет.
- `{id}` должен быть известным валидатором, иначе `404` (чтобы ручка не превращалась в публичный «узнай свой IP»).
- Адрес берётся только из `r.RemoteAddr`, приводится к каноничному виду (`::ffff:1.2.3.4` → `1.2.3.4`).
- Заголовки `X-Forwarded-For`/`X-Real-IP` **не учитываются**: подключение прямое, а доверие к заголовку позволило бы
валидатору подделать адрес и пройти проверку. Если появится обратный прокси, понадобится отдельная настройка
доверенных прокси — сейчас она не нужна (решение 1).
## Агент (`internal/agentcore`)
- Новая функция `detectViaControlAPI`: `GET /api/v1/agents/{id}/observed-ip` на `control_api_url`.
- **Новое TCP-соединение на каждый вызов** (отдельный `http.Transport` с `DisableKeepAlives`). Это ключевой момент:
соединение, открытое до привязки Floating IP (heartbeat, assignment), остаётся в старом NAT-состоянии и покажет
прежний адрес; общий клиент `apiclient` использовать нельзя.
- Токен агентов в этот запрос не нужен (ручка открытая) и не отправляется.
- Таймаут `self_check.timeout_seconds` (сейчас 10 с) действует **на каждый способ отдельно**: при общем дедлайне зависший
первый способ (приоритетный `control_api`) съел бы всё время, и запасной не успел бы ответить. Общий предел — таймаут × число способов.
- `detectPublicIP` заменяется перебором `methods` по приоритету; `fetchIPEcho` не меняется.
- Диагностика: если `control-api` вернул **частный** адрес (RFC 1918 и т. п.), в `detail` пишется подсказка: «control-api
доступен по внутренней сети, самопроверка через него невозможна; используйте ip_echo».
## Ограничение способа (важно для документации)
Способ `control_api` корректен только если соединение валидатора с `control-api` **выходит через внешнюю сеть**
(SNAT Floating IP). Если `control-api` достижим из облака по внутренней сети, он увидит частный адрес валидатора, и
этот способ всегда будет давать несовпадение (при `[control_api, ip_echo]` проверка пройдёт по `ip_echo`).
Если порт `control-api` опубликован через Docker, при выкладке проверить, что ручка показывает внешний адрес клиента, а
не адрес шлюза Docker.
## Откат и совместимость
- Новый агент + старый `control-api`: ручки нет (`404`), при `methods: [control_api, ip_echo]` агент переходит на `ip_echo`.
- Старый агент + новый `control-api`: ничего не меняется, ручка просто не вызывается.
- Откат: убрать `methods` из конфига агента (или вернуть `[ip_echo]`) и перезапустить агент.
- Схема БД и протокол `self-check` (`POST /agents/{id}/self-check`) не меняются.
## Затрагиваемые файлы
| Файл | Изменение |
|---|---|
| `internal/config/config.go` | `SelfCheckCfg.Methods`, дефолт `[ip_echo]` и проверка значений |
| `internal/httpapi/routes.go`, `handlers_agent.go` | маршрут и обработчик `observed-ip` |
| `internal/agentcore/agentcore.go` | `detectViaControlAPI`, перебор `methods` по приоритету |
| `configs/validator-agent.example.yaml` | пример и комментарии |
| `docs/API.md`, `docs/SETUP.md`, `docs/USAGE.md`, `README.md` | описание ручки, опции, ограничения |
| `bin/validator-agent`, `bin/control-api`, `SHA256SUMS` | пересборка |
## Тесты
- **config:** дефолт `[ip_echo]`; допустимые значения; ошибка на неизвестном способе.
- **httpapi:** прямой адрес; IPv4-mapped IPv6; заголовок `X-Forwarded-For` игнорируется; неизвестный валидатор → `404`.
- **agentcore:** порядок способов; успех второго способа после ошибки или несовпадения первого; провал, когда не
подтвердил ни один; каждый вызов открывает новое соединение (тестовый сервер считает соединения); подсказка про частный адрес.
- **e2e:** `scripts/run-local-e2e.sh` остаётся на `ip_echo` (проверка обратной совместимости).
## Выкладка
1. `control-api` с новой ручкой (поведение не меняется): пересборка образа, перезапуск.
2. Агенты на ВМ-валидаторах: новый `bin/validator-agent` и `methods: [control_api, ip_echo]` в их конфиге. Один валидатор
для начала, проверить `detail` в событиях (`matched (control_api)`), затем остальные.
## Решения пользователя (2026-10-02)
1. Подключение валидаторов к `control-api` — **напрямую** (без обратного прокси): `trusted_proxies` не нужен.
2. Ручка `observed-ip` — **открытая**.
3. Достаточно **одной успешной самопроверки любым из способов**; при внешнем размещении API способ через ручку API —
**в приоритете** (первый в списке).
@@ -0,0 +1,146 @@
# План: исправление оркестратора (двойная выдача валидатору) и очистки очереди
> Дата: 2026-10-02 09:02 UTC · Статус: **реализовано и проверено** (юнит-тесты с `-race`, локальный e2e; выкладка и приёмка на стенде — ниже). Решения пользователя: предел очистки 10 минут и 8 параллельных отвязок приняты как базовые
> Основание: [analysis/2026-10-02_08-56_1026-addresses_mass-check-analysis.md](../../analysis/2026-10-02_08-56_1026-addresses_mass-check-analysis.md)
## Context
Массовая проверка 2 октября остановилась на 1026 из 6440 адресов. 7 из 20 валидаторов «залипли»: v1, v12, v13 не взяли
ни одного задания после 07:17 и 07:33; v7, v16, v17, v3 залипали временно. Итог: 198 сбросов лизинга, 41 ошибка привязки
Floating IP (`409 fixed IP already has a floating IP`), **42 адреса в `fail` без единой выполненной проверки**, потеря
~27% пропускной способности. Отдельно «Очистить всё» не уложилась в таймаут клиента (256 с вместо секунд) и сначала
оборвалась с 600 ошибками `context canceled`.
Цель: валидатор в любой момент держит не больше одного адреса; потеря heartbeat не приводит к двойной выдаче;
«Очистить всё» выполняется за секунды и не прерывается разрывом соединения.
## Причины (по коду, подтверждены логами и БД)
| № | Причина | Где |
|---|---|---|
| 1 | Heartbeat возвращает `unreachable` → `idle`, не глядя на `current_ip_id`: занятый валидатор снова считается свободным | `internal/db/queries_validators.go:41` (`Heartbeat`), `:31` (`RegisterValidator`, возврат из `unreachable`) |
| 2 | Освобождение валидатора идёт **по имени**, а не по адресу, который он держит: завершение старого адреса освобождает валидатор, уже взявший новый. Так же `RequeueOrFail` (в т. ч. при сбросе лизинга) и `MarkFIPOccupied`, `FreeValidator`. Освобождённый ставится в `idle` даже если он `unreachable` — мёртвый валидатор получает новые адреса каждые 3 минуты | `internal/db/queries_ipqueue.go` (`ReleaseFIP`, `RequeueOrFail`, `MarkFIPOccupied`), `queries_validators.go:143` (`FreeValidator`), `internal/orchestrator/orchestrator.go:469` |
| 3 | Моя регрессия: защита от дублей привязки ключуется по валидатору (`assign:<validator>`). Вторая выдача того же валидатора не запускает привязку и стоит в `assigning_fip` до конца лизинга | `internal/orchestrator/orchestrator.go:149` |
| 4 | Агент шлёт heartbeat только между заданиями. Адрес с 3–4 таймаутами внешних проверок (~40 с) блокирует его дольше порога 30 с | `internal/agentcore/agentcore.go` (`Run`, `pollOnce`) |
| 5 | «Очистить всё» отвязывает FIP у **всех** строк с непустым `fip_id`, а он остаётся у `done`/`failed`. Больше 1000 последовательных вызовов OpenStack | `internal/db/queries_ipqueue.go` (`ListFIPRefs`, `ListFIPRefsByAddresses`), `internal/orchestrator/orchestrator.go` (`ClearQueue`, `DeleteIPs`) |
| 6 | Очистка работает на контексте HTTP-запроса: разрыв соединения клиентом обрывает её посреди дела (отвязано часть, БД не очищена) | `internal/httpapi/handlers_admin.go:322` |
## Инварианты, которые вводим
- **I1.** Валидатор держит не более одного адреса: `validators.current_ip_id = X` тогда и только тогда, когда у строки `X`
`owner_validator_id` равен этому валидатору и состояние не терминальное (`done`, `failed`, `occupied`).
- **I2.** `idle` означает `current_ip_id IS NULL`. Состояния `unreachable` и `unregistered` не затираются освобождением.
- **I3.** Валидатор освобождает только тот адрес, который он сейчас держит. Освобождение и возврат по лизингу чужого или
устаревшего адреса состояние валидатора не меняют.
## Изменения
### 1. База данных (без миграций, только запросы)
`internal/db/queries_validators.go`, `queries_ipqueue.go`:
- **`Heartbeat`:** `unreachable` → `assigned`, если `current_ip_id IS NOT NULL`, иначе `idle`.
То же в `RegisterValidator` (возврат из `unreachable`/`unregistered`).
- **Общая функция освобождения** `freeValidatorTx(tx, validatorID, ipID)`:
`UPDATE validators SET current_ip_id=NULL, state = CASE WHEN state='unreachable' THEN state ELSE 'idle' END WHERE validator_id=? AND current_ip_id=?`.
Используют `ReleaseFIP`, `RequeueOrFail`, `MarkFIPOccupied`, `FreeValidator` (получает второй аргумент — id адреса).
`deleteIPTx` (уже по `current_ip_id`) и `ClearAllIPs` переводятся на тот же `CASE`, чтобы не затирать `unreachable`.
- **`ClaimNextQueued`:** условие обновления валидатора дополняется `AND current_ip_id IS NULL`.
- **`ReconcileValidators`** (новая, вызывается из `Orchestrator.Tick`): лечит нарушение инвариантов, если они всё же возникли
(падение процесса, старые строки): валидатор с `current_ip_id`, чья строка не существует, терминальна или принадлежит другому
валидатору, освобождается; строка в `assigning_fip`/`awaiting_self_check`/`checking`, чей владелец не ссылается на неё,
не трогается (её вернёт сброс лизинга). Один короткий запрос на такт.
- **`ListFIPRefs` и `ListFIPRefsByAddresses`:** только строки в нетерминальных состояниях (`state NOT IN done, failed, occupied`)
с непустым `fip_id`: у терминальных FIP уже отвязан (агрегация, возврат по лизингу, отмена и `occupied` делают это до записи
состояния). Значения `fip_id` в строках не меняются (дашборд их показывает).
### 2. Оркестратор (`internal/orchestrator/orchestrator.go`)
- Ключ защиты от дублей привязки: `assign:<id адреса>`, а не `assign:<validator>` (строка 149). Дублирующий запуск привязки
одного и того же адреса по-прежнему исключён.
- `ForceCancel`, `sweepExpiredLeases`, `aggregateAndRelease`, `SelfCheckResult`: передают id адреса в освобождение (I3).
- **`ClearQueue`/`DeleteIPs`:** отвязка FIP из `ListFIPRefs` (теперь ≤ числа валидаторов) выполняется параллельно, не более
8 одновременных вызовов; затем прямой опрос портов (`releaseValidatorPorts`) как сейчас.
- `Tick`: вызывает `ReconcileValidators` первым шагом.
### 3. HTTP (`internal/httpapi/handlers_admin.go`)
- «Очистить всё», удаление списка и отмена: контекст отвязан от отмены запроса (`context.WithoutCancel`) с собственным пределом
времени (10 минут). Разрыв соединения клиентом (в т. ч. таймаут дашборда) больше не обрывает операцию на середине.
### 4. Агент (`internal/agentcore/agentcore.go`)
- Heartbeat уходит из `pollOnce` в **отдельную горутину** со своим тикером (период `poll_interval_seconds`); останавливается по
отмене контекста. Долгие внешние проверки больше не блокируют heartbeat. Ошибки heartbeat пишутся в лог (предупреждение).
- Протокол и конфигурация агента не меняются (старый агент с новым сервером и наоборот работают).
### 5. Документация
`docs/USAGE.md`/`docs/DIAGRAMS.md` (состояния валидатора и правила освобождения), запись в истории изменений `README.md`.
## Тесты
Пишу сам (агенты тесты не делают); каждый тест должен падать без исправления.
- **db:**
- `Heartbeat` из `unreachable` с адресом даёт `assigned`, без адреса `idle`; `RegisterValidator` аналогично;
- `ReleaseFIP`/`RequeueOrFail`/`MarkFIPOccupied` старого адреса не освобождают валидатор, который уже держит другой адрес;
- освобождение `unreachable`-валидатора оставляет `unreachable`;
- `ClaimNextQueued` не выдаёт адрес валидатору с `current_ip_id`;
- `ReconcileValidators` лечит битые строки и не трогает корректные;
- `ListFIPRefs`/`ListFIPRefsByAddresses` не содержат `done`/`failed`/`occupied`.
- **orchestrator:**
- сценарий инцидента: валидатор помечен `unreachable`, держа адрес A с идущими проверками → heartbeat → такт оркестратора
**не выдаёт** ему B; после завершения A валидатор получает B;
- мёртвый валидатор (`unreachable`, лизинг истёк) не получает новых адресов;
- `ClearQueue` при 1000 строк `done` с `fip_id` и 5 активных не вызывает `Disassociate` для `done` (счётчик вызовов в обёртке OpenStack);
- защита привязки по адресу (два адреса одного валидатора, искусственно, оба привязываются);
- **случайный сценарий под нагрузкой** (20 валидаторов, 300 адресов, случайная потеря heartbeat, долгие проверки, часть адресов
с отказом привязки): после каждого такта проверяются I1–I3; в конце все адреса завершены, сбросов лизинга нет.
- **httpapi:** «Очистить всё» доживает до конца при отмене контекста запроса посередине (БД очищена, FIP отвязаны).
- **agentcore:** во время долгой проверки (цель отвечает 3 с) heartbeat уходит по расписанию; остановка по отмене контекста.
- **Общий прогон:** `go vet`, `go test -race ./...`, `scripts/run-local-e2e.sh` (с `ip_echo` и `control_api`).
## Выкладка
1. Правки control-api: сборка `bin/control-api`, образ `civ-capi`, перезапуск (БД не затрагивается; миграций нет).
Одного этого достаточно, чтобы остановить двойные выдачи и бесконечное «залипание».
2. Агент: сборка `bin/validator-agent`, коммит, раскатка Ansible-сценарием `deploy/ansible` (запускает пользователь): heartbeat в
отдельном потоке убирает ложные `unreachable`.
3. Перепроверка 42 адресов: список выгружается из снимка анализа в файл `analysis/2026-10-02_08-56_failed-addresses.txt` (уже выгружен, 42 адреса); ставятся в очередь
через `POST /api/v1/admin/ips` (или «Перепроверка» в дашборде) после выкладки.
## Приёмка на стенде
Контрольная группа из 20 адресов, затем 400 адресов (при тех же внешних целях с долгими таймаутами) с наблюдением 30 минут:
| Показатель | Критерий |
|---|---|
| `lease_expired` | 0 |
| Ошибки привязки `409` | 0 |
| Двойные выдачи (два `claimed ip` одному валидатору за <10 с) | 0 |
| Завершено каждым валидатором | отклонение от среднего не больше 15% |
| `unreachable` при долгих проверках | нет (после раскатки нового агента) |
| «Очистить всё» при >1000 строк `done` | ответ не дольше 10 с |
| Адреса `fail` | только по существу (не из-за лизинга) |
## Риски и откат
- Изменения только в запросах и логике, без миграций; схема БД и протокол агента не меняются. Откат — предыдущий образ `civ-capi`
(`docker tag`/предыдущий коммит) и прежний бинарник агента.
- Риск: условные `UPDATE` по `current_ip_id` могут оставить валидатор занятым, если строка адреса пропала. Страхует `ReconcileValidators`.
- Валидатор `unreachable` теперь не получает адресов до первого heartbeat — это намеренно; если агент жив, но heartbeat по сети
не проходит, он простаивает (раньше брал адреса и терял их).
## Не входит в эту правку
- Параллельное выполнение внешних проверок в агенте (3–4 таймаута сейчас идут подряд): сократило бы цикл с ~84 с и снизило бы
нагрузку на heartbeat; отдельное решение.
- Судьба цели `packages.ubuntu.com` (71% `partial`) и причины недоступности проберов — отдельные вопросы из анализа.
- Сокращение `fip_settle_seconds` (30 с) ради темпа.
## Вопросы к согласованию
1. Предел времени «Очистить всё» (10 минут) и число параллельных отвязок (8) — подходят?
2. Выкладывать control-api сразу после тестов (до раскатки агента) — да, как в разделе «Выкладка»?
3. 42 адреса `fail` перепроверять сразу после выкладки или вместе со следующим большим прогоном?
+158 -15
View File
@@ -10,13 +10,16 @@ package agentcore
import (
"context"
"encoding/json"
"fmt"
"io"
"log/slog"
"net"
"net/http"
neturl "net/url"
"os"
"strings"
"sync/atomic"
"time"
"cloudipvalidator/internal/apiclient"
@@ -31,6 +34,10 @@ type Agent struct {
lastHandledIPID int64
// busy is true while an assignment is being worked on; it is reported in
// the heartbeat body (informational on the control-api side).
busy atomic.Bool
// registerRetryInitial/Max govern the backoff used while waiting for a
// successful registration (see registerWithRetry): control-api may not
// be up yet at agent boot, or may come and go across a redeploy, and the
@@ -68,6 +75,15 @@ func (a *Agent) Run(ctx context.Context) error {
}
interval := time.Duration(a.cfg.PollIntervalSeconds) * time.Second
// Heartbeats run on their own schedule. Sent from the poll loop they
// stopped for as long as a slow assignment took (an address whose
// outbound targets all time out keeps the loop busy for ~40 s), which
// control-api reads as a lost validator after heartbeat_timeout_seconds.
hbCtx, stopHeartbeat := context.WithCancel(ctx)
defer stopHeartbeat()
go a.heartbeatLoop(hbCtx, interval)
ticker := time.NewTicker(interval)
defer ticker.Stop()
@@ -146,12 +162,32 @@ type checkConfigDTO struct {
Targets []string `json:"targets"`
}
func (a *Agent) pollOnce(ctx context.Context) {
if _, err := a.client.Do(ctx, "POST", "/api/v1/agents/"+a.cfg.ValidatorID+"/heartbeat", heartbeatReq{LocalState: "idle"}, nil); err != nil {
a.log.Error("heartbeat", "err", err)
return
// heartbeatLoop sends a heartbeat now and then every interval until ctx is
// cancelled.
func (a *Agent) heartbeatLoop(ctx context.Context, interval time.Duration) {
ticker := time.NewTicker(interval)
defer ticker.Stop()
for {
a.sendHeartbeat(ctx)
select {
case <-ctx.Done():
return
case <-ticker.C:
}
}
}
func (a *Agent) sendHeartbeat(ctx context.Context) {
state := "idle"
if a.busy.Load() {
state = "checking"
}
if _, err := a.client.Do(ctx, "POST", "/api/v1/agents/"+a.cfg.ValidatorID+"/heartbeat", heartbeatReq{LocalState: state}, nil); err != nil && ctx.Err() == nil {
a.log.Error("heartbeat", "err", err)
}
}
func (a *Agent) pollOnce(ctx context.Context) {
var assignment assignmentResp
ok, err := a.client.Do(ctx, "GET", "/api/v1/agents/"+a.cfg.ValidatorID+"/assignment", nil, &assignment)
if err != nil {
@@ -167,6 +203,8 @@ func (a *Agent) pollOnce(ctx context.Context) {
return // already handled this IP's work this attempt
}
a.busy.Store(true)
defer a.busy.Store(false)
switch assignment.Phase {
case "awaiting_self_check":
a.handleSelfCheckAndRun(ctx, assignment)
@@ -181,18 +219,13 @@ func (a *Agent) pollOnce(ctx context.Context) {
func (a *Agent) handleSelfCheckAndRun(ctx context.Context, assignment assignmentResp) {
a.postEvent(ctx, assignment.IPID, "config_received", "")
timeout := time.Duration(a.cfg.SelfCheck.TimeoutSeconds) * time.Second
// Each method gets the full timeout (see runSelfCheckMethods), so the
// overall budget scales with the number of methods.
timeout := time.Duration(a.cfg.SelfCheck.TimeoutSeconds) * time.Second * time.Duration(len(a.selfCheckMethods()))
selfCtx, cancel := context.WithTimeout(ctx, timeout)
defer cancel()
detectedIP, err := a.detectPublicIP(selfCtx)
success := err == nil && detectedIP == assignment.IPAddress
detail := "matched"
if err != nil {
detail = "ip echo request failed: " + err.Error()
} else if !success {
detail = fmt.Sprintf("egress ip %q does not match assigned fip %q", detectedIP, assignment.IPAddress)
}
detectedIP, _, detail, success := a.runSelfCheckMethods(selfCtx, assignment.IPAddress)
a.postSelfCheck(ctx, assignment.IPID, detectedIP, success, detail)
a.postEvent(ctx, assignment.IPID, "self_check_result", fmt.Sprintf(`{"success":%t}`, success))
@@ -204,7 +237,117 @@ func (a *Agent) handleSelfCheckAndRun(ctx context.Context, assignment assignment
a.runChecks(ctx, assignment)
}
// detectPublicIP asks each configured IP-echo URL, in order, for the
// selfCheckMethods returns the configured methods in priority order. Configs
// built without the loader (tests) may leave the list empty; that means the
// historical behaviour, ip_echo only.
func (a *Agent) selfCheckMethods() []string {
if len(a.cfg.SelfCheck.Methods) == 0 {
return []string{config.SelfCheckIPEcho}
}
return a.cfg.SelfCheck.Methods
}
// runSelfCheckMethods tries the configured methods in priority order and
// stops at the first one that confirms assignedIP. A method that gives no
// answer and one that reports a different address are treated alike: the
// next method is tried, since either may be a limitation of that method
// (e.g. control-api reached over the internal network sees a private
// address) rather than proof the floating IP is not attached. The check
// fails only when no method confirms, and detail then carries the reason
// from every method. detectedIP is the matching address on success, else the
// last address any method reported (may be empty).
//
// Every method runs under its own self_check.timeout_seconds: with one shared
// deadline a hung first method (priority control_api) would use it all up and
// the fallback would never get a chance to answer.
func (a *Agent) runSelfCheckMethods(ctx context.Context, assignedIP string) (detectedIP, method, detail string, ok bool) {
var reasons []string
perMethod := time.Duration(a.cfg.SelfCheck.TimeoutSeconds) * time.Second
for _, m := range a.selfCheckMethods() {
mctx, cancel := ctx, context.CancelFunc(func() {})
if perMethod > 0 {
mctx, cancel = context.WithTimeout(ctx, perMethod)
}
var ip string
var err error
switch m {
case config.SelfCheckControlAPI:
ip, err = a.detectViaControlAPI(mctx)
case config.SelfCheckIPEcho:
ip, err = a.detectViaIPEcho(mctx)
if err != nil {
err = fmt.Errorf("ip echo request failed: %w", err)
}
default:
err = fmt.Errorf("unknown self-check method")
}
cancel()
if err != nil {
reasons = append(reasons, m+": "+err.Error())
continue
}
if ip == assignedIP {
return ip, m, fmt.Sprintf("matched (%s)", m), true
}
detectedIP = ip
reason := fmt.Sprintf("%s: egress ip %q does not match assigned fip %q", m, ip, assignedIP)
if m == config.SelfCheckControlAPI && isLocalAddr(ip) {
reason += " (control-api sees a private address; it is reachable over the internal network, self-check via control_api is not possible, use ip_echo)"
}
reasons = append(reasons, reason)
}
if len(reasons) == 0 {
reasons = append(reasons, "no self-check methods configured")
}
return detectedIP, "", strings.Join(reasons, "; "), false
}
// isLocalAddr reports whether ip is a private, loopback or link-local
// address, i.e. one that can never be a floating IP.
func isLocalAddr(ip string) bool {
parsed := net.ParseIP(ip)
return parsed != nil && (parsed.IsPrivate() || parsed.IsLoopback() || parsed.IsLinkLocalUnicast())
}
// detectViaControlAPI asks control-api which source address it sees for this
// validator. It is only meaningful when control-api is reached over the
// external network, where the floating IP is the visible source (see
// config.SelfCheckCfg).
//
// Every call dials a brand-new TCP connection through a dedicated transport
// with keep-alives off: a connection opened before the floating IP was
// attached (heartbeat, assignment polling) keeps its old NAT state and would
// keep reporting the previous address, so the shared apiclient must not be
// used. The endpoint is open, so no agent token is sent.
func (a *Agent) detectViaControlAPI(ctx context.Context) (string, error) {
url := strings.TrimRight(a.cfg.ControlAPIURL, "/") + "/api/v1/agents/" + neturl.PathEscape(a.cfg.ValidatorID) + "/observed-ip"
req, err := http.NewRequestWithContext(ctx, http.MethodGet, url, nil)
if err != nil {
return "", fmt.Errorf("build request: %w", err)
}
transport := &http.Transport{DisableKeepAlives: true}
defer transport.CloseIdleConnections()
resp, err := (&http.Client{Transport: transport}).Do(req)
if err != nil {
return "", err
}
defer resp.Body.Close()
if resp.StatusCode < 200 || resp.StatusCode > 299 {
return "", fmt.Errorf("unexpected status %d", resp.StatusCode)
}
var body struct {
IP string `json:"ip"`
}
if err := json.NewDecoder(io.LimitReader(resp.Body, 4096)).Decode(&body); err != nil {
return "", fmt.Errorf("decode response: %w", err)
}
if net.ParseIP(body.IP) == nil {
return "", fmt.Errorf("response is not a valid IP: %q", body.IP)
}
return body.IP, nil
}
// detectViaIPEcho asks each configured IP-echo URL, in order, for the
// address this validator is currently seen egressing from, returning the
// first one that answers with a parseable IP. These must be resources
// genuinely outside the cloud project (see config.SelfCheckCfg) — OpenStack
@@ -212,7 +355,7 @@ func (a *Agent) handleSelfCheckAndRun(ctx context.Context, assignment assignment
// network, so anything reachable over the project's internal network would
// report the validator's private address instead, regardless of whether
// the floating IP is correctly attached.
func (a *Agent) detectPublicIP(ctx context.Context) (string, error) {
func (a *Agent) detectViaIPEcho(ctx context.Context) (string, error) {
var lastErr error
for _, url := range a.cfg.SelfCheck.IPEchoURLs {
ip, err := fetchIPEcho(ctx, url)
+72
View File
@@ -0,0 +1,72 @@
package agentcore
import (
"context"
"fmt"
"net/http"
"net/http/httptest"
"sync/atomic"
"testing"
"time"
"cloudipvalidator/internal/config"
)
// While the agent is busy with a slow assignment (an address whose outbound
// targets time out keeps it occupied for tens of seconds) it must keep sending
// heartbeats; control-api marks a validator that stays silent for
// heartbeat_timeout_seconds as unreachable.
func TestHeartbeatContinuesDuringSlowChecks(t *testing.T) {
slowTarget := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
time.Sleep(3500 * time.Millisecond)
}))
defer slowTarget.Close()
var heartbeats, assignments int32
mux := http.NewServeMux()
mux.HandleFunc("POST /api/v1/agents/register", func(w http.ResponseWriter, r *http.Request) {
fmt.Fprint(w, `{"ok":true}`)
})
mux.HandleFunc("POST /api/v1/agents/val-1/heartbeat", func(w http.ResponseWriter, r *http.Request) {
atomic.AddInt32(&heartbeats, 1)
fmt.Fprint(w, `{"ok":true}`)
})
mux.HandleFunc("GET /api/v1/agents/val-1/assignment", func(w http.ResponseWriter, r *http.Request) {
if atomic.AddInt32(&assignments, 1) > 1 {
w.WriteHeader(http.StatusNoContent)
return
}
fmt.Fprintf(w, `{"ip_id":1,"ip_address":"1.1.1.1","phase":"checking","check_config":[{"type":"https","targets":[%q]}]}`, slowTarget.URL)
})
mux.HandleFunc("POST /api/v1/agents/val-1/results", func(w http.ResponseWriter, r *http.Request) { fmt.Fprint(w, `{"ok":true}`) })
mux.HandleFunc("POST /api/v1/agents/val-1/complete", func(w http.ResponseWriter, r *http.Request) { fmt.Fprint(w, `{"ok":true}`) })
capi := httptest.NewServer(mux)
defer capi.Close()
a := New(&config.ValidatorAgent{
ValidatorID: "val-1", ControlAPIURL: capi.URL, PollIntervalSeconds: 1,
Checks: config.AgentChecks{HTTPSTimeoutSeconds: 10, ICMPTimeoutSeconds: 1, ICMPCount: 1},
}, testLogger())
ctx, cancel := context.WithCancel(context.Background())
done := make(chan struct{})
go func() { _ = a.Run(ctx); close(done) }()
time.Sleep(3 * time.Second) // the slow check (3.5 s) is still running
during := atomic.LoadInt32(&heartbeats)
cancel()
select {
case <-done:
case <-time.After(10 * time.Second):
t.Fatal("Run did not stop after the context was cancelled")
}
// One per second plus the first: 3-4 in 3 s. With heartbeats in the poll
// loop there is exactly one, sent before the slow assignment started.
if during < 3 {
t.Fatalf("%d heartbeats in 3 s while a check was running, want at least 3", during)
}
if atomic.LoadInt32(&assignments) < 1 {
t.Fatal("the assignment was never fetched, the test did not exercise a busy agent")
}
}
+232
View File
@@ -0,0 +1,232 @@
package agentcore
import (
"context"
"fmt"
"net"
"net/http"
"net/http/httptest"
"strings"
"sync/atomic"
"testing"
"time"
"cloudipvalidator/internal/config"
)
// echoServer answers every request with body (an IP-echo stand-in).
func echoServer(t *testing.T, body string) *httptest.Server {
t.Helper()
ts := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
fmt.Fprint(w, body)
}))
t.Cleanup(ts.Close)
return ts
}
// controlAPIServer plays control-api's observed-ip route: it answers with ip
// (status 200) or with the given error status when ip is empty.
func controlAPIServer(t *testing.T, ip string, status int) *httptest.Server {
t.Helper()
ts := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
if r.URL.Path != "/api/v1/agents/val-1/observed-ip" {
http.NotFound(w, r)
return
}
if ip == "" {
w.WriteHeader(status)
return
}
fmt.Fprintf(w, `{"ip":%q,"source":"remote_addr"}`, ip)
}))
t.Cleanup(ts.Close)
return ts
}
func selfCheckAgent(controlAPIURL string, methods []string, echoURLs ...string) *Agent {
return &Agent{
cfg: &config.ValidatorAgent{
ValidatorID: "val-1",
ControlAPIURL: controlAPIURL,
SelfCheck: config.SelfCheckCfg{TimeoutSeconds: 2, Methods: methods, IPEchoURLs: echoURLs},
},
log: testLogger(),
}
}
func TestSelfCheckControlAPIFirstWins(t *testing.T) {
capi := controlAPIServer(t, "1.2.3.4", 0)
var echoCalls int32
echo := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
atomic.AddInt32(&echoCalls, 1)
fmt.Fprint(w, "1.2.3.4")
}))
defer echo.Close()
a := selfCheckAgent(capi.URL, []string{config.SelfCheckControlAPI, config.SelfCheckIPEcho}, echo.URL)
ip, method, detail, ok := a.runSelfCheckMethods(context.Background(), "1.2.3.4")
if !ok || ip != "1.2.3.4" || method != config.SelfCheckControlAPI || detail != "matched (control_api)" {
t.Fatalf("got ip=%q method=%q detail=%q ok=%v", ip, method, detail, ok)
}
if n := atomic.LoadInt32(&echoCalls); n != 0 {
t.Fatalf("ip_echo was called %d times although control_api already confirmed", n)
}
}
// Any one confirming method is enough: control-api answers with a different
// address, ip_echo confirms.
func TestSelfCheckFallsThroughOnMismatch(t *testing.T) {
capi := controlAPIServer(t, "5.5.5.5", 0)
echo := echoServer(t, "1.2.3.4")
a := selfCheckAgent(capi.URL, []string{config.SelfCheckControlAPI, config.SelfCheckIPEcho}, echo.URL)
ip, method, _, ok := a.runSelfCheckMethods(context.Background(), "1.2.3.4")
if !ok || ip != "1.2.3.4" || method != config.SelfCheckIPEcho {
t.Fatalf("got ip=%q method=%q ok=%v, want a pass via ip_echo", ip, method, ok)
}
}
// An old control-api without the route (404) or a failing one (5xx) must not
// stop the self-check: the next method decides.
func TestSelfCheckFallsThroughOnControlAPIError(t *testing.T) {
for _, status := range []int{http.StatusNotFound, http.StatusInternalServerError} {
capi := controlAPIServer(t, "", status)
echo := echoServer(t, "1.2.3.4")
a := selfCheckAgent(capi.URL, []string{config.SelfCheckControlAPI, config.SelfCheckIPEcho}, echo.URL)
_, method, _, ok := a.runSelfCheckMethods(context.Background(), "1.2.3.4")
if !ok || method != config.SelfCheckIPEcho {
t.Fatalf("status %d: method=%q ok=%v, want a pass via ip_echo", status, method, ok)
}
}
}
func TestSelfCheckFailsWhenNoMethodConfirms(t *testing.T) {
capi := controlAPIServer(t, "10.0.0.5", 0) // private address: internal-network case
echo := echoServer(t, "6.6.6.6")
a := selfCheckAgent(capi.URL, []string{config.SelfCheckControlAPI, config.SelfCheckIPEcho}, echo.URL)
ip, method, detail, ok := a.runSelfCheckMethods(context.Background(), "1.2.3.4")
if ok || method != "" {
t.Fatalf("expected a failure, got ok=%v method=%q", ok, method)
}
if ip != "6.6.6.6" {
t.Fatalf("detected ip = %q, want the last reported address 6.6.6.6", ip)
}
for _, want := range []string{
`control_api: egress ip "10.0.0.5" does not match assigned fip "1.2.3.4"`,
"private address", // the hint for the internal-network case
`ip_echo: egress ip "6.6.6.6" does not match`,
} {
if !strings.Contains(detail, want) {
t.Fatalf("detail %q does not contain %q", detail, want)
}
}
}
func TestSelfCheckPrivateHintOnlyForControlAPI(t *testing.T) {
echo := echoServer(t, "10.1.1.1")
a := selfCheckAgent("http://unused", []string{config.SelfCheckIPEcho}, echo.URL)
_, _, detail, ok := a.runSelfCheckMethods(context.Background(), "1.2.3.4")
if ok || strings.Contains(detail, "private address") {
t.Fatalf("ok=%v detail=%q: the control_api hint must not appear for ip_echo", ok, detail)
}
}
// With nothing configured (a config built without the loader) the agent
// behaves as before: ip_echo only, control-api is never asked.
func TestSelfCheckDefaultsToIPEcho(t *testing.T) {
var capiCalls int32
capi := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
atomic.AddInt32(&capiCalls, 1)
}))
defer capi.Close()
echo := echoServer(t, "1.2.3.4")
a := selfCheckAgent(capi.URL, nil, echo.URL)
_, method, _, ok := a.runSelfCheckMethods(context.Background(), "1.2.3.4")
if !ok || method != config.SelfCheckIPEcho || atomic.LoadInt32(&capiCalls) != 0 {
t.Fatalf("method=%q ok=%v control-api calls=%d", method, ok, atomic.LoadInt32(&capiCalls))
}
}
// A hung control-api must not use up the time of the fallback: each method
// has its own timeout.
func TestSelfCheckHungControlAPIDoesNotStarveFallback(t *testing.T) {
release := make(chan struct{})
capi := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
<-release
}))
defer capi.Close()
defer close(release)
echo := echoServer(t, "1.2.3.4")
a := selfCheckAgent(capi.URL, []string{config.SelfCheckControlAPI, config.SelfCheckIPEcho}, echo.URL)
a.cfg.SelfCheck.TimeoutSeconds = 1
start := time.Now()
// The outer context mirrors handleSelfCheckAndRun: timeout x methods.
ctx, cancel := context.WithTimeout(context.Background(), 2*time.Second)
defer cancel()
_, method, detail, ok := a.runSelfCheckMethods(ctx, "1.2.3.4")
if !ok || method != config.SelfCheckIPEcho {
t.Fatalf("method=%q ok=%v detail=%q, want a pass via ip_echo after control_api timed out", method, ok, detail)
}
if elapsed := time.Since(start); elapsed > 1900*time.Millisecond {
t.Fatalf("took %s: the hung method consumed the fallback's time", elapsed)
}
}
// Every control-api request must use a new TCP connection: a connection
// opened before the floating IP was attached would report the old address.
func TestDetectViaControlAPIDialsNewConnectionEachTime(t *testing.T) {
var conns int32
ts := httptest.NewUnstartedServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
fmt.Fprint(w, `{"ip":"1.2.3.4","source":"remote_addr"}`)
}))
ts.Config.ConnState = func(_ net.Conn, s http.ConnState) {
if s == http.StateNew {
atomic.AddInt32(&conns, 1)
}
}
ts.Start()
defer ts.Close()
a := selfCheckAgent(ts.URL, nil)
for i := 0; i < 3; i++ {
if _, err := a.detectViaControlAPI(context.Background()); err != nil {
t.Fatalf("call %d: %v", i, err)
}
}
if n := atomic.LoadInt32(&conns); n != 3 {
t.Fatalf("3 calls opened %d connections, want 3 (no keep-alive reuse)", n)
}
}
func TestDetectViaControlAPIRejectsBadAnswers(t *testing.T) {
for name, body := range map[string]string{"not json": "oops", "not an ip": `{"ip":"abc"}`, "empty": `{}`} {
ts := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) { fmt.Fprint(w, body) }))
a := selfCheckAgent(ts.URL, nil)
if ip, err := a.detectViaControlAPI(context.Background()); err == nil {
t.Fatalf("%s: expected an error, got %q", name, ip)
}
ts.Close()
}
}
// The agent token must never be sent on this request (the route is open and
// the token is meant for control-api writes only).
func TestDetectViaControlAPISendsNoToken(t *testing.T) {
var auth string
ts := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
auth = r.Header.Get("Authorization")
fmt.Fprint(w, `{"ip":"1.2.3.4"}`)
}))
defer ts.Close()
a := selfCheckAgent(ts.URL, nil)
if _, err := a.detectViaControlAPI(context.Background()); err != nil {
t.Fatal(err)
}
if auth != "" {
t.Fatalf("Authorization header sent: %q", auth)
}
}
+24 -3
View File
@@ -233,17 +233,30 @@ type ValidatorAgent struct {
Checks AgentChecks `yaml:"checks"`
}
// Self-check methods accepted in SelfCheckCfg.Methods.
const (
SelfCheckIPEcho = "ip_echo"
SelfCheckControlAPI = "control_api"
)
// SelfCheckCfg configures how the agent confirms its egress actually flows
// through the newly assigned floating IP. This must query a resource
// genuinely outside the cloud project: OpenStack only applies floating-IP
// SNAT to traffic leaving via the external/provider network, so any
// in-project resource (including control-api, if it's reachable over the
// project's internal network) would see the validator's private address
// instead — a false negative that never changes. IPEchoURLs are tried in
// order (falling through to the next on error/timeout, not on a genuine
// mismatch) until one returns a parseable IP.
// instead — a false negative that never changes. The control_api method is
// therefore only valid when control-api is reached over the external network
// (hosted outside the cloud); then it sees the floating IP as the source.
//
// Methods are tried in priority order and the self-check passes as soon as
// any one confirms the address; the next method is tried both when one gives
// no answer and when it reports a different address. IPEchoURLs are tried in
// order within ip_echo (falling through to the next on error/timeout, not on
// a genuine mismatch) until one returns a parseable IP.
type SelfCheckCfg struct {
TimeoutSeconds int `yaml:"timeout_seconds"`
Methods []string `yaml:"methods"`
IPEchoURLs []string `yaml:"ip_echo_urls"`
}
@@ -273,6 +286,14 @@ func LoadValidatorAgent(path string) (*ValidatorAgent, error) {
if c.SelfCheck.TimeoutSeconds == 0 {
c.SelfCheck.TimeoutSeconds = 10
}
if len(c.SelfCheck.Methods) == 0 {
c.SelfCheck.Methods = []string{SelfCheckIPEcho}
}
for _, m := range c.SelfCheck.Methods {
if m != SelfCheckIPEcho && m != SelfCheckControlAPI {
return nil, fmt.Errorf("self_check.methods: unknown method %q (allowed: %s, %s)", m, SelfCheckIPEcho, SelfCheckControlAPI)
}
}
if len(c.SelfCheck.IPEchoURLs) == 0 {
c.SelfCheck.IPEchoURLs = []string{"https://api.ipify.org", "https://ifconfig.me/ip"}
}
+55
View File
@@ -64,3 +64,58 @@ func TestControlAPIExampleConfigs(t *testing.T) {
t.Fatalf("load docker example: %v", err)
}
}
func writeAgentConfig(t *testing.T, selfCheck string) string {
t.Helper()
path := filepath.Join(t.TempDir(), "agent.yaml")
body := "validator_id: v1\ncontrol_api_url: http://x\nself_check:\n" + selfCheck
if err := os.WriteFile(path, []byte(body), 0o600); err != nil {
t.Fatal(err)
}
return path
}
func TestLoadValidatorAgentSelfCheckMethods(t *testing.T) {
// No methods key: the historical behaviour, ip_echo only.
c, err := LoadValidatorAgent(writeAgentConfig(t, " timeout_seconds: 5\n"))
if err != nil {
t.Fatalf("load: %v", err)
}
if len(c.SelfCheck.Methods) != 1 || c.SelfCheck.Methods[0] != SelfCheckIPEcho {
t.Fatalf("default methods = %v, want [ip_echo]", c.SelfCheck.Methods)
}
// Order is the priority and must be kept.
c, err = LoadValidatorAgent(writeAgentConfig(t, " methods: [control_api, ip_echo]\n"))
if err != nil {
t.Fatalf("load: %v", err)
}
if got := c.SelfCheck.Methods; len(got) != 2 || got[0] != SelfCheckControlAPI || got[1] != SelfCheckIPEcho {
t.Fatalf("methods = %v, want [control_api ip_echo]", got)
}
// An unknown method is a configuration error, not a silent no-op.
if _, err := LoadValidatorAgent(writeAgentConfig(t, " methods: [control_api, ipecho]\n")); err == nil {
t.Fatal("expected an error for an unknown self-check method")
}
}
// The shipped agent example must load, and the rxprod copy must stay
// byte-identical to it.
func TestValidatorAgentExampleConfig(t *testing.T) {
c, err := LoadValidatorAgent("../../configs/validator-agent.example.yaml")
if err != nil {
t.Fatalf("load example: %v", err)
}
if got := c.SelfCheck.Methods; len(got) != 2 || got[0] != SelfCheckControlAPI {
t.Fatalf("example methods = %v", got)
}
a, _ := os.ReadFile("../../configs/validator-agent.example.yaml")
b, err := os.ReadFile("../../rxprod-compose/sources/validator-agent.example.yaml")
if err != nil {
t.Fatal(err)
}
if !bytes.Equal(a, b) {
t.Fatal("rxprod-compose/sources/validator-agent.example.yaml differs from configs/validator-agent.example.yaml")
}
}
+77 -32
View File
@@ -101,7 +101,7 @@ func (d *DB) ClaimNextQueued(ctx context.Context, validatorID string, leaseTTL t
res, err = tx.ExecContext(ctx, `
UPDATE validators SET state=?, current_ip_id=?, updated_at=?
WHERE validator_id=? AND state=?
WHERE validator_id=? AND state=? AND current_ip_id IS NULL
`, ValidatorAssigned, item.ID, timeToDB(now), validatorID, ValidatorIdle)
if err != nil {
return nil, err
@@ -122,13 +122,59 @@ func (d *DB) ClaimNextQueued(ctx context.Context, validatorID string, leaseTTL t
return &item, nil
}
// SetFIPAssociated records that the floating IP is now attached. It only
// applies to a row still in assigning_fip: if the address was deleted or
// cancelled while the cloud call was in flight, it returns ErrInvalidState
// and the caller must detach the floating IP again.
func (d *DB) SetFIPAssociated(ctx context.Context, ipID int64, fipID string, leaseTTL time.Duration) error {
now := Now()
_, err := d.ExecContext(ctx, `
res, err := d.ExecContext(ctx, `
UPDATE ip_queue SET state=?, fip_id=?, fip_associated_at=?, lease_expires_at=?, updated_at=?
WHERE id=?
`, IPAwaitingSelfCheck, fipID, timeToDB(now), timeToDB(now.Add(leaseTTL)), timeToDB(now), ipID)
return err
WHERE id=? AND state=?
`, IPAwaitingSelfCheck, fipID, timeToDB(now), timeToDB(now.Add(leaseTTL)), timeToDB(now), ipID, IPAssigningFIP)
if err != nil {
return err
}
if n, _ := res.RowsAffected(); n == 0 {
return fmt.Errorf("ip_id %d is no longer assigning_fip: %w", ipID, ErrInvalidState)
}
return nil
}
// KnownAddresses returns the subset of addresses present in ip_registry,
// i.e. addresses this system has ever queued.
func (d *DB) KnownAddresses(ctx context.Context, addresses []string) (map[string]bool, error) {
const chunk = 500
known := make(map[string]bool, len(addresses))
for start := 0; start < len(addresses); start += chunk {
end := start + chunk
if end > len(addresses) {
end = len(addresses)
}
part := addresses[start:end]
args := make([]any, len(part))
for i, a := range part {
args[i] = a
}
rows, err := d.QueryContext(ctx, `SELECT ip_address FROM ip_registry WHERE ip_address IN (`+placeholders(len(part))+`)`, args...)
if err != nil {
return nil, err
}
for rows.Next() {
var a string
if err := rows.Scan(&a); err != nil {
rows.Close()
return nil, err
}
known[a] = true
}
if err := rows.Err(); err != nil {
rows.Close()
return nil, err
}
rows.Close()
}
return known, nil
}
func (d *DB) SetChecking(ctx context.Context, ipID int64, leaseTTL time.Duration) error {
@@ -161,7 +207,8 @@ func (d *DB) FinishIP(ctx context.Context, ipID int64, result string) error {
}
// ReleaseFIP records that the floating IP has been disassociated and frees
// the owning validator back to idle, in one transaction.
// the owning validator (if this address is still its current one, see
// freeValidatorSQL), in one transaction.
func (d *DB) ReleaseFIP(ctx context.Context, ipID int64, validatorID string) error {
tx, err := d.BeginTx(ctx, nil)
if err != nil {
@@ -173,10 +220,7 @@ func (d *DB) ReleaseFIP(ctx context.Context, ipID int64, validatorID string) err
if _, err := tx.ExecContext(ctx, `UPDATE ip_queue SET fip_released_at=?, updated_at=? WHERE id=?`, now, now, ipID); err != nil {
return err
}
if _, err := tx.ExecContext(ctx, `
UPDATE validators SET state=?, current_ip_id=NULL, updated_at=?
WHERE validator_id=?
`, ValidatorIdle, now, validatorID); err != nil {
if _, err := tx.ExecContext(ctx, freeValidatorSQL, now, validatorID, ipID); err != nil {
return err
}
return tx.Commit()
@@ -206,10 +250,7 @@ func (d *DB) MarkFIPOccupied(ctx context.Context, ipID int64, validatorID string
return err
}
if validatorID != "" {
if _, err := tx.ExecContext(ctx, `
UPDATE validators SET state=?, current_ip_id=NULL, updated_at=?
WHERE validator_id=?
`, ValidatorIdle, now, validatorID); err != nil {
if _, err := tx.ExecContext(ctx, freeValidatorSQL, now, validatorID, ipID); err != nil {
return err
}
}
@@ -269,10 +310,7 @@ func (d *DB) RequeueOrFail(ctx context.Context, ipID int64, validatorID string,
}
if validatorID != "" {
if _, err := tx.ExecContext(ctx, `
UPDATE validators SET state=?, current_ip_id=NULL, updated_at=?
WHERE validator_id=?
`, ValidatorIdle, now, validatorID); err != nil {
if _, err := tx.ExecContext(ctx, freeValidatorSQL, now, validatorID, ipID); err != nil {
return err
}
}
@@ -474,9 +512,11 @@ func (d *DB) DeleteIPs(ctx context.Context, addresses []string) (DeleteIPsResult
func deleteIPTx(ctx context.Context, tx *sql.Tx, ipID int64) error {
now := timeToDB(Now())
if _, err := tx.ExecContext(ctx, `
UPDATE validators SET state=?, current_ip_id=NULL, updated_at=?
UPDATE validators SET current_ip_id=NULL,
state = CASE WHEN state=? THEN state ELSE ? END,
updated_at=?
WHERE current_ip_id=?
`, ValidatorIdle, now, ipID); err != nil {
`, ValidatorUnreachable, ValidatorIdle, now, ipID); err != nil {
return fmt.Errorf("free owning validator: %w", err)
}
if _, err := tx.ExecContext(ctx, `UPDATE checks SET ip_id=NULL WHERE ip_id=?`, ipID); err != nil {
@@ -533,9 +573,11 @@ func (d *DB) ClearAllIPs(ctx context.Context) ([]string, error) {
now := timeToDB(Now())
if _, err := tx.ExecContext(ctx, `
UPDATE validators SET state=?, current_ip_id=NULL, updated_at=?
UPDATE validators SET current_ip_id=NULL,
state = CASE WHEN state=? THEN state ELSE ? END,
updated_at=?
WHERE current_ip_id IS NOT NULL
`, ValidatorIdle, now); err != nil {
`, ValidatorUnreachable, ValidatorIdle, now); err != nil {
return nil, fmt.Errorf("free owning validators: %w", err)
}
if _, err := tx.ExecContext(ctx, `UPDATE checks SET ip_id=NULL WHERE ip_id IS NOT NULL`); err != nil {
@@ -564,11 +606,15 @@ type FIPRef struct {
FIPID string
}
// ListFIPRefs returns every queue row with an attached floating IP (fip_id
// set) — typically at most one per validator — so a bulk clear can
// disassociate them without loading the whole queue.
// ListFIPRefs returns the queue rows that may still hold a floating IP: an
// fip_id is set and the row is not finished. A finished row (done, failed,
// occupied) keeps its fip_id for display, but its floating IP was already
// disassociated before the final state was written, so listing it would only
// make a bulk clear issue thousands of pointless cloud calls. At most one row
// per validator qualifies, so a clear does not need to load the whole queue.
func (d *DB) ListFIPRefs(ctx context.Context) ([]FIPRef, error) {
rows, err := d.QueryContext(ctx, `SELECT id, ip_address, fip_id FROM ip_queue WHERE fip_id<>'' ORDER BY id`)
rows, err := d.QueryContext(ctx, `SELECT id, ip_address, fip_id FROM ip_queue WHERE fip_id<>'' AND state NOT IN (?, ?, ?) ORDER BY id`,
IPDone, IPFailed, IPOccupied)
if err != nil {
return nil, err
}
@@ -585,8 +631,7 @@ func (d *DB) ListFIPRefs(ctx context.Context) ([]FIPRef, error) {
}
// ListFIPRefsByAddresses is ListFIPRefs restricted to the given addresses
// (unknown addresses and rows without an attached floating IP are simply
// absent), using a handful of IN (...) queries instead of one lookup per
// (unknown addresses and rows that hold no floating IP are simply absent), using a handful of IN (...) queries instead of one lookup per
// address.
func (d *DB) ListFIPRefsByAddresses(ctx context.Context, addresses []string) ([]FIPRef, error) {
const chunk = 500
@@ -597,12 +642,12 @@ func (d *DB) ListFIPRefsByAddresses(ctx context.Context, addresses []string) ([]
end = len(addresses)
}
part := addresses[start:end]
args := make([]any, len(part))
for i, a := range part {
args[i] = a
args := []any{IPDone, IPFailed, IPOccupied}
for _, a := range part {
args = append(args, a)
}
rows, err := d.QueryContext(ctx,
`SELECT id, ip_address, fip_id FROM ip_queue WHERE fip_id<>'' AND ip_address IN (`+placeholders(len(part))+`)`, args...)
`SELECT id, ip_address, fip_id FROM ip_queue WHERE fip_id<>'' AND state NOT IN (?, ?, ?) AND ip_address IN (`+placeholders(len(part))+`)`, args...)
if err != nil {
return nil, err
}
+186
View File
@@ -0,0 +1,186 @@
package db
import (
"context"
"fmt"
"testing"
"time"
)
func validatorState(t *testing.T, d *DB, id string) *Validator {
t.Helper()
v, err := d.GetValidator(testCtx(t), id)
if err != nil {
t.Fatalf("get validator %s: %v", id, err)
}
return v
}
func testCtx(t *testing.T) context.Context {
t.Helper()
return context.Background()
}
// claimFor2 seeds an address and returns its id without claiming it.
func claimFor2(t *testing.T, d *DB, addr string) int64 {
t.Helper()
ctx := testCtx(t)
if err := d.SeedQueue(ctx, []string{addr}); err != nil {
t.Fatalf("seed %s: %v", addr, err)
}
ip, err := d.GetIPByAddress(ctx, addr)
if err != nil {
t.Fatalf("get %s: %v", addr, err)
}
return ip.ID
}
// claimFor seeds an address and claims it for the validator.
func claimFor(t *testing.T, d *DB, addr, validatorID string) *IPQueueItem {
t.Helper()
ctx := testCtx(t)
if err := d.SeedQueue(ctx, []string{addr}); err != nil {
t.Fatalf("seed %s: %v", addr, err)
}
item, err := d.ClaimNextQueued(ctx, validatorID, time.Minute)
if err != nil || item == nil {
t.Fatalf("claim %s for %s: item=%v err=%v", addr, validatorID, item, err)
}
return item
}
func TestHeartbeatKeepsAssignedWhenValidatorStillHoldsAnAddress(t *testing.T) {
d, ctx := newTestDB(t)
_ = d.AdminCreateValidator(ctx, "v1", "p1")
_ = d.AdminCreateValidator(ctx, "v2", "p2")
item := claimFor(t, d, "1.1.1.1", "v1")
_ = d.MarkValidatorUnreachable(ctx, "v1")
_ = d.MarkValidatorUnreachable(ctx, "v2")
if err := d.Heartbeat(ctx, "v1"); err != nil {
t.Fatal(err)
}
if v := validatorState(t, d, "v1"); v.State != ValidatorAssigned || v.CurrentIPID == nil || *v.CurrentIPID != item.ID {
t.Fatalf("v1 after heartbeat: state=%s current_ip=%v, want assigned to %d", v.State, v.CurrentIPID, item.ID)
}
if err := d.Heartbeat(ctx, "v2"); err != nil {
t.Fatal(err)
}
if v := validatorState(t, d, "v2"); v.State != ValidatorIdle {
t.Fatalf("v2 after heartbeat: state=%s, want idle", v.State)
}
}
func TestRegisterValidatorReactivationKeepsAssignedAddress(t *testing.T) {
d, ctx := newTestDB(t)
_ = d.AdminCreateValidator(ctx, "v1", "p1")
item := claimFor(t, d, "1.1.1.1", "v1")
_ = d.MarkValidatorUnreachable(ctx, "v1")
if err := d.RegisterValidator(ctx, "v1", "host", "p1", "v"); err != nil {
t.Fatal(err)
}
if v := validatorState(t, d, "v1"); v.State != ValidatorAssigned || v.CurrentIPID == nil || *v.CurrentIPID != item.ID {
t.Fatalf("after re-register: state=%s current_ip=%v, want assigned to %d", v.State, v.CurrentIPID, item.ID)
}
}
func TestStaleReleasesLeaveTheCurrentAddressAlone(t *testing.T) {
cases := map[string]func(d *DB, ipID int64) error{
"ReleaseFIP": func(d *DB, id int64) error { return d.ReleaseFIP(context.Background(), id, "v1") },
"RequeueOrFail": func(d *DB, id int64) error { return d.RequeueOrFail(context.Background(), id, "v1", 3) },
"MarkFIPOccupied": func(d *DB, id int64) error { return d.MarkFIPOccupied(context.Background(), id, "v1") },
"FreeValidator": func(d *DB, id int64) error { return d.FreeValidator(context.Background(), "v1", id) },
}
for name, release := range cases {
t.Run(name, func(t *testing.T) {
d, ctx := newTestDB(t)
_ = d.AdminCreateValidator(ctx, "v1", "p1")
stale := claimFor(t, d, "1.1.1.1", "v1")
// v1 has moved on to another address (set directly: the claim path
// itself refuses a validator that is still busy).
cur := claimFor2(t, d, "2.2.2.2")
if _, err := d.ExecContext(ctx, `UPDATE ip_queue SET state='checking', owner_validator_id='v1' WHERE id=?`, cur); err != nil {
t.Fatal(err)
}
if _, err := d.ExecContext(ctx, `UPDATE validators SET state='assigned', current_ip_id=? WHERE validator_id='v1'`, cur); err != nil {
t.Fatal(err)
}
if err := release(d, stale.ID); err != nil {
t.Fatalf("%s: %v", name, err)
}
v := validatorState(t, d, "v1")
if v.State != ValidatorAssigned || v.CurrentIPID == nil || *v.CurrentIPID != cur {
t.Fatalf("%s freed a validator that holds another address: state=%s current_ip=%v", name, v.State, v.CurrentIPID)
}
})
}
}
// The right address frees the validator, but an unreachable one stays so.
func TestReleaseKeepsUnreachableValidatorUnreachable(t *testing.T) {
d, ctx := newTestDB(t)
_ = d.AdminCreateValidator(ctx, "v1", "p1")
_ = d.AdminCreateValidator(ctx, "v2", "p2")
a := claimFor(t, d, "1.1.1.1", "v1")
b := claimFor(t, d, "2.2.2.2", "v2")
_ = d.MarkValidatorUnreachable(ctx, "v1")
if err := d.ReleaseFIP(ctx, a.ID, "v1"); err != nil {
t.Fatal(err)
}
if v := validatorState(t, d, "v1"); v.State != ValidatorUnreachable || v.CurrentIPID != nil {
t.Fatalf("v1: state=%s current_ip=%v, want unreachable and empty", v.State, v.CurrentIPID)
}
if err := d.RequeueOrFail(ctx, b.ID, "v2", 3); err != nil {
t.Fatal(err)
}
if v := validatorState(t, d, "v2"); v.State != ValidatorIdle || v.CurrentIPID != nil {
t.Fatalf("v2: state=%s current_ip=%v, want idle and empty", v.State, v.CurrentIPID)
}
}
func TestClaimRefusesValidatorThatStillHoldsAnAddress(t *testing.T) {
d, ctx := newTestDB(t)
_ = d.AdminCreateValidator(ctx, "v1", "p1")
first := claimFor(t, d, "1.1.1.1", "v1")
// An old version could leave a validator idle while it still pointed at an address.
if _, err := d.ExecContext(ctx, `UPDATE validators SET state='idle' WHERE validator_id='v1'`); err != nil {
t.Fatal(err)
}
_ = d.SeedQueue(ctx, []string{"2.2.2.2"})
got, err := d.ClaimNextQueued(ctx, "v1", time.Minute)
if err != nil || got != nil {
t.Fatalf("claim for a validator that holds %d: item=%v err=%v, want nothing", first.ID, got, err)
}
if ip, _ := d.GetIPByAddress(ctx, "2.2.2.2"); ip.State != IPQueued {
t.Fatalf("2.2.2.2 is %s, want queued", ip.State)
}
}
func TestListFIPRefsSkipsFinishedAddresses(t *testing.T) {
d, ctx := newTestDB(t)
var addrs []string
for i := 0; i < 6; i++ {
addrs = append(addrs, fmt.Sprintf("10.0.0.%d", i+1))
}
_ = d.SeedQueue(ctx, addrs)
states := []string{"done", "failed", "occupied", "awaiting_self_check", "checking", "aggregating"}
for i, a := range addrs {
if _, err := d.ExecContext(ctx, `UPDATE ip_queue SET state=?, fip_id=? WHERE ip_address=?`, states[i], fmt.Sprintf("fip-%d", i), a); err != nil {
t.Fatal(err)
}
}
refs, err := d.ListFIPRefs(ctx)
if err != nil {
t.Fatal(err)
}
if len(refs) != 3 {
t.Fatalf("ListFIPRefs returned %d rows, want the 3 unfinished ones: %+v", len(refs), refs)
}
byAddr, err := d.ListFIPRefsByAddresses(ctx, addrs)
if err != nil || len(byAddr) != 3 {
t.Fatalf("ListFIPRefsByAddresses returned %d rows (err %v), want 3", len(byAddr), err)
}
}
+63 -12
View File
@@ -27,24 +27,32 @@ func (d *DB) RegisterValidator(ctx context.Context, validatorID, hostname, osPor
}
// A brand-new row already lands in ValidatorIdle via the INSERT branch;
// a re-registering validator that was 'unregistered' or 'unreachable'
// (but not mid-assignment) should also come back to idle.
// comes back: to idle when it holds no address, to assigned when it still
// does (its current_ip_id is kept, so it must not be handed another).
_, err = d.ExecContext(ctx, `
UPDATE validators SET state=?, updated_at=?
UPDATE validators SET
state = CASE WHEN current_ip_id IS NULL THEN ? ELSE ? END,
updated_at=?
WHERE validator_id=? AND state IN (?, ?)
`, ValidatorIdle, now, validatorID, ValidatorUnregistered, ValidatorUnreachable)
`, ValidatorIdle, ValidatorAssigned, now, validatorID, ValidatorUnregistered, ValidatorUnreachable)
if err != nil {
return fmt.Errorf("register validator (reactivate): %w", err)
}
return nil
}
// Heartbeat records a sign of life. A validator that was marked unreachable
// returns to idle if it holds no address, but to assigned if it still does:
// it was only silent (for example busy with slow checks), and handing it a
// second address while it works on the first would leave the second one
// without an owner that can ever pick it up.
func (d *DB) Heartbeat(ctx context.Context, validatorID string) error {
now := timeToDB(Now())
res, err := d.ExecContext(ctx, `
UPDATE validators SET last_heartbeat_at=?, updated_at=?,
state = CASE WHEN state=? THEN ? ELSE state END
state = CASE WHEN state=? THEN (CASE WHEN current_ip_id IS NULL THEN ? ELSE ? END) ELSE state END
WHERE validator_id=?
`, now, now, ValidatorUnreachable, ValidatorIdle, validatorID)
`, now, now, ValidatorUnreachable, ValidatorIdle, ValidatorAssigned, validatorID)
if err != nil {
return fmt.Errorf("heartbeat: %w", err)
}
@@ -138,16 +146,59 @@ func (d *DB) MarkValidatorUnreachable(ctx context.Context, validatorID string) e
return err
}
// FreeValidator returns a validator to idle with no assigned IP. Used after
// an IP finishes (success or failure) or is reclaimed by the lease sweep.
func (d *DB) FreeValidator(ctx context.Context, validatorID string) error {
_, err := d.ExecContext(ctx, `
UPDATE validators SET state=?, current_ip_id=NULL, updated_at=?
WHERE validator_id=?
`, ValidatorIdle, timeToDB(Now()), validatorID)
// freeValidatorSQL releases a validator from the address it holds. It only
// applies when the validator's current address is the one being released
// (args: now, validator id, ip id): a late release of an old address must not
// free a validator that has already moved on to another one. An unreachable
// validator stays unreachable until its next heartbeat, so a dead validator
// is not handed new addresses just because its lease was reclaimed.
var freeValidatorSQL = fmt.Sprintf(`
UPDATE validators SET current_ip_id=NULL,
state = CASE WHEN state='%s' THEN state ELSE '%s' END,
updated_at=?
WHERE validator_id=? AND current_ip_id=?`, ValidatorUnreachable, ValidatorIdle)
// FreeValidator releases a validator from the given address (see
// freeValidatorSQL). Used after an IP finishes (success or failure) or is
// reclaimed by the lease sweep.
func (d *DB) FreeValidator(ctx context.Context, validatorID string, ipID int64) error {
_, err := d.ExecContext(ctx, freeValidatorSQL, timeToDB(Now()), validatorID, ipID)
return err
}
// ReconcileValidators repairs validators whose state disagrees with the
// queue: a validator pointing at an address that no longer exists, is
// finished, or belongs to another validator is released; an "assigned"
// validator that holds nothing goes back to idle. The invariants normally
// hold by construction (every change is one transaction); this heals what a
// crash or an older version left behind. It returns the number of validators
// repaired.
func (d *DB) ReconcileValidators(ctx context.Context) (int64, error) {
now := timeToDB(Now())
res, err := d.ExecContext(ctx, fmt.Sprintf(`
UPDATE validators SET current_ip_id=NULL,
state = CASE WHEN state='%s' THEN state ELSE '%s' END,
updated_at=?
WHERE current_ip_id IS NOT NULL AND NOT EXISTS (
SELECT 1 FROM ip_queue q
WHERE q.id = validators.current_ip_id
AND q.owner_validator_id = validators.validator_id
AND q.state NOT IN ('%s','%s','%s'))`,
ValidatorUnreachable, ValidatorIdle, IPDone, IPFailed, IPOccupied), now)
if err != nil {
return 0, fmt.Errorf("reconcile validators: %w", err)
}
n, _ := res.RowsAffected()
res, err = d.ExecContext(ctx, `
UPDATE validators SET state=?, updated_at=?
WHERE state=? AND current_ip_id IS NULL`, ValidatorIdle, now, ValidatorAssigned)
if err != nil {
return n, fmt.Errorf("reconcile validators: %w", err)
}
m, _ := res.RowsAffected()
return n + m, nil
}
// AdminCreateValidator registers a brand-new validator via the admin API.
// Unlike RegisterValidator (used by the agent's self-registration call),
// this refuses to upsert over an existing row.
+2 -2
View File
@@ -97,8 +97,8 @@ func TestRouteTableIsClassified(t *testing.T) {
t.Fatalf("admin route %q is %s, want admin", rt.Pattern, rt.Access)
}
}
if counts["admin"] != 34 || counts["agent"] != 5 || counts["open"] != 7 {
t.Fatalf("access counts = %v, want admin=34 agent=5 open=7", counts)
if counts["admin"] != 34 || counts["agent"] != 5 || counts["open"] != 8 {
t.Fatalf("access counts = %v, want admin=34 agent=5 open=8", counts)
}
}
+5
View File
@@ -23,6 +23,11 @@ type okResponse struct {
OK bool `json:"ok"`
}
type observedIPResponse struct {
IP string `json:"ip"`
Source string `json:"source"`
}
type checkConfigDTO struct {
Type string `json:"type"`
Targets []string `json:"targets"`
+27 -4
View File
@@ -1,17 +1,32 @@
package httpapi
import (
"context"
"errors"
"fmt"
"net/http"
"net/url"
"strconv"
"strings"
"time"
"cloudipvalidator/internal/db"
"cloudipvalidator/internal/orchestrator"
)
// destructiveOpTimeout bounds a cancel, delete or clear. Such an operation
// talks to the cloud as well as the database, and abandoning it halfway
// leaves floating IPs detached from rows that still exist (or the reverse),
// so it must not die with the client connection: a client that gives up
// (a dashboard or curl timeout) only stops waiting for the answer.
const destructiveOpTimeout = 10 * time.Minute
// detachedContext returns a context that ignores cancellation of the request
// but keeps its values, with destructiveOpTimeout as the upper bound.
func detachedContext(r *http.Request) (context.Context, context.CancelFunc) {
return context.WithTimeout(context.WithoutCancel(r.Context()), destructiveOpTimeout)
}
func (s *Server) handleHealthz(w http.ResponseWriter, r *http.Request) {
writeJSON(w, http.StatusOK, okResponse{OK: true})
}
@@ -273,7 +288,9 @@ func toScanStatusDTO(st orchestrator.ScanStatus) scanStatusDTO {
// need to be disassociated in OpenStack.
func (s *Server) handleAdminCancelIP(w http.ResponseWriter, r *http.Request) {
address := r.PathValue("ip")
if err := s.Orch.ForceCancel(r.Context(), address); err != nil {
ctx, cancel := detachedContext(r)
defer cancel()
if err := s.Orch.ForceCancel(ctx, address); err != nil {
writeDBError(w, err)
return
}
@@ -286,7 +303,9 @@ func (s *Server) handleAdminCancelIP(w http.ResponseWriter, r *http.Request) {
// associated floating IP needs disassociating first.
func (s *Server) handleAdminDeleteIP(w http.ResponseWriter, r *http.Request) {
address := r.PathValue("ip")
if err := s.Orch.DeleteIP(r.Context(), address); err != nil {
ctx, cancel := detachedContext(r)
defer cancel()
if err := s.Orch.DeleteIP(ctx, address); err != nil {
writeDBError(w, err)
return
}
@@ -306,7 +325,9 @@ func (s *Server) handleAdminDeleteIPs(w http.ResponseWriter, r *http.Request) {
writeError(w, http.StatusBadRequest, "addresses must not be empty")
return
}
result, err := s.Orch.DeleteIPs(r.Context(), req.Addresses)
ctx, cancel := detachedContext(r)
defer cancel()
result, err := s.Orch.DeleteIPs(ctx, req.Addresses)
if err != nil {
writeDBError(w, err)
return
@@ -320,7 +341,9 @@ func (s *Server) handleAdminDeleteIPs(w http.ResponseWriter, r *http.Request) {
// handleAdminClearQueue permanently removes every address currently in the
// queue, including those actively being checked.
func (s *Server) handleAdminClearQueue(w http.ResponseWriter, r *http.Request) {
result, err := s.Orch.ClearQueue(r.Context())
ctx, cancel := detachedContext(r)
defer cancel()
result, err := s.Orch.ClearQueue(ctx)
if err != nil {
writeDBError(w, err)
return
+45
View File
@@ -1,7 +1,12 @@
package httpapi
import (
"database/sql"
"errors"
"fmt"
"net"
"net/http"
"net/netip"
"time"
"cloudipvalidator/internal/db"
@@ -64,6 +69,46 @@ func (s *Server) handleAgentAssignment(w http.ResponseWriter, r *http.Request) {
})
}
// handleAgentObservedIP tells a validator which source address control-api
// sees for its connection, so the agent's self-check can confirm the floating
// IP without a third-party IP-echo service. The route is open, so it is
// limited to known validators to keep it from becoming a public "what is my
// IP" service.
func (s *Server) handleAgentObservedIP(w http.ResponseWriter, r *http.Request) {
id := r.PathValue("id")
if _, err := s.DB.GetValidator(r.Context(), id); err != nil {
if errors.Is(err, sql.ErrNoRows) || errors.Is(err, db.ErrNotFound) {
writeError(w, http.StatusNotFound, "unknown validator: "+id)
return
}
writeError(w, http.StatusInternalServerError, err.Error())
return
}
ip, err := clientIP(r)
if err != nil {
writeError(w, http.StatusInternalServerError, err.Error())
return
}
writeJSON(w, http.StatusOK, observedIPResponse{IP: ip, Source: "remote_addr"})
}
// clientIP returns the peer address of the TCP connection in canonical form
// (IPv4-mapped IPv6 unmapped to plain IPv4). It deliberately ignores
// X-Forwarded-For / X-Real-IP: validators connect directly, and trusting a
// client-supplied header would let a validator forge the address and pass
// the self-check.
func clientIP(r *http.Request) (string, error) {
host, _, err := net.SplitHostPort(r.RemoteAddr)
if err != nil {
return "", fmt.Errorf("parse remote address %q: %w", r.RemoteAddr, err)
}
addr, err := netip.ParseAddr(host)
if err != nil {
return "", fmt.Errorf("parse remote address %q: %w", r.RemoteAddr, err)
}
return addr.Unmap().String(), nil
}
func (s *Server) handleAgentSelfCheck(w http.ResponseWriter, r *http.Request) {
id := r.PathValue("id")
var req selfCheckRequest
@@ -0,0 +1,65 @@
package httpapi
import (
"context"
"net/http"
"testing"
"time"
"cloudipvalidator/internal/openstack"
)
// slowDetachOS delays every Disassociate so the client can give up first.
type slowDetachOS struct {
*openstack.MockClient
delay time.Duration
}
func (s slowDetachOS) DisassociateFloatingIP(ctx context.Context, fipID string) error {
time.Sleep(s.delay)
return s.MockClient.DisassociateFloatingIP(ctx, fipID)
}
// A client that gives up (a dashboard or curl timeout) must not abort
// "clear queue" halfway: on 2026-10-02 the request context was cancelled after
// part of the floating IPs were detached, the database was left untouched and
// the checks kept running.
func TestClearQueueSurvivesClientDisconnect(t *testing.T) {
fc, d, orch, mock := newConfigTestHarness(t)
ctx := context.Background()
mock.Seed("fip-1", "1.2.3.4", "svc")
if err := d.RegisterValidator(ctx, "validator-1", "host", "port-1", "v"); err != nil {
t.Fatal(err)
}
if err := d.SeedQueue(ctx, []string{"1.2.3.4", "5.6.7.8"}); err != nil {
t.Fatal(err)
}
orch.Tick(ctx) // validator-1 takes 1.2.3.4 and attaches its floating IP
if fip, _ := mock.GetFloatingIPByAddress(ctx, "1.2.3.4"); fip.PortID != "port-1" {
t.Fatalf("setup: floating ip not attached (port %q)", fip.PortID)
}
orch.OS = slowDetachOS{MockClient: mock, delay: 400 * time.Millisecond}
reqCtx, cancel := context.WithTimeout(ctx, 100*time.Millisecond)
defer cancel()
// No body: the server only notices that the client has left (and cancels
// the request context) when it is not waiting for a request body.
req, _ := http.NewRequestWithContext(reqCtx, http.MethodPost, fc.base+"/api/v1/admin/ips/clear", nil)
if resp, err := fc.client.Do(req); err == nil {
resp.Body.Close()
t.Fatalf("the client was expected to time out, got status %d", resp.StatusCode)
}
deadline := time.Now().Add(5 * time.Second)
for {
ips, _ := d.ListIPs(ctx)
fip, _ := mock.GetFloatingIPByAddress(ctx, "1.2.3.4")
if len(ips) == 0 && fip.PortID == "" {
return
}
if time.Now().After(deadline) {
t.Fatalf("clear did not finish after the client left: %d rows, floating ip on port %q", len(ips), fip.PortID)
}
time.Sleep(50 * time.Millisecond)
}
}
@@ -0,0 +1,79 @@
package httpapi
import (
"context"
"encoding/json"
"net/http"
"net/http/httptest"
"testing"
)
func TestClientIP(t *testing.T) {
cases := []struct {
name string
remoteAddr string
headers map[string]string
want string
wantErr bool
}{
{name: "ipv4", remoteAddr: "90.156.213.5:51234", want: "90.156.213.5"},
{name: "ipv4 mapped in ipv6", remoteAddr: "[::ffff:90.156.213.5]:51234", want: "90.156.213.5"},
{name: "ipv6", remoteAddr: "[2001:db8::7]:443", want: "2001:db8::7"},
// A client-supplied header must never change the answer.
{name: "forwarded for is ignored", remoteAddr: "90.156.213.5:1",
headers: map[string]string{"X-Forwarded-For": "1.2.3.4", "X-Real-IP": "5.6.7.8"}, want: "90.156.213.5"},
{name: "no port", remoteAddr: "90.156.213.5", wantErr: true},
{name: "not an address", remoteAddr: "host:80", wantErr: true},
}
for _, tc := range cases {
t.Run(tc.name, func(t *testing.T) {
r := httptest.NewRequest(http.MethodGet, "/", nil)
r.RemoteAddr = tc.remoteAddr
for k, v := range tc.headers {
r.Header.Set(k, v)
}
got, err := clientIP(r)
if tc.wantErr {
if err == nil {
t.Fatalf("expected an error, got %q", got)
}
return
}
if err != nil || got != tc.want {
t.Fatalf("clientIP = %q, %v; want %q", got, err, tc.want)
}
})
}
}
// The route is open (no token), answers a known validator with the address
// control-api sees, and refuses unknown validators and a forged header.
func TestObservedIPEndpoint(t *testing.T) {
fc, d, _, _ := newConfigTestHarness(t)
if err := d.RegisterValidator(context.Background(), "validator-1", "host-1", "port-1", "v0.1"); err != nil {
t.Fatal(err)
}
req, _ := http.NewRequest(http.MethodGet, fc.base+"/api/v1/agents/validator-1/observed-ip", nil)
req.Header.Set("X-Forwarded-For", "8.8.8.8")
resp, err := fc.client.Do(req)
if err != nil {
t.Fatal(err)
}
defer resp.Body.Close()
if resp.StatusCode != http.StatusOK {
t.Fatalf("status = %d, want 200", resp.StatusCode)
}
var got observedIPResponse
if err := json.NewDecoder(resp.Body).Decode(&got); err != nil {
t.Fatal(err)
}
// httptest connects from loopback; the forged header must not leak through.
if got.IP != "127.0.0.1" || got.Source != "remote_addr" {
t.Fatalf("response = %+v, want 127.0.0.1 / remote_addr", got)
}
if resp, _ := fc.do(http.MethodGet, "/api/v1/agents/no-such-validator/observed-ip", nil); resp.StatusCode != http.StatusNotFound {
t.Fatalf("unknown validator: status = %d, want 404", resp.StatusCode)
}
}
+1
View File
@@ -39,6 +39,7 @@ func (s *Server) routeTable() []route {
{"POST /api/v1/agents/register", s.handleAgentRegister, accessOpen},
{"POST /api/v1/agents/{id}/heartbeat", s.handleAgentHeartbeat, accessOpen},
{"GET /api/v1/agents/{id}/assignment", s.handleAgentAssignment, accessOpen},
{"GET /api/v1/agents/{id}/observed-ip", s.handleAgentObservedIP, accessOpen},
{"POST /api/v1/agents/{id}/self-check", s.handleAgentSelfCheck, accessAgent},
{"POST /api/v1/agents/{id}/events", s.handleAgentEvent, accessAgent},
{"POST /api/v1/agents/{id}/results", s.handleAgentResults, accessAgent},
+21
View File
@@ -239,6 +239,27 @@ func (c *Client) ListFreeFloatingIPs(ctx context.Context, pageSize int, onPage f
return paginate(ctx, pageSize, c.retry, c.fetchPage, onPage)
}
// ListFloatingIPsByPort asks Neutron for the floating IPs attached to portID.
func (c *Client) ListFloatingIPsByPort(ctx context.Context, portID string) ([]FloatingIP, error) {
pages, err := floatingips.List(c.networking, floatingips.ListOpts{PortID: portID}).AllPages(ctx)
if err != nil {
return nil, fmt.Errorf("openstack: list floating ips of port %s: %w", portID, err)
}
list, err := floatingips.ExtractFloatingIPs(pages)
if err != nil {
return nil, fmt.Errorf("openstack: extract floating ips: %w", err)
}
out := make([]FloatingIP, 0, len(list))
for _, f := range list {
proj := f.TenantID
if proj == "" {
proj = f.ProjectID
}
out = append(out, FloatingIP{ID: f.ID, Address: f.FloatingIP, PortID: f.PortID, ProjectID: proj})
}
return out, nil
}
func (c *Client) ListFloatingIPs(ctx context.Context) ([]FloatingIP, error) {
return listAll(ctx, c)
}
+5
View File
@@ -39,6 +39,11 @@ type FloatingIPClient interface {
// onPage stops the walk and is returned as-is.
ListFreeFloatingIPs(ctx context.Context, pageSize int, onPage func(page []FloatingIP) error) (pages int, err error)
// ListFloatingIPsByPort returns the floating IPs currently attached to
// the given port — read from the cloud, not from our database, so it
// also finds attachments we lost track of.
ListFloatingIPsByPort(ctx context.Context, portID string) ([]FloatingIP, error)
// AssociateFloatingIP attaches the floating IP to the given Neutron
// port (the validator's primary NIC port).
AssociateFloatingIP(ctx context.Context, fipID, portID string) error
+12
View File
@@ -159,6 +159,18 @@ func (m *MockClient) ListFloatingIPs(ctx context.Context) ([]FloatingIP, error)
return listAll(ctx, m)
}
func (m *MockClient) ListFloatingIPsByPort(ctx context.Context, portID string) ([]FloatingIP, error) {
m.mu.Lock()
defer m.mu.Unlock()
var out []FloatingIP
for _, f := range m.fips {
if f.PortID == portID {
out = append(out, *f)
}
}
return out, nil
}
func (m *MockClient) AssociateFloatingIP(ctx context.Context, fipID, portID string) error {
m.mu.Lock()
defer m.mu.Unlock()
+222 -24
View File
@@ -49,6 +49,36 @@ type Orchestrator struct {
// scan is the background floating-IP scan job (see scanjob.go); its zero
// value is ready to use.
scan scanJob
// Async makes Tick return without waiting for the slow per-address work
// (OpenStack calls): each address is handled in its own goroutine, so one
// slow cloud call never delays another validator. With Async=false (the
// default, used by tests) Tick runs the same work in parallel but waits
// for all of it before returning.
Async bool
// inflight holds the keys of work already running in a goroutine, so the
// next Tick does not start the same work twice.
inflight sync.Map
wg sync.WaitGroup
}
// Wait blocks until all background work started by Tick has finished.
func (o *Orchestrator) Wait() { o.wg.Wait() }
// spawn runs fn in its own goroutine unless work with the same key is
// already running. It returns immediately; callers that need the result
// wait with o.wg (see runAll).
func (o *Orchestrator) spawn(key string, fn func()) {
if _, busy := o.inflight.LoadOrStore(key, struct{}{}); busy {
return
}
o.wg.Add(1)
go func() {
defer o.wg.Done()
defer o.inflight.Delete(key)
fn()
}()
}
// New constructs an Orchestrator. Egress check types/targets, prober sites,
@@ -79,7 +109,16 @@ func (o *Orchestrator) leaseTTL() time.Duration {
// validators, sweep the checking window for ready-to-aggregate IPs, and
// reclaim expired leases. Intended to be called on a fixed interval
// (Cfg.PollIntervalSeconds) by the caller (cmd/control-api/main.go).
//
// The slow part of every step (calls to OpenStack) runs per address in its
// own goroutine, so validators never wait for each other. In Async mode
// Tick does not wait for those goroutines; otherwise it waits for them.
func (o *Orchestrator) Tick(ctx context.Context) {
if n, err := o.DB.ReconcileValidators(ctx); err != nil {
o.Log.Error("reconcile validators", "err", err)
} else if n > 0 {
o.Log.Warn("repaired validators that disagreed with the queue", "count", n)
}
if err := o.assignIdleValidators(ctx); err != nil {
o.Log.Error("assign idle validators", "err", err)
}
@@ -89,6 +128,9 @@ func (o *Orchestrator) Tick(ctx context.Context) {
if err := o.sweepExpiredLeases(ctx); err != nil {
o.Log.Error("sweep expired leases", "err", err)
}
if !o.Async {
o.wg.Wait()
}
}
// assignIdleValidators claims the next queued IP for every currently idle
@@ -108,9 +150,16 @@ func (o *Orchestrator) assignIdleValidators(ctx context.Context) error {
continue // no work available for this validator right now
}
o.Log.Info("claimed ip", "validator", v.ValidatorID, "ip", item.IPAddress, "ip_id", item.ID)
if err := o.associateFIP(ctx, v.ValidatorID, v.OSPortID, item); err != nil {
o.Log.Error("associate fip", "validator", v.ValidatorID, "ip", item.IPAddress, "err", err)
}
v, item := v, item
// Keyed by address, not validator: the key only guards against
// starting the same association twice. A validator-wide key made a
// second address claimed while the first was still associating skip
// its association and wait for the lease to expire.
o.spawn(fmt.Sprintf("assign:%d", item.ID), func() {
if err := o.associateFIP(ctx, v.ValidatorID, v.OSPortID, item); err != nil {
o.Log.Error("associate fip", "validator", v.ValidatorID, "ip", item.IPAddress, "err", err)
}
})
}
return nil
}
@@ -147,6 +196,16 @@ func (o *Orchestrator) associateFIP(ctx context.Context, validatorID, osPortID s
return err
}
if err := o.DB.SetFIPAssociated(ctx, item.ID, fip.ID, o.leaseTTL()); err != nil {
if errors.Is(err, db.ErrInvalidState) {
// The address was deleted or cancelled while the cloud call was
// running: nobody owns this attachment any more, so detach it,
// otherwise it stays on the validator's port forever.
o.Log.Info("address removed during association, detaching floating ip", "ip", item.IPAddress, "fip_id", fip.ID)
if derr := o.OS.DisassociateFloatingIP(ctx, fip.ID); derr != nil {
o.Log.Error("disassociate fip of removed address", "ip", item.IPAddress, "fip_id", fip.ID, "err", derr)
}
return nil
}
return fmt.Errorf("set fip associated: %w", err)
}
o.event(ctx, "control-api", "", &item.ID, "fip_associated", fmt.Sprintf(`{"fip_id":%q,"validator_id":%q}`, fip.ID, validatorID))
@@ -172,6 +231,25 @@ func (o *Orchestrator) SelfCheckResult(ctx context.Context, validatorID string,
if err != nil {
return err
}
// A late report for an address this validator no longer owns (lease
// already reclaimed, address cancelled or deleted) must not touch
// the floating IP: it may be attached for another validator now.
if item.State != db.IPAwaitingSelfCheck || item.OwnerValidatorID == nil || *item.OwnerValidatorID != validatorID {
o.Log.Warn("ignoring failed self-check for an address the validator does not hold",
"validator", validatorID, "ip_id", ipID, "state", item.State)
return nil
}
// Detach the floating IP before the address goes back to the queue.
// requeueOrFail frees the validator in the database but knows nothing
// about the cloud: a floating IP left on the validator's port makes
// every later association on that port fail with 409 ("fixed IP
// already has a floating IP"). Best-effort, like the other release
// paths — the database state must be freed even if Neutron hiccups.
if item.FIPID != "" {
if err := o.OS.DisassociateFloatingIP(ctx, item.FIPID); err != nil {
o.Log.Error("disassociate fip after failed self-check", "ip_id", ipID, "fip_id", item.FIPID, "err", err)
}
}
if item.RetryCount+1 > o.Cfg.MaxSelfCheckRetries {
o.requeueOrFail(ctx, ipID, validatorID, "self-check failed: "+detail)
return nil
@@ -288,6 +366,87 @@ func (o *Orchestrator) MarkSiteComplete(ctx context.Context, ipID int64, siteInd
return o.DB.SetSiteComplete(ctx, ipID, siteIndex)
}
// releaseValidatorPorts detaches, in the cloud, every floating IP that sits on
// the given validators' ports and belongs to an address this system has
// queued. The cloud is asked directly because our database does not know
// about an attachment that is still being created (state assigning_fip has no
// fip_id yet) or that was lost earlier. Floating IPs this system never
// queued are left alone. Ports are processed in parallel.
func (o *Orchestrator) releaseValidatorPorts(ctx context.Context, validators []db.Validator) {
var wg sync.WaitGroup
for _, v := range validators {
if v.OSPortID == "" {
continue
}
wg.Add(1)
go func(v db.Validator) {
defer wg.Done()
fips, err := o.OS.ListFloatingIPsByPort(ctx, v.OSPortID)
if err != nil {
o.Log.Error("list floating ips of validator port", "validator", v.ValidatorID, "port", v.OSPortID, "err", err)
return
}
addrs := make([]string, len(fips))
for i, f := range fips {
addrs[i] = f.Address
}
known, err := o.DB.KnownAddresses(ctx, addrs)
if err != nil {
o.Log.Error("check known addresses", "validator", v.ValidatorID, "err", err)
return
}
for _, f := range fips {
if !known[f.Address] {
o.Log.Warn("floating ip on validator port was not queued by this system, left attached",
"validator", v.ValidatorID, "ip", f.Address)
continue
}
if err := o.OS.DisassociateFloatingIP(ctx, f.ID); err != nil {
o.Log.Error("detach floating ip from validator port", "validator", v.ValidatorID, "ip", f.Address, "err", err)
continue
}
o.Log.Info("detached floating ip from validator port", "validator", v.ValidatorID, "ip", f.Address)
}
}(v)
}
wg.Wait()
}
// validatorsByID returns the validators with the given ids (unknown ids are skipped).
func (o *Orchestrator) validatorsByID(ctx context.Context, ids ...string) []db.Validator {
var out []db.Validator
for _, id := range ids {
if v, err := o.DB.GetValidator(ctx, id); err == nil {
out = append(out, *v)
}
}
return out
}
// busyValidatorsFor returns the validators currently working on one of the
// given addresses.
func (o *Orchestrator) busyValidatorsFor(ctx context.Context, addresses []string) []db.Validator {
want := make(map[string]bool, len(addresses))
for _, a := range addresses {
want[a] = true
}
all, err := o.DB.ListValidators(ctx)
if err != nil {
o.Log.Error("list validators", "err", err)
return nil
}
var out []db.Validator
for _, v := range all {
if v.CurrentIPID == nil {
continue
}
if ip, err := o.DB.GetIP(ctx, *v.CurrentIPID); err == nil && want[ip.IPAddress] {
out = append(out, v)
}
}
return out
}
// ForceCancel stops an in-progress (or still-queued) check for the given
// address on admin request, even though it was never going to finish on
// its own within the checking window. If a floating IP is currently
@@ -316,9 +475,10 @@ func (o *Orchestrator) ForceCancel(ctx context.Context, ipAddress string) error
}
if item.OwnerValidatorID != nil {
if err := o.DB.FreeValidator(ctx, *item.OwnerValidatorID); err != nil {
if err := o.DB.FreeValidator(ctx, *item.OwnerValidatorID, item.ID); err != nil {
return fmt.Errorf("free validator: %w", err)
}
o.releaseValidatorPorts(ctx, o.validatorsByID(ctx, *item.OwnerValidatorID))
}
o.event(ctx, "control-api", "", &item.ID, "force_cancel", "")
@@ -347,9 +507,14 @@ func (o *Orchestrator) DeleteIP(ctx context.Context, ipAddress string) error {
}
}
var owners []db.Validator
if item.OwnerValidatorID != nil {
owners = o.validatorsByID(ctx, *item.OwnerValidatorID)
}
if err := o.DB.DeleteIP(ctx, item.ID); err != nil {
return err
}
o.releaseValidatorPorts(ctx, owners)
o.event(ctx, "control-api", "", nil, "ip_deleted", fmt.Sprintf(`{"ip_address":%q}`, ipAddress))
return nil
@@ -366,20 +531,42 @@ func (o *Orchestrator) DeleteIPs(ctx context.Context, addresses []string) (db.De
if err != nil {
o.Log.Error("list attached fips before delete", "err", err)
}
for _, ref := range refs {
if err := o.OS.DisassociateFloatingIP(ctx, ref.FIPID); err != nil {
o.Log.Error("disassociate fip on delete", "ip_id", ref.IPID, "fip_id", ref.FIPID, "err", err)
}
}
o.disassociateAll(ctx, refs, "delete")
owners := o.busyValidatorsFor(ctx, addresses)
result, err := o.DB.DeleteIPs(ctx, addresses)
if err != nil {
return result, err
}
o.releaseValidatorPorts(ctx, owners)
o.event(ctx, "control-api", "", nil, "ips_deleted", deletedAddressesPayload(result.Deleted))
return result, nil
}
// maxParallelDetach bounds how many floating IPs a bulk delete or clear
// detaches at the same time (each is one Neutron call).
const maxParallelDetach = 8
// disassociateAll detaches the given floating IPs best-effort, at most
// maxParallelDetach at a time; a failure is logged and does not stop the rest.
func (o *Orchestrator) disassociateAll(ctx context.Context, refs []db.FIPRef, what string) {
var wg sync.WaitGroup
sem := make(chan struct{}, maxParallelDetach)
for _, ref := range refs {
ref := ref
sem <- struct{}{}
wg.Add(1)
go func() {
defer wg.Done()
defer func() { <-sem }()
if err := o.OS.DisassociateFloatingIP(ctx, ref.FIPID); err != nil {
o.Log.Error("disassociate fip on "+what, "ip_id", ref.IPID, "fip_id", ref.FIPID, "err", err)
}
}()
}
wg.Wait()
}
// ClearQueue deletes every address currently in the queue, regardless of
// state — the "delete everything" operation. It is set-based (see
// db.ClearAllIPs): O(1) statements however many rows there are. Floating IPs
@@ -390,16 +577,21 @@ func (o *Orchestrator) ClearQueue(ctx context.Context) (db.DeleteIPsResult, erro
if err != nil {
return db.DeleteIPsResult{}, fmt.Errorf("list attached fips: %w", err)
}
for _, ref := range refs {
if err := o.OS.DisassociateFloatingIP(ctx, ref.FIPID); err != nil {
o.Log.Error("disassociate fip on clear queue", "ip_id", ref.IPID, "fip_id", ref.FIPID, "err", err)
}
}
o.disassociateAll(ctx, refs, "clear queue")
deleted, err := o.DB.ClearAllIPs(ctx)
if err != nil {
return db.DeleteIPsResult{}, err
}
// After the rows are gone, free every validator port in the cloud: this
// catches attachments that were still being created (no fip_id yet) and
// ones left over from earlier runs. An association that finishes after
// this point detaches itself (see associateFIP).
if validators, err := o.DB.ListValidators(ctx); err != nil {
o.Log.Error("list validators to release ports", "err", err)
} else {
o.releaseValidatorPorts(ctx, validators)
}
o.event(ctx, "control-api", "", nil, "queue_cleared", deletedAddressesPayload(deleted))
return db.DeleteIPsResult{Deleted: deleted}, nil
}
@@ -453,9 +645,12 @@ func (o *Orchestrator) sweepCheckingWindow(ctx context.Context) error {
if !ready {
continue
}
if err := o.aggregateAndRelease(ctx, item); err != nil {
o.Log.Error("aggregate and release", "ip_id", item.ID, "err", err)
}
item := item
o.spawn(fmt.Sprintf("aggregate:%d", item.ID), func() {
if err := o.aggregateAndRelease(ctx, item); err != nil {
o.Log.Error("aggregate and release", "ip_id", item.ID, "err", err)
}
})
}
return nil
}
@@ -606,14 +801,17 @@ func (o *Orchestrator) sweepExpiredLeases(ctx context.Context) error {
if item.OwnerValidatorID != nil {
validatorID = *item.OwnerValidatorID
}
o.Log.Info("lease expired, reclaiming", "ip_id", item.ID, "ip", item.IPAddress, "validator", validatorID)
if item.FIPID != "" {
if err := o.OS.DisassociateFloatingIP(ctx, item.FIPID); err != nil {
o.Log.Error("disassociate fip on lease reclaim", "ip_id", item.ID, "err", err)
item, validatorID := item, validatorID
o.spawn(fmt.Sprintf("lease:%d", item.ID), func() {
o.Log.Info("lease expired, reclaiming", "ip_id", item.ID, "ip", item.IPAddress, "validator", validatorID)
if item.FIPID != "" {
if err := o.OS.DisassociateFloatingIP(ctx, item.FIPID); err != nil {
o.Log.Error("disassociate fip on lease reclaim", "ip_id", item.ID, "err", err)
}
}
}
o.event(ctx, "control-api", "", &item.ID, "lease_expired", fmt.Sprintf(`{"validator_id":%q}`, validatorID))
o.requeueOrFail(ctx, item.ID, validatorID, "lease expired")
o.event(ctx, "control-api", "", &item.ID, "lease_expired", fmt.Sprintf(`{"validator_id":%q}`, validatorID))
o.requeueOrFail(ctx, item.ID, validatorID, "lease expired")
})
}
return nil
}
+220
View File
@@ -3,6 +3,7 @@ package orchestrator
import (
"context"
"errors"
"fmt"
"log/slog"
"os"
"path/filepath"
@@ -942,3 +943,222 @@ func TestScanFloatingIPsNoFreeAddressesIsNotAnError(t *testing.T) {
t.Fatalf("expected nothing added, got %+v", result)
}
}
// slowOS adds a fixed delay to every OpenStack call, like a loaded cloud.
type slowOS struct {
*openstack.MockClient
delay time.Duration
}
func (s slowOS) GetFloatingIPByAddress(ctx context.Context, a string) (*openstack.FloatingIP, error) {
time.Sleep(s.delay)
return s.MockClient.GetFloatingIPByAddress(ctx, a)
}
func (s slowOS) AssociateFloatingIP(ctx context.Context, fipID, portID string) error {
time.Sleep(s.delay)
return s.MockClient.AssociateFloatingIP(ctx, fipID, portID)
}
// All idle validators must start in the same tick: one slow cloud call may
// not delay the others (20 validators x 2 calls x 200ms = 8s if sequential).
func TestValidatorsStartInParallel(t *testing.T) {
ctx := context.Background()
o, d, mock := newTestOrchestrator(t, 180)
o.OS = slowOS{MockClient: mock, delay: 200 * time.Millisecond}
const n = 20
var addrs []string
for i := 1; i <= n; i++ {
addr := fmt.Sprintf("10.0.0.%d", i)
addrs = append(addrs, addr)
mock.Seed(fmt.Sprintf("fip-%d", i), addr, "svc-project")
if err := d.RegisterValidator(ctx, fmt.Sprintf("validator-%d", i), "h", fmt.Sprintf("port-%d", i), "v"); err != nil {
t.Fatalf("register validator: %v", err)
}
}
if err := d.SeedQueue(ctx, addrs); err != nil {
t.Fatalf("seed queue: %v", err)
}
start := time.Now()
o.Tick(ctx)
if elapsed := time.Since(start); elapsed > 2*time.Second {
t.Fatalf("tick took %s: validators are served one after another", elapsed)
}
for _, a := range addrs {
ip, err := d.GetIPByAddress(ctx, a)
if err != nil {
t.Fatalf("get ip %s: %v", a, err)
}
if ip.State != db.IPAwaitingSelfCheck {
t.Fatalf("ip %s: expected awaiting_self_check, got %s", a, ip.State)
}
}
}
// gatedOS blocks AssociateFloatingIP until release is closed, to hold an
// association "in flight" while the test clears the queue.
type gatedOS struct {
*openstack.MockClient
started chan struct{}
release chan struct{}
}
func (g gatedOS) AssociateFloatingIP(ctx context.Context, fipID, portID string) error {
g.started <- struct{}{}
<-g.release
return g.MockClient.AssociateFloatingIP(ctx, fipID, portID)
}
func portFIPs(t *testing.T, mock *openstack.MockClient, port string) []openstack.FloatingIP {
t.Helper()
got, err := mock.ListFloatingIPsByPort(context.Background(), port)
if err != nil {
t.Fatalf("list by port: %v", err)
}
return got
}
// Clear queue while an association is still running in the cloud: the
// validator port must end up free once the association finishes.
func TestClearQueueFreesPortOfInFlightAssociation(t *testing.T) {
ctx := context.Background()
o, d, mock := newTestOrchestrator(t, 180)
g := gatedOS{MockClient: mock, started: make(chan struct{}, 1), release: make(chan struct{})}
o.OS = g
o.Async = true
mock.Seed("fip-1", "1.2.3.4", "svc-project")
if err := d.RegisterValidator(ctx, "validator-1", "host-1", "port-1", "v0.1"); err != nil {
t.Fatal(err)
}
if err := d.SeedQueue(ctx, []string{"1.2.3.4"}); err != nil {
t.Fatal(err)
}
o.Tick(ctx) // claims the address and starts the association
<-g.started
if _, err := o.ClearQueue(ctx); err != nil {
t.Fatalf("clear queue: %v", err)
}
close(g.release) // the association completes after the clear
o.Wait()
if got := portFIPs(t, mock, "port-1"); len(got) != 0 {
t.Fatalf("port-1 still holds %d floating ip(s): %+v", len(got), got)
}
}
// Clear queue frees ports that already hold a queued address, including an
// attachment left over from earlier, but never touches a floating IP the
// system did not queue.
func TestClearQueueFreesAllValidatorPorts(t *testing.T) {
ctx := context.Background()
o, d, mock := newTestOrchestrator(t, 180)
mock.Seed("fip-1", "1.2.3.4", "svc-project")
mock.SeedWithPort("fip-leak", "1.2.3.9", "svc-project", "port-2") // leaked earlier
mock.SeedWithPort("fip-foreign", "9.9.9.9", "svc-project", "port-2")
if err := d.RegisterValidator(ctx, "validator-1", "h", "port-1", "v"); err != nil {
t.Fatal(err)
}
if err := d.SeedQueue(ctx, []string{"1.2.3.4", "1.2.3.9"}); err != nil {
t.Fatal(err)
}
o.Tick(ctx) // validator-1 gets 1.2.3.4 attached; 1.2.3.9 stays queued, its stale attachment on port-2 is unknown to the queue
if err := d.RegisterValidator(ctx, "validator-2", "h", "port-2", "v"); err != nil {
t.Fatal(err)
}
if got := portFIPs(t, mock, "port-1"); len(got) != 1 {
t.Fatalf("setup: expected 1 floating ip on port-1, got %d", len(got))
}
if _, err := o.ClearQueue(ctx); err != nil {
t.Fatalf("clear queue: %v", err)
}
if got := portFIPs(t, mock, "port-1"); len(got) != 0 {
t.Fatalf("port-1 not free: %+v", got)
}
got := portFIPs(t, mock, "port-2")
if len(got) != 1 || got[0].Address != "9.9.9.9" {
t.Fatalf("port-2: expected only the foreign 9.9.9.9 to remain, got %+v", got)
}
}
// A failed self-check must detach the floating IP in the cloud before the
// address returns to the queue; otherwise the validator's port keeps it and
// every later association on that port fails with 409.
func TestFailedSelfCheckDetachesFloatingIP(t *testing.T) {
ctx := context.Background()
o, d, mock := newTestOrchestrator(t, 180)
mock.Seed("fip-1", "1.2.3.4", "svc-project")
if err := d.RegisterValidator(ctx, "validator-1", "host-1", "port-1", "v0.1"); err != nil {
t.Fatal(err)
}
if err := d.SeedQueue(ctx, []string{"1.2.3.4"}); err != nil {
t.Fatal(err)
}
o.Tick(ctx)
if got := portFIPs(t, mock, "port-1"); len(got) != 1 {
t.Fatalf("setup: expected the floating ip on port-1, got %d", len(got))
}
ip, err := d.GetIPByAddress(ctx, "1.2.3.4")
if err != nil {
t.Fatal(err)
}
if err := o.SelfCheckResult(ctx, "validator-1", ip.ID, false, "ip echo timeout"); err != nil {
t.Fatalf("self-check result: %v", err)
}
if got := portFIPs(t, mock, "port-1"); len(got) != 0 {
t.Fatalf("port-1 still holds the floating ip after a failed self-check: %+v", got)
}
after, err := d.GetIPByAddress(ctx, "1.2.3.4")
if err != nil {
t.Fatal(err)
}
if after.State != db.IPQueued {
t.Fatalf("expected the address back in queued, got %s", after.State)
}
// The validator must be able to take the next claim: a second Tick
// associates the same address again without a conflict.
o.Tick(ctx)
if got := portFIPs(t, mock, "port-1"); len(got) != 1 {
t.Fatalf("expected the retry to attach the floating ip again, got %d", len(got))
}
}
// A failed self-check for an address the validator no longer holds is a late
// report: it must not detach a floating IP that now belongs to someone else.
func TestLateFailedSelfCheckIsIgnored(t *testing.T) {
ctx := context.Background()
o, d, mock := newTestOrchestrator(t, 180)
mock.Seed("fip-1", "1.2.3.4", "svc-project")
for i, id := range []string{"validator-1", "validator-2"} {
if err := d.RegisterValidator(ctx, id, "h", fmt.Sprintf("port-%d", i+1), "v"); err != nil {
t.Fatal(err)
}
}
if err := d.SeedQueue(ctx, []string{"1.2.3.4"}); err != nil {
t.Fatal(err)
}
o.Tick(ctx) // validator-1 (first by id) holds 1.2.3.4
ip, _ := d.GetIPByAddress(ctx, "1.2.3.4")
// validator-2 reports a failure for an address it does not own.
if err := o.SelfCheckResult(ctx, "validator-2", ip.ID, false, "late"); err != nil {
t.Fatalf("self-check result: %v", err)
}
if got := portFIPs(t, mock, "port-1"); len(got) != 1 {
t.Fatalf("a late report from another validator detached the floating ip: %+v", got)
}
cur, _ := d.GetIPByAddress(ctx, "1.2.3.4")
if cur.State != db.IPAwaitingSelfCheck {
t.Fatalf("a late report changed the address state to %s", cur.State)
}
}
@@ -0,0 +1,401 @@
package orchestrator
import (
"context"
"fmt"
"math/rand"
"sync/atomic"
"testing"
"time"
"cloudipvalidator/internal/db"
"cloudipvalidator/internal/openstack"
)
// finishChecks reports a successful egress run and all three inbound sites
// for the address, so the next Tick aggregates it and releases the validator.
func finishChecks(t *testing.T, o *Orchestrator, ip *db.IPQueueItem, validatorID string) {
t.Helper()
ctx := context.Background()
if err := o.RecordCheck(ctx, db.Check{
IPID: ip.ID, IPAddress: ip.IPAddress, AttemptNumber: ip.AttemptNumber,
ValidatorID: validatorID, Source: db.SourceEgress, CheckType: "https",
Target: "https://example.test", Success: true, CheckedAt: db.Now(),
}); err != nil {
t.Fatalf("record egress check: %v", err)
}
if err := o.MarkEgressComplete(ctx, ip.ID); err != nil {
t.Fatalf("mark egress complete: %v", err)
}
for site := 1; site <= 3; site++ {
for _, ct := range []string{"tcp-22", "ssh", "tcp-80", "icmp"} {
if err := o.RecordCheck(ctx, db.Check{
IPID: ip.ID, IPAddress: ip.IPAddress, AttemptNumber: ip.AttemptNumber,
Source: db.InboundSource(site), CheckType: ct, Target: ip.IPAddress,
Success: true, CheckedAt: db.Now(),
}); err != nil {
t.Fatalf("record inbound check: %v", err)
}
}
if err := o.MarkSiteComplete(ctx, ip.ID, site); err != nil {
t.Fatalf("mark site complete: %v", err)
}
}
}
// The incident of 2026-10-02: a validator busy with slow checks goes silent,
// is marked unreachable, and its next heartbeat used to return it to idle
// although it still held the address. It was then handed a second address,
// whose association never ran, and both stalled until their leases expired.
func TestSilentValidatorKeepsItsAddressAfterHeartbeat(t *testing.T) {
ctx := context.Background()
o, d, mock := newTestOrchestrator(t, 180)
mock.Seed("fip-a", "1.1.1.1", "svc")
mock.Seed("fip-b", "2.2.2.2", "svc")
if err := d.RegisterValidator(ctx, "validator-1", "host", "port-1", "v"); err != nil {
t.Fatal(err)
}
if err := d.SeedQueue(ctx, []string{"1.1.1.1", "2.2.2.2"}); err != nil {
t.Fatal(err)
}
o.Tick(ctx) // validator-1 claims 1.1.1.1
a, _ := d.GetIPByAddress(ctx, "1.1.1.1")
if err := o.SelfCheckResult(ctx, "validator-1", a.ID, true, "ok"); err != nil {
t.Fatal(err)
}
// The agent goes silent for longer than heartbeat_timeout_seconds ...
if err := d.MarkValidatorUnreachable(ctx, "validator-1"); err != nil {
t.Fatal(err)
}
// ... and then speaks again while the checks are still running.
if err := d.Heartbeat(ctx, "validator-1"); err != nil {
t.Fatal(err)
}
v, _ := d.GetValidator(ctx, "validator-1")
if v.State != db.ValidatorAssigned || v.CurrentIPID == nil || *v.CurrentIPID != a.ID {
t.Fatalf("after heartbeat: state=%s current_ip=%v, want assigned to %d", v.State, v.CurrentIPID, a.ID)
}
o.Tick(ctx)
if b, _ := d.GetIPByAddress(ctx, "2.2.2.2"); b.State != db.IPQueued {
t.Fatalf("the busy validator was given a second address: 2.2.2.2 is %s", b.State)
}
// The first address finishes: only now the validator may take the next one.
a, _ = d.GetIP(ctx, a.ID)
finishChecks(t, o, a, "validator-1")
o.Tick(ctx) // aggregates and releases 1.1.1.1
o.Tick(ctx) // validator-1 is idle again and claims 2.2.2.2
if b, _ := d.GetIPByAddress(ctx, "2.2.2.2"); b.State != db.IPAwaitingSelfCheck {
t.Fatalf("2.2.2.2 is %s, want awaiting_self_check once the validator is free", b.State)
}
}
// A dead validator whose lease was reclaimed must stay out of rotation until
// it speaks again; before, reclaiming set it idle and it was handed a fresh
// address every lease period, burning each address's retries.
func TestUnreachableValidatorGetsNoAddressesAfterLeaseReclaim(t *testing.T) {
ctx := context.Background()
o, d, mock := newTestOrchestrator(t, 1)
mock.Seed("fip-a", "1.1.1.1", "svc")
mock.Seed("fip-b", "2.2.2.2", "svc")
_ = d.RegisterValidator(ctx, "validator-1", "host", "port-1", "v")
_ = d.SeedQueue(ctx, []string{"1.1.1.1", "2.2.2.2"})
o.Tick(ctx) // claims 1.1.1.1, the agent never answers
if err := d.MarkValidatorUnreachable(ctx, "validator-1"); err != nil {
t.Fatal(err)
}
time.Sleep(1100 * time.Millisecond)
o.Cfg.LeaseTTLSeconds = 180
o.Tick(ctx) // lease sweep reclaims 1.1.1.1
a, _ := d.GetIPByAddress(ctx, "1.1.1.1")
if a.RetryCount != 1 {
t.Fatalf("retry_count = %d, want the lease to be reclaimed once", a.RetryCount)
}
v, _ := d.GetValidator(ctx, "validator-1")
if v.State != db.ValidatorUnreachable || v.CurrentIPID != nil {
t.Fatalf("validator state=%s current_ip=%v, want unreachable and empty", v.State, v.CurrentIPID)
}
o.Tick(ctx)
if a, _ := d.GetIPByAddress(ctx, "1.1.1.1"); a.State != db.IPQueued {
t.Fatalf("an unreachable validator was handed an address: 1.1.1.1 is %s", a.State)
}
if err := d.Heartbeat(ctx, "validator-1"); err != nil { // it is back
t.Fatal(err)
}
if v, _ := d.GetValidator(ctx, "validator-1"); v.State != db.ValidatorIdle {
t.Fatalf("after heartbeat state=%s, want idle (it holds nothing)", v.State)
}
o.Tick(ctx)
if a, _ := d.GetIPByAddress(ctx, "1.1.1.1"); a.State != db.IPAwaitingSelfCheck {
t.Fatalf("1.1.1.1 is %s, want it picked up again", a.State)
}
}
// countingOS counts Disassociate calls.
type countingOS struct {
*openstack.MockClient
disassociations int32
}
func (c *countingOS) DisassociateFloatingIP(ctx context.Context, fipID string) error {
atomic.AddInt32(&c.disassociations, 1)
return c.MockClient.DisassociateFloatingIP(ctx, fipID)
}
// Finished addresses keep their fip_id for display but their floating IP is
// already free; clearing the queue must not make a cloud call for each of
// them (on 2026-10-02 that was >1000 sequential calls: 256 s).
func TestClearQueueDoesNotTouchFinishedAddresses(t *testing.T) {
ctx := context.Background()
o, d, mock := newTestOrchestrator(t, 180)
cos := &countingOS{MockClient: mock}
o.OS = cos
const finished = 300
var addrs []string
for i := 0; i < finished; i++ {
addrs = append(addrs, fmt.Sprintf("10.9.%d.%d", i/200, i%200+1))
}
if err := d.SeedQueue(ctx, addrs); err != nil {
t.Fatal(err)
}
for i, a := range addrs {
state := []string{db.IPDone, db.IPFailed, db.IPOccupied}[i%3]
if _, err := d.ExecContext(ctx, `UPDATE ip_queue SET state=?, fip_id=? WHERE ip_address=?`, state, fmt.Sprintf("fip-old-%d", i), a); err != nil {
t.Fatal(err)
}
}
// One address really holds a floating IP.
mock.Seed("fip-live", "1.2.3.4", "svc")
_ = d.RegisterValidator(ctx, "validator-1", "host", "port-1", "v")
_ = d.SeedQueue(ctx, []string{"1.2.3.4"})
live, _ := d.GetIPByAddress(ctx, "1.2.3.4")
if _, err := d.ExecContext(ctx, `UPDATE ip_queue SET sequence=-1 WHERE id=?`, live.ID); err != nil {
t.Fatal(err)
}
o.Tick(ctx) // validator-1 takes the live address (lowest sequence) and attaches it
if got := portFIPs(t, mock, "port-1"); len(got) != 1 {
t.Fatalf("setup: the live address is not attached (%d floating ips on port-1)", len(got))
}
atomic.StoreInt32(&cos.disassociations, 0)
start := time.Now()
if _, err := o.ClearQueue(ctx); err != nil {
t.Fatalf("clear queue: %v", err)
}
// One call for the live address; the port sweep finds nothing left.
if n := atomic.LoadInt32(&cos.disassociations); n != 1 {
t.Fatalf("clear queue made %d disassociate calls, want 1 (finished addresses must be skipped)", n)
}
if got := portFIPs(t, mock, "port-1"); len(got) != 0 {
t.Fatalf("the live floating ip is still attached: %+v", got)
}
if left, _ := d.ListIPs(ctx); len(left) != 0 {
t.Fatalf("%d rows left after clear", len(left))
}
t.Logf("clear of %d rows took %s", finished+1, time.Since(start))
}
// A bulk detach runs in parallel but never above maxParallelDetach at once.
func TestDisassociateAllIsBoundedAndParallel(t *testing.T) {
o, _, mock := newTestOrchestrator(t, 180)
g := &gaugeOS{MockClient: mock, delay: 30 * time.Millisecond}
o.OS = g
var refs []db.FIPRef
for i := 0; i < 40; i++ {
id := fmt.Sprintf("fip-%d", i)
mock.Seed(id, fmt.Sprintf("10.8.0.%d", i+1), "svc")
refs = append(refs, db.FIPRef{IPID: int64(i), IPAddress: id, FIPID: id})
}
start := time.Now()
o.disassociateAll(context.Background(), refs, "test")
elapsed := time.Since(start)
if peak := atomic.LoadInt32(&g.peak); peak < 2 || peak > maxParallelDetach {
t.Fatalf("peak concurrency %d, want between 2 and %d", peak, maxParallelDetach)
}
if elapsed > 600*time.Millisecond { // 40 x 30 ms sequentially is 1.2 s
t.Fatalf("took %s: the detach is not parallel", elapsed)
}
}
type gaugeOS struct {
*openstack.MockClient
delay time.Duration
cur, peak int32
}
func (g *gaugeOS) DisassociateFloatingIP(ctx context.Context, fipID string) error {
n := atomic.AddInt32(&g.cur, 1)
for {
p := atomic.LoadInt32(&g.peak)
if n <= p || atomic.CompareAndSwapInt32(&g.peak, p, n) {
break
}
}
time.Sleep(g.delay)
atomic.AddInt32(&g.cur, -1)
return g.MockClient.DisassociateFloatingIP(ctx, fipID)
}
// Rows that disagree with the queue are repaired by ReconcileValidators (run on every Tick).
func TestReconcileRepairsInconsistentValidators(t *testing.T) {
ctx := context.Background()
_, d, _ := newTestOrchestrator(t, 180)
_ = d.RegisterValidator(ctx, "validator-1", "host", "port-1", "v")
_ = d.RegisterValidator(ctx, "validator-2", "host", "port-2", "v")
_ = d.SeedQueue(ctx, []string{"1.1.1.1", "2.2.2.2"})
finished, _ := d.GetIPByAddress(ctx, "1.1.1.1")
foreign, _ := d.GetIPByAddress(ctx, "2.2.2.2")
// validator-1 still points at a finished address; validator-2 at an
// address owned by somebody else (what an older version could leave behind).
for _, q := range []string{
fmt.Sprintf(`UPDATE ip_queue SET state='done' WHERE id=%d`, finished.ID),
fmt.Sprintf(`UPDATE ip_queue SET state='checking', owner_validator_id='validator-1' WHERE id=%d`, foreign.ID),
fmt.Sprintf(`UPDATE validators SET state='assigned', current_ip_id=%d WHERE validator_id='validator-1'`, finished.ID),
fmt.Sprintf(`UPDATE validators SET state='assigned', current_ip_id=%d WHERE validator_id='validator-2'`, foreign.ID),
} {
if _, err := d.ExecContext(ctx, q); err != nil {
t.Fatal(err)
}
}
n, err := d.ReconcileValidators(ctx)
if err != nil || n != 2 {
t.Fatalf("reconcile repaired %d validators (err %v), want 2", n, err)
}
for _, id := range []string{"validator-1", "validator-2"} {
v, _ := d.GetValidator(ctx, id)
if v.State != db.ValidatorIdle || v.CurrentIPID != nil {
t.Fatalf("%s: state=%s current_ip=%v, want idle", id, v.State, v.CurrentIPID)
}
}
// A consistent validator is left alone.
if n, _ := d.ReconcileValidators(ctx); n != 0 {
t.Fatalf("second reconcile repaired %d, want 0", n)
}
}
// Randomised run: many validators, random silences and association failures.
// After every tick each validator holds at most one address, the validator
// and the address agree on who holds what, and no lease is ever reclaimed.
func TestRandomFlowKeepsOneAddressPerValidator(t *testing.T) {
ctx := context.Background()
rng := rand.New(rand.NewSource(42))
o, d, mock := newTestOrchestrator(t, 180)
const validators, addresses = 12, 60
var addrs []string
for i := 0; i < validators; i++ {
_ = d.RegisterValidator(ctx, fmt.Sprintf("validator-%d", i+1), "host", fmt.Sprintf("port-%d", i+1), "v")
}
for i := 0; i < addresses; i++ {
a := fmt.Sprintf("10.7.%d.%d", i/200, i%200+1)
addrs = append(addrs, a)
mock.Seed(fmt.Sprintf("fip-%d", i), a, "svc")
}
if err := d.SeedQueue(ctx, addrs); err != nil {
t.Fatal(err)
}
checkInvariants := func(tick int) {
t.Helper()
vs, _ := d.ListValidators(ctx)
holders := map[int64]string{}
for _, v := range vs {
if v.CurrentIPID == nil {
if v.State == db.ValidatorAssigned {
t.Fatalf("tick %d: %s is assigned but holds nothing", tick, v.ValidatorID)
}
continue
}
ip, err := d.GetIP(ctx, *v.CurrentIPID)
if err != nil {
t.Fatalf("tick %d: %s points at a missing address %d", tick, v.ValidatorID, *v.CurrentIPID)
}
if ip.OwnerValidatorID == nil || *ip.OwnerValidatorID != v.ValidatorID {
t.Fatalf("tick %d: %s holds %s but its owner is %v", tick, v.ValidatorID, ip.IPAddress, ip.OwnerValidatorID)
}
if ip.State == db.IPDone || ip.State == db.IPFailed || ip.State == db.IPOccupied {
t.Fatalf("tick %d: %s still holds the finished address %s", tick, v.ValidatorID, ip.IPAddress)
}
if prev, ok := holders[ip.ID]; ok {
t.Fatalf("tick %d: %s is held by both %s and %s", tick, ip.IPAddress, prev, v.ValidatorID)
}
holders[ip.ID] = v.ValidatorID
}
// Every owned, unfinished address is the current one of its owner.
ips, _ := d.ListIPs(ctx)
perOwner := map[string]int{}
for _, ip := range ips {
if ip.OwnerValidatorID != nil && ip.State != db.IPDone && ip.State != db.IPFailed && ip.State != db.IPOccupied {
perOwner[*ip.OwnerValidatorID]++
}
}
for owner, n := range perOwner {
if n > 1 {
t.Fatalf("tick %d: %s owns %d unfinished addresses", tick, owner, n)
}
}
}
checkTicks := map[int64]int{} // address id -> ticks spent in checking
done := false
for tick := 1; tick <= 600 && !done; tick++ {
// Occasionally an association fails (409 on a port).
if rng.Intn(15) == 0 {
mock.AssociateFailures = map[string]error{fmt.Sprintf("fip-%d", rng.Intn(addresses)): fmt.Errorf("409 conflict")}
}
o.Tick(ctx)
vs, _ := d.ListValidators(ctx)
for _, v := range vs {
// A busy agent sometimes goes silent, then speaks again.
if v.CurrentIPID != nil && v.State == db.ValidatorAssigned && rng.Intn(10) == 0 {
_ = d.MarkValidatorUnreachable(ctx, v.ValidatorID)
}
if v.State == db.ValidatorUnreachable && rng.Intn(3) == 0 {
_ = d.Heartbeat(ctx, v.ValidatorID)
}
item, _, err := o.AssignmentForValidator(ctx, v.ValidatorID)
if err != nil || item == nil {
continue
}
if item.State == db.IPAwaitingSelfCheck {
_ = o.SelfCheckResult(ctx, v.ValidatorID, item.ID, true, "ok")
continue
}
checkTicks[item.ID]++
if checkTicks[item.ID] >= 1+rng.Intn(4) {
finishChecks(t, o, item, v.ValidatorID)
delete(checkTicks, item.ID)
}
}
checkInvariants(tick)
counts, _, _ := d.CountIPsByState(ctx)
unfinished := 0
for st, n := range counts {
if st != db.IPDone && st != db.IPFailed && st != db.IPOccupied {
unfinished += n
}
}
done = unfinished == 0
}
if !done {
counts, _, _ := d.CountIPsByState(ctx)
t.Fatalf("not all addresses finished: %v", counts)
}
var expired int
if err := d.QueryRowContext(ctx, `SELECT COUNT(*) FROM events WHERE event_type='lease_expired'`).Scan(&expired); err != nil {
t.Fatal(err)
}
if expired != 0 {
t.Fatalf("%d leases were reclaimed, want 0", expired)
}
}
@@ -13,7 +13,19 @@ poll_interval_seconds: 5
self_check:
timeout_seconds: 10
# Must be a resource genuinely outside the cloud project — OpenStack only
# Способы самопроверки в порядке приоритета (допустимо: ip_echo,
# control_api); по умолчанию [ip_echo]. Самопроверка успешна, если адрес
# подтвердил любой способ: пробуются по порядку, остановка на первом
# успешном, к следующему переходим и при отсутствии ответа, и при
# несовпадении адреса. Таймаут timeout_seconds действует на каждый способ
# отдельно (зависший первый способ не лишает второй времени). control_api спрашивает у control-api, с какого адреса он видит
# это соединение; рекомендуется [control_api, ip_echo], когда control-api
# стоит вне облака и валидатор ходит к нему напрямую (через внешнюю сеть).
# Ограничение: если control-api достижим по внутренней сети облака, он
# увидит частный адрес валидатора и control_api всегда даст несовпадение —
# тогда оставьте только ip_echo (или он сработает вторым в списке).
methods: [control_api, ip_echo]
# Used by the ip_echo method. Must be a resource genuinely outside the cloud project — OpenStack only
# applies floating-IP SNAT to traffic leaving via the external network,
# so anything reachable over the project's internal network (including
# control-api itself, if it's on the same internal network) would report
+3
View File
@@ -110,6 +110,9 @@ control_api_url: "http://127.0.0.1:28080"
poll_interval_seconds: 1
self_check:
timeout_seconds: 5
# ip_echo by default; E2E_SELF_CHECK_METHODS="[control_api]" runs the same
# scenario through control-api's observed-ip route (loopback sees 127.0.0.1).
methods: ${E2E_SELF_CHECK_METHODS:-[ip_echo]}
# Points at the local httpstub's /ip echo route rather than a real
# internet IP-echo service — self-check only produces a meaningful
# signal here because the "floating IP" under test (127.0.0.1) is