Compare commits
6
Commits
1ee5757004
...
main
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
1366ecbdea | ||
|
|
0532baff09 | ||
|
|
cf4a883363 | ||
|
|
49890ff5de | ||
|
|
abbee9a08a | ||
|
|
146259cabb |
No files matched your search
@@ -5,6 +5,8 @@
|
||||
!/deploy/docker/.env.example
|
||||
!/deploy/docker/.env.prod.example
|
||||
/deploy/docker/control-api/control-api.docker.yaml
|
||||
# Ansible: рабочий env-файл запуска validator-agent (шаблон .example остаётся в git).
|
||||
/deploy/ansible/env/*.env
|
||||
/rxprod-compose/.env
|
||||
|
||||
# Live runtime database for the rxprod-compose control-api container.
|
||||
|
||||
@@ -48,7 +48,7 @@ docker compose up -d --build # весь стенд на одной
|
||||
**Остальные компоненты**
|
||||
| Компонент | Ключи |
|
||||
|---|---|
|
||||
| `validator-agent` | `validator_id`, `control_api_url`, `control_api_token_env` (`CONTROL_API_AGENT_TOKEN`), `poll_interval_seconds`, `self_check.*` (таймаут, `ip_echo_urls`), `checks.*` (таймауты HTTPS/ICMP, число ICMP-пакетов, `ssh.*`) |
|
||||
| `validator-agent` | `validator_id`, `control_api_url`, `control_api_token_env` (`CONTROL_API_AGENT_TOKEN`), `poll_interval_seconds`, `self_check.*` (таймаут, `methods` — `ip_echo`/`control_api`, `ip_echo_urls`), `checks.*` (таймауты HTTPS/ICMP, число ICMP-пакетов, `ssh.*`) |
|
||||
| `prober` | `site_id`, `control_api_url`, `control_api_token_env` (`CONTROL_API_AGENT_TOKEN`), `poll_interval_seconds`, `checks.*` (таймауты TCP/ICMP, число ICMP-пакетов) |
|
||||
| `admin-dashboard` | `server.listen_addr` (`:8090`), `control_api.base_url`, `control_api.timeout_seconds`, `control_api.token_env` (`ADMIN_DASHBOARD_CONTROL_API_TOKEN`), `auth.username_env` / `password_env` / `session_secret_env` (`ADMIN_DASHBOARD_USERNAME` / `_PASSWORD` / `_SESSION_SECRET`), `auth.session_ttl_minutes` (480), `overview.last_completed_count` (20), `overview.poll_interval_seconds` (5) |
|
||||
|
||||
@@ -118,7 +118,7 @@ docs/ документация и планы доработок
|
||||
## API (`/api/v1`)
|
||||
| Область | Эндпоинты |
|
||||
|---|---|
|
||||
| Валидатор | `POST /agents/register`, `POST /agents/{id}/heartbeat`, `GET /agents/{id}/assignment`, `POST /agents/{id}/self-check\|events\|results\|complete` |
|
||||
| Валидатор | `POST /agents/register`, `POST /agents/{id}/heartbeat`, `GET /agents/{id}/assignment`, `GET /agents/{id}/observed-ip`, `POST /agents/{id}/self-check\|events\|results\|complete` |
|
||||
| Пробер | `POST /probers/register`, `POST /probers/{site_id}/heartbeat`, `GET /probers/{site_id}/assignments`, `POST /probers/{site_id}/results` |
|
||||
| Очередь | `GET /admin/status`, `GET\|POST /admin/ips` (`limit/offset/state/q/result/order` — постранично), `GET /admin/ips/{ip}`, `POST /admin/ips/{ip}/cancel`, `DELETE /admin/ips/{ip}`, `POST /admin/ips/delete\|clear`, `POST\|GET /admin/ips/scan` (фоновый скан: `202`, `dry_run`, `wait`; статус и прогресс) |
|
||||
| Реестр | `GET /admin/registry` (`limit/offset/q/last_result` — постранично), `GET /admin/registry/{ip}` |
|
||||
@@ -138,7 +138,7 @@ docs/ документация и планы доработок
|
||||
|---|---|---|
|
||||
| admin | `CONTROL_API_ADMIN_TOKEN` | все `/api/v1/admin/*` |
|
||||
| agent | `CONTROL_API_AGENT_TOKEN` | запись результатов валидатора и пробера: `self-check`, `events`, `results`, `complete` |
|
||||
| открыто | — | `GET /healthz`, `register`, `heartbeat` и получение задания (`GET assignment` / `assignments`) |
|
||||
| открыто | — | `GET /healthz`, `register`, `heartbeat`, получение задания (`GET assignment` / `assignments`) и `GET /agents/{id}/observed-ip` |
|
||||
Токены разные: административный не открывает методы агентов, и наоборот. Валидатор и пробер получают настройку и задание без токена, но не могут отправить результат без токена агентов.
|
||||
- **Пустой токен — уровень открыт** (обратная совместимость): `control-api` стартует с предупреждением в логе. На реальном стенде задайте оба токена и ограничьте доступ на уровне сети ([docs/SETUP.md](docs/SETUP.md#сетевые-доступы)); токены идут открытым текстом без TLS — публикуйте через reverse-proxy с TLS. Включать токен агентов нужно **после** его раздачи валидаторам и проберам ([порядок](docs/SETUP.md#5-аутентификация-токены-и-пароль-дашборда)).
|
||||
- **Дашборд закрыт логином и паролем** (один администратор; пароль и ключ сессии — из env). Сессия — подписанная cookie (`HttpOnly`, `SameSite=Strict`, без состояния на сервере), CSRF-защита по `Origin`, 5 неудачных входов за 10 минут с одного IP → `429`. Без заданных логина/пароля дашборд открыт (с предупреждением в логе). Дашборд ходит в API с токеном администратора. Подробности — [docs/DASHBOARD.md](docs/DASHBOARD.md#вход-и-сессия).
|
||||
@@ -149,7 +149,7 @@ docs/ документация и планы доработок
|
||||
## Публикация и эксплуатация
|
||||
- **Бинарники.** Готовые linux/amd64 лежат в `bin/` и **не обновляются автоматически**: после правок кода пересоберите их и обновите `SHA256SUMS` (команды — [docs/SETUP.md](docs/SETUP.md#вариант-b-сборка-из-исходников)); Dockerfile копируют именно `bin/*`.
|
||||
- **systemd.** Юниты в `deploy/systemd/`; у `control-api` — `EnvironmentFile` с учётными данными OpenStack.
|
||||
- **Docker.** `deploy/docker/docker-compose.yml` + `docker-compose.override.yml` (dev, mock) или `docker-compose.prod.yml` (без публикации портов, `restart: unless-stopped`). Какие сервисы поднимаются на хосте, задаёт `COMPOSE_PROFILES`: `control-plane`, `dashboard`, `prober`, `validator`. БД — volume `cloud-ip-validator-db`.
|
||||
- **Docker.** `deploy/docker/docker-compose.yml` + `docker-compose.override.yml` (dev, mock) или `docker-compose.prod.yml` (без публикации портов, `restart: unless-stopped`). Какие сервисы поднимаются на хосте, задаёт `COMPOSE_PROFILES`: `control-plane`, `dashboard`, `prober`, `validator`. БД — volume `cloud-ip-validator-db`. Массовая доставка `validator-agent` на валидаторы (сборка образа на каждом хосте, замена контейнера) — Ansible-сценарий [`deploy/ansible/`](deploy/ansible/README.md).
|
||||
- **Реальный стенд.** `rxprod-compose/` — compose с готовыми образами, собственным `control-api.yaml` и каталогом БД `capi-db/`; `.env` с учётными данными в репозиторий не входит.
|
||||
- **Миграции** применяются при старте `control-api`; версия схемы — `PRAGMA user_version`. Начальная загрузка (`validators`, `sites`, `targets`, `check_types`, `inbound_checks`) выполняется только в пустые таблицы.
|
||||
- Остановка (`SIGTERM`) корректно завершает HTTP-сервер и фоновые циклы. Состояние автоцикла и очереди сохраняется в БД.
|
||||
@@ -182,6 +182,7 @@ scripts/run-local-e2e.sh # сквозной прог
|
||||
| [docs/USAGE.md](docs/USAGE.md) | Повседневная работа: очередь, сканирование Floating IP, автоматический цикл, статус, реестр и история, разбор результатов |
|
||||
| [docs/API.md](docs/API.md) | Спецификация HTTP API `control-api` и примеры запросов (curl) |
|
||||
| [docs/DASHBOARD.md](docs/DASHBOARD.md) | Устройство `admin-dashboard`: страницы, поиск и фильтр, обработка ошибок |
|
||||
| [docs/ADMIN_CLEANUP.md](docs/ADMIN_CLEANUP.md) | Ручная очистка БД администратором (SQL): сброс данных перед новым прогоном, журнал событий, реестр проверок, отдельные адреса, резервная копия и восстановление |
|
||||
| [docs/DIAGRAMS.md](docs/DIAGRAMS.md) | Диаграммы потоков данных: control plane, egress-проверка, телеметрия |
|
||||
| [docs/LOCAL_E2E.md](docs/LOCAL_E2E.md) | Полностью офлайн-прогон всей системы одним скриптом |
|
||||
| [docs/changes/](docs/changes/) | Планы доработок и отчёты ревью с отметкой времени в имени файла (последняя: [скан при тысячах адресов](docs/changes/2026-10-01_18-59_fip-scan-at-scale-review.md)) |
|
||||
@@ -200,6 +201,8 @@ scripts/run-local-e2e.sh # сквозной прог
|
||||
|
||||
| Дата | Веха | Документ |
|
||||
|---|---|---|
|
||||
| 2026-10-02 | Валидатор не получает второй адрес при потере heartbeat (иначе адреса уходили в `fail` без проверок); heartbeat агента в отдельном потоке; быстрая очистка очереди, не зависящая от соединения клиента | [план](docs/changes/2026-10-02_09-02_orchestrator-validator-state-and-clear-plan.md) · [анализ инцидента](analysis/2026-10-02_08-56_1026-addresses_mass-check-analysis.md) · [USAGE](docs/USAGE.md#управление-валидаторами) |
|
||||
| 2026-10-02 | Самопроверка через ручку control-api (`self_check.methods`, способ `control_api` рядом с IP-echo); отвязка Floating IP при провале self-check | [план](docs/changes/2026-10-02_03-06_self-check-control-api-plan.md) · [API](docs/API.md#get-apiv1agentsidobserved-ip) |
|
||||
| 2026-10-01 | Скан Floating IP при тысячах адресов: фоновый постраничный скан с прогрессом, фаза `scanning` в автоцикле, постраничные `/ips` и `/registry`, «Обзор» на счётчиках | [план](docs/changes/2026-10-01_18-19_fip-scan-at-scale-plan.md) · [ревью и тесты](docs/changes/2026-10-01_18-59_fip-scan-at-scale-review.md) · [USAGE](docs/USAGE.md#сканирование-floating-ip-из-openstack) · [API](docs/API.md#post-apiv1adminipsscan) |
|
||||
| 2026-10-01 | Аутентификация: токены администратора и агентов для API, логин и пароль для дашборда | [план](docs/changes/2026-10-01_11-12_authentication-plan.md) · [ревью и тесты](docs/changes/2026-10-01_11-31_authentication-review.md) · [API](docs/API.md#аутентификация) |
|
||||
| 2026-10-01 | Автоматический цикл проверок по сценарию: очистка → скан FIP → проверка → пауза | [USAGE](docs/USAGE.md#автоматический-цикл-проверок) · [API](docs/API.md#автоматический-цикл-проверок) |
|
||||
|
||||
@@ -0,0 +1,167 @@
|
||||
# Аналитический разбор массовой проверки: 1026 адресов
|
||||
|
||||
> Время отчёта: 2026-10-02 08:56 UTC · Данные: снимок БД control-api на 08:46:53 UTC и лог control-api за 4 часа до остановки
|
||||
> Проверено адресов: **1026** (984 завершены `done` + 42 завершены `failed`) из 6440 в очереди
|
||||
> Окно прогона: 07:14:58 – 08:46:53 UTC (1 ч 32 мин). Остановлен вручную в 08:53:46 UTC.
|
||||
|
||||
## 1. Остановка проверок
|
||||
|
||||
- Проверки остановлены в 08:53:46 UTC операцией «Очистить всё» (`POST /api/v1/admin/ips/clear`).
|
||||
- После остановки: очередь пуста, все 20 валидаторов `idle`, последняя выдача адреса в 08:53:43, новых выдач нет.
|
||||
Реестр с историей адресов сохранён (6445 записей).
|
||||
- Отвязка Floating IP с портов валидаторов при очистке: 13 зависших привязок снято прямым опросом портов
|
||||
(`detached floating ip from validator port`), ещё 8 привязок, завершавшихся во время очистки, система сняла сама
|
||||
(`address removed during association`). Состояние портов в самом OpenStack на момент отчёта не проверялось.
|
||||
- **Первая попытка очистки не сработала.** Очистка шла дольше таймаута клиента (120 с) и оборвалась: 600 отвязок завершились
|
||||
с ошибкой `context canceled`, ничего не удалилось, проверки продолжались. Повторная очистка без таймаута выполнялась 256 с.
|
||||
- **Причина долгой очистки (дефект кода, не исправлен):** у завершённых адресов (`done`) в БД остаётся `fip_id`, и очистка
|
||||
последовательно отвязывает все такие FIP, хотя они уже свободны (более 1000 вызовов OpenStack по ~0,2 с).
|
||||
|
||||
## 2. Итоги прогона
|
||||
|
||||
| Показатель | Значение |
|
||||
|---|---|
|
||||
| Адресов в очереди | 6440 |
|
||||
| Завершено | 1026 (984 `done` + 42 `failed`) |
|
||||
| `pass` | 227 (23% от завершённых) |
|
||||
| `partial` | 757 (77%) |
|
||||
| `fail` | 42 (все `failed`, вердикта по существу нет, см. раздел 4) |
|
||||
| Остались в очереди / в работе на момент снимка | 5392 в очереди, 22 в работе |
|
||||
|
||||
Пропускная способность по 10-минутным окнам (завершено адресов за минуту): 07:10 — 5,8; 07:20 — 13,4; 07:30 — 11,6;
|
||||
07:40 — 10,1; 07:50 — 10,6; 08:00 — 9,4; 08:10 — 9,7; 08:20 — 10,5; 08:30 — 10,6; 08:40 — 6,7 (окно неполное).
|
||||
|
||||
Цикл одного адреса (от выдачи валидатору до итога): минимум 65 с, медиана 85 с, p90 100 с, p99 115 с, максимум 125 с,
|
||||
среднее 84 с (по 984 адресам `done`). Для 20 валидаторов это теоретически ~14 адресов в минуту; фактически ~10,4,
|
||||
потеря около 27% из-за залипших валидаторов (раздел 4).
|
||||
|
||||
## 3. Фактура по накопившимся ошибкам
|
||||
|
||||
### 3.1. Классы ошибок (с 07:10)
|
||||
|
||||
В логе control-api за 4 часа на уровне `ERROR` только один вид сообщений: 41 ошибка привязки Floating IP.
|
||||
|
||||
| Ошибка / событие | Число | Что это |
|
||||
|---|---|---|
|
||||
| `lease expired` (возврат адреса по истечении лизинга) | 198 | Валидатор не подхватил задание за время лизинга. Затронуто минимум 77 адресов (часть событий без привязки к адресу) |
|
||||
| Привязка FIP: `409 Cannot associate floating IP … fixed IP already has a floating IP` | 41 | На порту валидатора уже висит другой FIP. Порты: v1 — 13, v7 — 6, v13 — 6, v3 — 5, v12 — 5, v16 — 4, v17 — 2 |
|
||||
| `validator_unreachable` | 7 | По одному разу: v1 и v12 (07:18:03), v13 (07:33:53), v3 (07:34:33), v7 (07:40:03), v16 (07:42:33), v17 (08:17:53) |
|
||||
| `site_unreachable` | 5 | rxmsk (08:09, 08:45) и misha-v (08:09, 08:20, 08:45) |
|
||||
| «Очистить всё»: `context canceled` | 600 | Последствие обрыва первой очистки (08:49), см. раздел 1 |
|
||||
|
||||
Чего не было: **провалов self-check — 0 из 1015** результатов; все 1015 прошли способом `control_api` (запасной `ip_echo`
|
||||
не понадобился). Событий `fip_occupied` — 0, адресов `occupied` — 0.
|
||||
|
||||
Журнал событий за прогон: `self_check_result` 2030 (по две записи на адрес), `fip_associated` 1120, `config_received` 1015,
|
||||
`aggregated` 1004, `retry_or_fail` 239 (198 лизинг + 41 привязка), `lease_expired` 198, `validator_unreachable` 7,
|
||||
`site_unreachable` 5.
|
||||
|
||||
### 3.2. По валидаторам
|
||||
|
||||
| Валидатор | Завершено (`done`) | `failed` | Сбросов лизинга | `unreachable` |
|
||||
|---|---|---|---|---|
|
||||
| vkiplab-v1 | 2 | 17 | 49 | 1 |
|
||||
| vkiplab-v12 | 2 | 13 | 51 | 1 |
|
||||
| vkiplab-v13 | 13 | 10 | 41 | 1 |
|
||||
| vkiplab-v16 | 38 | 2 | 19 | 1 |
|
||||
| vkiplab-v7 | 36 | 0 | 21 | 1 |
|
||||
| vkiplab-v17 | 50 | 0 | 10 | 1 |
|
||||
| vkiplab-v3 | 50 | 0 | 7 | 1 |
|
||||
| остальные 13 (v2, v4–v6, v8–v11, v14, v15, v18–v20) | 60–63 | 0 | 0 | 0 |
|
||||
|
||||
(Сбросы лизинга и `unreachable` в таблице — только за прогон, с 07:10. У v14, v18 и v2 в истории БД есть сбросы лизинга
|
||||
за 1 октября, к этому прогону они не относятся.)
|
||||
|
||||
- **v1 и v12** не подхватили ни одного задания после 07:17 (последний подхват 07:17:29 и 07:17:26).
|
||||
- **v13** — после 07:33:16.
|
||||
- **v7, v16, v17, v3** залипали временно и затем восстановились; механизм восстановления не выяснен.
|
||||
|
||||
### 3.3. 42 адреса `fail`
|
||||
|
||||
- По валидаторам: v1 — 17, v12 — 13, v13 — 10, v16 — 2. Все 42 исчерпали повторы: `retry_count` = 4 и `attempt_number` = 4 у каждого.
|
||||
- Причины неудачных попыток по этим адресам: 143 сброса лизинга и 25 ошибок привязки `409`.
|
||||
- **Ни одной проверки по ним не выполнено** (в таблице проверок у этих адресов 0 записей): адреса не получили вердикта,
|
||||
и `fail` здесь не характеризует сами адреса.
|
||||
- Появлялись равномерно с 07:31 до 08:46 (4–10 за 10 минут).
|
||||
- Подсети: 37.139.x, 79.137.x, 83.166.x и другие, без концентрации.
|
||||
|
||||
## 4. Причины
|
||||
|
||||
Подтверждены по БД и логам control-api. Логи агентов на самих валидаторах не изучались.
|
||||
|
||||
**Причина 1. Валидатор получает два адреса сразу и застревает.** Два дефекта вместе:
|
||||
- Heartbeat (`queries_validators.go`) возвращает валидатор из `unreachable` в `idle`, не проверяя, что за ним числится адрес.
|
||||
Адрес с долгими внешними проверками (3–4 таймаута по 10 с) блокирует агента больше 30 с (порог
|
||||
`heartbeat_timeout_seconds`), control-api помечает валидатор недоступным, затем возвращает в `idle` занятым.
|
||||
- При завершении старого адреса `ReleaseFIP` (`queries_ipqueue.go`) освобождает валидатор по его имени, а не по адресу.
|
||||
- Подтверждение: **35 двойных выдач** (два адреса одному валидатору с интервалом ~5 с) в логе: v1 — 10, v12 — 9, v7 — 6,
|
||||
v13 — 4, v17 — 3, v16 — 2, v3 — 1. Из 1265 выдач в логе.
|
||||
|
||||
**Причина 2. Регрессия моей правки с параллельной привязкой.** Защита от дублей в `orchestrator.go:149` ключуется по
|
||||
валидатору (`assign:<validator>`). Вторая выдача того же валидатора пропускает привязку, и адрес стоит в `assigning_fip`
|
||||
до истечения лизинга. В БД у валидатора одно поле `current_ip_id`, оно указывает на последний выданный адрес, агент по нему
|
||||
получает пустое задание, лизинг истекает, валидатор берёт новый адрес, и круг повторяется. Ошибки `409` — следствие: на
|
||||
порту остаётся FIP первого адреса, привязать второй нельзя.
|
||||
|
||||
**Причина 3. Агент молчит во время долгих проверок.** Heartbeat отправляется только между заданиями; адреса с несколькими
|
||||
таймаутами ведут к `unreachable` (пусковой механизм причины 1). Пример: v1 проверял `37.139.32.1`, v12 — `37.139.32.4`;
|
||||
у обоих проваливались все четыре внешних HTTPS-цели (~40 с таймаутов), в 07:18:03 оба помечены `unreachable`.
|
||||
|
||||
**Дефект очистки.** См. раздел 1: отвязка всех `done`-адресов последовательно.
|
||||
|
||||
## 5. Результаты проверок самих адресов
|
||||
|
||||
**Почему 77% `partial`:** 735 адресов проваливают только исходящие HTTPS, ещё 22 — исходящие и входящие.
|
||||
|
||||
| Цель (egress HTTPS) | Адресов с провалом (из 984) |
|
||||
|---|---|
|
||||
| `packages.ubuntu.com` | 698 (71%) |
|
||||
| `repo.almalinux.org/almalinux/` | 304 (31%) |
|
||||
| `github.com` | 251 (26%) |
|
||||
| `hub.docker.com` | 17 (2%) |
|
||||
|
||||
Число провалов на один `partial`-адрес: 1 — 392 адреса, 2 — 202, 3 — 136, 4 и больше — 27.
|
||||
|
||||
По подсетям (доля адресов с провалом цели):
|
||||
|
||||
| Подсеть | Адресов | `packages.ubuntu.com` | `repo.almalinux.org` | `github.com` | `hub.docker.com` | Доля `partial` |
|
||||
|---|---|---|---|---|---|---|
|
||||
| 83.166.x | 435 | 80% | 53% | 42% | 0% | 92% |
|
||||
| 37.139.x | 418 | 62% | 11% | 10% | 4% | 63% |
|
||||
| 79.137.x | 109 | 66% | 19% | 18% | 0% | 68% |
|
||||
| 5.188.x | 22 | 59% | 0% | 0% | 0% | 59% |
|
||||
|
||||
- `packages.ubuntu.com` проваливается у всех подсетей и во все окна. Доля проваленных строк egress для этой цели по 10-минутным
|
||||
окнам выросла с 26% до 40% (строки включают HTTPS и ICMP, поэтому реальная доля HTTPS вдвое выше).
|
||||
- Доля `partial` почти одинакова у всех валидаторов (72–84%; v1 — 100% по 2 адресам, v12 — 50% по 2): причина в самом адресе
|
||||
или во внешнем ресурсе, а не в валидаторе.
|
||||
- Провалы `repo.almalinux.org` и `github.com` сильно зависят от подсети (83.166.x — 53% и 42%, 37.139.x — 11% и 10%).
|
||||
- Успешные HTTPS-проверки: медиана 200 мс, p90 6707 мс (много ответов близко к таймауту 10 с).
|
||||
|
||||
**Входящие проверки.** Провалы только у `inbound-site-1` (33 пробы из 2934) и `inbound-site-3` (50 из 2931); `inbound-site-2` — 0 из 2952,
|
||||
`inbound-site-4` — 1 из 2952. Провалы всплесками в 08:00, 08:20 и 08:40 по ICMP, SSH и TCP 22; у 20 адресов упали все пробы
|
||||
одной площадки. Пробер rxmsk и misha-v в эти же минуты отмечены `unreachable`. Причина недоступности проберов не выяснена.
|
||||
|
||||
## 6. Рекомендации
|
||||
|
||||
1. **Исправить оркестратор** (до повторного запуска): защита от дублей по адресу, `unreachable` → `assigned` вместо `idle`,
|
||||
освобождение валидатора только по текущему адресу, heartbeat агента в отдельном потоке.
|
||||
2. **Исправить очистку:** отвязывать только действительно привязанные FIP (без `fip_released_at`) и параллельно.
|
||||
3. **Перепроверить 42 адреса** после исправления: вердикта по существу они не получили. Все `fail` с причиной «lease expired» считать недействительными.
|
||||
4. **Решить по `packages.ubuntu.com`:** 71% `partial` и 10 с таймаута на каждый такой адрес; если ресурс нестабилен, убрать или заменить.
|
||||
5. **Проверить причины `site_unreachable`** у проберов (возможна перегрузка при большой очереди).
|
||||
6. Проверить в OpenStack, что порты валидаторов свободны от Floating IP.
|
||||
|
||||
## 7. Что не проверено
|
||||
|
||||
- Состояние портов и Floating IP в самом OpenStack.
|
||||
- Логи агентов `validator-agent` на валидаторах и логи проберов (причины блокировок heartbeat подтверждены по времени и
|
||||
событиям control-api, не по логам агентов).
|
||||
- Почему залипшие v7, v16, v17, v3 восстановились.
|
||||
- Причины провалов внешних HTTPS-целей (ресурс, сеть облака или фильтрация) и недоступности проберов.
|
||||
|
||||
## 8. Источники данных
|
||||
|
||||
- Снимок БД control-api: `.backup` на 08:46:53 UTC (таблицы `ip_queue`, `checks`, `events`, `validators`, `sites`).
|
||||
- Лог контейнера control-api за 4 часа до остановки (1504 строки) и за период очистки.
|
||||
- Статус и результат очистки: HTTP 200 за 256,7 с.
|
||||
@@ -0,0 +1,42 @@
|
||||
37.139.32.70
|
||||
37.139.32.71
|
||||
37.139.32.72
|
||||
37.139.32.151
|
||||
37.139.33.219
|
||||
37.139.33.227
|
||||
37.139.33.230
|
||||
37.139.34.75
|
||||
37.139.34.94
|
||||
37.139.34.128
|
||||
37.139.40.122
|
||||
37.139.41.103
|
||||
37.139.41.129
|
||||
37.139.41.157
|
||||
37.139.41.202
|
||||
37.139.42.126
|
||||
37.139.43.9
|
||||
79.137.174.150
|
||||
79.137.174.172
|
||||
79.137.174.205
|
||||
79.137.174.210
|
||||
79.137.175.14
|
||||
79.137.175.51
|
||||
79.137.175.71
|
||||
79.137.175.162
|
||||
83.166.232.116
|
||||
83.166.232.117
|
||||
83.166.232.159
|
||||
83.166.233.37
|
||||
83.166.233.75
|
||||
83.166.233.78
|
||||
83.166.234.136
|
||||
83.166.235.75
|
||||
83.166.235.132
|
||||
83.166.235.247
|
||||
83.166.237.59
|
||||
83.166.237.228
|
||||
83.166.238.19
|
||||
83.166.248.57
|
||||
83.166.248.60
|
||||
83.166.248.92
|
||||
83.166.248.188
|
||||
+2
-2
@@ -1,4 +1,4 @@
|
||||
4302c5d936b0f6567e4b3b7af3de0123afd862b3cdda53859fad0c7eec443c33 control-api
|
||||
091ad94b5706b1778b181542de7a4b251cd16c421144269d0955769603d51e06 validator-agent
|
||||
9447699d9eed4b9fb5d417769fc868986a10357ec633ec3742883b3002aaafc3 control-api
|
||||
9fb6608b84143f7c4f318f3cc92dcd9f95c7831d627b23d67cce5a5908ced704 validator-agent
|
||||
3e9e14dbb361ee76aaad7c1da6864b3ea111e0ed151403f904b12485631bbf75 prober
|
||||
d72234688eb1954dbff420ca1c8e83b83ceab561b18a0336af0f73fe4d9ac8be admin-dashboard
|
||||
Binary file not shown.
Binary file not shown.
@@ -13,7 +13,19 @@ poll_interval_seconds: 5
|
||||
|
||||
self_check:
|
||||
timeout_seconds: 10
|
||||
# Must be a resource genuinely outside the cloud project — OpenStack only
|
||||
# Способы самопроверки в порядке приоритета (допустимо: ip_echo,
|
||||
# control_api); по умолчанию [ip_echo]. Самопроверка успешна, если адрес
|
||||
# подтвердил любой способ: пробуются по порядку, остановка на первом
|
||||
# успешном, к следующему переходим и при отсутствии ответа, и при
|
||||
# несовпадении адреса. Таймаут timeout_seconds действует на каждый способ
|
||||
# отдельно (зависший первый способ не лишает второй времени). control_api спрашивает у control-api, с какого адреса он видит
|
||||
# это соединение; рекомендуется [control_api, ip_echo], когда control-api
|
||||
# стоит вне облака и валидатор ходит к нему напрямую (через внешнюю сеть).
|
||||
# Ограничение: если control-api достижим по внутренней сети облака, он
|
||||
# увидит частный адрес валидатора и control_api всегда даст несовпадение —
|
||||
# тогда оставьте только ip_echo (или он сработает вторым в списке).
|
||||
methods: [control_api, ip_echo]
|
||||
# Used by the ip_echo method. Must be a resource genuinely outside the cloud project — OpenStack only
|
||||
# applies floating-IP SNAT to traffic leaving via the external network,
|
||||
# so anything reachable over the project's internal network (including
|
||||
# control-api itself, if it's on the same internal network) would report
|
||||
|
||||
@@ -0,0 +1,110 @@
|
||||
# Доставка validator-agent на валидаторы (Ansible)
|
||||
|
||||
Сценарий запускается с jump-хоста и на каждой ВМ-валидаторе: обновляет git-клон в `/opt/cloud-ip-validator`,
|
||||
**собирает образ на самом хосте**, останавливает и удаляет старый контейнер `validator-agent` (образ — `cloud-ip-validator-validator-agent`)
|
||||
и поднимает на его месте новый. Параметры запуска агента вынесены в env-файл.
|
||||
|
||||
Порядок безопасен: сначала проверки, обновление кода и сборка образа, и только потом замена контейнера. Если что-то
|
||||
упало до замены, старый контейнер продолжает работать. Простой валидатора — секунды (`stop` + `rm` + `run`).
|
||||
|
||||
## Требования
|
||||
|
||||
| Где | Что |
|
||||
|---|---|
|
||||
| jump-хост (Debian 13) | `apt install ansible-core`; SSH-доступ по ключу ко всем валидаторам |
|
||||
| валидаторы | Docker, git, клон репозитория в `/opt/cloud-ip-validator`, `sudo` без пароля для SSH-пользователя, доступ к Gitea и Docker Hub (`alpine:3.20`), архитектура x86_64 |
|
||||
|
||||
Только модули `ansible.builtin`: Python Docker SDK и дополнительные коллекции на валидаторах не нужны.
|
||||
Go на валидаторах не нужен: образ копирует закоммиченный `bin/validator-agent` (его сумма проверяется по `bin/SHA256SUMS`).
|
||||
|
||||
## Подготовка (один раз)
|
||||
|
||||
```bash
|
||||
cd deploy/ansible
|
||||
cp env/validator-agent.env.example env/validator-agent.env
|
||||
chmod 600 env/validator-agent.env
|
||||
$EDITOR env/validator-agent.env # адрес control-api, токен, способы самопроверки, таймауты
|
||||
ansible validators -m ping # проверка связи (20 хостов: validator-1 ... validator-20)
|
||||
```
|
||||
|
||||
- В `inventory/hosts.yml` 20 хостов `validator-1 … validator-20` с адресами (`ansible_host`) и переменной `validator_id`
|
||||
(`vkiplab-v1 … vkiplab-v20`, как в control-api). Соответствие `validator-N` → `vkiplab-vN` задано по порядку номеров;
|
||||
**сценарий сверяет `validator_id` с тем, что записано в работающем контейнере на хосте, и при расхождении останавливается до замены
|
||||
контейнера**. `VALIDATOR_AGENT_VALIDATOR_ID` в env-файл писать не нужно: сценарий подставляет его из inventory.
|
||||
- Отпечатки хостов запоминаются при первом подключении (`StrictHostKeyChecking=accept-new` в `ansible.cfg`); сменившийся отпечаток известного хоста — ошибка.
|
||||
- SSH: пользователь `debian` и ключ `~/.ssh/vk_cloud_priv.key` на jump-хосте (права 0600) — одинаковые на всех хостах; меняются в
|
||||
`inventory/group_vars/validators.yml` (`ansible_user`, `ansible_ssh_private_key_file`). Для docker и записи env-файла сценарий
|
||||
повышает права через `sudo`.
|
||||
- Токен в env-файле можно зашифровать: `ansible-vault encrypt env/validator-agent.env`, запускать с `--ask-vault-pass`.
|
||||
Рабочий `env/validator-agent.env` в git не попадает (`.gitignore`).
|
||||
|
||||
## Запуск
|
||||
|
||||
```bash
|
||||
cd deploy/ansible
|
||||
ansible-playbook playbooks/deploy-validator-agent.yml --limit vkiplab-v1 # канарейка: один валидатор
|
||||
ansible-playbook playbooks/deploy-validator-agent.yml # все валидаторы волнами
|
||||
ansible-playbook playbooks/deploy-validator-agent.yml --check # только проверки (preflight), без изменений
|
||||
```
|
||||
|
||||
Волны по умолчанию: 1 хост, затем 4, затем все остальные (`deploy_serial: [1, 4, "100%"]`). Любой сбой в волне останавливает
|
||||
прогон: следующая волна не начнётся. Все сразу: `-e '{"deploy_serial": ["100%"]}'`.
|
||||
|
||||
Что выкатывается — `deploy_ref` (по умолчанию `main`): ветка, тег или коммит. Выкатить нужно **запушенный** коммит:
|
||||
валидаторы берут код из репозитория, а не с jump-хоста. После правок кода сначала пересоберите `bin/validator-agent`
|
||||
(см. [SETUP.md](../../docs/SETUP.md#обновление-образов-после-изменения-кода)) и закоммитьте его вместе с `bin/SHA256SUMS`.
|
||||
|
||||
### Откат
|
||||
|
||||
```bash
|
||||
ansible-playbook playbooks/deploy-validator-agent.yml -e deploy_ref=<предыдущий коммит или тег>
|
||||
```
|
||||
|
||||
Для тега или коммита клон переходит в detached HEAD; следующий запуск с `deploy_ref=main` возвращает его на ветку.
|
||||
|
||||
## Что делает сценарий на каждом хосте
|
||||
|
||||
1. **preflight** — env-файл на jump-хосте существует и в нём задан `VALIDATOR_AGENT_CONTROL_API_URL` (адрес-пример
|
||||
`example.com` не принимается); на валидаторе отвечает Docker, есть git и клон, архитектура x86_64; `validator_id` из inventory совпадает
|
||||
с `validator_id` работающего контейнера; на хосте нет другого контейнера этого агента (по имени или образу) — иначе рядом со старым
|
||||
запустился бы второй с тем же `validator_id`. Показывает текущий контейнер.
|
||||
2. **git** — от имени `git_user` (на валидаторах `root`: клон принадлежит ему) `fetch`, затем клон сбрасывается на `deploy_ref` (`checkout --force`). Через `git`, а не модуль `git`:
|
||||
учётные данные, уже настроенные в клоне, не трогаются. Локальные правки отслеживаемых файлов в клоне будут сброшены;
|
||||
неотслеживаемые и игнорируемые (`.env.*`) — нет.
|
||||
3. **build** — проверка `bin/validator-agent` по `bin/SHA256SUMS`; `docker build --platform linux/amd64` с контекстом в корне репозитория,
|
||||
образ получает метку ревизии (`cloud-ip-validator-validator-agent:<хеш>`) и `latest`.
|
||||
4. **replace** — env-файл копируется в `/opt/cloud-ip-validator/deploy/docker/.env.validator` (0600; путь закрыт `.gitignore`),
|
||||
затем `docker stop` → `docker rm -f` → `docker run -d --restart unless-stopped --cap-add NET_RAW --env-file …`.
|
||||
5. **verify** — ждёт строку `registered` в логе агента (регистрация в control-api), проверяет, что контейнер запущен, без перезапусков
|
||||
и на только что собранном образе. При неудаче хост падает с последними строками лога, следующие волны не стартуют.
|
||||
Затем остаются 3 последних образа с метками ревизий (`keep_images`), остальные и «висячие» удаляются.
|
||||
|
||||
Все шаги можно запускать по тегам: `--tags preflight|git|build|replace|verify`.
|
||||
|
||||
## Параметры
|
||||
|
||||
**`env/validator-agent.env`** — запуск агента (читает `deploy/docker/validator-agent/docker-entrypoint.sh`):
|
||||
`VALIDATOR_AGENT_CONTROL_API_URL`, `CONTROL_API_AGENT_TOKEN`, `VALIDATOR_AGENT_SELF_CHECK_METHODS`,
|
||||
`VALIDATOR_AGENT_SELF_CHECK_TIMEOUT_SECONDS`, `VALIDATOR_AGENT_POLL_INTERVAL_SECONDS`, `VALIDATOR_AGENT_HTTPS_TIMEOUT_SECONDS`,
|
||||
`VALIDATOR_AGENT_ICMP_TIMEOUT_SECONDS`, `VALIDATOR_AGENT_ICMP_COUNT`, `VALIDATOR_AGENT_SSH_ENABLED`, `VALIDATOR_AGENT_SSH_TIMEOUT_SECONDS`.
|
||||
Смена только env-файла не требует изменения кода: достаточно запустить сценарий (контейнер пересоздаётся с новыми значениями).
|
||||
|
||||
**`inventory/group_vars/validators.yml`** — доставка: `repo_dir`, `repo_remote`, `deploy_ref`, `git_user` (пользователь, у которого в клоне
|
||||
настроен доступ к репозиторию; пусто — SSH-пользователь), `image_name`, `container_name`, `platform`, `keep_images`, `restart_policy`,
|
||||
`capabilities`, `log_max_size`, `log_max_file`, `stop_timeout`, `verify_retries`, `verify_delay`, `local_env_file`. Любой параметр
|
||||
переопределяется ключом `-e`.
|
||||
|
||||
## Разбор сбоев
|
||||
|
||||
- **«Нет env-файла» / «не задан VALIDATOR_AGENT_CONTROL_API_URL»** — см. «Подготовка».
|
||||
- **«в запущенном контейнере validator_id=…, а в inventory …»** — соответствие `validator-N` и `validator_id` в `inventory/hosts.yml`
|
||||
неверно для этого хоста: исправьте inventory (контейнер при этом не тронут).
|
||||
- **Контейнер не прошёл проверку** — в сообщении последние строки лога. `registration failed` — control-api недоступен с валидатора
|
||||
или `validator_id` не заведён в control-api. Контейнер остаётся на хосте для разбора (`docker logs`).
|
||||
- **`Permission denied` на `.git/FETCH_HEAD`, `dubious ownership`, `could not read Username`** — git запущен не от владельца клона или
|
||||
учётные данные есть у другого пользователя: задайте `git_user` (на валидаторах — `root`).
|
||||
- **«найден другой контейнер агента»** — на хосте есть контейнер с похожим именем или образом, не совпадающий с `container_name`:
|
||||
проверьте имя (`docker ps -a`) и при необходимости удалите лишний контейнер вручную.
|
||||
- **`bin/validator-agent` не совпадает с суммой** — в репозитории устарел `bin/SHA256SUMS`: пересоберите бинарник и обновите сумму.
|
||||
- **Сборка падает на `FROM alpine:3.20` / `apk add`** — с валидатора нет доступа к Docker Hub / репозиториям Alpine.
|
||||
- **Старый образ другого имени остаётся** — сценарий чистит только образы `image_name`; образ с прежним именем удалите вручную (`docker rmi`).
|
||||
@@ -0,0 +1,15 @@
|
||||
# Запускать из каталога deploy/ansible (там же лежит этот файл).
|
||||
[defaults]
|
||||
inventory = inventory/hosts.yml
|
||||
roles_path = roles
|
||||
forks = 20
|
||||
retry_files_enabled = False
|
||||
interpreter_python = auto_silent
|
||||
callback_result_format = yaml
|
||||
|
||||
[ssh_connection]
|
||||
pipelining = True
|
||||
# accept-new: отпечаток нового хоста запоминается при первом подключении (без
|
||||
# интерактивного вопроса); изменившийся отпечаток известного хоста по-прежнему
|
||||
# приводит к ошибке.
|
||||
ssh_args = -o ControlMaster=auto -o ControlPersist=60s -o StrictHostKeyChecking=accept-new
|
||||
+29
@@ -0,0 +1,29 @@
|
||||
# Параметры запуска validator-agent. Скопируйте в validator-agent.env и заполните:
|
||||
# cp validator-agent.env.example validator-agent.env && chmod 600 validator-agent.env
|
||||
# Файл передаётся контейнеру как `docker run --env-file` (формат KEY=VALUE, без
|
||||
# кавычек и пробелов вокруг "="). Читает их deploy/docker/validator-agent/docker-entrypoint.sh.
|
||||
# VALIDATOR_AGENT_VALIDATOR_ID сюда НЕ пишется: сценарий подставляет имя хоста.
|
||||
|
||||
# Адрес control-api (обязательно). Для внешнего размещения — адрес, доступный
|
||||
# с валидаторов напрямую (через него же работает способ самопроверки control_api).
|
||||
VALIDATOR_AGENT_CONTROL_API_URL=https://control-api.example.com
|
||||
|
||||
# Токен агентов (CONTROL_API_AGENT_TOKEN на стороне control-api). Пусто — если
|
||||
# токен на control-api ещё не включён.
|
||||
CONTROL_API_AGENT_TOKEN=
|
||||
|
||||
# Способы самопроверки в порядке приоритета: ip_echo, control_api.
|
||||
# Самопроверка проходит, если адрес подтвердил любой способ. Без пробелов.
|
||||
VALIDATOR_AGENT_SELF_CHECK_METHODS=[control_api,ip_echo]
|
||||
# Таймаут одного способа, секунд.
|
||||
VALIDATOR_AGENT_SELF_CHECK_TIMEOUT_SECONDS=10
|
||||
|
||||
# Период опроса control-api, секунд.
|
||||
VALIDATOR_AGENT_POLL_INTERVAL_SECONDS=5
|
||||
|
||||
# Исходящие проверки.
|
||||
VALIDATOR_AGENT_HTTPS_TIMEOUT_SECONDS=10
|
||||
VALIDATOR_AGENT_ICMP_TIMEOUT_SECONDS=5
|
||||
VALIDATOR_AGENT_ICMP_COUNT=3
|
||||
VALIDATOR_AGENT_SSH_ENABLED=false
|
||||
VALIDATOR_AGENT_SSH_TIMEOUT_SECONDS=5
|
||||
@@ -0,0 +1,50 @@
|
||||
---
|
||||
# Параметры доставки validator-agent (не секреты). Любой из них можно
|
||||
# переопределить в командной строке: -e deploy_ref=<коммит|тег>.
|
||||
# Параметры запуска самого агента (адрес control-api, токен, способы
|
||||
# самопроверки, таймауты) лежат в env-файле, см. local_env_file.
|
||||
|
||||
# SSH: пользователь и ключ одинаковы на jump-хосте и на валидаторах. Путь к
|
||||
# ключу — на jump-хосте (ключ с правами 0600). Для docker и записи env-файла
|
||||
# сценарий повышает права через sudo (become).
|
||||
ansible_user: debian
|
||||
ansible_ssh_private_key_file: ~/.ssh/vk_cloud_priv.key
|
||||
|
||||
# --- git-клон на валидаторе ---------------------------------------------
|
||||
repo_dir: /opt/cloud-ip-validator
|
||||
repo_remote: origin
|
||||
# Ветка, тег или коммит, который нужно выкатить (откат: -e deploy_ref=<коммит>).
|
||||
deploy_ref: main
|
||||
# Пользователь, от имени которого выполняется git в клоне (через sudo): владелец
|
||||
# клона. На валидаторах клон принадлежит root, репозиторий читается без учётных
|
||||
# данных. Пусто — git работает от SSH-пользователя без sudo.
|
||||
git_user: root
|
||||
|
||||
# --- образ и контейнер --------------------------------------------------
|
||||
image_name: cloud-ip-validator-validator-agent
|
||||
# Имя контейнера на валидаторах (имя образа — другое, см. image_name).
|
||||
container_name: validator-agent
|
||||
platform: linux/amd64
|
||||
dockerfile: deploy/docker/validator-agent/Dockerfile
|
||||
# Сколько образов с метками ревизий хранить (кроме latest); старые удаляются.
|
||||
keep_images: 3
|
||||
|
||||
# --- запуск контейнера (как в deploy/docker/RUN.txt) --------------------
|
||||
restart_policy: unless-stopped
|
||||
capabilities: [NET_RAW]
|
||||
log_max_size: 10m
|
||||
log_max_file: "3"
|
||||
stop_timeout: 10
|
||||
|
||||
# --- проверка после запуска ---------------------------------------------
|
||||
verify_retries: 10
|
||||
verify_delay: 3
|
||||
|
||||
# --- env-файл на jump-хосте ---------------------------------------------
|
||||
# Параметры запуска агента. Рабочий файл создаётся из validator-agent.env.example
|
||||
# и в git не попадает. Файл можно зашифровать: ansible-vault encrypt <файл>
|
||||
# (тогда запускайте с --ask-vault-pass или --vault-password-file).
|
||||
local_env_file: "{{ playbook_dir }}/../env/validator-agent.env"
|
||||
# Куда файл копируется на валидатор (путь закрыт .gitignore репозитория,
|
||||
# переживает git reset).
|
||||
remote_env_file: "{{ repo_dir }}/deploy/docker/.env.validator"
|
||||
@@ -0,0 +1,69 @@
|
||||
# Валидаторы. Имя хоста (validator-N) — это имя ВМ; validator_id в control-api
|
||||
# другой (vkiplab-vN) и задаётся переменной validator_id. Соответствие
|
||||
# validator-N -> vkiplab-vN предполагается по порядку номеров; сценарий
|
||||
# сверяет validator_id с тем, что записано в работающем контейнере на хосте,
|
||||
# и при расхождении останавливается ДО замены контейнера.
|
||||
all:
|
||||
children:
|
||||
validators:
|
||||
hosts:
|
||||
validator-1:
|
||||
ansible_host: 10.11.12.161
|
||||
validator_id: vkiplab-v1
|
||||
validator-2:
|
||||
ansible_host: 10.11.12.177
|
||||
validator_id: vkiplab-v2
|
||||
validator-3:
|
||||
ansible_host: 10.11.12.33
|
||||
validator_id: vkiplab-v3
|
||||
validator-4:
|
||||
ansible_host: 10.11.12.41
|
||||
validator_id: vkiplab-v4
|
||||
validator-5:
|
||||
ansible_host: 10.11.12.193
|
||||
validator_id: vkiplab-v5
|
||||
validator-6:
|
||||
ansible_host: 10.11.12.197
|
||||
validator_id: vkiplab-v6
|
||||
validator-7:
|
||||
ansible_host: 10.11.12.198
|
||||
validator_id: vkiplab-v7
|
||||
validator-8:
|
||||
ansible_host: 10.11.12.169
|
||||
validator_id: vkiplab-v8
|
||||
validator-9:
|
||||
ansible_host: 10.11.12.185
|
||||
validator_id: vkiplab-v9
|
||||
validator-10:
|
||||
ansible_host: 10.11.12.186
|
||||
validator_id: vkiplab-v10
|
||||
validator-11:
|
||||
ansible_host: 10.11.12.199
|
||||
validator_id: vkiplab-v11
|
||||
validator-12:
|
||||
ansible_host: 10.11.12.196
|
||||
validator_id: vkiplab-v12
|
||||
validator-13:
|
||||
ansible_host: 10.11.12.194
|
||||
validator_id: vkiplab-v13
|
||||
validator-14:
|
||||
ansible_host: 10.11.12.195
|
||||
validator_id: vkiplab-v14
|
||||
validator-15:
|
||||
ansible_host: 10.11.12.65
|
||||
validator_id: vkiplab-v15
|
||||
validator-16:
|
||||
ansible_host: 10.11.12.189
|
||||
validator_id: vkiplab-v16
|
||||
validator-17:
|
||||
ansible_host: 10.11.12.190
|
||||
validator_id: vkiplab-v17
|
||||
validator-18:
|
||||
ansible_host: 10.11.12.168
|
||||
validator_id: vkiplab-v18
|
||||
validator-19:
|
||||
ansible_host: 10.11.12.191
|
||||
validator_id: vkiplab-v19
|
||||
validator-20:
|
||||
ansible_host: 10.11.12.73
|
||||
validator_id: vkiplab-v20
|
||||
@@ -0,0 +1,17 @@
|
||||
---
|
||||
# Доставка validator-agent на все валидаторы: обновить git-клон, собрать образ
|
||||
# на каждом хосте, остановить и удалить старый контейнер, поднять новый.
|
||||
# Запуск (из deploy/ansible): ansible-playbook playbooks/deploy-validator-agent.yml
|
||||
- name: Deliver validator-agent to the validators
|
||||
hosts: validators
|
||||
become: true
|
||||
gather_facts: false
|
||||
# Волны: один хост (канарейка), затем четыре, затем все остальные. Любой
|
||||
# сбой останавливает прогон — следующая волна не начнётся.
|
||||
# Все сразу: -e '{"deploy_serial": ["100%"]}'.
|
||||
serial: "{{ deploy_serial }}"
|
||||
max_fail_percentage: 0
|
||||
vars:
|
||||
deploy_serial: [1, 4, "100%"]
|
||||
roles:
|
||||
- validator_agent
|
||||
@@ -0,0 +1,40 @@
|
||||
---
|
||||
# Dockerfile ничего не компилирует: в образ копируется закоммиченный
|
||||
# bin/validator-agent. Не даём выкатить бинарник, не совпадающий с суммой.
|
||||
- name: Check bin/validator-agent against SHA256SUMS
|
||||
ansible.builtin.shell: |
|
||||
set -o pipefail
|
||||
grep -E '[[:space:]]validator-agent$' SHA256SUMS | sha256sum -c -
|
||||
args:
|
||||
chdir: "{{ repo_dir }}/bin"
|
||||
executable: /bin/bash
|
||||
changed_when: false
|
||||
|
||||
# Контекст сборки — корень репозитория (так и в SETUP.md). Слои кэшируются,
|
||||
# при смене bin/ образ пересобирается сам.
|
||||
- name: Build the image
|
||||
ansible.builtin.command:
|
||||
argv:
|
||||
- docker
|
||||
- build
|
||||
- --platform
|
||||
- "{{ platform }}"
|
||||
- --label
|
||||
- "git.rev={{ rev_after.stdout }}"
|
||||
- --label
|
||||
- deployed.by=ansible
|
||||
- -t
|
||||
- "{{ image_ref }}"
|
||||
- -f
|
||||
- "{{ dockerfile }}"
|
||||
- .
|
||||
chdir: "{{ repo_dir }}"
|
||||
|
||||
- name: Tag the image as latest
|
||||
ansible.builtin.command: "docker tag {{ image_ref }} {{ image_name }}:latest"
|
||||
|
||||
- name: Read the image id
|
||||
ansible.builtin.command:
|
||||
argv: [docker, image, inspect, --format, "{% raw %}{{.Id}}{% endraw %}", "{{ image_ref }}"]
|
||||
changed_when: false
|
||||
register: built_image
|
||||
@@ -0,0 +1,80 @@
|
||||
---
|
||||
# Команды git вместо модуля git: модуль переписывает URL remote и может
|
||||
# затереть учётные данные, уже настроенные в клоне. Работаем от git_user
|
||||
# (или от SSH-пользователя, если он не задан).
|
||||
- name: Remember the current revision
|
||||
ansible.builtin.command: "{{ git_cmd }} rev-parse HEAD"
|
||||
become: "{{ git_user | length > 0 }}"
|
||||
become_user: "{{ git_user }}"
|
||||
changed_when: false
|
||||
register: rev_before
|
||||
|
||||
- name: Fetch the remote
|
||||
ansible.builtin.command: "{{ git_cmd }} fetch --prune --tags {{ repo_remote }}"
|
||||
become: "{{ git_user | length > 0 }}"
|
||||
become_user: "{{ git_user }}"
|
||||
changed_when: false
|
||||
|
||||
# deploy_ref — ветка, тег или коммит. Ветка берётся из remote (свежая),
|
||||
# тег и коммит — как есть.
|
||||
- name: Resolve deploy_ref as a remote branch
|
||||
ansible.builtin.command: >-
|
||||
{{ git_cmd }} rev-parse --verify --quiet
|
||||
refs/remotes/{{ repo_remote }}/{{ deploy_ref }}^{commit}
|
||||
become: "{{ git_user | length > 0 }}"
|
||||
become_user: "{{ git_user }}"
|
||||
changed_when: false
|
||||
failed_when: false
|
||||
register: ref_branch
|
||||
|
||||
- name: Resolve deploy_ref as a tag or commit
|
||||
ansible.builtin.command: "{{ git_cmd }} rev-parse --verify --quiet {{ deploy_ref }}^{commit}"
|
||||
become: "{{ git_user | length > 0 }}"
|
||||
become_user: "{{ git_user }}"
|
||||
changed_when: false
|
||||
failed_when: false
|
||||
register: ref_other
|
||||
when: ref_branch.rc != 0
|
||||
|
||||
- name: Fail if deploy_ref does not exist
|
||||
ansible.builtin.assert:
|
||||
that: ref_branch.rc == 0 or (ref_other.rc | default(1)) == 0
|
||||
fail_msg: "deploy_ref={{ deploy_ref }} не найден в {{ repo_dir }} ({{ repo_remote }})."
|
||||
quiet: true
|
||||
|
||||
- name: Fix the target revision
|
||||
ansible.builtin.set_fact:
|
||||
target_rev: "{{ ref_branch.stdout if ref_branch.rc == 0 else ref_other.stdout }}"
|
||||
|
||||
# Ветка: остаёмся на локальной ветке (клон не уходит в detached HEAD),
|
||||
# сброс на remote. Тег или коммит: detached HEAD.
|
||||
- name: Check out the branch
|
||||
ansible.builtin.command: "{{ git_cmd }} checkout --force -B {{ deploy_ref }} {{ target_rev }}"
|
||||
become: "{{ git_user | length > 0 }}"
|
||||
become_user: "{{ git_user }}"
|
||||
when: ref_branch.rc == 0
|
||||
changed_when: rev_before.stdout != target_rev
|
||||
|
||||
- name: Check out the tag or commit
|
||||
ansible.builtin.command: "{{ git_cmd }} checkout --force --detach {{ target_rev }}"
|
||||
become: "{{ git_user | length > 0 }}"
|
||||
become_user: "{{ git_user }}"
|
||||
when: ref_branch.rc != 0
|
||||
changed_when: rev_before.stdout != target_rev
|
||||
|
||||
- name: Read the deployed revision
|
||||
ansible.builtin.command: "{{ git_cmd }} rev-parse HEAD"
|
||||
become: "{{ git_user | length > 0 }}"
|
||||
become_user: "{{ git_user }}"
|
||||
changed_when: false
|
||||
register: rev_after
|
||||
|
||||
- name: Check that the clone is at the target revision
|
||||
ansible.builtin.assert:
|
||||
that: rev_after.stdout == target_rev
|
||||
fail_msg: "Клон на {{ rev_after.stdout }}, ожидалось {{ target_rev }}."
|
||||
quiet: true
|
||||
|
||||
- name: Remember the short revision
|
||||
ansible.builtin.set_fact:
|
||||
deploy_rev: "{{ rev_after.stdout[:12] }}"
|
||||
@@ -0,0 +1,33 @@
|
||||
---
|
||||
# Порядок важен: сначала всё, что не трогает работающий контейнер (проверки,
|
||||
# обновление кода, сборка образа), и только потом замена контейнера. Если
|
||||
# что-то упало до replace, старый контейнер продолжает работать.
|
||||
- name: Preflight checks
|
||||
ansible.builtin.import_tasks: preflight.yml
|
||||
tags: [preflight]
|
||||
|
||||
- name: Dry run stops after preflight
|
||||
ansible.builtin.debug:
|
||||
msg: "check mode: git, build, replace and verify are skipped"
|
||||
when: ansible_check_mode
|
||||
tags: [always]
|
||||
|
||||
- name: Update the git clone
|
||||
ansible.builtin.import_tasks: git.yml
|
||||
when: not ansible_check_mode
|
||||
tags: [git]
|
||||
|
||||
- name: Build the image
|
||||
ansible.builtin.import_tasks: build.yml
|
||||
when: not ansible_check_mode
|
||||
tags: [build]
|
||||
|
||||
- name: Replace the container
|
||||
ansible.builtin.import_tasks: replace.yml
|
||||
when: not ansible_check_mode
|
||||
tags: [replace]
|
||||
|
||||
- name: Verify the new container
|
||||
ansible.builtin.import_tasks: verify.yml
|
||||
when: not ansible_check_mode
|
||||
tags: [verify]
|
||||
@@ -0,0 +1,150 @@
|
||||
---
|
||||
# --- на jump-хосте (один раз) -------------------------------------------
|
||||
- name: Check that the env file exists on the jump host
|
||||
ansible.builtin.stat:
|
||||
path: "{{ local_env_file }}"
|
||||
delegate_to: localhost
|
||||
become: false
|
||||
run_once: true
|
||||
check_mode: false
|
||||
register: env_file_stat
|
||||
|
||||
- name: Fail early without an env file
|
||||
ansible.builtin.assert:
|
||||
that: env_file_stat.stat.exists
|
||||
fail_msg: >-
|
||||
Нет env-файла {{ local_env_file }}. Создайте его:
|
||||
cp env/validator-agent.env.example env/validator-agent.env и заполните.
|
||||
quiet: true
|
||||
run_once: true
|
||||
|
||||
# Содержимое файла (в нём токен) не выводится: разбор идёт в задаче с no_log,
|
||||
# а проверка и её сообщение — по готовым булевым значениям.
|
||||
- name: Inspect the env file without printing it
|
||||
ansible.builtin.set_fact:
|
||||
env_url_set: "{{ env_file_text is regex('(?m)^VALIDATOR_AGENT_CONTROL_API_URL=\\S+') }}"
|
||||
env_url_is_example: "{{ env_file_text is regex('(?m)^VALIDATOR_AGENT_CONTROL_API_URL=\\S*example\\.com') }}"
|
||||
vars:
|
||||
env_file_text: "{{ lookup('ansible.builtin.file', local_env_file) }}"
|
||||
run_once: true
|
||||
no_log: true
|
||||
|
||||
- name: Check that the env file sets the control-api address
|
||||
ansible.builtin.assert:
|
||||
that:
|
||||
- env_url_set | bool
|
||||
- not (env_url_is_example | bool)
|
||||
fail_msg: >-
|
||||
В {{ local_env_file }} не задан VALIDATOR_AGENT_CONTROL_API_URL
|
||||
(или остался адрес-пример example.com).
|
||||
quiet: true
|
||||
run_once: true
|
||||
|
||||
# --- на каждом валидаторе -----------------------------------------------
|
||||
- name: Check that Docker answers
|
||||
ansible.builtin.command: docker version --format {% raw %}'{{.Server.Version}}'{% endraw %}
|
||||
changed_when: false
|
||||
check_mode: false
|
||||
|
||||
- name: Check that git is installed
|
||||
ansible.builtin.command: git --version
|
||||
changed_when: false
|
||||
check_mode: false
|
||||
|
||||
- name: Check that the git clone exists
|
||||
ansible.builtin.stat:
|
||||
path: "{{ repo_dir }}/.git"
|
||||
check_mode: false
|
||||
register: clone_stat
|
||||
|
||||
- name: Fail without a clone
|
||||
ansible.builtin.assert:
|
||||
that: clone_stat.stat.exists
|
||||
fail_msg: "Нет git-клона {{ repo_dir }} на {{ inventory_hostname }}."
|
||||
quiet: true
|
||||
|
||||
- name: Read the CPU architecture
|
||||
ansible.builtin.command: uname -m
|
||||
changed_when: false
|
||||
check_mode: false
|
||||
register: arch
|
||||
|
||||
- name: The image is linux/amd64 only
|
||||
ansible.builtin.assert:
|
||||
that: arch.stdout in ['x86_64', 'amd64']
|
||||
fail_msg: "Архитектура {{ arch.stdout }}: образ {{ platform }} здесь не запустится (exec format error)."
|
||||
quiet: true
|
||||
|
||||
- name: Look at the current container
|
||||
ansible.builtin.command: >-
|
||||
docker container inspect --format
|
||||
{% raw %}'{{.Config.Image}} {{.State.Status}}'{% endraw %}
|
||||
{{ container_name }}
|
||||
register: current_container
|
||||
changed_when: false
|
||||
failed_when: false
|
||||
check_mode: false
|
||||
|
||||
# Защита от второго агента: если на хосте уже есть другой контейнер этого
|
||||
# агента (по имени или по образу), сценарий остановится, а не запустит
|
||||
# рядом ещё один с тем же validator_id.
|
||||
- name: List containers on the validator
|
||||
ansible.builtin.command: docker ps -a --format {% raw %}'{{.Names}}|{{.Image}}'{% endraw %}
|
||||
register: all_containers
|
||||
changed_when: false
|
||||
check_mode: false
|
||||
|
||||
- name: Check that there is no other agent container
|
||||
ansible.builtin.assert:
|
||||
that: (other_agents | from_json) | length == 0
|
||||
fail_msg: >-
|
||||
{{ inventory_hostname }}: найден другой контейнер агента: {{ (other_agents | from_json) | join(', ') }}.
|
||||
Сценарий заменяет только контейнер {{ container_name }}. Проверьте container_name в
|
||||
inventory/group_vars/validators.yml или удалите лишний контейнер вручную.
|
||||
quiet: true
|
||||
vars:
|
||||
other_agents: >-
|
||||
{%- set found = [] -%}
|
||||
{%- for line in all_containers.stdout_lines -%}
|
||||
{%- set row = line.split('|') -%}
|
||||
{%- if row[0] != container_name and ('validator-agent' in row[0] or row[1] == image_name or row[1].startswith(image_name ~ ':')) -%}
|
||||
{%- set _ = found.append(row[0] ~ ' (' ~ row[1] ~ ')') -%}
|
||||
{%- endif -%}
|
||||
{%- endfor -%}
|
||||
{{- found | to_json -}}
|
||||
|
||||
# validator_id работающего контейнера — эталон: если он отличается от
|
||||
# inventory, заменять контейнер нельзя (агент зарегистрировался бы под чужим
|
||||
# именем, адреса привязывались бы к порту другой ВМ). Выводится только он,
|
||||
# а не все переменные окружения (там токен).
|
||||
- name: Read validator_id of the running container
|
||||
ansible.builtin.shell: |
|
||||
set -o pipefail
|
||||
docker container inspect --format '{% raw %}{{range .Config.Env}}{{println .}}{{end}}{% endraw %}' {{ container_name }} \
|
||||
| sed -n 's/^VALIDATOR_AGENT_VALIDATOR_ID=//p'
|
||||
args:
|
||||
executable: /bin/bash
|
||||
register: running_validator_id
|
||||
changed_when: false
|
||||
failed_when: false
|
||||
check_mode: false
|
||||
when: current_container.rc == 0
|
||||
|
||||
- name: Check validator_id against the running container
|
||||
ansible.builtin.assert:
|
||||
that: >-
|
||||
current_container.rc != 0
|
||||
or (running_validator_id.stdout | trim) == ''
|
||||
or (running_validator_id.stdout | trim) == effective_validator_id
|
||||
fail_msg: >-
|
||||
{{ inventory_hostname }}: в запущенном контейнере validator_id={{ running_validator_id.stdout | default('') | trim }},
|
||||
а в inventory {{ effective_validator_id }}. Проверьте соответствие имени ВМ и validator_id
|
||||
в inventory/hosts.yml; контейнер не тронут.
|
||||
quiet: true
|
||||
|
||||
- name: Report the current container
|
||||
ansible.builtin.debug:
|
||||
msg: >-
|
||||
{{ container_name }}:
|
||||
{{ current_container.stdout if current_container.rc == 0 else 'контейнера нет (будет создан)' }};
|
||||
validator_id для запуска: {{ effective_validator_id }}
|
||||
@@ -0,0 +1,45 @@
|
||||
---
|
||||
# Образ уже собран: простой валидатора — только stop + rm + run.
|
||||
- name: Copy the env file to the validator
|
||||
ansible.builtin.copy:
|
||||
src: "{{ local_env_file }}"
|
||||
dest: "{{ remote_env_file }}"
|
||||
owner: root
|
||||
group: root
|
||||
mode: "0600"
|
||||
no_log: true
|
||||
|
||||
- name: Check whether the container exists
|
||||
ansible.builtin.command: "docker container inspect {{ container_name }}"
|
||||
register: container_exists
|
||||
changed_when: false
|
||||
failed_when: false
|
||||
|
||||
- name: Stop the current container
|
||||
ansible.builtin.command: "docker stop -t {{ stop_timeout }} {{ container_name }}"
|
||||
when: container_exists.rc == 0
|
||||
|
||||
- name: Remove the current container
|
||||
ansible.builtin.command: "docker rm -f {{ container_name }}"
|
||||
when: container_exists.rc == 0
|
||||
|
||||
# --restart нужен: агент завершается, если регистрация в control-api не
|
||||
# удалась, и должен подняться снова. validator_id берётся из inventory
|
||||
# (validator_id) и перекрывает env-файл.
|
||||
- name: Start the new container
|
||||
ansible.builtin.command:
|
||||
argv: >-
|
||||
{{ ['docker', 'run', '-d',
|
||||
'--name', container_name,
|
||||
'--restart', restart_policy,
|
||||
'--platform', platform,
|
||||
'--env-file', remote_env_file,
|
||||
'-e', 'VALIDATOR_AGENT_VALIDATOR_ID=' ~ effective_validator_id,
|
||||
'--log-driver', 'json-file',
|
||||
'--log-opt', 'max-size=' ~ log_max_size,
|
||||
'--log-opt', 'max-file=' ~ log_max_file,
|
||||
'--label', 'git.rev=' ~ rev_after.stdout,
|
||||
'--label', 'deployed.by=ansible']
|
||||
+ (capabilities | map('regex_replace', '^(.*)$', '--cap-add=\1') | list)
|
||||
+ [image_ref] }}
|
||||
register: started
|
||||
@@ -0,0 +1,65 @@
|
||||
---
|
||||
- name: Verify the new container
|
||||
block:
|
||||
# Агент пишет "registered" после успешной регистрации в control-api.
|
||||
- name: Wait for the agent to register in control-api
|
||||
ansible.builtin.command: "docker logs --tail 200 {{ container_name }}"
|
||||
register: agent_logs
|
||||
changed_when: false
|
||||
until: agent_logs.stdout is search('msg=registered validator_id=' ~ effective_validator_id ~ '(\s|$)') or agent_logs.stderr is search('msg=registered validator_id=' ~ effective_validator_id ~ '(\s|$)')
|
||||
retries: "{{ verify_retries | int }}"
|
||||
delay: "{{ verify_delay | int }}"
|
||||
|
||||
- name: Inspect the container
|
||||
ansible.builtin.command:
|
||||
argv: [docker, inspect, --format, "{% raw %}{{.State.Running}} {{.RestartCount}} {{.Image}}{% endraw %}", "{{ container_name }}"]
|
||||
register: container_state
|
||||
changed_when: false
|
||||
|
||||
- name: Check the container state
|
||||
ansible.builtin.assert:
|
||||
that:
|
||||
- container_state.stdout.split()[0] == 'true'
|
||||
- container_state.stdout.split()[1] == '0'
|
||||
- container_state.stdout.split()[2] == built_image.stdout
|
||||
fail_msg: >-
|
||||
Контейнер {{ container_name }} в состоянии «{{ container_state.stdout }}»
|
||||
(ожидалось: запущен, 0 перезапусков, образ {{ built_image.stdout }}).
|
||||
quiet: true
|
||||
rescue:
|
||||
- name: Collect the container log
|
||||
ansible.builtin.command: "docker logs --tail 30 {{ container_name }}"
|
||||
register: failed_logs
|
||||
changed_when: false
|
||||
failed_when: false
|
||||
|
||||
- name: Fail the host and stop the next waves
|
||||
ansible.builtin.fail:
|
||||
msg: |-
|
||||
{{ inventory_hostname }}: контейнер не прошёл проверку после запуска.
|
||||
Последние строки лога:
|
||||
{{ failed_logs.stdout }}{{ failed_logs.stderr }}
|
||||
|
||||
# Старые образы с метками ревизий: оставляем keep_images последних (docker
|
||||
# выводит от новых к старым), образ работающего контейнера docker не удалит.
|
||||
- name: List image tags
|
||||
ansible.builtin.command:
|
||||
argv: [docker, images, "{{ image_name }}", --format, "{% raw %}{{.Tag}}{% endraw %}"]
|
||||
register: image_tags
|
||||
changed_when: false
|
||||
|
||||
- name: Remove old revision images
|
||||
ansible.builtin.command: "docker rmi {{ image_name }}:{{ item }}"
|
||||
loop: "{{ (image_tags.stdout_lines | reject('equalto', 'latest') | list)[keep_images | int:] }}"
|
||||
changed_when: true
|
||||
failed_when: false
|
||||
|
||||
- name: Remove dangling images
|
||||
ansible.builtin.command: docker image prune -f
|
||||
changed_when: false
|
||||
|
||||
- name: Summary
|
||||
ansible.builtin.debug:
|
||||
msg: >-
|
||||
{{ inventory_hostname }} ({{ effective_validator_id }}): {{ deploy_rev }} ({{ deploy_ref }}),
|
||||
образ {{ built_image.stdout[:19] }}, контейнер {{ container_name }} запущен
|
||||
@@ -0,0 +1,7 @@
|
||||
---
|
||||
# git с явным safe.directory: клон может принадлежать другому пользователю.
|
||||
git_cmd: "git -c safe.directory={{ repo_dir }} -C {{ repo_dir }}"
|
||||
# validator_id в control-api: из inventory, иначе имя хоста.
|
||||
effective_validator_id: "{{ validator_id | default(inventory_hostname) }}"
|
||||
# Образ с меткой ревизии, собранный в этом прогоне.
|
||||
image_ref: "{{ image_name }}:{{ deploy_rev | default('unknown') }}"
|
||||
@@ -100,6 +100,7 @@ services:
|
||||
VALIDATOR_AGENT_CONTROL_API_URL: "${VALIDATOR_AGENT_CONTROL_API_URL:-http://control-api:8080}"
|
||||
VALIDATOR_AGENT_POLL_INTERVAL_SECONDS: "${VALIDATOR_AGENT_POLL_INTERVAL_SECONDS:-5}"
|
||||
VALIDATOR_AGENT_SELF_CHECK_TIMEOUT_SECONDS: "${VALIDATOR_AGENT_SELF_CHECK_TIMEOUT_SECONDS:-10}"
|
||||
VALIDATOR_AGENT_SELF_CHECK_METHODS: "${VALIDATOR_AGENT_SELF_CHECK_METHODS:-[ip_echo]}"
|
||||
VALIDATOR_AGENT_HTTPS_TIMEOUT_SECONDS: "${VALIDATOR_AGENT_HTTPS_TIMEOUT_SECONDS:-10}"
|
||||
VALIDATOR_AGENT_ICMP_TIMEOUT_SECONDS: "${VALIDATOR_AGENT_ICMP_TIMEOUT_SECONDS:-5}"
|
||||
VALIDATOR_AGENT_ICMP_COUNT: "${VALIDATOR_AGENT_ICMP_COUNT:-3}"
|
||||
|
||||
@@ -6,6 +6,7 @@ set -eu
|
||||
|
||||
export VALIDATOR_AGENT_POLL_INTERVAL_SECONDS="${VALIDATOR_AGENT_POLL_INTERVAL_SECONDS:-5}"
|
||||
export VALIDATOR_AGENT_SELF_CHECK_TIMEOUT_SECONDS="${VALIDATOR_AGENT_SELF_CHECK_TIMEOUT_SECONDS:-10}"
|
||||
export VALIDATOR_AGENT_SELF_CHECK_METHODS="${VALIDATOR_AGENT_SELF_CHECK_METHODS:-[ip_echo]}"
|
||||
export VALIDATOR_AGENT_HTTPS_TIMEOUT_SECONDS="${VALIDATOR_AGENT_HTTPS_TIMEOUT_SECONDS:-10}"
|
||||
export VALIDATOR_AGENT_ICMP_TIMEOUT_SECONDS="${VALIDATOR_AGENT_ICMP_TIMEOUT_SECONDS:-5}"
|
||||
export VALIDATOR_AGENT_ICMP_COUNT="${VALIDATOR_AGENT_ICMP_COUNT:-3}"
|
||||
@@ -16,7 +17,7 @@ export VALIDATOR_AGENT_SSH_TIMEOUT_SECONDS="${VALIDATOR_AGENT_SSH_TIMEOUT_SECOND
|
||||
# NOT templated into the YAML: the binary reads them straight from this
|
||||
# container's environment (names default in the config loader), so secrets
|
||||
# never land in a file inside the container.
|
||||
envsubst '${VALIDATOR_AGENT_VALIDATOR_ID} ${VALIDATOR_AGENT_CONTROL_API_URL} ${VALIDATOR_AGENT_POLL_INTERVAL_SECONDS} ${VALIDATOR_AGENT_SELF_CHECK_TIMEOUT_SECONDS} ${VALIDATOR_AGENT_HTTPS_TIMEOUT_SECONDS} ${VALIDATOR_AGENT_ICMP_TIMEOUT_SECONDS} ${VALIDATOR_AGENT_ICMP_COUNT} ${VALIDATOR_AGENT_SSH_ENABLED} ${VALIDATOR_AGENT_SSH_TIMEOUT_SECONDS}' \
|
||||
envsubst '${VALIDATOR_AGENT_VALIDATOR_ID} ${VALIDATOR_AGENT_CONTROL_API_URL} ${VALIDATOR_AGENT_POLL_INTERVAL_SECONDS} ${VALIDATOR_AGENT_SELF_CHECK_TIMEOUT_SECONDS} ${VALIDATOR_AGENT_SELF_CHECK_METHODS} ${VALIDATOR_AGENT_HTTPS_TIMEOUT_SECONDS} ${VALIDATOR_AGENT_ICMP_TIMEOUT_SECONDS} ${VALIDATOR_AGENT_ICMP_COUNT} ${VALIDATOR_AGENT_SSH_ENABLED} ${VALIDATOR_AGENT_SSH_TIMEOUT_SECONDS}' \
|
||||
< /etc/validator-agent/validator-agent.yaml.tmpl > /etc/validator-agent/validator-agent.yaml
|
||||
|
||||
exec /usr/local/bin/validator-agent -config /etc/validator-agent/validator-agent.yaml
|
||||
@@ -4,6 +4,7 @@ poll_interval_seconds: ${VALIDATOR_AGENT_POLL_INTERVAL_SECONDS}
|
||||
|
||||
self_check:
|
||||
timeout_seconds: ${VALIDATOR_AGENT_SELF_CHECK_TIMEOUT_SECONDS}
|
||||
methods: ${VALIDATOR_AGENT_SELF_CHECK_METHODS}
|
||||
|
||||
checks:
|
||||
https_timeout_seconds: ${VALIDATOR_AGENT_HTTPS_TIMEOUT_SECONDS}
|
||||
|
||||
@@ -0,0 +1,248 @@
|
||||
# Ручная очистка базы данных control-api (SQL)
|
||||
|
||||
Когда нужна: подготовка к новому полному прогону, разбор после инцидента, освобождение места, удаление отдельных адресов.
|
||||
Все примеры проверены 2026-10-02 на копии боевой БД (20 валидаторов, 4 площадки, 6445 адресов в реестре, 29 889 проверок,
|
||||
19 897 событий): без ошибок, `integrity_check` = `ok`, нарушений внешних ключей нет.
|
||||
|
||||
> **Сначала API.** Если в очереди есть адреса в работе, очищайте очередь штатно: кнопка «Очистить всё» на `/ips` или
|
||||
> `POST /api/v1/admin/ips/clear` (см. [USAGE.md](USAGE.md#удаление-адресов-из-очереди)). Только так control-api отвяжет Floating IP от портов
|
||||
> валидаторов. SQL ниже работает с базой «в покое»: он не обращается к OpenStack.
|
||||
|
||||
## 1. Что в базе и что нельзя трогать
|
||||
|
||||
| Группа | Таблицы | Можно чистить |
|
||||
|---|---|---|
|
||||
| Данные прогона | `ip_queue` (очередь), `ip_registry` (реестр адресов), `checks` (реестр проверок), `ip_site_checks` (признаки площадок по адресам в работе), `events` (журнал событий) | да |
|
||||
| Настройки (не трогать) | `validators`, `sites`, `target_groups` (цели), `check_types`, `inbound_checks_settings`, `settings`, `auto_cycle` | **нет** |
|
||||
| Служебное | `sqlite_sequence` (нумерация записей), `PRAGMA user_version` (версия схемы) | нумерацию можно сбросить, версию не менять |
|
||||
|
||||
Связи (внешние ключи): `checks`, `events`, `ip_site_checks` ссылаются на `ip_queue`; `checks`, `events`, `ip_queue` — на `ip_registry`;
|
||||
`validators.current_ip_id` — на `ip_queue`. Поэтому порядок удаления всегда такой: сначала `validators.current_ip_id` в `NULL`, затем
|
||||
`ip_site_checks`, `checks`, `events`, `ip_queue` и в конце `ip_registry`. Времена в БД хранятся строками `2026-10-02T07:14:58.857Z` (UTC).
|
||||
|
||||
## 2. Подготовка (всегда, перед любой очисткой)
|
||||
|
||||
**2.1. Где лежит база и чем её открывать.** Нужен клиент `sqlite3` на хосте (`apt install sqlite3`).
|
||||
|
||||
| Развёртывание | Файл базы |
|
||||
|---|---|
|
||||
| `rxprod-compose/` (боевой стенд) | `rxprod-compose/capi-db/control-api.db` |
|
||||
| Docker (`deploy/docker`) | в томе `cloud-ip-validator-db`: `docker volume inspect cloud-ip-validator-db --format '{{.Mountpoint}}'`, файл `control-api.db` в этом каталоге (нужен root) |
|
||||
| systemd | `/var/lib/cloud-ip-validator/control-api.db` (`database.path` в `control-api.yaml`) |
|
||||
|
||||
Дальше в примерах `DB=путь/к/control-api.db`.
|
||||
|
||||
**2.2. Проверить, что в очереди ничего не в работе:**
|
||||
|
||||
```bash
|
||||
sqlite3 -readonly "$DB" "
|
||||
SELECT state, COUNT(*) FROM ip_queue
|
||||
WHERE state NOT IN ('done', 'failed', 'occupied') GROUP BY state;"
|
||||
```
|
||||
|
||||
Пустой вывод — всё завершено. Адреса в работе есть — сначала «Очистить всё» через API (см. выше). Затем проверьте в OpenStack, что на портах
|
||||
валидаторов нет лишних Floating IP; по базе видно только то, что control-api считает привязанным:
|
||||
|
||||
```bash
|
||||
sqlite3 -readonly "$DB" "
|
||||
SELECT ip_address, state, fip_id FROM ip_queue
|
||||
WHERE fip_id <> '' AND state NOT IN ('done', 'failed', 'occupied');"
|
||||
```
|
||||
|
||||
**2.3. Остановить control-api.** Он держит базу открытой и пишет в неё на каждом такте; ручные правки поверх работающего процесса
|
||||
ненадёжны.
|
||||
|
||||
```bash
|
||||
cd rxprod-compose && docker compose stop control-api # compose-развёртывание
|
||||
sudo systemctl stop control-api # systemd
|
||||
```
|
||||
|
||||
**2.4. Сделать резервную копию** (штатной командой SQLite, не `cp`: у базы есть журнал `-wal`):
|
||||
|
||||
```bash
|
||||
TS=$(date -u +%Y-%m-%d_%H-%M)
|
||||
sqlite3 "$DB" ".backup '$(dirname "$DB")/backup-$TS-before-cleanup.db'"
|
||||
sqlite3 -readonly "$(dirname "$DB")/backup-$TS-before-cleanup.db" "PRAGMA integrity_check;" # должно быть: ok
|
||||
```
|
||||
|
||||
Лог контейнера до чистки при необходимости сохраните отдельно: `docker logs <контейнер> > backup-$TS.container.log 2>&1`.
|
||||
|
||||
**2.5. Посмотреть, что и сколько лежит** (до и после чистки):
|
||||
|
||||
```bash
|
||||
sqlite3 -readonly "$DB" "
|
||||
SELECT 'ip_queue' AS tbl, COUNT(*) AS n FROM ip_queue
|
||||
UNION ALL SELECT 'ip_registry', COUNT(*) FROM ip_registry
|
||||
UNION ALL SELECT 'checks', COUNT(*) FROM checks
|
||||
UNION ALL SELECT 'ip_site_checks', COUNT(*) FROM ip_site_checks
|
||||
UNION ALL SELECT 'events', COUNT(*) FROM events
|
||||
UNION ALL SELECT 'validators', COUNT(*) FROM validators
|
||||
UNION ALL SELECT 'sites', COUNT(*) FROM sites;"
|
||||
```
|
||||
|
||||
## 3. Сценарии
|
||||
|
||||
Команды выполняются так: `sqlite3 "$DB"` и вставить блок, либо сохранить блок в файл и выполнить `sqlite3 "$DB" < файл.sql`.
|
||||
|
||||
### 3.1. Полный сброс данных прогона (перед новым полным прогоном)
|
||||
|
||||
Очищает очередь, реестр адресов, реестр проверок, журнал событий. Валидаторы, площадки, цели, типы проверок и все настройки остаются.
|
||||
|
||||
```sql
|
||||
PRAGMA foreign_keys = ON;
|
||||
BEGIN;
|
||||
-- валидаторы больше не ссылаются на адреса очереди
|
||||
UPDATE validators SET current_ip_id = NULL WHERE current_ip_id IS NOT NULL;
|
||||
UPDATE validators SET state = 'idle' WHERE state = 'assigned';
|
||||
-- порядок важен: сначала зависимые таблицы
|
||||
DELETE FROM ip_site_checks;
|
||||
DELETE FROM checks;
|
||||
DELETE FROM events;
|
||||
DELETE FROM ip_queue;
|
||||
DELETE FROM ip_registry;
|
||||
-- нумерация снова с 1 (необязательно)
|
||||
DELETE FROM sqlite_sequence WHERE name IN ('ip_registry', 'checks', 'ip_queue', 'events');
|
||||
COMMIT;
|
||||
```
|
||||
|
||||
Затем освободите место (отдельной командой, не внутри транзакции):
|
||||
|
||||
```sql
|
||||
PRAGMA wal_checkpoint(TRUNCATE);
|
||||
VACUUM;
|
||||
```
|
||||
|
||||
Файл сжимается до сотен килобайт (на проверочной копии: 20 МБ → 128 КБ).
|
||||
|
||||
### 3.2. Только журнал событий
|
||||
|
||||
Очередь, реестр и проверки не затрагиваются.
|
||||
|
||||
Весь журнал:
|
||||
|
||||
```sql
|
||||
DELETE FROM events;
|
||||
DELETE FROM sqlite_sequence WHERE name = 'events';
|
||||
```
|
||||
|
||||
Только старше 7 дней (число дней меняйте в `'-7 days'`):
|
||||
|
||||
```sql
|
||||
DELETE FROM events
|
||||
WHERE occurred_at < strftime('%Y-%m-%dT%H:%M:%fZ', 'now', '-7 days');
|
||||
```
|
||||
|
||||
### 3.3. Только реестр проверок (история проверок)
|
||||
|
||||
Адреса и их итоги в очереди остаются; пропадает подробная история проверок (на странице «Реестр» исчезнут результаты).
|
||||
|
||||
Вся история:
|
||||
|
||||
```sql
|
||||
DELETE FROM checks;
|
||||
DELETE FROM sqlite_sequence WHERE name = 'checks';
|
||||
```
|
||||
|
||||
Оставить последние 3 цикла каждого адреса (число `3` меняйте):
|
||||
|
||||
```sql
|
||||
DELETE FROM checks
|
||||
WHERE cycle_id <= (SELECT MAX(c2.cycle_id) FROM checks c2 WHERE c2.registry_id = checks.registry_id) - 3;
|
||||
```
|
||||
|
||||
Постоянное ограничение глубины истории лучше задать настройкой `history_retention_cycles` (страница `/settings` или
|
||||
`PUT /api/v1/admin/config/orchestrator`): control-api сам подрезает историю при завершении каждого адреса. SQL выше нужен для разовой чистки.
|
||||
|
||||
### 3.4. Удалить конкретные адреса целиком
|
||||
|
||||
Удаляет адрес из очереди и реестра вместе со всей его историей (проверки и события). Список адресов подставьте в первую команду
|
||||
`CREATE TEMP TABLE doomed_reg`. Адрес, который сейчас проверяется, удалять этим способом нельзя: используйте API (раздел 5).
|
||||
|
||||
```sql
|
||||
PRAGMA foreign_keys = ON;
|
||||
BEGIN;
|
||||
CREATE TEMP TABLE doomed_reg AS
|
||||
SELECT id FROM ip_registry WHERE ip_address IN ('5.188.140.6', '5.188.140.62'); -- ваши адреса
|
||||
CREATE TEMP TABLE doomed_ip AS
|
||||
SELECT id FROM ip_queue WHERE registry_id IN (SELECT id FROM doomed_reg);
|
||||
UPDATE validators SET current_ip_id = NULL WHERE current_ip_id IN (SELECT id FROM doomed_ip);
|
||||
DELETE FROM ip_site_checks WHERE ip_id IN (SELECT id FROM doomed_ip);
|
||||
DELETE FROM checks WHERE registry_id IN (SELECT id FROM doomed_reg);
|
||||
DELETE FROM events WHERE registry_id IN (SELECT id FROM doomed_reg) OR ip_id IN (SELECT id FROM doomed_ip);
|
||||
DELETE FROM ip_queue WHERE id IN (SELECT id FROM doomed_ip);
|
||||
DELETE FROM ip_registry WHERE id IN (SELECT id FROM doomed_reg);
|
||||
DROP TABLE doomed_ip;
|
||||
DROP TABLE doomed_reg;
|
||||
COMMIT;
|
||||
```
|
||||
|
||||
### 3.5. Только освободить место
|
||||
|
||||
Если удалили много, а файл не уменьшился (SQLite не отдаёт место ОС до `VACUUM`):
|
||||
|
||||
```sql
|
||||
PRAGMA wal_checkpoint(TRUNCATE);
|
||||
VACUUM;
|
||||
```
|
||||
|
||||
Размер и свободные страницы:
|
||||
|
||||
```sql
|
||||
SELECT page_count * page_size / 1024 AS size_kb, freelist_count * page_size / 1024 AS free_kb
|
||||
FROM pragma_page_count(), pragma_page_size(), pragma_freelist_count();
|
||||
```
|
||||
|
||||
## 4. После очистки: запуск и проверка
|
||||
|
||||
```bash
|
||||
cd rxprod-compose && docker compose up -d --no-deps control-api # или: sudo systemctl start control-api
|
||||
```
|
||||
|
||||
Если нужно, чтобы и лог контейнера начался с нуля, пересоздайте контейнер: `docker compose up -d --force-recreate --no-deps control-api`
|
||||
(старый лог сохраните заранее, п. 2.4).
|
||||
|
||||
Проверка базы (до запуска или на копии):
|
||||
|
||||
```bash
|
||||
sqlite3 -readonly "$DB" "
|
||||
PRAGMA integrity_check;
|
||||
PRAGMA foreign_key_check;
|
||||
SELECT state, COUNT(*) FROM validators GROUP BY state;
|
||||
SELECT COUNT(*) AS validators_with_address FROM validators WHERE current_ip_id IS NOT NULL;
|
||||
SELECT COUNT(*) AS queue_rows FROM ip_queue;"
|
||||
```
|
||||
|
||||
Ожидается: `ok`; пустой результат `foreign_key_check`; у валидаторов состояние `idle`; `validators_with_address` = 0; для полного сброса `queue_rows` = 0.
|
||||
Через API: `GET /api/v1/admin/status` (`total_ips` = 0, `total_validators` = число валидаторов) и `GET /api/v1/admin/validators`
|
||||
(через 10–15 секунд после запуска у всех свежий `last_heartbeat_at`).
|
||||
|
||||
## 5. Что делать через API, а не через SQL
|
||||
|
||||
| Задача | Как |
|
||||
|---|---|
|
||||
| Остановить проверку адреса, удалить адрес в работе | `POST /api/v1/admin/ips/{ip}/cancel`, `DELETE /api/v1/admin/ips/{ip}` |
|
||||
| Очистить всю очередь с отвязкой Floating IP | `POST /api/v1/admin/ips/clear` |
|
||||
| Перепроверить завершённые адреса | `POST /api/v1/admin/ips` со списком адресов |
|
||||
| Валидатор «завис» с адресом | ничего не править: лизинг истечёт, адрес вернётся в очередь, валидатор освободится сам |
|
||||
| Добавить/убрать валидатор, площадку, цель | `/api/v1/admin/config/*` или страницы дашборда |
|
||||
|
||||
## 6. Восстановление из копии
|
||||
|
||||
```bash
|
||||
cd rxprod-compose && docker compose stop control-api
|
||||
cp capi-db/backup-<метка>-before-cleanup.db capi-db/control-api.db
|
||||
rm -f capi-db/control-api.db-wal capi-db/control-api.db-shm # старый журнал к новой копии не относится
|
||||
docker compose up -d --no-deps control-api
|
||||
```
|
||||
|
||||
## 7. Ловушки
|
||||
|
||||
- Не выполняйте `DELETE` при работающем control-api: он пишет в ту же базу.
|
||||
- Не удаляйте строки настроек (`validators`, `sites`, `target_groups`, `check_types`, `inbound_checks_settings`, `settings`, `auto_cycle`):
|
||||
после этого control-api либо не стартует, либо работает без площадок и целей.
|
||||
- Не нарушайте порядок удаления (раздел 1) и не отключайте `PRAGMA foreign_keys = ON` в блоках выше: база сама остановит ошибочное удаление.
|
||||
- `VACUUM` нельзя вызывать внутри транзакции и пока control-api запущен.
|
||||
- Не копируйте файл базы командой `cp` при работающем процессе: журнал `-wal` останется в неконсистентном состоянии. Используйте `.backup`.
|
||||
- Не меняйте `PRAGMA user_version`: по нему control-api применяет миграции схемы.
|
||||
- Не правьте состояние адресов и валидаторов вручную (`state`, `owner_validator_id`, `lease_expires_at`) вместо API: control-api сверяет их на каждом такте и исправит расхождение,
|
||||
но до этого результат непредсказуем.
|
||||
+32
-4
@@ -41,10 +41,10 @@ JSON, базовый префикс прикладных методов — `/ap
|
||||
|---|---|---|
|
||||
| **admin** | `CONTROL_API_ADMIN_TOKEN` | все `/api/v1/admin/*` (очередь, реестр, автоцикл, `config/*`) |
|
||||
| **agent** | `CONTROL_API_AGENT_TOKEN` | запись результатов: `POST /agents/{id}/self-check`, `/events`, `/results`, `/complete` и `POST /probers/{site_id}/results` |
|
||||
| **открыто** | — | `GET /healthz`; `POST /agents/register`, `POST /agents/{id}/heartbeat`, `GET /agents/{id}/assignment`; `POST /probers/register`, `POST /probers/{site_id}/heartbeat`, `GET /probers/{site_id}/assignments` |
|
||||
| **открыто** | — | `GET /healthz`; `POST /agents/register`, `POST /agents/{id}/heartbeat`, `GET /agents/{id}/assignment`, `GET /agents/{id}/observed-ip`; `POST /probers/register`, `POST /probers/{site_id}/heartbeat`, `GET /probers/{site_id}/assignments` |
|
||||
|
||||
- Токены разные: токен администратора **не** подходит для методов агентов, и наоборот.
|
||||
- Валидатор и пробер могут без токена зарегистрироваться, слать heartbeat и забирать задание (настройку); отправка результатов без токена агентов — `401`.
|
||||
- Валидатор и пробер могут без токена зарегистрироваться, слать heartbeat и забирать задание (настройку); валидатор также может спросить, с какого адреса его видит control-api (`observed-ip`); отправка результатов без токена агентов — `401`.
|
||||
- Имена переменных меняются в секции `auth` конфига control-api (`admin_token_env`, `agent_token_env`); сами значения в YAML не хранятся.
|
||||
- Ответ при отказе: `401 {"error": "unauthorized"}` с заголовком `WWW-Authenticate: Bearer`.
|
||||
- Токен не задан (пустая переменная) — уровень открыт; в логе control-api при старте предупреждение. Токены нужно генерировать случайными: `openssl rand -hex 32`.
|
||||
@@ -145,6 +145,30 @@ curl -s -H "Authorization: Bearer $ADMIN_TOKEN" http://<control-api>:8080/api/v1
|
||||
`check_config` — уже развёрнутая конфигурация проверок (тип + список
|
||||
целей), агенту не нужно самому сопоставлять группы целей.
|
||||
|
||||
### `GET /api/v1/agents/{id}/observed-ip`
|
||||
|
||||
С какого адреса control-api видит соединение валидатора. Используется
|
||||
self-check способом `control_api` (`self_check.methods` в
|
||||
`validator-agent.yaml`) как альтернатива внешнему IP-echo сервису. Уровень
|
||||
доступа — открыто: отдаётся только адрес самого вызывающего.
|
||||
|
||||
Ответ `200`:
|
||||
```json
|
||||
{"ip": "203.0.113.10", "source": "remote_addr"}
|
||||
```
|
||||
|
||||
- Адрес берётся только из адреса TCP-соединения (`RemoteAddr`), приведённого
|
||||
к каноничному виду (`::ffff:1.2.3.4` → `1.2.3.4`). Заголовки
|
||||
`X-Forwarded-For` / `X-Real-IP` **не учитываются**: иначе валидатор мог бы
|
||||
подделать адрес и пройти проверку. Метод рассчитан на прямое подключение
|
||||
без обратного прокси.
|
||||
- `404`, если `validator_id` не зарегистрирован.
|
||||
- Способ корректен, только если соединение выходит через внешнюю сеть
|
||||
(SNAT Floating IP). Если control-api достижим из облака по внутренней
|
||||
сети, он увидит приватный адрес валидатора. Если порт control-api
|
||||
опубликован через Docker, проверьте, что ручка показывает внешний адрес
|
||||
клиента, а не адрес шлюза Docker.
|
||||
|
||||
### `POST /api/v1/agents/{id}/self-check`
|
||||
|
||||
Отчёт о результате self-check — подтверждение, что исходящий трафик
|
||||
@@ -152,7 +176,11 @@ curl -s -H "Authorization: Bearer $ADMIN_TOKEN" http://<control-api>:8080/api/v1
|
||||
определяет это **сам**, обращаясь к внешнему (снаружи облака) IP-echo
|
||||
сервису (`self_check.ip_echo_urls` в `validator-agent.yaml`, например
|
||||
`api.ipify.org`) и сравнивая ответ с `ip_address` из задания — control-api
|
||||
в этом определении не участвует. Важно, что ресурс должен быть именно
|
||||
в этом определении не участвует. Дополнительно можно включить способ
|
||||
`control_api` (`self_check.methods`): агент спрашивает у control-api через
|
||||
`GET /agents/{id}/observed-ip`, с какого адреса тот его видит. Способы
|
||||
пробуются по приоритету, достаточно подтверждения любым из них; без
|
||||
настройки работает только IP-echo. Важно, что ресурс должен быть именно
|
||||
внешним: OpenStack применяет SNAT через Floating IP только к трафику,
|
||||
уходящему через внешнюю сеть, поэтому обращение к чему-либо внутри
|
||||
проекта (в том числе к самому control-api, если он в той же внутренней
|
||||
@@ -165,7 +193,7 @@ curl -s -H "Authorization: Bearer $ADMIN_TOKEN" http://<control-api>:8080/api/v1
|
||||
"ip_id": 42,
|
||||
"detected_egress_ip": "203.0.113.10",
|
||||
"success": true,
|
||||
"detail": "matched"
|
||||
"detail": "matched (control_api)"
|
||||
}
|
||||
```
|
||||
|
||||
|
||||
+18
-5
@@ -701,10 +701,17 @@ docker run -d --platform linux/amd64 --cap-add NET_RAW --name validator-agent \
|
||||
| `VALIDATOR_AGENT_SSH_TIMEOUT_SECONDS` | нет | `5` |
|
||||
|
||||
`--cap-add NET_RAW` обязателен для ICMP-проверок, как и у `prober`.
|
||||
`self_check.ip_echo_urls` в переменные не вынесен — при отсутствии в
|
||||
конфиге агент сам подставляет дефолт (`api.ipify.org`, `ifconfig.me`);
|
||||
свой список задавайте через смонтированный конфиг вместо шаблона, если
|
||||
нужно переопределить.
|
||||
`self_check.ip_echo_urls` и `self_check.methods` в переменные не вынесены —
|
||||
при отсутствии в конфиге агент сам подставляет дефолты (`api.ipify.org`,
|
||||
`ifconfig.me` и `methods: [ip_echo]`); свои значения задавайте через
|
||||
смонтированный конфиг вместо шаблона, если нужно переопределить.
|
||||
`methods` — способы самопроверки в порядке приоритета (`ip_echo`,
|
||||
`control_api`), достаточно подтверждения любым. Способ `control_api`
|
||||
спрашивает у control-api, с какого адреса он видит валидатора
|
||||
(`GET /agents/{id}/observed-ip`); при внешнем размещении control-api
|
||||
рекомендуется `[control_api, ip_echo]`. Ограничение: если control-api
|
||||
достижим из облака по внутренней сети, он увидит приватный адрес
|
||||
валидатора и этот способ всегда даст несовпадение — используйте `ip_echo`.
|
||||
|
||||
### Обновление образов после изменения кода
|
||||
|
||||
@@ -732,6 +739,10 @@ docker compose up -d --build
|
||||
сборке) и `docker rm -f <имя> && docker run ... ` (или `docker restart`,
|
||||
если менялись только переменные окружения, а не сам бинарник/образ).
|
||||
|
||||
Для массового обновления валидаторов (git-клон, сборка образа на хосте,
|
||||
замена контейнера, проверка регистрации) есть Ansible-сценарий:
|
||||
[`deploy/ansible/`](../deploy/ansible/README.md).
|
||||
|
||||
### Диагностика Docker-развёртывания
|
||||
|
||||
- **Контейнер сразу падает, в логах `exec format error`** — образ собран
|
||||
@@ -804,7 +815,9 @@ curl -s http://<control-api>:8080/api/v1/admin/validators | python3 -m json.tool
|
||||
через floating IP, который в данный момент привязан к валидатору.
|
||||
- `validator-agent` → внешние IP-echo сервисы из `self_check.ip_echo_urls`
|
||||
(по умолчанию `api.ipify.org`, `ifconfig.me`) — **обязательно вне
|
||||
облака**: это и есть механизм self-check (см.
|
||||
облака**: это и есть механизм self-check способом `ip_echo` (при
|
||||
`self_check.methods` с `control_api` достаточно ещё и доступа к control-api
|
||||
по внешней сети; см.
|
||||
[DIAGRAMS.md](DIAGRAMS.md#2-поток-данных-от-валидатора-к-целевому-серверу-egress-проверка)).
|
||||
Если валидатор не может достучаться ни до одного из этих адресов,
|
||||
self-check никогда не пройдёт и IP будет бесконечно возвращаться в
|
||||
|
||||
+26
-2
@@ -408,6 +408,17 @@ curl -s http://<control-api>:8080/api/v1/admin/validators | python3 -m json.tool
|
||||
(занят), `unreachable` (пропустил heartbeat дольше
|
||||
`orchestrator.heartbeat_timeout_seconds`).
|
||||
|
||||
Правила, которые держат состояние валидатора согласованным:
|
||||
- валидатор держит **не более одного адреса**; адрес освобождает валидатор
|
||||
только пока он остаётся его текущим (запоздалое завершение старого адреса
|
||||
чужого валидатора не освобождает);
|
||||
- после пропущенного heartbeat валидатор, у которого есть адрес, возвращается
|
||||
в `assigned`, а не в `idle`, и не получает второй адрес; без адреса — в `idle`;
|
||||
- `unreachable`-валидатор не получает адресов, пока не пришлёт heartbeat
|
||||
(даже если лизинг его адреса истёк и адрес вернулся в очередь);
|
||||
- на каждом такте оркестратор сверяет валидаторы с очередью и исправляет
|
||||
расхождения (в логе `repaired validators that disagreed with the queue`).
|
||||
|
||||
**Добавление нового валидатора (без перезапуска control-api):**
|
||||
1. Поднимите новую ВМ в сервисном проекте облака, узнайте её Neutron
|
||||
`port_id`.
|
||||
@@ -690,6 +701,12 @@ curl -s -X POST http://<control-api>:8080/api/v1/admin/ips/delete \
|
||||
curl -s -X POST http://<control-api>:8080/api/v1/admin/ips/clear
|
||||
```
|
||||
|
||||
Очистка отвязывает Floating IP только у адресов, которые ещё в работе
|
||||
(завершённые уже свободны), затем сама опрашивает порты валидаторов в облаке и
|
||||
снимает оставшиеся привязки адресов из реестра. Занимает секунды. Она не
|
||||
прерывается разрывом соединения (таймаутом клиента или дашборда): операция
|
||||
доводится до конца на стороне control-api, предел — 10 минут.
|
||||
|
||||
В `admin-dashboard` то же самое доступно на странице `/ips`: чекбоксы у
|
||||
каждой строки + кнопка «Удалить выбранные» для точечного/массового
|
||||
удаления, кнопка «Удалить» в каждой строке, и отдельная кнопка «Очистить
|
||||
@@ -742,7 +759,9 @@ curl -s -X PUT http://<control-api>:8080/api/v1/admin/config/orchestrator \
|
||||
дело обычно в self-check: он запрашивает внешние (вне облака) сервисы из
|
||||
`self_check.ip_echo_urls` в конфиге валидатора (по умолчанию
|
||||
`api.ipify.org`, `ifconfig.me`) — если у ВМ-валидатора нет исходящего
|
||||
доступа в интернет к этим адресам, запрос не проходит вообще, и агент
|
||||
доступа в интернет к этим адресам, запрос не проходит вообще (при
|
||||
`self_check.methods: [control_api, ip_echo]` агент сперва спросит адрес у
|
||||
control-api, и проверка может пройти и без IP-echo), и агент
|
||||
даже не может *сообщить* результат control-api (ни успешный, ни
|
||||
неуспешный) — тогда статус реально зависает до истечения
|
||||
`orchestrator.lease_ttl_seconds`, после чего адрес возвращается в
|
||||
@@ -763,7 +782,12 @@ https://api.ipify.org`) и логи `journalctl -u validator-agent` на пре
|
||||
облака (см. `self_check.ip_echo_urls`) — запрос к чему-либо внутри
|
||||
проекта (в том числе к самому control-api, если он в той же внутренней
|
||||
сети) покажет приватный адрес валидатора независимо от того, правильно
|
||||
ли привязан FIP, и всегда будет давать ложный провал.
|
||||
ли привязан FIP, и всегда будет давать ложный провал. Это относится и к
|
||||
способу `control_api` (`self_check.methods`): он корректен только когда
|
||||
валидатор ходит к control-api через внешнюю сеть; при внутреннем доступе в
|
||||
`detail` будет подсказка про приватный адрес — оставьте `ip_echo`. В
|
||||
`detail` события `self_check_result` указан сработавший способ
|
||||
(`matched (control_api)`) либо причина по каждому способу.
|
||||
|
||||
**Площадка (`site-N`) никогда не отчитывается по конкретному IP.**
|
||||
Сперва проверьте статус самой площадки — `GET
|
||||
|
||||
@@ -0,0 +1,110 @@
|
||||
# План: самопроверка через control-api (дополнительный способ сверки публичного IP)
|
||||
|
||||
> Дата: 2026-10-02 03:06 · Статус: **реализовано и проверено** (юнит-тесты, локальный e2e с `ip_echo` и с `control_api`) (решения пользователя — в конце)
|
||||
|
||||
## Зачем
|
||||
|
||||
Самопроверка агента подтверждает, что исходящий трафик валидатора идёт через выданный Floating IP: агент спрашивает свой
|
||||
публичный адрес у внешнего IP-echo сервиса и сверяет с назначенным адресом. Сейчас это единственный способ, и он хрупкий:
|
||||
1 октября таймауты `https://ifconfig.me/ip` дали волну провалов self-check (33 случая за вечер).
|
||||
|
||||
`control-api` в текущем развёртывании стоит **во внешнем окружении** (вне облака), валидаторы подключаются к нему **напрямую**.
|
||||
Значит, соединение валидатора с ним выходит наружу через Floating IP, и `control-api` сам видит публичный адрес источника.
|
||||
Это даёт второй способ сверки без сторонних сервисов: агент спрашивает у `control-api`, с какого адреса тот его видит.
|
||||
|
||||
Требование: **существующий способ (IP-echo) сохраняется**, новый добавляется как опция агента.
|
||||
|
||||
## Решение в двух строках
|
||||
|
||||
1. `control-api` получает ручку «с какого адреса ты меня видишь».
|
||||
2. Агент получает настройку `self_check.methods` — список способов в порядке приоритета; по умолчанию `[ip_echo]` (всё как сейчас).
|
||||
|
||||
## Конфигурация агента
|
||||
|
||||
```yaml
|
||||
self_check:
|
||||
timeout_seconds: 10
|
||||
methods: [control_api, ip_echo] # по умолчанию [ip_echo]
|
||||
ip_echo_urls: [...] # без изменений
|
||||
```
|
||||
|
||||
- Допустимые значения: `ip_echo` (текущий способ), `control_api` (новый). Неизвестное значение — ошибка при старте агента.
|
||||
- **Самопроверка успешна, если её подтвердил любой из способов.** Способы пробуются по порядку приоритета (первый —
|
||||
главный); остановка на первом успешном. К следующему способу переходим и при отсутствии ответа (ошибка, таймаут,
|
||||
404/5xx), и при несовпадении адреса. Провал — только если не подтвердил ни один способ; в `detail` попадает причина по
|
||||
каждому способу.
|
||||
- Внутри способа `ip_echo` поведение прежнее: URL перебираются по порядку, переход к следующему URL только при ошибке.
|
||||
- Пустой список или отсутствие ключа → `[ip_echo]`. Агент без новой настройки ведёт себя ровно как раньше.
|
||||
- Для внешнего размещения `control-api` в примерах и рекомендациях стоит `[control_api, ip_echo]`: способ через API в приоритете.
|
||||
- Итог в `detail`: `detected_egress_ip=<ip> matched (control_api)`; видно в событиях адреса и в дашборде.
|
||||
|
||||
## Control API
|
||||
|
||||
**Новая ручка:** `GET /api/v1/agents/{id}/observed-ip` → `200 {"ip":"90.156.213.5","source":"remote_addr"}`.
|
||||
|
||||
- Уровень доступа — **открыто**, как `heartbeat` и `assignment`: ручка отдаёт только адрес самого вызывающего, секретов нет.
|
||||
- `{id}` должен быть известным валидатором, иначе `404` (чтобы ручка не превращалась в публичный «узнай свой IP»).
|
||||
- Адрес берётся только из `r.RemoteAddr`, приводится к каноничному виду (`::ffff:1.2.3.4` → `1.2.3.4`).
|
||||
- Заголовки `X-Forwarded-For`/`X-Real-IP` **не учитываются**: подключение прямое, а доверие к заголовку позволило бы
|
||||
валидатору подделать адрес и пройти проверку. Если появится обратный прокси, понадобится отдельная настройка
|
||||
доверенных прокси — сейчас она не нужна (решение 1).
|
||||
|
||||
## Агент (`internal/agentcore`)
|
||||
|
||||
- Новая функция `detectViaControlAPI`: `GET /api/v1/agents/{id}/observed-ip` на `control_api_url`.
|
||||
- **Новое TCP-соединение на каждый вызов** (отдельный `http.Transport` с `DisableKeepAlives`). Это ключевой момент:
|
||||
соединение, открытое до привязки Floating IP (heartbeat, assignment), остаётся в старом NAT-состоянии и покажет
|
||||
прежний адрес; общий клиент `apiclient` использовать нельзя.
|
||||
- Токен агентов в этот запрос не нужен (ручка открытая) и не отправляется.
|
||||
- Таймаут `self_check.timeout_seconds` (сейчас 10 с) действует **на каждый способ отдельно**: при общем дедлайне зависший
|
||||
первый способ (приоритетный `control_api`) съел бы всё время, и запасной не успел бы ответить. Общий предел — таймаут × число способов.
|
||||
- `detectPublicIP` заменяется перебором `methods` по приоритету; `fetchIPEcho` не меняется.
|
||||
- Диагностика: если `control-api` вернул **частный** адрес (RFC 1918 и т. п.), в `detail` пишется подсказка: «control-api
|
||||
доступен по внутренней сети, самопроверка через него невозможна; используйте ip_echo».
|
||||
|
||||
## Ограничение способа (важно для документации)
|
||||
|
||||
Способ `control_api` корректен только если соединение валидатора с `control-api` **выходит через внешнюю сеть**
|
||||
(SNAT Floating IP). Если `control-api` достижим из облака по внутренней сети, он увидит частный адрес валидатора, и
|
||||
этот способ всегда будет давать несовпадение (при `[control_api, ip_echo]` проверка пройдёт по `ip_echo`).
|
||||
Если порт `control-api` опубликован через Docker, при выкладке проверить, что ручка показывает внешний адрес клиента, а
|
||||
не адрес шлюза Docker.
|
||||
|
||||
## Откат и совместимость
|
||||
|
||||
- Новый агент + старый `control-api`: ручки нет (`404`), при `methods: [control_api, ip_echo]` агент переходит на `ip_echo`.
|
||||
- Старый агент + новый `control-api`: ничего не меняется, ручка просто не вызывается.
|
||||
- Откат: убрать `methods` из конфига агента (или вернуть `[ip_echo]`) и перезапустить агент.
|
||||
- Схема БД и протокол `self-check` (`POST /agents/{id}/self-check`) не меняются.
|
||||
|
||||
## Затрагиваемые файлы
|
||||
|
||||
| Файл | Изменение |
|
||||
|---|---|
|
||||
| `internal/config/config.go` | `SelfCheckCfg.Methods`, дефолт `[ip_echo]` и проверка значений |
|
||||
| `internal/httpapi/routes.go`, `handlers_agent.go` | маршрут и обработчик `observed-ip` |
|
||||
| `internal/agentcore/agentcore.go` | `detectViaControlAPI`, перебор `methods` по приоритету |
|
||||
| `configs/validator-agent.example.yaml` | пример и комментарии |
|
||||
| `docs/API.md`, `docs/SETUP.md`, `docs/USAGE.md`, `README.md` | описание ручки, опции, ограничения |
|
||||
| `bin/validator-agent`, `bin/control-api`, `SHA256SUMS` | пересборка |
|
||||
|
||||
## Тесты
|
||||
|
||||
- **config:** дефолт `[ip_echo]`; допустимые значения; ошибка на неизвестном способе.
|
||||
- **httpapi:** прямой адрес; IPv4-mapped IPv6; заголовок `X-Forwarded-For` игнорируется; неизвестный валидатор → `404`.
|
||||
- **agentcore:** порядок способов; успех второго способа после ошибки или несовпадения первого; провал, когда не
|
||||
подтвердил ни один; каждый вызов открывает новое соединение (тестовый сервер считает соединения); подсказка про частный адрес.
|
||||
- **e2e:** `scripts/run-local-e2e.sh` остаётся на `ip_echo` (проверка обратной совместимости).
|
||||
|
||||
## Выкладка
|
||||
|
||||
1. `control-api` с новой ручкой (поведение не меняется): пересборка образа, перезапуск.
|
||||
2. Агенты на ВМ-валидаторах: новый `bin/validator-agent` и `methods: [control_api, ip_echo]` в их конфиге. Один валидатор
|
||||
для начала, проверить `detail` в событиях (`matched (control_api)`), затем остальные.
|
||||
|
||||
## Решения пользователя (2026-10-02)
|
||||
|
||||
1. Подключение валидаторов к `control-api` — **напрямую** (без обратного прокси): `trusted_proxies` не нужен.
|
||||
2. Ручка `observed-ip` — **открытая**.
|
||||
3. Достаточно **одной успешной самопроверки любым из способов**; при внешнем размещении API способ через ручку API —
|
||||
**в приоритете** (первый в списке).
|
||||
@@ -0,0 +1,146 @@
|
||||
# План: исправление оркестратора (двойная выдача валидатору) и очистки очереди
|
||||
|
||||
> Дата: 2026-10-02 09:02 UTC · Статус: **реализовано и проверено** (юнит-тесты с `-race`, локальный e2e; выкладка и приёмка на стенде — ниже). Решения пользователя: предел очистки 10 минут и 8 параллельных отвязок приняты как базовые
|
||||
> Основание: [analysis/2026-10-02_08-56_1026-addresses_mass-check-analysis.md](../../analysis/2026-10-02_08-56_1026-addresses_mass-check-analysis.md)
|
||||
|
||||
## Context
|
||||
|
||||
Массовая проверка 2 октября остановилась на 1026 из 6440 адресов. 7 из 20 валидаторов «залипли»: v1, v12, v13 не взяли
|
||||
ни одного задания после 07:17 и 07:33; v7, v16, v17, v3 залипали временно. Итог: 198 сбросов лизинга, 41 ошибка привязки
|
||||
Floating IP (`409 fixed IP already has a floating IP`), **42 адреса в `fail` без единой выполненной проверки**, потеря
|
||||
~27% пропускной способности. Отдельно «Очистить всё» не уложилась в таймаут клиента (256 с вместо секунд) и сначала
|
||||
оборвалась с 600 ошибками `context canceled`.
|
||||
|
||||
Цель: валидатор в любой момент держит не больше одного адреса; потеря heartbeat не приводит к двойной выдаче;
|
||||
«Очистить всё» выполняется за секунды и не прерывается разрывом соединения.
|
||||
|
||||
## Причины (по коду, подтверждены логами и БД)
|
||||
|
||||
| № | Причина | Где |
|
||||
|---|---|---|
|
||||
| 1 | Heartbeat возвращает `unreachable` → `idle`, не глядя на `current_ip_id`: занятый валидатор снова считается свободным | `internal/db/queries_validators.go:41` (`Heartbeat`), `:31` (`RegisterValidator`, возврат из `unreachable`) |
|
||||
| 2 | Освобождение валидатора идёт **по имени**, а не по адресу, который он держит: завершение старого адреса освобождает валидатор, уже взявший новый. Так же `RequeueOrFail` (в т. ч. при сбросе лизинга) и `MarkFIPOccupied`, `FreeValidator`. Освобождённый ставится в `idle` даже если он `unreachable` — мёртвый валидатор получает новые адреса каждые 3 минуты | `internal/db/queries_ipqueue.go` (`ReleaseFIP`, `RequeueOrFail`, `MarkFIPOccupied`), `queries_validators.go:143` (`FreeValidator`), `internal/orchestrator/orchestrator.go:469` |
|
||||
| 3 | Моя регрессия: защита от дублей привязки ключуется по валидатору (`assign:<validator>`). Вторая выдача того же валидатора не запускает привязку и стоит в `assigning_fip` до конца лизинга | `internal/orchestrator/orchestrator.go:149` |
|
||||
| 4 | Агент шлёт heartbeat только между заданиями. Адрес с 3–4 таймаутами внешних проверок (~40 с) блокирует его дольше порога 30 с | `internal/agentcore/agentcore.go` (`Run`, `pollOnce`) |
|
||||
| 5 | «Очистить всё» отвязывает FIP у **всех** строк с непустым `fip_id`, а он остаётся у `done`/`failed`. Больше 1000 последовательных вызовов OpenStack | `internal/db/queries_ipqueue.go` (`ListFIPRefs`, `ListFIPRefsByAddresses`), `internal/orchestrator/orchestrator.go` (`ClearQueue`, `DeleteIPs`) |
|
||||
| 6 | Очистка работает на контексте HTTP-запроса: разрыв соединения клиентом обрывает её посреди дела (отвязано часть, БД не очищена) | `internal/httpapi/handlers_admin.go:322` |
|
||||
|
||||
## Инварианты, которые вводим
|
||||
|
||||
- **I1.** Валидатор держит не более одного адреса: `validators.current_ip_id = X` тогда и только тогда, когда у строки `X`
|
||||
`owner_validator_id` равен этому валидатору и состояние не терминальное (`done`, `failed`, `occupied`).
|
||||
- **I2.** `idle` означает `current_ip_id IS NULL`. Состояния `unreachable` и `unregistered` не затираются освобождением.
|
||||
- **I3.** Валидатор освобождает только тот адрес, который он сейчас держит. Освобождение и возврат по лизингу чужого или
|
||||
устаревшего адреса состояние валидатора не меняют.
|
||||
|
||||
## Изменения
|
||||
|
||||
### 1. База данных (без миграций, только запросы)
|
||||
|
||||
`internal/db/queries_validators.go`, `queries_ipqueue.go`:
|
||||
|
||||
- **`Heartbeat`:** `unreachable` → `assigned`, если `current_ip_id IS NOT NULL`, иначе `idle`.
|
||||
То же в `RegisterValidator` (возврат из `unreachable`/`unregistered`).
|
||||
- **Общая функция освобождения** `freeValidatorTx(tx, validatorID, ipID)`:
|
||||
`UPDATE validators SET current_ip_id=NULL, state = CASE WHEN state='unreachable' THEN state ELSE 'idle' END WHERE validator_id=? AND current_ip_id=?`.
|
||||
Используют `ReleaseFIP`, `RequeueOrFail`, `MarkFIPOccupied`, `FreeValidator` (получает второй аргумент — id адреса).
|
||||
`deleteIPTx` (уже по `current_ip_id`) и `ClearAllIPs` переводятся на тот же `CASE`, чтобы не затирать `unreachable`.
|
||||
- **`ClaimNextQueued`:** условие обновления валидатора дополняется `AND current_ip_id IS NULL`.
|
||||
- **`ReconcileValidators`** (новая, вызывается из `Orchestrator.Tick`): лечит нарушение инвариантов, если они всё же возникли
|
||||
(падение процесса, старые строки): валидатор с `current_ip_id`, чья строка не существует, терминальна или принадлежит другому
|
||||
валидатору, освобождается; строка в `assigning_fip`/`awaiting_self_check`/`checking`, чей владелец не ссылается на неё,
|
||||
не трогается (её вернёт сброс лизинга). Один короткий запрос на такт.
|
||||
- **`ListFIPRefs` и `ListFIPRefsByAddresses`:** только строки в нетерминальных состояниях (`state NOT IN done, failed, occupied`)
|
||||
с непустым `fip_id`: у терминальных FIP уже отвязан (агрегация, возврат по лизингу, отмена и `occupied` делают это до записи
|
||||
состояния). Значения `fip_id` в строках не меняются (дашборд их показывает).
|
||||
|
||||
### 2. Оркестратор (`internal/orchestrator/orchestrator.go`)
|
||||
|
||||
- Ключ защиты от дублей привязки: `assign:<id адреса>`, а не `assign:<validator>` (строка 149). Дублирующий запуск привязки
|
||||
одного и того же адреса по-прежнему исключён.
|
||||
- `ForceCancel`, `sweepExpiredLeases`, `aggregateAndRelease`, `SelfCheckResult`: передают id адреса в освобождение (I3).
|
||||
- **`ClearQueue`/`DeleteIPs`:** отвязка FIP из `ListFIPRefs` (теперь ≤ числа валидаторов) выполняется параллельно, не более
|
||||
8 одновременных вызовов; затем прямой опрос портов (`releaseValidatorPorts`) как сейчас.
|
||||
- `Tick`: вызывает `ReconcileValidators` первым шагом.
|
||||
|
||||
### 3. HTTP (`internal/httpapi/handlers_admin.go`)
|
||||
|
||||
- «Очистить всё», удаление списка и отмена: контекст отвязан от отмены запроса (`context.WithoutCancel`) с собственным пределом
|
||||
времени (10 минут). Разрыв соединения клиентом (в т. ч. таймаут дашборда) больше не обрывает операцию на середине.
|
||||
|
||||
### 4. Агент (`internal/agentcore/agentcore.go`)
|
||||
|
||||
- Heartbeat уходит из `pollOnce` в **отдельную горутину** со своим тикером (период `poll_interval_seconds`); останавливается по
|
||||
отмене контекста. Долгие внешние проверки больше не блокируют heartbeat. Ошибки heartbeat пишутся в лог (предупреждение).
|
||||
- Протокол и конфигурация агента не меняются (старый агент с новым сервером и наоборот работают).
|
||||
|
||||
### 5. Документация
|
||||
|
||||
`docs/USAGE.md`/`docs/DIAGRAMS.md` (состояния валидатора и правила освобождения), запись в истории изменений `README.md`.
|
||||
|
||||
## Тесты
|
||||
|
||||
Пишу сам (агенты тесты не делают); каждый тест должен падать без исправления.
|
||||
|
||||
- **db:**
|
||||
- `Heartbeat` из `unreachable` с адресом даёт `assigned`, без адреса `idle`; `RegisterValidator` аналогично;
|
||||
- `ReleaseFIP`/`RequeueOrFail`/`MarkFIPOccupied` старого адреса не освобождают валидатор, который уже держит другой адрес;
|
||||
- освобождение `unreachable`-валидатора оставляет `unreachable`;
|
||||
- `ClaimNextQueued` не выдаёт адрес валидатору с `current_ip_id`;
|
||||
- `ReconcileValidators` лечит битые строки и не трогает корректные;
|
||||
- `ListFIPRefs`/`ListFIPRefsByAddresses` не содержат `done`/`failed`/`occupied`.
|
||||
- **orchestrator:**
|
||||
- сценарий инцидента: валидатор помечен `unreachable`, держа адрес A с идущими проверками → heartbeat → такт оркестратора
|
||||
**не выдаёт** ему B; после завершения A валидатор получает B;
|
||||
- мёртвый валидатор (`unreachable`, лизинг истёк) не получает новых адресов;
|
||||
- `ClearQueue` при 1000 строк `done` с `fip_id` и 5 активных не вызывает `Disassociate` для `done` (счётчик вызовов в обёртке OpenStack);
|
||||
- защита привязки по адресу (два адреса одного валидатора, искусственно, оба привязываются);
|
||||
- **случайный сценарий под нагрузкой** (20 валидаторов, 300 адресов, случайная потеря heartbeat, долгие проверки, часть адресов
|
||||
с отказом привязки): после каждого такта проверяются I1–I3; в конце все адреса завершены, сбросов лизинга нет.
|
||||
- **httpapi:** «Очистить всё» доживает до конца при отмене контекста запроса посередине (БД очищена, FIP отвязаны).
|
||||
- **agentcore:** во время долгой проверки (цель отвечает 3 с) heartbeat уходит по расписанию; остановка по отмене контекста.
|
||||
- **Общий прогон:** `go vet`, `go test -race ./...`, `scripts/run-local-e2e.sh` (с `ip_echo` и `control_api`).
|
||||
|
||||
## Выкладка
|
||||
|
||||
1. Правки control-api: сборка `bin/control-api`, образ `civ-capi`, перезапуск (БД не затрагивается; миграций нет).
|
||||
Одного этого достаточно, чтобы остановить двойные выдачи и бесконечное «залипание».
|
||||
2. Агент: сборка `bin/validator-agent`, коммит, раскатка Ansible-сценарием `deploy/ansible` (запускает пользователь): heartbeat в
|
||||
отдельном потоке убирает ложные `unreachable`.
|
||||
3. Перепроверка 42 адресов: список выгружается из снимка анализа в файл `analysis/2026-10-02_08-56_failed-addresses.txt` (уже выгружен, 42 адреса); ставятся в очередь
|
||||
через `POST /api/v1/admin/ips` (или «Перепроверка» в дашборде) после выкладки.
|
||||
|
||||
## Приёмка на стенде
|
||||
|
||||
Контрольная группа из 20 адресов, затем 400 адресов (при тех же внешних целях с долгими таймаутами) с наблюдением 30 минут:
|
||||
|
||||
| Показатель | Критерий |
|
||||
|---|---|
|
||||
| `lease_expired` | 0 |
|
||||
| Ошибки привязки `409` | 0 |
|
||||
| Двойные выдачи (два `claimed ip` одному валидатору за <10 с) | 0 |
|
||||
| Завершено каждым валидатором | отклонение от среднего не больше 15% |
|
||||
| `unreachable` при долгих проверках | нет (после раскатки нового агента) |
|
||||
| «Очистить всё» при >1000 строк `done` | ответ не дольше 10 с |
|
||||
| Адреса `fail` | только по существу (не из-за лизинга) |
|
||||
|
||||
## Риски и откат
|
||||
|
||||
- Изменения только в запросах и логике, без миграций; схема БД и протокол агента не меняются. Откат — предыдущий образ `civ-capi`
|
||||
(`docker tag`/предыдущий коммит) и прежний бинарник агента.
|
||||
- Риск: условные `UPDATE` по `current_ip_id` могут оставить валидатор занятым, если строка адреса пропала. Страхует `ReconcileValidators`.
|
||||
- Валидатор `unreachable` теперь не получает адресов до первого heartbeat — это намеренно; если агент жив, но heartbeat по сети
|
||||
не проходит, он простаивает (раньше брал адреса и терял их).
|
||||
|
||||
## Не входит в эту правку
|
||||
|
||||
- Параллельное выполнение внешних проверок в агенте (3–4 таймаута сейчас идут подряд): сократило бы цикл с ~84 с и снизило бы
|
||||
нагрузку на heartbeat; отдельное решение.
|
||||
- Судьба цели `packages.ubuntu.com` (71% `partial`) и причины недоступности проберов — отдельные вопросы из анализа.
|
||||
- Сокращение `fip_settle_seconds` (30 с) ради темпа.
|
||||
|
||||
## Вопросы к согласованию
|
||||
|
||||
1. Предел времени «Очистить всё» (10 минут) и число параллельных отвязок (8) — подходят?
|
||||
2. Выкладывать control-api сразу после тестов (до раскатки агента) — да, как в разделе «Выкладка»?
|
||||
3. 42 адреса `fail` перепроверять сразу после выкладки или вместе со следующим большим прогоном?
|
||||
+158
-15
@@ -10,13 +10,16 @@ package agentcore
|
||||
|
||||
import (
|
||||
"context"
|
||||
"encoding/json"
|
||||
"fmt"
|
||||
"io"
|
||||
"log/slog"
|
||||
"net"
|
||||
"net/http"
|
||||
neturl "net/url"
|
||||
"os"
|
||||
"strings"
|
||||
"sync/atomic"
|
||||
"time"
|
||||
|
||||
"cloudipvalidator/internal/apiclient"
|
||||
@@ -31,6 +34,10 @@ type Agent struct {
|
||||
|
||||
lastHandledIPID int64
|
||||
|
||||
// busy is true while an assignment is being worked on; it is reported in
|
||||
// the heartbeat body (informational on the control-api side).
|
||||
busy atomic.Bool
|
||||
|
||||
// registerRetryInitial/Max govern the backoff used while waiting for a
|
||||
// successful registration (see registerWithRetry): control-api may not
|
||||
// be up yet at agent boot, or may come and go across a redeploy, and the
|
||||
@@ -68,6 +75,15 @@ func (a *Agent) Run(ctx context.Context) error {
|
||||
}
|
||||
|
||||
interval := time.Duration(a.cfg.PollIntervalSeconds) * time.Second
|
||||
|
||||
// Heartbeats run on their own schedule. Sent from the poll loop they
|
||||
// stopped for as long as a slow assignment took (an address whose
|
||||
// outbound targets all time out keeps the loop busy for ~40 s), which
|
||||
// control-api reads as a lost validator after heartbeat_timeout_seconds.
|
||||
hbCtx, stopHeartbeat := context.WithCancel(ctx)
|
||||
defer stopHeartbeat()
|
||||
go a.heartbeatLoop(hbCtx, interval)
|
||||
|
||||
ticker := time.NewTicker(interval)
|
||||
defer ticker.Stop()
|
||||
|
||||
@@ -146,12 +162,32 @@ type checkConfigDTO struct {
|
||||
Targets []string `json:"targets"`
|
||||
}
|
||||
|
||||
func (a *Agent) pollOnce(ctx context.Context) {
|
||||
if _, err := a.client.Do(ctx, "POST", "/api/v1/agents/"+a.cfg.ValidatorID+"/heartbeat", heartbeatReq{LocalState: "idle"}, nil); err != nil {
|
||||
a.log.Error("heartbeat", "err", err)
|
||||
return
|
||||
// heartbeatLoop sends a heartbeat now and then every interval until ctx is
|
||||
// cancelled.
|
||||
func (a *Agent) heartbeatLoop(ctx context.Context, interval time.Duration) {
|
||||
ticker := time.NewTicker(interval)
|
||||
defer ticker.Stop()
|
||||
for {
|
||||
a.sendHeartbeat(ctx)
|
||||
select {
|
||||
case <-ctx.Done():
|
||||
return
|
||||
case <-ticker.C:
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
func (a *Agent) sendHeartbeat(ctx context.Context) {
|
||||
state := "idle"
|
||||
if a.busy.Load() {
|
||||
state = "checking"
|
||||
}
|
||||
if _, err := a.client.Do(ctx, "POST", "/api/v1/agents/"+a.cfg.ValidatorID+"/heartbeat", heartbeatReq{LocalState: state}, nil); err != nil && ctx.Err() == nil {
|
||||
a.log.Error("heartbeat", "err", err)
|
||||
}
|
||||
}
|
||||
|
||||
func (a *Agent) pollOnce(ctx context.Context) {
|
||||
var assignment assignmentResp
|
||||
ok, err := a.client.Do(ctx, "GET", "/api/v1/agents/"+a.cfg.ValidatorID+"/assignment", nil, &assignment)
|
||||
if err != nil {
|
||||
@@ -167,6 +203,8 @@ func (a *Agent) pollOnce(ctx context.Context) {
|
||||
return // already handled this IP's work this attempt
|
||||
}
|
||||
|
||||
a.busy.Store(true)
|
||||
defer a.busy.Store(false)
|
||||
switch assignment.Phase {
|
||||
case "awaiting_self_check":
|
||||
a.handleSelfCheckAndRun(ctx, assignment)
|
||||
@@ -181,18 +219,13 @@ func (a *Agent) pollOnce(ctx context.Context) {
|
||||
func (a *Agent) handleSelfCheckAndRun(ctx context.Context, assignment assignmentResp) {
|
||||
a.postEvent(ctx, assignment.IPID, "config_received", "")
|
||||
|
||||
timeout := time.Duration(a.cfg.SelfCheck.TimeoutSeconds) * time.Second
|
||||
// Each method gets the full timeout (see runSelfCheckMethods), so the
|
||||
// overall budget scales with the number of methods.
|
||||
timeout := time.Duration(a.cfg.SelfCheck.TimeoutSeconds) * time.Second * time.Duration(len(a.selfCheckMethods()))
|
||||
selfCtx, cancel := context.WithTimeout(ctx, timeout)
|
||||
defer cancel()
|
||||
|
||||
detectedIP, err := a.detectPublicIP(selfCtx)
|
||||
success := err == nil && detectedIP == assignment.IPAddress
|
||||
detail := "matched"
|
||||
if err != nil {
|
||||
detail = "ip echo request failed: " + err.Error()
|
||||
} else if !success {
|
||||
detail = fmt.Sprintf("egress ip %q does not match assigned fip %q", detectedIP, assignment.IPAddress)
|
||||
}
|
||||
detectedIP, _, detail, success := a.runSelfCheckMethods(selfCtx, assignment.IPAddress)
|
||||
|
||||
a.postSelfCheck(ctx, assignment.IPID, detectedIP, success, detail)
|
||||
a.postEvent(ctx, assignment.IPID, "self_check_result", fmt.Sprintf(`{"success":%t}`, success))
|
||||
@@ -204,7 +237,117 @@ func (a *Agent) handleSelfCheckAndRun(ctx context.Context, assignment assignment
|
||||
a.runChecks(ctx, assignment)
|
||||
}
|
||||
|
||||
// detectPublicIP asks each configured IP-echo URL, in order, for the
|
||||
// selfCheckMethods returns the configured methods in priority order. Configs
|
||||
// built without the loader (tests) may leave the list empty; that means the
|
||||
// historical behaviour, ip_echo only.
|
||||
func (a *Agent) selfCheckMethods() []string {
|
||||
if len(a.cfg.SelfCheck.Methods) == 0 {
|
||||
return []string{config.SelfCheckIPEcho}
|
||||
}
|
||||
return a.cfg.SelfCheck.Methods
|
||||
}
|
||||
|
||||
// runSelfCheckMethods tries the configured methods in priority order and
|
||||
// stops at the first one that confirms assignedIP. A method that gives no
|
||||
// answer and one that reports a different address are treated alike: the
|
||||
// next method is tried, since either may be a limitation of that method
|
||||
// (e.g. control-api reached over the internal network sees a private
|
||||
// address) rather than proof the floating IP is not attached. The check
|
||||
// fails only when no method confirms, and detail then carries the reason
|
||||
// from every method. detectedIP is the matching address on success, else the
|
||||
// last address any method reported (may be empty).
|
||||
//
|
||||
// Every method runs under its own self_check.timeout_seconds: with one shared
|
||||
// deadline a hung first method (priority control_api) would use it all up and
|
||||
// the fallback would never get a chance to answer.
|
||||
func (a *Agent) runSelfCheckMethods(ctx context.Context, assignedIP string) (detectedIP, method, detail string, ok bool) {
|
||||
var reasons []string
|
||||
perMethod := time.Duration(a.cfg.SelfCheck.TimeoutSeconds) * time.Second
|
||||
for _, m := range a.selfCheckMethods() {
|
||||
mctx, cancel := ctx, context.CancelFunc(func() {})
|
||||
if perMethod > 0 {
|
||||
mctx, cancel = context.WithTimeout(ctx, perMethod)
|
||||
}
|
||||
var ip string
|
||||
var err error
|
||||
switch m {
|
||||
case config.SelfCheckControlAPI:
|
||||
ip, err = a.detectViaControlAPI(mctx)
|
||||
case config.SelfCheckIPEcho:
|
||||
ip, err = a.detectViaIPEcho(mctx)
|
||||
if err != nil {
|
||||
err = fmt.Errorf("ip echo request failed: %w", err)
|
||||
}
|
||||
default:
|
||||
err = fmt.Errorf("unknown self-check method")
|
||||
}
|
||||
cancel()
|
||||
if err != nil {
|
||||
reasons = append(reasons, m+": "+err.Error())
|
||||
continue
|
||||
}
|
||||
if ip == assignedIP {
|
||||
return ip, m, fmt.Sprintf("matched (%s)", m), true
|
||||
}
|
||||
detectedIP = ip
|
||||
reason := fmt.Sprintf("%s: egress ip %q does not match assigned fip %q", m, ip, assignedIP)
|
||||
if m == config.SelfCheckControlAPI && isLocalAddr(ip) {
|
||||
reason += " (control-api sees a private address; it is reachable over the internal network, self-check via control_api is not possible, use ip_echo)"
|
||||
}
|
||||
reasons = append(reasons, reason)
|
||||
}
|
||||
if len(reasons) == 0 {
|
||||
reasons = append(reasons, "no self-check methods configured")
|
||||
}
|
||||
return detectedIP, "", strings.Join(reasons, "; "), false
|
||||
}
|
||||
|
||||
// isLocalAddr reports whether ip is a private, loopback or link-local
|
||||
// address, i.e. one that can never be a floating IP.
|
||||
func isLocalAddr(ip string) bool {
|
||||
parsed := net.ParseIP(ip)
|
||||
return parsed != nil && (parsed.IsPrivate() || parsed.IsLoopback() || parsed.IsLinkLocalUnicast())
|
||||
}
|
||||
|
||||
// detectViaControlAPI asks control-api which source address it sees for this
|
||||
// validator. It is only meaningful when control-api is reached over the
|
||||
// external network, where the floating IP is the visible source (see
|
||||
// config.SelfCheckCfg).
|
||||
//
|
||||
// Every call dials a brand-new TCP connection through a dedicated transport
|
||||
// with keep-alives off: a connection opened before the floating IP was
|
||||
// attached (heartbeat, assignment polling) keeps its old NAT state and would
|
||||
// keep reporting the previous address, so the shared apiclient must not be
|
||||
// used. The endpoint is open, so no agent token is sent.
|
||||
func (a *Agent) detectViaControlAPI(ctx context.Context) (string, error) {
|
||||
url := strings.TrimRight(a.cfg.ControlAPIURL, "/") + "/api/v1/agents/" + neturl.PathEscape(a.cfg.ValidatorID) + "/observed-ip"
|
||||
req, err := http.NewRequestWithContext(ctx, http.MethodGet, url, nil)
|
||||
if err != nil {
|
||||
return "", fmt.Errorf("build request: %w", err)
|
||||
}
|
||||
transport := &http.Transport{DisableKeepAlives: true}
|
||||
defer transport.CloseIdleConnections()
|
||||
resp, err := (&http.Client{Transport: transport}).Do(req)
|
||||
if err != nil {
|
||||
return "", err
|
||||
}
|
||||
defer resp.Body.Close()
|
||||
if resp.StatusCode < 200 || resp.StatusCode > 299 {
|
||||
return "", fmt.Errorf("unexpected status %d", resp.StatusCode)
|
||||
}
|
||||
var body struct {
|
||||
IP string `json:"ip"`
|
||||
}
|
||||
if err := json.NewDecoder(io.LimitReader(resp.Body, 4096)).Decode(&body); err != nil {
|
||||
return "", fmt.Errorf("decode response: %w", err)
|
||||
}
|
||||
if net.ParseIP(body.IP) == nil {
|
||||
return "", fmt.Errorf("response is not a valid IP: %q", body.IP)
|
||||
}
|
||||
return body.IP, nil
|
||||
}
|
||||
|
||||
// detectViaIPEcho asks each configured IP-echo URL, in order, for the
|
||||
// address this validator is currently seen egressing from, returning the
|
||||
// first one that answers with a parseable IP. These must be resources
|
||||
// genuinely outside the cloud project (see config.SelfCheckCfg) — OpenStack
|
||||
@@ -212,7 +355,7 @@ func (a *Agent) handleSelfCheckAndRun(ctx context.Context, assignment assignment
|
||||
// network, so anything reachable over the project's internal network would
|
||||
// report the validator's private address instead, regardless of whether
|
||||
// the floating IP is correctly attached.
|
||||
func (a *Agent) detectPublicIP(ctx context.Context) (string, error) {
|
||||
func (a *Agent) detectViaIPEcho(ctx context.Context) (string, error) {
|
||||
var lastErr error
|
||||
for _, url := range a.cfg.SelfCheck.IPEchoURLs {
|
||||
ip, err := fetchIPEcho(ctx, url)
|
||||
|
||||
@@ -0,0 +1,72 @@
|
||||
package agentcore
|
||||
|
||||
import (
|
||||
"context"
|
||||
"fmt"
|
||||
"net/http"
|
||||
"net/http/httptest"
|
||||
"sync/atomic"
|
||||
"testing"
|
||||
"time"
|
||||
|
||||
"cloudipvalidator/internal/config"
|
||||
)
|
||||
|
||||
// While the agent is busy with a slow assignment (an address whose outbound
|
||||
// targets time out keeps it occupied for tens of seconds) it must keep sending
|
||||
// heartbeats; control-api marks a validator that stays silent for
|
||||
// heartbeat_timeout_seconds as unreachable.
|
||||
func TestHeartbeatContinuesDuringSlowChecks(t *testing.T) {
|
||||
slowTarget := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
|
||||
time.Sleep(3500 * time.Millisecond)
|
||||
}))
|
||||
defer slowTarget.Close()
|
||||
|
||||
var heartbeats, assignments int32
|
||||
mux := http.NewServeMux()
|
||||
mux.HandleFunc("POST /api/v1/agents/register", func(w http.ResponseWriter, r *http.Request) {
|
||||
fmt.Fprint(w, `{"ok":true}`)
|
||||
})
|
||||
mux.HandleFunc("POST /api/v1/agents/val-1/heartbeat", func(w http.ResponseWriter, r *http.Request) {
|
||||
atomic.AddInt32(&heartbeats, 1)
|
||||
fmt.Fprint(w, `{"ok":true}`)
|
||||
})
|
||||
mux.HandleFunc("GET /api/v1/agents/val-1/assignment", func(w http.ResponseWriter, r *http.Request) {
|
||||
if atomic.AddInt32(&assignments, 1) > 1 {
|
||||
w.WriteHeader(http.StatusNoContent)
|
||||
return
|
||||
}
|
||||
fmt.Fprintf(w, `{"ip_id":1,"ip_address":"1.1.1.1","phase":"checking","check_config":[{"type":"https","targets":[%q]}]}`, slowTarget.URL)
|
||||
})
|
||||
mux.HandleFunc("POST /api/v1/agents/val-1/results", func(w http.ResponseWriter, r *http.Request) { fmt.Fprint(w, `{"ok":true}`) })
|
||||
mux.HandleFunc("POST /api/v1/agents/val-1/complete", func(w http.ResponseWriter, r *http.Request) { fmt.Fprint(w, `{"ok":true}`) })
|
||||
capi := httptest.NewServer(mux)
|
||||
defer capi.Close()
|
||||
|
||||
a := New(&config.ValidatorAgent{
|
||||
ValidatorID: "val-1", ControlAPIURL: capi.URL, PollIntervalSeconds: 1,
|
||||
Checks: config.AgentChecks{HTTPSTimeoutSeconds: 10, ICMPTimeoutSeconds: 1, ICMPCount: 1},
|
||||
}, testLogger())
|
||||
|
||||
ctx, cancel := context.WithCancel(context.Background())
|
||||
done := make(chan struct{})
|
||||
go func() { _ = a.Run(ctx); close(done) }()
|
||||
|
||||
time.Sleep(3 * time.Second) // the slow check (3.5 s) is still running
|
||||
during := atomic.LoadInt32(&heartbeats)
|
||||
cancel()
|
||||
select {
|
||||
case <-done:
|
||||
case <-time.After(10 * time.Second):
|
||||
t.Fatal("Run did not stop after the context was cancelled")
|
||||
}
|
||||
|
||||
// One per second plus the first: 3-4 in 3 s. With heartbeats in the poll
|
||||
// loop there is exactly one, sent before the slow assignment started.
|
||||
if during < 3 {
|
||||
t.Fatalf("%d heartbeats in 3 s while a check was running, want at least 3", during)
|
||||
}
|
||||
if atomic.LoadInt32(&assignments) < 1 {
|
||||
t.Fatal("the assignment was never fetched, the test did not exercise a busy agent")
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,232 @@
|
||||
package agentcore
|
||||
|
||||
import (
|
||||
"context"
|
||||
"fmt"
|
||||
"net"
|
||||
"net/http"
|
||||
"net/http/httptest"
|
||||
"strings"
|
||||
"sync/atomic"
|
||||
"testing"
|
||||
"time"
|
||||
|
||||
"cloudipvalidator/internal/config"
|
||||
)
|
||||
|
||||
// echoServer answers every request with body (an IP-echo stand-in).
|
||||
func echoServer(t *testing.T, body string) *httptest.Server {
|
||||
t.Helper()
|
||||
ts := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
|
||||
fmt.Fprint(w, body)
|
||||
}))
|
||||
t.Cleanup(ts.Close)
|
||||
return ts
|
||||
}
|
||||
|
||||
// controlAPIServer plays control-api's observed-ip route: it answers with ip
|
||||
// (status 200) or with the given error status when ip is empty.
|
||||
func controlAPIServer(t *testing.T, ip string, status int) *httptest.Server {
|
||||
t.Helper()
|
||||
ts := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
|
||||
if r.URL.Path != "/api/v1/agents/val-1/observed-ip" {
|
||||
http.NotFound(w, r)
|
||||
return
|
||||
}
|
||||
if ip == "" {
|
||||
w.WriteHeader(status)
|
||||
return
|
||||
}
|
||||
fmt.Fprintf(w, `{"ip":%q,"source":"remote_addr"}`, ip)
|
||||
}))
|
||||
t.Cleanup(ts.Close)
|
||||
return ts
|
||||
}
|
||||
|
||||
func selfCheckAgent(controlAPIURL string, methods []string, echoURLs ...string) *Agent {
|
||||
return &Agent{
|
||||
cfg: &config.ValidatorAgent{
|
||||
ValidatorID: "val-1",
|
||||
ControlAPIURL: controlAPIURL,
|
||||
SelfCheck: config.SelfCheckCfg{TimeoutSeconds: 2, Methods: methods, IPEchoURLs: echoURLs},
|
||||
},
|
||||
log: testLogger(),
|
||||
}
|
||||
}
|
||||
|
||||
func TestSelfCheckControlAPIFirstWins(t *testing.T) {
|
||||
capi := controlAPIServer(t, "1.2.3.4", 0)
|
||||
var echoCalls int32
|
||||
echo := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
|
||||
atomic.AddInt32(&echoCalls, 1)
|
||||
fmt.Fprint(w, "1.2.3.4")
|
||||
}))
|
||||
defer echo.Close()
|
||||
|
||||
a := selfCheckAgent(capi.URL, []string{config.SelfCheckControlAPI, config.SelfCheckIPEcho}, echo.URL)
|
||||
ip, method, detail, ok := a.runSelfCheckMethods(context.Background(), "1.2.3.4")
|
||||
if !ok || ip != "1.2.3.4" || method != config.SelfCheckControlAPI || detail != "matched (control_api)" {
|
||||
t.Fatalf("got ip=%q method=%q detail=%q ok=%v", ip, method, detail, ok)
|
||||
}
|
||||
if n := atomic.LoadInt32(&echoCalls); n != 0 {
|
||||
t.Fatalf("ip_echo was called %d times although control_api already confirmed", n)
|
||||
}
|
||||
}
|
||||
|
||||
// Any one confirming method is enough: control-api answers with a different
|
||||
// address, ip_echo confirms.
|
||||
func TestSelfCheckFallsThroughOnMismatch(t *testing.T) {
|
||||
capi := controlAPIServer(t, "5.5.5.5", 0)
|
||||
echo := echoServer(t, "1.2.3.4")
|
||||
|
||||
a := selfCheckAgent(capi.URL, []string{config.SelfCheckControlAPI, config.SelfCheckIPEcho}, echo.URL)
|
||||
ip, method, _, ok := a.runSelfCheckMethods(context.Background(), "1.2.3.4")
|
||||
if !ok || ip != "1.2.3.4" || method != config.SelfCheckIPEcho {
|
||||
t.Fatalf("got ip=%q method=%q ok=%v, want a pass via ip_echo", ip, method, ok)
|
||||
}
|
||||
}
|
||||
|
||||
// An old control-api without the route (404) or a failing one (5xx) must not
|
||||
// stop the self-check: the next method decides.
|
||||
func TestSelfCheckFallsThroughOnControlAPIError(t *testing.T) {
|
||||
for _, status := range []int{http.StatusNotFound, http.StatusInternalServerError} {
|
||||
capi := controlAPIServer(t, "", status)
|
||||
echo := echoServer(t, "1.2.3.4")
|
||||
a := selfCheckAgent(capi.URL, []string{config.SelfCheckControlAPI, config.SelfCheckIPEcho}, echo.URL)
|
||||
_, method, _, ok := a.runSelfCheckMethods(context.Background(), "1.2.3.4")
|
||||
if !ok || method != config.SelfCheckIPEcho {
|
||||
t.Fatalf("status %d: method=%q ok=%v, want a pass via ip_echo", status, method, ok)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
func TestSelfCheckFailsWhenNoMethodConfirms(t *testing.T) {
|
||||
capi := controlAPIServer(t, "10.0.0.5", 0) // private address: internal-network case
|
||||
echo := echoServer(t, "6.6.6.6")
|
||||
|
||||
a := selfCheckAgent(capi.URL, []string{config.SelfCheckControlAPI, config.SelfCheckIPEcho}, echo.URL)
|
||||
ip, method, detail, ok := a.runSelfCheckMethods(context.Background(), "1.2.3.4")
|
||||
if ok || method != "" {
|
||||
t.Fatalf("expected a failure, got ok=%v method=%q", ok, method)
|
||||
}
|
||||
if ip != "6.6.6.6" {
|
||||
t.Fatalf("detected ip = %q, want the last reported address 6.6.6.6", ip)
|
||||
}
|
||||
for _, want := range []string{
|
||||
`control_api: egress ip "10.0.0.5" does not match assigned fip "1.2.3.4"`,
|
||||
"private address", // the hint for the internal-network case
|
||||
`ip_echo: egress ip "6.6.6.6" does not match`,
|
||||
} {
|
||||
if !strings.Contains(detail, want) {
|
||||
t.Fatalf("detail %q does not contain %q", detail, want)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
func TestSelfCheckPrivateHintOnlyForControlAPI(t *testing.T) {
|
||||
echo := echoServer(t, "10.1.1.1")
|
||||
a := selfCheckAgent("http://unused", []string{config.SelfCheckIPEcho}, echo.URL)
|
||||
_, _, detail, ok := a.runSelfCheckMethods(context.Background(), "1.2.3.4")
|
||||
if ok || strings.Contains(detail, "private address") {
|
||||
t.Fatalf("ok=%v detail=%q: the control_api hint must not appear for ip_echo", ok, detail)
|
||||
}
|
||||
}
|
||||
|
||||
// With nothing configured (a config built without the loader) the agent
|
||||
// behaves as before: ip_echo only, control-api is never asked.
|
||||
func TestSelfCheckDefaultsToIPEcho(t *testing.T) {
|
||||
var capiCalls int32
|
||||
capi := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
|
||||
atomic.AddInt32(&capiCalls, 1)
|
||||
}))
|
||||
defer capi.Close()
|
||||
echo := echoServer(t, "1.2.3.4")
|
||||
|
||||
a := selfCheckAgent(capi.URL, nil, echo.URL)
|
||||
_, method, _, ok := a.runSelfCheckMethods(context.Background(), "1.2.3.4")
|
||||
if !ok || method != config.SelfCheckIPEcho || atomic.LoadInt32(&capiCalls) != 0 {
|
||||
t.Fatalf("method=%q ok=%v control-api calls=%d", method, ok, atomic.LoadInt32(&capiCalls))
|
||||
}
|
||||
}
|
||||
|
||||
// A hung control-api must not use up the time of the fallback: each method
|
||||
// has its own timeout.
|
||||
func TestSelfCheckHungControlAPIDoesNotStarveFallback(t *testing.T) {
|
||||
release := make(chan struct{})
|
||||
capi := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
|
||||
<-release
|
||||
}))
|
||||
defer capi.Close()
|
||||
defer close(release)
|
||||
echo := echoServer(t, "1.2.3.4")
|
||||
|
||||
a := selfCheckAgent(capi.URL, []string{config.SelfCheckControlAPI, config.SelfCheckIPEcho}, echo.URL)
|
||||
a.cfg.SelfCheck.TimeoutSeconds = 1
|
||||
|
||||
start := time.Now()
|
||||
// The outer context mirrors handleSelfCheckAndRun: timeout x methods.
|
||||
ctx, cancel := context.WithTimeout(context.Background(), 2*time.Second)
|
||||
defer cancel()
|
||||
_, method, detail, ok := a.runSelfCheckMethods(ctx, "1.2.3.4")
|
||||
if !ok || method != config.SelfCheckIPEcho {
|
||||
t.Fatalf("method=%q ok=%v detail=%q, want a pass via ip_echo after control_api timed out", method, ok, detail)
|
||||
}
|
||||
if elapsed := time.Since(start); elapsed > 1900*time.Millisecond {
|
||||
t.Fatalf("took %s: the hung method consumed the fallback's time", elapsed)
|
||||
}
|
||||
}
|
||||
|
||||
// Every control-api request must use a new TCP connection: a connection
|
||||
// opened before the floating IP was attached would report the old address.
|
||||
func TestDetectViaControlAPIDialsNewConnectionEachTime(t *testing.T) {
|
||||
var conns int32
|
||||
ts := httptest.NewUnstartedServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
|
||||
fmt.Fprint(w, `{"ip":"1.2.3.4","source":"remote_addr"}`)
|
||||
}))
|
||||
ts.Config.ConnState = func(_ net.Conn, s http.ConnState) {
|
||||
if s == http.StateNew {
|
||||
atomic.AddInt32(&conns, 1)
|
||||
}
|
||||
}
|
||||
ts.Start()
|
||||
defer ts.Close()
|
||||
|
||||
a := selfCheckAgent(ts.URL, nil)
|
||||
for i := 0; i < 3; i++ {
|
||||
if _, err := a.detectViaControlAPI(context.Background()); err != nil {
|
||||
t.Fatalf("call %d: %v", i, err)
|
||||
}
|
||||
}
|
||||
if n := atomic.LoadInt32(&conns); n != 3 {
|
||||
t.Fatalf("3 calls opened %d connections, want 3 (no keep-alive reuse)", n)
|
||||
}
|
||||
}
|
||||
|
||||
func TestDetectViaControlAPIRejectsBadAnswers(t *testing.T) {
|
||||
for name, body := range map[string]string{"not json": "oops", "not an ip": `{"ip":"abc"}`, "empty": `{}`} {
|
||||
ts := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) { fmt.Fprint(w, body) }))
|
||||
a := selfCheckAgent(ts.URL, nil)
|
||||
if ip, err := a.detectViaControlAPI(context.Background()); err == nil {
|
||||
t.Fatalf("%s: expected an error, got %q", name, ip)
|
||||
}
|
||||
ts.Close()
|
||||
}
|
||||
}
|
||||
|
||||
// The agent token must never be sent on this request (the route is open and
|
||||
// the token is meant for control-api writes only).
|
||||
func TestDetectViaControlAPISendsNoToken(t *testing.T) {
|
||||
var auth string
|
||||
ts := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
|
||||
auth = r.Header.Get("Authorization")
|
||||
fmt.Fprint(w, `{"ip":"1.2.3.4"}`)
|
||||
}))
|
||||
defer ts.Close()
|
||||
a := selfCheckAgent(ts.URL, nil)
|
||||
if _, err := a.detectViaControlAPI(context.Background()); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if auth != "" {
|
||||
t.Fatalf("Authorization header sent: %q", auth)
|
||||
}
|
||||
}
|
||||
@@ -233,17 +233,30 @@ type ValidatorAgent struct {
|
||||
Checks AgentChecks `yaml:"checks"`
|
||||
}
|
||||
|
||||
// Self-check methods accepted in SelfCheckCfg.Methods.
|
||||
const (
|
||||
SelfCheckIPEcho = "ip_echo"
|
||||
SelfCheckControlAPI = "control_api"
|
||||
)
|
||||
|
||||
// SelfCheckCfg configures how the agent confirms its egress actually flows
|
||||
// through the newly assigned floating IP. This must query a resource
|
||||
// genuinely outside the cloud project: OpenStack only applies floating-IP
|
||||
// SNAT to traffic leaving via the external/provider network, so any
|
||||
// in-project resource (including control-api, if it's reachable over the
|
||||
// project's internal network) would see the validator's private address
|
||||
// instead — a false negative that never changes. IPEchoURLs are tried in
|
||||
// order (falling through to the next on error/timeout, not on a genuine
|
||||
// mismatch) until one returns a parseable IP.
|
||||
// instead — a false negative that never changes. The control_api method is
|
||||
// therefore only valid when control-api is reached over the external network
|
||||
// (hosted outside the cloud); then it sees the floating IP as the source.
|
||||
//
|
||||
// Methods are tried in priority order and the self-check passes as soon as
|
||||
// any one confirms the address; the next method is tried both when one gives
|
||||
// no answer and when it reports a different address. IPEchoURLs are tried in
|
||||
// order within ip_echo (falling through to the next on error/timeout, not on
|
||||
// a genuine mismatch) until one returns a parseable IP.
|
||||
type SelfCheckCfg struct {
|
||||
TimeoutSeconds int `yaml:"timeout_seconds"`
|
||||
Methods []string `yaml:"methods"`
|
||||
IPEchoURLs []string `yaml:"ip_echo_urls"`
|
||||
}
|
||||
|
||||
@@ -273,6 +286,14 @@ func LoadValidatorAgent(path string) (*ValidatorAgent, error) {
|
||||
if c.SelfCheck.TimeoutSeconds == 0 {
|
||||
c.SelfCheck.TimeoutSeconds = 10
|
||||
}
|
||||
if len(c.SelfCheck.Methods) == 0 {
|
||||
c.SelfCheck.Methods = []string{SelfCheckIPEcho}
|
||||
}
|
||||
for _, m := range c.SelfCheck.Methods {
|
||||
if m != SelfCheckIPEcho && m != SelfCheckControlAPI {
|
||||
return nil, fmt.Errorf("self_check.methods: unknown method %q (allowed: %s, %s)", m, SelfCheckIPEcho, SelfCheckControlAPI)
|
||||
}
|
||||
}
|
||||
if len(c.SelfCheck.IPEchoURLs) == 0 {
|
||||
c.SelfCheck.IPEchoURLs = []string{"https://api.ipify.org", "https://ifconfig.me/ip"}
|
||||
}
|
||||
|
||||
@@ -64,3 +64,58 @@ func TestControlAPIExampleConfigs(t *testing.T) {
|
||||
t.Fatalf("load docker example: %v", err)
|
||||
}
|
||||
}
|
||||
|
||||
func writeAgentConfig(t *testing.T, selfCheck string) string {
|
||||
t.Helper()
|
||||
path := filepath.Join(t.TempDir(), "agent.yaml")
|
||||
body := "validator_id: v1\ncontrol_api_url: http://x\nself_check:\n" + selfCheck
|
||||
if err := os.WriteFile(path, []byte(body), 0o600); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
return path
|
||||
}
|
||||
|
||||
func TestLoadValidatorAgentSelfCheckMethods(t *testing.T) {
|
||||
// No methods key: the historical behaviour, ip_echo only.
|
||||
c, err := LoadValidatorAgent(writeAgentConfig(t, " timeout_seconds: 5\n"))
|
||||
if err != nil {
|
||||
t.Fatalf("load: %v", err)
|
||||
}
|
||||
if len(c.SelfCheck.Methods) != 1 || c.SelfCheck.Methods[0] != SelfCheckIPEcho {
|
||||
t.Fatalf("default methods = %v, want [ip_echo]", c.SelfCheck.Methods)
|
||||
}
|
||||
|
||||
// Order is the priority and must be kept.
|
||||
c, err = LoadValidatorAgent(writeAgentConfig(t, " methods: [control_api, ip_echo]\n"))
|
||||
if err != nil {
|
||||
t.Fatalf("load: %v", err)
|
||||
}
|
||||
if got := c.SelfCheck.Methods; len(got) != 2 || got[0] != SelfCheckControlAPI || got[1] != SelfCheckIPEcho {
|
||||
t.Fatalf("methods = %v, want [control_api ip_echo]", got)
|
||||
}
|
||||
|
||||
// An unknown method is a configuration error, not a silent no-op.
|
||||
if _, err := LoadValidatorAgent(writeAgentConfig(t, " methods: [control_api, ipecho]\n")); err == nil {
|
||||
t.Fatal("expected an error for an unknown self-check method")
|
||||
}
|
||||
}
|
||||
|
||||
// The shipped agent example must load, and the rxprod copy must stay
|
||||
// byte-identical to it.
|
||||
func TestValidatorAgentExampleConfig(t *testing.T) {
|
||||
c, err := LoadValidatorAgent("../../configs/validator-agent.example.yaml")
|
||||
if err != nil {
|
||||
t.Fatalf("load example: %v", err)
|
||||
}
|
||||
if got := c.SelfCheck.Methods; len(got) != 2 || got[0] != SelfCheckControlAPI {
|
||||
t.Fatalf("example methods = %v", got)
|
||||
}
|
||||
a, _ := os.ReadFile("../../configs/validator-agent.example.yaml")
|
||||
b, err := os.ReadFile("../../rxprod-compose/sources/validator-agent.example.yaml")
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if !bytes.Equal(a, b) {
|
||||
t.Fatal("rxprod-compose/sources/validator-agent.example.yaml differs from configs/validator-agent.example.yaml")
|
||||
}
|
||||
}
|
||||
@@ -101,7 +101,7 @@ func (d *DB) ClaimNextQueued(ctx context.Context, validatorID string, leaseTTL t
|
||||
|
||||
res, err = tx.ExecContext(ctx, `
|
||||
UPDATE validators SET state=?, current_ip_id=?, updated_at=?
|
||||
WHERE validator_id=? AND state=?
|
||||
WHERE validator_id=? AND state=? AND current_ip_id IS NULL
|
||||
`, ValidatorAssigned, item.ID, timeToDB(now), validatorID, ValidatorIdle)
|
||||
if err != nil {
|
||||
return nil, err
|
||||
@@ -207,7 +207,8 @@ func (d *DB) FinishIP(ctx context.Context, ipID int64, result string) error {
|
||||
}
|
||||
|
||||
// ReleaseFIP records that the floating IP has been disassociated and frees
|
||||
// the owning validator back to idle, in one transaction.
|
||||
// the owning validator (if this address is still its current one, see
|
||||
// freeValidatorSQL), in one transaction.
|
||||
func (d *DB) ReleaseFIP(ctx context.Context, ipID int64, validatorID string) error {
|
||||
tx, err := d.BeginTx(ctx, nil)
|
||||
if err != nil {
|
||||
@@ -219,10 +220,7 @@ func (d *DB) ReleaseFIP(ctx context.Context, ipID int64, validatorID string) err
|
||||
if _, err := tx.ExecContext(ctx, `UPDATE ip_queue SET fip_released_at=?, updated_at=? WHERE id=?`, now, now, ipID); err != nil {
|
||||
return err
|
||||
}
|
||||
if _, err := tx.ExecContext(ctx, `
|
||||
UPDATE validators SET state=?, current_ip_id=NULL, updated_at=?
|
||||
WHERE validator_id=?
|
||||
`, ValidatorIdle, now, validatorID); err != nil {
|
||||
if _, err := tx.ExecContext(ctx, freeValidatorSQL, now, validatorID, ipID); err != nil {
|
||||
return err
|
||||
}
|
||||
return tx.Commit()
|
||||
@@ -252,10 +250,7 @@ func (d *DB) MarkFIPOccupied(ctx context.Context, ipID int64, validatorID string
|
||||
return err
|
||||
}
|
||||
if validatorID != "" {
|
||||
if _, err := tx.ExecContext(ctx, `
|
||||
UPDATE validators SET state=?, current_ip_id=NULL, updated_at=?
|
||||
WHERE validator_id=?
|
||||
`, ValidatorIdle, now, validatorID); err != nil {
|
||||
if _, err := tx.ExecContext(ctx, freeValidatorSQL, now, validatorID, ipID); err != nil {
|
||||
return err
|
||||
}
|
||||
}
|
||||
@@ -315,10 +310,7 @@ func (d *DB) RequeueOrFail(ctx context.Context, ipID int64, validatorID string,
|
||||
}
|
||||
|
||||
if validatorID != "" {
|
||||
if _, err := tx.ExecContext(ctx, `
|
||||
UPDATE validators SET state=?, current_ip_id=NULL, updated_at=?
|
||||
WHERE validator_id=?
|
||||
`, ValidatorIdle, now, validatorID); err != nil {
|
||||
if _, err := tx.ExecContext(ctx, freeValidatorSQL, now, validatorID, ipID); err != nil {
|
||||
return err
|
||||
}
|
||||
}
|
||||
@@ -520,9 +512,11 @@ func (d *DB) DeleteIPs(ctx context.Context, addresses []string) (DeleteIPsResult
|
||||
func deleteIPTx(ctx context.Context, tx *sql.Tx, ipID int64) error {
|
||||
now := timeToDB(Now())
|
||||
if _, err := tx.ExecContext(ctx, `
|
||||
UPDATE validators SET state=?, current_ip_id=NULL, updated_at=?
|
||||
UPDATE validators SET current_ip_id=NULL,
|
||||
state = CASE WHEN state=? THEN state ELSE ? END,
|
||||
updated_at=?
|
||||
WHERE current_ip_id=?
|
||||
`, ValidatorIdle, now, ipID); err != nil {
|
||||
`, ValidatorUnreachable, ValidatorIdle, now, ipID); err != nil {
|
||||
return fmt.Errorf("free owning validator: %w", err)
|
||||
}
|
||||
if _, err := tx.ExecContext(ctx, `UPDATE checks SET ip_id=NULL WHERE ip_id=?`, ipID); err != nil {
|
||||
@@ -579,9 +573,11 @@ func (d *DB) ClearAllIPs(ctx context.Context) ([]string, error) {
|
||||
|
||||
now := timeToDB(Now())
|
||||
if _, err := tx.ExecContext(ctx, `
|
||||
UPDATE validators SET state=?, current_ip_id=NULL, updated_at=?
|
||||
UPDATE validators SET current_ip_id=NULL,
|
||||
state = CASE WHEN state=? THEN state ELSE ? END,
|
||||
updated_at=?
|
||||
WHERE current_ip_id IS NOT NULL
|
||||
`, ValidatorIdle, now); err != nil {
|
||||
`, ValidatorUnreachable, ValidatorIdle, now); err != nil {
|
||||
return nil, fmt.Errorf("free owning validators: %w", err)
|
||||
}
|
||||
if _, err := tx.ExecContext(ctx, `UPDATE checks SET ip_id=NULL WHERE ip_id IS NOT NULL`); err != nil {
|
||||
@@ -610,11 +606,15 @@ type FIPRef struct {
|
||||
FIPID string
|
||||
}
|
||||
|
||||
// ListFIPRefs returns every queue row with an attached floating IP (fip_id
|
||||
// set) — typically at most one per validator — so a bulk clear can
|
||||
// disassociate them without loading the whole queue.
|
||||
// ListFIPRefs returns the queue rows that may still hold a floating IP: an
|
||||
// fip_id is set and the row is not finished. A finished row (done, failed,
|
||||
// occupied) keeps its fip_id for display, but its floating IP was already
|
||||
// disassociated before the final state was written, so listing it would only
|
||||
// make a bulk clear issue thousands of pointless cloud calls. At most one row
|
||||
// per validator qualifies, so a clear does not need to load the whole queue.
|
||||
func (d *DB) ListFIPRefs(ctx context.Context) ([]FIPRef, error) {
|
||||
rows, err := d.QueryContext(ctx, `SELECT id, ip_address, fip_id FROM ip_queue WHERE fip_id<>'' ORDER BY id`)
|
||||
rows, err := d.QueryContext(ctx, `SELECT id, ip_address, fip_id FROM ip_queue WHERE fip_id<>'' AND state NOT IN (?, ?, ?) ORDER BY id`,
|
||||
IPDone, IPFailed, IPOccupied)
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
@@ -631,8 +631,7 @@ func (d *DB) ListFIPRefs(ctx context.Context) ([]FIPRef, error) {
|
||||
}
|
||||
|
||||
// ListFIPRefsByAddresses is ListFIPRefs restricted to the given addresses
|
||||
// (unknown addresses and rows without an attached floating IP are simply
|
||||
// absent), using a handful of IN (...) queries instead of one lookup per
|
||||
// (unknown addresses and rows that hold no floating IP are simply absent), using a handful of IN (...) queries instead of one lookup per
|
||||
// address.
|
||||
func (d *DB) ListFIPRefsByAddresses(ctx context.Context, addresses []string) ([]FIPRef, error) {
|
||||
const chunk = 500
|
||||
@@ -643,12 +642,12 @@ func (d *DB) ListFIPRefsByAddresses(ctx context.Context, addresses []string) ([]
|
||||
end = len(addresses)
|
||||
}
|
||||
part := addresses[start:end]
|
||||
args := make([]any, len(part))
|
||||
for i, a := range part {
|
||||
args[i] = a
|
||||
args := []any{IPDone, IPFailed, IPOccupied}
|
||||
for _, a := range part {
|
||||
args = append(args, a)
|
||||
}
|
||||
rows, err := d.QueryContext(ctx,
|
||||
`SELECT id, ip_address, fip_id FROM ip_queue WHERE fip_id<>'' AND ip_address IN (`+placeholders(len(part))+`)`, args...)
|
||||
`SELECT id, ip_address, fip_id FROM ip_queue WHERE fip_id<>'' AND state NOT IN (?, ?, ?) AND ip_address IN (`+placeholders(len(part))+`)`, args...)
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
|
||||
@@ -0,0 +1,186 @@
|
||||
package db
|
||||
|
||||
import (
|
||||
"context"
|
||||
"fmt"
|
||||
"testing"
|
||||
"time"
|
||||
)
|
||||
|
||||
func validatorState(t *testing.T, d *DB, id string) *Validator {
|
||||
t.Helper()
|
||||
v, err := d.GetValidator(testCtx(t), id)
|
||||
if err != nil {
|
||||
t.Fatalf("get validator %s: %v", id, err)
|
||||
}
|
||||
return v
|
||||
}
|
||||
|
||||
func testCtx(t *testing.T) context.Context {
|
||||
t.Helper()
|
||||
return context.Background()
|
||||
}
|
||||
|
||||
// claimFor2 seeds an address and returns its id without claiming it.
|
||||
func claimFor2(t *testing.T, d *DB, addr string) int64 {
|
||||
t.Helper()
|
||||
ctx := testCtx(t)
|
||||
if err := d.SeedQueue(ctx, []string{addr}); err != nil {
|
||||
t.Fatalf("seed %s: %v", addr, err)
|
||||
}
|
||||
ip, err := d.GetIPByAddress(ctx, addr)
|
||||
if err != nil {
|
||||
t.Fatalf("get %s: %v", addr, err)
|
||||
}
|
||||
return ip.ID
|
||||
}
|
||||
|
||||
// claimFor seeds an address and claims it for the validator.
|
||||
func claimFor(t *testing.T, d *DB, addr, validatorID string) *IPQueueItem {
|
||||
t.Helper()
|
||||
ctx := testCtx(t)
|
||||
if err := d.SeedQueue(ctx, []string{addr}); err != nil {
|
||||
t.Fatalf("seed %s: %v", addr, err)
|
||||
}
|
||||
item, err := d.ClaimNextQueued(ctx, validatorID, time.Minute)
|
||||
if err != nil || item == nil {
|
||||
t.Fatalf("claim %s for %s: item=%v err=%v", addr, validatorID, item, err)
|
||||
}
|
||||
return item
|
||||
}
|
||||
|
||||
func TestHeartbeatKeepsAssignedWhenValidatorStillHoldsAnAddress(t *testing.T) {
|
||||
d, ctx := newTestDB(t)
|
||||
_ = d.AdminCreateValidator(ctx, "v1", "p1")
|
||||
_ = d.AdminCreateValidator(ctx, "v2", "p2")
|
||||
item := claimFor(t, d, "1.1.1.1", "v1")
|
||||
_ = d.MarkValidatorUnreachable(ctx, "v1")
|
||||
_ = d.MarkValidatorUnreachable(ctx, "v2")
|
||||
|
||||
if err := d.Heartbeat(ctx, "v1"); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if v := validatorState(t, d, "v1"); v.State != ValidatorAssigned || v.CurrentIPID == nil || *v.CurrentIPID != item.ID {
|
||||
t.Fatalf("v1 after heartbeat: state=%s current_ip=%v, want assigned to %d", v.State, v.CurrentIPID, item.ID)
|
||||
}
|
||||
if err := d.Heartbeat(ctx, "v2"); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if v := validatorState(t, d, "v2"); v.State != ValidatorIdle {
|
||||
t.Fatalf("v2 after heartbeat: state=%s, want idle", v.State)
|
||||
}
|
||||
}
|
||||
|
||||
func TestRegisterValidatorReactivationKeepsAssignedAddress(t *testing.T) {
|
||||
d, ctx := newTestDB(t)
|
||||
_ = d.AdminCreateValidator(ctx, "v1", "p1")
|
||||
item := claimFor(t, d, "1.1.1.1", "v1")
|
||||
_ = d.MarkValidatorUnreachable(ctx, "v1")
|
||||
|
||||
if err := d.RegisterValidator(ctx, "v1", "host", "p1", "v"); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if v := validatorState(t, d, "v1"); v.State != ValidatorAssigned || v.CurrentIPID == nil || *v.CurrentIPID != item.ID {
|
||||
t.Fatalf("after re-register: state=%s current_ip=%v, want assigned to %d", v.State, v.CurrentIPID, item.ID)
|
||||
}
|
||||
}
|
||||
|
||||
func TestStaleReleasesLeaveTheCurrentAddressAlone(t *testing.T) {
|
||||
cases := map[string]func(d *DB, ipID int64) error{
|
||||
"ReleaseFIP": func(d *DB, id int64) error { return d.ReleaseFIP(context.Background(), id, "v1") },
|
||||
"RequeueOrFail": func(d *DB, id int64) error { return d.RequeueOrFail(context.Background(), id, "v1", 3) },
|
||||
"MarkFIPOccupied": func(d *DB, id int64) error { return d.MarkFIPOccupied(context.Background(), id, "v1") },
|
||||
"FreeValidator": func(d *DB, id int64) error { return d.FreeValidator(context.Background(), "v1", id) },
|
||||
}
|
||||
for name, release := range cases {
|
||||
t.Run(name, func(t *testing.T) {
|
||||
d, ctx := newTestDB(t)
|
||||
_ = d.AdminCreateValidator(ctx, "v1", "p1")
|
||||
stale := claimFor(t, d, "1.1.1.1", "v1")
|
||||
// v1 has moved on to another address (set directly: the claim path
|
||||
// itself refuses a validator that is still busy).
|
||||
cur := claimFor2(t, d, "2.2.2.2")
|
||||
if _, err := d.ExecContext(ctx, `UPDATE ip_queue SET state='checking', owner_validator_id='v1' WHERE id=?`, cur); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if _, err := d.ExecContext(ctx, `UPDATE validators SET state='assigned', current_ip_id=? WHERE validator_id='v1'`, cur); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
|
||||
if err := release(d, stale.ID); err != nil {
|
||||
t.Fatalf("%s: %v", name, err)
|
||||
}
|
||||
v := validatorState(t, d, "v1")
|
||||
if v.State != ValidatorAssigned || v.CurrentIPID == nil || *v.CurrentIPID != cur {
|
||||
t.Fatalf("%s freed a validator that holds another address: state=%s current_ip=%v", name, v.State, v.CurrentIPID)
|
||||
}
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
// The right address frees the validator, but an unreachable one stays so.
|
||||
func TestReleaseKeepsUnreachableValidatorUnreachable(t *testing.T) {
|
||||
d, ctx := newTestDB(t)
|
||||
_ = d.AdminCreateValidator(ctx, "v1", "p1")
|
||||
_ = d.AdminCreateValidator(ctx, "v2", "p2")
|
||||
a := claimFor(t, d, "1.1.1.1", "v1")
|
||||
b := claimFor(t, d, "2.2.2.2", "v2")
|
||||
_ = d.MarkValidatorUnreachable(ctx, "v1")
|
||||
|
||||
if err := d.ReleaseFIP(ctx, a.ID, "v1"); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if v := validatorState(t, d, "v1"); v.State != ValidatorUnreachable || v.CurrentIPID != nil {
|
||||
t.Fatalf("v1: state=%s current_ip=%v, want unreachable and empty", v.State, v.CurrentIPID)
|
||||
}
|
||||
if err := d.RequeueOrFail(ctx, b.ID, "v2", 3); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if v := validatorState(t, d, "v2"); v.State != ValidatorIdle || v.CurrentIPID != nil {
|
||||
t.Fatalf("v2: state=%s current_ip=%v, want idle and empty", v.State, v.CurrentIPID)
|
||||
}
|
||||
}
|
||||
|
||||
func TestClaimRefusesValidatorThatStillHoldsAnAddress(t *testing.T) {
|
||||
d, ctx := newTestDB(t)
|
||||
_ = d.AdminCreateValidator(ctx, "v1", "p1")
|
||||
first := claimFor(t, d, "1.1.1.1", "v1")
|
||||
// An old version could leave a validator idle while it still pointed at an address.
|
||||
if _, err := d.ExecContext(ctx, `UPDATE validators SET state='idle' WHERE validator_id='v1'`); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
_ = d.SeedQueue(ctx, []string{"2.2.2.2"})
|
||||
got, err := d.ClaimNextQueued(ctx, "v1", time.Minute)
|
||||
if err != nil || got != nil {
|
||||
t.Fatalf("claim for a validator that holds %d: item=%v err=%v, want nothing", first.ID, got, err)
|
||||
}
|
||||
if ip, _ := d.GetIPByAddress(ctx, "2.2.2.2"); ip.State != IPQueued {
|
||||
t.Fatalf("2.2.2.2 is %s, want queued", ip.State)
|
||||
}
|
||||
}
|
||||
|
||||
func TestListFIPRefsSkipsFinishedAddresses(t *testing.T) {
|
||||
d, ctx := newTestDB(t)
|
||||
var addrs []string
|
||||
for i := 0; i < 6; i++ {
|
||||
addrs = append(addrs, fmt.Sprintf("10.0.0.%d", i+1))
|
||||
}
|
||||
_ = d.SeedQueue(ctx, addrs)
|
||||
states := []string{"done", "failed", "occupied", "awaiting_self_check", "checking", "aggregating"}
|
||||
for i, a := range addrs {
|
||||
if _, err := d.ExecContext(ctx, `UPDATE ip_queue SET state=?, fip_id=? WHERE ip_address=?`, states[i], fmt.Sprintf("fip-%d", i), a); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
}
|
||||
refs, err := d.ListFIPRefs(ctx)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if len(refs) != 3 {
|
||||
t.Fatalf("ListFIPRefs returned %d rows, want the 3 unfinished ones: %+v", len(refs), refs)
|
||||
}
|
||||
byAddr, err := d.ListFIPRefsByAddresses(ctx, addrs)
|
||||
if err != nil || len(byAddr) != 3 {
|
||||
t.Fatalf("ListFIPRefsByAddresses returned %d rows (err %v), want 3", len(byAddr), err)
|
||||
}
|
||||
}
|
||||
@@ -27,24 +27,32 @@ func (d *DB) RegisterValidator(ctx context.Context, validatorID, hostname, osPor
|
||||
}
|
||||
// A brand-new row already lands in ValidatorIdle via the INSERT branch;
|
||||
// a re-registering validator that was 'unregistered' or 'unreachable'
|
||||
// (but not mid-assignment) should also come back to idle.
|
||||
// comes back: to idle when it holds no address, to assigned when it still
|
||||
// does (its current_ip_id is kept, so it must not be handed another).
|
||||
_, err = d.ExecContext(ctx, `
|
||||
UPDATE validators SET state=?, updated_at=?
|
||||
UPDATE validators SET
|
||||
state = CASE WHEN current_ip_id IS NULL THEN ? ELSE ? END,
|
||||
updated_at=?
|
||||
WHERE validator_id=? AND state IN (?, ?)
|
||||
`, ValidatorIdle, now, validatorID, ValidatorUnregistered, ValidatorUnreachable)
|
||||
`, ValidatorIdle, ValidatorAssigned, now, validatorID, ValidatorUnregistered, ValidatorUnreachable)
|
||||
if err != nil {
|
||||
return fmt.Errorf("register validator (reactivate): %w", err)
|
||||
}
|
||||
return nil
|
||||
}
|
||||
|
||||
// Heartbeat records a sign of life. A validator that was marked unreachable
|
||||
// returns to idle if it holds no address, but to assigned if it still does:
|
||||
// it was only silent (for example busy with slow checks), and handing it a
|
||||
// second address while it works on the first would leave the second one
|
||||
// without an owner that can ever pick it up.
|
||||
func (d *DB) Heartbeat(ctx context.Context, validatorID string) error {
|
||||
now := timeToDB(Now())
|
||||
res, err := d.ExecContext(ctx, `
|
||||
UPDATE validators SET last_heartbeat_at=?, updated_at=?,
|
||||
state = CASE WHEN state=? THEN ? ELSE state END
|
||||
state = CASE WHEN state=? THEN (CASE WHEN current_ip_id IS NULL THEN ? ELSE ? END) ELSE state END
|
||||
WHERE validator_id=?
|
||||
`, now, now, ValidatorUnreachable, ValidatorIdle, validatorID)
|
||||
`, now, now, ValidatorUnreachable, ValidatorIdle, ValidatorAssigned, validatorID)
|
||||
if err != nil {
|
||||
return fmt.Errorf("heartbeat: %w", err)
|
||||
}
|
||||
@@ -138,16 +146,59 @@ func (d *DB) MarkValidatorUnreachable(ctx context.Context, validatorID string) e
|
||||
return err
|
||||
}
|
||||
|
||||
// FreeValidator returns a validator to idle with no assigned IP. Used after
|
||||
// an IP finishes (success or failure) or is reclaimed by the lease sweep.
|
||||
func (d *DB) FreeValidator(ctx context.Context, validatorID string) error {
|
||||
_, err := d.ExecContext(ctx, `
|
||||
UPDATE validators SET state=?, current_ip_id=NULL, updated_at=?
|
||||
WHERE validator_id=?
|
||||
`, ValidatorIdle, timeToDB(Now()), validatorID)
|
||||
// freeValidatorSQL releases a validator from the address it holds. It only
|
||||
// applies when the validator's current address is the one being released
|
||||
// (args: now, validator id, ip id): a late release of an old address must not
|
||||
// free a validator that has already moved on to another one. An unreachable
|
||||
// validator stays unreachable until its next heartbeat, so a dead validator
|
||||
// is not handed new addresses just because its lease was reclaimed.
|
||||
var freeValidatorSQL = fmt.Sprintf(`
|
||||
UPDATE validators SET current_ip_id=NULL,
|
||||
state = CASE WHEN state='%s' THEN state ELSE '%s' END,
|
||||
updated_at=?
|
||||
WHERE validator_id=? AND current_ip_id=?`, ValidatorUnreachable, ValidatorIdle)
|
||||
|
||||
// FreeValidator releases a validator from the given address (see
|
||||
// freeValidatorSQL). Used after an IP finishes (success or failure) or is
|
||||
// reclaimed by the lease sweep.
|
||||
func (d *DB) FreeValidator(ctx context.Context, validatorID string, ipID int64) error {
|
||||
_, err := d.ExecContext(ctx, freeValidatorSQL, timeToDB(Now()), validatorID, ipID)
|
||||
return err
|
||||
}
|
||||
|
||||
// ReconcileValidators repairs validators whose state disagrees with the
|
||||
// queue: a validator pointing at an address that no longer exists, is
|
||||
// finished, or belongs to another validator is released; an "assigned"
|
||||
// validator that holds nothing goes back to idle. The invariants normally
|
||||
// hold by construction (every change is one transaction); this heals what a
|
||||
// crash or an older version left behind. It returns the number of validators
|
||||
// repaired.
|
||||
func (d *DB) ReconcileValidators(ctx context.Context) (int64, error) {
|
||||
now := timeToDB(Now())
|
||||
res, err := d.ExecContext(ctx, fmt.Sprintf(`
|
||||
UPDATE validators SET current_ip_id=NULL,
|
||||
state = CASE WHEN state='%s' THEN state ELSE '%s' END,
|
||||
updated_at=?
|
||||
WHERE current_ip_id IS NOT NULL AND NOT EXISTS (
|
||||
SELECT 1 FROM ip_queue q
|
||||
WHERE q.id = validators.current_ip_id
|
||||
AND q.owner_validator_id = validators.validator_id
|
||||
AND q.state NOT IN ('%s','%s','%s'))`,
|
||||
ValidatorUnreachable, ValidatorIdle, IPDone, IPFailed, IPOccupied), now)
|
||||
if err != nil {
|
||||
return 0, fmt.Errorf("reconcile validators: %w", err)
|
||||
}
|
||||
n, _ := res.RowsAffected()
|
||||
res, err = d.ExecContext(ctx, `
|
||||
UPDATE validators SET state=?, updated_at=?
|
||||
WHERE state=? AND current_ip_id IS NULL`, ValidatorIdle, now, ValidatorAssigned)
|
||||
if err != nil {
|
||||
return n, fmt.Errorf("reconcile validators: %w", err)
|
||||
}
|
||||
m, _ := res.RowsAffected()
|
||||
return n + m, nil
|
||||
}
|
||||
|
||||
// AdminCreateValidator registers a brand-new validator via the admin API.
|
||||
// Unlike RegisterValidator (used by the agent's self-registration call),
|
||||
// this refuses to upsert over an existing row.
|
||||
|
||||
@@ -97,8 +97,8 @@ func TestRouteTableIsClassified(t *testing.T) {
|
||||
t.Fatalf("admin route %q is %s, want admin", rt.Pattern, rt.Access)
|
||||
}
|
||||
}
|
||||
if counts["admin"] != 34 || counts["agent"] != 5 || counts["open"] != 7 {
|
||||
t.Fatalf("access counts = %v, want admin=34 agent=5 open=7", counts)
|
||||
if counts["admin"] != 34 || counts["agent"] != 5 || counts["open"] != 8 {
|
||||
t.Fatalf("access counts = %v, want admin=34 agent=5 open=8", counts)
|
||||
}
|
||||
}
|
||||
|
||||
|
||||
@@ -23,6 +23,11 @@ type okResponse struct {
|
||||
OK bool `json:"ok"`
|
||||
}
|
||||
|
||||
type observedIPResponse struct {
|
||||
IP string `json:"ip"`
|
||||
Source string `json:"source"`
|
||||
}
|
||||
|
||||
type checkConfigDTO struct {
|
||||
Type string `json:"type"`
|
||||
Targets []string `json:"targets"`
|
||||
|
||||
@@ -1,17 +1,32 @@
|
||||
package httpapi
|
||||
|
||||
import (
|
||||
"context"
|
||||
"errors"
|
||||
"fmt"
|
||||
"net/http"
|
||||
"net/url"
|
||||
"strconv"
|
||||
"strings"
|
||||
"time"
|
||||
|
||||
"cloudipvalidator/internal/db"
|
||||
"cloudipvalidator/internal/orchestrator"
|
||||
)
|
||||
|
||||
// destructiveOpTimeout bounds a cancel, delete or clear. Such an operation
|
||||
// talks to the cloud as well as the database, and abandoning it halfway
|
||||
// leaves floating IPs detached from rows that still exist (or the reverse),
|
||||
// so it must not die with the client connection: a client that gives up
|
||||
// (a dashboard or curl timeout) only stops waiting for the answer.
|
||||
const destructiveOpTimeout = 10 * time.Minute
|
||||
|
||||
// detachedContext returns a context that ignores cancellation of the request
|
||||
// but keeps its values, with destructiveOpTimeout as the upper bound.
|
||||
func detachedContext(r *http.Request) (context.Context, context.CancelFunc) {
|
||||
return context.WithTimeout(context.WithoutCancel(r.Context()), destructiveOpTimeout)
|
||||
}
|
||||
|
||||
func (s *Server) handleHealthz(w http.ResponseWriter, r *http.Request) {
|
||||
writeJSON(w, http.StatusOK, okResponse{OK: true})
|
||||
}
|
||||
@@ -273,7 +288,9 @@ func toScanStatusDTO(st orchestrator.ScanStatus) scanStatusDTO {
|
||||
// need to be disassociated in OpenStack.
|
||||
func (s *Server) handleAdminCancelIP(w http.ResponseWriter, r *http.Request) {
|
||||
address := r.PathValue("ip")
|
||||
if err := s.Orch.ForceCancel(r.Context(), address); err != nil {
|
||||
ctx, cancel := detachedContext(r)
|
||||
defer cancel()
|
||||
if err := s.Orch.ForceCancel(ctx, address); err != nil {
|
||||
writeDBError(w, err)
|
||||
return
|
||||
}
|
||||
@@ -286,7 +303,9 @@ func (s *Server) handleAdminCancelIP(w http.ResponseWriter, r *http.Request) {
|
||||
// associated floating IP needs disassociating first.
|
||||
func (s *Server) handleAdminDeleteIP(w http.ResponseWriter, r *http.Request) {
|
||||
address := r.PathValue("ip")
|
||||
if err := s.Orch.DeleteIP(r.Context(), address); err != nil {
|
||||
ctx, cancel := detachedContext(r)
|
||||
defer cancel()
|
||||
if err := s.Orch.DeleteIP(ctx, address); err != nil {
|
||||
writeDBError(w, err)
|
||||
return
|
||||
}
|
||||
@@ -306,7 +325,9 @@ func (s *Server) handleAdminDeleteIPs(w http.ResponseWriter, r *http.Request) {
|
||||
writeError(w, http.StatusBadRequest, "addresses must not be empty")
|
||||
return
|
||||
}
|
||||
result, err := s.Orch.DeleteIPs(r.Context(), req.Addresses)
|
||||
ctx, cancel := detachedContext(r)
|
||||
defer cancel()
|
||||
result, err := s.Orch.DeleteIPs(ctx, req.Addresses)
|
||||
if err != nil {
|
||||
writeDBError(w, err)
|
||||
return
|
||||
@@ -320,7 +341,9 @@ func (s *Server) handleAdminDeleteIPs(w http.ResponseWriter, r *http.Request) {
|
||||
// handleAdminClearQueue permanently removes every address currently in the
|
||||
// queue, including those actively being checked.
|
||||
func (s *Server) handleAdminClearQueue(w http.ResponseWriter, r *http.Request) {
|
||||
result, err := s.Orch.ClearQueue(r.Context())
|
||||
ctx, cancel := detachedContext(r)
|
||||
defer cancel()
|
||||
result, err := s.Orch.ClearQueue(ctx)
|
||||
if err != nil {
|
||||
writeDBError(w, err)
|
||||
return
|
||||
|
||||
@@ -1,7 +1,12 @@
|
||||
package httpapi
|
||||
|
||||
import (
|
||||
"database/sql"
|
||||
"errors"
|
||||
"fmt"
|
||||
"net"
|
||||
"net/http"
|
||||
"net/netip"
|
||||
"time"
|
||||
|
||||
"cloudipvalidator/internal/db"
|
||||
@@ -64,6 +69,46 @@ func (s *Server) handleAgentAssignment(w http.ResponseWriter, r *http.Request) {
|
||||
})
|
||||
}
|
||||
|
||||
// handleAgentObservedIP tells a validator which source address control-api
|
||||
// sees for its connection, so the agent's self-check can confirm the floating
|
||||
// IP without a third-party IP-echo service. The route is open, so it is
|
||||
// limited to known validators to keep it from becoming a public "what is my
|
||||
// IP" service.
|
||||
func (s *Server) handleAgentObservedIP(w http.ResponseWriter, r *http.Request) {
|
||||
id := r.PathValue("id")
|
||||
if _, err := s.DB.GetValidator(r.Context(), id); err != nil {
|
||||
if errors.Is(err, sql.ErrNoRows) || errors.Is(err, db.ErrNotFound) {
|
||||
writeError(w, http.StatusNotFound, "unknown validator: "+id)
|
||||
return
|
||||
}
|
||||
writeError(w, http.StatusInternalServerError, err.Error())
|
||||
return
|
||||
}
|
||||
ip, err := clientIP(r)
|
||||
if err != nil {
|
||||
writeError(w, http.StatusInternalServerError, err.Error())
|
||||
return
|
||||
}
|
||||
writeJSON(w, http.StatusOK, observedIPResponse{IP: ip, Source: "remote_addr"})
|
||||
}
|
||||
|
||||
// clientIP returns the peer address of the TCP connection in canonical form
|
||||
// (IPv4-mapped IPv6 unmapped to plain IPv4). It deliberately ignores
|
||||
// X-Forwarded-For / X-Real-IP: validators connect directly, and trusting a
|
||||
// client-supplied header would let a validator forge the address and pass
|
||||
// the self-check.
|
||||
func clientIP(r *http.Request) (string, error) {
|
||||
host, _, err := net.SplitHostPort(r.RemoteAddr)
|
||||
if err != nil {
|
||||
return "", fmt.Errorf("parse remote address %q: %w", r.RemoteAddr, err)
|
||||
}
|
||||
addr, err := netip.ParseAddr(host)
|
||||
if err != nil {
|
||||
return "", fmt.Errorf("parse remote address %q: %w", r.RemoteAddr, err)
|
||||
}
|
||||
return addr.Unmap().String(), nil
|
||||
}
|
||||
|
||||
func (s *Server) handleAgentSelfCheck(w http.ResponseWriter, r *http.Request) {
|
||||
id := r.PathValue("id")
|
||||
var req selfCheckRequest
|
||||
|
||||
@@ -0,0 +1,65 @@
|
||||
package httpapi
|
||||
|
||||
import (
|
||||
"context"
|
||||
"net/http"
|
||||
"testing"
|
||||
"time"
|
||||
|
||||
"cloudipvalidator/internal/openstack"
|
||||
)
|
||||
|
||||
// slowDetachOS delays every Disassociate so the client can give up first.
|
||||
type slowDetachOS struct {
|
||||
*openstack.MockClient
|
||||
delay time.Duration
|
||||
}
|
||||
|
||||
func (s slowDetachOS) DisassociateFloatingIP(ctx context.Context, fipID string) error {
|
||||
time.Sleep(s.delay)
|
||||
return s.MockClient.DisassociateFloatingIP(ctx, fipID)
|
||||
}
|
||||
|
||||
// A client that gives up (a dashboard or curl timeout) must not abort
|
||||
// "clear queue" halfway: on 2026-10-02 the request context was cancelled after
|
||||
// part of the floating IPs were detached, the database was left untouched and
|
||||
// the checks kept running.
|
||||
func TestClearQueueSurvivesClientDisconnect(t *testing.T) {
|
||||
fc, d, orch, mock := newConfigTestHarness(t)
|
||||
ctx := context.Background()
|
||||
mock.Seed("fip-1", "1.2.3.4", "svc")
|
||||
if err := d.RegisterValidator(ctx, "validator-1", "host", "port-1", "v"); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if err := d.SeedQueue(ctx, []string{"1.2.3.4", "5.6.7.8"}); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
orch.Tick(ctx) // validator-1 takes 1.2.3.4 and attaches its floating IP
|
||||
if fip, _ := mock.GetFloatingIPByAddress(ctx, "1.2.3.4"); fip.PortID != "port-1" {
|
||||
t.Fatalf("setup: floating ip not attached (port %q)", fip.PortID)
|
||||
}
|
||||
orch.OS = slowDetachOS{MockClient: mock, delay: 400 * time.Millisecond}
|
||||
|
||||
reqCtx, cancel := context.WithTimeout(ctx, 100*time.Millisecond)
|
||||
defer cancel()
|
||||
// No body: the server only notices that the client has left (and cancels
|
||||
// the request context) when it is not waiting for a request body.
|
||||
req, _ := http.NewRequestWithContext(reqCtx, http.MethodPost, fc.base+"/api/v1/admin/ips/clear", nil)
|
||||
if resp, err := fc.client.Do(req); err == nil {
|
||||
resp.Body.Close()
|
||||
t.Fatalf("the client was expected to time out, got status %d", resp.StatusCode)
|
||||
}
|
||||
|
||||
deadline := time.Now().Add(5 * time.Second)
|
||||
for {
|
||||
ips, _ := d.ListIPs(ctx)
|
||||
fip, _ := mock.GetFloatingIPByAddress(ctx, "1.2.3.4")
|
||||
if len(ips) == 0 && fip.PortID == "" {
|
||||
return
|
||||
}
|
||||
if time.Now().After(deadline) {
|
||||
t.Fatalf("clear did not finish after the client left: %d rows, floating ip on port %q", len(ips), fip.PortID)
|
||||
}
|
||||
time.Sleep(50 * time.Millisecond)
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,79 @@
|
||||
package httpapi
|
||||
|
||||
import (
|
||||
"context"
|
||||
"encoding/json"
|
||||
"net/http"
|
||||
"net/http/httptest"
|
||||
"testing"
|
||||
)
|
||||
|
||||
func TestClientIP(t *testing.T) {
|
||||
cases := []struct {
|
||||
name string
|
||||
remoteAddr string
|
||||
headers map[string]string
|
||||
want string
|
||||
wantErr bool
|
||||
}{
|
||||
{name: "ipv4", remoteAddr: "90.156.213.5:51234", want: "90.156.213.5"},
|
||||
{name: "ipv4 mapped in ipv6", remoteAddr: "[::ffff:90.156.213.5]:51234", want: "90.156.213.5"},
|
||||
{name: "ipv6", remoteAddr: "[2001:db8::7]:443", want: "2001:db8::7"},
|
||||
// A client-supplied header must never change the answer.
|
||||
{name: "forwarded for is ignored", remoteAddr: "90.156.213.5:1",
|
||||
headers: map[string]string{"X-Forwarded-For": "1.2.3.4", "X-Real-IP": "5.6.7.8"}, want: "90.156.213.5"},
|
||||
{name: "no port", remoteAddr: "90.156.213.5", wantErr: true},
|
||||
{name: "not an address", remoteAddr: "host:80", wantErr: true},
|
||||
}
|
||||
for _, tc := range cases {
|
||||
t.Run(tc.name, func(t *testing.T) {
|
||||
r := httptest.NewRequest(http.MethodGet, "/", nil)
|
||||
r.RemoteAddr = tc.remoteAddr
|
||||
for k, v := range tc.headers {
|
||||
r.Header.Set(k, v)
|
||||
}
|
||||
got, err := clientIP(r)
|
||||
if tc.wantErr {
|
||||
if err == nil {
|
||||
t.Fatalf("expected an error, got %q", got)
|
||||
}
|
||||
return
|
||||
}
|
||||
if err != nil || got != tc.want {
|
||||
t.Fatalf("clientIP = %q, %v; want %q", got, err, tc.want)
|
||||
}
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
// The route is open (no token), answers a known validator with the address
|
||||
// control-api sees, and refuses unknown validators and a forged header.
|
||||
func TestObservedIPEndpoint(t *testing.T) {
|
||||
fc, d, _, _ := newConfigTestHarness(t)
|
||||
if err := d.RegisterValidator(context.Background(), "validator-1", "host-1", "port-1", "v0.1"); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
|
||||
req, _ := http.NewRequest(http.MethodGet, fc.base+"/api/v1/agents/validator-1/observed-ip", nil)
|
||||
req.Header.Set("X-Forwarded-For", "8.8.8.8")
|
||||
resp, err := fc.client.Do(req)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
defer resp.Body.Close()
|
||||
if resp.StatusCode != http.StatusOK {
|
||||
t.Fatalf("status = %d, want 200", resp.StatusCode)
|
||||
}
|
||||
var got observedIPResponse
|
||||
if err := json.NewDecoder(resp.Body).Decode(&got); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
// httptest connects from loopback; the forged header must not leak through.
|
||||
if got.IP != "127.0.0.1" || got.Source != "remote_addr" {
|
||||
t.Fatalf("response = %+v, want 127.0.0.1 / remote_addr", got)
|
||||
}
|
||||
|
||||
if resp, _ := fc.do(http.MethodGet, "/api/v1/agents/no-such-validator/observed-ip", nil); resp.StatusCode != http.StatusNotFound {
|
||||
t.Fatalf("unknown validator: status = %d, want 404", resp.StatusCode)
|
||||
}
|
||||
}
|
||||
@@ -39,6 +39,7 @@ func (s *Server) routeTable() []route {
|
||||
{"POST /api/v1/agents/register", s.handleAgentRegister, accessOpen},
|
||||
{"POST /api/v1/agents/{id}/heartbeat", s.handleAgentHeartbeat, accessOpen},
|
||||
{"GET /api/v1/agents/{id}/assignment", s.handleAgentAssignment, accessOpen},
|
||||
{"GET /api/v1/agents/{id}/observed-ip", s.handleAgentObservedIP, accessOpen},
|
||||
{"POST /api/v1/agents/{id}/self-check", s.handleAgentSelfCheck, accessAgent},
|
||||
{"POST /api/v1/agents/{id}/events", s.handleAgentEvent, accessAgent},
|
||||
{"POST /api/v1/agents/{id}/results", s.handleAgentResults, accessAgent},
|
||||
|
||||
@@ -114,6 +114,11 @@ func (o *Orchestrator) leaseTTL() time.Duration {
|
||||
// own goroutine, so validators never wait for each other. In Async mode
|
||||
// Tick does not wait for those goroutines; otherwise it waits for them.
|
||||
func (o *Orchestrator) Tick(ctx context.Context) {
|
||||
if n, err := o.DB.ReconcileValidators(ctx); err != nil {
|
||||
o.Log.Error("reconcile validators", "err", err)
|
||||
} else if n > 0 {
|
||||
o.Log.Warn("repaired validators that disagreed with the queue", "count", n)
|
||||
}
|
||||
if err := o.assignIdleValidators(ctx); err != nil {
|
||||
o.Log.Error("assign idle validators", "err", err)
|
||||
}
|
||||
@@ -146,7 +151,11 @@ func (o *Orchestrator) assignIdleValidators(ctx context.Context) error {
|
||||
}
|
||||
o.Log.Info("claimed ip", "validator", v.ValidatorID, "ip", item.IPAddress, "ip_id", item.ID)
|
||||
v, item := v, item
|
||||
o.spawn("assign:"+v.ValidatorID, func() {
|
||||
// Keyed by address, not validator: the key only guards against
|
||||
// starting the same association twice. A validator-wide key made a
|
||||
// second address claimed while the first was still associating skip
|
||||
// its association and wait for the lease to expire.
|
||||
o.spawn(fmt.Sprintf("assign:%d", item.ID), func() {
|
||||
if err := o.associateFIP(ctx, v.ValidatorID, v.OSPortID, item); err != nil {
|
||||
o.Log.Error("associate fip", "validator", v.ValidatorID, "ip", item.IPAddress, "err", err)
|
||||
}
|
||||
@@ -222,6 +231,25 @@ func (o *Orchestrator) SelfCheckResult(ctx context.Context, validatorID string,
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
// A late report for an address this validator no longer owns (lease
|
||||
// already reclaimed, address cancelled or deleted) must not touch
|
||||
// the floating IP: it may be attached for another validator now.
|
||||
if item.State != db.IPAwaitingSelfCheck || item.OwnerValidatorID == nil || *item.OwnerValidatorID != validatorID {
|
||||
o.Log.Warn("ignoring failed self-check for an address the validator does not hold",
|
||||
"validator", validatorID, "ip_id", ipID, "state", item.State)
|
||||
return nil
|
||||
}
|
||||
// Detach the floating IP before the address goes back to the queue.
|
||||
// requeueOrFail frees the validator in the database but knows nothing
|
||||
// about the cloud: a floating IP left on the validator's port makes
|
||||
// every later association on that port fail with 409 ("fixed IP
|
||||
// already has a floating IP"). Best-effort, like the other release
|
||||
// paths — the database state must be freed even if Neutron hiccups.
|
||||
if item.FIPID != "" {
|
||||
if err := o.OS.DisassociateFloatingIP(ctx, item.FIPID); err != nil {
|
||||
o.Log.Error("disassociate fip after failed self-check", "ip_id", ipID, "fip_id", item.FIPID, "err", err)
|
||||
}
|
||||
}
|
||||
if item.RetryCount+1 > o.Cfg.MaxSelfCheckRetries {
|
||||
o.requeueOrFail(ctx, ipID, validatorID, "self-check failed: "+detail)
|
||||
return nil
|
||||
@@ -447,7 +475,7 @@ func (o *Orchestrator) ForceCancel(ctx context.Context, ipAddress string) error
|
||||
}
|
||||
|
||||
if item.OwnerValidatorID != nil {
|
||||
if err := o.DB.FreeValidator(ctx, *item.OwnerValidatorID); err != nil {
|
||||
if err := o.DB.FreeValidator(ctx, *item.OwnerValidatorID, item.ID); err != nil {
|
||||
return fmt.Errorf("free validator: %w", err)
|
||||
}
|
||||
o.releaseValidatorPorts(ctx, o.validatorsByID(ctx, *item.OwnerValidatorID))
|
||||
@@ -503,11 +531,7 @@ func (o *Orchestrator) DeleteIPs(ctx context.Context, addresses []string) (db.De
|
||||
if err != nil {
|
||||
o.Log.Error("list attached fips before delete", "err", err)
|
||||
}
|
||||
for _, ref := range refs {
|
||||
if err := o.OS.DisassociateFloatingIP(ctx, ref.FIPID); err != nil {
|
||||
o.Log.Error("disassociate fip on delete", "ip_id", ref.IPID, "fip_id", ref.FIPID, "err", err)
|
||||
}
|
||||
}
|
||||
o.disassociateAll(ctx, refs, "delete")
|
||||
|
||||
owners := o.busyValidatorsFor(ctx, addresses)
|
||||
result, err := o.DB.DeleteIPs(ctx, addresses)
|
||||
@@ -519,6 +543,30 @@ func (o *Orchestrator) DeleteIPs(ctx context.Context, addresses []string) (db.De
|
||||
return result, nil
|
||||
}
|
||||
|
||||
// maxParallelDetach bounds how many floating IPs a bulk delete or clear
|
||||
// detaches at the same time (each is one Neutron call).
|
||||
const maxParallelDetach = 8
|
||||
|
||||
// disassociateAll detaches the given floating IPs best-effort, at most
|
||||
// maxParallelDetach at a time; a failure is logged and does not stop the rest.
|
||||
func (o *Orchestrator) disassociateAll(ctx context.Context, refs []db.FIPRef, what string) {
|
||||
var wg sync.WaitGroup
|
||||
sem := make(chan struct{}, maxParallelDetach)
|
||||
for _, ref := range refs {
|
||||
ref := ref
|
||||
sem <- struct{}{}
|
||||
wg.Add(1)
|
||||
go func() {
|
||||
defer wg.Done()
|
||||
defer func() { <-sem }()
|
||||
if err := o.OS.DisassociateFloatingIP(ctx, ref.FIPID); err != nil {
|
||||
o.Log.Error("disassociate fip on "+what, "ip_id", ref.IPID, "fip_id", ref.FIPID, "err", err)
|
||||
}
|
||||
}()
|
||||
}
|
||||
wg.Wait()
|
||||
}
|
||||
|
||||
// ClearQueue deletes every address currently in the queue, regardless of
|
||||
// state — the "delete everything" operation. It is set-based (see
|
||||
// db.ClearAllIPs): O(1) statements however many rows there are. Floating IPs
|
||||
@@ -529,11 +577,7 @@ func (o *Orchestrator) ClearQueue(ctx context.Context) (db.DeleteIPsResult, erro
|
||||
if err != nil {
|
||||
return db.DeleteIPsResult{}, fmt.Errorf("list attached fips: %w", err)
|
||||
}
|
||||
for _, ref := range refs {
|
||||
if err := o.OS.DisassociateFloatingIP(ctx, ref.FIPID); err != nil {
|
||||
o.Log.Error("disassociate fip on clear queue", "ip_id", ref.IPID, "fip_id", ref.FIPID, "err", err)
|
||||
}
|
||||
}
|
||||
o.disassociateAll(ctx, refs, "clear queue")
|
||||
|
||||
deleted, err := o.DB.ClearAllIPs(ctx)
|
||||
if err != nil {
|
||||
|
||||
@@ -1086,3 +1086,79 @@ func TestClearQueueFreesAllValidatorPorts(t *testing.T) {
|
||||
t.Fatalf("port-2: expected only the foreign 9.9.9.9 to remain, got %+v", got)
|
||||
}
|
||||
}
|
||||
|
||||
// A failed self-check must detach the floating IP in the cloud before the
|
||||
// address returns to the queue; otherwise the validator's port keeps it and
|
||||
// every later association on that port fails with 409.
|
||||
func TestFailedSelfCheckDetachesFloatingIP(t *testing.T) {
|
||||
ctx := context.Background()
|
||||
o, d, mock := newTestOrchestrator(t, 180)
|
||||
|
||||
mock.Seed("fip-1", "1.2.3.4", "svc-project")
|
||||
if err := d.RegisterValidator(ctx, "validator-1", "host-1", "port-1", "v0.1"); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if err := d.SeedQueue(ctx, []string{"1.2.3.4"}); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
o.Tick(ctx)
|
||||
if got := portFIPs(t, mock, "port-1"); len(got) != 1 {
|
||||
t.Fatalf("setup: expected the floating ip on port-1, got %d", len(got))
|
||||
}
|
||||
ip, err := d.GetIPByAddress(ctx, "1.2.3.4")
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
|
||||
if err := o.SelfCheckResult(ctx, "validator-1", ip.ID, false, "ip echo timeout"); err != nil {
|
||||
t.Fatalf("self-check result: %v", err)
|
||||
}
|
||||
if got := portFIPs(t, mock, "port-1"); len(got) != 0 {
|
||||
t.Fatalf("port-1 still holds the floating ip after a failed self-check: %+v", got)
|
||||
}
|
||||
after, err := d.GetIPByAddress(ctx, "1.2.3.4")
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if after.State != db.IPQueued {
|
||||
t.Fatalf("expected the address back in queued, got %s", after.State)
|
||||
}
|
||||
|
||||
// The validator must be able to take the next claim: a second Tick
|
||||
// associates the same address again without a conflict.
|
||||
o.Tick(ctx)
|
||||
if got := portFIPs(t, mock, "port-1"); len(got) != 1 {
|
||||
t.Fatalf("expected the retry to attach the floating ip again, got %d", len(got))
|
||||
}
|
||||
}
|
||||
|
||||
// A failed self-check for an address the validator no longer holds is a late
|
||||
// report: it must not detach a floating IP that now belongs to someone else.
|
||||
func TestLateFailedSelfCheckIsIgnored(t *testing.T) {
|
||||
ctx := context.Background()
|
||||
o, d, mock := newTestOrchestrator(t, 180)
|
||||
|
||||
mock.Seed("fip-1", "1.2.3.4", "svc-project")
|
||||
for i, id := range []string{"validator-1", "validator-2"} {
|
||||
if err := d.RegisterValidator(ctx, id, "h", fmt.Sprintf("port-%d", i+1), "v"); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
}
|
||||
if err := d.SeedQueue(ctx, []string{"1.2.3.4"}); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
o.Tick(ctx) // validator-1 (first by id) holds 1.2.3.4
|
||||
ip, _ := d.GetIPByAddress(ctx, "1.2.3.4")
|
||||
|
||||
// validator-2 reports a failure for an address it does not own.
|
||||
if err := o.SelfCheckResult(ctx, "validator-2", ip.ID, false, "late"); err != nil {
|
||||
t.Fatalf("self-check result: %v", err)
|
||||
}
|
||||
if got := portFIPs(t, mock, "port-1"); len(got) != 1 {
|
||||
t.Fatalf("a late report from another validator detached the floating ip: %+v", got)
|
||||
}
|
||||
cur, _ := d.GetIPByAddress(ctx, "1.2.3.4")
|
||||
if cur.State != db.IPAwaitingSelfCheck {
|
||||
t.Fatalf("a late report changed the address state to %s", cur.State)
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,401 @@
|
||||
package orchestrator
|
||||
|
||||
import (
|
||||
"context"
|
||||
"fmt"
|
||||
"math/rand"
|
||||
"sync/atomic"
|
||||
"testing"
|
||||
"time"
|
||||
|
||||
"cloudipvalidator/internal/db"
|
||||
"cloudipvalidator/internal/openstack"
|
||||
)
|
||||
|
||||
// finishChecks reports a successful egress run and all three inbound sites
|
||||
// for the address, so the next Tick aggregates it and releases the validator.
|
||||
func finishChecks(t *testing.T, o *Orchestrator, ip *db.IPQueueItem, validatorID string) {
|
||||
t.Helper()
|
||||
ctx := context.Background()
|
||||
if err := o.RecordCheck(ctx, db.Check{
|
||||
IPID: ip.ID, IPAddress: ip.IPAddress, AttemptNumber: ip.AttemptNumber,
|
||||
ValidatorID: validatorID, Source: db.SourceEgress, CheckType: "https",
|
||||
Target: "https://example.test", Success: true, CheckedAt: db.Now(),
|
||||
}); err != nil {
|
||||
t.Fatalf("record egress check: %v", err)
|
||||
}
|
||||
if err := o.MarkEgressComplete(ctx, ip.ID); err != nil {
|
||||
t.Fatalf("mark egress complete: %v", err)
|
||||
}
|
||||
for site := 1; site <= 3; site++ {
|
||||
for _, ct := range []string{"tcp-22", "ssh", "tcp-80", "icmp"} {
|
||||
if err := o.RecordCheck(ctx, db.Check{
|
||||
IPID: ip.ID, IPAddress: ip.IPAddress, AttemptNumber: ip.AttemptNumber,
|
||||
Source: db.InboundSource(site), CheckType: ct, Target: ip.IPAddress,
|
||||
Success: true, CheckedAt: db.Now(),
|
||||
}); err != nil {
|
||||
t.Fatalf("record inbound check: %v", err)
|
||||
}
|
||||
}
|
||||
if err := o.MarkSiteComplete(ctx, ip.ID, site); err != nil {
|
||||
t.Fatalf("mark site complete: %v", err)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// The incident of 2026-10-02: a validator busy with slow checks goes silent,
|
||||
// is marked unreachable, and its next heartbeat used to return it to idle
|
||||
// although it still held the address. It was then handed a second address,
|
||||
// whose association never ran, and both stalled until their leases expired.
|
||||
func TestSilentValidatorKeepsItsAddressAfterHeartbeat(t *testing.T) {
|
||||
ctx := context.Background()
|
||||
o, d, mock := newTestOrchestrator(t, 180)
|
||||
mock.Seed("fip-a", "1.1.1.1", "svc")
|
||||
mock.Seed("fip-b", "2.2.2.2", "svc")
|
||||
if err := d.RegisterValidator(ctx, "validator-1", "host", "port-1", "v"); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if err := d.SeedQueue(ctx, []string{"1.1.1.1", "2.2.2.2"}); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
|
||||
o.Tick(ctx) // validator-1 claims 1.1.1.1
|
||||
a, _ := d.GetIPByAddress(ctx, "1.1.1.1")
|
||||
if err := o.SelfCheckResult(ctx, "validator-1", a.ID, true, "ok"); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
|
||||
// The agent goes silent for longer than heartbeat_timeout_seconds ...
|
||||
if err := d.MarkValidatorUnreachable(ctx, "validator-1"); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
// ... and then speaks again while the checks are still running.
|
||||
if err := d.Heartbeat(ctx, "validator-1"); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
v, _ := d.GetValidator(ctx, "validator-1")
|
||||
if v.State != db.ValidatorAssigned || v.CurrentIPID == nil || *v.CurrentIPID != a.ID {
|
||||
t.Fatalf("after heartbeat: state=%s current_ip=%v, want assigned to %d", v.State, v.CurrentIPID, a.ID)
|
||||
}
|
||||
|
||||
o.Tick(ctx)
|
||||
if b, _ := d.GetIPByAddress(ctx, "2.2.2.2"); b.State != db.IPQueued {
|
||||
t.Fatalf("the busy validator was given a second address: 2.2.2.2 is %s", b.State)
|
||||
}
|
||||
|
||||
// The first address finishes: only now the validator may take the next one.
|
||||
a, _ = d.GetIP(ctx, a.ID)
|
||||
finishChecks(t, o, a, "validator-1")
|
||||
o.Tick(ctx) // aggregates and releases 1.1.1.1
|
||||
o.Tick(ctx) // validator-1 is idle again and claims 2.2.2.2
|
||||
if b, _ := d.GetIPByAddress(ctx, "2.2.2.2"); b.State != db.IPAwaitingSelfCheck {
|
||||
t.Fatalf("2.2.2.2 is %s, want awaiting_self_check once the validator is free", b.State)
|
||||
}
|
||||
}
|
||||
|
||||
// A dead validator whose lease was reclaimed must stay out of rotation until
|
||||
// it speaks again; before, reclaiming set it idle and it was handed a fresh
|
||||
// address every lease period, burning each address's retries.
|
||||
func TestUnreachableValidatorGetsNoAddressesAfterLeaseReclaim(t *testing.T) {
|
||||
ctx := context.Background()
|
||||
o, d, mock := newTestOrchestrator(t, 1)
|
||||
mock.Seed("fip-a", "1.1.1.1", "svc")
|
||||
mock.Seed("fip-b", "2.2.2.2", "svc")
|
||||
_ = d.RegisterValidator(ctx, "validator-1", "host", "port-1", "v")
|
||||
_ = d.SeedQueue(ctx, []string{"1.1.1.1", "2.2.2.2"})
|
||||
|
||||
o.Tick(ctx) // claims 1.1.1.1, the agent never answers
|
||||
if err := d.MarkValidatorUnreachable(ctx, "validator-1"); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
time.Sleep(1100 * time.Millisecond)
|
||||
o.Cfg.LeaseTTLSeconds = 180
|
||||
o.Tick(ctx) // lease sweep reclaims 1.1.1.1
|
||||
|
||||
a, _ := d.GetIPByAddress(ctx, "1.1.1.1")
|
||||
if a.RetryCount != 1 {
|
||||
t.Fatalf("retry_count = %d, want the lease to be reclaimed once", a.RetryCount)
|
||||
}
|
||||
v, _ := d.GetValidator(ctx, "validator-1")
|
||||
if v.State != db.ValidatorUnreachable || v.CurrentIPID != nil {
|
||||
t.Fatalf("validator state=%s current_ip=%v, want unreachable and empty", v.State, v.CurrentIPID)
|
||||
}
|
||||
o.Tick(ctx)
|
||||
if a, _ := d.GetIPByAddress(ctx, "1.1.1.1"); a.State != db.IPQueued {
|
||||
t.Fatalf("an unreachable validator was handed an address: 1.1.1.1 is %s", a.State)
|
||||
}
|
||||
|
||||
if err := d.Heartbeat(ctx, "validator-1"); err != nil { // it is back
|
||||
t.Fatal(err)
|
||||
}
|
||||
if v, _ := d.GetValidator(ctx, "validator-1"); v.State != db.ValidatorIdle {
|
||||
t.Fatalf("after heartbeat state=%s, want idle (it holds nothing)", v.State)
|
||||
}
|
||||
o.Tick(ctx)
|
||||
if a, _ := d.GetIPByAddress(ctx, "1.1.1.1"); a.State != db.IPAwaitingSelfCheck {
|
||||
t.Fatalf("1.1.1.1 is %s, want it picked up again", a.State)
|
||||
}
|
||||
}
|
||||
|
||||
// countingOS counts Disassociate calls.
|
||||
type countingOS struct {
|
||||
*openstack.MockClient
|
||||
disassociations int32
|
||||
}
|
||||
|
||||
func (c *countingOS) DisassociateFloatingIP(ctx context.Context, fipID string) error {
|
||||
atomic.AddInt32(&c.disassociations, 1)
|
||||
return c.MockClient.DisassociateFloatingIP(ctx, fipID)
|
||||
}
|
||||
|
||||
// Finished addresses keep their fip_id for display but their floating IP is
|
||||
// already free; clearing the queue must not make a cloud call for each of
|
||||
// them (on 2026-10-02 that was >1000 sequential calls: 256 s).
|
||||
func TestClearQueueDoesNotTouchFinishedAddresses(t *testing.T) {
|
||||
ctx := context.Background()
|
||||
o, d, mock := newTestOrchestrator(t, 180)
|
||||
cos := &countingOS{MockClient: mock}
|
||||
o.OS = cos
|
||||
|
||||
const finished = 300
|
||||
var addrs []string
|
||||
for i := 0; i < finished; i++ {
|
||||
addrs = append(addrs, fmt.Sprintf("10.9.%d.%d", i/200, i%200+1))
|
||||
}
|
||||
if err := d.SeedQueue(ctx, addrs); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
for i, a := range addrs {
|
||||
state := []string{db.IPDone, db.IPFailed, db.IPOccupied}[i%3]
|
||||
if _, err := d.ExecContext(ctx, `UPDATE ip_queue SET state=?, fip_id=? WHERE ip_address=?`, state, fmt.Sprintf("fip-old-%d", i), a); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
}
|
||||
// One address really holds a floating IP.
|
||||
mock.Seed("fip-live", "1.2.3.4", "svc")
|
||||
_ = d.RegisterValidator(ctx, "validator-1", "host", "port-1", "v")
|
||||
_ = d.SeedQueue(ctx, []string{"1.2.3.4"})
|
||||
live, _ := d.GetIPByAddress(ctx, "1.2.3.4")
|
||||
if _, err := d.ExecContext(ctx, `UPDATE ip_queue SET sequence=-1 WHERE id=?`, live.ID); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
o.Tick(ctx) // validator-1 takes the live address (lowest sequence) and attaches it
|
||||
if got := portFIPs(t, mock, "port-1"); len(got) != 1 {
|
||||
t.Fatalf("setup: the live address is not attached (%d floating ips on port-1)", len(got))
|
||||
}
|
||||
atomic.StoreInt32(&cos.disassociations, 0)
|
||||
|
||||
start := time.Now()
|
||||
if _, err := o.ClearQueue(ctx); err != nil {
|
||||
t.Fatalf("clear queue: %v", err)
|
||||
}
|
||||
// One call for the live address; the port sweep finds nothing left.
|
||||
if n := atomic.LoadInt32(&cos.disassociations); n != 1 {
|
||||
t.Fatalf("clear queue made %d disassociate calls, want 1 (finished addresses must be skipped)", n)
|
||||
}
|
||||
if got := portFIPs(t, mock, "port-1"); len(got) != 0 {
|
||||
t.Fatalf("the live floating ip is still attached: %+v", got)
|
||||
}
|
||||
if left, _ := d.ListIPs(ctx); len(left) != 0 {
|
||||
t.Fatalf("%d rows left after clear", len(left))
|
||||
}
|
||||
t.Logf("clear of %d rows took %s", finished+1, time.Since(start))
|
||||
}
|
||||
|
||||
// A bulk detach runs in parallel but never above maxParallelDetach at once.
|
||||
func TestDisassociateAllIsBoundedAndParallel(t *testing.T) {
|
||||
o, _, mock := newTestOrchestrator(t, 180)
|
||||
g := &gaugeOS{MockClient: mock, delay: 30 * time.Millisecond}
|
||||
o.OS = g
|
||||
var refs []db.FIPRef
|
||||
for i := 0; i < 40; i++ {
|
||||
id := fmt.Sprintf("fip-%d", i)
|
||||
mock.Seed(id, fmt.Sprintf("10.8.0.%d", i+1), "svc")
|
||||
refs = append(refs, db.FIPRef{IPID: int64(i), IPAddress: id, FIPID: id})
|
||||
}
|
||||
start := time.Now()
|
||||
o.disassociateAll(context.Background(), refs, "test")
|
||||
elapsed := time.Since(start)
|
||||
if peak := atomic.LoadInt32(&g.peak); peak < 2 || peak > maxParallelDetach {
|
||||
t.Fatalf("peak concurrency %d, want between 2 and %d", peak, maxParallelDetach)
|
||||
}
|
||||
if elapsed > 600*time.Millisecond { // 40 x 30 ms sequentially is 1.2 s
|
||||
t.Fatalf("took %s: the detach is not parallel", elapsed)
|
||||
}
|
||||
}
|
||||
|
||||
type gaugeOS struct {
|
||||
*openstack.MockClient
|
||||
delay time.Duration
|
||||
cur, peak int32
|
||||
}
|
||||
|
||||
func (g *gaugeOS) DisassociateFloatingIP(ctx context.Context, fipID string) error {
|
||||
n := atomic.AddInt32(&g.cur, 1)
|
||||
for {
|
||||
p := atomic.LoadInt32(&g.peak)
|
||||
if n <= p || atomic.CompareAndSwapInt32(&g.peak, p, n) {
|
||||
break
|
||||
}
|
||||
}
|
||||
time.Sleep(g.delay)
|
||||
atomic.AddInt32(&g.cur, -1)
|
||||
return g.MockClient.DisassociateFloatingIP(ctx, fipID)
|
||||
}
|
||||
|
||||
// Rows that disagree with the queue are repaired by ReconcileValidators (run on every Tick).
|
||||
func TestReconcileRepairsInconsistentValidators(t *testing.T) {
|
||||
ctx := context.Background()
|
||||
_, d, _ := newTestOrchestrator(t, 180)
|
||||
_ = d.RegisterValidator(ctx, "validator-1", "host", "port-1", "v")
|
||||
_ = d.RegisterValidator(ctx, "validator-2", "host", "port-2", "v")
|
||||
_ = d.SeedQueue(ctx, []string{"1.1.1.1", "2.2.2.2"})
|
||||
|
||||
finished, _ := d.GetIPByAddress(ctx, "1.1.1.1")
|
||||
foreign, _ := d.GetIPByAddress(ctx, "2.2.2.2")
|
||||
// validator-1 still points at a finished address; validator-2 at an
|
||||
// address owned by somebody else (what an older version could leave behind).
|
||||
for _, q := range []string{
|
||||
fmt.Sprintf(`UPDATE ip_queue SET state='done' WHERE id=%d`, finished.ID),
|
||||
fmt.Sprintf(`UPDATE ip_queue SET state='checking', owner_validator_id='validator-1' WHERE id=%d`, foreign.ID),
|
||||
fmt.Sprintf(`UPDATE validators SET state='assigned', current_ip_id=%d WHERE validator_id='validator-1'`, finished.ID),
|
||||
fmt.Sprintf(`UPDATE validators SET state='assigned', current_ip_id=%d WHERE validator_id='validator-2'`, foreign.ID),
|
||||
} {
|
||||
if _, err := d.ExecContext(ctx, q); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
}
|
||||
n, err := d.ReconcileValidators(ctx)
|
||||
if err != nil || n != 2 {
|
||||
t.Fatalf("reconcile repaired %d validators (err %v), want 2", n, err)
|
||||
}
|
||||
for _, id := range []string{"validator-1", "validator-2"} {
|
||||
v, _ := d.GetValidator(ctx, id)
|
||||
if v.State != db.ValidatorIdle || v.CurrentIPID != nil {
|
||||
t.Fatalf("%s: state=%s current_ip=%v, want idle", id, v.State, v.CurrentIPID)
|
||||
}
|
||||
}
|
||||
// A consistent validator is left alone.
|
||||
if n, _ := d.ReconcileValidators(ctx); n != 0 {
|
||||
t.Fatalf("second reconcile repaired %d, want 0", n)
|
||||
}
|
||||
}
|
||||
|
||||
// Randomised run: many validators, random silences and association failures.
|
||||
// After every tick each validator holds at most one address, the validator
|
||||
// and the address agree on who holds what, and no lease is ever reclaimed.
|
||||
func TestRandomFlowKeepsOneAddressPerValidator(t *testing.T) {
|
||||
ctx := context.Background()
|
||||
rng := rand.New(rand.NewSource(42))
|
||||
o, d, mock := newTestOrchestrator(t, 180)
|
||||
|
||||
const validators, addresses = 12, 60
|
||||
var addrs []string
|
||||
for i := 0; i < validators; i++ {
|
||||
_ = d.RegisterValidator(ctx, fmt.Sprintf("validator-%d", i+1), "host", fmt.Sprintf("port-%d", i+1), "v")
|
||||
}
|
||||
for i := 0; i < addresses; i++ {
|
||||
a := fmt.Sprintf("10.7.%d.%d", i/200, i%200+1)
|
||||
addrs = append(addrs, a)
|
||||
mock.Seed(fmt.Sprintf("fip-%d", i), a, "svc")
|
||||
}
|
||||
if err := d.SeedQueue(ctx, addrs); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
|
||||
checkInvariants := func(tick int) {
|
||||
t.Helper()
|
||||
vs, _ := d.ListValidators(ctx)
|
||||
holders := map[int64]string{}
|
||||
for _, v := range vs {
|
||||
if v.CurrentIPID == nil {
|
||||
if v.State == db.ValidatorAssigned {
|
||||
t.Fatalf("tick %d: %s is assigned but holds nothing", tick, v.ValidatorID)
|
||||
}
|
||||
continue
|
||||
}
|
||||
ip, err := d.GetIP(ctx, *v.CurrentIPID)
|
||||
if err != nil {
|
||||
t.Fatalf("tick %d: %s points at a missing address %d", tick, v.ValidatorID, *v.CurrentIPID)
|
||||
}
|
||||
if ip.OwnerValidatorID == nil || *ip.OwnerValidatorID != v.ValidatorID {
|
||||
t.Fatalf("tick %d: %s holds %s but its owner is %v", tick, v.ValidatorID, ip.IPAddress, ip.OwnerValidatorID)
|
||||
}
|
||||
if ip.State == db.IPDone || ip.State == db.IPFailed || ip.State == db.IPOccupied {
|
||||
t.Fatalf("tick %d: %s still holds the finished address %s", tick, v.ValidatorID, ip.IPAddress)
|
||||
}
|
||||
if prev, ok := holders[ip.ID]; ok {
|
||||
t.Fatalf("tick %d: %s is held by both %s and %s", tick, ip.IPAddress, prev, v.ValidatorID)
|
||||
}
|
||||
holders[ip.ID] = v.ValidatorID
|
||||
}
|
||||
// Every owned, unfinished address is the current one of its owner.
|
||||
ips, _ := d.ListIPs(ctx)
|
||||
perOwner := map[string]int{}
|
||||
for _, ip := range ips {
|
||||
if ip.OwnerValidatorID != nil && ip.State != db.IPDone && ip.State != db.IPFailed && ip.State != db.IPOccupied {
|
||||
perOwner[*ip.OwnerValidatorID]++
|
||||
}
|
||||
}
|
||||
for owner, n := range perOwner {
|
||||
if n > 1 {
|
||||
t.Fatalf("tick %d: %s owns %d unfinished addresses", tick, owner, n)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
checkTicks := map[int64]int{} // address id -> ticks spent in checking
|
||||
done := false
|
||||
for tick := 1; tick <= 600 && !done; tick++ {
|
||||
// Occasionally an association fails (409 on a port).
|
||||
if rng.Intn(15) == 0 {
|
||||
mock.AssociateFailures = map[string]error{fmt.Sprintf("fip-%d", rng.Intn(addresses)): fmt.Errorf("409 conflict")}
|
||||
}
|
||||
o.Tick(ctx)
|
||||
|
||||
vs, _ := d.ListValidators(ctx)
|
||||
for _, v := range vs {
|
||||
// A busy agent sometimes goes silent, then speaks again.
|
||||
if v.CurrentIPID != nil && v.State == db.ValidatorAssigned && rng.Intn(10) == 0 {
|
||||
_ = d.MarkValidatorUnreachable(ctx, v.ValidatorID)
|
||||
}
|
||||
if v.State == db.ValidatorUnreachable && rng.Intn(3) == 0 {
|
||||
_ = d.Heartbeat(ctx, v.ValidatorID)
|
||||
}
|
||||
item, _, err := o.AssignmentForValidator(ctx, v.ValidatorID)
|
||||
if err != nil || item == nil {
|
||||
continue
|
||||
}
|
||||
if item.State == db.IPAwaitingSelfCheck {
|
||||
_ = o.SelfCheckResult(ctx, v.ValidatorID, item.ID, true, "ok")
|
||||
continue
|
||||
}
|
||||
checkTicks[item.ID]++
|
||||
if checkTicks[item.ID] >= 1+rng.Intn(4) {
|
||||
finishChecks(t, o, item, v.ValidatorID)
|
||||
delete(checkTicks, item.ID)
|
||||
}
|
||||
}
|
||||
checkInvariants(tick)
|
||||
|
||||
counts, _, _ := d.CountIPsByState(ctx)
|
||||
unfinished := 0
|
||||
for st, n := range counts {
|
||||
if st != db.IPDone && st != db.IPFailed && st != db.IPOccupied {
|
||||
unfinished += n
|
||||
}
|
||||
}
|
||||
done = unfinished == 0
|
||||
}
|
||||
if !done {
|
||||
counts, _, _ := d.CountIPsByState(ctx)
|
||||
t.Fatalf("not all addresses finished: %v", counts)
|
||||
}
|
||||
var expired int
|
||||
if err := d.QueryRowContext(ctx, `SELECT COUNT(*) FROM events WHERE event_type='lease_expired'`).Scan(&expired); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if expired != 0 {
|
||||
t.Fatalf("%d leases were reclaimed, want 0", expired)
|
||||
}
|
||||
}
|
||||
@@ -13,7 +13,19 @@ poll_interval_seconds: 5
|
||||
|
||||
self_check:
|
||||
timeout_seconds: 10
|
||||
# Must be a resource genuinely outside the cloud project — OpenStack only
|
||||
# Способы самопроверки в порядке приоритета (допустимо: ip_echo,
|
||||
# control_api); по умолчанию [ip_echo]. Самопроверка успешна, если адрес
|
||||
# подтвердил любой способ: пробуются по порядку, остановка на первом
|
||||
# успешном, к следующему переходим и при отсутствии ответа, и при
|
||||
# несовпадении адреса. Таймаут timeout_seconds действует на каждый способ
|
||||
# отдельно (зависший первый способ не лишает второй времени). control_api спрашивает у control-api, с какого адреса он видит
|
||||
# это соединение; рекомендуется [control_api, ip_echo], когда control-api
|
||||
# стоит вне облака и валидатор ходит к нему напрямую (через внешнюю сеть).
|
||||
# Ограничение: если control-api достижим по внутренней сети облака, он
|
||||
# увидит частный адрес валидатора и control_api всегда даст несовпадение —
|
||||
# тогда оставьте только ip_echo (или он сработает вторым в списке).
|
||||
methods: [control_api, ip_echo]
|
||||
# Used by the ip_echo method. Must be a resource genuinely outside the cloud project — OpenStack only
|
||||
# applies floating-IP SNAT to traffic leaving via the external network,
|
||||
# so anything reachable over the project's internal network (including
|
||||
# control-api itself, if it's on the same internal network) would report
|
||||
|
||||
@@ -110,6 +110,9 @@ control_api_url: "http://127.0.0.1:28080"
|
||||
poll_interval_seconds: 1
|
||||
self_check:
|
||||
timeout_seconds: 5
|
||||
# ip_echo by default; E2E_SELF_CHECK_METHODS="[control_api]" runs the same
|
||||
# scenario through control-api's observed-ip route (loopback sees 127.0.0.1).
|
||||
methods: ${E2E_SELF_CHECK_METHODS:-[ip_echo]}
|
||||
# Points at the local httpstub's /ip echo route rather than a real
|
||||
# internet IP-echo service — self-check only produces a meaningful
|
||||
# signal here because the "floating IP" under test (127.0.0.1) is
|
||||
|
||||
Reference in new issue
Block a user