Retry a failed self-check on another validator; add the self-check failure ceiling

A validator that failed the self-check of an address no longer gets that address
again in the current round (ClaimNextQueued skips it); the validator itself stays
in service and takes all other addresses. The verdict fail is set when the number
of failed self-checks of an address reaches settings.self_check_max_attempts
(1..50, default 5, independent of the number of validators); max_retries and
retry_count are no longer used for self-check. If every working validator has
already failed the address, a new round starts and the exclusions lapse.

Migration 0012: ip_self_check_failures (permanent history per registry address),
ip_queue.sc_failures and sc_round_start_cycle (cycle_id is used instead of
attempt_number, which restarts when a queue row is recreated), the setting.
db.FailSelfCheck does it in one transaction; re-submission starts a new series.
API: self_check_max_attempts in GET/PUT /admin/config/orchestrator,
self_check_failed_on in /admin/ips/{ip} and /admin/registry/{ip}. Dashboard: the
field on /settings and the line "Self-check не прошёл на: ..." on the address
pages. Docs, plan and summary in docs/changes/.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
This commit is contained in:
ayurishchevandClaude Sonnet 5.5 committed 2026-10-04 09:50:12 +03:00
1 parent b7669c9e41
commit e95b5eb7d5
34 files changed
+1085 -75

No files matched your search

+3 -2
View File
@@ -32,7 +32,7 @@ docker compose up -d --build # весь стенд на одной
| `openstack.auth_method` | `token` (готовый токен проекта из `OS_TOKEN`, сам не обновляется) или `password` (логин/пароль Keystone, токен перевыпускается автоматически) | | `openstack.auth_method` | `token` (готовый токен проекта из `OS_TOKEN`, сам не обновляется) или `password` (логин/пароль Keystone, токен перевыпускается автоматически) |
| `openstack.*_env` | Имена переменных окружения: `OS_AUTH_URL`, `OS_PROJECT_ID`, `OS_REGION_NAME`, `OS_INTERFACE`, `OS_TOKEN` либо `OS_USERNAME`/`OS_USER_DOMAIN_NAME`/`OS_PASSWORD` | | `openstack.*_env` | Имена переменных окружения: `OS_AUTH_URL`, `OS_PROJECT_ID`, `OS_REGION_NAME`, `OS_INTERFACE`, `OS_TOKEN` либо `OS_USERNAME`/`OS_USER_DOMAIN_NAME`/`OS_PASSWORD` |
| `orchestrator.poll_interval_seconds` | Период такта оркестратора (5) | | `orchestrator.poll_interval_seconds` | Период такта оркестратора (5) |
| `orchestrator.self_check_timeout_seconds`, `max_self_check_retries` | Ожидание self-check валидатора (60) и число его повторов (3) | | `orchestrator.self_check_timeout_seconds`, `max_self_check_retries` | Ожидание self-check валидатора (60). `max_self_check_retries` **устарело** и не используется: потолок провалов self-check задаётся настройкой `self_check_max_attempts` (по умолчанию 5, `/settings`) |
| `orchestrator.checking_window_seconds` | Сколько ждать результаты проверок, прежде чем подвести итог; отсчёт от начала проверки (120) | | `orchestrator.checking_window_seconds` | Сколько ждать результаты проверок, прежде чем подвести итог; отсчёт от начала проверки (120) |
| `orchestrator.lease_ttl_seconds`, `max_retries` | Лизинг адреса за валидатором (180) и число возвратов в очередь при его истечении (3) | | `orchestrator.lease_ttl_seconds`, `max_retries` | Лизинг адреса за валидатором (180) и число возвратов в очередь при его истечении (3) |
| `orchestrator.heartbeat_timeout_seconds` | После скольких секунд тишины валидатор или площадка считаются потерянными (30) | | `orchestrator.heartbeat_timeout_seconds` | После скольких секунд тишины валидатор или площадка считаются потерянными (30) |
@@ -169,7 +169,7 @@ docs/ документация и планы доработок
| `/registry`, `/registry/{ip}` | Реестр всех адресов (постранично) и полная история проверок адреса; поиск, фильтр и страница сохраняются в адресной строке. Последний результат разделён на уровни Egress и Ingress: «успешно из всего» по каждому и по типам проверок (icmp, ssh, tcp, https…) | | `/registry`, `/registry/{ip}` | Реестр всех адресов (постранично) и полная история проверок адреса; поиск, фильтр и страница сохраняются в адресной строке. Последний результат разделён на уровни Egress и Ingress: «успешно из всего» по каждому и по типам проверок (icmp, ssh, tcp, https…) |
| `/analytics` | Аналитика одного завершённого запуска: показатели, причины `partial`, подсети, провалы по целям и типам проверок, ingress по площадкам, классы ошибок, валидаторы; выбор запуска; списки адресов с выгрузкой в CSV | | `/analytics` | Аналитика одного завершённого запуска: показатели, причины `partial`, подсети, провалы по целям и типам проверок, ingress по площадкам, классы ошибок, валидаторы; выбор запуска; списки адресов с выгрузкой в CSV |
| `/validators`, `/sites`, `/targets`, `/check-types` | Управление валидаторами, внешними площадками, группами целей и типами проверок | | `/validators`, `/sites`, `/targets`, `/check-types` | Управление валидаторами, внешними площадками, группами целей и типами проверок |
| `/settings` | Панель «Автоматический цикл», пауза перед self-check, глубина истории, TCP-порты и ICMP для inbound-проверок | | `/settings` | Панель «Автоматический цикл», пауза перед self-check, потолок провалов self-check на адрес, глубина истории, TCP-порты и ICMP для inbound-проверок |
- Порядок блоков на `/overview` фиксирован: статистика → фильтр → таблицы; поллится только блок таблиц, поэтому набранный в фильтре текст не сбрасывается. - Порядок блоков на `/overview` фиксирован: статистика → фильтр → таблицы; поллится только блок таблиц, поэтому набранный в фильтре текст не сбрасывается.
- Ошибки control-api показываются баннером; при недоступном API страница остаётся рабочей. - Ошибки control-api показываются баннером; при недоступном API страница остаётся рабочей.
- Тёмная и светлая темы, переключатель в шапке. - Тёмная и светлая темы, переключатель в шапке.
@@ -208,6 +208,7 @@ scripts/run-local-e2e.sh # сквозной прог
| Дата | Веха | Документ | | Дата | Веха | Документ |
|---|---|---| |---|---|---|
| 2026-10-04 | Повтор после сбоя self-check — на другом валидаторе: валидатор, проваливший self-check, этому адресу больше не выдаётся; потолок провалов `self_check_max_attempts` (по умолчанию 5, миграция `0012`); поле `self_check_failed_on` | [план](docs/changes/2026-10-04_08-01_self-check-exclude-validator-plan.md) · [итог](docs/changes/2026-10-04_08-01_self-check-exclude-validator-summary.md) |
| 2026-10-03 | Раздел «Аналитика»: запуски проверки (миграция `0011`), показатели и списки адресов по запуску, подсети, CSV; сайдбар: связь с control-api и выход наверху, группы разделов; фильтры реестра по запуску и подсети | [план](docs/changes/2026-10-03_16-39_analytics-section-plan.md) · [итог](docs/changes/2026-10-03_18-41_analytics-section-summary.md) · [макет](docs/mockups/analytics-mockup.html) · [USAGE](docs/USAGE.md#аналитика-запусков) · [API](docs/API.md#аналитика-запусков) | | 2026-10-03 | Раздел «Аналитика»: запуски проверки (миграция `0011`), показатели и списки адресов по запуску, подсети, CSV; сайдбар: связь с control-api и выход наверху, группы разделов; фильтры реестра по запуску и подсети | [план](docs/changes/2026-10-03_16-39_analytics-section-plan.md) · [итог](docs/changes/2026-10-03_18-41_analytics-section-summary.md) · [макет](docs/mockups/analytics-mockup.html) · [USAGE](docs/USAGE.md#аналитика-запусков) · [API](docs/API.md#аналитика-запусков) |
| 2026-10-03 | Вердикт без опоздавших результатов: проверки фиксируются в момент вердикта, площадка зондирует адрес один раз за попытку, окно проверки считается от её начала, время записи по часам сервера (`checks.recorded_at`) | [план](docs/changes/2026-10-03_17-21_verdict-no-late-results-plan.md) · [итог](docs/changes/2026-10-03_17-21_verdict-no-late-results-summary.md) · [API](docs/API.md#результаты-после-вердикта) · [USAGE](docs/USAGE.md#просмотр-деталей-и-истории-по-конкретному-адресу) | | 2026-10-03 | Вердикт без опоздавших результатов: проверки фиксируются в момент вердикта, площадка зондирует адрес один раз за попытку, окно проверки считается от её начала, время записи по часам сервера (`checks.recorded_at`) | [план](docs/changes/2026-10-03_17-21_verdict-no-late-results-plan.md) · [итог](docs/changes/2026-10-03_17-21_verdict-no-late-results-summary.md) · [API](docs/API.md#результаты-после-вердикта) · [USAGE](docs/USAGE.md#просмотр-деталей-и-истории-по-конкретному-адресу) |
| 2026-10-03 | Реестр: последний результат по уровням Egress и Ingress, «успешно из всего» и разбивка по типам проверок (поля `egress`, `ingress`, `last_cycle_id` в `GET /admin/registry`) | [план](docs/changes/2026-10-03_16-24_registry-egress-ingress-levels-plan.md) · [итог](docs/changes/2026-10-03_16-24_registry-egress-ingress-levels-summary.md) · [USAGE](docs/USAGE.md#реестр-адресов-и-глубина-истории) · [API](docs/API.md#get-apiv1adminregistry) | | 2026-10-03 | Реестр: последний результат по уровням Egress и Ingress, «успешно из всего» и разбивка по типам проверок (поля `egress`, `ingress`, `last_cycle_id` в `GET /admin/registry`) | [план](docs/changes/2026-10-03_16-24_registry-egress-ingress-levels-plan.md) · [итог](docs/changes/2026-10-03_16-24_registry-egress-ingress-levels-summary.md) · [USAGE](docs/USAGE.md#реестр-адресов-и-глубина-истории) · [API](docs/API.md#get-apiv1adminregistry) |
+2 -2
View File
@@ -1,4 +1,4 @@
7a5227edddb89763f8db8ea56be5a1703626998498c3969b69272bfdb9d670b9 control-api 0ce9e3b5637b892420a61c8c01843a2348faeaddd11a601ed88e0b1ef25f04c8 control-api
9fb6608b84143f7c4f318f3cc92dcd9f95c7831d627b23d67cce5a5908ced704 validator-agent 9fb6608b84143f7c4f318f3cc92dcd9f95c7831d627b23d67cce5a5908ced704 validator-agent
3e9e14dbb361ee76aaad7c1da6864b3ea111e0ed151403f904b12485631bbf75 prober 3e9e14dbb361ee76aaad7c1da6864b3ea111e0ed151403f904b12485631bbf75 prober
1e5c0c3e2857aa2e856d89a6d72db7f402179b9d912177040c6e1fc9217b86d6 admin-dashboard b5dce1ecee2ebd03603cf8574ccea7129958923ca3db3b29fa4dfa87ac14f5ab admin-dashboard
Binary file not shown.
BIN
View File
Binary file not shown.
+6 -4
View File
@@ -12,7 +12,7 @@
| Группа | Таблицы | Можно чистить | | Группа | Таблицы | Можно чистить |
|---|---|---| |---|---|---|
| Данные прогона | `ip_queue` (очередь), `ip_registry` (реестр адресов), `checks` (реестр проверок), `ip_site_checks` (признаки площадок по адресам в работе), `check_runs` и `run_results` (запуски и итоги адресов в них), `events` (журнал событий) | да | | Данные прогона | `ip_queue` (очередь), `ip_registry` (реестр адресов), `checks` (реестр проверок), `ip_site_checks` (признаки площадок по адресам в работе), `check_runs` и `run_results` (запуски и итоги адресов в них), `ip_self_check_failures` (история сбоев self-check по адресам), `events` (журнал событий) | да |
| Настройки (не трогать) | `validators`, `sites`, `target_groups` (цели), `check_types`, `inbound_checks_settings`, `settings`, `auto_cycle`, `subnets` (список подсетей для аналитики) | **нет** | | Настройки (не трогать) | `validators`, `sites`, `target_groups` (цели), `check_types`, `inbound_checks_settings`, `settings`, `auto_cycle`, `subnets` (список подсетей для аналитики) | **нет** |
| Служебное | `sqlite_sequence` (нумерация записей), `PRAGMA user_version` (версия схемы) | нумерацию можно сбросить, версию не менять | | Служебное | `sqlite_sequence` (нумерация записей), `PRAGMA user_version` (версия схемы) | нумерацию можно сбросить, версию не менять |
@@ -86,7 +86,7 @@ UNION ALL SELECT 'sites', COUNT(*) FROM sites;"
### 3.1. Полный сброс данных прогона (перед новым полным прогоном) ### 3.1. Полный сброс данных прогона (перед новым полным прогоном)
Очищает очередь, реестр адресов, реестр проверок, журнал событий. Валидаторы, площадки, цели, типы проверок и все настройки остаются. Очищает очередь, реестр адресов, реестр проверок, историю сбоев self-check, журнал событий. Валидаторы, площадки, цели, типы проверок и все настройки остаются.
```sql ```sql
PRAGMA foreign_keys = ON; PRAGMA foreign_keys = ON;
@@ -99,11 +99,12 @@ DELETE FROM ip_site_checks;
DELETE FROM run_results; DELETE FROM run_results;
DELETE FROM check_runs; DELETE FROM check_runs;
DELETE FROM checks; DELETE FROM checks;
DELETE FROM ip_self_check_failures;
DELETE FROM events; DELETE FROM events;
DELETE FROM ip_queue; DELETE FROM ip_queue;
DELETE FROM ip_registry; DELETE FROM ip_registry;
-- нумерация снова с 1 (необязательно) -- нумерация снова с 1 (необязательно)
DELETE FROM sqlite_sequence WHERE name IN ('ip_registry', 'checks', 'ip_queue', 'events', 'check_runs'); DELETE FROM sqlite_sequence WHERE name IN ('ip_registry', 'checks', 'ip_queue', 'events', 'check_runs', 'ip_self_check_failures');
COMMIT; COMMIT;
``` ```
@@ -157,7 +158,7 @@ WHERE cycle_id <= (SELECT MAX(c2.cycle_id) FROM checks c2 WHERE c2.registry_id =
### 3.4. Удалить конкретные адреса целиком ### 3.4. Удалить конкретные адреса целиком
Удаляет адрес из очереди и реестра вместе со всей его историей (проверки и события). Список адресов подставьте в первую команду Удаляет адрес из очереди и реестра вместе со всей его историей (проверки, сбои self-check и события). Список адресов подставьте в первую команду
`CREATE TEMP TABLE doomed_reg`. Адрес, который сейчас проверяется, удалять этим способом нельзя: используйте API (раздел 5). `CREATE TEMP TABLE doomed_reg`. Адрес, который сейчас проверяется, удалять этим способом нельзя: используйте API (раздел 5).
```sql ```sql
@@ -171,6 +172,7 @@ UPDATE validators SET current_ip_id = NULL WHERE current_ip_id IN (SELECT id FRO
DELETE FROM ip_site_checks WHERE ip_id IN (SELECT id FROM doomed_ip); DELETE FROM ip_site_checks WHERE ip_id IN (SELECT id FROM doomed_ip);
DELETE FROM run_results WHERE registry_id IN (SELECT id FROM doomed_reg); DELETE FROM run_results WHERE registry_id IN (SELECT id FROM doomed_reg);
DELETE FROM checks WHERE registry_id IN (SELECT id FROM doomed_reg); DELETE FROM checks WHERE registry_id IN (SELECT id FROM doomed_reg);
DELETE FROM ip_self_check_failures WHERE registry_id IN (SELECT id FROM doomed_reg);
DELETE FROM events WHERE registry_id IN (SELECT id FROM doomed_reg) OR ip_id IN (SELECT id FROM doomed_ip); DELETE FROM events WHERE registry_id IN (SELECT id FROM doomed_reg) OR ip_id IN (SELECT id FROM doomed_ip);
DELETE FROM ip_queue WHERE id IN (SELECT id FROM doomed_ip); DELETE FROM ip_queue WHERE id IN (SELECT id FROM doomed_ip);
DELETE FROM ip_registry WHERE id IN (SELECT id FROM doomed_reg); DELETE FROM ip_registry WHERE id IN (SELECT id FROM doomed_reg);
+28 -7
View File
@@ -198,8 +198,13 @@ self-check способом `control_api` (`self_check.methods` в
``` ```
Ответ: `{"ok": true}`. При `success: false` control-api сам решает — Ответ: `{"ok": true}`. При `success: false` control-api сам решает —
повторить попытку назначения FIP или пометить IP как `failed` (после вернуть адрес в очередь или пометить IP как `failed`. Сбой записывается в
исчерпания `orchestrator.max_self_check_retries`). историю адреса; повтор **не отдаётся этому валидатору** (он остаётся в
работе и берёт остальные адреса). Итог `fail` ставится, когда число
проваленных self-check у адреса достигло потолка `self_check_max_attempts`
(по умолчанию 5, см. [«Настройки оркестратора»](#настройки-оркестратора-apiv1adminconfigorchestrator));
`orchestrator.max_self_check_retries` не используется. Сбой привязки FIP и
истечение лизинга идут по `max_retries`, как раньше.
### `POST /api/v1/agents/{id}/events` ### `POST /api/v1/agents/{id}/events`
@@ -415,10 +420,14 @@ IP на данном проходе". До этого момента control-api
{ {
"ip": { "ID": 42, "IPAddress": "203.0.113.10", "State": "done", "OverallResult": "pass", "...": "..." }, "ip": { "ID": 42, "IPAddress": "203.0.113.10", "State": "done", "OverallResult": "pass", "...": "..." },
"checks": [ {"Source": "egress", "CheckType": "https", "Target": "https://github.com", "Success": true, "...": "..."} ], "checks": [ {"Source": "egress", "CheckType": "https", "Target": "https://github.com", "Success": true, "...": "..."} ],
"events": [ {"EventType": "fip_associated", "OccurredAt": "...", "...": "..."} ] "events": [ {"EventType": "fip_associated", "OccurredAt": "...", "...": "..."} ],
"self_check_failed_on": ["vkiplab-v17"]
} }
``` ```
`self_check_failed_on` — валидаторы, у которых self-check на этом адресе не
прошёл в текущем запуске (по алфавиту; `[]`, если сбоев не было).
> Обратите внимание: вложенные объекты `ip`/`checks`/`events` сериализуются > Обратите внимание: вложенные объекты `ip`/`checks`/`events` сериализуются
> без переопределения имён полей (используются имена Go-структур, например > без переопределения имён полей (используются имена Go-структур, например
> `IPAddress`, `State`, `Success`) — в отличие от методов для > `IPAddress`, `State`, `Success`) — в отличие от методов для
@@ -722,10 +731,14 @@ curl -s -X POST http://<control-api>:8080/api/v1/admin/auto-cycle/stop
"checks": [ "checks": [
{"CycleID": 4, "Source": "egress", "CheckType": "https", "Success": true, "...": "..."}, {"CycleID": 4, "Source": "egress", "CheckType": "https", "Success": true, "...": "..."},
{"CycleID": 3, "Source": "egress", "CheckType": "https", "Success": false, "...": "..."} {"CycleID": 3, "Source": "egress", "CheckType": "https", "Success": false, "...": "..."}
] ],
"self_check_failed_on": ["vkiplab-v17"]
} }
``` ```
`self_check_failed_on` — валидаторы, у которых self-check на этом адресе не
прошёл, по всем запускам (по алфавиту; `[]`, если сбоев не было).
`404`, если адрес никогда не ставился на проверку. `404`, если адрес никогда не ставился на проверку.
**Глубина хранения.** Сколько последних циклов на адрес хранится в **Глубина хранения.** Сколько последних циклов на адрес хранится в
@@ -797,7 +810,7 @@ curl -s -X POST "$BASE/api/v1/admin/ips/clear"
### Настройки оркестратора: `/api/v1/admin/config/orchestrator` ### Настройки оркестратора: `/api/v1/admin/config/orchestrator`
Два параметра: Три параметра:
- `fip_settle_seconds` — пауза между привязкой Floating IP к валидатору и - `fip_settle_seconds` — пауза между привязкой Floating IP к валидатору и
моментом, когда self-check по этому адресу становится доступен агенту моментом, когда self-check по этому адресу становится доступен агенту
@@ -811,11 +824,19 @@ curl -s -X POST "$BASE/api/v1/admin/ips/clear"
на адрес в реестре (`GET /api/v1/admin/registry/{ip}`, см. на адрес в реестре (`GET /api/v1/admin/registry/{ip}`, см.
[«Реестр адресов»](#реестр-адресов-и-история-проверок)). `0` — без [«Реестр адресов»](#реестр-адресов-и-история-проверок)). `0` — без
ограничения (поведение по умолчанию). ограничения (поведение по умолчанию).
- `self_check_max_attempts` — потолок провалов self-check на один адрес
(1…50, по умолчанию 5). Когда у адреса провалено столько self-check,
он получает итог `fail`. Не зависит от числа валидаторов. Валидатор,
проваливший self-check на адресе, этому адресу больше не выдаётся (на
остальные адреса это не влияет); если все рабочие валидаторы уже
провалили адрес, исключения сбрасываются и повторы продолжаются до
потолка. Значение действует на следующих повторах без перезапуска.
В `PUT` поле необязательно: если не передано, не меняется.
| Метод | Путь | Тело | Успех | Ошибки | | Метод | Путь | Тело | Успех | Ошибки |
|---|---|---|---|---| |---|---|---|---|---|
| GET | `/api/v1/admin/config/orchestrator` | — | `{"fip_settle_seconds":N,"history_retention_cycles":M}` | | | GET | `/api/v1/admin/config/orchestrator` | — | `{"fip_settle_seconds":N,"history_retention_cycles":M,"self_check_max_attempts":K}` | |
| PUT | `/api/v1/admin/config/orchestrator` | `{"fip_settle_seconds":N,"history_retention_cycles":M}` | `200` | `400`, если `N < 0` или `M < 0`, или если `fip_settle_seconds + self_check_timeout_seconds >= lease_ttl_seconds` (пауза не должна съедать весь лизинг адреса — иначе self-check не успеет пройти до истечения `lease_ttl_seconds`, и адрес будет вечно возвращаться в очередь) | | PUT | `/api/v1/admin/config/orchestrator` | `{"fip_settle_seconds":N,"history_retention_cycles":M,"self_check_max_attempts":K}` | `200` | `400`, если `N < 0`, `M < 0` или `K` вне 1…50, или если `fip_settle_seconds + self_check_timeout_seconds >= lease_ttl_seconds` (пауза не должна съедать весь лизинг адреса — иначе self-check не успеет пройти до истечения `lease_ttl_seconds`, и адрес будет вечно возвращаться в очередь) |
Как и остальные разделы этой группы, YAML-поле `orchestrator. Как и остальные разделы этой группы, YAML-поле `orchestrator.
fip_settle_seconds` в `control-api.yaml` — только одноразовый bootstrap fip_settle_seconds` в `control-api.yaml` — только одноразовый bootstrap
+1 -1
View File
@@ -63,7 +63,7 @@ admin-dashboard -config /etc/cloud-ip-validator/admin-dashboard.yaml
| `/sites` | Площадки — число слотов не ограничено, форма сверху добавляет новый слот, назначить/сменить/освободить `site_id` в каждой строке; колонка «Статус» показывает бейдж подключения пробера (`unregistered`/`idle`/`unreachable`, по аналогии с `/validators`), см. [USAGE.md](USAGE.md#состояния-площадки). | | `/sites` | Площадки — число слотов не ограничено, форма сверху добавляет новый слот, назначить/сменить/освободить `site_id` в каждой строке; колонка «Статус» показывает бейдж подключения пробера (`unregistered`/`idle`/`unreachable`, по аналогии с `/validators`), см. [USAGE.md](USAGE.md#состояния-площадки). |
| `/targets` | Группы целей для egress-проверок — создание/редактирование/удаление. | | `/targets` | Группы целей для egress-проверок — создание/редактирование/удаление. |
| `/check-types` | Типы проверок (`https`/`icmp`/`ssh`/...), включение/выключение, привязка к группам целей. | | `/check-types` | Типы проверок (`https`/`icmp`/`ssh`/...), включение/выключение, привязка к группам целей. |
| `/settings` | Четыре блока. Первый — панель **«Автоматический цикл»**: статус и фаза, время последнего/следующего запуска, результат последнего цикла, поля «Интервал между циклами (мин)» и «Максимальная длительность проверки (мин, 0 = без лимита)» с кнопкой «Сохранить» и кнопка «Включить»/«Выключить» (показывается та, что сейчас применима). Значения вводятся в минутах (допустимы дробные), в control-api уходят секундами; минимум интервала — 1 минута (`60` с), нарушение приходит предупреждением в баннере. Подробности — [USAGE.md](USAGE.md#автоматический-цикл-проверок), API — [API.md](API.md#автоматический-цикл-проверок). Далее три формы: `fip_settle_seconds` — пауза (в секундах) между привязкой Floating IP и началом self-check («прогрев» дата-плейна OpenStack, см. [USAGE.md](USAGE.md#пауза-перед-self-check-fip_settle_seconds)); `history_retention_cycles` — сколько последних циклов проверки хранить на адрес в реестре (0 — без ограничения); и типы проверок пробера — TCP-порты (через запятую) + чекбокс ICMP, общие для всех площадок (см. [USAGE.md](USAGE.md#управление-типами-проверок-пробера)). | | `/settings` | Четыре блока. Первый — панель **«Автоматический цикл»**: статус и фаза, время последнего/следующего запуска, результат последнего цикла, поля «Интервал между циклами (мин)» и «Максимальная длительность проверки (мин, 0 = без лимита)» с кнопкой «Сохранить» и кнопка «Включить»/«Выключить» (показывается та, что сейчас применима). Значения вводятся в минутах (допустимы дробные), в control-api уходят секундами; минимум интервала — 1 минута (`60` с), нарушение приходит предупреждением в баннере. Подробности — [USAGE.md](USAGE.md#автоматический-цикл-проверок), API — [API.md](API.md#автоматический-цикл-проверок). Далее формы: `fip_settle_seconds` — пауза (в секундах) между привязкой Floating IP и началом self-check («прогрев» дата-плейна OpenStack, см. [USAGE.md](USAGE.md#пауза-перед-self-check-fip_settle_seconds)); `self_check_max_attempts` — «Потолок провалов self-check на адрес» (1…50, по умолчанию 5; повторы идут на других валидаторах, см. [USAGE.md](USAGE.md#повтор-после-сбоя-self-check)); `history_retention_cycles` — сколько последних циклов проверки хранить на адрес в реестре (0 — без ограничения); и типы проверок пробера — TCP-порты (через запятую) + чекбокс ICMP, общие для всех площадок (см. [USAGE.md](USAGE.md#управление-типами-проверок-пробера)). |
### «В работе», «в очереди» и «последняя завершённая» проверка ### «В работе», «в очереди» и «последняя завершённая» проверка
+36 -2
View File
@@ -33,6 +33,7 @@
- [Принудительная остановка проверки](#принудительная-остановка-проверки) - [Принудительная остановка проверки](#принудительная-остановка-проверки)
- [Удаление адресов из очереди](#удаление-адресов-из-очереди) - [Удаление адресов из очереди](#удаление-адресов-из-очереди)
- [Пауза перед self-check (fip_settle_seconds)](#пауза-перед-self-check-fip_settle_seconds) - [Пауза перед self-check (fip_settle_seconds)](#пауза-перед-self-check-fip_settle_seconds)
- [Повтор после сбоя self-check](#повтор-после-сбоя-self-check)
- [Частые проблемы и что с ними делать](#частые-проблемы-и-что-с-ними-делать) - [Частые проблемы и что с ними делать](#частые-проблемы-и-что-с-ними-делать)
## Как устроена работа с системой ## Как устроена работа с системой
@@ -306,8 +307,9 @@ curl -s http://<control-api>:8080/api/v1/admin/ips \
оператора: годится ли адрес для данного случая использования. оператора: годится ли адрес для данного случая использования.
- **`fail`** — либо ни одна проверка не прошла, либо адрес вообще не - **`fail`** — либо ни одна проверка не прошла, либо адрес вообще не
дошёл до стадии проверок (например, self-check не подтвердился — дошёл до стадии проверок (например, self-check не подтвердился —
трафик валидатора не пошёл через назначенный FIP — и попытки трафик валидатора не пошёл через назначенный FIP — и число провалов
исчерпались). Смотрите `events` по этому адресу (см. ниже), чтобы достигло потолка `self_check_max_attempts`, см.
[«Повтор после сбоя self-check»](#повтор-после-сбоя-self-check)). Смотрите `events` по этому адресу (см. ниже), чтобы
понять, на каком шаге и почему. понять, на каком шаге и почему.
- **`cancelled`** — проверку остановил оператор через `POST - **`cancelled`** — проверку остановил оператор через `POST
/api/v1/admin/ips/{ip}/cancel` (`State` при этом — `failed`), а не /api/v1/admin/ips/{ip}/cancel` (`State` при этом — `failed`), а не
@@ -774,6 +776,37 @@ curl -s -X POST http://<control-api>:8080/api/v1/admin/ips/clear
всё» — каждая с подтверждением, явно предупреждающим о необратимости всё» — каждая с подтверждением, явно предупреждающим о необратимости
(см. [DASHBOARD.md](DASHBOARD.md)). (см. [DASHBOARD.md](DASHBOARD.md)).
## Повтор после сбоя self-check
Если self-check адреса не прошёл на валидаторе N, адрес возвращается в
очередь, но **этому же валидатору больше не выдаётся**: повтор достаётся
другому. Правило действует только для этого адреса. Валидатор N не
блокируется и не помечается неисправным, он продолжает брать остальные
адреса. Так один временно неисправный валидатор (например, у него не
работает трансляция плавающего IP) не расходует все попытки адреса.
Итог `fail` ставится, когда число проваленных self-check у адреса достигло
потолка **«Потолок провалов self-check на адрес»** (`self_check_max_attempts`,
по умолчанию 5; `/settings` или `PUT /api/v1/admin/config/orchestrator`,
допустимо 1…50). Потолок не зависит от числа валидаторов и действует на
следующих повторах без перезапуска. Если все рабочие валидаторы
(`idle`, `assigned`, `checking`) уже провалили адрес, а потолок не достигнут
(валидаторов меньше потолка, в том числе один), исключения сбрасываются и
повторы продолжаются до потолка.
Исключение создаёт только проваленный self-check. Ошибка привязки
плавающего IP, потеря валидатора и истечение лизинга исключений не создают и
идут по `orchestrator.max_retries`, как раньше. Поле
`orchestrator.max_self_check_retries` в YAML больше не используется.
Ручная перепроверка или повторная постановка адреса начинает серию заново
(счётчик сбоев обнуляется). История сбоев хранится постоянно (таблица
`ip_self_check_failures`) и отдаётся в поле `self_check_failed_on` ответов
`GET /admin/ips/{ip}` и `GET /admin/registry/{ip}`; на страницах `/ips/{ip}` и
`/registry/{ip}` показана строка «Self-check не прошёл на: …». В журнале
событий при повторе появляется `validator_excluded`. Показ на странице
аналитики — отдельный следующий шаг.
## Пауза перед self-check (fip_settle_seconds) ## Пауза перед self-check (fip_settle_seconds)
Как только Floating IP привязывается к валидатору, control-api по Как только Floating IP привязывается к валидатору, control-api по
@@ -833,6 +866,7 @@ https://api.ipify.org`) и логи `journalctl -u validator-agent` на пре
**Адрес постоянно проваливает self-check (не зависает, а именно **Адрес постоянно проваливает self-check (не зависает, а именно
возвращается в очередь снова и снова).** возвращается в очередь снова и снова).**
См. также [«Повтор после сбоя self-check»](#повтор-после-сбоя-self-check): повторы идут на других валидаторах, а список «Self-check не прошёл на: …» показан на странице адреса.
Смотрите `events` по адресу (`GET /api/v1/admin/ips/{ip}`) — в детали Смотрите `events` по адресу (`GET /api/v1/admin/ips/{ip}`) — в детали
события `self_check_result` будет указан обнаруженный исходящий адрес. события `self_check_result` будет указан обнаруженный исходящий адрес.
Если он не совпадает с ожидаемым — вероятно, на ВМ-валидаторе есть другой Если он не совпадает с ожидаемым — вероятно, на ВМ-валидаторе есть другой
@@ -0,0 +1,100 @@
# План: повтор после сбоя self-check — на другом валидаторе
Статус: реализовано, см. [итог](2026-10-04_08-01_self-check-exclude-validator-summary.md).
## 1. Проблема
В запуске 2 два адреса получили `fail`: `89.208.220.90` и `89.208.221.74`. Оба раза self-check провалился четыре раза подряд на одном валидаторе `vkiplab-v17` (он выходил в сеть с общего адреса облака `109.120.182.224`, а не с плавающего IP), с интервалом около 70 секунд.
Причина в логике повтора: после сбоя адрес возвращается в начало очереди, а освободившийся `v17` берёт следующий адрес с наименьшим `sequence`, то есть тот же. Один временно неисправный валидатор расходует все попытки адреса. Сам адрес исправен: `89.208.216.141`, попавший на `v17` дважды, после перехода на `v13` прошёл проверки.
## 2. Требования и принятые решения
1. Адрес, не прошедший self-check на валидаторе N, при повторе **не отдаётся валидатору N**. Правило действует только для этого адреса.
2. Валидатор N остаётся в системе как был: не блокируется, не помечается неисправным, продолжает брать и проверять все остальные адреса.
3. **Итог `fail`** ставится, когда число проваленных self-check у адреса достигло **потолка**. Потолок не зависит от числа валидаторов. Это **настройка** `self_check_max_attempts`: она меняется в разделе настроек `/settings` и через API; **значение по умолчанию для текущего окружения — 5**. Исключение валидаторов (п.1) гарантирует, что первые попытки идут на разных валидаторах: адрес получает `fail` после 5 провалов на 5 разных валидаторах (а не на одном).
4. Исключение создаёт **только** проваленный self-check; на ошибки привязки плавающего IP, потерю валидатора и истечение аренды оно не распространяется.
5. Сведения «Self-check не прошёл на: …» сохраняются и отдаются в API для последующего показа на странице аналитики.
## 3. Решение
### 3.1. Данные (миграция `0012_self_check_failures.sql`)
- Таблица `ip_self_check_failures(registry_id, run_id, cycle_id, attempt_number, validator_id, failed_at, detail)`. Постоянная история сбоев self-check по адресу; не очищается при перепроверках, живёт вместе с реестром (не удаляется вместе со строкой очереди) и служит источником для аналитики.
- Колонки `ip_queue.sc_failures` (сколько self-check провалено в текущей серии попыток) и `ip_queue.sc_round_start_attempt` (номер попытки, с которой действуют исключения; «раунд»).
- Настройка `self_check_max_attempts` в таблице `settings`: миграция записывает **5** (допустимо 1…50), дальше значение меняется в `/settings` или через API. Поле `orchestrator.max_self_check_retries` в YAML для self-check больше не используется (остаётся в файле для совместимости, в документации помечается устаревшим).
Исключение для валидатора N на адресе действует, если в `ip_self_check_failures` есть сбой этого валидатора на этом адресе с `attempt_number >= sc_round_start_attempt`.
### 3.2. Обработка сбоя (`Orchestrator.SelfCheckResult`, ветка «не прошёл»)
После проверки, что адрес принадлежит этому валидатору и ждёт self-check, и отвязки плавающего IP (как сейчас):
1. Записать сбой в `ip_self_check_failures`, увеличить `sc_failures`.
2. Решение:
- **`sc_failures >= self_check_max_attempts`** → `fail` (причина «self-check не прошёл N раз», в событии перечислены валидаторы);
- иначе → вернуть адрес в очередь. Если среди рабочих валидаторов (`idle`, `assigned`, `checking`; `unreachable` и `unregistered` не считаются) не осталось ни одного без сбоя в текущем раунде, начинается новый раунд (`sc_round_start_attempt` = следующая попытка): исключения теряют силу, повтор идёт по обычным правилам. Так при числе валидаторов меньше потолка (в том числе при единственном) повторы продолжаются до потолка, а адрес не застревает в очереди.
Лимит `max_retries` и счётчик `retry_count` для self-check больше не используются: возврат в очередь после сбоя self-check не увеличивает `retry_count` и не может привести к `fail` по нему. Для остальных причин возврата (ошибка привязки, истечение аренды) всё остаётся как сейчас.
### 3.3. Выдача адресов (`DB.ClaimNextQueued`)
Валидатор получает ближайший по `sequence` адрес, **кроме** тех, где он исключён (по 3.1). Исключённый адрес остаётся в очереди для остальных валидаторов и очередь не блокирует. Запрос по-прежнему одной транзакцией.
### 3.4. Жизненный цикл
- Ручная перепроверка или повторная постановка адреса (`SubmitIPs`) начинает новую серию: `sc_failures = 0`, `sc_round_start_attempt` = текущая попытка.
- История в `ip_self_check_failures` остаётся.
- Удаление адреса и очистка очереди историю не трогают; `docs/ADMIN_CLEANUP.md` дополняется новой таблицей.
### 3.5. Данные для аналитики и видимость
- Событие `self_check_result` остаётся как сейчас; добавляется событие `validator_excluded` (`{"validator_id": …, "failures": N}`).
- `GET /admin/ips/{ip}` и `GET /admin/registry/{ip}` получают поле `self_check_failed_on` (список валидаторов, у которых self-check на этом адресе не прошёл; для реестра — по всем запускам, для запуска — фильтр по `run_id`).
- Строка «Self-check не прошёл на: v17, …» на странице адреса (`/ips/{ip}`, `/registry/{ip}`).
- **Отображение на странице аналитики не входит в эту доработку:** данные и API будут готовы, показ (колонка в списках адресов, блок «Self-check по валидаторам») делается следующим шагом отдельным планом.
### 3.6. Настройка в интерфейсе (`/settings`)
Изменения отражаются в разделе настроек явно:
- В форме с `fip_settle_seconds` и `history_retention_cycles` добавляется поле **«Потолок провалов self-check на адрес»** (`self_check_max_attempts`), целое число, по умолчанию 5, допустимо 1…50.
- Под полем пояснение: «Сколько раз self-check может не пройти у одного адреса, прежде чем адрес получит итог `fail`. Повторы идут на других валидаторах: валидатор, на котором адрес не прошёл self-check, этому адресу больше не выдаётся (на остальные адреса это не влияет)».
- Значение сохраняется той же кнопкой «Сохранить», действует на следующих повторах без перезапуска; неверное значение (не число, вне 1…50) показывается предупреждением в баннере, форма остаётся прежней.
- В API: `GET/PUT /admin/config/orchestrator` получает поле `self_check_max_attempts` (ошибка 400 при неверном значении).
- Описание поля добавляется в `docs/DASHBOARD.md` (строка `/settings`) и `docs/USAGE.md`.
## 4. Файлы
- БД: `internal/db/migrations/0012_self_check_failures.sql`, `db.go`, `queries_ipqueue.go` (`ClaimNextQueued`, запись сбоя, новый возврат в очередь без `retry_count`, принудительный `fail`, сброс серии в `SubmitIPsAs`), `queries_settings.go` (настройка), запрос истории для API.
- Оркестратор: `orchestrator.go` (`SelfCheckResult`), `config.go` (начальное значение настройки).
- API и дашборд: `handlers_admin.go`, `handlers_config.go`, DTO, `handlers_settings.go`, `settings.html`, `ip_detail.html`, `registry_detail.html`.
- Документы: `USAGE.md` («Как читать итоговый результат», повторы self-check, настройка), `API.md`, `ADMIN_CLEANUP.md`, `README.md`, итог `docs/changes/…-summary.md`.
## 5. Тесты
- **БД:** валидатор, на котором адрес исключён, пропускает его и берёт следующий; другой валидатор берёт исключённый адрес; исключения действуют только для этого адреса; новый раунд снимает исключения; перепостановка сбрасывает серию, но не историю; настройка (значения по умолчанию, границы).
- **Оркестратор (сценарий инцидента на трёх валидаторах):** v1 всегда проваливает self-check — адрес не возвращается на v1, проходит на v2, а v1 в это время берёт и проходит другие адреса; потолок: адрес, провалившийся `self_check_max_attempts` раз, получает `fail`, раньше — нет; при потолке меньше числа валидаторов `fail` наступает на потолке без перебора всех; при потолке больше числа валидаторов начинается новый раунд; единственный валидатор повторяет на себе до потолка; `unreachable` не считается валидатором; сбой привязки и истечение аренды исключений и счётчика self-check не создают и по-прежнему идут по `max_retries`; изменение настройки действует на следующем повторе.
- **API и дашборд:** поле `self_check_failed_on`; поле «Потолок провалов self-check на адрес» отображается в `/settings` со значением 5, сохраняется, неверные значения показывают предупреждение; поле в `GET/PUT /admin/config/orchestrator`; строка «Self-check не прошёл на: …» на странице адреса.
- Миграция `0012` на копии боевой БД (настройка равна 5), повторный запуск безопасен. Полный `go build ./... && go vet ./... && go test ./...`.
## 6. Выкладка и проверка
Пересборка и перезапуск `control-api` и `admin-dashboard` (меняется форма настроек и страница адреса); агенты и prober без изменений. Порядок: проверка пустой очереди, копия БД, тег отката образа, миграция `0012`. Сбой на стенде воспроизвести нельзя, поэтому проверка там — штатный небольшой прогон без зависших `queued`; логику повторов покрывает тест на трёх валидаторах.
## 7. Риски и последствия принятого решения
- **Время до `fail` у неисправного адреса** ограничено потолком: 5 попыток по ~70 секунд, около 6 минут вместо ~5 минут сейчас (4 попытки). Очередь и другие валидаторы это не блокирует.
- **Потолок меньше числа валидаторов.** Адрес, которому не повезло с 5 неисправными валидаторами подряд, получит `fail`, хотя на остальных 15 прошёл бы. Это сознательный компромисс в пользу ограниченного времени; потолок меняется в настройках.
- Адрес, исключённый на части валидаторов, может чуть дольше ждать подходящего валидатора.
- Если сбой общий (проблема облака для всех валидаторов), адрес в итоге получит `fail`; это верное поведение.
- Не решается: причина сбоя `v17` в облаке (трансляция плавающего IP) — её нужно смотреть в OpenStack.
## 8. Решения по согласованию (все приняты)
1. Потолок попыток задаётся независимо от числа валидаторов; значение меняется в меню настроек (`/settings`) и через API.
2. Значение по умолчанию для текущего окружения — 5.
3. Исключение действует только для проваленного self-check; на ошибки привязки не распространяется.
4. «Self-check не прошёл на: …» сохраняется и отдаётся в API; показ на странице аналитики — следующим шагом.
5. Изменения отражаются в разделе настроек UI (п.3.6).
Открытых вопросов нет. Жду команды начать реализацию.
@@ -0,0 +1,32 @@
# Итог: повтор после сбоя self-check — на другом валидаторе
План: [2026-10-04_08-01_self-check-exclude-validator-plan.md](2026-10-04_08-01_self-check-exclude-validator-plan.md).
Статус: код написан и проверен (gofmt, build, vet, test, миграция на копии боевой БД); стенд **не пересобирался**.
## Что изменено
- **Миграция `0012_self_check_failures.sql`** (версия схемы 12): таблица `ip_self_check_failures` (история сбоев по адресу, по `registry_id`, без внешних ключей); колонки `ip_queue.sc_failures` и `ip_queue.sc_round_start_cycle`; настройка `settings.self_check_max_attempts` (`DEFAULT 5`, допустимо 1…50).
- **`db.ClaimNextQueued`** пропускает адреса, на которых этот валидатор провалил self-check в текущем раунде; адрес остаётся в очереди для других валидаторов.
- **`db.FailSelfCheck`** (одна транзакция): записывает сбой, увеличивает `sc_failures`; при `sc_failures >= потолка` — `fail`; иначе возвращает адрес в очередь без изменения `retry_count`; если рабочих валидаторов без сбоя в раунде не осталось — новый раунд. Позднее сообщение (адрес не у этого валидатора) — `ErrInvalidState`.
- **Оркестратор** (`SelfCheckResult` → `failSelfCheck`): потолок читается из настроек на каждом сбое. События: `retry_or_fail` (при `fail` — со списком валидаторов) и `validator_excluded`. `max_self_check_retries` не используется (поле в YAML осталось).
- **`SubmitIPsAs`/`SeedQueue`**: перепостановка начинает новую серию (`sc_failures = 0`, новый раунд); история остаётся.
- **API**: `GET/PUT /admin/config/orchestrator` — поле `self_check_max_attempts` (400 вне 1…50; в PUT необязательно); `self_check_failed_on` в `GET /admin/ips/{ip}` (фильтр по запуску) и `GET /admin/registry/{ip}` (все запуски).
- **Дашборд**: поле «Потолок провалов self-check на адрес» на `/settings`; строка «Self-check не прошёл на: …» на `/ips/{ip}` и `/registry/{ip}`.
- **Документы**: `USAGE.md` (новый раздел «Повтор после сбоя self-check»), `API.md`, `DASHBOARD.md`, `ADMIN_CLEANUP.md` (новая таблица в сценариях 3.1 и 3.4), `README.md`.
## Отступления от плана
1. Раунд хранится как `sc_round_start_cycle` (номер `cycle_id`), а не `sc_round_start_attempt`: `attempt_number` сбрасывается при удалении и повторном создании строки очереди, а история остаётся — новая строка получила бы чужие исключения. `cycle_id` по адресу не повторяется. Поведение для оператора то же.
2. Значение 5 задано `DEFAULT 5` колонки (покрывает и существующую строку `settings`, и чистую установку), `config.go` не менялся.
3. В `PUT` потолок необязателен — старые клиенты без поля не получают 400.
4. `validator_excluded` не пишется, если начался новый раунд (исключение сразу теряет силу).
## Проверки
- `gofmt -l` — пусто; `go build ./...`, `go vet ./...` — без ошибок; `go test ./...` — все пакеты `ok`. Две проверки версии схемы в старых тестах миграций (`TestMigration0010…`, `TestMigration0011…`) обновлены с 11 на 12.
- Новые тесты: БД (`queries_selfcheck_test.go`), оркестратор (сценарий на трёх валидаторах, потолок, один и два валидатора, смена настройки), API, дашборд.
- Миграция `0012` на копии боевой БД: версия 11 → 12, `self_check_max_attempts = 5`, повторное открытие без ошибок, таблица сбоев пуста.
## Выкладка
Не выполнена. Нужны пересборка и перезапуск `control-api` и `admin-dashboard` по процедуре (проверка пустой очереди, копия БД, тег отката `pre-self-check-exclude`, миграция `0012`); на стенде в очереди сейчас 6489 адресов `done`, `queued` нет.
+2 -2
View File
@@ -356,10 +356,10 @@ func (c *client) GetOrchestratorSettings(ctx context.Context) (orchestratorSetti
return out, err return out, err
} }
func (c *client) PutOrchestratorSettings(ctx context.Context, fipSettleSeconds, historyRetentionCycles int) (orchestratorSettingsDTO, error) { func (c *client) PutOrchestratorSettings(ctx context.Context, fipSettleSeconds, historyRetentionCycles, selfCheckMaxAttempts int) (orchestratorSettingsDTO, error) {
var out orchestratorSettingsDTO var out orchestratorSettingsDTO
err := c.do(ctx, http.MethodPut, "/api/v1/admin/config/orchestrator", err := c.do(ctx, http.MethodPut, "/api/v1/admin/config/orchestrator",
orchestratorSettingsDTO{FIPSettleSeconds: fipSettleSeconds, HistoryRetentionCycles: historyRetentionCycles}, &out) orchestratorSettingsDTO{FIPSettleSeconds: fipSettleSeconds, HistoryRetentionCycles: historyRetentionCycles, SelfCheckMaxAttempts: selfCheckMaxAttempts}, &out)
return out, err return out, err
} }
+15 -2
View File
@@ -37,6 +37,10 @@ type fakeControlAPI struct {
inboundICMP bool inboundICMP bool
historyRetentionCycles int historyRetentionCycles int
selfCheckMaxAttempts int
// selfCheckFailedOn is served as self_check_failed_on of the address
// detail and registry history endpoints.
selfCheckFailedOn []string
// Analytics: the run selector, the report JSON per run, the lists per // Analytics: the run selector, the report JSON per run, the lists per
// "run/kind[/class]" and the subnet list of /settings. // "run/kind[/class]" and the subnet list of /settings.
@@ -94,6 +98,8 @@ func newFakeControlAPI(t *testing.T) (*fakeControlAPI, string) {
registry: map[string]registryItem{}, registry: map[string]registryItem{},
registryChecks: map[string][]check{}, registryChecks: map[string][]check{},
autoCycle: autoCycleDTO{IntervalSeconds: 3600, Phase: "idle"}, autoCycle: autoCycleDTO{IntervalSeconds: 3600, Phase: "idle"},
selfCheckMaxAttempts: 5,
} }
ts := httptest.NewServer(f.handler()) ts := httptest.NewServer(f.handler())
t.Cleanup(ts.Close) t.Cleanup(ts.Close)
@@ -191,7 +197,7 @@ func (f *fakeControlAPI) handler() http.Handler {
addr := r.PathValue("ip") addr := r.PathValue("ip")
for _, ip := range f.ips { for _, ip := range f.ips {
if ip.IPAddress == addr { if ip.IPAddress == addr {
writeJSON(w, http.StatusOK, ipDetailResponse{IP: ip, Checks: []check{}, Events: []event{}}) writeJSON(w, http.StatusOK, ipDetailResponse{IP: ip, Checks: []check{}, Events: []event{}, SelfCheckFailedOn: f.selfCheckFailedOn})
return return
} }
} }
@@ -310,6 +316,7 @@ func (f *fakeControlAPI) handler() http.Handler {
writeJSON(w, http.StatusOK, orchestratorSettingsDTO{ writeJSON(w, http.StatusOK, orchestratorSettingsDTO{
FIPSettleSeconds: f.fipSettleSeconds, FIPSettleSeconds: f.fipSettleSeconds,
HistoryRetentionCycles: f.historyRetentionCycles, HistoryRetentionCycles: f.historyRetentionCycles,
SelfCheckMaxAttempts: f.selfCheckMaxAttempts,
}) })
}) })
mux.HandleFunc("PUT /api/v1/admin/config/orchestrator", func(w http.ResponseWriter, r *http.Request) { mux.HandleFunc("PUT /api/v1/admin/config/orchestrator", func(w http.ResponseWriter, r *http.Request) {
@@ -325,11 +332,17 @@ func (f *fakeControlAPI) handler() http.Handler {
writeAPIErr(w, http.StatusBadRequest, "history_retention_cycles must be >= 0") writeAPIErr(w, http.StatusBadRequest, "history_retention_cycles must be >= 0")
return return
} }
if req.SelfCheckMaxAttempts < 1 || req.SelfCheckMaxAttempts > 50 {
writeAPIErr(w, http.StatusBadRequest, "self_check_max_attempts must be in 1..50")
return
}
f.fipSettleSeconds = req.FIPSettleSeconds f.fipSettleSeconds = req.FIPSettleSeconds
f.historyRetentionCycles = req.HistoryRetentionCycles f.historyRetentionCycles = req.HistoryRetentionCycles
f.selfCheckMaxAttempts = req.SelfCheckMaxAttempts
writeJSON(w, http.StatusOK, orchestratorSettingsDTO{ writeJSON(w, http.StatusOK, orchestratorSettingsDTO{
FIPSettleSeconds: f.fipSettleSeconds, FIPSettleSeconds: f.fipSettleSeconds,
HistoryRetentionCycles: f.historyRetentionCycles, HistoryRetentionCycles: f.historyRetentionCycles,
SelfCheckMaxAttempts: f.selfCheckMaxAttempts,
}) })
}) })
@@ -474,7 +487,7 @@ func (f *fakeControlAPI) handler() http.Handler {
writeAPIErr(w, http.StatusNotFound, "unknown ip: "+addr) writeAPIErr(w, http.StatusNotFound, "unknown ip: "+addr)
return return
} }
writeJSON(w, http.StatusOK, registryHistoryResponse{Registry: item, Checks: f.registryChecks[addr]}) writeJSON(w, http.StatusOK, registryHistoryResponse{Registry: item, Checks: f.registryChecks[addr], SelfCheckFailedOn: f.selfCheckFailedOn})
}) })
mux.HandleFunc("GET /api/v1/admin/config/subnets", func(w http.ResponseWriter, r *http.Request) { mux.HandleFunc("GET /api/v1/admin/config/subnets", func(w http.ResponseWriter, r *http.Request) {
+7
View File
@@ -95,6 +95,9 @@ type ipDetailResponse struct {
IP ipQueueItem `json:"ip"` IP ipQueueItem `json:"ip"`
Checks []check `json:"checks"` Checks []check `json:"checks"`
Events []event `json:"events"` Events []event `json:"events"`
// SelfCheckFailedOn lists the validators whose self-check of this address
// failed (in its current run).
SelfCheckFailedOn []string `json:"self_check_failed_on"`
} }
type validator struct { type validator struct {
@@ -308,6 +311,9 @@ func (t typeStat) Class() string { return statClass(t.OK, t.Total) }
type registryHistoryResponse struct { type registryHistoryResponse struct {
Registry registryItem `json:"registry"` Registry registryItem `json:"registry"`
Checks []check `json:"checks"` Checks []check `json:"checks"`
// SelfCheckFailedOn lists the validators whose self-check of this address
// failed, over every run.
SelfCheckFailedOn []string `json:"self_check_failed_on"`
} }
type validatorDTO struct { type validatorDTO struct {
@@ -344,6 +350,7 @@ type errorResponse struct {
type orchestratorSettingsDTO struct { type orchestratorSettingsDTO struct {
FIPSettleSeconds int `json:"fip_settle_seconds"` FIPSettleSeconds int `json:"fip_settle_seconds"`
HistoryRetentionCycles int `json:"history_retention_cycles"` HistoryRetentionCycles int `json:"history_retention_cycles"`
SelfCheckMaxAttempts int `json:"self_check_max_attempts"`
} }
type inboundChecksDTO struct { type inboundChecksDTO struct {
+6 -1
View File
@@ -77,7 +77,12 @@ func (s *Server) handleSettingsPut(w http.ResponseWriter, r *http.Request) {
s.renderSettingsForm(w, r, &apiErr{Status: http.StatusBadRequest, Message: "глубина истории должна быть целым числом циклов"}) s.renderSettingsForm(w, r, &apiErr{Status: http.StatusBadRequest, Message: "глубина истории должна быть целым числом циклов"})
return return
} }
_, err = s.CA.PutOrchestratorSettings(r.Context(), seconds, retentionCycles) selfCheckMax, err := strconv.Atoi(r.PostFormValue("self_check_max_attempts"))
if err != nil {
s.renderSettingsForm(w, r, &apiErr{Status: http.StatusBadRequest, Message: "потолок провалов self-check должен быть целым числом"})
return
}
_, err = s.CA.PutOrchestratorSettings(r.Context(), seconds, retentionCycles, selfCheckMax)
s.renderSettingsForm(w, r, err) s.renderSettingsForm(w, r, err)
} }
+48 -4
View File
@@ -307,9 +307,12 @@ func TestSettingsGetAndPut(t *testing.T) {
if !strings.Contains(page, `value="0"`) { if !strings.Contains(page, `value="0"`) {
t.Fatalf("expected default 0 in the form, got:\n%s", page) t.Fatalf("expected default 0 in the form, got:\n%s", page)
} }
if !strings.Contains(page, "Потолок провалов self-check на адрес") || !strings.Contains(page, `id="self_check_max_attempts" name="self_check_max_attempts" min="1" max="50" step="1" value="5"`) {
t.Fatalf("expected the self-check ceiling field with the default 5, got:\n%s", page)
}
body := postForm(t, ts, "PUT", "/settings", map[string][]string{ body := postForm(t, ts, "PUT", "/settings", map[string][]string{
"fip_settle_seconds": {"15"}, "history_retention_cycles": {"10"}, "fip_settle_seconds": {"15"}, "history_retention_cycles": {"10"}, "self_check_max_attempts": {"8"},
}) })
if !strings.Contains(body, `value="15"`) || !strings.Contains(body, `value="10"`) { if !strings.Contains(body, `value="15"`) || !strings.Contains(body, `value="10"`) {
t.Fatalf("expected updated values 15/10 in re-rendered form, got:\n%s", body) t.Fatalf("expected updated values 15/10 in re-rendered form, got:\n%s", body)
@@ -320,11 +323,14 @@ func TestSettingsGetAndPut(t *testing.T) {
if fake.historyRetentionCycles != 10 { if fake.historyRetentionCycles != 10 {
t.Fatalf("expected fake control-api history_retention_cycles updated, got %d", fake.historyRetentionCycles) t.Fatalf("expected fake control-api history_retention_cycles updated, got %d", fake.historyRetentionCycles)
} }
if fake.selfCheckMaxAttempts != 8 || !strings.Contains(body, `name="self_check_max_attempts" min="1" max="50" step="1" value="8"`) {
t.Fatalf("expected the self-check ceiling 8 saved and shown, got %d:\n%s", fake.selfCheckMaxAttempts, body)
}
// A control-api validation error (negative value here) surfaces via // A control-api validation error (negative value here) surfaces via
// the banner, not a crash. // the banner, not a crash.
body = postForm(t, ts, "PUT", "/settings", map[string][]string{ body = postForm(t, ts, "PUT", "/settings", map[string][]string{
"fip_settle_seconds": {"-1"}, "history_retention_cycles": {"10"}, "fip_settle_seconds": {"-1"}, "history_retention_cycles": {"10"}, "self_check_max_attempts": {"8"},
}) })
if !strings.Contains(body, "alert-warning") { if !strings.Contains(body, "alert-warning") {
t.Fatalf("expected client error banner for invalid value, got:\n%s", body) t.Fatalf("expected client error banner for invalid value, got:\n%s", body)
@@ -333,7 +339,7 @@ func TestSettingsGetAndPut(t *testing.T) {
// A non-numeric value is caught by the dashboard itself before it ever // A non-numeric value is caught by the dashboard itself before it ever
// reaches control-api. // reaches control-api.
body = postForm(t, ts, "PUT", "/settings", map[string][]string{ body = postForm(t, ts, "PUT", "/settings", map[string][]string{
"fip_settle_seconds": {"not-a-number"}, "history_retention_cycles": {"10"}, "fip_settle_seconds": {"not-a-number"}, "history_retention_cycles": {"10"}, "self_check_max_attempts": {"8"},
}) })
if !strings.Contains(body, "alert-warning") { if !strings.Contains(body, "alert-warning") {
t.Fatalf("expected client error banner for non-numeric value, got:\n%s", body) t.Fatalf("expected client error banner for non-numeric value, got:\n%s", body)
@@ -341,11 +347,49 @@ func TestSettingsGetAndPut(t *testing.T) {
// Same for a non-numeric retention value. // Same for a non-numeric retention value.
body = postForm(t, ts, "PUT", "/settings", map[string][]string{ body = postForm(t, ts, "PUT", "/settings", map[string][]string{
"fip_settle_seconds": {"15"}, "history_retention_cycles": {"not-a-number"}, "fip_settle_seconds": {"15"}, "history_retention_cycles": {"not-a-number"}, "self_check_max_attempts": {"8"},
}) })
if !strings.Contains(body, "alert-warning") { if !strings.Contains(body, "alert-warning") {
t.Fatalf("expected client error banner for non-numeric retention value, got:\n%s", body) t.Fatalf("expected client error banner for non-numeric retention value, got:\n%s", body)
} }
// A ceiling outside 1..50 is refused by control-api and shown as a
// warning; the form keeps the stored value.
body = postForm(t, ts, "PUT", "/settings", map[string][]string{
"fip_settle_seconds": {"15"}, "history_retention_cycles": {"10"}, "self_check_max_attempts": {"51"},
})
if !strings.Contains(body, "alert-warning") || fake.selfCheckMaxAttempts != 8 {
t.Fatalf("expected a warning and the stored ceiling 8 kept, got %d:\n%s", fake.selfCheckMaxAttempts, body)
}
// A non-numeric ceiling is caught by the dashboard itself.
body = postForm(t, ts, "PUT", "/settings", map[string][]string{
"fip_settle_seconds": {"15"}, "history_retention_cycles": {"10"}, "self_check_max_attempts": {"many"},
})
if !strings.Contains(body, "alert-warning") {
t.Fatalf("expected client error banner for non-numeric ceiling, got:\n%s", body)
}
}
// The address pages say on which validators the self-check failed.
func TestAddressPagesShowSelfCheckFailedOn(t *testing.T) {
fake, caURL := newFakeControlAPI(t)
now := time.Now()
fake.ips = []ipQueueItem{{IPAddress: "5.5.5.5", State: "checking", UpdatedAt: now, CreatedAt: now}}
fake.registry["5.5.5.5"] = registryItem{IPAddress: "5.5.5.5", FirstSeenAt: now, LastSeenAt: now}
ts := newTestServer(t, caURL)
for _, path := range []string{"/ips/5.5.5.5", "/registry/5.5.5.5"} {
if page := get(t, ts, path); strings.Contains(page, "Self-check не прошёл на") {
t.Fatalf("%s: no failures, but the line is shown:\n%s", path, page)
}
}
fake.selfCheckFailedOn = []string{"vkiplab-v17", "vkiplab-v13"}
for _, path := range []string{"/ips/5.5.5.5", "/registry/5.5.5.5"} {
if page := get(t, ts, path); !strings.Contains(page, "Self-check не прошёл на: vkiplab-v17, vkiplab-v13") {
t.Fatalf("%s: expected the failed validators line, got:\n%s", path, page)
}
}
} }
func TestSettingsPageShowsInboundChecks(t *testing.T) { func TestSettingsPageShowsInboundChecks(t *testing.T) {
@@ -34,6 +34,7 @@
создан: {{fmtTime .Detail.IP.CreatedAt}} · назначен: {{fmtTime .Detail.IP.AssignedAt}} · создан: {{fmtTime .Detail.IP.CreatedAt}} · назначен: {{fmtTime .Detail.IP.AssignedAt}} ·
агрегирован: {{fmtTime .Detail.IP.AggregatedAt}} · FIP освобождён: {{fmtTime .Detail.IP.FIPReleasedAt}} агрегирован: {{fmtTime .Detail.IP.AggregatedAt}} · FIP освобождён: {{fmtTime .Detail.IP.FIPReleasedAt}}
</p> </p>
{{if .Detail.SelfCheckFailedOn}}<p class="muted">Self-check не прошёл на: {{join .Detail.SelfCheckFailedOn ", "}}</p>{{end}}
<h2 class="section-title">Проверки (попытка {{.Detail.IP.AttemptNumber}})</h2> <h2 class="section-title">Проверки (попытка {{.Detail.IP.AttemptNumber}})</h2>
{{if .Detail.Checks}} {{if .Detail.Checks}}
@@ -34,6 +34,7 @@
впервые замечен: {{fmtTime .History.Registry.FirstSeenAt}} · последний раз замечен: {{fmtTime .History.Registry.LastSeenAt}} впервые замечен: {{fmtTime .History.Registry.FirstSeenAt}} · последний раз замечен: {{fmtTime .History.Registry.LastSeenAt}}
· всего циклов проверки: {{.History.Registry.TotalCycles}} · всего циклов проверки: {{.History.Registry.TotalCycles}}
</p> </p>
{{if .History.SelfCheckFailedOn}}<p class="muted">Self-check не прошёл на: {{join .History.SelfCheckFailedOn ", "}}</p>{{end}}
<h2 class="section-title">История проверок (все сохранённые циклы)</h2> <h2 class="section-title">История проверок (все сохранённые циклы)</h2>
{{if .History.Checks}} {{if .History.Checks}}
@@ -80,6 +80,10 @@
<a href="/registry">реестре</a> для каждого адреса. Адрес и его накопленная статистика остаются в реестре даже <a href="/registry">реестре</a> для каждого адреса. Адрес и его накопленная статистика остаются в реестре даже
после удаления из очереди — эта настройка ограничивает только глубину истории конкретных проверок, не сам реестр. после удаления из очереди — эта настройка ограничивает только глубину истории конкретных проверок, не сам реестр.
<code>0</code> — хранить без ограничения.</p> <code>0</code> — хранить без ограничения.</p>
<p class="muted" style="margin-bottom:16px">Сколько раз self-check может не пройти у одного адреса, прежде чем адрес
получит итог <code>fail</code> (допустимо 1–50, по умолчанию 5). Повторы идут на других валидаторах: валидатор, на
котором адрес не прошёл self-check, этому адресу больше не выдаётся (на остальные адреса это не влияет).
Значение действует на следующих повторах, без перезапуска.</p>
<form hx-put="/settings" hx-target="#settings-form-wrap" hx-swap="innerHTML"> <form hx-put="/settings" hx-target="#settings-form-wrap" hx-swap="innerHTML">
<div class="field-row"> <div class="field-row">
<div class="field"> <div class="field">
@@ -90,6 +94,10 @@
<label for="history_retention_cycles">Глубина истории проверок (циклов на адрес, 0 = не ограничено)</label> <label for="history_retention_cycles">Глубина истории проверок (циклов на адрес, 0 = не ограничено)</label>
<input type="number" id="history_retention_cycles" name="history_retention_cycles" min="0" step="1" value="{{.Settings.HistoryRetentionCycles}}" required> <input type="number" id="history_retention_cycles" name="history_retention_cycles" min="0" step="1" value="{{.Settings.HistoryRetentionCycles}}" required>
</div> </div>
<div class="field">
<label for="self_check_max_attempts">Потолок провалов self-check на адрес</label>
<input type="number" id="self_check_max_attempts" name="self_check_max_attempts" min="1" max="50" step="1" value="{{.Settings.SelfCheckMaxAttempts}}" required>
</div>
<button type="submit" class="btn btn-primary">Сохранить</button> <button type="submit" class="btn btn-primary">Сохранить</button>
</div> </div>
</form> </form>
+4
View File
@@ -46,6 +46,9 @@ var verdictIntegritySchema string
//go:embed migrations/0011_check_runs.sql //go:embed migrations/0011_check_runs.sql
var checkRunsSchema string var checkRunsSchema string
//go:embed migrations/0012_self_check_failures.sql
var selfCheckFailuresSchema string
// migrations is the ordered list of schema versions. Each entry's SQL is // migrations is the ordered list of schema versions. Each entry's SQL is
// applied, in order, for any version greater than the database's current // applied, in order, for any version greater than the database's current
// PRAGMA user_version — so a fresh database walks the whole list and an // PRAGMA user_version — so a fresh database walks the whole list and an
@@ -65,6 +68,7 @@ var migrations = []struct {
{9, scaleIndexesSchema}, {9, scaleIndexesSchema},
{10, verdictIntegritySchema}, {10, verdictIntegritySchema},
{11, checkRunsSchema}, {11, checkRunsSchema},
{12, selfCheckFailuresSchema},
} }
type DB struct { type DB struct {
@@ -0,0 +1,38 @@
-- Self-check failures per address (see
-- docs/changes/2026-10-04_08-01_self-check-exclude-validator-plan.md).
--
-- A failed self-check no longer sends the address back to the validator that
-- failed it: ClaimNextQueued skips an address for every validator that failed
-- it in the current round. ip_self_check_failures is the permanent history of
-- those failures; it is keyed by registry_id (like checks and events), so it
-- outlives the ip_queue row and is not touched by re-checks. No foreign keys
-- on purpose: the manual cleanup in docs/ADMIN_CLEANUP.md deletes freely.
--
-- A round is the part of a series of failures in which validators that failed
-- stay excluded. ip_queue.sc_round_start_cycle is the first cycle_id of the
-- current round: a failure excludes its validator only if its cycle_id is not
-- below it. cycle_id (not attempt_number) is used because it never repeats for
-- an address, even when the ip_queue row is deleted and created again.
--
-- ip_queue.sc_failures counts self-check failures of the current series; a
-- manual re-check or re-submission starts a new series.
--
-- settings.self_check_max_attempts is the ceiling of self-check failures per
-- address, after which the address gets the verdict fail (1..50, default 5).
CREATE TABLE ip_self_check_failures (
id INTEGER PRIMARY KEY AUTOINCREMENT,
registry_id INTEGER NOT NULL,
run_id INTEGER,
cycle_id INTEGER NOT NULL,
attempt_number INTEGER NOT NULL,
validator_id TEXT NOT NULL,
failed_at TIMESTAMP NOT NULL,
detail TEXT NOT NULL DEFAULT ''
);
CREATE INDEX idx_sc_failures_registry ON ip_self_check_failures(registry_id, validator_id);
ALTER TABLE ip_queue ADD COLUMN sc_failures INTEGER NOT NULL DEFAULT 0;
ALTER TABLE ip_queue ADD COLUMN sc_round_start_cycle INTEGER NOT NULL DEFAULT 0;
ALTER TABLE settings ADD COLUMN self_check_max_attempts INTEGER NOT NULL DEFAULT 5;
+10
View File
@@ -276,10 +276,20 @@ type Settings struct {
// HistoryRetentionCycles caps how many recent check cycles are kept per // HistoryRetentionCycles caps how many recent check cycles are kept per
// registry address (see PruneRegistryHistory); 0 means unlimited. // registry address (see PruneRegistryHistory); 0 means unlimited.
HistoryRetentionCycles int HistoryRetentionCycles int
// SelfCheckMaxAttempts is the ceiling of failed self-checks per address
// (1..MaxSelfCheckMaxAttempts); reaching it gives the address the verdict
// fail (see FailSelfCheck).
SelfCheckMaxAttempts int
CreatedAt time.Time CreatedAt time.Time
UpdatedAt time.Time UpdatedAt time.Time
} }
// Bounds of Settings.SelfCheckMaxAttempts.
const (
MinSelfCheckMaxAttempts = 1
MaxSelfCheckMaxAttempts = 50
)
// InboundChecksSettings is the singleton row describing what the prober // InboundChecksSettings is the singleton row describing what the prober
// checks on every site for every in-flight IP (TCP ports + optional ICMP). // checks on every site for every in-flight IP (TCP ports + optional ICMP).
// Admin-configurable at runtime (see queries_inbound.go). // Admin-configurable at runtime (see queries_inbound.go).
+194 -14
View File
@@ -44,10 +44,10 @@ func (d *DB) SeedQueue(ctx context.Context, addresses []string) error {
} }
} }
if _, err := tx.ExecContext(ctx, ` if _, err := tx.ExecContext(ctx, `
INSERT INTO ip_queue (ip_address, sequence, state, registry_id, cycle_id, run_id, created_at, updated_at) INSERT INTO ip_queue (ip_address, sequence, state, registry_id, cycle_id, run_id, sc_round_start_cycle, created_at, updated_at)
VALUES (?, ?, ?, ?, ?, ?, ?, ?) VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?)
ON CONFLICT(ip_address) DO NOTHING ON CONFLICT(ip_address) DO NOTHING
`, addr, i, IPQueued, registryID, cycle, runID, now, now); err != nil { `, addr, i, IPQueued, registryID, cycle, runID, cycle, now, now); err != nil {
return fmt.Errorf("seed %s: %w", addr, err) return fmt.Errorf("seed %s: %w", addr, err)
} }
} }
@@ -55,8 +55,11 @@ func (d *DB) SeedQueue(ctx context.Context, addresses []string) error {
} }
// ClaimNextQueued atomically hands the next queued IP (lowest sequence) to // ClaimNextQueued atomically hands the next queued IP (lowest sequence) to
// the given idle validator. It returns (nil, nil) if the validator isn't // the given idle validator, skipping addresses this validator failed the
// idle or no IP is queued. The DB connection pool is capped at one physical // self-check of in the current round (see FailSelfCheck): such an address
// stays queued for the other validators and does not block the ones behind
// it. It returns (nil, nil) if the validator isn't idle or no IP is
// claimable for it. The DB connection pool is capped at one physical
// connection (see Open), so this transaction already has exclusive access // connection (see Open), so this transaction already has exclusive access
// to the database for its duration — no other claim, requeue, or update can // to the database for its duration — no other claim, requeue, or update can
// interleave — which combined with the conditional UPDATEs (checked via // interleave — which combined with the conditional UPDATEs (checked via
@@ -82,9 +85,12 @@ func (d *DB) ClaimNextQueued(ctx context.Context, validatorID string, leaseTTL t
var item IPQueueItem var item IPQueueItem
err = tx.QueryRowContext(ctx, ` err = tx.QueryRowContext(ctx, `
SELECT id, ip_address, sequence, attempt_number, retry_count SELECT q.id, q.ip_address, q.sequence, q.attempt_number, q.retry_count
FROM ip_queue WHERE state=? ORDER BY sequence LIMIT 1 FROM ip_queue q WHERE q.state=? AND NOT EXISTS (
`, IPQueued).Scan(&item.ID, &item.IPAddress, &item.Sequence, &item.AttemptNumber, &item.RetryCount) SELECT 1 FROM ip_self_check_failures f
WHERE f.registry_id=q.registry_id AND f.validator_id=? AND f.cycle_id>=q.sc_round_start_cycle)
ORDER BY q.sequence LIMIT 1
`, IPQueued, validatorID).Scan(&item.ID, &item.IPAddress, &item.Sequence, &item.AttemptNumber, &item.RetryCount)
if err == sql.ErrNoRows { if err == sql.ErrNoRows {
return nil, nil return nil, nil
} }
@@ -352,6 +358,178 @@ func (d *DB) RequeueOrFail(ctx context.Context, ipID int64, validatorID string,
return tx.Commit() return tx.Commit()
} }
// SelfCheckFailure is the outcome of FailSelfCheck.
type SelfCheckFailure struct {
// Failures is the number of failed self-checks in the address's current
// series, this one included.
Failures int
// Failed: the ceiling was reached and the address got the verdict fail.
Failed bool
// Validators lists the validators that failed the self-check in this
// series, oldest first (with repeats if one failed it more than once).
Validators []string
// NewRound: the address was queued again, and every working validator had
// already failed it in the round, so the round was reset — the exclusions
// no longer apply (always false when Failed).
NewRound bool
}
// FailSelfCheck handles a failed self-check of the address ipID held by
// validatorID, in one transaction: it records the failure in
// ip_self_check_failures and in ip_queue.sc_failures, then
//
// - if sc_failures reached maxAttempts: marks the address failed with the
// verdict fail;
// - otherwise sends it back to the queue without touching retry_count (a
// failed self-check has its own ceiling, unlike a failed association or
// an expired lease, see RequeueOrFail). The validator is excluded from
// this address for the rest of the round (ClaimNextQueued). If no
// working validator (idle, assigned, checking) is left without a failure
// in the round, a new round starts: the exclusions lapse and the retry
// follows the usual rules, so with fewer validators than maxAttempts the
// address never gets stuck in the queue.
//
// The validator is freed in the same transaction. Returns ErrInvalidState if
// the address is not awaiting_self_check on this validator (a late report).
func (d *DB) FailSelfCheck(ctx context.Context, ipID int64, validatorID, detail string, maxAttempts int) (SelfCheckFailure, error) {
var out SelfCheckFailure
tx, err := d.BeginTx(ctx, nil)
if err != nil {
return out, err
}
defer tx.Rollback()
var registryID, cycle, roundStart int64
var runID sql.NullInt64
var attempt, failures int
err = tx.QueryRowContext(ctx, `
SELECT registry_id, run_id, cycle_id, attempt_number, sc_failures, sc_round_start_cycle
FROM ip_queue WHERE id=? AND state=? AND owner_validator_id=?
`, ipID, IPAwaitingSelfCheck, validatorID).Scan(&registryID, &runID, &cycle, &attempt, &failures, &roundStart)
if err == sql.ErrNoRows {
return out, fmt.Errorf("ip_id %d is not awaiting self-check on %s: %w", ipID, validatorID, ErrInvalidState)
}
if err != nil {
return out, err
}
now := timeToDB(Now())
if _, err := tx.ExecContext(ctx, `
INSERT INTO ip_self_check_failures (registry_id, run_id, cycle_id, attempt_number, validator_id, failed_at, detail)
VALUES (?, ?, ?, ?, ?, ?, ?)
`, registryID, runID, cycle, attempt, validatorID, now, detail); err != nil {
return out, fmt.Errorf("record self-check failure: %w", err)
}
failures++
out.Failures = failures
rows, err := tx.QueryContext(ctx, `
SELECT validator_id FROM (
SELECT id, validator_id FROM ip_self_check_failures WHERE registry_id=? ORDER BY id DESC LIMIT ?
) ORDER BY id
`, registryID, failures)
if err != nil {
return out, err
}
for rows.Next() {
var v string
if err := rows.Scan(&v); err != nil {
rows.Close()
return out, err
}
out.Validators = append(out.Validators, v)
}
if err := rows.Err(); err != nil {
rows.Close()
return out, err
}
rows.Close()
if failures >= maxAttempts {
out.Failed = true
if _, err := tx.ExecContext(ctx, `
UPDATE ip_queue SET
state=?, sc_failures=?, overall_result=?, aggregated_at=?, updated_at=?
WHERE id=?
`, IPFailed, failures, ResultFail, now, now, ipID); err != nil {
return out, err
}
if err := upsertRunResultTx(ctx, tx, ipID, ResultFail, -1, now); err != nil {
return out, err
}
if err := finalizeRunsTx(ctx, tx, now); err != nil {
return out, err
}
} else {
newCycle, err := nextRegistryCycleTx(ctx, tx, registryID, now)
if err != nil {
return out, err
}
var left int
if err := tx.QueryRowContext(ctx, `
SELECT COUNT(*) FROM validators v
WHERE v.state IN (?, ?, ?) AND NOT EXISTS (
SELECT 1 FROM ip_self_check_failures f
WHERE f.registry_id=? AND f.validator_id=v.validator_id AND f.cycle_id>=?)
`, ValidatorIdle, ValidatorAssigned, ValidatorChecking, registryID, roundStart).Scan(&left); err != nil {
return out, err
}
if left == 0 {
out.NewRound = true
roundStart = int64(newCycle)
}
if _, err := tx.ExecContext(ctx, `
UPDATE ip_queue SET
state=?, owner_validator_id=NULL, fip_id='', attempt_number=attempt_number+1,
cycle_id=?, sc_failures=?, sc_round_start_cycle=?, lease_expires_at=NULL, egress_complete=0,
overall_result='', assigned_at=NULL, checking_started_at=NULL, fip_associated_at=NULL, updated_at=?
WHERE id=?
`, IPQueued, newCycle, failures, roundStart, now, ipID); err != nil {
return out, err
}
}
if _, err := tx.ExecContext(ctx, freeValidatorSQL, now, validatorID, ipID); err != nil {
return out, err
}
return out, tx.Commit()
}
// GetIPRunID returns the check run the queue row belongs to (0 if none).
func (d *DB) GetIPRunID(ctx context.Context, ipID int64) (int64, error) {
var runID sql.NullInt64
if err := d.QueryRowContext(ctx, `SELECT run_id FROM ip_queue WHERE id=?`, ipID).Scan(&runID); err != nil {
return 0, err
}
return runID.Int64, nil
}
// ListSelfCheckFailedOn returns the distinct validators whose self-check of
// the address failed, in alphabetical order: over the whole history of the
// address, or, if runID > 0, only the failures that happened in that run.
func (d *DB) ListSelfCheckFailedOn(ctx context.Context, registryID, runID int64) ([]string, error) {
q := `SELECT DISTINCT validator_id FROM ip_self_check_failures WHERE registry_id=?`
args := []any{registryID}
if runID > 0 {
q += ` AND run_id=?`
args = append(args, runID)
}
rows, err := d.QueryContext(ctx, q+` ORDER BY validator_id`, args...)
if err != nil {
return nil, err
}
defer rows.Close()
out := []string{}
for rows.Next() {
var v string
if err := rows.Scan(&v); err != nil {
return nil, err
}
out = append(out, v)
}
return out, rows.Err()
}
// SubmitIPs is the single admin entry point for both "add new addresses to // SubmitIPs is the single admin entry point for both "add new addresses to
// the queue" and "force a re-check of an already-finished address" — the // the queue" and "force a re-check of an already-finished address" — the
// same list can freely mix both. Addresses are processed in one // same list can freely mix both. Addresses are processed in one
@@ -360,7 +538,9 @@ func (d *DB) RequeueOrFail(ctx context.Context, ipID int64, validatorID string,
// - unknown address: inserted as a new queued row. // - unknown address: inserted as a new queued row.
// - address currently done/failed/occupied: reset to queued (new attempt, // - address currently done/failed/occupied: reset to queued (new attempt,
// retry_count cleared — this is a deliberate admin-triggered restart, // retry_count cleared — this is a deliberate admin-triggered restart,
// not a system retry). // not a system retry). It also starts a new series of self-check
// failures (sc_failures=0, validators that failed it before are no
// longer excluded); the history in ip_self_check_failures stays.
// - address currently queued (not yet claimed): left in state=queued, // - address currently queued (not yet claimed): left in state=queued,
// only its sequence is updated. // only its sequence is updated.
// - address currently mid-check (assigning_fip / awaiting_self_check / // - address currently mid-check (assigning_fip / awaiting_self_check /
@@ -426,9 +606,9 @@ func (d *DB) SubmitIPsAs(ctx context.Context, addresses []string, kind string) (
return result, rErr return result, rErr
} }
if _, err := tx.ExecContext(ctx, ` if _, err := tx.ExecContext(ctx, `
INSERT INTO ip_queue (ip_address, sequence, state, registry_id, cycle_id, run_id, created_at, updated_at) INSERT INTO ip_queue (ip_address, sequence, state, registry_id, cycle_id, run_id, sc_round_start_cycle, created_at, updated_at)
VALUES (?, ?, ?, ?, ?, ?, ?, ?) VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?)
`, addr, seq, IPQueued, registryID, cycle, rid, now, now); err != nil { `, addr, seq, IPQueued, registryID, cycle, rid, cycle, now, now); err != nil {
return result, fmt.Errorf("insert %s: %w", addr, err) return result, fmt.Errorf("insert %s: %w", addr, err)
} }
result.Added = append(result.Added, addr) result.Added = append(result.Added, addr)
@@ -453,10 +633,10 @@ func (d *DB) SubmitIPsAs(ctx context.Context, addresses []string, kind string) (
UPDATE ip_queue SET UPDATE ip_queue SET
state=?, sequence=?, owner_validator_id=NULL, fip_id='', retry_count=0, state=?, sequence=?, owner_validator_id=NULL, fip_id='', retry_count=0,
attempt_number=attempt_number+1, cycle_id=?, lease_expires_at=NULL, egress_complete=0, attempt_number=attempt_number+1, cycle_id=?, lease_expires_at=NULL, egress_complete=0,
overall_result='', run_id=?, overall_result='', run_id=?, sc_failures=0, sc_round_start_cycle=?,
assigned_at=NULL, checking_started_at=NULL, fip_associated_at=NULL, aggregated_at=NULL, fip_released_at=NULL, updated_at=? assigned_at=NULL, checking_started_at=NULL, fip_associated_at=NULL, aggregated_at=NULL, fip_released_at=NULL, updated_at=?
WHERE ip_address=? WHERE ip_address=?
`, IPQueued, seq, cycle, rid, now, addr); err != nil { `, IPQueued, seq, cycle, rid, cycle, now, addr); err != nil {
return result, fmt.Errorf("requeue %s: %w", addr, err) return result, fmt.Errorf("requeue %s: %w", addr, err)
} }
result.Requeued = append(result.Requeued, addr) result.Requeued = append(result.Requeued, addr)
+1 -1
View File
@@ -383,7 +383,7 @@ func TestMigration0011BuildsRunsFromExistingData(t *testing.T) {
} }
var ver int var ver int
d.QueryRowContext(ctx, `PRAGMA user_version`).Scan(&ver) d.QueryRowContext(ctx, `PRAGMA user_version`).Scan(&ver)
if ver != 11 { if ver != 12 {
t.Errorf("user_version = %d", ver) t.Errorf("user_version = %d", ver)
} }
} }
+190
View File
@@ -0,0 +1,190 @@
package db
import (
"context"
"errors"
"testing"
"time"
"cloudipvalidator/internal/config"
)
// claimAwaitingSelfCheck claims the next address for the validator and brings
// it to awaiting_self_check, as the orchestrator does after the association.
func claimAwaitingSelfCheck(t *testing.T, ctx context.Context, d *DB, validatorID string) *IPQueueItem {
t.Helper()
item, err := d.ClaimNextQueued(ctx, validatorID, time.Minute)
if err != nil || item == nil {
t.Fatalf("claim for %s: item=%+v err=%v", validatorID, item, err)
}
if err := d.SetFIPAssociated(ctx, item.ID, "fip-"+item.IPAddress, time.Minute); err != nil {
t.Fatalf("set fip associated: %v", err)
}
return item
}
// A validator that failed the self-check of an address is not given that
// address again, takes the next one instead, and another validator takes the
// excluded one; the exclusion covers that address only.
func TestClaimSkipsAddressExcludedForValidator(t *testing.T) {
d, ctx := newTestDB(t)
if err := d.SeedQueue(ctx, []string{"1.1.1.1", "2.2.2.2"}); err != nil {
t.Fatal(err)
}
for _, v := range []string{"v1", "v2"} {
if err := d.AdminCreateValidator(ctx, v, "port-"+v); err != nil {
t.Fatal(err)
}
}
first := claimAwaitingSelfCheck(t, ctx, d, "v1")
if first.IPAddress != "1.1.1.1" {
t.Fatalf("expected 1.1.1.1 first, got %s", first.IPAddress)
}
res, err := d.FailSelfCheck(ctx, first.ID, "v1", "ip echo timeout", 5)
if err != nil {
t.Fatalf("fail self-check: %v", err)
}
if res.Failed || res.NewRound || res.Failures != 1 {
t.Fatalf("expected a plain retry after the first failure, got %+v", res)
}
back, _ := d.GetIP(ctx, first.ID)
if back.State != IPQueued || back.RetryCount != 0 {
t.Fatalf("expected queued with retry_count untouched, got state=%s retry_count=%d", back.State, back.RetryCount)
}
if v, _ := d.GetValidator(ctx, "v1"); v.State != ValidatorIdle {
t.Fatalf("expected v1 freed to idle, got %s", v.State)
}
// v1 skips 1.1.1.1 (still ahead in the queue) and takes 2.2.2.2.
second, err := d.ClaimNextQueued(ctx, "v1", time.Minute)
if err != nil || second == nil || second.IPAddress != "2.2.2.2" {
t.Fatalf("expected v1 to take 2.2.2.2, got %+v err=%v", second, err)
}
// v2 takes the excluded address.
other, err := d.ClaimNextQueued(ctx, "v2", time.Minute)
if err != nil || other == nil || other.IPAddress != "1.1.1.1" {
t.Fatalf("expected v2 to take 1.1.1.1, got %+v err=%v", other, err)
}
}
// When every working validator has failed the address in the round, a new
// round starts and the exclusions lapse; an unreachable validator does not
// count as working.
func TestSelfCheckNewRoundLiftsExclusions(t *testing.T) {
d, ctx := newTestDB(t)
if err := d.SeedQueue(ctx, []string{"1.1.1.1"}); err != nil {
t.Fatal(err)
}
for _, v := range []string{"v1", "v2"} {
if err := d.AdminCreateValidator(ctx, v, "port-"+v); err != nil {
t.Fatal(err)
}
}
if _, err := d.Exec(`UPDATE validators SET state=? WHERE validator_id='v2'`, ValidatorUnreachable); err != nil {
t.Fatal(err)
}
item := claimAwaitingSelfCheck(t, ctx, d, "v1")
res, err := d.FailSelfCheck(ctx, item.ID, "v1", "x", 5)
if err != nil {
t.Fatal(err)
}
if !res.NewRound {
t.Fatalf("expected a new round (the only working validator failed it), got %+v", res)
}
again, err := d.ClaimNextQueued(ctx, "v1", time.Minute)
if err != nil || again == nil || again.IPAddress != "1.1.1.1" {
t.Fatalf("expected v1 to get the address again in the new round, got %+v err=%v", again, err)
}
}
// Reaching the ceiling gives the verdict fail; a re-submission starts a new
// series (counter and exclusions) but keeps the failure history.
func TestSelfCheckCeilingAndResubmit(t *testing.T) {
d, ctx := newTestDB(t)
if err := d.SeedQueue(ctx, []string{"1.1.1.1"}); err != nil {
t.Fatal(err)
}
for _, v := range []string{"v1", "v2"} {
if err := d.AdminCreateValidator(ctx, v, "port-"+v); err != nil {
t.Fatal(err)
}
}
item := claimAwaitingSelfCheck(t, ctx, d, "v1")
if res, err := d.FailSelfCheck(ctx, item.ID, "v1", "x", 2); err != nil || res.Failed {
t.Fatalf("below the ceiling: res=%+v err=%v", res, err)
}
item = claimAwaitingSelfCheck(t, ctx, d, "v2")
res, err := d.FailSelfCheck(ctx, item.ID, "v2", "y", 2)
if err != nil {
t.Fatal(err)
}
if !res.Failed || res.Failures != 2 || len(res.Validators) != 2 || res.Validators[0] != "v1" || res.Validators[1] != "v2" {
t.Fatalf("expected fail at the ceiling on v1, v2, got %+v", res)
}
ip, _ := d.GetIP(ctx, item.ID)
if ip.State != IPFailed || ip.OverallResult != ResultFail {
t.Fatalf("expected failed/fail, got %s/%s", ip.State, ip.OverallResult)
}
if _, err := d.SubmitIPs(ctx, []string{"1.1.1.1"}); err != nil {
t.Fatal(err)
}
var failures int
if err := d.QueryRow(`SELECT sc_failures FROM ip_queue WHERE id=?`, item.ID).Scan(&failures); err != nil || failures != 0 {
t.Fatalf("expected the series reset to 0, got %d err=%v", failures, err)
}
if got, err := d.ListSelfCheckFailedOn(ctx, ip.RegistryID, 0); err != nil || len(got) != 2 || got[0] != "v1" || got[1] != "v2" {
t.Fatalf("expected the history kept (v1, v2), got %v err=%v", got, err)
}
// v1 failed it before the re-submission, but is no longer excluded.
if again, err := d.ClaimNextQueued(ctx, "v1", time.Minute); err != nil || again == nil {
t.Fatalf("expected v1 to claim the re-submitted address, got %+v err=%v", again, err)
}
}
// A late report for an address the validator does not hold is refused.
func TestFailSelfCheckRequiresAwaitingSelfCheckOnValidator(t *testing.T) {
d, ctx := newTestDB(t)
if err := d.SeedQueue(ctx, []string{"1.1.1.1"}); err != nil {
t.Fatal(err)
}
for _, v := range []string{"v1", "v2"} {
if err := d.AdminCreateValidator(ctx, v, "port-"+v); err != nil {
t.Fatal(err)
}
}
item := claimAwaitingSelfCheck(t, ctx, d, "v1")
if _, err := d.FailSelfCheck(ctx, item.ID, "v2", "late", 5); !errors.Is(err, ErrInvalidState) {
t.Fatalf("expected ErrInvalidState for another validator, got %v", err)
}
cur, _ := d.GetIP(ctx, item.ID)
if got, _ := d.ListSelfCheckFailedOn(ctx, cur.RegistryID, 0); len(got) != 0 {
t.Fatalf("a refused report must leave no history, got %v", got)
}
}
func TestSelfCheckMaxAttemptsSetting(t *testing.T) {
d, ctx := newTestDB(t)
if err := d.BootstrapFromConfig(ctx, &config.ControlAPI{}); err != nil {
t.Fatal(err)
}
if s, err := d.GetSettings(ctx); err != nil || s.SelfCheckMaxAttempts != 5 {
t.Fatalf("expected default 5, got %+v err=%v", s, err)
}
for _, bad := range []int{0, -1, 51} {
if err := d.SetSelfCheckMaxAttempts(ctx, bad); !errors.Is(err, ErrValidation) {
t.Fatalf("expected ErrValidation for %d, got %v", bad, err)
}
}
for _, ok := range []int{1, 50} {
if err := d.SetSelfCheckMaxAttempts(ctx, ok); err != nil {
t.Fatalf("set %d: %v", ok, err)
}
if s, _ := d.GetSettings(ctx); s.SelfCheckMaxAttempts != ok {
t.Fatalf("expected %d, got %d", ok, s.SelfCheckMaxAttempts)
}
}
}
+16 -2
View File
@@ -13,8 +13,8 @@ func (d *DB) GetSettings(ctx context.Context) (Settings, error) {
var s Settings var s Settings
var createdAt, updatedAt string var createdAt, updatedAt string
err := d.QueryRowContext(ctx, ` err := d.QueryRowContext(ctx, `
SELECT fip_settle_seconds, history_retention_cycles, created_at, updated_at FROM settings WHERE id=1 SELECT fip_settle_seconds, history_retention_cycles, self_check_max_attempts, created_at, updated_at FROM settings WHERE id=1
`).Scan(&s.FIPSettleSeconds, &s.HistoryRetentionCycles, &createdAt, &updatedAt) `).Scan(&s.FIPSettleSeconds, &s.HistoryRetentionCycles, &s.SelfCheckMaxAttempts, &createdAt, &updatedAt)
if err != nil { if err != nil {
return Settings{}, err return Settings{}, err
} }
@@ -58,3 +58,17 @@ func (d *DB) SetHistoryRetentionCycles(ctx context.Context, cycles int) error {
`, cycles, now) `, cycles, now)
return err return err
} }
// SetSelfCheckMaxAttempts persists the ceiling of failed self-checks per
// address (1..MaxSelfCheckMaxAttempts). It applies from the next failed
// self-check, without a restart.
func (d *DB) SetSelfCheckMaxAttempts(ctx context.Context, attempts int) error {
if attempts < MinSelfCheckMaxAttempts || attempts > MaxSelfCheckMaxAttempts {
return fmt.Errorf("self_check_max_attempts must be in %d..%d: %w", MinSelfCheckMaxAttempts, MaxSelfCheckMaxAttempts, ErrValidation)
}
now := timeToDB(Now())
_, err := d.ExecContext(ctx, `
UPDATE settings SET self_check_max_attempts=?, updated_at=? WHERE id=1
`, attempts, now)
return err
}
@@ -240,7 +240,7 @@ func TestMigration0010MarksRowsAfterVerdict(t *testing.T) {
t.Errorf("ssh: after_verdict=%d recorded=%s created=%s", a, rec, cr) t.Errorf("ssh: after_verdict=%d recorded=%s created=%s", a, rec, cr)
} }
var ver int var ver int
if err := d.QueryRowContext(ctx, `PRAGMA user_version`).Scan(&ver); err != nil || ver != 11 { if err := d.QueryRowContext(ctx, `PRAGMA user_version`).Scan(&ver); err != nil || ver != 12 {
t.Errorf("user_version=%d err=%v", ver, err) t.Errorf("user_version=%d err=%v", ver, err)
} }
} }
+12 -3
View File
@@ -162,12 +162,21 @@ type putCheckTypeRequest struct {
Targets []string `json:"targets"` Targets []string `json:"targets"`
} }
// orchestratorSettingsDTO doubles as both the GET response and the PUT // orchestratorSettingsDTO is the GET response for
// request body for /api/v1/admin/config/orchestrator — a single-field DTO, // /api/v1/admin/config/orchestrator.
// same shape both ways, like putSiteRequest/siteDTO.
type orchestratorSettingsDTO struct { type orchestratorSettingsDTO struct {
FIPSettleSeconds int `json:"fip_settle_seconds"` FIPSettleSeconds int `json:"fip_settle_seconds"`
HistoryRetentionCycles int `json:"history_retention_cycles"` HistoryRetentionCycles int `json:"history_retention_cycles"`
SelfCheckMaxAttempts int `json:"self_check_max_attempts"`
}
// putOrchestratorSettingsRequest is the PUT body. SelfCheckMaxAttempts is a
// pointer so that a client written before the field existed (it sends only
// the first two) leaves the ceiling unchanged instead of failing validation.
type putOrchestratorSettingsRequest struct {
FIPSettleSeconds int `json:"fip_settle_seconds"`
HistoryRetentionCycles int `json:"history_retention_cycles"`
SelfCheckMaxAttempts *int `json:"self_check_max_attempts"`
} }
// inboundChecksDTO doubles as both the GET response and the PUT request // inboundChecksDTO doubles as both the GET response and the PUT request
+13 -1
View File
@@ -178,11 +178,23 @@ func (s *Server) handleAdminIPDetail(w http.ResponseWriter, r *http.Request) {
writeError(w, http.StatusInternalServerError, err.Error()) writeError(w, http.StatusInternalServerError, err.Error())
return return
} }
// Self-check failures of this address in its current run.
runID, err := s.DB.GetIPRunID(r.Context(), item.ID)
if err != nil {
writeError(w, http.StatusInternalServerError, err.Error())
return
}
failedOn, err := s.DB.ListSelfCheckFailedOn(r.Context(), item.RegistryID, runID)
if err != nil {
writeError(w, http.StatusInternalServerError, err.Error())
return
}
writeJSON(w, http.StatusOK, struct { writeJSON(w, http.StatusOK, struct {
IP *db.IPQueueItem `json:"ip"` IP *db.IPQueueItem `json:"ip"`
Checks []db.Check `json:"checks"` Checks []db.Check `json:"checks"`
Events []db.Event `json:"events"` Events []db.Event `json:"events"`
}{item, checks, events}) SelfCheckFailedOn []string `json:"self_check_failed_on"`
}{item, checks, events, failedOn})
} }
func (s *Server) handleAdminValidators(w http.ResponseWriter, r *http.Request) { func (s *Server) handleAdminValidators(w http.ResponseWriter, r *http.Request) {
+24 -3
View File
@@ -1,6 +1,7 @@
package httpapi package httpapi
import ( import (
"fmt"
"net/http" "net/http"
"strconv" "strconv"
@@ -203,6 +204,7 @@ func (s *Server) handleConfigGetOrchestratorSettings(w http.ResponseWriter, r *h
writeJSON(w, http.StatusOK, orchestratorSettingsDTO{ writeJSON(w, http.StatusOK, orchestratorSettingsDTO{
FIPSettleSeconds: settings.FIPSettleSeconds, FIPSettleSeconds: settings.FIPSettleSeconds,
HistoryRetentionCycles: settings.HistoryRetentionCycles, HistoryRetentionCycles: settings.HistoryRetentionCycles,
SelfCheckMaxAttempts: settings.SelfCheckMaxAttempts,
}) })
} }
@@ -211,11 +213,17 @@ func (s *Server) handleConfigGetOrchestratorSettings(w http.ResponseWriter, r *h
// needs cross-field validation against the static lease_ttl_seconds/ // needs cross-field validation against the static lease_ttl_seconds/
// self_check_timeout_seconds config, which only the orchestrator has. // self_check_timeout_seconds config, which only the orchestrator has.
func (s *Server) handleConfigPutOrchestratorSettings(w http.ResponseWriter, r *http.Request) { func (s *Server) handleConfigPutOrchestratorSettings(w http.ResponseWriter, r *http.Request) {
var req orchestratorSettingsDTO var req putOrchestratorSettingsRequest
if err := readJSON(r, &req); err != nil { if err := readJSON(r, &req); err != nil {
writeError(w, http.StatusBadRequest, "invalid body: "+err.Error()) writeError(w, http.StatusBadRequest, "invalid body: "+err.Error())
return return
} }
// Checked before anything is saved, so a bad ceiling does not leave the
// other two settings half-applied. An omitted field keeps the current value.
if a := req.SelfCheckMaxAttempts; a != nil && (*a < db.MinSelfCheckMaxAttempts || *a > db.MaxSelfCheckMaxAttempts) {
writeError(w, http.StatusBadRequest, fmt.Sprintf("self_check_max_attempts must be in %d..%d", db.MinSelfCheckMaxAttempts, db.MaxSelfCheckMaxAttempts))
return
}
if err := s.Orch.SetFIPSettleSeconds(r.Context(), req.FIPSettleSeconds); err != nil { if err := s.Orch.SetFIPSettleSeconds(r.Context(), req.FIPSettleSeconds); err != nil {
writeDBError(w, err) writeDBError(w, err)
return return
@@ -224,9 +232,22 @@ func (s *Server) handleConfigPutOrchestratorSettings(w http.ResponseWriter, r *h
writeDBError(w, err) writeDBError(w, err)
return return
} }
if req.SelfCheckMaxAttempts != nil {
if err := s.DB.SetSelfCheckMaxAttempts(r.Context(), *req.SelfCheckMaxAttempts); err != nil {
writeDBError(w, err)
return
}
}
// Answer with what is stored now, so an omitted ceiling shows its current value.
settings, err := s.DB.GetSettings(r.Context())
if err != nil {
writeDBError(w, err)
return
}
writeJSON(w, http.StatusOK, orchestratorSettingsDTO{ writeJSON(w, http.StatusOK, orchestratorSettingsDTO{
FIPSettleSeconds: req.FIPSettleSeconds, FIPSettleSeconds: settings.FIPSettleSeconds,
HistoryRetentionCycles: req.HistoryRetentionCycles, HistoryRetentionCycles: settings.HistoryRetentionCycles,
SelfCheckMaxAttempts: settings.SelfCheckMaxAttempts,
}) })
} }
+23 -10
View File
@@ -426,19 +426,32 @@ func TestOrchestratorSettingsGetPut(t *testing.T) {
if err := json.Unmarshal(body, &got); err != nil { if err := json.Unmarshal(body, &got); err != nil {
t.Fatalf("unmarshal get response: %v", err) t.Fatalf("unmarshal get response: %v", err)
} }
if got.FIPSettleSeconds != 0 { if got.FIPSettleSeconds != 0 || got.SelfCheckMaxAttempts != 5 {
t.Fatalf("expected default 0, got %+v", got) t.Fatalf("expected defaults 0 and 5, got %+v", got)
} }
resp, body = fc.do(http.MethodPut, "/api/v1/admin/config/orchestrator", orchestratorSettingsDTO{FIPSettleSeconds: 20}) resp, body = fc.do(http.MethodPut, "/api/v1/admin/config/orchestrator", putOrchestratorSettingsRequest{FIPSettleSeconds: 20})
if resp.StatusCode != http.StatusOK { if resp.StatusCode != http.StatusOK {
t.Fatalf("put settings: status=%d body=%s", resp.StatusCode, body) t.Fatalf("put settings: status=%d body=%s", resp.StatusCode, body)
} }
if err := json.Unmarshal(body, &got); err != nil { if err := json.Unmarshal(body, &got); err != nil {
t.Fatalf("unmarshal put response: %v", err) t.Fatalf("unmarshal put response: %v", err)
} }
if got.FIPSettleSeconds != 20 { if got.FIPSettleSeconds != 20 || got.SelfCheckMaxAttempts != 5 {
t.Fatalf("expected 20, got %+v", got) t.Fatalf("expected 20 with the omitted ceiling kept at 5, got %+v", got)
}
eight := 8
resp, body = fc.do(http.MethodPut, "/api/v1/admin/config/orchestrator", putOrchestratorSettingsRequest{FIPSettleSeconds: 20, SelfCheckMaxAttempts: &eight})
if resp.StatusCode != http.StatusOK {
t.Fatalf("put settings with ceiling: status=%d body=%s", resp.StatusCode, body)
}
for _, bad := range []int{0, 51} {
bad := bad
resp, body = fc.do(http.MethodPut, "/api/v1/admin/config/orchestrator", putOrchestratorSettingsRequest{FIPSettleSeconds: 20, SelfCheckMaxAttempts: &bad})
if resp.StatusCode != http.StatusBadRequest {
t.Fatalf("expected 400 for self_check_max_attempts=%d, status=%d body=%s", bad, resp.StatusCode, body)
}
} }
resp, body = fc.do(http.MethodGet, "/api/v1/admin/config/orchestrator", nil) resp, body = fc.do(http.MethodGet, "/api/v1/admin/config/orchestrator", nil)
@@ -448,8 +461,8 @@ func TestOrchestratorSettingsGetPut(t *testing.T) {
if err := json.Unmarshal(body, &got); err != nil { if err := json.Unmarshal(body, &got); err != nil {
t.Fatalf("unmarshal get-after-put response: %v", err) t.Fatalf("unmarshal get-after-put response: %v", err)
} }
if got.FIPSettleSeconds != 20 { if got.FIPSettleSeconds != 20 || got.SelfCheckMaxAttempts != 8 {
t.Fatalf("expected 20 to persist, got %+v", got) t.Fatalf("expected 20 and 8 to persist, got %+v", got)
} }
} }
@@ -459,12 +472,12 @@ func TestOrchestratorSettingsPutValidation(t *testing.T) {
fc, _, _, _ := newConfigTestHarness(t) fc, _, _, _ := newConfigTestHarness(t)
// newConfigTestHarness: LeaseTTLSeconds=180, SelfCheckTimeoutSeconds=10. // newConfigTestHarness: LeaseTTLSeconds=180, SelfCheckTimeoutSeconds=10.
resp, body := fc.do(http.MethodPut, "/api/v1/admin/config/orchestrator", orchestratorSettingsDTO{FIPSettleSeconds: 175}) resp, body := fc.do(http.MethodPut, "/api/v1/admin/config/orchestrator", putOrchestratorSettingsRequest{FIPSettleSeconds: 175})
if resp.StatusCode != http.StatusBadRequest { if resp.StatusCode != http.StatusBadRequest {
t.Fatalf("expected 400 for settle seconds too close to lease ttl, status=%d body=%s", resp.StatusCode, body) t.Fatalf("expected 400 for settle seconds too close to lease ttl, status=%d body=%s", resp.StatusCode, body)
} }
resp, body = fc.do(http.MethodPut, "/api/v1/admin/config/orchestrator", orchestratorSettingsDTO{FIPSettleSeconds: -1}) resp, body = fc.do(http.MethodPut, "/api/v1/admin/config/orchestrator", putOrchestratorSettingsRequest{FIPSettleSeconds: -1})
if resp.StatusCode != http.StatusBadRequest { if resp.StatusCode != http.StatusBadRequest {
t.Fatalf("expected 400 for negative settle seconds, status=%d body=%s", resp.StatusCode, body) t.Fatalf("expected 400 for negative settle seconds, status=%d body=%s", resp.StatusCode, body)
} }
@@ -479,7 +492,7 @@ func TestFIPSettleDelayGatesAssignmentEndpoint(t *testing.T) {
ctx := context.Background() ctx := context.Background()
mock.Seed("fip-1", "9.9.9.9", "svc-project") mock.Seed("fip-1", "9.9.9.9", "svc-project")
resp, body := fc.do(http.MethodPut, "/api/v1/admin/config/orchestrator", orchestratorSettingsDTO{FIPSettleSeconds: 1}) resp, body := fc.do(http.MethodPut, "/api/v1/admin/config/orchestrator", putOrchestratorSettingsRequest{FIPSettleSeconds: 1})
if resp.StatusCode != http.StatusOK { if resp.StatusCode != http.StatusOK {
t.Fatalf("put settings: status=%d body=%s", resp.StatusCode, body) t.Fatalf("put settings: status=%d body=%s", resp.StatusCode, body)
} }
+8 -1
View File
@@ -81,10 +81,17 @@ func (s *Server) handleAdminRegistryHistory(w http.ResponseWriter, r *http.Reque
writeError(w, http.StatusInternalServerError, err.Error()) writeError(w, http.StatusInternalServerError, err.Error())
return return
} }
// Self-check failures over every run of the address.
failedOn, err := s.DB.ListSelfCheckFailedOn(r.Context(), summary.ID, 0)
if err != nil {
writeError(w, http.StatusInternalServerError, err.Error())
return
}
writeJSON(w, http.StatusOK, struct { writeJSON(w, http.StatusOK, struct {
Registry registryDTO `json:"registry"` Registry registryDTO `json:"registry"`
Checks []db.Check `json:"checks"` Checks []db.Check `json:"checks"`
}{registrySummaryToDTO(*summary), checks}) SelfCheckFailedOn []string `json:"self_check_failed_on"`
}{registrySummaryToDTO(*summary), checks, failedOn})
} }
func registrySummaryToDTO(s db.RegistrySummary) registryDTO { func registrySummaryToDTO(s db.RegistrySummary) registryDTO {
@@ -219,3 +219,51 @@ func TestAdminRegistryLevels(t *testing.T) {
t.Errorf("empty level must serialise by_type as [], got %s", body) t.Errorf("empty level must serialise by_type as [], got %s", body)
} }
} }
// TestSelfCheckFailedOnInAddressEndpoints proves GET /admin/ips/{ip} and GET
// /admin/registry/{ip} list the validators whose self-check of the address
// failed ([] when none).
func TestSelfCheckFailedOnInAddressEndpoints(t *testing.T) {
fc, d, orch, mock := newConfigTestHarness(t)
ctx := context.Background()
mock.Seed("fip-1", "9.9.9.9", "svc-project")
fc.do(http.MethodPost, "/api/v1/admin/config/validators", createValidatorRequest{ValidatorID: "validator-1", OSPortID: "port-1"})
fc.do(http.MethodPost, "/api/v1/agents/register", registerAgentRequest{ValidatorID: "validator-1"})
fc.do(http.MethodPost, "/api/v1/admin/ips", submitIPsRequest{Addresses: []string{"9.9.9.9"}})
failedOn := func(path string) []string {
t.Helper()
resp, body := fc.do(http.MethodGet, path, nil)
if resp.StatusCode != http.StatusOK {
t.Fatalf("get %s: status=%d body=%s", path, resp.StatusCode, body)
}
var got struct {
SelfCheckFailedOn []string `json:"self_check_failed_on"`
}
if err := json.Unmarshal(body, &got); err != nil {
t.Fatalf("unmarshal %s: %v", path, err)
}
if got.SelfCheckFailedOn == nil {
t.Fatalf("%s: self_check_failed_on must be [] rather than null, body=%s", path, body)
}
return got.SelfCheckFailedOn
}
if got := failedOn("/api/v1/admin/ips/9.9.9.9"); len(got) != 0 {
t.Fatalf("expected no failures yet, got %v", got)
}
orch.Tick(ctx) // claim + associate -> awaiting_self_check
ip, err := d.GetIPByAddress(ctx, "9.9.9.9")
if err != nil {
t.Fatal(err)
}
if err := orch.SelfCheckResult(ctx, "validator-1", ip.ID, false, "ip echo timeout"); err != nil {
t.Fatalf("self-check result: %v", err)
}
for _, path := range []string{"/api/v1/admin/ips/9.9.9.9", "/api/v1/admin/registry/9.9.9.9"} {
if got := failedOn(path); !reflect.DeepEqual(got, []string{"validator-1"}) {
t.Fatalf("%s: expected [validator-1], got %v", path, got)
}
}
}
+41 -12
View File
@@ -14,6 +14,7 @@ import (
"errors" "errors"
"fmt" "fmt"
"log/slog" "log/slog"
"strings"
"sync" "sync"
"time" "time"
@@ -243,9 +244,9 @@ func (o *Orchestrator) SelfCheckResult(ctx context.Context, validatorID string,
return nil return nil
} }
// Detach the floating IP before the address goes back to the queue. // Detach the floating IP before the address goes back to the queue.
// requeueOrFail frees the validator in the database but knows nothing // The database side (db.FailSelfCheck) frees the validator but knows
// about the cloud: a floating IP left on the validator's port makes // nothing about the cloud: a floating IP left on the validator's port
// every later association on that port fail with 409 ("fixed IP // makes every later association on that port fail with 409 ("fixed IP
// already has a floating IP"). Best-effort, like the other release // already has a floating IP"). Best-effort, like the other release
// paths — the database state must be freed even if Neutron hiccups. // paths — the database state must be freed even if Neutron hiccups.
if item.FIPID != "" { if item.FIPID != "" {
@@ -253,20 +254,48 @@ func (o *Orchestrator) SelfCheckResult(ctx context.Context, validatorID string,
o.Log.Error("disassociate fip after failed self-check", "ip_id", ipID, "fip_id", item.FIPID, "err", err) o.Log.Error("disassociate fip after failed self-check", "ip_id", ipID, "fip_id", item.FIPID, "err", err)
} }
} }
if item.RetryCount+1 > o.Cfg.MaxSelfCheckRetries { return o.failSelfCheck(ctx, item, validatorID, detail)
o.requeueOrFail(ctx, ipID, validatorID, "self-check failed: "+detail)
return nil
}
// Retry association without fully requeuing: re-drive the same
// claim by cycling back through requeue/claim keeps the logic in
// one place at the cost of the IP briefly returning to `queued`.
o.requeueOrFail(ctx, ipID, validatorID, "self-check failed, retrying: "+detail)
return nil
} }
return o.DB.SetChecking(ctx, ipID, o.leaseTTL()) return o.DB.SetChecking(ctx, ipID, o.leaseTTL())
} }
// failSelfCheck records a failed self-check and decides what happens to the
// address (see db.FailSelfCheck): the ceiling self_check_max_attempts is read
// from the settings on every failure, so a change applies to the next one. The
// retry_count / max_retries pair is not involved: a failed self-check has its
// own ceiling, and the validator that failed it is not handed this address
// again in the current round (db.ClaimNextQueued). The validator itself stays
// in service.
func (o *Orchestrator) failSelfCheck(ctx context.Context, item *db.IPQueueItem, validatorID, detail string) error {
settings, err := o.DB.GetSettings(ctx)
if err != nil {
return err
}
res, err := o.DB.FailSelfCheck(ctx, item.ID, validatorID, detail, settings.SelfCheckMaxAttempts)
if errors.Is(err, db.ErrInvalidState) {
o.Log.Warn("ignoring failed self-check for an address the validator does not hold",
"validator", validatorID, "ip_id", item.ID)
return nil
}
if err != nil {
return err
}
if res.Failed {
o.event(ctx, "control-api", "", &item.ID, "retry_or_fail",
fmt.Sprintf(`{"reason":%q}`, fmt.Sprintf("self-check failed %d times (on %s), giving up: %s",
res.Failures, strings.Join(res.Validators, ", "), detail)))
return nil
}
o.event(ctx, "control-api", "", &item.ID, "retry_or_fail",
fmt.Sprintf(`{"reason":%q}`, "self-check failed, retrying: "+detail))
if !res.NewRound {
o.event(ctx, "control-api", "", &item.ID, "validator_excluded",
fmt.Sprintf(`{"validator_id":%q,"failures":%d}`, validatorID, res.Failures))
}
return nil
}
// AssignmentForValidator returns the check config for a validator's current // AssignmentForValidator returns the check config for a validator's current
// IP if it's ready to be worked on (awaiting_self_check or checking), // IP if it's ready to be worked on (awaiting_self_check or checking),
// or nil if the validator has nothing to do right now. The check config is // or nil if the validator has nothing to do right now. The check config is
+166
View File
@@ -1162,3 +1162,169 @@ func TestLateFailedSelfCheckIsIgnored(t *testing.T) {
t.Fatalf("a late report changed the address state to %s", cur.State) t.Fatalf("a late report changed the address state to %s", cur.State)
} }
} }
// reportSelfChecks answers the self-check of every address that is waiting for
// one: a validator in failing reports a failure, any other a success.
func reportSelfChecks(t *testing.T, ctx context.Context, o *Orchestrator, d *db.DB, failing ...string) {
t.Helper()
fails := map[string]bool{}
for _, v := range failing {
fails[v] = true
}
validators, err := d.ListValidators(ctx)
if err != nil {
t.Fatal(err)
}
for _, v := range validators {
if v.CurrentIPID == nil {
continue
}
ip, err := d.GetIP(ctx, *v.CurrentIPID)
if err != nil || ip.State != db.IPAwaitingSelfCheck {
continue
}
if err := o.SelfCheckResult(ctx, v.ValidatorID, ip.ID, !fails[v.ValidatorID], "ip echo timeout"); err != nil {
t.Fatalf("self-check result of %s: %v", v.ValidatorID, err)
}
}
}
func registerValidators(t *testing.T, ctx context.Context, d *db.DB, ids ...string) {
t.Helper()
for i, id := range ids {
if err := d.RegisterValidator(ctx, id, "h", fmt.Sprintf("port-%d", i+1), "v"); err != nil {
t.Fatal(err)
}
}
}
func eventTypes(t *testing.T, ctx context.Context, d *db.DB, ipID int64) map[string]int {
t.Helper()
events, err := d.ListEventsForIP(ctx, ipID)
if err != nil {
t.Fatal(err)
}
out := map[string]int{}
for _, e := range events {
out[e.EventType]++
}
return out
}
// The incident: v1 fails the self-check of 1.1.1.1. The address must not go
// back to v1; another validator takes it and passes, while v1 keeps taking
// other addresses.
func TestSelfCheckFailureExcludesValidatorForThatAddress(t *testing.T) {
ctx := context.Background()
o, d, mock := newTestOrchestrator(t, 180)
for i, a := range []string{"1.1.1.1", "2.2.2.2", "3.3.3.3"} {
mock.Seed(fmt.Sprintf("fip-%d", i+1), a, "svc-project")
}
registerValidators(t, ctx, d, "v1", "v2", "v3")
if err := d.SeedQueue(ctx, []string{"1.1.1.1", "2.2.2.2"}); err != nil {
t.Fatal(err)
}
o.Tick(ctx) // v1 -> 1.1.1.1, v2 -> 2.2.2.2, v3 stays idle
reportSelfChecks(t, ctx, o, d, "v1")
a1, _ := d.GetIPByAddress(ctx, "1.1.1.1")
if a1.State != db.IPQueued || a1.RetryCount != 0 {
t.Fatalf("expected 1.1.1.1 back in the queue with retry_count 0, got %s/%d", a1.State, a1.RetryCount)
}
if v1, _ := d.GetValidator(ctx, "v1"); v1.State != db.ValidatorIdle {
t.Fatalf("v1 must stay in service, got %s", v1.State)
}
if _, err := d.SubmitIPs(ctx, []string{"3.3.3.3"}); err != nil {
t.Fatal(err)
}
o.Tick(ctx) // v1 skips 1.1.1.1 and takes 3.3.3.3; v3 takes 1.1.1.1
a1, _ = d.GetIPByAddress(ctx, "1.1.1.1")
a3, _ := d.GetIPByAddress(ctx, "3.3.3.3")
if a1.OwnerValidatorID == nil || *a1.OwnerValidatorID != "v3" {
t.Fatalf("expected 1.1.1.1 on v3, got %v", a1.OwnerValidatorID)
}
if a3.OwnerValidatorID == nil || *a3.OwnerValidatorID != "v1" {
t.Fatalf("expected 3.3.3.3 on v1, got %v", a3.OwnerValidatorID)
}
reportSelfChecks(t, ctx, o, d) // v1 passes this time: the others are fine
a1, _ = d.GetIPByAddress(ctx, "1.1.1.1")
a3, _ = d.GetIPByAddress(ctx, "3.3.3.3")
if a1.State != db.IPChecking || a3.State != db.IPChecking {
t.Fatalf("expected both addresses to pass the self-check, got %s and %s", a1.State, a3.State)
}
if got, _ := d.ListSelfCheckFailedOn(ctx, a1.RegistryID, 0); len(got) != 1 || got[0] != "v1" {
t.Fatalf("expected self-check failures on v1 only, got %v", got)
}
if ev := eventTypes(t, ctx, d, a1.ID); ev["validator_excluded"] != 1 {
t.Fatalf("expected one validator_excluded event, got %v", ev)
}
}
// The ceiling does not depend on the number of validators and is read at each
// failure: the address fails once it is reached, not earlier, even though the
// third validator never tried it.
func TestSelfCheckCeilingGivesFailAndFollowsSetting(t *testing.T) {
ctx := context.Background()
o, d, mock := newTestOrchestrator(t, 180)
mock.Seed("fip-1", "1.1.1.1", "svc-project")
registerValidators(t, ctx, d, "v1", "v2", "v3")
if err := d.SeedQueue(ctx, []string{"1.1.1.1"}); err != nil {
t.Fatal(err)
}
o.Tick(ctx)
reportSelfChecks(t, ctx, o, d, "v1", "v2", "v3")
if ip, _ := d.GetIPByAddress(ctx, "1.1.1.1"); ip.State != db.IPQueued {
t.Fatalf("expected queued after the first failure (ceiling 5), got %s", ip.State)
}
if err := d.SetSelfCheckMaxAttempts(ctx, 2); err != nil {
t.Fatal(err)
}
o.Tick(ctx)
reportSelfChecks(t, ctx, o, d, "v1", "v2", "v3")
ip, _ := d.GetIPByAddress(ctx, "1.1.1.1")
if ip.State != db.IPFailed || ip.OverallResult != db.ResultFail || ip.RetryCount != 0 {
t.Fatalf("expected failed/fail at the lowered ceiling 2, got %s/%s retry_count=%d", ip.State, ip.OverallResult, ip.RetryCount)
}
if got, _ := d.ListSelfCheckFailedOn(ctx, ip.RegistryID, 0); len(got) != 2 {
t.Fatalf("expected failures on two validators, got %v", got)
}
}
// With fewer working validators than the ceiling (down to a single one) the
// retries continue on the validators in a new round, up to the ceiling; the
// address never gets stuck in the queue.
func TestSelfCheckFewerValidatorsThanCeilingRetriesUntilCeiling(t *testing.T) {
for _, validators := range [][]string{{"v1"}, {"v1", "v2"}} {
ctx := context.Background()
o, d, mock := newTestOrchestrator(t, 180)
mock.Seed("fip-1", "1.1.1.1", "svc-project")
registerValidators(t, ctx, d, validators...)
if err := d.SeedQueue(ctx, []string{"1.1.1.1"}); err != nil {
t.Fatal(err)
}
if err := d.SetSelfCheckMaxAttempts(ctx, 5); err != nil {
t.Fatal(err)
}
for attempt := 1; attempt <= 5; attempt++ {
o.Tick(ctx)
ip, _ := d.GetIPByAddress(ctx, "1.1.1.1")
if ip.State != db.IPAwaitingSelfCheck {
t.Fatalf("%d validators, attempt %d: expected the address to be handed out, got %s", len(validators), attempt, ip.State)
}
reportSelfChecks(t, ctx, o, d, validators...)
ip, _ = d.GetIPByAddress(ctx, "1.1.1.1")
want := db.IPQueued
if attempt == 5 {
want = db.IPFailed
}
if ip.State != want {
t.Fatalf("%d validators, after failure %d: expected %s, got %s", len(validators), attempt, want, ip.State)
}
}
}
}