Skip the check cycle for a Floating IP already occupied by another port

The cloud is live: the address list submitted as "free" (bootstrap config
or POST /api/v1/admin/ips) can drift by the time the orchestrator claims
it, or an operator can queue an already-occupied address by mistake.
Neutron's floating-IP association is a blind "last write wins" PUT with
no conflict error to catch, so associateFIP now checks the FIP's PortID
(already fetched via GetFloatingIPByAddress) before associating, guarded
against the false-positive of the FIP already belonging to this same
validator's own port.

A match routes the address straight to a new terminal ip_queue.state
("occupied", distinct from failed/fail) via db.MarkFIPOccupied — no
retries, since Neutron won't free it on its own and requeuing would let
it be reclaimed again next tick, starving the rest of the queue — plus a
dedicated fip_occupied audit event. Resubmitting the address later (once
the conflict is resolved) resets it to queued via the existing
POST /api/v1/admin/ips resubmit path (CancelIP/ListExpiredLeases updated
to treat occupied as terminal too). admin-dashboard gets its own "занят"
badge, distinct from fail/partial/cancelled.

Rebuilt bin/{control-api,admin-dashboard,prober,validator-agent} and
bin/SHA256SUMS per docs/SETUP.md's documented build recipe, since
control-api and admin-dashboard source changed.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NeVbMVEiE7XQAkBd7HQgj6
This commit is contained in:
ayurishchevandClaude Sonnet 5 committed 2026-09-13 23:54:35 +03:00
1 parent 2246369b64
commit 93c79b63ea
17 files changed
+326 -17

No files matched your search

+16 -1
View File
@@ -117,7 +117,7 @@ curl -s http://<control-api>:8080/api/v1/admin/ips \
|---|---|
| `IPAddress` | Проверяемый адрес |
| `Sequence` | Позиция в очереди (порядок постановки — из конфига при первом старте либо из последнего вызова `POST /api/v1/admin/ips`) |
| `State` | Текущий этап: `queued`, `assigning_fip`, `awaiting_self_check`, `checking`, `aggregating`, `done`, `failed` |
| `State` | Текущий этап: `queued`, `assigning_fip`, `awaiting_self_check`, `checking`, `aggregating`, `done`, `failed`, `occupied` (см. ниже) |
| `OwnerValidatorID` | Какой валидатор сейчас (или последним) занимался этим адресом |
| `FIPID` | Идентификатор Floating IP в OpenStack, к которому привязан адрес (пусто, если ещё/уже не привязан) |
| `AttemptNumber` | Номер попытки — растёт при каждом requeue (сбой привязки, сбой self-check, реклейм по таймауту) |
@@ -154,6 +154,21 @@ curl -s http://<control-api>:8080/api/v1/admin/ips \
система по итогам проверок. Отличать от обычного `fail` полезно, чтобы
не путать «адрес не прошёл проверку» с «проверку прервали вручную».
Отдельно от `OverallResult` стоит состояние **`State: "occupied"`** —
облако живое, и список адресов, переданный как «свободные» (из конфига
или через `POST /api/v1/admin/ips`), мог с тех пор разойтись с
реальностью, либо адрес мог быть передан на проверку по ошибке уже
занятым. Если при попытке привязки Floating IP control-api видит, что тот
уже привязан к чужому порту, адрес переводится в `occupied` **до начала**
цикла проверки — `OverallResult` при этом остаётся пустым, это не `fail`:
`fail` означает «проверка стартовала и не прошла», `occupied` — «проверка
не стартовала, адрес занят кем-то другим». В `events` по адресу
появляется строка `fip_occupied`. Автоматических повторных попыток нет
(Neutron сам не освобождает адрес) — верните адрес в работу вручную через
`POST /api/v1/admin/ips`, когда убедитесь, что конфликт в облаке
разрешился; в дашборде для таких адресов также показывается кнопка
«Перепроверить» вместо «Отменить».
Отсутствие ответа от источника (площадка не прислала результат до
истечения `checking_window_seconds`) засчитывается как провал — это
управляется настройкой `aggregation.missing_counts_as_fail` в конфиге