Skip the check cycle for a Floating IP already occupied by another port

The cloud is live: the address list submitted as "free" (bootstrap config
or POST /api/v1/admin/ips) can drift by the time the orchestrator claims
it, or an operator can queue an already-occupied address by mistake.
Neutron's floating-IP association is a blind "last write wins" PUT with
no conflict error to catch, so associateFIP now checks the FIP's PortID
(already fetched via GetFloatingIPByAddress) before associating, guarded
against the false-positive of the FIP already belonging to this same
validator's own port.

A match routes the address straight to a new terminal ip_queue.state
("occupied", distinct from failed/fail) via db.MarkFIPOccupied — no
retries, since Neutron won't free it on its own and requeuing would let
it be reclaimed again next tick, starving the rest of the queue — plus a
dedicated fip_occupied audit event. Resubmitting the address later (once
the conflict is resolved) resets it to queued via the existing
POST /api/v1/admin/ips resubmit path (CancelIP/ListExpiredLeases updated
to treat occupied as terminal too). admin-dashboard gets its own "занят"
badge, distinct from fail/partial/cancelled.

Rebuilt bin/{control-api,admin-dashboard,prober,validator-agent} and
bin/SHA256SUMS per docs/SETUP.md's documented build recipe, since
control-api and admin-dashboard source changed.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NeVbMVEiE7XQAkBd7HQgj6
This commit is contained in:
ayurishchevandClaude Sonnet 5 committed 2026-09-13 23:54:35 +03:00
1 parent 2246369b64
commit 93c79b63ea
17 files changed
+326 -17

No files matched your search

+17 -2
View File
@@ -598,6 +598,21 @@ queued ──(control-api сам, без вызова API)──▶ assigning_fi
`orchestrator.poll_interval_seconds`), явного HTTP-метода для их запуска
нет — это фоновый цикл (`Tick`), а не запрос/ответ.
Из `assigning_fip` есть и второй, терминальный исход: если на момент
попытки ассоциации Floating IP уже привязан к чужому порту (облако живое —
список адресов мог разойтись с реальностью с момента постановки в
очередь, либо адрес был ошибочно передан занятым), control-api переводит
адрес в состояние `occupied` вместо продолжения в `awaiting_self_check` —
цикл проверки для этой попытки не запускается вовсе. Это отдельное
терминальное состояние, а не `failed`: `failed` означает «проверка
стартовала и не прошла», `occupied` — «проверка не стартовала, потому что
адрес занят кем-то другим». В аудит-логе адреса (`events`) фиксируется
строка `fip_occupied`. `overall_result` для этого состояния остаётся
пустым. Как и `done`/`failed`, `occupied` сбрасывается обратно в `queued`
повторной постановкой через `POST /api/v1/admin/ips` — этим способом
оператор возвращает адрес в работу, убедившись, что конфликт в облаке
разрешился.
Вход в `awaiting_self_check` не означает мгновенную видимость агенту: если
настроена пауза (`fip_settle_seconds`, см.
[«Настройки оркестратора»](#настройки-оркестратора-apiv1adminconfigorchestrator)
@@ -616,8 +631,8 @@ queued ──(control-api сам, без вызова API)──▶ assigning_fi
`/api/v1/admin/ips*`, а не самим оркестратором:
- **любое нетерминальное состояние → `failed` (`overall_result:
"cancelled"`)** — `POST /api/v1/admin/ips/{ip}/cancel`;
- **`done`/`failed` → `queued` (новая попытка)** — `POST
/api/v1/admin/ips` с уже завершённым адресом в списке;
- **`done`/`failed`/`occupied` → `queued` (новая попытка)** — `POST
/api/v1/admin/ips` с уже завершённым (или занятым) адресом в списке;
- **любое состояние → адрес физически исчезает из очереди**, вместе со
всей историей — `DELETE /api/v1/admin/ips/{ip}`, `POST
/api/v1/admin/ips/delete`, `POST /api/v1/admin/ips/clear` (см.