Skip the check cycle for a Floating IP already occupied by another port
The cloud is live: the address list submitted as "free" (bootstrap config
or POST /api/v1/admin/ips) can drift by the time the orchestrator claims
it, or an operator can queue an already-occupied address by mistake.
Neutron's floating-IP association is a blind "last write wins" PUT with
no conflict error to catch, so associateFIP now checks the FIP's PortID
(already fetched via GetFloatingIPByAddress) before associating, guarded
against the false-positive of the FIP already belonging to this same
validator's own port.
A match routes the address straight to a new terminal ip_queue.state
("occupied", distinct from failed/fail) via db.MarkFIPOccupied — no
retries, since Neutron won't free it on its own and requeuing would let
it be reclaimed again next tick, starving the rest of the queue — plus a
dedicated fip_occupied audit event. Resubmitting the address later (once
the conflict is resolved) resets it to queued via the existing
POST /api/v1/admin/ips resubmit path (CancelIP/ListExpiredLeases updated
to treat occupied as terminal too). admin-dashboard gets its own "занят"
badge, distinct from fail/partial/cancelled.
Rebuilt bin/{control-api,admin-dashboard,prober,validator-agent} and
bin/SHA256SUMS per docs/SETUP.md's documented build recipe, since
control-api and admin-dashboard source changed.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NeVbMVEiE7XQAkBd7HQgj6
This commit is contained in:
1 parent
2246369b64
commit
93c79b63ea
17 files changed
+326
-17
No files matched your search
+17
-2
@@ -598,6 +598,21 @@ queued ──(control-api сам, без вызова API)──▶ assigning_fi
|
||||
`orchestrator.poll_interval_seconds`), явного HTTP-метода для их запуска
|
||||
нет — это фоновый цикл (`Tick`), а не запрос/ответ.
|
||||
|
||||
Из `assigning_fip` есть и второй, терминальный исход: если на момент
|
||||
попытки ассоциации Floating IP уже привязан к чужому порту (облако живое —
|
||||
список адресов мог разойтись с реальностью с момента постановки в
|
||||
очередь, либо адрес был ошибочно передан занятым), control-api переводит
|
||||
адрес в состояние `occupied` вместо продолжения в `awaiting_self_check` —
|
||||
цикл проверки для этой попытки не запускается вовсе. Это отдельное
|
||||
терминальное состояние, а не `failed`: `failed` означает «проверка
|
||||
стартовала и не прошла», `occupied` — «проверка не стартовала, потому что
|
||||
адрес занят кем-то другим». В аудит-логе адреса (`events`) фиксируется
|
||||
строка `fip_occupied`. `overall_result` для этого состояния остаётся
|
||||
пустым. Как и `done`/`failed`, `occupied` сбрасывается обратно в `queued`
|
||||
повторной постановкой через `POST /api/v1/admin/ips` — этим способом
|
||||
оператор возвращает адрес в работу, убедившись, что конфликт в облаке
|
||||
разрешился.
|
||||
|
||||
Вход в `awaiting_self_check` не означает мгновенную видимость агенту: если
|
||||
настроена пауза (`fip_settle_seconds`, см.
|
||||
[«Настройки оркестратора»](#настройки-оркестратора-apiv1adminconfigorchestrator)
|
||||
@@ -616,8 +631,8 @@ queued ──(control-api сам, без вызова API)──▶ assigning_fi
|
||||
`/api/v1/admin/ips*`, а не самим оркестратором:
|
||||
- **любое нетерминальное состояние → `failed` (`overall_result:
|
||||
"cancelled"`)** — `POST /api/v1/admin/ips/{ip}/cancel`;
|
||||
- **`done`/`failed` → `queued` (новая попытка)** — `POST
|
||||
/api/v1/admin/ips` с уже завершённым адресом в списке;
|
||||
- **`done`/`failed`/`occupied` → `queued` (новая попытка)** — `POST
|
||||
/api/v1/admin/ips` с уже завершённым (или занятым) адресом в списке;
|
||||
- **любое состояние → адрес физически исчезает из очереди**, вместе со
|
||||
всей историей — `DELETE /api/v1/admin/ips/{ip}`, `POST
|
||||
/api/v1/admin/ips/delete`, `POST /api/v1/admin/ips/clear` (см.
|
||||
|
||||
Reference in new issue
Block a user