Keep one address per validator; fix heartbeat handling and queue clear
A mass check on 2026-10-02 stalled 7 of 20 validators and sent 42 addresses to fail without a single check. A validator busy with slow checks went silent, was marked unreachable, and its next heartbeat put it back to idle while it still held the address; it was handed a second one, whose association never ran (the in-flight guard was keyed by validator), and both waited for their leases to expire. - Heartbeat/re-register return an unreachable validator to assigned when it still holds an address, else idle. - A validator is released only from the address it currently holds (ReleaseFIP, RequeueOrFail, MarkFIPOccupied, FreeValidator); an unreachable validator stays unreachable until its next heartbeat, so a dead validator is no longer handed a new address every lease period. - ClaimNextQueued refuses a validator that still has an address; a ReconcileValidators pass on every tick repairs rows that disagree with the queue. - Association guard is keyed by address, not validator. - The agent sends heartbeats from their own goroutine. - Clear queue / delete: detach only floating IPs of unfinished rows (done, failed and occupied rows kept their fip_id and made a clear issue >1000 sequential cloud calls: 256 s), at most 8 in parallel; the operation no longer dies with the client connection (10 minute limit). Includes the incident analysis and the plan under analysis/ and docs/changes/, and rebuilt bin/control-api and bin/validator-agent. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
This commit is contained in:
1 parent
cf4a883363
commit
0532baff09
17 files changed
+1293
-62
No files matched your search
@@ -408,6 +408,17 @@ curl -s http://<control-api>:8080/api/v1/admin/validators | python3 -m json.tool
|
||||
(занят), `unreachable` (пропустил heartbeat дольше
|
||||
`orchestrator.heartbeat_timeout_seconds`).
|
||||
|
||||
Правила, которые держат состояние валидатора согласованным:
|
||||
- валидатор держит **не более одного адреса**; адрес освобождает валидатор
|
||||
только пока он остаётся его текущим (запоздалое завершение старого адреса
|
||||
чужого валидатора не освобождает);
|
||||
- после пропущенного heartbeat валидатор, у которого есть адрес, возвращается
|
||||
в `assigned`, а не в `idle`, и не получает второй адрес; без адреса — в `idle`;
|
||||
- `unreachable`-валидатор не получает адресов, пока не пришлёт heartbeat
|
||||
(даже если лизинг его адреса истёк и адрес вернулся в очередь);
|
||||
- на каждом такте оркестратор сверяет валидаторы с очередью и исправляет
|
||||
расхождения (в логе `repaired validators that disagreed with the queue`).
|
||||
|
||||
**Добавление нового валидатора (без перезапуска control-api):**
|
||||
1. Поднимите новую ВМ в сервисном проекте облака, узнайте её Neutron
|
||||
`port_id`.
|
||||
@@ -690,6 +701,12 @@ curl -s -X POST http://<control-api>:8080/api/v1/admin/ips/delete \
|
||||
curl -s -X POST http://<control-api>:8080/api/v1/admin/ips/clear
|
||||
```
|
||||
|
||||
Очистка отвязывает Floating IP только у адресов, которые ещё в работе
|
||||
(завершённые уже свободны), затем сама опрашивает порты валидаторов в облаке и
|
||||
снимает оставшиеся привязки адресов из реестра. Занимает секунды. Она не
|
||||
прерывается разрывом соединения (таймаутом клиента или дашборда): операция
|
||||
доводится до конца на стороне control-api, предел — 10 минут.
|
||||
|
||||
В `admin-dashboard` то же самое доступно на странице `/ips`: чекбоксы у
|
||||
каждой строки + кнопка «Удалить выбранные» для точечного/массового
|
||||
удаления, кнопка «Удалить» в каждой строке, и отдельная кнопка «Очистить
|
||||
|
||||
Reference in new issue
Block a user