A mass check on 2026-10-02 stalled 7 of 20 validators and sent 42
addresses to fail without a single check. A validator busy with slow
checks went silent, was marked unreachable, and its next heartbeat put it
back to idle while it still held the address; it was handed a second one,
whose association never ran (the in-flight guard was keyed by validator),
and both waited for their leases to expire.
- Heartbeat/re-register return an unreachable validator to assigned when
it still holds an address, else idle.
- A validator is released only from the address it currently holds
(ReleaseFIP, RequeueOrFail, MarkFIPOccupied, FreeValidator); an
unreachable validator stays unreachable until its next heartbeat, so a
dead validator is no longer handed a new address every lease period.
- ClaimNextQueued refuses a validator that still has an address; a
ReconcileValidators pass on every tick repairs rows that disagree with
the queue.
- Association guard is keyed by address, not validator.
- The agent sends heartbeats from their own goroutine.
- Clear queue / delete: detach only floating IPs of unfinished rows (done,
failed and occupied rows kept their fip_id and made a clear issue >1000
sequential cloud calls: 256 s), at most 8 in parallel; the operation no
longer dies with the client connection (10 minute limit).
Includes the incident analysis and the plan under analysis/ and
docs/changes/, and rebuilt bin/control-api and bin/validator-agent.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>