16 Commits
Author SHA1 Message Date
ayurishchevandClaude Sonnet 5.5 0532baff09 Keep one address per validator; fix heartbeat handling and queue clear
A mass check on 2026-10-02 stalled 7 of 20 validators and sent 42
addresses to fail without a single check. A validator busy with slow
checks went silent, was marked unreachable, and its next heartbeat put it
back to idle while it still held the address; it was handed a second one,
whose association never ran (the in-flight guard was keyed by validator),
and both waited for their leases to expire.

- Heartbeat/re-register return an unreachable validator to assigned when
  it still holds an address, else idle.
- A validator is released only from the address it currently holds
  (ReleaseFIP, RequeueOrFail, MarkFIPOccupied, FreeValidator); an
  unreachable validator stays unreachable until its next heartbeat, so a
  dead validator is no longer handed a new address every lease period.
- ClaimNextQueued refuses a validator that still has an address; a
  ReconcileValidators pass on every tick repairs rows that disagree with
  the queue.
- Association guard is keyed by address, not validator.
- The agent sends heartbeats from their own goroutine.
- Clear queue / delete: detach only floating IPs of unfinished rows (done,
  failed and occupied rows kept their fip_id and made a clear issue >1000
  sequential cloud calls: 256 s), at most 8 in parallel; the operation no
  longer dies with the client connection (10 minute limit).

Includes the incident analysis and the plan under analysis/ and
docs/changes/, and rebuilt bin/control-api and bin/validator-agent.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-10-02 14:42:12 +03:00
ayurishchevandClaude Sonnet 5.5 146259cabb Detach the floating IP when a self-check fails
A failed self-check returned the address to the queue and freed the
validator in the database, but left the floating IP attached to the
validator's port. Every later association on that port then failed with
409 ("fixed IP already has a floating IP"), so one failed self-check
poisoned a validator for good; on 2026-10-01 all 20 validators were
poisoned within 23 minutes after ifconfig.me timeouts.

SelfCheckResult now disassociates the floating IP before requeueing, and
ignores a late failed report for an address the validator no longer
holds (it could belong to another validator by then).

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-10-02 03:24:19 +03:00
ayurishchevandClaude Sonnet 5.5 1ee5757004 Free validator ports in the cloud on clear/cancel/delete
An address still being associated (assigning_fip) has no fip_id in the
database, so "clear queue" did not detach it, and the association then
finished after the row was gone, leaving the floating IP on the validator
port for good.

- After clear/cancel/delete, ask the cloud which floating IPs sit on the
  affected validator ports (new ListFloatingIPsByPort) and detach those
  that this system queued (known in ip_registry); foreign ones are left.
- SetFIPAssociated applies only to a row still in assigning_fip; if the
  address was removed meanwhile, associateFIP detaches the floating IP.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-10-01 21:00:31 +03:00
ayurishchevandClaude Sonnet 5.5 c0e300f71c Run per-validator OpenStack work in parallel so all validators start at once
The orchestrator claimed idle validators one by one and associated each
floating IP synchronously (~30 s per address), so 20 validators started
about 30 s apart. Aggregation/disassociation and lease reclaim were
sequential in the same way.

The slow OpenStack calls now run in one goroutine per address, guarded by
an in-flight set against duplicates. control-api runs Tick in Async mode
(Tick does not wait); tests keep the waiting mode.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-10-01 20:37:32 +03:00
ayurishchevandClaude Sonnet 5.5 aff8fe38b5 Scan floating IPs in the background, page by page, so thousands of addresses work
The "Scan Floating IP" button failed with a client timeout: the project now
holds ~6.4k floating IPs and the scan listed them all in one unpaginated,
timeout-less Neutron request on the HTTP request context.

openstack: ListFreeFloatingIPs reads marker-based pages (fields= keeps them
small) with per-page retry/backoff on transport errors, 5xx and 429, and every
request now has a timeout (also ends hangs inside the orchestrator tick).

orchestrator: the scan is a single-flight background job on the process
context with progress (clearing/listing/enqueuing/done/error), dry_run, full
discovery before anything is enqueued, then SubmitIPs in chunks of 500 in
ascending IP order; a failed read leaves the queue untouched. The auto-cycle
gets a "scanning" phase that polls the job, so the control loop and
autoCycleMu are never held across OpenStack/DB work; it recovers after a
restart and waits for (instead of adopting) a scan started by someone else.

db: migration 0009 (indexes), paged ListIPsPage/ListRegistryPage, GROUP BY
counters, EXISTS completion check, set-based ClearAllIPs.

API: POST /admin/ips/scan -> 202 (dry_run, wait), GET /admin/ips/scan, paging
and filters on /admin/ips and /admin/registry (bare arrays without limit),
results_by_overall in /admin/status.

dashboard: scan progress panel and dry-run button, paginated /ips and
/registry with server-side filters, Overview on counters and capped lists
with progress/ETA, "select all N by filter", hx-params fix for per-row
buttons, real counts in confirmations.

Also: docs (API, USAGE, DASHBOARD, README), plan and review under
docs/changes/, bin/ rebuilt with new SHA256SUMS.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-10-01 19:31:11 +03:00
ayurishchevandClaude Sonnet 5.5 cd37b10f3b Add optional automatic check cycle (clear queue -> scan FIPs -> wait -> repeat)
An admin-controlled scenario that repeats what the operator does by hand:
clear the IP queue, scan and enqueue all free Floating IPs, wait until every
queued address reaches a terminal state (so results are in the Registry),
then wait a configurable interval and start over.

- control-api: new auto_cycle singleton table (migration 0008) holding
  enabled/interval/max-run settings and persisted phase state, so the cycle
  survives restarts; engine in orchestrator/autocycle.go driven from the
  existing loop tick with an injectable "now" for deterministic tests.
- Interval (default 1h, min 60s) and max wait (default unlimited, timeout
  outcome) are runtime settings, never hardcoded.
- The periodic fip_scan_interval_seconds scan is skipped while the cycle is
  enabled. An emptied queue mid-cycle counts as finished; stopping during
  the pause keeps the last cycle's outcome.
- API: GET/PUT /api/v1/admin/auto-cycle, POST .../start, POST .../stop.
- admin-dashboard: "Автоматический цикл" panel on /settings and an
  "Автоцикл активен" indicator on /overview.
- Tests for db, orchestrator, httpapi and dashboard; run-local-e2e.sh now
  exercises a full auto cycle; docs updated; bin/ rebuilt with refreshed
  SHA256SUMS.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-10-01 10:28:53 +03:00
ayurishchevandClaude Sonnet 5 582b44f314 Add Floating IP scanning and a durable address registry with configurable history depth
Adds POST /api/v1/admin/ips/scan (plus an optional periodic ticker) to
discover free Floating IPs in the OpenStack project and feed them straight
into the check queue. More importantly, decouples check/event history from
ip_queue's lifecycle: a new ip_registry table (migration 0007) gives every
address ever submitted a durable identity, so deleting it from the queue no
longer destroys its history — it's still reachable via the new
GET /api/v1/admin/registry[/{ip}] endpoints and the dashboard's /registry
pages, with retention depth configurable in check cycles per address
(history_retention_cycles, 0 = unlimited).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-23 09:52:01 +03:00
ayurishchevandClaude Sonnet 5 93c79b63ea Skip the check cycle for a Floating IP already occupied by another port
The cloud is live: the address list submitted as "free" (bootstrap config
or POST /api/v1/admin/ips) can drift by the time the orchestrator claims
it, or an operator can queue an already-occupied address by mistake.
Neutron's floating-IP association is a blind "last write wins" PUT with
no conflict error to catch, so associateFIP now checks the FIP's PortID
(already fetched via GetFloatingIPByAddress) before associating, guarded
against the false-positive of the FIP already belonging to this same
validator's own port.

A match routes the address straight to a new terminal ip_queue.state
("occupied", distinct from failed/fail) via db.MarkFIPOccupied — no
retries, since Neutron won't free it on its own and requeuing would let
it be reclaimed again next tick, starving the rest of the queue — plus a
dedicated fip_occupied audit event. Resubmitting the address later (once
the conflict is resolved) resets it to queued via the existing
POST /api/v1/admin/ips resubmit path (CancelIP/ListExpiredLeases updated
to treat occupied as terminal too). admin-dashboard gets its own "занят"
badge, distinct from fail/partial/cancelled.

Rebuilt bin/{control-api,admin-dashboard,prober,validator-agent} and
bin/SHA256SUMS per docs/SETUP.md's documented build recipe, since
control-api and admin-dashboard source changed.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NeVbMVEiE7XQAkBd7HQgj6
2026-09-13 23:54:35 +03:00
ayurishchev 550ce3fec4 Enable SSH and TLS handshake on prober for TCP22 and TCP443 ports 2026-08-26 23:52:31 +03:00
ayurishchev ef24cc9858 New external site management. New external prober heartbeat feature. 2026-08-26 20:47:54 +03:00
ayurishchev 42f584dd3c Manage external site-prober actions va API and Dashboard 2026-08-26 19:45:48 +03:00
ayurishchev f22ad569b6 added FIP Cooldown before start 2026-08-24 10:29:08 +03:00
ayurishchev f8336740ad feature: handle delete operation for IPs 2026-08-23 22:24:55 +03:00
ayurishchev 37910e410b admin control features and admin dashboard 2026-08-23 20:39:22 +03:00
ayurishchev c630f13c57 make inbound check optional 2026-08-21 11:25:52 +03:00
ayurishchev 7e44db87b2 repo init 2026-08-21 07:34:45 +03:00