control-api is hosted outside the cloud and validators reach it directly,
so it sees the floating IP as the connection's source address. New open
route GET /api/v1/agents/{id}/observed-ip returns that address (taken only
from the TCP peer; forwarding headers are ignored so a validator cannot
forge it).
The agent gets self_check.methods, a priority-ordered list of ip_echo
(unchanged) and control_api; the default stays [ip_echo]. The self-check
passes when any method confirms the address; the next method is tried on
no answer and on a mismatch. Each method has its own timeout so a hung
first method cannot starve the fallback, and control_api uses a new TCP
connection per call (a connection opened before the floating IP was
attached would keep reporting the old address).
Also: docker agent template/env, example config, docs, plan in
docs/changes, e2e script switch E2E_SELF_CHECK_METHODS, rebuilt
bin/control-api and bin/validator-agent.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
An address still being associated (assigning_fip) has no fip_id in the
database, so "clear queue" did not detach it, and the association then
finished after the row was gone, leaving the floating IP on the validator
port for good.
- After clear/cancel/delete, ask the cloud which floating IPs sit on the
affected validator ports (new ListFloatingIPsByPort) and detach those
that this system queued (known in ip_registry); foreign ones are left.
- SetFIPAssociated applies only to a row still in assigning_fip; if the
address was removed meanwhile, associateFIP detaches the floating IP.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
The orchestrator claimed idle validators one by one and associated each
floating IP synchronously (~30 s per address), so 20 validators started
about 30 s apart. Aggregation/disassociation and lease reclaim were
sequential in the same way.
The slow OpenStack calls now run in one goroutine per address, guarded by
an in-flight set against duplicates. control-api runs Tick in Async mode
(Tick does not wait); tests keep the waiting mode.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
The "Scan Floating IP" button failed with a client timeout: the project now
holds ~6.4k floating IPs and the scan listed them all in one unpaginated,
timeout-less Neutron request on the HTTP request context.
openstack: ListFreeFloatingIPs reads marker-based pages (fields= keeps them
small) with per-page retry/backoff on transport errors, 5xx and 429, and every
request now has a timeout (also ends hangs inside the orchestrator tick).
orchestrator: the scan is a single-flight background job on the process
context with progress (clearing/listing/enqueuing/done/error), dry_run, full
discovery before anything is enqueued, then SubmitIPs in chunks of 500 in
ascending IP order; a failed read leaves the queue untouched. The auto-cycle
gets a "scanning" phase that polls the job, so the control loop and
autoCycleMu are never held across OpenStack/DB work; it recovers after a
restart and waits for (instead of adopting) a scan started by someone else.
db: migration 0009 (indexes), paged ListIPsPage/ListRegistryPage, GROUP BY
counters, EXISTS completion check, set-based ClearAllIPs.
API: POST /admin/ips/scan -> 202 (dry_run, wait), GET /admin/ips/scan, paging
and filters on /admin/ips and /admin/registry (bare arrays without limit),
results_by_overall in /admin/status.
dashboard: scan progress panel and dry-run button, paginated /ips and
/registry with server-side filters, Overview on counters and capped lists
with progress/ETA, "select all N by filter", hx-params fix for per-row
buttons, real counts in confirmations.
Also: docs (API, USAGE, DASHBOARD, README), plan and review under
docs/changes/, bin/ rebuilt with new SHA256SUMS.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
control-api: every route now carries a mandatory access level (admin / agent /
open) in a route table. All /api/v1/admin/* require the admin token; the
write calls of validator-agent and prober (self-check, events, results,
complete) require a separate static agent token; register, heartbeat and
fetching the assignment stay open. Tokens come from env vars, are compared in
constant time and never logged. An empty token leaves that level open with a
startup warning (backward compatible).
validator-agent / prober: apiclient sends the agent token only to control-api.
admin-dashboard: login/password (from env) with a stateless HMAC session
cookie, Origin-based CSRF check, per-IP brute-force throttle, HX-Redirect for
htmx polls, logout in the sidebar; the dashboard calls control-api with the
admin token. Login page layout fixed after review.
Also: env plumbing in docker-compose/rxprod-compose/systemd/config examples,
e2e script with token assertions, tests, docs (API, SETUP, USAGE, DASHBOARD,
README), plan and review under docs/changes/, bin/ rebuilt with new
SHA256SUMS.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
An admin-controlled scenario that repeats what the operator does by hand:
clear the IP queue, scan and enqueue all free Floating IPs, wait until every
queued address reaches a terminal state (so results are in the Registry),
then wait a configurable interval and start over.
- control-api: new auto_cycle singleton table (migration 0008) holding
enabled/interval/max-run settings and persisted phase state, so the cycle
survives restarts; engine in orchestrator/autocycle.go driven from the
existing loop tick with an injectable "now" for deterministic tests.
- Interval (default 1h, min 60s) and max wait (default unlimited, timeout
outcome) are runtime settings, never hardcoded.
- The periodic fip_scan_interval_seconds scan is skipped while the cycle is
enabled. An emptied queue mid-cycle counts as finished; stopping during
the pause keeps the last cycle's outcome.
- API: GET/PUT /api/v1/admin/auto-cycle, POST .../start, POST .../stop.
- admin-dashboard: "Автоматический цикл" panel on /settings and an
"Автоцикл активен" indicator on /overview.
- Tests for db, orchestrator, httpapi and dashboard; run-local-e2e.sh now
exercises a full auto cycle; docs updated; bin/ rebuilt with refreshed
SHA256SUMS.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Placing the filter form outside the polled div (so the poll can't wipe out
typed/selected values) put it before the stat-grid, since both used to
live inside that same polled block — stats ended up after the filter
instead of before it, as it always was.
Splits the stat-grid out into its own #overview-stats div, positioned
before the filter form; the actual poll target is now #overview-tables
(just the two tables). Since #overview-stats no longer polls directly,
/overview/fragment now also renders it as an out-of-band swap alongside
the main #overview-tables response — the same hx-swap-oob idiom already
used for the shared error banner — so the stat counts still refresh every
tick even though they're outside the polled element.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Both pages rendered their full lists with no way to narrow them. Adds a
?q=&status= filter (substring match on address, exact match on result
status) applied dashboard-side, in Go, over the already-fetched list —
no control-api/db changes needed.
Overview: the filter form lives outside the polling target (#overview-live)
so the recurring poll never wipes out what's typed/selected; the poll and
both filter inputs share hx-sync="#overview-live:queue last" (the same
fix that resolved the earlier abandoned /ips auto-refresh races) and the
poll now carries hx-include="#overview-filter" so it keeps honoring the
current filter on every tick. Applies uniformly to both the "Текущая
проверка" and "Последние N завершённых" tables, per the confirmed design:
picking a specific status naturally hides in-progress rows, since they
have no result yet.
Registry: no polling exists there, so the filter form reuses the full page
via hx-select/hx-replace-url — simpler than adding a parallel fragment
endpoint, and gives a bookmarkable/shareable filtered URL.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
fillRegistrySummary derived LastResult from whichever single checks row
happened to have the latest checked_at, not the cycle's actual aggregated
result — a cycle with a mix of passing and failing checks (e.g. one egress
target timed out while the rest, including the chronologically-last check,
succeeded) rendered as a green "pass" badge on /registry, disagreeing with
the correct "partial" badge already shown on /ips for the same address.
Now prefers ip_queue.overall_result (the orchestrator's own aggregation)
when a live queue row has a finished cycle, leaves the badge blank while a
cycle is still in progress, and only falls back to classifying the most
recent cycle's own checks (pass/fail/partial) once the address has been
deleted from the queue and overall_result is no longer available.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Adds POST /api/v1/admin/ips/scan (plus an optional periodic ticker) to
discover free Floating IPs in the OpenStack project and feed them straight
into the check queue. More importantly, decouples check/event history from
ip_queue's lifecycle: a new ip_registry table (migration 0007) gives every
address ever submitted a durable identity, so deleting it from the queue no
longer destroys its history — it's still reachable via the new
GET /api/v1/admin/registry[/{ip}] endpoints and the dashboard's /registry
pages, with retention depth configurable in check cycles per address
(history_retention_cycles, 0 = unlimited).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Both binaries registered with control-api exactly once at startup and
exited (os.Exit(1)) on any failure — including control-api simply not
being up yet (no ordering guarantee between the two at boot/redeploy) or
the admin not having added this validator_id/site_id to the config yet.
Run() now retries registration with capped exponential backoff (3s->30s)
until it succeeds or the process is asked to shut down, instead of
crashing; registerWithRetry is identical in agentcore and probercore
since their Run/register shape already was.
Separately, the admin dashboard's Validators page had no hostname column
even though the agent already reports one on register (mirroring the
prober) and control-api already persists it — only the admin-config read
DTO (validatorDTO in httpapi and dashboard) dropped it before it reached
the template. Added hostname + last_heartbeat_at to that DTO end-to-end
and a Хост/Heartbeat column to validators.html, matching sites.html.
Rebuilt bin/{control-api,admin-dashboard,prober,validator-agent} and
bin/SHA256SUMS per docs/SETUP.md's documented build recipe.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
The cloud is live: the address list submitted as "free" (bootstrap config
or POST /api/v1/admin/ips) can drift by the time the orchestrator claims
it, or an operator can queue an already-occupied address by mistake.
Neutron's floating-IP association is a blind "last write wins" PUT with
no conflict error to catch, so associateFIP now checks the FIP's PortID
(already fetched via GetFloatingIPByAddress) before associating, guarded
against the false-positive of the FIP already belonging to this same
validator's own port.
A match routes the address straight to a new terminal ip_queue.state
("occupied", distinct from failed/fail) via db.MarkFIPOccupied — no
retries, since Neutron won't free it on its own and requeuing would let
it be reclaimed again next tick, starving the rest of the queue — plus a
dedicated fip_occupied audit event. Resubmitting the address later (once
the conflict is resolved) resets it to queued via the existing
POST /api/v1/admin/ips resubmit path (CancelIP/ListExpiredLeases updated
to treat occupied as terminal too). admin-dashboard gets its own "занят"
badge, distinct from fail/partial/cancelled.
Rebuilt bin/{control-api,admin-dashboard,prober,validator-agent} and
bin/SHA256SUMS per docs/SETUP.md's documented build recipe, since
control-api and admin-dashboard source changed.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NeVbMVEiE7XQAkBd7HQgj6