The 5s auto-refresh (ec44d54) and the hx-sync fix on top of it (5a53705)
didn't resolve the issues seen in manual testing. Rather than keep
debugging htmx's polling/preserve/sync interaction, drop auto-refresh
entirely: handleIPsFragment, GET /ips/fragment, and the poll
attributes/PollSeconds plumbing are all removed. The table now only
updates when a button action re-renders it, as it did before auto-refresh
was added — the bulk-recheck feature itself (handleIPsRecheckSelected,
POST /ips/recheck, "Перепроверить выбранные") is untouched.
Also drops hx-preserve/id from the row checkboxes: it existed solely to
survive the auto-poll wiping a selection mid-task, so it has no purpose
left, and it was actively wrong for one case — after a successful
"Перепроверить выбранные", it kept the just-submitted addresses checked
instead of clearing them. Since the checkbox's checked state was never
server-rendered to begin with, removing hx-preserve alone makes every
table swap (including the recheck button's own) render fresh, unchecked
boxes, which is exactly the desired "selection clears once the action has
been applied" behavior.
Left hx-sync="#ips-table-wrap:queue last" on the action buttons/form —
still cheap protection against a double-click race between two real user
actions, independent of the now-removed polling.
Rebuilt bin/admin-dashboard and bin/SHA256SUMS per docs/SETUP.md.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
The 5s auto-poll (hx-trigger="every Ns" on #ips-table-wrap) and every
button/form that also swaps #ips-table-wrap (bulk recheck/delete/clear,
per-row recheck/cancel/delete, the add-address form) each fired
independent, uncoordinated htmx requests against the same target. With no
hx-sync, whichever response landed last won — including a poll's
in-flight GET landing *after* a slower mutation's own response and
silently reverting the just-applied change with stale data. This matched
every symptom reported: buttons needing several clicks before they
"took", "Перепроверить выбранные" appearing to do nothing with many rows
selected (more DB writes -> wider race window for a poll to land after
and clobber it), the page "blinking" back to a stale queued state a few
seconds after a bulk recheck actually succeeded, and auto-refresh working
"every other time".
Every element that targets #ips-table-wrap now shares
hx-sync="#ips-table-wrap:queue last", so at most one request affecting it
is ever in flight: a trigger that fires while another is pending gets
queued (never aborted mid-write) and only the most recent queued trigger
actually runs once the current one finishes, guaranteeing responses are
always applied in the order they actually resolve.
Rebuilt bin/admin-dashboard (only internal/dashboard changed) and
bin/SHA256SUMS per docs/SETUP.md's documented build recipe.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Per docs/SETUP.md's documented build recipe. control-api/prober/
validator-agent are unaffected (only internal/dashboard changed) and keep
their existing hashes in bin/SHA256SUMS.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Both binaries registered with control-api exactly once at startup and
exited (os.Exit(1)) on any failure — including control-api simply not
being up yet (no ordering guarantee between the two at boot/redeploy) or
the admin not having added this validator_id/site_id to the config yet.
Run() now retries registration with capped exponential backoff (3s->30s)
until it succeeds or the process is asked to shut down, instead of
crashing; registerWithRetry is identical in agentcore and probercore
since their Run/register shape already was.
Separately, the admin dashboard's Validators page had no hostname column
even though the agent already reports one on register (mirroring the
prober) and control-api already persists it — only the admin-config read
DTO (validatorDTO in httpapi and dashboard) dropped it before it reached
the template. Added hostname + last_heartbeat_at to that DTO end-to-end
and a Хост/Heartbeat column to validators.html, matching sites.html.
Rebuilt bin/{control-api,admin-dashboard,prober,validator-agent} and
bin/SHA256SUMS per docs/SETUP.md's documented build recipe.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
The cloud is live: the address list submitted as "free" (bootstrap config
or POST /api/v1/admin/ips) can drift by the time the orchestrator claims
it, or an operator can queue an already-occupied address by mistake.
Neutron's floating-IP association is a blind "last write wins" PUT with
no conflict error to catch, so associateFIP now checks the FIP's PortID
(already fetched via GetFloatingIPByAddress) before associating, guarded
against the false-positive of the FIP already belonging to this same
validator's own port.
A match routes the address straight to a new terminal ip_queue.state
("occupied", distinct from failed/fail) via db.MarkFIPOccupied — no
retries, since Neutron won't free it on its own and requeuing would let
it be reclaimed again next tick, starving the rest of the queue — plus a
dedicated fip_occupied audit event. Resubmitting the address later (once
the conflict is resolved) resets it to queued via the existing
POST /api/v1/admin/ips resubmit path (CancelIP/ListExpiredLeases updated
to treat occupied as terminal too). admin-dashboard gets its own "занят"
badge, distinct from fail/partial/cancelled.
Rebuilt bin/{control-api,admin-dashboard,prober,validator-agent} and
bin/SHA256SUMS per docs/SETUP.md's documented build recipe, since
control-api and admin-dashboard source changed.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NeVbMVEiE7XQAkBd7HQgj6