docs/ADMIN_CLEANUP.md: what is in the control-api database and what must
not be touched, preparation (stop, backup, checks), ready SQL for a full
reset before a new run, the event log, the check registry, single
addresses and compaction, verification after the cleanup, restore from a
backup, and what to do through the API instead. Every SQL block was run
on a copy of the production backup. Linked from the README.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
A mass check on 2026-10-02 stalled 7 of 20 validators and sent 42
addresses to fail without a single check. A validator busy with slow
checks went silent, was marked unreachable, and its next heartbeat put it
back to idle while it still held the address; it was handed a second one,
whose association never ran (the in-flight guard was keyed by validator),
and both waited for their leases to expire.
- Heartbeat/re-register return an unreachable validator to assigned when
it still holds an address, else idle.
- A validator is released only from the address it currently holds
(ReleaseFIP, RequeueOrFail, MarkFIPOccupied, FreeValidator); an
unreachable validator stays unreachable until its next heartbeat, so a
dead validator is no longer handed a new address every lease period.
- ClaimNextQueued refuses a validator that still has an address; a
ReconcileValidators pass on every tick repairs rows that disagree with
the queue.
- Association guard is keyed by address, not validator.
- The agent sends heartbeats from their own goroutine.
- Clear queue / delete: detach only floating IPs of unfinished rows (done,
failed and occupied rows kept their fip_id and made a clear issue >1000
sequential cloud calls: 256 s), at most 8 in parallel; the operation no
longer dies with the client connection (10 minute limit).
Includes the incident analysis and the plan under analysis/ and
docs/changes/, and rebuilt bin/control-api and bin/validator-agent.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
The real container on the validators is named validator-agent (the image
is cloud-ip-validator-validator-agent); the playbook used the image name
as the container name, so it would have started a second agent next to
the old one with the same validator_id. The clone on the validators is
owned by root, so git must run as root (git_user), otherwise fetch fails
with "cannot open .git/FETCH_HEAD: Permission denied".
Preflight now stops when the host has another container of this agent
(by name or image) besides container_name.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Run from the jump host: on each validator it updates the git clone in
/opt/cloud-ip-validator, builds the image there, stops and removes the
current container and starts a new one from the new image. Run
parameters live in an env file (deploy/ansible/env/validator-agent.env,
git-ignored, template committed).
The image is built before the running container is touched, so a failed
build leaves the old container running. Hosts are updated in waves
(1, 4, rest) and any failure stops the run. validator_id comes from the
inventory and is checked against the running container before it is
replaced. Only ansible.builtin modules are used, so the validators need
no extra packages.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
control-api is hosted outside the cloud and validators reach it directly,
so it sees the floating IP as the connection's source address. New open
route GET /api/v1/agents/{id}/observed-ip returns that address (taken only
from the TCP peer; forwarding headers are ignored so a validator cannot
forge it).
The agent gets self_check.methods, a priority-ordered list of ip_echo
(unchanged) and control_api; the default stays [ip_echo]. The self-check
passes when any method confirms the address; the next method is tried on
no answer and on a mismatch. Each method has its own timeout so a hung
first method cannot starve the fallback, and control_api uses a new TCP
connection per call (a connection opened before the floating IP was
attached would keep reporting the old address).
Also: docker agent template/env, example config, docs, plan in
docs/changes, e2e script switch E2E_SELF_CHECK_METHODS, rebuilt
bin/control-api and bin/validator-agent.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
A failed self-check returned the address to the queue and freed the
validator in the database, but left the floating IP attached to the
validator's port. Every later association on that port then failed with
409 ("fixed IP already has a floating IP"), so one failed self-check
poisoned a validator for good; on 2026-10-01 all 20 validators were
poisoned within 23 minutes after ifconfig.me timeouts.
SelfCheckResult now disassociates the floating IP before requeueing, and
ignores a late failed report for an address the validator no longer
holds (it could belong to another validator by then).
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
An address still being associated (assigning_fip) has no fip_id in the
database, so "clear queue" did not detach it, and the association then
finished after the row was gone, leaving the floating IP on the validator
port for good.
- After clear/cancel/delete, ask the cloud which floating IPs sit on the
affected validator ports (new ListFloatingIPsByPort) and detach those
that this system queued (known in ip_registry); foreign ones are left.
- SetFIPAssociated applies only to a row still in assigning_fip; if the
address was removed meanwhile, associateFIP detaches the floating IP.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
The orchestrator claimed idle validators one by one and associated each
floating IP synchronously (~30 s per address), so 20 validators started
about 30 s apart. Aggregation/disassociation and lease reclaim were
sequential in the same way.
The slow OpenStack calls now run in one goroutine per address, guarded by
an in-flight set against duplicates. control-api runs Tick in Async mode
(Tick does not wait); tests keep the waiting mode.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
The "Scan Floating IP" button failed with a client timeout: the project now
holds ~6.4k floating IPs and the scan listed them all in one unpaginated,
timeout-less Neutron request on the HTTP request context.
openstack: ListFreeFloatingIPs reads marker-based pages (fields= keeps them
small) with per-page retry/backoff on transport errors, 5xx and 429, and every
request now has a timeout (also ends hangs inside the orchestrator tick).
orchestrator: the scan is a single-flight background job on the process
context with progress (clearing/listing/enqueuing/done/error), dry_run, full
discovery before anything is enqueued, then SubmitIPs in chunks of 500 in
ascending IP order; a failed read leaves the queue untouched. The auto-cycle
gets a "scanning" phase that polls the job, so the control loop and
autoCycleMu are never held across OpenStack/DB work; it recovers after a
restart and waits for (instead of adopting) a scan started by someone else.
db: migration 0009 (indexes), paged ListIPsPage/ListRegistryPage, GROUP BY
counters, EXISTS completion check, set-based ClearAllIPs.
API: POST /admin/ips/scan -> 202 (dry_run, wait), GET /admin/ips/scan, paging
and filters on /admin/ips and /admin/registry (bare arrays without limit),
results_by_overall in /admin/status.
dashboard: scan progress panel and dry-run button, paginated /ips and
/registry with server-side filters, Overview on counters and capped lists
with progress/ETA, "select all N by filter", hx-params fix for per-row
buttons, real counts in confirmations.
Also: docs (API, USAGE, DASHBOARD, README), plan and review under
docs/changes/, bin/ rebuilt with new SHA256SUMS.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
control-api: every route now carries a mandatory access level (admin / agent /
open) in a route table. All /api/v1/admin/* require the admin token; the
write calls of validator-agent and prober (self-check, events, results,
complete) require a separate static agent token; register, heartbeat and
fetching the assignment stay open. Tokens come from env vars, are compared in
constant time and never logged. An empty token leaves that level open with a
startup warning (backward compatible).
validator-agent / prober: apiclient sends the agent token only to control-api.
admin-dashboard: login/password (from env) with a stateless HMAC session
cookie, Origin-based CSRF check, per-IP brute-force throttle, HX-Redirect for
htmx polls, logout in the sidebar; the dashboard calls control-api with the
admin token. Login page layout fixed after review.
Also: env plumbing in docker-compose/rxprod-compose/systemd/config examples,
e2e script with token assertions, tests, docs (API, SETUP, USAGE, DASHBOARD,
README), plan and review under docs/changes/, bin/ rebuilt with new
SHA256SUMS.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Follow the section layout used in the ipam_control project: quick start,
configuration tables, architecture with a directory tree, data-model rules
(queue states, results, registry, auto-cycle), API table, security notes,
operations, dashboard pages, tests, docs index and change history. Details
stay in docs/ and are linked rather than duplicated.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
An admin-controlled scenario that repeats what the operator does by hand:
clear the IP queue, scan and enqueue all free Floating IPs, wait until every
queued address reaches a terminal state (so results are in the Registry),
then wait a configurable interval and start over.
- control-api: new auto_cycle singleton table (migration 0008) holding
enabled/interval/max-run settings and persisted phase state, so the cycle
survives restarts; engine in orchestrator/autocycle.go driven from the
existing loop tick with an injectable "now" for deterministic tests.
- Interval (default 1h, min 60s) and max wait (default unlimited, timeout
outcome) are runtime settings, never hardcoded.
- The periodic fip_scan_interval_seconds scan is skipped while the cycle is
enabled. An emptied queue mid-cycle counts as finished; stopping during
the pause keeps the last cycle's outcome.
- API: GET/PUT /api/v1/admin/auto-cycle, POST .../start, POST .../stop.
- admin-dashboard: "Автоматический цикл" panel on /settings and an
"Автоцикл активен" indicator on /overview.
- Tests for db, orchestrator, httpapi and dashboard; run-local-e2e.sh now
exercises a full auto cycle; docs updated; bin/ rebuilt with refreshed
SHA256SUMS.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Replaced the duplicated Docker walkthrough (now fully covered in
docs/SETUP.md) with a single link, and turned the prose description of
each component into an explicit role/state/host table so the architecture
reads at a glance. Cuts the file roughly in half.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Adds a dedicated DASHBOARD.md subsection covering what the q/status
filter applies to on each page, why it's dashboard-side only, and the
/overview layout constraint (stats panel outside the polled block, kept
in sync via an out-of-band swap) so future layout changes don't
reintroduce the polling-wipes-the-filter bug. Cross-links added from
USAGE.md's queue-observation and registry sections.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Placing the filter form outside the polled div (so the poll can't wipe out
typed/selected values) put it before the stat-grid, since both used to
live inside that same polled block — stats ended up after the filter
instead of before it, as it always was.
Splits the stat-grid out into its own #overview-stats div, positioned
before the filter form; the actual poll target is now #overview-tables
(just the two tables). Since #overview-stats no longer polls directly,
/overview/fragment now also renders it as an out-of-band swap alongside
the main #overview-tables response — the same hx-swap-oob idiom already
used for the shared error banner — so the stat counts still refresh every
tick even though they're outside the polled element.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Both pages rendered their full lists with no way to narrow them. Adds a
?q=&status= filter (substring match on address, exact match on result
status) applied dashboard-side, in Go, over the already-fetched list —
no control-api/db changes needed.
Overview: the filter form lives outside the polling target (#overview-live)
so the recurring poll never wipes out what's typed/selected; the poll and
both filter inputs share hx-sync="#overview-live:queue last" (the same
fix that resolved the earlier abandoned /ips auto-refresh races) and the
poll now carries hx-include="#overview-filter" so it keeps honoring the
current filter on every tick. Applies uniformly to both the "Текущая
проверка" and "Последние N завершённых" tables, per the confirmed design:
picking a specific status naturally hides in-progress rows, since they
have no result yet.
Registry: no polling exists there, so the filter form reuses the full page
via hx-select/hx-replace-url — simpler than adding a parallel fragment
endpoint, and gives a bookmarkable/shareable filtered URL.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
fillRegistrySummary derived LastResult from whichever single checks row
happened to have the latest checked_at, not the cycle's actual aggregated
result — a cycle with a mix of passing and failing checks (e.g. one egress
target timed out while the rest, including the chronologically-last check,
succeeded) rendered as a green "pass" badge on /registry, disagreeing with
the correct "partial" badge already shown on /ips for the same address.
Now prefers ip_queue.overall_result (the orchestrator's own aggregation)
when a live queue row has a finished cycle, leaves the badge blank while a
cycle is still in progress, and only falls back to classifying the most
recent cycle's own checks (pass/fail/partial) once the address has been
deleted from the queue and overall_result is no longer available.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
deploy/docker/RUN.txt only documented docker build/run per component with
no explanation of docker-compose, multi-host profiles, or how the images
relate to bin/*. Adds a full "Развёртывание в Docker" section to SETUP.md
covering requirements, single-host and multi-host docker-compose flows,
per-component docker build/run (absorbing RUN.txt's content with more
context), rebuilding images after code changes, and troubleshooting.
README.md now points here as the primary source, with RUN.txt kept as a
quick command reference.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Adds POST /api/v1/admin/ips/scan (plus an optional periodic ticker) to
discover free Floating IPs in the OpenStack project and feed them straight
into the check queue. More importantly, decouples check/event history from
ip_queue's lifecycle: a new ip_registry table (migration 0007) gives every
address ever submitted a durable identity, so deleting it from the queue no
longer destroys its history — it's still reachable via the new
GET /api/v1/admin/registry[/{ip}] endpoints and the dashboard's /registry
pages, with retention depth configurable in check cycles per address
(history_retention_cycles, 0 = unlimited).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
The 5s auto-refresh (ec44d54) and the hx-sync fix on top of it (5a53705)
didn't resolve the issues seen in manual testing. Rather than keep
debugging htmx's polling/preserve/sync interaction, drop auto-refresh
entirely: handleIPsFragment, GET /ips/fragment, and the poll
attributes/PollSeconds plumbing are all removed. The table now only
updates when a button action re-renders it, as it did before auto-refresh
was added — the bulk-recheck feature itself (handleIPsRecheckSelected,
POST /ips/recheck, "Перепроверить выбранные") is untouched.
Also drops hx-preserve/id from the row checkboxes: it existed solely to
survive the auto-poll wiping a selection mid-task, so it has no purpose
left, and it was actively wrong for one case — after a successful
"Перепроверить выбранные", it kept the just-submitted addresses checked
instead of clearing them. Since the checkbox's checked state was never
server-rendered to begin with, removing hx-preserve alone makes every
table swap (including the recheck button's own) render fresh, unchecked
boxes, which is exactly the desired "selection clears once the action has
been applied" behavior.
Left hx-sync="#ips-table-wrap:queue last" on the action buttons/form —
still cheap protection against a double-click race between two real user
actions, independent of the now-removed polling.
Rebuilt bin/admin-dashboard and bin/SHA256SUMS per docs/SETUP.md.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
The 5s auto-poll (hx-trigger="every Ns" on #ips-table-wrap) and every
button/form that also swaps #ips-table-wrap (bulk recheck/delete/clear,
per-row recheck/cancel/delete, the add-address form) each fired
independent, uncoordinated htmx requests against the same target. With no
hx-sync, whichever response landed last won — including a poll's
in-flight GET landing *after* a slower mutation's own response and
silently reverting the just-applied change with stale data. This matched
every symptom reported: buttons needing several clicks before they
"took", "Перепроверить выбранные" appearing to do nothing with many rows
selected (more DB writes -> wider race window for a poll to land after
and clobber it), the page "blinking" back to a stale queued state a few
seconds after a bulk recheck actually succeeded, and auto-refresh working
"every other time".
Every element that targets #ips-table-wrap now shares
hx-sync="#ips-table-wrap:queue last", so at most one request affecting it
is ever in flight: a trigger that fires while another is pending gets
queued (never aborted mid-write) and only the most recent queued trigger
actually runs once the current one finishes, guaranteeing responses are
always applied in the order they actually resolve.
Rebuilt bin/admin-dashboard (only internal/dashboard changed) and
bin/SHA256SUMS per docs/SETUP.md's documented build recipe.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Per docs/SETUP.md's documented build recipe. control-api/prober/
validator-agent are unaffected (only internal/dashboard changed) and keep
their existing hashes in bin/SHA256SUMS.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
"Перепроверить выбранные" joins the existing "Удалить выбранные" / "Очистить
всё" bulk actions, using the checked-row selection the same way delete
already does — no control-api changes needed, since forcing a recheck of a
batch of addresses (new, finished, or already-queued, skipping anything
mid-check) is exactly what POST /api/v1/admin/ips (db.SubmitIPs) already
does, and the dashboard's own SubmitIPs client method already backs both
the top add/recheck form and the single-row recheck button.
The page also now auto-refreshes every 5s (handleIPsFragment + GET
/ips/fragment), mirroring the overview page's existing hx-trigger="every
Ns" polling and reusing the same config-driven interval
(Cfg.OverviewPollIntervalS) rather than adding a duplicate knob. Since the
table now polls itself, each row's selection checkbox gets a stable id
plus hx-preserve so a checked box survives the refresh (its own DOM node
is kept) while the rest of the row still updates live.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
docker-compose.yml now runs from prebuilt images (civ-capi/civ-adash/
civ-prober) instead of building from source, mounts this host's own
rxprod-compose/control-api.yaml and capi-db/ (both local, gitignored
except control-api.yaml itself which is now committed), and moves
control-api off port 8080 onto 8081.
rxprod-compose/sources/ carries local copies of the example configs
(admin-dashboard/control-api/prober/validator-agent) plus .env.example,
moved here from rxprod-compose/.env.example, for reference alongside
this specific deployment's docker-compose.yml.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
control-api.yaml is the actual bootstrap config the rxprod-compose stack's
control-api container mounts (per its docker-compose.yml), previously
undocumented/uncommitted — credentials stay out of it entirely (only the
env var names to read them from, per its own header comment).
.gitignore now also excludes rxprod-compose/capi-db/ (the stack's live
SQLite database — runtime state, not source) and graphify-out/ (this
session's /graphify knowledge-graph output — regenerable, not source).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Each target group's relationship to the check types that reference it was
previously invisible without cross-referencing /check-types by hand.
handleTargetsPage/renderTargetsTable now resolve that server-side via
loadTargetsPage (group -> referencing check types, enabled or not) and
render it as a status pill per group, plus a stats row (group count,
total targets, check types using them, unused groups) kept live via an
out-of-band swap on every create/update/delete.
Also auto-sizes the targets textarea as you type (targets.html), and
lets the stat-card grid collapse to however many state cards actually
exist instead of leaving empty background where a fixed repeat(6, ...)
had no card to fill.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Both binaries registered with control-api exactly once at startup and
exited (os.Exit(1)) on any failure — including control-api simply not
being up yet (no ordering guarantee between the two at boot/redeploy) or
the admin not having added this validator_id/site_id to the config yet.
Run() now retries registration with capped exponential backoff (3s->30s)
until it succeeds or the process is asked to shut down, instead of
crashing; registerWithRetry is identical in agentcore and probercore
since their Run/register shape already was.
Separately, the admin dashboard's Validators page had no hostname column
even though the agent already reports one on register (mirroring the
prober) and control-api already persists it — only the admin-config read
DTO (validatorDTO in httpapi and dashboard) dropped it before it reached
the template. Added hostname + last_heartbeat_at to that DTO end-to-end
and a Хост/Heartbeat column to validators.html, matching sites.html.
Rebuilt bin/{control-api,admin-dashboard,prober,validator-agent} and
bin/SHA256SUMS per docs/SETUP.md's documented build recipe.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
The existing deploy/docker/docker-compose.yml (+ override/prod, Compose
profiles) is flexible but requires understanding profiles and file
layering. For a server that only ever runs these three fixed roles
against a real OpenStack (VK Cloud) deployment, rxprod-compose/ adds a
single flat docker-compose.yml with no profiles — control-api,
admin-dashboard and prober wired together directly, both ports published
on the host (control-api needs to be reachable by validator-agent running
separately on real cloud VMs). OpenStack credentials and the prober's
site_id come from a local .env (gitignored; .env.example is the
template).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NeVbMVEiE7XQAkBd7HQgj6
The cloud is live: the address list submitted as "free" (bootstrap config
or POST /api/v1/admin/ips) can drift by the time the orchestrator claims
it, or an operator can queue an already-occupied address by mistake.
Neutron's floating-IP association is a blind "last write wins" PUT with
no conflict error to catch, so associateFIP now checks the FIP's PortID
(already fetched via GetFloatingIPByAddress) before associating, guarded
against the false-positive of the FIP already belonging to this same
validator's own port.
A match routes the address straight to a new terminal ip_queue.state
("occupied", distinct from failed/fail) via db.MarkFIPOccupied — no
retries, since Neutron won't free it on its own and requeuing would let
it be reclaimed again next tick, starving the rest of the queue — plus a
dedicated fip_occupied audit event. Resubmitting the address later (once
the conflict is resolved) resets it to queued via the existing
POST /api/v1/admin/ips resubmit path (CancelIP/ListExpiredLeases updated
to treat occupied as terminal too). admin-dashboard gets its own "занят"
badge, distinct from fail/partial/cancelled.
Rebuilt bin/{control-api,admin-dashboard,prober,validator-agent} and
bin/SHA256SUMS per docs/SETUP.md's documented build recipe, since
control-api and admin-dashboard source changed.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NeVbMVEiE7XQAkBd7HQgj6
Standalone HTML page (docs/CONTROL_DATA_PLANE.html) showing the
docker-compose deployment topology (hosts, COMPOSE_PROFILES) and the
egress/inbound check traffic, refined through several presentation
review passes. Linked from README.md's docs table and docs/DIAGRAMS.md
alongside the existing Mermaid diagrams.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NeVbMVEiE7XQAkBd7HQgj6
control-api was the only component deployable via a bare Dockerfile alone
(no entrypoint/template, no default mountable config), and there was no
compose file wiring the four services' network, healthcheck ordering, or
DB volume together. Adds a base docker-compose.yml plus dev (override,
auto-loaded) and prod overlays, split by Compose profiles matching the
real deployment topology (control-plane/dashboard/prober/validator), a
ready-to-copy mock config for control-api so `docker compose up` works
out of the box, and the repo's first .gitignore for the local env/config
copies developers create from the committed examples.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NeVbMVEiE7XQAkBd7HQgj6