The "Scan Floating IP" button failed with a client timeout: the project now
holds ~6.4k floating IPs and the scan listed them all in one unpaginated,
timeout-less Neutron request on the HTTP request context.
openstack: ListFreeFloatingIPs reads marker-based pages (fields= keeps them
small) with per-page retry/backoff on transport errors, 5xx and 429, and every
request now has a timeout (also ends hangs inside the orchestrator tick).
orchestrator: the scan is a single-flight background job on the process
context with progress (clearing/listing/enqueuing/done/error), dry_run, full
discovery before anything is enqueued, then SubmitIPs in chunks of 500 in
ascending IP order; a failed read leaves the queue untouched. The auto-cycle
gets a "scanning" phase that polls the job, so the control loop and
autoCycleMu are never held across OpenStack/DB work; it recovers after a
restart and waits for (instead of adopting) a scan started by someone else.
db: migration 0009 (indexes), paged ListIPsPage/ListRegistryPage, GROUP BY
counters, EXISTS completion check, set-based ClearAllIPs.
API: POST /admin/ips/scan -> 202 (dry_run, wait), GET /admin/ips/scan, paging
and filters on /admin/ips and /admin/registry (bare arrays without limit),
results_by_overall in /admin/status.
dashboard: scan progress panel and dry-run button, paginated /ips and
/registry with server-side filters, Overview on counters and capped lists
with progress/ETA, "select all N by filter", hx-params fix for per-row
buttons, real counts in confirmations.
Also: docs (API, USAGE, DASHBOARD, README), plan and review under
docs/changes/, bin/ rebuilt with new SHA256SUMS.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
control-api: every route now carries a mandatory access level (admin / agent /
open) in a route table. All /api/v1/admin/* require the admin token; the
write calls of validator-agent and prober (self-check, events, results,
complete) require a separate static agent token; register, heartbeat and
fetching the assignment stay open. Tokens come from env vars, are compared in
constant time and never logged. An empty token leaves that level open with a
startup warning (backward compatible).
validator-agent / prober: apiclient sends the agent token only to control-api.
admin-dashboard: login/password (from env) with a stateless HMAC session
cookie, Origin-based CSRF check, per-IP brute-force throttle, HX-Redirect for
htmx polls, logout in the sidebar; the dashboard calls control-api with the
admin token. Login page layout fixed after review.
Also: env plumbing in docker-compose/rxprod-compose/systemd/config examples,
e2e script with token assertions, tests, docs (API, SETUP, USAGE, DASHBOARD,
README), plan and review under docs/changes/, bin/ rebuilt with new
SHA256SUMS.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
An admin-controlled scenario that repeats what the operator does by hand:
clear the IP queue, scan and enqueue all free Floating IPs, wait until every
queued address reaches a terminal state (so results are in the Registry),
then wait a configurable interval and start over.
- control-api: new auto_cycle singleton table (migration 0008) holding
enabled/interval/max-run settings and persisted phase state, so the cycle
survives restarts; engine in orchestrator/autocycle.go driven from the
existing loop tick with an injectable "now" for deterministic tests.
- Interval (default 1h, min 60s) and max wait (default unlimited, timeout
outcome) are runtime settings, never hardcoded.
- The periodic fip_scan_interval_seconds scan is skipped while the cycle is
enabled. An emptied queue mid-cycle counts as finished; stopping during
the pause keeps the last cycle's outcome.
- API: GET/PUT /api/v1/admin/auto-cycle, POST .../start, POST .../stop.
- admin-dashboard: "Автоматический цикл" panel on /settings and an
"Автоцикл активен" indicator on /overview.
- Tests for db, orchestrator, httpapi and dashboard; run-local-e2e.sh now
exercises a full auto cycle; docs updated; bin/ rebuilt with refreshed
SHA256SUMS.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Placing the filter form outside the polled div (so the poll can't wipe out
typed/selected values) put it before the stat-grid, since both used to
live inside that same polled block — stats ended up after the filter
instead of before it, as it always was.
Splits the stat-grid out into its own #overview-stats div, positioned
before the filter form; the actual poll target is now #overview-tables
(just the two tables). Since #overview-stats no longer polls directly,
/overview/fragment now also renders it as an out-of-band swap alongside
the main #overview-tables response — the same hx-swap-oob idiom already
used for the shared error banner — so the stat counts still refresh every
tick even though they're outside the polled element.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Both pages rendered their full lists with no way to narrow them. Adds a
?q=&status= filter (substring match on address, exact match on result
status) applied dashboard-side, in Go, over the already-fetched list —
no control-api/db changes needed.
Overview: the filter form lives outside the polling target (#overview-live)
so the recurring poll never wipes out what's typed/selected; the poll and
both filter inputs share hx-sync="#overview-live:queue last" (the same
fix that resolved the earlier abandoned /ips auto-refresh races) and the
poll now carries hx-include="#overview-filter" so it keeps honoring the
current filter on every tick. Applies uniformly to both the "Текущая
проверка" and "Последние N завершённых" tables, per the confirmed design:
picking a specific status naturally hides in-progress rows, since they
have no result yet.
Registry: no polling exists there, so the filter form reuses the full page
via hx-select/hx-replace-url — simpler than adding a parallel fragment
endpoint, and gives a bookmarkable/shareable filtered URL.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>