Commit Graph
11 Commits
Author SHA1 Message Date
ayurishchevandClaude Sonnet 5.5 aff8fe38b5 Scan floating IPs in the background, page by page, so thousands of addresses work
The "Scan Floating IP" button failed with a client timeout: the project now
holds ~6.4k floating IPs and the scan listed them all in one unpaginated,
timeout-less Neutron request on the HTTP request context.

openstack: ListFreeFloatingIPs reads marker-based pages (fields= keeps them
small) with per-page retry/backoff on transport errors, 5xx and 429, and every
request now has a timeout (also ends hangs inside the orchestrator tick).

orchestrator: the scan is a single-flight background job on the process
context with progress (clearing/listing/enqueuing/done/error), dry_run, full
discovery before anything is enqueued, then SubmitIPs in chunks of 500 in
ascending IP order; a failed read leaves the queue untouched. The auto-cycle
gets a "scanning" phase that polls the job, so the control loop and
autoCycleMu are never held across OpenStack/DB work; it recovers after a
restart and waits for (instead of adopting) a scan started by someone else.

db: migration 0009 (indexes), paged ListIPsPage/ListRegistryPage, GROUP BY
counters, EXISTS completion check, set-based ClearAllIPs.

API: POST /admin/ips/scan -> 202 (dry_run, wait), GET /admin/ips/scan, paging
and filters on /admin/ips and /admin/registry (bare arrays without limit),
results_by_overall in /admin/status.

dashboard: scan progress panel and dry-run button, paginated /ips and
/registry with server-side filters, Overview on counters and capped lists
with progress/ETA, "select all N by filter", hx-params fix for per-row
buttons, real counts in confirmations.

Also: docs (API, USAGE, DASHBOARD, README), plan and review under
docs/changes/, bin/ rebuilt with new SHA256SUMS.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-10-01 19:31:11 +03:00
ayurishchevandClaude Sonnet 5.5 debf2afed2 Add authentication: admin/agent bearer tokens for the API, login for the dashboard
control-api: every route now carries a mandatory access level (admin / agent /
open) in a route table. All /api/v1/admin/* require the admin token; the
write calls of validator-agent and prober (self-check, events, results,
complete) require a separate static agent token; register, heartbeat and
fetching the assignment stay open. Tokens come from env vars, are compared in
constant time and never logged. An empty token leaves that level open with a
startup warning (backward compatible).

validator-agent / prober: apiclient sends the agent token only to control-api.

admin-dashboard: login/password (from env) with a stateless HMAC session
cookie, Origin-based CSRF check, per-IP brute-force throttle, HX-Redirect for
htmx polls, logout in the sidebar; the dashboard calls control-api with the
admin token. Login page layout fixed after review.

Also: env plumbing in docker-compose/rxprod-compose/systemd/config examples,
e2e script with token assertions, tests, docs (API, SETUP, USAGE, DASHBOARD,
README), plan and review under docs/changes/, bin/ rebuilt with new
SHA256SUMS.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-10-01 11:35:24 +03:00
ayurishchevandClaude Sonnet 5.5 cd37b10f3b Add optional automatic check cycle (clear queue -> scan FIPs -> wait -> repeat)
An admin-controlled scenario that repeats what the operator does by hand:
clear the IP queue, scan and enqueue all free Floating IPs, wait until every
queued address reaches a terminal state (so results are in the Registry),
then wait a configurable interval and start over.

- control-api: new auto_cycle singleton table (migration 0008) holding
  enabled/interval/max-run settings and persisted phase state, so the cycle
  survives restarts; engine in orchestrator/autocycle.go driven from the
  existing loop tick with an injectable "now" for deterministic tests.
- Interval (default 1h, min 60s) and max wait (default unlimited, timeout
  outcome) are runtime settings, never hardcoded.
- The periodic fip_scan_interval_seconds scan is skipped while the cycle is
  enabled. An emptied queue mid-cycle counts as finished; stopping during
  the pause keeps the last cycle's outcome.
- API: GET/PUT /api/v1/admin/auto-cycle, POST .../start, POST .../stop.
- admin-dashboard: "Автоматический цикл" panel on /settings and an
  "Автоцикл активен" indicator on /overview.
- Tests for db, orchestrator, httpapi and dashboard; run-local-e2e.sh now
  exercises a full auto cycle; docs updated; bin/ rebuilt with refreshed
  SHA256SUMS.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-10-01 10:28:53 +03:00
ayurishchevandClaude Sonnet 5 582b44f314 Add Floating IP scanning and a durable address registry with configurable history depth
Adds POST /api/v1/admin/ips/scan (plus an optional periodic ticker) to
discover free Floating IPs in the OpenStack project and feed them straight
into the check queue. More importantly, decouples check/event history from
ip_queue's lifecycle: a new ip_registry table (migration 0007) gives every
address ever submitted a durable identity, so deleting it from the queue no
longer destroys its history — it's still reachable via the new
GET /api/v1/admin/registry[/{ip}] endpoints and the dashboard's /registry
pages, with retention depth configurable in check cycles per address
(history_retention_cycles, 0 = unlimited).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-23 09:52:01 +03:00
ayurishchevandClaude Sonnet 5 78b20fa5be Remove IPs page auto-refresh; clear checkboxes after bulk recheck
The 5s auto-refresh (ec44d54) and the hx-sync fix on top of it (5a53705)
didn't resolve the issues seen in manual testing. Rather than keep
debugging htmx's polling/preserve/sync interaction, drop auto-refresh
entirely: handleIPsFragment, GET /ips/fragment, and the poll
attributes/PollSeconds plumbing are all removed. The table now only
updates when a button action re-renders it, as it did before auto-refresh
was added — the bulk-recheck feature itself (handleIPsRecheckSelected,
POST /ips/recheck, "Перепроверить выбранные") is untouched.

Also drops hx-preserve/id from the row checkboxes: it existed solely to
survive the auto-poll wiping a selection mid-task, so it has no purpose
left, and it was actively wrong for one case — after a successful
"Перепроверить выбранные", it kept the just-submitted addresses checked
instead of clearing them. Since the checkbox's checked state was never
server-rendered to begin with, removing hx-preserve alone makes every
table swap (including the recheck button's own) render fresh, unchecked
boxes, which is exactly the desired "selection clears once the action has
been applied" behavior.

Left hx-sync="#ips-table-wrap:queue last" on the action buttons/form —
still cheap protection against a double-click race between two real user
actions, independent of the now-removed polling.

Rebuilt bin/admin-dashboard and bin/SHA256SUMS per docs/SETUP.md.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-18 19:29:26 +03:00
ayurishchevandClaude Sonnet 5 ec44d54ac1 Add bulk recheck and auto-refresh to the IPs dashboard page
"Перепроверить выбранные" joins the existing "Удалить выбранные" / "Очистить
всё" bulk actions, using the checked-row selection the same way delete
already does — no control-api changes needed, since forcing a recheck of a
batch of addresses (new, finished, or already-queued, skipping anything
mid-check) is exactly what POST /api/v1/admin/ips (db.SubmitIPs) already
does, and the dashboard's own SubmitIPs client method already backs both
the top add/recheck form and the single-row recheck button.

The page also now auto-refreshes every 5s (handleIPsFragment + GET
/ips/fragment), mirroring the overview page's existing hx-trigger="every
Ns" polling and reusing the same config-driven interval
(Cfg.OverviewPollIntervalS) rather than adding a duplicate knob. Since the
table now polls itself, each row's selection checkbox gets a stable id
plus hx-preserve so a checked box survives the refresh (its own DOM node
is kept) while the rest of the row still updates live.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-18 18:47:11 +03:00
ayurishchev ef24cc9858 New external site management. New external prober heartbeat feature. 2026-08-26 20:47:54 +03:00
ayurishchev 42f584dd3c Manage external site-prober actions va API and Dashboard 2026-08-26 19:45:48 +03:00
ayurishchev f22ad569b6 added FIP Cooldown before start 2026-08-24 10:29:08 +03:00
ayurishchev f8336740ad feature: handle delete operation for IPs 2026-08-23 22:24:55 +03:00
ayurishchev 37910e410b admin control features and admin dashboard 2026-08-23 20:39:22 +03:00