New page /analytics/compare and API GET /admin/analytics/compare (+ /lists/{group}):
the administrator picks an old (A) and a new (B) run; the report shows the new
addresses (only in B), the ones that left (only in A) and the common ones whose
membership in the seven indicators (pass, partial, fail, egress https any/all,
ingress ssh any/all) differs, with a "what changed" summary per address; the
dynamics of each indicator (delta = new - left + entered - exited) and a verdict
transition matrix. Every number opens a list with CSV. Cancelled addresses are not
part of a run. The list dialog moved to a shared analytics-dialog.js and template;
/analytics got a "compare with another run" button.
Docs, plan and summary in docs/changes/.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
The three verdict cards on /analytics now open the same dialog as the https/ssh
cards, with the addresses of the run that got this verdict (CSV and copy
included). New list kinds verdict_pass, verdict_partial and verdict_fail in
GET /admin/analytics/runs/{id}/lists/{kind}: address, subnet, validator,
egress and ingress "ok of all", checks stored of expected; partial adds the
reason, the same names as the "Why partial" block. Cancelled addresses are not
listed; the row count equals summary.pass/partial/fail. The fail card stays
inert at zero. addr.incomplete() is shared by the list and the report.
Docs, plan and summary in docs/changes/.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Runs (migration 0011): a run groups the cycles of one launch. It opens when an
address enters an idle queue, takes everything submitted or re-checked while it
is open and is finalized when all its addresses are done; a re-check after that
opens a new run, so results of different runs never mix. check_runs,
run_results (one result per address and run, with the verdict and the expected
and stored check counts), subnets, run_id on ip_queue and checks. Existing data
is split into runs at pauses of more than an hour; ingress checks get the
validator that held the address (also at write time from now on).
Analytics (internal/analytics): figures computed from the stored checks of the
latest cycle of each address in the run, as facts next to the verdict: summary,
reasons of partial, data quality, subnets, targets and the subnet x target
matrix by check type, ingress by site, error classes, validators, and the
address lists behind the indicators and error classes. API: analytics runs,
report, lists (JSON or CSV), subnet list; run and subnet filters for the
registry.
Dashboard: /analytics matching the approved mockup (run selector, indicators
with address lists and CSV, error-class dialogs, drill-down to the registry),
subnet list on /settings. Sidebar: the control-api link state, theme toggle and
logout moved to the top, the three dots next to the logo removed, sections
grouped.
Rebuilt bin/control-api and bin/admin-dashboard to match. Plan, summary and the
updated README, API, USAGE, DASHBOARD and ADMIN_CLEANUP docs are in docs/.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Registry: the "last result" column now also shows, per level (egress,
ingress), how many of the recorded checks of the latest cycle succeeded, split
by check family (tcp-22 and tcp-443 are both "tcp"). One grouped query per
chunk of addresses; new fields last_cycle_id, egress, ingress in
GET /admin/registry; the dashboard renders them under the verdict.
Verdict integrity (migration 0010):
- the prober is handed an address once per site and attempt, not on every
poll, so results are no longer overwritten by later probe rounds;
- UpsertCheckIfOpen refuses writes once the address is aggregating or has its
verdict, or for an older attempt; senders get {"ok":true,"ignored":N} and a
result_dropped event is recorded;
- the checking window counts from checking_started_at, not from assigned_at;
- checks.recorded_at (server clock) and checks.after_verdict (flag for rows
written after the verdict in existing data);
- the verdict rule is a pure function (computeVerdict) and the aggregated
event carries the egress/ingress check counts.
Rebuilt bin/control-api and bin/admin-dashboard to match. Plans and summaries
are in docs/changes; README, API, USAGE, DASHBOARD and DIAGRAMS are updated.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
The "Scan Floating IP" button failed with a client timeout: the project now
holds ~6.4k floating IPs and the scan listed them all in one unpaginated,
timeout-less Neutron request on the HTTP request context.
openstack: ListFreeFloatingIPs reads marker-based pages (fields= keeps them
small) with per-page retry/backoff on transport errors, 5xx and 429, and every
request now has a timeout (also ends hangs inside the orchestrator tick).
orchestrator: the scan is a single-flight background job on the process
context with progress (clearing/listing/enqueuing/done/error), dry_run, full
discovery before anything is enqueued, then SubmitIPs in chunks of 500 in
ascending IP order; a failed read leaves the queue untouched. The auto-cycle
gets a "scanning" phase that polls the job, so the control loop and
autoCycleMu are never held across OpenStack/DB work; it recovers after a
restart and waits for (instead of adopting) a scan started by someone else.
db: migration 0009 (indexes), paged ListIPsPage/ListRegistryPage, GROUP BY
counters, EXISTS completion check, set-based ClearAllIPs.
API: POST /admin/ips/scan -> 202 (dry_run, wait), GET /admin/ips/scan, paging
and filters on /admin/ips and /admin/registry (bare arrays without limit),
results_by_overall in /admin/status.
dashboard: scan progress panel and dry-run button, paginated /ips and
/registry with server-side filters, Overview on counters and capped lists
with progress/ETA, "select all N by filter", hx-params fix for per-row
buttons, real counts in confirmations.
Also: docs (API, USAGE, DASHBOARD, README), plan and review under
docs/changes/, bin/ rebuilt with new SHA256SUMS.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
control-api: every route now carries a mandatory access level (admin / agent /
open) in a route table. All /api/v1/admin/* require the admin token; the
write calls of validator-agent and prober (self-check, events, results,
complete) require a separate static agent token; register, heartbeat and
fetching the assignment stay open. Tokens come from env vars, are compared in
constant time and never logged. An empty token leaves that level open with a
startup warning (backward compatible).
validator-agent / prober: apiclient sends the agent token only to control-api.
admin-dashboard: login/password (from env) with a stateless HMAC session
cookie, Origin-based CSRF check, per-IP brute-force throttle, HX-Redirect for
htmx polls, logout in the sidebar; the dashboard calls control-api with the
admin token. Login page layout fixed after review.
Also: env plumbing in docker-compose/rxprod-compose/systemd/config examples,
e2e script with token assertions, tests, docs (API, SETUP, USAGE, DASHBOARD,
README), plan and review under docs/changes/, bin/ rebuilt with new
SHA256SUMS.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Each target group's relationship to the check types that reference it was
previously invisible without cross-referencing /check-types by hand.
handleTargetsPage/renderTargetsTable now resolve that server-side via
loadTargetsPage (group -> referencing check types, enabled or not) and
render it as a status pill per group, plus a stats row (group count,
total targets, check types using them, unused groups) kept live via an
out-of-band swap on every create/update/delete.
Also auto-sizes the targets textarea as you type (targets.html), and
lets the stat-card grid collapse to however many state cards actually
exist instead of leaving empty background where a fixed repeat(6, ...)
had no card to fill.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
The cloud is live: the address list submitted as "free" (bootstrap config
or POST /api/v1/admin/ips) can drift by the time the orchestrator claims
it, or an operator can queue an already-occupied address by mistake.
Neutron's floating-IP association is a blind "last write wins" PUT with
no conflict error to catch, so associateFIP now checks the FIP's PortID
(already fetched via GetFloatingIPByAddress) before associating, guarded
against the false-positive of the FIP already belonging to this same
validator's own port.
A match routes the address straight to a new terminal ip_queue.state
("occupied", distinct from failed/fail) via db.MarkFIPOccupied — no
retries, since Neutron won't free it on its own and requeuing would let
it be reclaimed again next tick, starving the rest of the queue — plus a
dedicated fip_occupied audit event. Resubmitting the address later (once
the conflict is resolved) resets it to queued via the existing
POST /api/v1/admin/ips resubmit path (CancelIP/ListExpiredLeases updated
to treat occupied as terminal too). admin-dashboard gets its own "занят"
badge, distinct from fail/partial/cancelled.
Rebuilt bin/{control-api,admin-dashboard,prober,validator-agent} and
bin/SHA256SUMS per docs/SETUP.md's documented build recipe, since
control-api and admin-dashboard source changed.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NeVbMVEiE7XQAkBd7HQgj6