Commit Graph
18 Commits
Author SHA1 Message Date
ayurishchevandClaude Sonnet 5.5 e95b5eb7d5 Retry a failed self-check on another validator; add the self-check failure ceiling
A validator that failed the self-check of an address no longer gets that address
again in the current round (ClaimNextQueued skips it); the validator itself stays
in service and takes all other addresses. The verdict fail is set when the number
of failed self-checks of an address reaches settings.self_check_max_attempts
(1..50, default 5, independent of the number of validators); max_retries and
retry_count are no longer used for self-check. If every working validator has
already failed the address, a new round starts and the exclusions lapse.

Migration 0012: ip_self_check_failures (permanent history per registry address),
ip_queue.sc_failures and sc_round_start_cycle (cycle_id is used instead of
attempt_number, which restarts when a queue row is recreated), the setting.
db.FailSelfCheck does it in one transaction; re-submission starts a new series.
API: self_check_max_attempts in GET/PUT /admin/config/orchestrator,
self_check_failed_on in /admin/ips/{ip} and /admin/registry/{ip}. Dashboard: the
field on /settings and the line "Self-check не прошёл на: ..." on the address
pages. Docs, plan and summary in docs/changes/.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-10-04 09:50:12 +03:00
ayurishchevandClaude Sonnet 5.5 b7669c9e41 Add the Analytics section: check runs, analytics API and page
Runs (migration 0011): a run groups the cycles of one launch. It opens when an
address enters an idle queue, takes everything submitted or re-checked while it
is open and is finalized when all its addresses are done; a re-check after that
opens a new run, so results of different runs never mix. check_runs,
run_results (one result per address and run, with the verdict and the expected
and stored check counts), subnets, run_id on ip_queue and checks. Existing data
is split into runs at pauses of more than an hour; ingress checks get the
validator that held the address (also at write time from now on).

Analytics (internal/analytics): figures computed from the stored checks of the
latest cycle of each address in the run, as facts next to the verdict: summary,
reasons of partial, data quality, subnets, targets and the subnet x target
matrix by check type, ingress by site, error classes, validators, and the
address lists behind the indicators and error classes. API: analytics runs,
report, lists (JSON or CSV), subnet list; run and subnet filters for the
registry.

Dashboard: /analytics matching the approved mockup (run selector, indicators
with address lists and CSV, error-class dialogs, drill-down to the registry),
subnet list on /settings. Sidebar: the control-api link state, theme toggle and
logout moved to the top, the three dots next to the logo removed, sections
grouped.

Rebuilt bin/control-api and bin/admin-dashboard to match. Plan, summary and the
updated README, API, USAGE, DASHBOARD and ADMIN_CLEANUP docs are in docs/.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-10-03 18:36:03 +03:00
ayurishchevandClaude Sonnet 5.5 864208238f Show egress/ingress levels in the registry; freeze checks at the verdict
Registry: the "last result" column now also shows, per level (egress,
ingress), how many of the recorded checks of the latest cycle succeeded, split
by check family (tcp-22 and tcp-443 are both "tcp"). One grouped query per
chunk of addresses; new fields last_cycle_id, egress, ingress in
GET /admin/registry; the dashboard renders them under the verdict.

Verdict integrity (migration 0010):
- the prober is handed an address once per site and attempt, not on every
  poll, so results are no longer overwritten by later probe rounds;
- UpsertCheckIfOpen refuses writes once the address is aggregating or has its
  verdict, or for an older attempt; senders get {"ok":true,"ignored":N} and a
  result_dropped event is recorded;
- the checking window counts from checking_started_at, not from assigned_at;
- checks.recorded_at (server clock) and checks.after_verdict (flag for rows
  written after the verdict in existing data);
- the verdict rule is a pure function (computeVerdict) and the aggregated
  event carries the egress/ingress check counts.

Rebuilt bin/control-api and bin/admin-dashboard to match. Plans and summaries
are in docs/changes; README, API, USAGE, DASHBOARD and DIAGRAMS are updated.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-10-03 17:59:52 +03:00
ayurishchevandClaude Sonnet 5.5 0532baff09 Keep one address per validator; fix heartbeat handling and queue clear
A mass check on 2026-10-02 stalled 7 of 20 validators and sent 42
addresses to fail without a single check. A validator busy with slow
checks went silent, was marked unreachable, and its next heartbeat put it
back to idle while it still held the address; it was handed a second one,
whose association never ran (the in-flight guard was keyed by validator),
and both waited for their leases to expire.

- Heartbeat/re-register return an unreachable validator to assigned when
  it still holds an address, else idle.
- A validator is released only from the address it currently holds
  (ReleaseFIP, RequeueOrFail, MarkFIPOccupied, FreeValidator); an
  unreachable validator stays unreachable until its next heartbeat, so a
  dead validator is no longer handed a new address every lease period.
- ClaimNextQueued refuses a validator that still has an address; a
  ReconcileValidators pass on every tick repairs rows that disagree with
  the queue.
- Association guard is keyed by address, not validator.
- The agent sends heartbeats from their own goroutine.
- Clear queue / delete: detach only floating IPs of unfinished rows (done,
  failed and occupied rows kept their fip_id and made a clear issue >1000
  sequential cloud calls: 256 s), at most 8 in parallel; the operation no
  longer dies with the client connection (10 minute limit).

Includes the incident analysis and the plan under analysis/ and
docs/changes/, and rebuilt bin/control-api and bin/validator-agent.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-10-02 14:42:12 +03:00
ayurishchevandClaude Sonnet 5.5 abbee9a08a Add self-check via control-api (self_check.methods)
control-api is hosted outside the cloud and validators reach it directly,
so it sees the floating IP as the connection's source address. New open
route GET /api/v1/agents/{id}/observed-ip returns that address (taken only
from the TCP peer; forwarding headers are ignored so a validator cannot
forge it).

The agent gets self_check.methods, a priority-ordered list of ip_echo
(unchanged) and control_api; the default stays [ip_echo]. The self-check
passes when any method confirms the address; the next method is tried on
no answer and on a mismatch. Each method has its own timeout so a hung
first method cannot starve the fallback, and control_api uses a new TCP
connection per call (a connection opened before the floating IP was
attached would keep reporting the old address).

Also: docker agent template/env, example config, docs, plan in
docs/changes, e2e script switch E2E_SELF_CHECK_METHODS, rebuilt
bin/control-api and bin/validator-agent.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-10-02 03:24:20 +03:00
ayurishchevandClaude Sonnet 5.5 aff8fe38b5 Scan floating IPs in the background, page by page, so thousands of addresses work
The "Scan Floating IP" button failed with a client timeout: the project now
holds ~6.4k floating IPs and the scan listed them all in one unpaginated,
timeout-less Neutron request on the HTTP request context.

openstack: ListFreeFloatingIPs reads marker-based pages (fields= keeps them
small) with per-page retry/backoff on transport errors, 5xx and 429, and every
request now has a timeout (also ends hangs inside the orchestrator tick).

orchestrator: the scan is a single-flight background job on the process
context with progress (clearing/listing/enqueuing/done/error), dry_run, full
discovery before anything is enqueued, then SubmitIPs in chunks of 500 in
ascending IP order; a failed read leaves the queue untouched. The auto-cycle
gets a "scanning" phase that polls the job, so the control loop and
autoCycleMu are never held across OpenStack/DB work; it recovers after a
restart and waits for (instead of adopting) a scan started by someone else.

db: migration 0009 (indexes), paged ListIPsPage/ListRegistryPage, GROUP BY
counters, EXISTS completion check, set-based ClearAllIPs.

API: POST /admin/ips/scan -> 202 (dry_run, wait), GET /admin/ips/scan, paging
and filters on /admin/ips and /admin/registry (bare arrays without limit),
results_by_overall in /admin/status.

dashboard: scan progress panel and dry-run button, paginated /ips and
/registry with server-side filters, Overview on counters and capped lists
with progress/ETA, "select all N by filter", hx-params fix for per-row
buttons, real counts in confirmations.

Also: docs (API, USAGE, DASHBOARD, README), plan and review under
docs/changes/, bin/ rebuilt with new SHA256SUMS.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-10-01 19:31:11 +03:00
ayurishchevandClaude Sonnet 5.5 debf2afed2 Add authentication: admin/agent bearer tokens for the API, login for the dashboard
control-api: every route now carries a mandatory access level (admin / agent /
open) in a route table. All /api/v1/admin/* require the admin token; the
write calls of validator-agent and prober (self-check, events, results,
complete) require a separate static agent token; register, heartbeat and
fetching the assignment stay open. Tokens come from env vars, are compared in
constant time and never logged. An empty token leaves that level open with a
startup warning (backward compatible).

validator-agent / prober: apiclient sends the agent token only to control-api.

admin-dashboard: login/password (from env) with a stateless HMAC session
cookie, Origin-based CSRF check, per-IP brute-force throttle, HX-Redirect for
htmx polls, logout in the sidebar; the dashboard calls control-api with the
admin token. Login page layout fixed after review.

Also: env plumbing in docker-compose/rxprod-compose/systemd/config examples,
e2e script with token assertions, tests, docs (API, SETUP, USAGE, DASHBOARD,
README), plan and review under docs/changes/, bin/ rebuilt with new
SHA256SUMS.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-10-01 11:35:24 +03:00
ayurishchevandClaude Sonnet 5.5 cd37b10f3b Add optional automatic check cycle (clear queue -> scan FIPs -> wait -> repeat)
An admin-controlled scenario that repeats what the operator does by hand:
clear the IP queue, scan and enqueue all free Floating IPs, wait until every
queued address reaches a terminal state (so results are in the Registry),
then wait a configurable interval and start over.

- control-api: new auto_cycle singleton table (migration 0008) holding
  enabled/interval/max-run settings and persisted phase state, so the cycle
  survives restarts; engine in orchestrator/autocycle.go driven from the
  existing loop tick with an injectable "now" for deterministic tests.
- Interval (default 1h, min 60s) and max wait (default unlimited, timeout
  outcome) are runtime settings, never hardcoded.
- The periodic fip_scan_interval_seconds scan is skipped while the cycle is
  enabled. An emptied queue mid-cycle counts as finished; stopping during
  the pause keeps the last cycle's outcome.
- API: GET/PUT /api/v1/admin/auto-cycle, POST .../start, POST .../stop.
- admin-dashboard: "Автоматический цикл" panel on /settings and an
  "Автоцикл активен" indicator on /overview.
- Tests for db, orchestrator, httpapi and dashboard; run-local-e2e.sh now
  exercises a full auto cycle; docs updated; bin/ rebuilt with refreshed
  SHA256SUMS.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-10-01 10:28:53 +03:00
ayurishchevandClaude Sonnet 5 582b44f314 Add Floating IP scanning and a durable address registry with configurable history depth
Adds POST /api/v1/admin/ips/scan (plus an optional periodic ticker) to
discover free Floating IPs in the OpenStack project and feed them straight
into the check queue. More importantly, decouples check/event history from
ip_queue's lifecycle: a new ip_registry table (migration 0007) gives every
address ever submitted a durable identity, so deleting it from the queue no
longer destroys its history — it's still reachable via the new
GET /api/v1/admin/registry[/{ip}] endpoints and the dashboard's /registry
pages, with retention depth configurable in check cycles per address
(history_retention_cycles, 0 = unlimited).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-23 09:52:01 +03:00
ayurishchevandClaude Sonnet 5 95f8066eed Retry validator-agent/prober registration; show validator hostname
Both binaries registered with control-api exactly once at startup and
exited (os.Exit(1)) on any failure — including control-api simply not
being up yet (no ordering guarantee between the two at boot/redeploy) or
the admin not having added this validator_id/site_id to the config yet.
Run() now retries registration with capped exponential backoff (3s->30s)
until it succeeds or the process is asked to shut down, instead of
crashing; registerWithRetry is identical in agentcore and probercore
since their Run/register shape already was.

Separately, the admin dashboard's Validators page had no hostname column
even though the agent already reports one on register (mirroring the
prober) and control-api already persists it — only the admin-config read
DTO (validatorDTO in httpapi and dashboard) dropped it before it reached
the template. Added hostname + last_heartbeat_at to that DTO end-to-end
and a Хост/Heartbeat column to validators.html, matching sites.html.

Rebuilt bin/{control-api,admin-dashboard,prober,validator-agent} and
bin/SHA256SUMS per docs/SETUP.md's documented build recipe.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-18 10:43:00 +03:00
ayurishchev 550ce3fec4 Enable SSH and TLS handshake on prober for TCP22 and TCP443 ports 2026-08-26 23:52:31 +03:00
ayurishchev ef24cc9858 New external site management. New external prober heartbeat feature. 2026-08-26 20:47:54 +03:00
ayurishchev 42f584dd3c Manage external site-prober actions va API and Dashboard 2026-08-26 19:45:48 +03:00
ayurishchev f22ad569b6 added FIP Cooldown before start 2026-08-24 10:29:08 +03:00
ayurishchev f8336740ad feature: handle delete operation for IPs 2026-08-23 22:24:55 +03:00
ayurishchev 37910e410b admin control features and admin dashboard 2026-08-23 20:39:22 +03:00
ayurishchev 67adc6670e fix public ip selfcheck logic 2026-08-21 11:04:49 +03:00
ayurishchev 7e44db87b2 repo init 2026-08-21 07:34:45 +03:00