Files
ayurishchevandClaude Sonnet 5.5 aff8fe38b5 Scan floating IPs in the background, page by page, so thousands of addresses work
The "Scan Floating IP" button failed with a client timeout: the project now
holds ~6.4k floating IPs and the scan listed them all in one unpaginated,
timeout-less Neutron request on the HTTP request context.

openstack: ListFreeFloatingIPs reads marker-based pages (fields= keeps them
small) with per-page retry/backoff on transport errors, 5xx and 429, and every
request now has a timeout (also ends hangs inside the orchestrator tick).

orchestrator: the scan is a single-flight background job on the process
context with progress (clearing/listing/enqueuing/done/error), dry_run, full
discovery before anything is enqueued, then SubmitIPs in chunks of 500 in
ascending IP order; a failed read leaves the queue untouched. The auto-cycle
gets a "scanning" phase that polls the job, so the control loop and
autoCycleMu are never held across OpenStack/DB work; it recovers after a
restart and waits for (instead of adopting) a scan started by someone else.

db: migration 0009 (indexes), paged ListIPsPage/ListRegistryPage, GROUP BY
counters, EXISTS completion check, set-based ClearAllIPs.

API: POST /admin/ips/scan -> 202 (dry_run, wait), GET /admin/ips/scan, paging
and filters on /admin/ips and /admin/registry (bare arrays without limit),
results_by_overall in /admin/status.

dashboard: scan progress panel and dry-run button, paginated /ips and
/registry with server-side filters, Overview on counters and capped lists
with progress/ETA, "select all N by filter", hx-params fix for per-row
buttons, real counts in confirmations.

Also: docs (API, USAGE, DASHBOARD, README), plan and review under
docs/changes/, bin/ rebuilt with new SHA256SUMS.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-10-01 19:31:11 +03:00

5.3 KiB

Local end-to-end smoke test

scripts/run-local-e2e.sh runs the full system as local processes with no real OpenStack cloud and no real internet access:

  • control-api with openstack.mode: mock — the in-memory openstack.MockClient stands in for Neutron, pre-seeded with one synthetic floating IP per configured address (see cmd/control-api/main.go's newOpenStackClient).
  • 1 validator-agent (validator_01), started with -stub-ports 22022,28081,28443,28888 — trivial accept-and-close TCP listeners standing in for the base-minimum services (22/80/443/8080) a real validator would run. ICMP needs no stub: the kernel answers echo requests to any local address (127.0.0.0/8) on its own.
  • 3 probers (site-1/site-2/site-3), all probing 127.0.0.1.
  • 3 httpstub instances (scripts/httpstub) — / always returns 200 OK, standing in for the real outbound egress targets (hub.docker.com / github.com / packages.ubuntu.com); /ip echoes the caller's remote address as plain text, standing in for the external IP-echo service the validator-agent's self-check normally queries in production (see internal/agentcore.detectPublicIP and config.SelfCheckCfg.IPEchoURLs).

The generated config uses a short lease_ttl_seconds: 8 and checking_window_seconds: 15 so the whole run finishes in well under a minute instead of using the (much longer) production defaults.

Why only one IP address (127.0.0.1), and why it must be exactly that one: self-check compares the source address the validator-agent's own request to the IP-echo target arrives with against the address it was just assigned. In a real deployment that target is a real external IP-echo service, and Neutron actually SNATs the validator's egress traffic through whichever floating IP is attached, so any configured address self-checks correctly. This offline harness points self_check.ip_echo_urls at the local httpstub's /ip route instead (see validator-agent.yaml generated by the script) — there's no real network-level SNAT here, the agent's traffic to that stub always really originates from 127.0.0.1, so only that literal loopback address can ever pass self-check in this harness. This is a limitation of the harness's fidelity, not of the self-check mechanism itself.

Running it

scripts/run-local-e2e.sh

It will:

  1. Build all three binaries plus httpstub into a temp workdir.
  2. Start the 3 stub HTTP targets, control-api, the validator-agent, and the 3 probers.
  3. Wait for GET /healthz to come up.
  4. A few seconds in, kill the validator-agent mid-run and wait past the 8s lease TTL, to demonstrate that control-api's lease sweep reclaims the in-flight IP (moves it back to queued, bumps retry_count) without any special crash-recovery code — it's the same sweep that runs every tick. It then restarts the validator-agent so the queue can finish.
  5. Poll GET /api/v1/admin/status until every configured IP has reached a terminal state (done or failed).
  6. Print the final /api/v1/admin/status and /api/v1/admin/ips output.
  7. Force a re-check of the finished address via POST /api/v1/admin/ips and wait for it to drain (attempt_number advances).
  8. Exercise the automatic cycle (/api/v1/admin/auto-cycle): set the smallest allowed interval (60s) and a 120s run limit, start it, and wait for the first cycle to finish. The script then asserts that the outcome is completed, the phase is waiting with runs_total=1, and that the registry's total_cycles for 127.0.0.1 grew (the cycle cleared the queue, re-scanned the mock floating IP and re-checked it). The cycle now passes through the background scan (scanning phase) before the checks run. Finally it stops the cycle and asserts it is idle. The second cycle (the interval wait) is covered by unit tests, so the script does not sit through the 60s pause. The script exits non-zero if any assertion fails.

Expect to see 127.0.0.1 end with "state":"done" and "overall_result":"pass" (all egress checks against the stub targets succeed, and all 3 probers can reach the stub TCP listeners and get ICMP replies from loopback).

Inspecting a run

The workdir (printed at the end, /tmp/cloud-ip-validator-e2e.XXXXXX) is not deleted automatically, so you can inspect:

  • logs/control-api.log, logs/validator-agent.log, logs/prober-site-*.log
  • control-api.db — open with sqlite3 to inspect the checks and events tables directly, e.g.:
    sqlite3 /tmp/cloud-ip-validator-e2e.XXXXXX/control-api.db \
      "select ip_address, source, check_type, success from checks order by id"
    

What this does not cover

This harness proves the orchestration, HTTP protocol, and check-running logic all work together correctly. It does not exercise the real internal/openstack/client.go (gophercloud) path — that only runs against openstack.mode: real with actual OpenStack credentials. That path has its own read-only smoke test, internal/openstack/client_live_test.go, skipped by default and gated behind OPENSTACK_LIVE_TEST=1:

OPENSTACK_LIVE_TEST=1 \
OS_AUTH_URL=https://keystone.example:5000/v3 \
OS_TOKEN=... \
OS_PROJECT_ID=... \
OS_TEST_FLOATING_IP=203.0.113.10 \
go test ./internal/openstack/... -run TestClientLive -v