The "Scan Floating IP" button failed with a client timeout: the project now holds ~6.4k floating IPs and the scan listed them all in one unpaginated, timeout-less Neutron request on the HTTP request context. openstack: ListFreeFloatingIPs reads marker-based pages (fields= keeps them small) with per-page retry/backoff on transport errors, 5xx and 429, and every request now has a timeout (also ends hangs inside the orchestrator tick). orchestrator: the scan is a single-flight background job on the process context with progress (clearing/listing/enqueuing/done/error), dry_run, full discovery before anything is enqueued, then SubmitIPs in chunks of 500 in ascending IP order; a failed read leaves the queue untouched. The auto-cycle gets a "scanning" phase that polls the job, so the control loop and autoCycleMu are never held across OpenStack/DB work; it recovers after a restart and waits for (instead of adopting) a scan started by someone else. db: migration 0009 (indexes), paged ListIPsPage/ListRegistryPage, GROUP BY counters, EXISTS completion check, set-based ClearAllIPs. API: POST /admin/ips/scan -> 202 (dry_run, wait), GET /admin/ips/scan, paging and filters on /admin/ips and /admin/registry (bare arrays without limit), results_by_overall in /admin/status. dashboard: scan progress panel and dry-run button, paginated /ips and /registry with server-side filters, Overview on counters and capped lists with progress/ETA, "select all N by filter", hx-params fix for per-row buttons, real counts in confirmations. Also: docs (API, USAGE, DASHBOARD, README), plan and review under docs/changes/, bin/ rebuilt with new SHA256SUMS. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
5.3 KiB
Local end-to-end smoke test
scripts/run-local-e2e.sh runs the full system as local processes with no
real OpenStack cloud and no real internet access:
- control-api with
openstack.mode: mock— the in-memoryopenstack.MockClientstands in for Neutron, pre-seeded with one synthetic floating IP per configured address (seecmd/control-api/main.go'snewOpenStackClient). - 1 validator-agent (
validator_01), started with-stub-ports 22022,28081,28443,28888— trivial accept-and-close TCP listeners standing in for the base-minimum services (22/80/443/8080) a real validator would run. ICMP needs no stub: the kernel answers echo requests to any local address (127.0.0.0/8) on its own. - 3 probers (
site-1/site-2/site-3), all probing127.0.0.1. - 3
httpstubinstances (scripts/httpstub) —/always returns200 OK, standing in for the real outbound egress targets (hub.docker.com / github.com / packages.ubuntu.com);/ipechoes the caller's remote address as plain text, standing in for the external IP-echo service the validator-agent's self-check normally queries in production (seeinternal/agentcore.detectPublicIPandconfig.SelfCheckCfg.IPEchoURLs).
The generated config uses a short lease_ttl_seconds: 8 and
checking_window_seconds: 15 so the whole run finishes in well under a
minute instead of using the (much longer) production defaults.
Why only one IP address (127.0.0.1), and why it must be exactly that
one: self-check compares the source address the validator-agent's own
request to the IP-echo target arrives with against the address it was
just assigned. In a real deployment that target is a real external
IP-echo service, and Neutron actually SNATs the validator's egress
traffic through whichever floating IP is attached, so any configured
address self-checks correctly. This offline harness points
self_check.ip_echo_urls at the local httpstub's /ip route instead
(see validator-agent.yaml generated by the script) — there's no real
network-level SNAT here, the agent's traffic to that stub always really
originates from 127.0.0.1, so only that literal loopback address can
ever pass self-check in this harness. This is a limitation of the
harness's fidelity, not of the self-check mechanism itself.
Running it
scripts/run-local-e2e.sh
It will:
- Build all three binaries plus
httpstubinto a temp workdir. - Start the 3 stub HTTP targets, control-api, the validator-agent, and the 3 probers.
- Wait for
GET /healthzto come up. - A few seconds in, kill the validator-agent mid-run and wait past the
8s lease TTL, to demonstrate that control-api's lease sweep reclaims the
in-flight IP (moves it back to
queued, bumpsretry_count) without any special crash-recovery code — it's the same sweep that runs every tick. It then restarts the validator-agent so the queue can finish. - Poll
GET /api/v1/admin/statusuntil every configured IP has reached a terminal state (doneorfailed). - Print the final
/api/v1/admin/statusand/api/v1/admin/ipsoutput. - Force a re-check of the finished address via
POST /api/v1/admin/ipsand wait for it to drain (attempt_numberadvances). - Exercise the automatic cycle (
/api/v1/admin/auto-cycle): set the smallest allowed interval (60s) and a 120s run limit,startit, and wait for the first cycle to finish. The script then asserts that the outcome iscompleted, the phase iswaitingwithruns_total=1, and that the registry'stotal_cyclesfor127.0.0.1grew (the cycle cleared the queue, re-scanned the mock floating IP and re-checked it). The cycle now passes through the background scan (scanningphase) before the checks run. Finally itstops the cycle and asserts it isidle. The second cycle (the interval wait) is covered by unit tests, so the script does not sit through the 60s pause. The script exits non-zero if any assertion fails.
Expect to see 127.0.0.1 end with "state":"done" and
"overall_result":"pass" (all egress checks against the stub targets
succeed, and all 3 probers can reach the stub TCP listeners and get ICMP
replies from loopback).
Inspecting a run
The workdir (printed at the end, /tmp/cloud-ip-validator-e2e.XXXXXX) is
not deleted automatically, so you can inspect:
logs/control-api.log,logs/validator-agent.log,logs/prober-site-*.logcontrol-api.db— open withsqlite3to inspect thechecksandeventstables directly, e.g.:sqlite3 /tmp/cloud-ip-validator-e2e.XXXXXX/control-api.db \ "select ip_address, source, check_type, success from checks order by id"
What this does not cover
This harness proves the orchestration, HTTP protocol, and check-running
logic all work together correctly. It does not exercise the real
internal/openstack/client.go (gophercloud) path — that only runs against
openstack.mode: real with actual OpenStack credentials. That path has its
own read-only smoke test, internal/openstack/client_live_test.go, skipped
by default and gated behind OPENSTACK_LIVE_TEST=1:
OPENSTACK_LIVE_TEST=1 \
OS_AUTH_URL=https://keystone.example:5000/v3 \
OS_TOKEN=... \
OS_PROJECT_ID=... \
OS_TEST_FLOATING_IP=203.0.113.10 \
go test ./internal/openstack/... -run TestClientLive -v