Files
ayurishchevandClaude Sonnet 5.5 aff8fe38b5 Scan floating IPs in the background, page by page, so thousands of addresses work
The "Scan Floating IP" button failed with a client timeout: the project now
holds ~6.4k floating IPs and the scan listed them all in one unpaginated,
timeout-less Neutron request on the HTTP request context.

openstack: ListFreeFloatingIPs reads marker-based pages (fields= keeps them
small) with per-page retry/backoff on transport errors, 5xx and 429, and every
request now has a timeout (also ends hangs inside the orchestrator tick).

orchestrator: the scan is a single-flight background job on the process
context with progress (clearing/listing/enqueuing/done/error), dry_run, full
discovery before anything is enqueued, then SubmitIPs in chunks of 500 in
ascending IP order; a failed read leaves the queue untouched. The auto-cycle
gets a "scanning" phase that polls the job, so the control loop and
autoCycleMu are never held across OpenStack/DB work; it recovers after a
restart and waits for (instead of adopting) a scan started by someone else.

db: migration 0009 (indexes), paged ListIPsPage/ListRegistryPage, GROUP BY
counters, EXISTS completion check, set-based ClearAllIPs.

API: POST /admin/ips/scan -> 202 (dry_run, wait), GET /admin/ips/scan, paging
and filters on /admin/ips and /admin/registry (bare arrays without limit),
results_by_overall in /admin/status.

dashboard: scan progress panel and dry-run button, paginated /ips and
/registry with server-side filters, Overview on counters and capped lists
with progress/ETA, "select all N by filter", hx-params fix for per-row
buttons, real counts in confirmations.

Also: docs (API, USAGE, DASHBOARD, README), plan and review under
docs/changes/, bin/ rebuilt with new SHA256SUMS.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-10-01 19:31:11 +03:00

109 lines
5.3 KiB
Markdown

# Local end-to-end smoke test
`scripts/run-local-e2e.sh` runs the full system as local processes with no
real OpenStack cloud and no real internet access:
- **control-api** with `openstack.mode: mock` — the in-memory
`openstack.MockClient` stands in for Neutron, pre-seeded with one
synthetic floating IP per configured address (see
`cmd/control-api/main.go`'s `newOpenStackClient`).
- **1 validator-agent** (`validator_01`), started with
`-stub-ports 22022,28081,28443,28888` — trivial accept-and-close TCP
listeners standing in for the base-minimum services (22/80/443/8080) a
real validator would run. ICMP needs no stub: the kernel answers echo
requests to any local address (127.0.0.0/8) on its own.
- **3 probers** (`site-1`/`site-2`/`site-3`), all probing `127.0.0.1`.
- **3 `httpstub` instances** (`scripts/httpstub`) — `/` always returns
`200 OK`, standing in for the real outbound egress targets (hub.docker.com
/ github.com / packages.ubuntu.com); `/ip` echoes the caller's remote
address as plain text, standing in for the external IP-echo service the
validator-agent's self-check normally queries in production (see
`internal/agentcore.detectPublicIP` and `config.SelfCheckCfg.IPEchoURLs`).
The generated config uses a short `lease_ttl_seconds: 8` and
`checking_window_seconds: 15` so the whole run finishes in well under a
minute instead of using the (much longer) production defaults.
**Why only one IP address (`127.0.0.1`), and why it must be exactly that
one:** self-check compares the source address the validator-agent's own
request to the IP-echo target arrives with against the address it was
just assigned. In a real deployment that target is a real external
IP-echo service, and Neutron actually SNATs the validator's egress
traffic through whichever floating IP is attached, so any configured
address self-checks correctly. This offline harness points
`self_check.ip_echo_urls` at the local `httpstub`'s `/ip` route instead
(see `validator-agent.yaml` generated by the script) — there's no real
network-level SNAT here, the agent's traffic to that stub always really
originates from `127.0.0.1`, so only that literal loopback address can
ever pass self-check in this harness. This is a limitation of the
harness's fidelity, not of the self-check mechanism itself.
## Running it
```
scripts/run-local-e2e.sh
```
It will:
1. Build all three binaries plus `httpstub` into a temp workdir.
2. Start the 3 stub HTTP targets, control-api, the validator-agent, and the
3 probers.
3. Wait for `GET /healthz` to come up.
4. A few seconds in, **kill the validator-agent mid-run** and wait past the
8s lease TTL, to demonstrate that control-api's lease sweep reclaims the
in-flight IP (moves it back to `queued`, bumps `retry_count`) without
any special crash-recovery code — it's the same sweep that runs every
tick. It then restarts the validator-agent so the queue can finish.
5. Poll `GET /api/v1/admin/status` until every configured IP has reached a
terminal state (`done` or `failed`).
6. Print the final `/api/v1/admin/status` and `/api/v1/admin/ips` output.
7. Force a re-check of the finished address via `POST /api/v1/admin/ips`
and wait for it to drain (`attempt_number` advances).
8. Exercise the **automatic cycle** (`/api/v1/admin/auto-cycle`): set the
smallest allowed interval (60s) and a 120s run limit, `start` it, and
wait for the first cycle to finish. The script then asserts that the
outcome is `completed`, the phase is `waiting` with `runs_total=1`, and
that the registry's `total_cycles` for `127.0.0.1` grew (the cycle
cleared the queue, re-scanned the mock floating IP and re-checked it).
The cycle now passes through the background scan (`scanning` phase) before the checks run. Finally it `stop`s the cycle and asserts it is `idle`. The second cycle
(the interval wait) is covered by unit tests, so the script does not
sit through the 60s pause. The script exits non-zero if any assertion
fails.
Expect to see `127.0.0.1` end with `"state":"done"` and
`"overall_result":"pass"` (all egress checks against the stub targets
succeed, and all 3 probers can reach the stub TCP listeners and get ICMP
replies from loopback).
## Inspecting a run
The workdir (printed at the end, `/tmp/cloud-ip-validator-e2e.XXXXXX`) is
**not** deleted automatically, so you can inspect:
- `logs/control-api.log`, `logs/validator-agent.log`, `logs/prober-site-*.log`
- `control-api.db` — open with `sqlite3` to inspect the `checks` and
`events` tables directly, e.g.:
```
sqlite3 /tmp/cloud-ip-validator-e2e.XXXXXX/control-api.db \
"select ip_address, source, check_type, success from checks order by id"
```
## What this does *not* cover
This harness proves the orchestration, HTTP protocol, and check-running
logic all work together correctly. It does **not** exercise the real
`internal/openstack/client.go` (gophercloud) path — that only runs against
`openstack.mode: real` with actual OpenStack credentials. That path has its
own read-only smoke test, `internal/openstack/client_live_test.go`, skipped
by default and gated behind `OPENSTACK_LIVE_TEST=1`:
```
OPENSTACK_LIVE_TEST=1 \
OS_AUTH_URL=https://keystone.example:5000/v3 \
OS_TOKEN=... \
OS_PROJECT_ID=... \
OS_TEST_FLOATING_IP=203.0.113.10 \
go test ./internal/openstack/... -run TestClientLive -v
```