Files
cloud-ip-validator/docs/LOCAL_E2E.md
T
2026-08-21 07:34:45 +03:00

4.0 KiB

Local end-to-end smoke test

scripts/run-local-e2e.sh runs the full system as local processes with no real OpenStack cloud and no real internet access:

  • control-api with openstack.mode: mock — the in-memory openstack.MockClient stands in for Neutron, pre-seeded with one synthetic floating IP per configured address (see cmd/control-api/main.go's newOpenStackClient).
  • 1 validator-agent (validator_01), started with -stub-ports 12022,18081,18443,18888 — trivial accept-and-close TCP listeners standing in for the base-minimum services (22/80/443/8080) a real validator would run. ICMP needs no stub: the kernel answers echo requests to any local address (127.0.0.0/8) on its own.
  • 3 probers (site-1/site-2/site-3), all probing 127.0.0.1.
  • 3 httpstub instances (scripts/httpstub) standing in for the real outbound targets (hub.docker.com / github.com / packages.ubuntu.com), always returning 200 OK.

The generated config uses a short lease_ttl_seconds: 8 and checking_window_seconds: 15 so the whole run finishes in well under a minute instead of using the (much longer) production defaults.

Why only one IP address (127.0.0.1), and why it must be exactly that one: the self-check mechanism (GET /whatsmyip, see internal/httpapi/server.go's remoteIP) compares the TCP source address the validator-agent's own outbound connection to control-api arrives with against the address it was just assigned. In a real deployment, Neutron actually SNATs the validator's egress traffic through whichever floating IP is attached, so any configured address self-checks correctly. This offline harness has no real network-level SNAT — the agent's traffic to control-api always really originates from 127.0.0.1 — so only that literal loopback address can ever pass self-check here. This is a limitation of the harness's fidelity, not of the self-check mechanism itself.

Running it

scripts/run-local-e2e.sh

It will:

  1. Build all three binaries plus httpstub into a temp workdir.
  2. Start the 3 stub HTTP targets, control-api, the validator-agent, and the 3 probers.
  3. Wait for GET /healthz to come up.
  4. A few seconds in, kill the validator-agent mid-run and wait past the 8s lease TTL, to demonstrate that control-api's lease sweep reclaims the in-flight IP (moves it back to queued, bumps retry_count) without any special crash-recovery code — it's the same sweep that runs every tick. It then restarts the validator-agent so the queue can finish.
  5. Poll GET /api/v1/admin/status until every configured IP has reached a terminal state (done or failed).
  6. Print the final /api/v1/admin/status and /api/v1/admin/ips output.

Expect to see 127.0.0.1 end with "state":"done" and "overall_result":"pass" (all egress checks against the stub targets succeed, and all 3 probers can reach the stub TCP listeners and get ICMP replies from loopback).

Inspecting a run

The workdir (printed at the end, /tmp/cloud-ip-validator-e2e.XXXXXX) is not deleted automatically, so you can inspect:

  • logs/control-api.log, logs/validator-agent.log, logs/prober-site-*.log
  • control-api.db — open with sqlite3 to inspect the checks and events tables directly, e.g.:
    sqlite3 /tmp/cloud-ip-validator-e2e.XXXXXX/control-api.db \
      "select ip_address, source, check_type, success from checks order by id"
    

What this does not cover

This harness proves the orchestration, HTTP protocol, and check-running logic all work together correctly. It does not exercise the real internal/openstack/client.go (gophercloud) path — that only runs against openstack.mode: real with actual OpenStack credentials. That path has its own read-only smoke test, internal/openstack/client_live_test.go, skipped by default and gated behind OPENSTACK_LIVE_TEST=1:

OPENSTACK_LIVE_TEST=1 \
OS_AUTH_URL=https://keystone.example:5000/v3 \
OS_TOKEN=... \
OS_PROJECT_ID=... \
OS_TEST_FLOATING_IP=203.0.113.10 \
go test ./internal/openstack/... -run TestClientLive -v