Scan floating IPs in the background, page by page, so thousands of addresses work
The "Scan Floating IP" button failed with a client timeout: the project now holds ~6.4k floating IPs and the scan listed them all in one unpaginated, timeout-less Neutron request on the HTTP request context. openstack: ListFreeFloatingIPs reads marker-based pages (fields= keeps them small) with per-page retry/backoff on transport errors, 5xx and 429, and every request now has a timeout (also ends hangs inside the orchestrator tick). orchestrator: the scan is a single-flight background job on the process context with progress (clearing/listing/enqueuing/done/error), dry_run, full discovery before anything is enqueued, then SubmitIPs in chunks of 500 in ascending IP order; a failed read leaves the queue untouched. The auto-cycle gets a "scanning" phase that polls the job, so the control loop and autoCycleMu are never held across OpenStack/DB work; it recovers after a restart and waits for (instead of adopting) a scan started by someone else. db: migration 0009 (indexes), paged ListIPsPage/ListRegistryPage, GROUP BY counters, EXISTS completion check, set-based ClearAllIPs. API: POST /admin/ips/scan -> 202 (dry_run, wait), GET /admin/ips/scan, paging and filters on /admin/ips and /admin/registry (bare arrays without limit), results_by_overall in /admin/status. dashboard: scan progress panel and dry-run button, paginated /ips and /registry with server-side filters, Overview on counters and capped lists with progress/ETA, "select all N by filter", hx-params fix for per-row buttons, real counts in confirmations. Also: docs (API, USAGE, DASHBOARD, README), plan and review under docs/changes/, bin/ rebuilt with new SHA256SUMS. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
This commit is contained in:
1 parent
debf2afed2
commit
aff8fe38b5
61 files changed
+5802
-505
No files matched your search
@@ -38,6 +38,8 @@ docker compose up -d --build # весь стенд на одной
|
||||
| `orchestrator.heartbeat_timeout_seconds` | После скольких секунд тишины валидатор или площадка считаются потерянными (30) |
|
||||
| `orchestrator.fip_settle_seconds` | Только начальное значение: пауза между привязкой FIP и self-check; далее управляется на лету. Должно выполняться `fip_settle_seconds + self_check_timeout_seconds < lease_ttl_seconds` |
|
||||
| `orchestrator.fip_scan_interval_seconds` | Периодический скан Floating IP (0 — выключен). Не выполняется, пока включён автоматический цикл |
|
||||
| `orchestrator.fip_scan_timeout_seconds` | Предел всего сканирования Floating IP (1800) |
|
||||
| `openstack.list_page_size`, `list_page_retries`, `request_timeout_seconds` | Скан читает Floating IP страницами (200), повторяя страницу при обрыве/5xx/429 (5 раз); таймаут одного запроса к OpenStack (60 с) |
|
||||
| `auth.admin_token_env`, `auth.agent_token_env` | Имена переменных окружения с токеном администратора (`CONTROL_API_ADMIN_TOKEN`) и токеном агентов (`CONTROL_API_AGENT_TOKEN`). Значения в YAML не хранятся; пустой токен — соответствующий уровень API открыт (с предупреждением в логе) |
|
||||
| `aggregation.missing_counts_as_fail` | Отсутствие ответа источника засчитывается как провал (`true`) |
|
||||
| `validators`, `sites`, `check_types`, `targets`, `inbound_checks` | Начальная загрузка пустой БД: валидаторы (`validator_id` + `os_port_id`), внешние площадки, типы и цели egress-проверок, порты inbound-проверок. Дальше источник истины — БД, правки через API/UI |
|
||||
@@ -103,22 +105,23 @@ docs/ документация и планы доработок
|
||||
- История проверок хранится по циклам; `history_retention_cycles` ограничивает глубину (0 — без ограничения), сама запись реестра остаётся.
|
||||
|
||||
**Автоматический цикл** (опционально, по умолчанию выключен)
|
||||
- Один цикл: очистить очередь → просканировать Floating IP → дождаться, пока все адреса станут терминальными → пауза `interval_seconds` → заново. Пауза считается от завершения цикла.
|
||||
- Один цикл: очистить очередь → просканировать Floating IP (фаза `scanning`, фоновое задание) → дождаться, пока все адреса станут терминальными → пауза `interval_seconds` → заново. Пауза считается от завершения цикла. Сканирование не блокирует оркестратор; чужое идущее сканирование цикл дожидается, а не присоединяется к нему.
|
||||
- `interval_seconds` — по умолчанию 3600, минимум 60; `max_run_seconds` — максимальное ожидание проверок (0 — без лимита), по истечении исход `timeout`.
|
||||
- Исходы цикла: `completed`, `no_free_ips`, `timeout`, `error`, `stopped`. Опустевшая очередь посреди цикла считается завершением.
|
||||
- Состояние хранится в БД и переживает перезапуск; выключение не прерывает идущие проверки. Подробности — [docs/USAGE.md](docs/USAGE.md#автоматический-цикл-проверок).
|
||||
|
||||
**Удаление и сканирование**
|
||||
- Удаление (точечное, списком, «очистить всё») убирает строку очереди, но не историю в реестре.
|
||||
- Скан Floating IP ставит в очередь только свободные адреса (не привязанные ни к одному порту); уже идущие проверки не трогаются.
|
||||
- Скан Floating IP ставит в очередь только свободные адреса (не привязанные ни к одному порту); уже идущие проверки не трогаются. Он идёт **в фоне и читает облако страницами**: подходит и для тысяч адресов (на стенде 6441 Floating IP читаются ≈ 1,5–2 мин). Адреса ставятся в очередь только после полного обнаружения (кусками по 500, по возрастанию IP); при сбое чтения очередь не меняется. `dry_run=true` / «Пробное сканирование» считает адреса, не меняя очередь.
|
||||
- Пропускная способность: ≈ 50 с на адрес на валидатор — очередь из 6440 адресов это ≈ 18 ч на 5 валидаторах, ≈ 9 ч на 10 (рычаги: число валидаторов и `fip_settle_seconds`).
|
||||
|
||||
## API (`/api/v1`)
|
||||
| Область | Эндпоинты |
|
||||
|---|---|
|
||||
| Валидатор | `POST /agents/register`, `POST /agents/{id}/heartbeat`, `GET /agents/{id}/assignment`, `POST /agents/{id}/self-check\|events\|results\|complete` |
|
||||
| Пробер | `POST /probers/register`, `POST /probers/{site_id}/heartbeat`, `GET /probers/{site_id}/assignments`, `POST /probers/{site_id}/results` |
|
||||
| Очередь | `GET /admin/status`, `GET\|POST /admin/ips`, `GET /admin/ips/{ip}`, `POST /admin/ips/{ip}/cancel`, `DELETE /admin/ips/{ip}`, `POST /admin/ips/delete\|clear\|scan` |
|
||||
| Реестр | `GET /admin/registry`, `GET /admin/registry/{ip}` |
|
||||
| Очередь | `GET /admin/status`, `GET\|POST /admin/ips` (`limit/offset/state/q/result/order` — постранично), `GET /admin/ips/{ip}`, `POST /admin/ips/{ip}/cancel`, `DELETE /admin/ips/{ip}`, `POST /admin/ips/delete\|clear`, `POST\|GET /admin/ips/scan` (фоновый скан: `202`, `dry_run`, `wait`; статус и прогресс) |
|
||||
| Реестр | `GET /admin/registry` (`limit/offset/q/last_result` — постранично), `GET /admin/registry/{ip}` |
|
||||
| Автоцикл | `GET\|PUT /admin/auto-cycle`, `POST /admin/auto-cycle/start\|stop` |
|
||||
| Конфигурация | `/admin/config/validators`, `/sites`, `/targets`, `/check-types`, `GET\|PUT /admin/config/orchestrator`, `GET\|PUT /admin/config/inbound-checks` |
|
||||
| Служебное | `GET /admin/validators`, `GET /healthz` |
|
||||
@@ -155,9 +158,9 @@ docs/ документация и планы доработок
|
||||
Страницы `admin-dashboard` (подробно — [docs/DASHBOARD.md](docs/DASHBOARD.md)):
|
||||
| Страница | Назначение |
|
||||
|---|---|
|
||||
| `/overview` | Счётчики по состояниям, «текущая» и «последние завершённые» проверки, поиск по IP и фильтр по статусу, индикатор автоцикла; обновляется без перезагрузки |
|
||||
| `/ips`, `/ips/{ip}` | Очередь: добавление адресов, «Сканировать Floating IP», перепроверка, отмена, удаление (в том числе списком и «Очистить всё»); детали и события адреса |
|
||||
| `/registry`, `/registry/{ip}` | Реестр всех адресов и полная история проверок адреса; поиск и фильтр сохраняются в адресной строке |
|
||||
| `/overview` | Счётчики и прогресс («Готово D из T», оценка времени), «в работе», «в очереди: Q», «последние завершённые», поиск по IP и фильтр по статусу, индикатор скана и автоцикла; работает на счётчиках и ограниченных списках, поэтому быстрый и при тысячах адресов |
|
||||
| `/ips`, `/ips/{ip}` | Очередь **постранично** с поиском и фильтром на сервере: добавление адресов, «Сканировать Floating IP» (панель прогресса) и «Пробное сканирование», перепроверка, отмена, удаление (страница или «все N по фильтру», «Очистить всё»); детали и события адреса |
|
||||
| `/registry`, `/registry/{ip}` | Реестр всех адресов (постранично) и полная история проверок адреса; поиск, фильтр и страница сохраняются в адресной строке |
|
||||
| `/validators`, `/sites`, `/targets`, `/check-types` | Управление валидаторами, внешними площадками, группами целей и типами проверок |
|
||||
| `/settings` | Панель «Автоматический цикл», пауза перед self-check, глубина истории, TCP-порты и ICMP для inbound-проверок |
|
||||
- Порядок блоков на `/overview` фиксирован: статистика → фильтр → таблицы; поллится только блок таблиц, поэтому набранный в фильтре текст не сбрасывается.
|
||||
@@ -181,7 +184,7 @@ scripts/run-local-e2e.sh # сквозной прог
|
||||
| [docs/DASHBOARD.md](docs/DASHBOARD.md) | Устройство `admin-dashboard`: страницы, поиск и фильтр, обработка ошибок |
|
||||
| [docs/DIAGRAMS.md](docs/DIAGRAMS.md) | Диаграммы потоков данных: control plane, egress-проверка, телеметрия |
|
||||
| [docs/LOCAL_E2E.md](docs/LOCAL_E2E.md) | Полностью офлайн-прогон всей системы одним скриптом |
|
||||
| [docs/changes/](docs/changes/) | Планы доработок и отчёты ревью с отметкой времени в имени файла (последняя: [аутентификация](docs/changes/2026-10-01_11-31_authentication-review.md)) |
|
||||
| [docs/changes/](docs/changes/) | Планы доработок и отчёты ревью с отметкой времени в имени файла (последняя: [скан при тысячах адресов](docs/changes/2026-10-01_18-59_fip-scan-at-scale-review.md)) |
|
||||
| [docs/CONTROL_DATA_PLANE.html](docs/CONTROL_DATA_PLANE.html) | Презентационные схемы control/data plane для docker-compose-деплоя — открыть в браузере |
|
||||
|
||||
## История изменений
|
||||
@@ -197,6 +200,7 @@ scripts/run-local-e2e.sh # сквозной прог
|
||||
|
||||
| Дата | Веха | Документ |
|
||||
|---|---|---|
|
||||
| 2026-10-01 | Скан Floating IP при тысячах адресов: фоновый постраничный скан с прогрессом, фаза `scanning` в автоцикле, постраничные `/ips` и `/registry`, «Обзор» на счётчиках | [план](docs/changes/2026-10-01_18-19_fip-scan-at-scale-plan.md) · [ревью и тесты](docs/changes/2026-10-01_18-59_fip-scan-at-scale-review.md) · [USAGE](docs/USAGE.md#сканирование-floating-ip-из-openstack) · [API](docs/API.md#post-apiv1adminipsscan) |
|
||||
| 2026-10-01 | Аутентификация: токены администратора и агентов для API, логин и пароль для дашборда | [план](docs/changes/2026-10-01_11-12_authentication-plan.md) · [ревью и тесты](docs/changes/2026-10-01_11-31_authentication-review.md) · [API](docs/API.md#аутентификация) |
|
||||
| 2026-10-01 | Автоматический цикл проверок по сценарию: очистка → скан FIP → проверка → пауза | [USAGE](docs/USAGE.md#автоматический-цикл-проверок) · [API](docs/API.md#автоматический-цикл-проверок) |
|
||||
| 2026-09-23 | Сканирование Floating IP и устойчивый реестр адресов с настраиваемой глубиной истории | [USAGE](docs/USAGE.md#реестр-адресов-и-глубина-истории) |
|
||||
|
||||
+4
-4
@@ -1,4 +1,4 @@
|
||||
a0eac5409a2929d641ef2a217c31f1b6a974a8679da866ebbb50ff1d5b57df8d control-api
|
||||
5d7fb2f476871843ae4179dddb61e77d3f71e8eb73a00e331f0119d8893ba813 validator-agent
|
||||
23c687a988350c61f486d400af654bd563865fff59421c437eb679e0e4e10500 prober
|
||||
90921ad25f16a922124368770e439d1228c5b0bf7126ba6c726f3fb98bb22f7a admin-dashboard
|
||||
26ff297b66a0c974b7143179e44a58f476265482cd1929bc01a95ad29e50ca7a control-api
|
||||
091ad94b5706b1778b181542de7a4b251cd16c421144269d0955769603d51e06 validator-agent
|
||||
3e9e14dbb361ee76aaad7c1da6864b3ea111e0ed151403f904b12485631bbf75 prober
|
||||
d72234688eb1954dbff420ca1c8e83b83ceab561b18a0336af0f73fe4d9ac8be admin-dashboard
|
||||
Binary file not shown.
Binary file not shown.
BIN
Binary file not shown.
Binary file not shown.
+14
-2
@@ -59,6 +59,9 @@ func run(configPath string, log *slog.Logger) error {
|
||||
}
|
||||
|
||||
orch := orchestrator.New(database, osClient, cfg, log)
|
||||
// Background jobs (the floating-IP scan) live as long as the process, not
|
||||
// as long as the HTTP request or loop iteration that started them.
|
||||
orch.SetContext(ctx)
|
||||
|
||||
adminToken := os.Getenv(cfg.Auth.AdminTokenEnv)
|
||||
agentToken := os.Getenv(cfg.Auth.AgentTokenEnv)
|
||||
@@ -135,8 +138,10 @@ func runOrchestratorLoop(ctx context.Context, orch *orchestrator.Orchestrator, c
|
||||
} else if ac.Enabled {
|
||||
continue
|
||||
}
|
||||
if _, _, err := orch.ScanFloatingIPs(ctx); err != nil {
|
||||
log.Error("scan floating ips", "err", err)
|
||||
// Non-blocking: the scan runs in the background (single-flight, so
|
||||
// a still-running scan is simply joined) and must not stall Tick.
|
||||
if st, started := orch.StartScan(orchestrator.ScanOptions{}); !started {
|
||||
log.Info("periodic floating ip scan skipped: a scan is already running", "state", st.State)
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -158,11 +163,18 @@ func newOpenStackClient(ctx context.Context, cfg *config.ControlAPI) (openstack.
|
||||
}
|
||||
|
||||
func newRealOpenStackClient(ctx context.Context, cfg *config.ControlAPI) (openstack.FloatingIPClient, error) {
|
||||
retries := cfg.OpenStack.ListPageRetries
|
||||
if retries < 0 {
|
||||
retries = 0 // negative in the config disables retries
|
||||
}
|
||||
clientCfg := openstack.ClientConfig{
|
||||
AuthURL: os.Getenv(cfg.OpenStack.AuthURLEnv),
|
||||
ProjectID: os.Getenv(cfg.OpenStack.ProjectIDEnv),
|
||||
Region: os.Getenv(cfg.OpenStack.RegionEnv),
|
||||
Interface: os.Getenv(cfg.OpenStack.InterfaceEnv),
|
||||
|
||||
RequestTimeout: time.Duration(cfg.OpenStack.RequestTimeoutSeconds) * time.Second,
|
||||
ListPageRetries: retries,
|
||||
}
|
||||
|
||||
switch cfg.OpenStack.AuthMethod {
|
||||
|
||||
@@ -37,6 +37,20 @@ openstack:
|
||||
user_domain_name_env: "OS_USER_DOMAIN_NAME"
|
||||
password_env: "OS_PASSWORD"
|
||||
|
||||
# Постраничное чтение Floating IP (скан при тысячах адресов): сколько
|
||||
# адресов запрашивать у Neutron за один запрос. Default 200.
|
||||
# Page size of the paged floating-IP listing. Default 200.
|
||||
list_page_size: 200
|
||||
# Таймаут каждого HTTP-запроса к Keystone/Neutron, секунд. Default 60.
|
||||
# Per-request HTTP timeout (also protects the orchestrator tick from a
|
||||
# hung Neutron call). Default 60.
|
||||
request_timeout_seconds: 60
|
||||
# Сколько раз повторять неудавшуюся страницу (сетевая ошибка, EOF/
|
||||
# RemoteDisconnected, 5xx, 429) с паузами 1,2,4,8,16 с. Default 5;
|
||||
# отрицательное значение отключает повторы.
|
||||
# Retries per failed listing page. Default 5; negative disables retries.
|
||||
list_page_retries: 5
|
||||
|
||||
# Аутентификация API: здесь только ИМЕНА переменных окружения, значения
|
||||
# (статические bearer-токены) задаются окружением процесса — см.
|
||||
# deploy/systemd/control-api.service (EnvironmentFile=). Генерация:
|
||||
@@ -76,6 +90,10 @@ orchestrator:
|
||||
# scan on demand via POST /api/v1/admin/ips/scan or the dashboard's
|
||||
# "Scan Floating IPs" button.
|
||||
fip_scan_interval_seconds: 0
|
||||
# Общий таймаут одного фонового скана Floating IP (очистка + чтение всех
|
||||
# страниц + постановка в очередь), секунд. Default 1800.
|
||||
# Overall deadline of one background floating-IP scan. Default 1800.
|
||||
fip_scan_timeout_seconds: 1800
|
||||
|
||||
aggregation:
|
||||
missing_counts_as_fail: true
|
||||
|
||||
@@ -16,6 +16,8 @@ database:
|
||||
|
||||
openstack:
|
||||
mode: "mock"
|
||||
# Постраничное чтение Floating IP / page size of the paged listing (default 200)
|
||||
list_page_size: 200
|
||||
|
||||
# Статические bearer-токены читаются из переменных окружения контейнера с
|
||||
# этими именами (задаются в .env, см. deploy/docker/.env.example). Пусто =
|
||||
@@ -34,6 +36,7 @@ orchestrator:
|
||||
heartbeat_timeout_seconds: 30
|
||||
fip_settle_seconds: 0
|
||||
fip_scan_interval_seconds: 0
|
||||
fip_scan_timeout_seconds: 1800 # общий таймаут фонового скана / scan deadline
|
||||
|
||||
aggregation:
|
||||
missing_counts_as_fail: true
|
||||
|
||||
+66
-22
@@ -330,15 +330,37 @@ IP на данном проходе". До этого момента control-api
|
||||
{
|
||||
"total_ips": 25,
|
||||
"ips_by_state": {"queued": 10, "checking": 3, "done": 11, "failed": 1},
|
||||
"results_by_overall": {"pass": 8, "partial": 3, "fail": 0, "cancelled": 0},
|
||||
"total_validators": 4
|
||||
}
|
||||
```
|
||||
|
||||
`results_by_overall` — сколько адресов с каким итогом (всегда все четыре ключа). Счётчики считаются
|
||||
запросами `GROUP BY` на стороне БД, а не загрузкой всей очереди, поэтому метод быстрый и при тысячах адресов.
|
||||
|
||||
### `GET /api/v1/admin/ips`
|
||||
|
||||
Полный список всех IP из очереди со всеми полями (см.
|
||||
Список IP из очереди со всеми полями (см.
|
||||
[USAGE.md](USAGE.md#значения-полей-ip) — расшифровка полей и статусов).
|
||||
|
||||
**Без параметров** — как раньше: весь список одним массивом (при тысячах адресов это мегабайты — для больших очередей
|
||||
используйте постраничный режим). **С `limit`** — постраничный режим: ответ — конверт
|
||||
|
||||
```json
|
||||
{"items": [ ... ], "total": 6440, "limit": 50, "offset": 0}
|
||||
```
|
||||
|
||||
| Параметр | Значение |
|
||||
|---|---|
|
||||
| `limit` | размер страницы, `1`…`1000` (иначе `400`); включает постраничный режим |
|
||||
| `offset` | смещение, `>= 0` (без `limit` — `400`) |
|
||||
| `state` | одно или несколько состояний через запятую (`queued`, `assigning_fip`, `awaiting_self_check`, `checking`, `aggregating`, `done`, `failed`, `occupied`) |
|
||||
| `q` | подстрока адреса |
|
||||
| `result` | итог: `pass`, `partial`, `fail`, `cancelled` |
|
||||
| `order` | `sequence` (по умолчанию, порядок очереди) или `aggregated_at_desc` (последние завершённые) |
|
||||
|
||||
`total` — число записей после фильтров. Параметры фильтров без `limit` возвращают отфильтрованный массив.
|
||||
|
||||
### `GET /api/v1/admin/ips/{ip}`
|
||||
|
||||
Детали по одному адресу: сам объект IP, все проверки текущей попытки и
|
||||
@@ -470,33 +492,52 @@ YAML для этой секции больше не перечитывается
|
||||
|
||||
### `POST /api/v1/admin/ips/scan`
|
||||
|
||||
Сканирует текущий проект OpenStack на предмет свободных (не привязанных ни
|
||||
к одному порту) Floating IP и сразу передаёт найденный список в `POST
|
||||
/api/v1/admin/ips` — тот же add/requeue/reorder-вызов, как если бы
|
||||
оператор ввёл эти адреса вручную. Не принимает тело запроса.
|
||||
Запускает **фоновое** сканирование проекта OpenStack: находит все свободные (не привязанные ни к одному порту) Floating IP
|
||||
и ставит их в очередь — тот же add/requeue/reorder, что и `POST /api/v1/admin/ips`. Не принимает тело запроса и **сразу отвечает**
|
||||
`202` со статусом задания; ход сканирования смотрите через `GET /api/v1/admin/ips/scan`.
|
||||
|
||||
Почему в фоне: в проекте может быть тысячи Floating IP (на стенде — около 6,4 тыс.), Neutron отдаёт такой список минуты. Control-api читает
|
||||
его **страницами** (по `openstack.list_page_size`, по умолчанию 200, по `marker`), повторяет страницу при обрыве соединения/5xx/429,
|
||||
сначала обнаруживает **все** адреса и только потом ставит их в очередь кусками по 500 в порядке возрастания IP. Если чтение не удалось
|
||||
(после повторов), в очередь не попадает ничего — очередь остаётся как была, а статус задания — `error`.
|
||||
|
||||
| Параметр | Значение |
|
||||
|---|---|
|
||||
| `dry_run=true` | только найти и посчитать свободные адреса; очередь не меняется (безопасная проверка, итог — в статусе) |
|
||||
| `wait=true` | дождаться окончания и ответить `200` прежним телом `{scanned_free, added[], requeued[], reordered[], skipped_in_progress[]}` (для curl и скриптов; при ошибке `502`) |
|
||||
|
||||
Одновременно идёт одно сканирование: повторный запрос во время работы **присоединяется** к текущему и тоже отвечает `202` с его статусом.
|
||||
|
||||
### `GET /api/v1/admin/ips/scan`
|
||||
|
||||
Статус и прогресс сканирования (admin-токен).
|
||||
|
||||
Ответ (`200`):
|
||||
```json
|
||||
{
|
||||
"scanned_free": 3,
|
||||
"added": ["203.0.113.20"],
|
||||
"requeued": [],
|
||||
"reordered": ["203.0.113.10", "203.0.113.11"],
|
||||
"skipped_in_progress": []
|
||||
"state": "listing",
|
||||
"running": true,
|
||||
"dry_run": false,
|
||||
"pages": 12,
|
||||
"discovered": 2400,
|
||||
"free": 2399,
|
||||
"added": 0,
|
||||
"requeued": 0,
|
||||
"reordered": 0,
|
||||
"skipped_in_progress": 0,
|
||||
"started_at": "2026-10-01T15:47:40.759Z",
|
||||
"finished_at": null,
|
||||
"error": ""
|
||||
}
|
||||
```
|
||||
|
||||
`scanned_free` — сколько свободных Floating IP нашлось в проекте всего
|
||||
(включая уже стоящие в очереди — они попадут в `reordered`, а не
|
||||
`added`). Если свободных адресов нет вообще, это не ошибка: ответ будет
|
||||
`{"scanned_free": 0, "added": [], ...}`.
|
||||
`state`: `idle` (в этом процессе сканирования ещё не было), `clearing` (очистка очереди — только в автоцикле), `listing` (чтение страниц),
|
||||
`enqueuing` (постановка в очередь), `done`, `error` (причина в `error`), `cancelled`. `discovered` — сколько Floating IP прочитано
|
||||
(свободных и занятых), `free` — из них свободных, `added`/`requeued`/`reordered`/`skipped_in_progress` — итог постановки в очередь
|
||||
(как в `POST /admin/ips`). Статус хранится в памяти процесса: после перезапуска control-api он снова `idle`.
|
||||
|
||||
Помимо ручного вызова, сканирование можно включить по расписанию —
|
||||
`orchestrator.fip_scan_interval_seconds` в `control-api.yaml` (0, по
|
||||
умолчанию, — только по запросу через эту ручку или кнопку «Сканировать
|
||||
Floating IP» в дашборде). Пока включён
|
||||
[автоматический цикл](#автоматический-цикл-проверок), периодический скан
|
||||
не выполняется.
|
||||
Помимо ручного вызова, сканирование можно включить по расписанию — `orchestrator.fip_scan_interval_seconds` в `control-api.yaml`
|
||||
(0, по умолчанию, — только по запросу через эту ручку или кнопку «Сканировать Floating IP» в дашборде). Пока включён
|
||||
[автоматический цикл](#автоматический-цикл-проверок), периодический скан не выполняется.
|
||||
|
||||
## Автоматический цикл проверок
|
||||
|
||||
@@ -575,7 +616,10 @@ curl -s -X POST http://<control-api>:8080/api/v1/admin/auto-cycle/stop
|
||||
|
||||
### `GET /api/v1/admin/registry`
|
||||
|
||||
Список всех адресов реестра с краткой сводкой по каждому.
|
||||
Список адресов реестра с краткой сводкой по каждому. Без параметров — все адреса одним массивом; **с `limit`** (`1`…`1000`) —
|
||||
постраничный конверт `{"items": [...], "total": N, "limit": L, "offset": O}`, параметры `offset`, `q` (подстрока адреса) и
|
||||
`last_result` (`pass`/`partial`/`fail`/`cancelled`). Страница и фильтры применяются в SQL до расчёта сводки, поэтому
|
||||
реестр из тысяч адресов отдаётся за доли секунды.
|
||||
|
||||
```json
|
||||
[
|
||||
|
||||
+40
-27
@@ -49,7 +49,7 @@ admin-dashboard -config /etc/cloud-ip-validator/admin-dashboard.yaml
|
||||
| Страница | Назначение |
|
||||
|---|---|
|
||||
| `/overview` | Сводная статистика: счётчики по состояниям, «текущая проверка» (live-снимок всех IP не в терминальном состоянии) и «последние N завершённых» (по умолчанию 20, `overview.last_completed_count`) с разбивкой pass/partial/fail/cancelled. Обновляется каждые `overview.poll_interval_seconds` секунд без перезагрузки страницы. Поиск по IP и фильтр по статусу (`pass`/`partial`/`fail`/`cancelled`) над обеими таблицами — набранное/выбранное не сбрасывается очередным обновлением. Пока включён [автоматический цикл](USAGE.md#автоматический-цикл-проверок), под счётчиками показывается индикатор «Автоцикл активен» с текущей фазой и временем следующего запуска; управляется цикл на `/settings`. |
|
||||
| `/ips` | Полная очередь. Форма сверху принимает список адресов (по одному на строке или через запятую) и отправляет их в `POST /api/v1/admin/ips` — **один и тот же вызов** добавляет новые адреса и принудительно перезапускает уже завершённые (см. ниже). Кнопка «Сканировать Floating IP» делает то же самое автоматически: находит в проекте OpenStack все свободные (не привязанные к порту) Floating IP и сразу ставит их в очередь (`POST /api/v1/admin/ips/scan`, см. [API.md](API.md#post-apiv1adminipsscan)) — то же сканирование можно включить по расписанию через `orchestrator.fip_scan_interval_seconds`. У каждого адреса — кнопка «Перепроверить» (для `done`/`failed`) или «Отменить» (для активных состояний), и всегда — «Удалить» (безвозвратно убирает адрес из очереди, но не из реестра — см. ниже). Чекбоксы у строк + кнопка «Удалить выбранные» удаляют список одним вызовом; «Очистить всё» удаляет вообще всё, включая активные проверки — обе операции требуют явного подтверждения. Пока не истекла настроенная на `/settings` пауза (`fip_settle_seconds`), только что привязавший Floating IP адрес показывает отдельный бейдж «прогрев FIP» вместо обычного статуса. Если на момент попытки привязки Floating IP оказался уже занят другим портом (дрейф состояния облака или ошибочно переданный адрес), цикл проверки для него не запускается — адрес показывает отдельный бейдж «занят» (отличный от «fail») и строку `fip_occupied` в списке событий на его странице; кнопка «Перепроверить» ставит его в очередь заново. |
|
||||
| `/ips` | Очередь **постранично** (по 50 адресов; 25/50/100/200) с поиском по IP и фильтром по состоянию/итогу на сервере; кнопка «Сканировать Floating IP» запускает фоновое сканирование с панелью прогресса, «Пробное сканирование» ничего не ставит в очередь (подробности — «Очередь из тысяч адресов» ниже). Форма сверху принимает список адресов (по одному на строке или через запятую) и отправляет их в `POST /api/v1/admin/ips` — **один и тот же вызов** добавляет новые адреса и принудительно перезапускает уже завершённые (см. ниже). Кнопка «Сканировать Floating IP» делает то же самое автоматически: находит в проекте OpenStack все свободные (не привязанные к порту) Floating IP и сразу ставит их в очередь (`POST /api/v1/admin/ips/scan`, см. [API.md](API.md#post-apiv1adminipsscan)) — то же сканирование можно включить по расписанию через `orchestrator.fip_scan_interval_seconds`. У каждого адреса — кнопка «Перепроверить» (для `done`/`failed`) или «Отменить» (для активных состояний), и всегда — «Удалить» (безвозвратно убирает адрес из очереди, но не из реестра — см. ниже). Чекбоксы у строк + кнопка «Удалить выбранные» удаляют список одним вызовом; «Очистить всё» удаляет вообще всё, включая активные проверки — обе операции требуют явного подтверждения. Пока не истекла настроенная на `/settings` пауза (`fip_settle_seconds`), только что привязавший Floating IP адрес показывает отдельный бейдж «прогрев FIP» вместо обычного статуса. Если на момент попытки привязки Floating IP оказался уже занят другим портом (дрейф состояния облака или ошибочно переданный адрес), цикл проверки для него не запускается — адрес показывает отдельный бейдж «занят» (отличный от «fail») и строку `fip_occupied` в списке событий на его странице; кнопка «Перепроверить» ставит его в очередь заново. |
|
||||
| `/ips/{ip}` | Детали одного адреса, пока он в очереди: все проверки текущей попытки и вся история событий, плюс ссылка на полную историю в реестре (см. ниже). |
|
||||
| `/registry` | **Реестр** — все адреса, когда-либо поставленные на проверку, независимо от того, стоят ли они сейчас в очереди. Переживает удаление адреса из `/ips` и повторное добавление того же адреса позже (см. «Реестр адресов» ниже). Поиск по IP и фильтр по статусу — то же самое, что на `/overview`, плюс отражается в адресной строке (`?q=&status=`), так что отфильтрованную ссылку можно сохранить/переслать. |
|
||||
| `/registry/{ip}` | Полная сохранённая история проверок одного адреса по всем циклам (не только текущему) — в отличие от `/ips/{ip}`, которая показывает только текущую попытку. |
|
||||
@@ -59,41 +59,37 @@ admin-dashboard -config /etc/cloud-ip-validator/admin-dashboard.yaml
|
||||
| `/check-types` | Типы проверок (`https`/`icmp`/`ssh`/...), включение/выключение, привязка к группам целей. |
|
||||
| `/settings` | Четыре блока. Первый — панель **«Автоматический цикл»**: статус и фаза, время последнего/следующего запуска, результат последнего цикла, поля «Интервал между циклами (мин)» и «Максимальная длительность проверки (мин, 0 = без лимита)» с кнопкой «Сохранить» и кнопка «Включить»/«Выключить» (показывается та, что сейчас применима). Значения вводятся в минутах (допустимы дробные), в control-api уходят секундами; минимум интервала — 1 минута (`60` с), нарушение приходит предупреждением в баннере. Подробности — [USAGE.md](USAGE.md#автоматический-цикл-проверок), API — [API.md](API.md#автоматический-цикл-проверок). Далее три формы: `fip_settle_seconds` — пауза (в секундах) между привязкой Floating IP и началом self-check («прогрев» дата-плейна OpenStack, см. [USAGE.md](USAGE.md#пауза-перед-self-check-fip_settle_seconds)); `history_retention_cycles` — сколько последних циклов проверки хранить на адрес в реестре (0 — без ограничения); и типы проверок пробера — TCP-порты (через запятую) + чекбокс ICMP, общие для всех площадок (см. [USAGE.md](USAGE.md#управление-типами-проверок-пробера)). |
|
||||
|
||||
### «Текущая» и «последняя завершённая» проверка
|
||||
### «В работе», «в очереди» и «последняя завершённая» проверка
|
||||
|
||||
В `control-api` нет понятия «запуска»/«цикла проверки» как отдельной
|
||||
сущности — есть только общая очередь IP-адресов
|
||||
(`docs/PLAN_ADMIN_DASHBOARD.md`). Дашборд ничего не меняет в этом
|
||||
устройстве и не заводит своего состояния:
|
||||
устройстве и не заводит своего состояния. Очередь может содержать тысячи
|
||||
адресов, поэтому `/overview` **никогда не загружает её целиком** — на каждое
|
||||
обновление запрашиваются счётчики и несколько ограниченных списков:
|
||||
|
||||
- **Текущая проверка** — все адреса, которые прямо сейчас не в
|
||||
состоянии `done`/`failed` (`queued`, `assigning_fip`,
|
||||
`awaiting_self_check`, `checking`, `aggregating`), вычисляется заново на
|
||||
каждый запрос из `GET /api/v1/admin/status` + `GET /api/v1/admin/ips`.
|
||||
- **Последняя завершённая проверка** — последние N адресов, перешедших в
|
||||
`done`/`failed`, отсортированные по `AggregatedAt` по убыванию (не
|
||||
«последний запуск», а именно скользящее окно последних по времени
|
||||
завершений).
|
||||
- **Счётчики** — `GET /api/v1/admin/status` (по состояниям и `results_by_overall`).
|
||||
- **В работе** — адреса в `assigning_fip`, `awaiting_self_check`, `checking`, `aggregating`
|
||||
(не более 100; естественный предел — число валидаторов).
|
||||
- **В очереди: Q** — счётчик `queued` со ссылкой на `/ips?state=queued` и несколько ближайших адресов.
|
||||
- **Последние N завершённых** — последние N адресов в `done`/`failed` по `AggregatedAt` по убыванию
|
||||
(не «последний запуск», а скользящее окно последних по времени завершений).
|
||||
- **Прогресс** в блоке статистики: «Готово D из T (P%) · в работе A · в очереди Q» с полосой и оценкой
|
||||
оставшегося времени (по скорости последних завершений, когда их не меньше пяти). `occupied` считается
|
||||
завершённым состоянием.
|
||||
|
||||
### Поиск по IP и фильтр по статусу
|
||||
|
||||
На `/overview` и `/registry` есть форма из двух полей — поиск по IP
|
||||
(подстрока, без учёта регистра) и выпадающий список статуса
|
||||
(`pass`/`partial`/`fail`/`cancelled`). Оба поля работают вместе (И, а не
|
||||
ИЛИ) и применяются целиком на стороне дашборда — `client.ListIPs`/
|
||||
`client.ListRegistry` всегда получают от `control-api` полный список,
|
||||
`internal/httpapi`/`internal/db` про фильтр вообще не знают.
|
||||
На `/overview`, `/ips` и `/registry` есть поиск по IP (подстрока) и фильтр по статусу/итогу (`pass`/`partial`/`fail`/`cancelled`;
|
||||
на `/ips` — ещё по состоянию очереди). Поля работают вместе (И, а не ИЛИ). Фильтрация выполняется **на стороне `control-api`**
|
||||
(параметры `q`, `state`, `result`/`last_result` у `GET /admin/ips` и `GET /admin/registry`), а дашборд получает только нужную страницу, поэтому
|
||||
фильтр работает быстро при любом размере очереди.
|
||||
|
||||
- **`/overview`** — фильтр действует на обе таблицы сразу («Текущая
|
||||
проверка» и «Последние N завершённых»). Статус — это фильтр по
|
||||
итоговому результату (`OverallResult`), поэтому выбор конкретного
|
||||
статуса скрывает «Текущую проверку» целиком: у ещё идущих проверок
|
||||
результата попросту нет. Панель статистики (счётчики сверху) фильтру не
|
||||
подчиняется — это агрегаты по всей очереди, а не по видимым строкам.
|
||||
- **`/registry`** — тот же принцип, но по одной таблице (`LastResult`), и
|
||||
значения полей отражаются в адресной строке (`?q=&status=`) через
|
||||
`hx-replace-url` — отфильтрованную ссылку можно сохранить или переслать,
|
||||
а обновление страницы (F5) сохраняет применённый фильтр.
|
||||
- **`/overview`** — `q` и статус передаются в списки «В работе» и «Последние N завершённых». Статус — это фильтр по
|
||||
итоговому результату (`OverallResult`), поэтому выбор конкретного статуса скрывает «В работе» и «В очереди»: у ещё идущих проверок
|
||||
результата попросту нет. Панель статистики (счётчики сверху) фильтру не подчиняется — это агрегаты по всей очереди, а не по видимым строкам.
|
||||
- **`/registry` и `/ips`** — таблица постраничная; значения полей и страница отражаются в адресной строке (`?q=&status=&page=`) через
|
||||
`hx-replace-url` — отфильтрованную ссылку можно сохранить или переслать, а обновление страницы (F5) сохраняет применённый фильтр.
|
||||
|
||||
**Раскладка `/overview` сверху вниз**: панель статистики → форма
|
||||
фильтра → таблицы. Панель статистики и форма фильтра физически лежат
|
||||
@@ -157,6 +153,23 @@ auto-refresh на `/ips`, см. git-историю). Опрашивается т
|
||||
циклов) остаётся всегда. Подробнее —
|
||||
[API.md](API.md#реестр-адресов-и-история-проверок).
|
||||
|
||||
## Очередь из тысяч адресов
|
||||
|
||||
После сканирования проекта в очереди может оказаться несколько тысяч адресов, поэтому тяжёлые страницы работают постранично:
|
||||
|
||||
- **`/ips` и `/registry`** — параметры `page` и `per_page` (по умолчанию 50; допустимо 25/50/100/200), «Показано a–b из N» и кнопки ‹ ›.
|
||||
Поиск (`q`), состояние/итог и размер страницы применяются **на стороне control-api** (`GET /admin/ips?limit=…`, `GET /admin/registry?limit=…`),
|
||||
поэтому страница весит десятки килобайт независимо от длины очереди. Фильтры и страница отражены в адресной строке.
|
||||
- **Массовые операции.** Чекбоксы выбирают строки текущей страницы (счётчик «Выбрано на странице: k из P»). Если отмечен заголовок таблицы и записей
|
||||
больше страницы, появляется ссылка «Выбрать все N по фильтру»: тогда «Перепроверить»/«Удалить» применяются ко **всем** адресам по текущему фильтру
|
||||
(адреса разрешаются на сервере и отправляются кусками по 500). Подтверждения показывают реальное число: «Удалить ВСЕ 6440 адресов…».
|
||||
«Очистить всё» очищает очередь целиком одной быстрой операцией.
|
||||
- **Сканирование.** Кнопка «Сканировать Floating IP» мгновенно возвращает панель прогресса под кнопкой; пока задание идёт, панель сама
|
||||
обновляется каждые 2 секунды, по окончании опрос прекращается и таблица перезагружается. Во время сканирования кнопки заблокированы; повторное
|
||||
нажатие присоединяется к идущему заданию. Ошибка (например, OpenStack недоступен) показывается в панели с причиной.
|
||||
Саму таблицу `/ips` по таймеру по-прежнему не обновляем — она не сбрасывает ввод оператора.
|
||||
- Долгие операции (`Очистить всё`, массовое удаление/перепроверка) выполняются с увеличенным таймаутом (120 с), остальные запросы к control-api — с `control_api.timeout_seconds`.
|
||||
|
||||
## Вход и сессия
|
||||
|
||||
Если заданы `ADMIN_DASHBOARD_USERNAME` и `ADMIN_DASHBOARD_PASSWORD`, все страницы, кроме `/login` и `/static/*`, требуют входа.
|
||||
|
||||
+1
-1
@@ -66,7 +66,7 @@ It will:
|
||||
outcome is `completed`, the phase is `waiting` with `runs_total=1`, and
|
||||
that the registry's `total_cycles` for `127.0.0.1` grew (the cycle
|
||||
cleared the queue, re-scanned the mock floating IP and re-checked it).
|
||||
Finally it `stop`s the cycle and asserts it is `idle`. The second cycle
|
||||
The cycle now passes through the background scan (`scanning` phase) before the checks run. Finally it `stop`s the cycle and asserts it is `idle`. The second cycle
|
||||
(the interval wait) is covered by unit tests, so the script does not
|
||||
sit through the 60s pause. The script exits non-zero if any assertion
|
||||
fails.
|
||||
|
||||
+38
-7
@@ -90,7 +90,8 @@ curl -s -X POST http://<control-api>:8080/api/v1/admin/ips \
|
||||
самому найти их в облаке:
|
||||
|
||||
```bash
|
||||
curl -s -X POST http://<control-api>:8080/api/v1/admin/ips/scan
|
||||
curl -s -X POST http://<control-api>:8080/api/v1/admin/ips/scan # 202: сканирование запущено в фоне
|
||||
curl -s http://<control-api>:8080/api/v1/admin/ips/scan # ход и результат
|
||||
```
|
||||
|
||||
Сканируются все Floating IP текущего проекта OpenStack, но в очередь
|
||||
@@ -101,8 +102,23 @@ curl -s -X POST http://<control-api>:8080/api/v1/admin/ips/scan
|
||||
встают в очередь, уже завершённые перезапускаются, активно проверяемые не
|
||||
трогаются (см. [выше](#добавление-новых-ip-в-очередь)).
|
||||
|
||||
В `admin-dashboard` то же самое — кнопка «Сканировать Floating IP» на
|
||||
странице `/ips`.
|
||||
**Сколько адресов — не важно.** Сканирование работает в фоне и читает список из OpenStack
|
||||
**страницами** (по 200 адресов, с повторами при обрывах), поэтому подходит и для проекта с
|
||||
тысячами Floating IP: на стенде с 6441 адресом чтение занимает около 1,5–2 минут. Сначала
|
||||
обнаруживаются **все** адреса, и только потом они ставятся в очередь (кусками по 500, в порядке
|
||||
возрастания IP); сразу после этого начинаются проверки. Если чтение сорвалось даже после повторов,
|
||||
в очередь не попадает ничего — очередь остаётся как была, а на панели виден `error` и причина.
|
||||
Одновременно идёт одно сканирование: повторное нажатие присоединяется к текущему.
|
||||
|
||||
В `admin-dashboard` — кнопка «Сканировать Floating IP» на странице `/ips`: она сразу отвечает, а под
|
||||
кнопкой появляется панель прогресса (читаются страницы → ставятся в очередь → готово: прочитано
|
||||
страниц, найдено, свободных, добавлено, время), по окончании таблица обновляется сама. Рядом —
|
||||
«Пробное сканирование»: оно проходит все страницы и показывает, сколько свободных адресов нашлось,
|
||||
**не меняя очередь** (удобно проверить, что облако отвечает и сколько адресов будет поставлено).
|
||||
|
||||
Параметры чтения (`control-api.yaml`): `openstack.list_page_size` (200), `openstack.request_timeout_seconds`
|
||||
(60 — таймаут одного запроса к OpenStack), `openstack.list_page_retries` (5), `orchestrator.fip_scan_timeout_seconds`
|
||||
(1800 — предел всего сканирования).
|
||||
|
||||
Если хочется, чтобы сканирование происходило само по расписанию, а не
|
||||
только по запросу — задайте `orchestrator.fip_scan_interval_seconds`
|
||||
@@ -111,6 +127,11 @@ curl -s -X POST http://<control-api>:8080/api/v1/admin/ips/scan
|
||||
[автоматический цикл](#автоматический-цикл-проверок), это периодическое
|
||||
сканирование не выполняется — цикл сам управляет очередью.
|
||||
|
||||
> **Сколько займут проверки.** Один адрес занимает около 50 секунд на валидаторе (из них 30 с — пауза
|
||||
> `fip_settle_seconds`). Поэтому очередь из 6440 адресов — примерно 18 часов на 5 валидаторах,
|
||||
> 9 часов на 10, 4,5 часа на 20. Ускорить можно числом валидаторов и (осторожно) `fip_settle_seconds`;
|
||||
> на странице «Обзор» виден прогресс «Готово D из T» и оценка оставшегося времени.
|
||||
|
||||
## Автоматический цикл проверок
|
||||
|
||||
Опциональный режим, который сам повторяет то, что оператор делает руками:
|
||||
@@ -122,7 +143,9 @@ curl -s -X POST http://<control-api>:8080/api/v1/admin/ips/scan
|
||||
в [реестре](#реестр-адресов-и-глубина-истории) при этом сохраняется.
|
||||
2. Control-api находит все свободные Floating IP и ставит их в очередь —
|
||||
то же, что «Сканировать Floating IP» (см.
|
||||
[выше](#сканирование-floating-ip-из-openstack)).
|
||||
[выше](#сканирование-floating-ip-из-openstack)). Шаги 1–2 выполняются одним
|
||||
фоновым заданием, поэтому долгое чтение тысяч адресов не блокирует работу
|
||||
оркестратора (назначение валидаторов, лизинги, heartbeat).
|
||||
3. Проверки запускаются сами — как для любого адреса в очереди.
|
||||
4. Цикл ждёт, пока **все** адреса очереди дойдут до конечного состояния
|
||||
(`done`, `failed` или `occupied`). К этому моменту результат каждого
|
||||
@@ -136,7 +159,7 @@ curl -s -X POST http://<control-api>:8080/api/v1/admin/ips/scan
|
||||
| Параметр | По умолчанию | Смысл |
|
||||
|---|---|---|
|
||||
| `interval_seconds` | `3600` (1 час) | Пауза между циклами. Не меньше `60`: слишком частые сканы нагружают API OpenStack. |
|
||||
| `max_run_seconds` | `0` (без лимита) | Сколько максимум ждать на шаге 4. По истечении цикл фиксирует `timeout` и переходит к паузе — защита от зависания (нет свободных валидаторов, недоступна площадка). Очередь при этом не трогается: следующий цикл её очистит, а до тех пор видно, что именно не дошло до конца. |
|
||||
| `max_run_seconds` | `0` (без лимита) | Сколько максимум ждать на шаге 4 (отсчёт — **от конца сканирования**). По истечении цикл фиксирует `timeout` и переходит к паузе — защита от зависания (нет свободных валидаторов, недоступна площадка). Очередь при этом не трогается: следующий цикл её очистит, а до тех пор видно, что именно не дошло до конца. **При тысячах адресов оставьте `0`** (или задайте больше расчётного времени: 6440 адресов — часы). |
|
||||
|
||||
Параметры хранятся в базе и меняются на лету, без перезапуска; в `control-api.yaml`
|
||||
ничего задавать не нужно. Новый `interval_seconds` применяется к паузе
|
||||
@@ -170,7 +193,8 @@ curl -s http://<control-api>:8080/api/v1/admin/auto-cycle
|
||||
### Фазы и результат последнего цикла
|
||||
|
||||
`phase` показывает, что происходит сейчас: `idle` (автоцикл выключен или ещё не
|
||||
стартовал), `running` (идут проверки — шаги 3–4) и `waiting` (пауза между циклами,
|
||||
стартовал), `scanning` (шаги 1–2: очистка очереди и чтение/постановка Floating IP),
|
||||
`running` (идут проверки — шаги 3–4) и `waiting` (пауза между циклами,
|
||||
шаг 5; время следующего запуска — `next_run_at`). Результат последнего цикла
|
||||
(`last_outcome`):
|
||||
|
||||
@@ -185,7 +209,8 @@ curl -s http://<control-api>:8080/api/v1/admin/auto-cycle
|
||||
### Что важно знать
|
||||
|
||||
- **Выключение не прерывает проверки**, которые уже идут: они закончатся и попадут
|
||||
в реестр, остановится только повторение.
|
||||
в реестр, остановится только повторение. Если выключить цикл во время фазы `scanning`,
|
||||
сканирование отменяется (уже поставленные в очередь куски остаются).
|
||||
- Автоцикл **владеет очередью**: каждый цикл начинается с её полной очистки,
|
||||
поэтому адреса, добавленные вручную, будут удалены (их история в реестре
|
||||
остаётся). Ручные «Очистить всё» и «Сканировать Floating IP» во время цикла
|
||||
@@ -195,6 +220,12 @@ curl -s http://<control-api>:8080/api/v1/admin/auto-cycle
|
||||
- События цикла (`auto_cycle_started`, `auto_cycle_completed`, `auto_cycle_timeout`,
|
||||
`auto_cycle_error`, `auto_cycle_stopped`) пишутся в журнал событий вместе с
|
||||
`queue_cleared` и `fip_scan`.
|
||||
- Если в момент старта цикла уже идёт чужое сканирование (ручное, пробное или по расписанию),
|
||||
цикл **дожидается** его окончания и запускает собственное (с очисткой очереди) — присоединяться
|
||||
к чужому нельзя: оно могло ничего не поставить в очередь.
|
||||
- Если сканирование завершилось ошибкой, очередь уже очищена (шаг 1) и остаётся пустой до следующего
|
||||
цикла (`interval_seconds`); исход цикла — `error` с причиной в `last_error`.
|
||||
- Цикл на тысячах адресов длится часы; пауза `interval_seconds` отсчитывается после его завершения.
|
||||
- В реальном OpenStack отвязка Floating IP после очистки очереди может
|
||||
отразиться с задержкой; если скан сразу после неё не увидел свободных адресов,
|
||||
цикл завершится с `no_free_ips` и повторится через `interval_seconds`.
|
||||
|
||||
@@ -0,0 +1,166 @@
|
||||
# План: сканирование Floating IP и автоцикл при тысячах адресов
|
||||
|
||||
> Дата: 2026-10-01 18:19 MSK · Статус: **реализовано** — результаты ревью и тестов: [2026-10-01_18-59_fip-scan-at-scale-review.md](2026-10-01_18-59_fip-scan-at-scale-review.md)
|
||||
|
||||
## Context
|
||||
|
||||
Нажатие «Сканировать Floating IP» на живом стенде падает: `control-api недоступен: … context deadline exceeded`. Расследование
|
||||
(2026-10-01) показало: в проекте OpenStack **6441 Floating IP, 6440 свободны** (раньше было 5). `ListFloatingIPs` запрашивает весь
|
||||
список одним запросом без `limit` и без таймаута, Neutron отвечает >60 с (постранично: 200 адресов ≈ 2,4 с, весь список ≈ 70–80 с),
|
||||
а дашборд ждёт 10 с. Запрос к тому же привязан к `r.Context()`: при обрыве соединения скан отменяется и не может завершиться.
|
||||
|
||||
Цель пользователя прежняя: **одной кнопкой подключить к проверке все доступные в проекте адреса, даже если их тысячи**, после чего
|
||||
проверка запускается автоматически; автоматический цикл (очистка → скан → проверка → пауза) должен работать в этих условиях.
|
||||
Решение пользователя по объёму: **полная адаптация** — фоновый постраничный скан + автоцикл + постраничные страницы дашборда.
|
||||
|
||||
Ожидание по времени (оценка по реальным данным стенда: слот на адрес ≈ 50 с, из них 30 с — `fip_settle_seconds`):
|
||||
6440 адресов ≈ 18 ч на 5 валидаторах, ≈ 9 ч на 10, ≈ 4,5 ч на 20. Это ограничение пропускной способности, а не кода; рычаги —
|
||||
число валидаторов и `fip_settle_seconds`. Для автоцикла это значит: `max_run_seconds` должен быть `0` (без лимита) или > 20 ч.
|
||||
|
||||
## Карта кода (по графу и разведке)
|
||||
|
||||
- `openstack.FloatingIPClient` (`internal/openstack/interface.go`): `GetFloatingIPByAddress`, `ListFloatingIPs`, `Associate…`, `Disassociate…`;
|
||||
реализации — реальный `Client` (`client.go`, **нет таймаутов и ретраев**) и `MockClient` (`mock.go`, есть `ListFailure`, нет пагинации).
|
||||
- `Orchestrator.ScanFloatingIPs` (`internal/orchestrator/orchestrator.go:412`) — синхронный: `OS.ListFloatingIPs` → фильтр `PortID==""` →
|
||||
один `DB.SubmitIPs`. Вызывается из `handleAdminScanFloatingIPs` (`internal/httpapi/handlers_admin.go:107`, на `r.Context()`),
|
||||
из `autoCycleStartRun` (`autocycle.go`, **в горутине цикла оркестратора под `autoCycleMu`** — минутный скан остановит `Tick`
|
||||
и sweeps) и из периодического `scanTickerC` (`cmd/control-api/main.go`).
|
||||
- БД: SQLite, `SetMaxOpenConns(1)`; `SubmitIPs` — одна транзакция, ~6 запросов на новый адрес (6440 ≈ 38 тыс. запросов);
|
||||
`DeleteIPs`/`ClearQueue` — одна транзакция, ~6 запросов на адрес; `ListRegistry` — N+1 и **O(n²)**: нет индекса `ip_queue(registry_id)`.
|
||||
- Full-table загрузчики: `GET /admin/ips`, `GET /admin/status` (грузит все строки ради счёта), `GET /admin/registry`, `autoCycleCheckRun`
|
||||
(`ListIPs` каждый тик), дашборд `/overview` (опрос каждые 5 с: ~3 МБ JSON и таблица «текущая проверка» из ~6000 `queued`-строк),
|
||||
`/ips` (~8 МБ HTML), `/registry`. Пагинации и фильтров на сервере нет.
|
||||
- Дашборд: `client.do` с таймаутом 10 с; фрагменты и опрос htmx (`overview.html`, `overview_fragment.html`), GET-фильтр
|
||||
`registry.html` (`hx-select` + `hx-replace-url`) — идиома для переиспользования. Per-row кнопки `hx-delete` при «выбрать все»
|
||||
кладут все отмеченные адреса в URL (уже сейчас дефект, при тысячах — фатальный).
|
||||
- Тик оркестратора не читает всю очередь (`ClaimNextQueued`, `ListChecking` — по индексу), пропускная способность не зависит от размера очереди.
|
||||
|
||||
## Дизайн
|
||||
|
||||
### 1. OpenStack: постраничное чтение, таймауты, ретраи (`internal/openstack`, `internal/config`)
|
||||
- Новый метод интерфейса `ListFreeFloatingIPs(ctx, pageSize int, onPage func(page []FloatingIP) error) (pages int, err error)`:
|
||||
цикл «страница → `onPage`»; для каждой страницы один запрос `floatingips.List(ListOpts{Limit, Marker=lastID})` с `EachPage`
|
||||
(возврат `false` после первой страницы) — собственная пагинация по `marker`, а не `next`-ссылка (за прокси она может указывать на
|
||||
внутренний хост). Свой `ListOptsBuilder`, добавляющий `fields=id&fields=floating_ip_address&fields=port_id&fields=project_id`
|
||||
(в gophercloud `ListOpts.Fields` нет; на стенде проверено: `fields` + `marker` работают). Фильтр свободных — на клиенте
|
||||
(`PortID==""`); серверный `status=DOWN` не используем (надмножество, возможны гонки статуса).
|
||||
- **Ретраи страницы** с backoff (по умолчанию 5 попыток, 1→2→4→8→16 с) на сетевые ошибки, `EOF/RemoteDisconnected`, 5xx и 429
|
||||
(на стенде уже наблюдался `RemoteDisconnected` на второй странице); 4xx (кроме 429) — без ретрая. Контекст отменяет ретраи.
|
||||
- **Таймаут на запрос**: `provider.HTTPClient.Timeout` (`openstack.request_timeout_seconds`, 60) — закрывает и вечные зависания в `Tick`
|
||||
(`GetFloatingIPByAddress`/`Associate`/`Disassociate`), ключевой побочный эффект.
|
||||
- Конфиг: `openstack.list_page_size` (200), `openstack.request_timeout_seconds` (60), `orchestrator.fip_scan_timeout_seconds` (1800),
|
||||
`openstack.list_page_retries` (5); дефолты в `LoadControlAPI`, примеры в `configs/*.example.yaml` и `rxprod-compose/sources/`.
|
||||
- `ListFloatingIPs` (полный список) остаётся для совместимости и тестов (реализован поверх нового метода).
|
||||
- `MockClient`: пагинация (`PageSize`), счётчик вызовов/страниц, очередь ошибок `ListFailures []error` (по одной на запрос),
|
||||
опциональная задержка страницы, `SeedMany(n)` для тестов на тысячи.
|
||||
|
||||
### 2. Фоновое задание скана (`internal/orchestrator/scanjob.go`)
|
||||
- `ScanJob` в `Orchestrator`, **нулевое значение пригодно** (тесты строят `&Orchestrator{…}` литералом): `sync.Mutex`, текущий прогресс,
|
||||
`cancel`. Метод `StartScan(opts) (ScanStatus, started bool)` — single-flight: если скан уже идёт, возвращает его статус
|
||||
(`started=false`). Горутина работает на контексте жизни процесса (хранится в `Orchestrator`, задаётся из `main`, по умолчанию
|
||||
`context.Background()`), **не** на `r.Context()`; общий дедлайн `fip_scan_timeout_seconds`.
|
||||
- Опции: `ClearFirst bool` (для автоцикла), `DryRun bool` (только обнаружить и посчитать, очередь не трогать — безопасная проверка
|
||||
на живом стенде и полезная функция для оператора).
|
||||
- Фазы и прогресс: `idle → clearing → listing → enqueuing → done|error|cancelled`; поля `pages`, `discovered`, `free`, `added`,
|
||||
`requeued`, `reordered`, `skipped_in_progress`, `started_at`, `finished_at`, `error`.
|
||||
- **Алгоритм:** (1) `ClearFirst` → `ClearQueue`; (2) чтение всех страниц в память (6440 строк — килобайты), `free = PortID==""`;
|
||||
(3) сортировка по IPv4 по возрастанию (детерминированный порядок очереди); (4) **только после полного обнаружения** — `SubmitIPs`
|
||||
кусками по 500 в этом порядке (`base=MAX+1` пересчитывается на вызов ⇒ порядок сохраняется; транзакции короткие, единственное
|
||||
соединение освобождается между кусками); проверки стартуют, как только появляются первые `queued`; (5) одно событие `fip_scan`
|
||||
с итоговыми счётчиками. Ошибка чтения после ретраев ⇒ **ничего не ставится в очередь** (для ручного скана очередь не меняется),
|
||||
статус `error` с причиной; повтор — кнопкой или следующим циклом. Ошибка БД посередине ⇒ уже поставленные куски остаются
|
||||
(повтор идемпотентен: `SubmitIPs` переупорядочивает/пропускает).
|
||||
- `ScanFloatingIPs(ctx)` остаётся тонкой синхронной обёрткой («запустить и дождаться») для существующих тестов/скриптов.
|
||||
- Периодический `scanTickerC` вызывает неблокирующий `StartScan`.
|
||||
|
||||
### 3. Масштабирование БД и запросов (`internal/db`, миграция `0009`)
|
||||
- Миграция `0009_scale_indexes.sql`: `idx_ip_queue_registry ON ip_queue(registry_id)` (убирает O(n²) в реестре),
|
||||
`idx_ip_queue_state_aggregated ON ip_queue(state, aggregated_at)` (список «последние завершённые»).
|
||||
- Новые запросы: `CountIPsByState`, `CountIPsByResult` (GROUP BY — вместо загрузки всех строк в `/admin/status`),
|
||||
`AnyNonTerminalIP` (`SELECT EXISTS … state NOT IN (done,failed,occupied)`), `ListIPsPage(filter{states[], q, result, order},
|
||||
limit, offset) → (items, total)`, `ListRegistryPage(filter{q, lastResult}, limit, offset) → (items, total)` — **LIMIT/OFFSET до**
|
||||
`fillRegistrySummary`, поэтому 3–4 запроса на строку платят только строки страницы. Фильтр `lastResult` реализуется одним SQL:
|
||||
`ip_registry r LEFT JOIN ip_queue q ON q.registry_id=r.id`, условие `(q.id IS NOT NULL AND q.overall_result=?) OR (q.id IS NULL AND
|
||||
<подзапрос по checks последнего цикла: pass/fail/partial>=?)` — та же семантика, что `fillRegistrySummary`/`lastCycleResultFromChecks`,
|
||||
без денормализации и миграции данных.
|
||||
- `ClearQueue`: set-based очистка без цикла по адресам — `UPDATE validators SET current_ip_id=NULL…`, `UPDATE checks SET ip_id=NULL`,
|
||||
`UPDATE events SET ip_id=NULL`, `DELETE ip_site_checks`, `DELETE ip_queue` (5 запросов, O(n)); disassociate FIP только для строк с `FIPID`;
|
||||
событие `queue_cleared` — счётчик и усечённый список (не 6440 адресов). `Orchestrator.DeleteIPs` — выбор строк без N `GetIPByAddress`.
|
||||
|
||||
### 4. HTTP API (`internal/httpapi`) — обратная совместимость сохраняется
|
||||
| Метод | Путь | Изменение |
|
||||
|---|---|---|
|
||||
| POST | `/admin/ips/scan` | `202 {state, started_at, …}` (запуск или уже идущий скан — `202` с текущим статусом); `?dry_run=true`; `?wait=true` — старая синхронная семантика (`200` + счётчики) для curl/скриптов |
|
||||
| GET | `/admin/ips/scan` | **новый**: статус и прогресс скана (admin-токен) |
|
||||
| GET | `/admin/ips` | без параметров — как раньше (массив); с `limit` — конверт `{items,total,limit,offset}`; фильтры `state` (csv), `q`, `result`, `order` |
|
||||
| GET | `/admin/registry` | то же: `limit/offset/q/last_result` → конверт с `total` |
|
||||
| GET | `/admin/status` | + `results_by_overall`, счёт через `GROUP BY` |
|
||||
| GET | `/admin/overview` | **опционально** одним запросом: счётчики, активные (≤100), последние завершённые (N), ближайшие в очереди (≤10), статус скана и автоцикла |
|
||||
|
||||
Новые admin-маршруты попадают в таблицу `routes.go` с `accessAdmin`; `TestRouteTableClassification` (`auth_test.go`) обновить (+1–2 admin).
|
||||
|
||||
### 5. Автоцикл (`internal/orchestrator/autocycle.go`, `queries_autocycle.go`, `dashboard/dto.go`)
|
||||
- Новая фаза **`scanning`** (миграция не нужна — валидатор фаз в `UpdateAutoCycleState` расширить; подписи `PhaseLabel` — «сканирование Floating IP»).
|
||||
- `autoCycleStartRun` перестаёт блокировать цикл: запускает `StartScan{ClearFirst:true}` (очистка + скан целиком в фоне, **литерал
|
||||
сценария пользователя сохранён: очистка → скан**) и сразу переводит фазу в `scanning`; `autoCycleMu` держится только на время
|
||||
чтения/записи состояния, а не на всё время скана ⇒ `Tick` и `Start/Stop` не блокируются.
|
||||
- Шаг `scanning`: опрос статуса задания. `running` → выход; `error` → существующий путь `fail()` (исход `error`, повтор через
|
||||
`interval_seconds`); `done` и `free==0` → `no_free_ips`; `done` → фаза `running`, `last_scanned_free`, **`run_started_at` = конец скана**
|
||||
(лимит `max_run_seconds` считается от конца скана). Таймаут самого скана — `fip_scan_timeout_seconds`.
|
||||
- **Восстановление после рестарта:** фаза `scanning` без живого задания ⇒ заново `StartScan{ClearFirst:true}` (идемпотентно).
|
||||
- `Stop` отменяет задание скана (исход `stopped`, если шёл скан или проверка).
|
||||
- `autoCycleCheckRun`: проверка завершения — `AnyNonTerminalIP` вместо `ListIPs` каждый тик; `COUNT` только при завершении.
|
||||
- Документировать: при тысячах адресов `max_run_seconds=0`; цикл длится часы; интервал отсчитывается от завершения.
|
||||
|
||||
### 6. Дашборд (`internal/dashboard`, шаблоны, CSS)
|
||||
- **Скан-кнопка и прогресс:** `hx-post="/ips/scan"` возвращает панель `scan_progress` (вне `#ips-form`): стадия, `<progress>`,
|
||||
прочитано/свободных/добавлено, время, ошибка; пока `running` панель сама опрашивает `GET /ips/scan/status` (`hx-trigger="every 2s"`),
|
||||
по завершении — без триггера и с `HX-Trigger: scan-finished`, по которому таблица перезагружается (`hx-get` + `hx-select`);
|
||||
кнопка блокируется на время скана; `409`/ошибки — штатным баннером (`bannerFor`). Скан-старт возвращается мгновенно, таймаут
|
||||
10 с больше не проблема.
|
||||
- **Пагинация `/ips` и `/registry`:** `page`, `per_page` (50 по умолчанию; 25/50/100/200), partial `pager` («Показано a–b из N», ‹ ›,
|
||||
`url.Values` для экранирования); фильтры на сервере: `/registry` — `q`, `status`; `/ips` — новая форма `q` + состояние
|
||||
(все / в очереди / в работе / done / failed / occupied / результат), идиома GET + `hx-select` + `hx-replace-url`.
|
||||
Скрытые `page/q/state` внутри `#ips-form`, чтобы мутации возвращали ту же страницу.
|
||||
- **Массовые операции:** чекбоксы — только строки страницы + счётчик «Выбрано на странице k из 50»; при отмеченном заголовке и
|
||||
`total > per_page` — ссылка «Выбрать все N по фильтру» (`scope=all`): адреса разрешаются на сервере постранично и уходят в
|
||||
`DeleteIPs`/`SubmitIPs` кусками по ~500. Per-row и scan/clear-кнопки получают `hx-params="page,q,state"` (чинит URL из тысяч адресов).
|
||||
`hx-confirm` с реальным числом (`Удалить ВСЕ {{.Total}} адресов…`).
|
||||
- **«Обзор»:** без `ListIPs`; `Status` (+`results_by_overall`) и ограниченные списки: «В работе» (активные состояния), «В очереди: Q»
|
||||
(счётчик + ссылка на `/ips?state=queued`, ≤10 ближайших), «Последние N завершённых»; `occupied` — терминальное состояние.
|
||||
Индикатор в блоке статистики (OOB-обновление, как сейчас): «Готово D из T (P%) · в работе A · в очереди Q» + `<progress>`,
|
||||
оценка времени по скорости последних завершённых; статус скана («Сканирование: прочитано X») и автоцикла.
|
||||
- `client.go`: `ListIPsPage`, `ListRegistryPage`, `ScanStatus`; длинный таймаут/отдельный клиент для clear и массовых операций.
|
||||
|
||||
### 7. Прочее
|
||||
- `routes`, `dto_admin.go`, `docs/API.md` (202/статус/пагинация/`dry_run`/`wait`), `docs/USAGE.md` (скан, автоцикл: фаза `scanning`,
|
||||
ожидание ≈ N×50 с/валидаторов, рычаги), `docs/DASHBOARD.md`, `docs/LOCAL_E2E.md`, `README.md`, `configs/*.example.yaml`
|
||||
(+ копии в `rxprod-compose/sources/` и docker-примеры).
|
||||
- `rxprod-compose/control-api.yaml` (живой конфиг) — при необходимости задать `max_run_seconds` автоцикла = 0 (уже 0) и
|
||||
`openstack.list_page_size`; по умолчанию достаточно дефолтов.
|
||||
- Опционально (не входит): переупорядочить `Tick` (сначала sweeps, потом `assignIdleValidators`) — экономит до 5 с на адрес (~10 %).
|
||||
|
||||
## Тесты (минимальные, в стиле существующих)
|
||||
- `internal/openstack`: пагинация по marker на моке (N=2500, `PageSize=200`), ретрай страницы при `ListFailures`, отмена контекста;
|
||||
классификация ретраемых ошибок.
|
||||
- `internal/orchestrator`: задание скана — single-flight, прогресс, `DryRun`, ошибка чтения ⇒ очередь не изменена, куски по 500,
|
||||
порядок по возрастанию IP, 6440 адресов за разумное время; автоцикл: фаза `scanning`, шаги с явным `now`, ошибка скана,
|
||||
`no_free_ips`, рестарт в фазе `scanning`, `Stop` отменяет скан, `Tick` не блокируется во время долгого скана (мок с задержкой).
|
||||
- `internal/db`: `ListIPsPage`/`ListRegistryPage` (фильтры, total, LIMIT до summary), `CountIPsBy*`, `AnyNonTerminalIP`,
|
||||
set-based `ClearQueue`, тест на 6440 строк (время и отсутствие N+1).
|
||||
- `internal/httpapi`: 202/409-семантика скана, `?wait=true`, `?dry_run=true`, `GET /ips/scan`, конверт пагинации, обратная
|
||||
совместимость без `limit`; обновить `TestScanFloatingIPsEndpoint` и счётчики в `auth_test.go`.
|
||||
- `internal/dashboard`: панель прогресса и остановка опроса, пагинация/pager, фильтры `/ips`, `scope=all`, Overview на ограниченных
|
||||
списках при тысячах `queued`, `hx-params`; обновить `fakeControlAPI` и затронутые тесты (`TestIPsScan` и др.).
|
||||
- `scripts/run-local-e2e.sh`: автоцикл проходит через фазу `scanning` (мок-пагинация), полный сценарий остаётся зелёным за 90 с.
|
||||
|
||||
## Верификация
|
||||
1. `go build ./... && go vet ./... && go test ./...` и `go test -race` для `orchestrator`, `openstack`, `httpapi`, `dashboard`, `db`.
|
||||
2. `scripts/run-local-e2e.sh` (с токенами) — зелёный, автоцикл проходит `scanning → running → waiting`.
|
||||
3. **Живой стенд, безопасно (чтение):** `POST /admin/ips/scan?dry_run=true` — скан проходит все страницы реального Neutron, прогресс
|
||||
идёт, итог ≈ 6440 свободных, очередь не меняется, дашборд не получает таймаута. Замер времени скана и нагрузки.
|
||||
4. **Живой стенд, по согласованию:** кнопка «Сканировать Floating IP» — адреса поставлены в очередь, проверки идут, `/ips`,
|
||||
`/registry`, «Обзор» остаются быстрыми (проверка размера ответов и времени), «Очистить всё» отрабатывает за секунды.
|
||||
Решение о реальной постановке 6440 адресов (≈18 ч проверок на 5 валидаторах) принимает пользователь; перед этим — бэкап БД.
|
||||
5. Автоцикл на живом стенде — только после п. 4 и с согласия пользователя; наблюдать фазу `scanning`, отсутствие блокировки `Tick`.
|
||||
6. Реальный браузер (Playwright из venv): прогресс скана, пагинация, фильтры, выбор «все N по фильтру», отсутствие JS-ошибок.
|
||||
@@ -0,0 +1,81 @@
|
||||
# Ревью и тестирование: скан Floating IP и автоцикл при тысячах адресов
|
||||
|
||||
> Дата: 2026-10-01 18:59 MSK · План: [2026-10-01_18-19_fip-scan-at-scale-plan.md](2026-10-01_18-19_fip-scan-at-scale-plan.md)
|
||||
> Статус: **реализовано и проверено на живом окружении**; реальная постановка 6440 адресов в очередь и включение автоцикла на живом стенде **не выполнялись** — ждут решения пользователя.
|
||||
|
||||
## Итог
|
||||
|
||||
Исходная ошибка («control-api недоступен: context deadline exceeded» при нажатии «Сканировать Floating IP») устранена: на живом стенде кнопка
|
||||
возвращает панель прогресса за 0,7 с, а сканирование реального Neutron (6441 Floating IP, 6440 свободных) проходит за ≈ 1,5–2 минуты в фоне
|
||||
с видимым прогрессом. Код написан двумя агентами параллельно (Sonnet 5.5): серверная часть и дашборд, ревью и все проверки — независимо (Sonnet 5.5, high).
|
||||
Найден и исправлен один дефект автоцикла и одна мелочь в клиенте OpenStack; остальные замечания — ограничения дизайна, перечислены ниже.
|
||||
|
||||
## Причина исходной ошибки
|
||||
|
||||
`ListFloatingIPs` запрашивал весь список одним запросом без `limit` и без таймаута: при ≈ 6,4 тыс. адресов Neutron отвечал дольше минуты, дашборд ждал 10 с,
|
||||
а запрос был привязан к `r.Context()` — при обрыве соединения скан отменялся и не мог завершиться. Ошибка нигде не логировалась.
|
||||
Побочные открытия: у клиента OpenStack вообще не было таймаутов и ретраев; `ListRegistry` был O(n²) (нет индекса `ip_queue(registry_id)`);
|
||||
`/ips` отдавал ≈ 8 МБ HTML, а «Обзор» каждые 5 с тянул ≈ 3 МБ JSON.
|
||||
|
||||
## Что реализовано
|
||||
|
||||
| Область | Изменения |
|
||||
|---|---|
|
||||
| OpenStack | `ListFreeFloatingIPs` — постраничное чтение по `marker` (200 на страницу, `fields=` сокращает ответ), повтор страницы с backoff на обрывы/`RemoteDisconnected`/5xx/429, таймаут запроса (`openstack.request_timeout_seconds`, 60 с — закрывает и зависания в тике оркестратора); `MockClient` с пагинацией, `ListFailures`, `PageDelay`, `SeedMany` |
|
||||
| Скан-задание | `orchestrator/scanjob.go`: single-flight фоновое задание на контексте процесса (не `r.Context()`), фазы `clearing → listing → enqueuing → done/error/cancelled`, прогресс; сначала полное обнаружение, затем `SubmitIPs` кусками по 500 по возрастанию IP; сбой чтения ⇒ очередь не меняется; `dry_run`; синхронная обёртка `ScanFloatingIPs` сохранена |
|
||||
| Автоцикл | фаза `scanning`: очистка и скан — одно фоновое задание, цикл оркестратора и `autoCycleMu` не блокируются; восстановление после рестарта; `Stop` отменяет скан; завершение определяется `EXISTS`, а не чтением всей очереди каждый тик |
|
||||
| БД | миграция `0009` (индексы `ip_queue(registry_id)`, `ip_queue(state, aggregated_at)`); `ListIPsPage`, `ListRegistryPage` (LIMIT/OFFSET до расчёта сводки, фильтр итога одним SQL), `CountIPsByState/Result`, `AnyNonTerminalIP`; `ClearAllIPs` — 5 запросов вместо цикла по адресам |
|
||||
| API | `POST /admin/ips/scan` → `202` (`dry_run`, `wait`), новый `GET /admin/ips/scan`, пагинация и фильтры у `GET /admin/ips` и `/admin/registry` (без `limit` — прежний массив), `results_by_overall` в `/admin/status`, `count` у `clear` |
|
||||
| Дашборд | панель прогресса скана и «Пробное сканирование»; постраничные `/ips` и `/registry` с серверными фильтрами; «Обзор» на счётчиках и ограниченных списках (прогресс «Готово D из T», оценка времени); «Выбрать все N по фильтру»; `hx-params` на кнопках (исправлен дефект — отмеченные адреса попадали в URL `hx-delete`); подтверждения с реальным числом; длинный таймаут для массовых операций |
|
||||
| Конфиг и документы | `openstack.list_page_size/request_timeout_seconds/list_page_retries`, `orchestrator.fip_scan_timeout_seconds`; `API.md`, `USAGE.md`, `DASHBOARD.md`, `README.md`, примеры конфигов |
|
||||
|
||||
## Результаты проверок
|
||||
|
||||
| Проверка | Результат |
|
||||
|---|---|
|
||||
| `gofmt`, `go build ./...`, `go vet ./...` | чисто |
|
||||
| `go test ./...` (включая тесты на 6440 адресов) | все пакеты зелёные |
|
||||
| `go test -race -short` (openstack, orchestrator, db, httpapi, dashboard, config) | зелёные (тесты на 6440 адресов под `-race` слишком долгие, пропускаются по `-short`; агент прогонял их полностью) |
|
||||
| `scripts/run-local-e2e.sh` (с токенами) | exit 0: автоцикл прошёл через `scanning`, проверки реальными агентом и пробером — `pass` |
|
||||
| **Живой стенд — пробный скан реального Neutron** (`POST …/scan?dry_run=true`) | 33 страницы, **6441 найдено / 6440 свободных за 93 с**, `202` за 2 мс, повторный `POST` присоединился к идущему заданию, **API отвечал за 2–4 мс всё время скана**, очередь осталась пустой |
|
||||
| **Живой дашборд (настоящий Chromium)**, кнопка «Пробное сканирование» | панель за 0,7 с без баннера ошибки, прогресс по страницам, итог «готово: 33 страницы, 6441, 6440, время 1 мин 48 с» |
|
||||
| Изолированный mock-стенд на **6440 адресах**, настоящий Chromium (17 из 19 автопроверок, 2 — ложные, см. ниже) | `/ips` **68 КБ за 0,18 с** (было ≈ 8 МБ), фрагмент «Обзора» **3 КБ за 0,05 с**, `/registry` 34 КБ за 0,12 с; пагинация и серверный поиск; прогресс скана (читаются страницы → ставятся в очередь → готово); «Очистить всё» с реальным числом в подтверждении — **0,5 с**; «Выбрать все 6440 по фильтру» + массовое удаление — **8 с**; JS-ошибок нет; ошибок в логе control-api нет |
|
||||
| Миграция `0009` на живой БД | `user_version = 9`, оба индекса созданы, данные не тронуты |
|
||||
|
||||
Примечание: два «FAIL» в браузерном скрипте mock-стенда — ошибка самого скрипта: панель показывает состояние заглавными («ГОТОВО», CSS), а скрипт искал строчные.
|
||||
Выведенный текст панели подтверждает успех (`добавлено6440`, `найдено адресов6440`). Реальных провалов нет.
|
||||
|
||||
## Замечания ревью
|
||||
|
||||
| № | Серьёзность | Замечание | Статус |
|
||||
|---|---|---|---|
|
||||
| 1 | средняя | **Автоцикл «усыновлял» чужое сканирование.** Если в момент старта цикла уже шло ручное/периодическое/**пробное** сканирование, `StartScan` возвращал `started=false`, а цикл переходил в `scanning` и ждал чужое задание. Пробное ничего не ставит в очередь, ручное не очищает очередь ⇒ цикл переходил в `running` над пустой/нетронутой очередью и сразу отчитывался `completed` (`runs_total+1`) без единой проверки | **исправлено**: если собственный скан не стартовал, цикл ничего не меняет и пробует снова на следующем такте (после окончания чужого); добавлен тест `TestAutoCycleWaitsForForeignScanInsteadOfFollowingIt` (падал до правки) |
|
||||
| 2 | низкая | Клиент OpenStack заполнял `ProjectID` только из `tenant_id`; при `fields=` Neutron может вернуть лишь `project_id` | **исправлено** (запасной вариант `project_id`); поле нигде не влияет на логику |
|
||||
| 3 | низкая | `aggregated_at_desc` сортирует по `strftime(...)` — временные метки хранятся как RFC3339Nano, и сырая сортировка текстом неверна (поймал тест агента); индекс `(state, aggregated_at)` помогает фильтру по состоянию, но не сортировке | принято; на 6440 строк незаметно |
|
||||
| 4 | низкая | `StopAutoCycle` в фазе `scanning` вызывает `CancelScan`, который ждёт до 5 с под `autoCycleMu`: «Выключить» может занять до 5 с, а следующий шаг цикла — подождать | принято |
|
||||
| 5 | низкая | Если скан завершился ошибкой после «Очистить» (шаг 1 цикла), очередь остаётся пустой до следующего цикла (`interval_seconds`); исход — `error` с причиной | принято, описано в `USAGE.md`; при желании — отдельная доработка (повтор скана сразу) |
|
||||
| 6 | низкая | Статус скана хранится в памяти: после рестарта control-api он `idle`; автоцикл в фазе `scanning` при этом корректно перезапускает скан | принято |
|
||||
| 7 | инфо | Отмена сканирования (`CancelScan`) вызывается только из `Stop` автоцикла; ручной кнопки/эндпоинта отмены нет | не входило в план |
|
||||
| 8 | инфо | Дашборд: мутации заменяют `#ips-table-wrap` целиком (`outerHTML`), чтобы `hx-get` обёртки всегда указывал на текущую страницу/фильтр; убраны функции `filterQueueItems/filterRegistryItems/currentlyChecking/lastCompleted` вместе с тестами (фильтрация перенесена на сервер); сводка «последние N» считается по показанному (возможно, отфильтрованному) окну, а общие итоги — отдельной строкой | принято |
|
||||
| 9 | инфо | Тесты на 6440 адресов под `-race` занимают 40–90 с на пакет — пропускаются по `-short` | принято |
|
||||
|
||||
## Пропускная способность (важно для автоцикла)
|
||||
|
||||
Проверка не стала быстрее — стало возможным её запустить. По фактическим данным стенда слот на адрес ≈ 50 с на валидатор (из них 30 с — `fip_settle_seconds`):
|
||||
**6440 адресов ≈ 18 ч на 5 валидаторах, ≈ 9 ч на 10, ≈ 4,5 ч на 20.** Для автоцикла `max_run_seconds` должен оставаться `0`. Рычаги — число валидаторов и (осторожно)
|
||||
`fip_settle_seconds`. На странице «Обзор» виден прогресс и оценка времени.
|
||||
|
||||
## Состояние живого стенда (`rxprod-compose`)
|
||||
|
||||
- Развёрнуты новые образы `civ-capi`, `civ-adash`, `civ-prober` (и пересобран `civ-agent`); миграция `0009` применена. Предыдущие образы сохранены под тегом `:pre-scale`
|
||||
(откат: `docker tag civ-capi:pre-scale civ-capi:latest` и `docker compose up -d`; на `0009` откат БД не нужен — это только индексы). Бэкап БД перед обновлением — в каталоге scratchpad сессии (`/tmp`, временный).
|
||||
- Очередь пуста (5 адресов прежней работы остались в реестре). **Реальная постановка 6440 адресов и автоцикл на живом стенде не запускались**: это ≈ 18 ч реальных проверок,
|
||||
решение за пользователем. Перед запуском рекомендую «Пробное сканирование» (уже отработало штатно) и бэкап БД.
|
||||
- Токен агентов по-прежнему не включён (внешние валидаторы и пробер `rxyc` со старыми бинарниками) — см. [ревью аутентификации](2026-10-01_11-31_authentication-review.md).
|
||||
Новые бинарники для внешних валидаторов — в `bin/` (после раскатки токена агентов их можно обновить одновременно).
|
||||
- `bin/` пересобран (`CGO_ENABLED=0`, `-trimpath -ldflags="-s -w"`), `SHA256SUMS` обновлён. Временные контейнеры `civ-scale-*` удалены.
|
||||
|
||||
## Что осталось
|
||||
|
||||
- Решение пользователя: поставить 6440 адресов в очередь («Сканировать Floating IP») и/или включить автоцикл на живом стенде.
|
||||
- По желанию: кнопка/эндпоинт отмены скана (замечание 7), повтор скана внутри цикла при ошибке (замечание 5), `Cache-Control: no-store` (из ревью аутентификации).
|
||||
@@ -69,6 +69,17 @@ type OpenStackConfig struct {
|
||||
UsernameEnv string `yaml:"username_env"` // default OS_USERNAME — used when auth_method: password
|
||||
UserDomainNameEnv string `yaml:"user_domain_name_env"` // default OS_USER_DOMAIN_NAME
|
||||
PasswordEnv string `yaml:"password_env"` // default OS_PASSWORD
|
||||
|
||||
// ListPageSize is how many floating IPs one Neutron list request asks
|
||||
// for (the scan reads the project page by page). Default 200.
|
||||
ListPageSize int `yaml:"list_page_size"`
|
||||
// RequestTimeoutSeconds bounds every single HTTP request to Keystone and
|
||||
// Neutron. Default 60.
|
||||
RequestTimeoutSeconds int `yaml:"request_timeout_seconds"`
|
||||
// ListPageRetries is how many times one failed page of the listing is
|
||||
// retried (backoff 1s,2s,4s,...) on network errors, 5xx and 429. Default
|
||||
// 5; a negative value disables retries.
|
||||
ListPageRetries int `yaml:"list_page_retries"`
|
||||
}
|
||||
|
||||
type OrchestratorConfig struct {
|
||||
@@ -93,6 +104,9 @@ type OrchestratorConfig struct {
|
||||
// POST /api/v1/admin/ips/scan or the dashboard's "Scan Floating IPs"
|
||||
// button either way.
|
||||
FIPScanIntervalSeconds int `yaml:"fip_scan_interval_seconds"`
|
||||
// FIPScanTimeoutSeconds is the overall deadline of one background
|
||||
// floating-IP scan (clear + paged read + enqueue). Default 1800.
|
||||
FIPScanTimeoutSeconds int `yaml:"fip_scan_timeout_seconds"`
|
||||
}
|
||||
|
||||
type AggregationConfig struct {
|
||||
@@ -164,6 +178,15 @@ func LoadControlAPI(path string) (*ControlAPI, error) {
|
||||
if c.OpenStack.PasswordEnv == "" {
|
||||
c.OpenStack.PasswordEnv = "OS_PASSWORD"
|
||||
}
|
||||
if c.OpenStack.ListPageSize == 0 {
|
||||
c.OpenStack.ListPageSize = 200
|
||||
}
|
||||
if c.OpenStack.RequestTimeoutSeconds == 0 {
|
||||
c.OpenStack.RequestTimeoutSeconds = 60
|
||||
}
|
||||
if c.OpenStack.ListPageRetries == 0 {
|
||||
c.OpenStack.ListPageRetries = 5
|
||||
}
|
||||
if c.Auth.AdminTokenEnv == "" {
|
||||
c.Auth.AdminTokenEnv = "CONTROL_API_ADMIN_TOKEN"
|
||||
}
|
||||
@@ -191,6 +214,9 @@ func LoadControlAPI(path string) (*ControlAPI, error) {
|
||||
if c.Orchestrator.HeartbeatTimeoutSeconds == 0 {
|
||||
c.Orchestrator.HeartbeatTimeoutSeconds = 30
|
||||
}
|
||||
if c.Orchestrator.FIPScanTimeoutSeconds == 0 {
|
||||
c.Orchestrator.FIPScanTimeoutSeconds = 1800
|
||||
}
|
||||
return &c, nil
|
||||
}
|
||||
|
||||
|
||||
@@ -0,0 +1,66 @@
|
||||
package config
|
||||
|
||||
import (
|
||||
"bytes"
|
||||
"os"
|
||||
"path/filepath"
|
||||
"testing"
|
||||
)
|
||||
|
||||
func TestLoadControlAPIScanDefaults(t *testing.T) {
|
||||
path := filepath.Join(t.TempDir(), "c.yaml")
|
||||
if err := os.WriteFile(path, []byte("server:\n listen_addr: \":8080\"\n"), 0o600); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
c, err := LoadControlAPI(path)
|
||||
if err != nil {
|
||||
t.Fatalf("load: %v", err)
|
||||
}
|
||||
if c.OpenStack.ListPageSize != 200 || c.OpenStack.RequestTimeoutSeconds != 60 ||
|
||||
c.OpenStack.ListPageRetries != 5 || c.Orchestrator.FIPScanTimeoutSeconds != 1800 {
|
||||
t.Fatalf("unexpected defaults: openstack=%+v orchestrator=%+v", c.OpenStack, c.Orchestrator)
|
||||
}
|
||||
}
|
||||
|
||||
func TestLoadControlAPIScanOverrides(t *testing.T) {
|
||||
path := filepath.Join(t.TempDir(), "c.yaml")
|
||||
yaml := "openstack:\n list_page_size: 50\n request_timeout_seconds: 10\n list_page_retries: -1\n" +
|
||||
"orchestrator:\n fip_scan_timeout_seconds: 99\n"
|
||||
if err := os.WriteFile(path, []byte(yaml), 0o600); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
c, err := LoadControlAPI(path)
|
||||
if err != nil {
|
||||
t.Fatalf("load: %v", err)
|
||||
}
|
||||
if c.OpenStack.ListPageSize != 50 || c.OpenStack.RequestTimeoutSeconds != 10 ||
|
||||
c.OpenStack.ListPageRetries != -1 || c.Orchestrator.FIPScanTimeoutSeconds != 99 {
|
||||
t.Fatalf("overrides lost: openstack=%+v orchestrator=%+v", c.OpenStack, c.Orchestrator)
|
||||
}
|
||||
}
|
||||
|
||||
// The shipped example must load and carry the scan settings, and the rxprod
|
||||
// copy must stay byte-identical to it.
|
||||
func TestControlAPIExampleConfigs(t *testing.T) {
|
||||
c, err := LoadControlAPI("../../configs/control-api.example.yaml")
|
||||
if err != nil {
|
||||
t.Fatalf("load example: %v", err)
|
||||
}
|
||||
if c.OpenStack.ListPageSize != 200 || c.Orchestrator.FIPScanTimeoutSeconds != 1800 {
|
||||
t.Fatalf("example scan settings: %+v %+v", c.OpenStack, c.Orchestrator)
|
||||
}
|
||||
a, err := os.ReadFile("../../configs/control-api.example.yaml")
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
b, err := os.ReadFile("../../rxprod-compose/sources/control-api.example.yaml")
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if !bytes.Equal(a, b) {
|
||||
t.Fatalf("rxprod-compose/sources/control-api.example.yaml differs from configs/control-api.example.yaml")
|
||||
}
|
||||
if _, err := LoadControlAPI("../../deploy/docker/control-api/control-api.docker.example.yaml"); err != nil {
|
||||
t.Fatalf("load docker example: %v", err)
|
||||
}
|
||||
}
|
||||
+130
-20
@@ -8,6 +8,8 @@ import (
|
||||
"io"
|
||||
"net/http"
|
||||
"net/url"
|
||||
"strconv"
|
||||
"strings"
|
||||
"time"
|
||||
)
|
||||
|
||||
@@ -39,15 +41,44 @@ func (e *apiErr) Error() string {
|
||||
type client struct {
|
||||
baseURL string
|
||||
http *http.Client
|
||||
// long is used for clear / bulk operations, which legitimately take far
|
||||
// longer than a plain read (thousands of rows in one transaction): same
|
||||
// transport, but a longer whole-call timeout.
|
||||
long *http.Client
|
||||
// token, when non-empty, is sent to control-api as a Bearer credential.
|
||||
token string
|
||||
}
|
||||
|
||||
// longCallTimeout is the minimum whole-call timeout for clear and bulk
|
||||
// operations (ClearQueue, DeleteIPs, SubmitIPs).
|
||||
const longCallTimeout = 120 * time.Second
|
||||
|
||||
func newClient(baseURL string, timeout time.Duration) *client {
|
||||
return &client{baseURL: baseURL, http: &http.Client{Timeout: timeout}}
|
||||
longT := longCallTimeout
|
||||
if timeout > longT {
|
||||
longT = timeout
|
||||
}
|
||||
return &client{
|
||||
baseURL: baseURL,
|
||||
http: &http.Client{Timeout: timeout},
|
||||
long: &http.Client{Timeout: longT},
|
||||
}
|
||||
}
|
||||
|
||||
func (c *client) do(ctx context.Context, method, path string, body, out interface{}) error {
|
||||
return c.doWith(ctx, c.http, method, path, body, out)
|
||||
}
|
||||
|
||||
// doLong is do with the long (clear/bulk) timeout.
|
||||
func (c *client) doLong(ctx context.Context, method, path string, body, out interface{}) error {
|
||||
hc := c.long
|
||||
if hc == nil {
|
||||
hc = c.http
|
||||
}
|
||||
return c.doWith(ctx, hc, method, path, body, out)
|
||||
}
|
||||
|
||||
func (c *client) doWith(ctx context.Context, hc *http.Client, method, path string, body, out interface{}) error {
|
||||
var reader io.Reader
|
||||
if body != nil {
|
||||
b, err := json.Marshal(body)
|
||||
@@ -67,7 +98,7 @@ func (c *client) do(ctx context.Context, method, path string, body, out interfac
|
||||
req.Header.Set("Authorization", "Bearer "+c.token)
|
||||
}
|
||||
|
||||
resp, err := c.http.Do(req)
|
||||
resp, err := hc.Do(req)
|
||||
if err != nil {
|
||||
return &apiErr{Status: 0, Message: err.Error()}
|
||||
}
|
||||
@@ -96,9 +127,58 @@ func (c *client) Status(ctx context.Context) (statusResponse, error) {
|
||||
return out, err
|
||||
}
|
||||
|
||||
func (c *client) ListIPs(ctx context.Context) ([]ipQueueItem, error) {
|
||||
var out []ipQueueItem
|
||||
err := c.do(ctx, http.MethodGet, "/api/v1/admin/ips", nil, &out)
|
||||
// maxPageLimit is control-api's cap on `limit`.
|
||||
const maxPageLimit = 1000
|
||||
|
||||
// clampLimit keeps limit within 1..maxPageLimit: a request without `limit`
|
||||
// would make control-api answer with the legacy unbounded bare array.
|
||||
func clampLimit(limit int) int {
|
||||
if limit < 1 {
|
||||
return 1
|
||||
}
|
||||
if limit > maxPageLimit {
|
||||
return maxPageLimit
|
||||
}
|
||||
return limit
|
||||
}
|
||||
|
||||
// ipsQuery selects one page of GET /admin/ips: server-side filters plus
|
||||
// limit/offset. Order is "sequence" (default) or "aggregated_at_desc".
|
||||
type ipsQuery struct {
|
||||
States []string
|
||||
Q string
|
||||
Result string
|
||||
Order string
|
||||
Limit int
|
||||
Offset int
|
||||
}
|
||||
|
||||
func (q ipsQuery) values() url.Values {
|
||||
v := url.Values{}
|
||||
v.Set("limit", strconv.Itoa(clampLimit(q.Limit)))
|
||||
if q.Offset > 0 {
|
||||
v.Set("offset", strconv.Itoa(q.Offset))
|
||||
}
|
||||
if len(q.States) > 0 {
|
||||
v.Set("state", strings.Join(q.States, ","))
|
||||
}
|
||||
if q.Q != "" {
|
||||
v.Set("q", q.Q)
|
||||
}
|
||||
if q.Result != "" {
|
||||
v.Set("result", q.Result)
|
||||
}
|
||||
if q.Order != "" {
|
||||
v.Set("order", q.Order)
|
||||
}
|
||||
return v
|
||||
}
|
||||
|
||||
// ListIPsPage returns one page of the check queue plus the total number of
|
||||
// rows matching the filter. Never loads the whole queue.
|
||||
func (c *client) ListIPsPage(ctx context.Context, q ipsQuery) (ipsPage, error) {
|
||||
var out ipsPage
|
||||
err := c.do(ctx, http.MethodGet, "/api/v1/admin/ips?"+q.values().Encode(), nil, &out)
|
||||
return out, err
|
||||
}
|
||||
|
||||
@@ -112,7 +192,7 @@ func (c *client) GetIP(ctx context.Context, ip string) (ipDetailResponse, error)
|
||||
// forcing a recheck of already-finished ones — see docs/API.md.
|
||||
func (c *client) SubmitIPs(ctx context.Context, addresses []string) (submitIPsResponse, error) {
|
||||
var out submitIPsResponse
|
||||
err := c.do(ctx, http.MethodPost, "/api/v1/admin/ips", map[string][]string{"addresses": addresses}, &out)
|
||||
err := c.doLong(ctx, http.MethodPost, "/api/v1/admin/ips", map[string][]string{"addresses": addresses}, &out)
|
||||
return out, err
|
||||
}
|
||||
|
||||
@@ -129,7 +209,7 @@ func (c *client) DeleteIP(ctx context.Context, ip string) error {
|
||||
// DeleteIPs permanently removes a specific list of addresses in one call.
|
||||
func (c *client) DeleteIPs(ctx context.Context, addresses []string) (deleteIPsResponse, error) {
|
||||
var out deleteIPsResponse
|
||||
err := c.do(ctx, http.MethodPost, "/api/v1/admin/ips/delete", map[string][]string{"addresses": addresses}, &out)
|
||||
err := c.doLong(ctx, http.MethodPost, "/api/v1/admin/ips/delete", map[string][]string{"addresses": addresses}, &out)
|
||||
return out, err
|
||||
}
|
||||
|
||||
@@ -137,25 +217,55 @@ func (c *client) DeleteIPs(ctx context.Context, addresses []string) (deleteIPsRe
|
||||
// including those actively being checked.
|
||||
func (c *client) ClearQueue(ctx context.Context) (clearQueueResponse, error) {
|
||||
var out clearQueueResponse
|
||||
err := c.do(ctx, http.MethodPost, "/api/v1/admin/ips/clear", nil, &out)
|
||||
err := c.doLong(ctx, http.MethodPost, "/api/v1/admin/ips/clear", nil, &out)
|
||||
return out, err
|
||||
}
|
||||
|
||||
// ScanFloatingIPs lists the OpenStack project's free (unassociated)
|
||||
// floating IPs and submits them to the check queue — see
|
||||
// orchestrator.ScanFloatingIPs.
|
||||
func (c *client) ScanFloatingIPs(ctx context.Context) (scanIPsResponse, error) {
|
||||
var out scanIPsResponse
|
||||
err := c.do(ctx, http.MethodPost, "/api/v1/admin/ips/scan", nil, &out)
|
||||
// StartScan starts control-api's background floating-IP scan (or joins the
|
||||
// one already running) and returns immediately with its current status.
|
||||
func (c *client) StartScan(ctx context.Context, dryRun bool) (scanStatusDTO, error) {
|
||||
var out scanStatusDTO
|
||||
path := "/api/v1/admin/ips/scan"
|
||||
if dryRun {
|
||||
path += "?dry_run=true"
|
||||
}
|
||||
err := c.do(ctx, http.MethodPost, path, nil, &out)
|
||||
return out, err
|
||||
}
|
||||
|
||||
// ListRegistry returns every address ever submitted to the check queue,
|
||||
// each with a summary of its accumulated check history — survives an
|
||||
// address being deleted from the queue and later re-added.
|
||||
func (c *client) ListRegistry(ctx context.Context) ([]registryItem, error) {
|
||||
var out []registryItem
|
||||
err := c.do(ctx, http.MethodGet, "/api/v1/admin/registry", nil, &out)
|
||||
// ScanStatus returns the progress of the background scan job.
|
||||
func (c *client) ScanStatus(ctx context.Context) (scanStatusDTO, error) {
|
||||
var out scanStatusDTO
|
||||
err := c.do(ctx, http.MethodGet, "/api/v1/admin/ips/scan", nil, &out)
|
||||
return out, err
|
||||
}
|
||||
|
||||
// registryQuery selects one page of GET /admin/registry.
|
||||
type registryQuery struct {
|
||||
Q string
|
||||
LastResult string
|
||||
Limit int
|
||||
Offset int
|
||||
}
|
||||
|
||||
// ListRegistryPage returns one page of the registry (every address ever
|
||||
// submitted to the check queue, each with a summary of its accumulated check
|
||||
// history — survives an address being deleted from the queue and later
|
||||
// re-added) plus the total number of rows matching the filter.
|
||||
func (c *client) ListRegistryPage(ctx context.Context, q registryQuery) (registryPage, error) {
|
||||
v := url.Values{}
|
||||
v.Set("limit", strconv.Itoa(clampLimit(q.Limit)))
|
||||
if q.Offset > 0 {
|
||||
v.Set("offset", strconv.Itoa(q.Offset))
|
||||
}
|
||||
if q.Q != "" {
|
||||
v.Set("q", q.Q)
|
||||
}
|
||||
if q.LastResult != "" {
|
||||
v.Set("last_result", q.LastResult)
|
||||
}
|
||||
var out registryPage
|
||||
err := c.do(ctx, http.MethodGet, "/api/v1/admin/registry?"+v.Encode(), nil, &out)
|
||||
return out, err
|
||||
}
|
||||
|
||||
|
||||
@@ -9,6 +9,8 @@ import (
|
||||
"net/http/httptest"
|
||||
"net/url"
|
||||
"os"
|
||||
"sort"
|
||||
"strconv"
|
||||
"strings"
|
||||
"sync"
|
||||
"testing"
|
||||
@@ -42,6 +44,27 @@ type fakeControlAPI struct {
|
||||
registry map[string]registryItem
|
||||
registryChecks map[string][]check
|
||||
|
||||
// Scan job state machine (see the scan handlers): POST starts a job that
|
||||
// stays "running" for scanRunPolls GET polls (0 = finishes at once), then
|
||||
// ends as done — or as error when scanFinalError is set. scan is the status
|
||||
// served by GET; tests may also set it directly (with scanPollsLeft == 0 it
|
||||
// stays as is). scanStartStatus != 0 makes POST fail with that HTTP status.
|
||||
scan scanStatusDTO
|
||||
scanRunPolls int
|
||||
scanPollsLeft int
|
||||
scanFinalError string
|
||||
scanStartStatus int
|
||||
|
||||
// Requests seen on the list endpoints, for "never loads everything" checks:
|
||||
// the raw query of every GET /ips (ipsQueries) and the number of GET
|
||||
// /registry calls without `limit` (bare), plus the sizes of the DeleteIPs /
|
||||
// SubmitIPs bulk calls.
|
||||
ipsQueries []string
|
||||
registryQueries []string
|
||||
bareIPsCalls int
|
||||
deleteChunks []int
|
||||
submitChunks []int
|
||||
|
||||
// autoCycle is the state served by /api/v1/admin/auto-cycle*;
|
||||
// autoCycleDown makes all four endpoints answer 500 (unavailable API).
|
||||
autoCycle autoCycleDTO
|
||||
@@ -86,16 +109,72 @@ func (f *fakeControlAPI) handler() http.Handler {
|
||||
f.mu.Lock()
|
||||
defer f.mu.Unlock()
|
||||
byState := map[string]int{}
|
||||
results := map[string]int{"pass": 0, "partial": 0, "fail": 0, "cancelled": 0}
|
||||
for _, ip := range f.ips {
|
||||
byState[ip.State]++
|
||||
if ip.OverallResult != "" {
|
||||
results[ip.OverallResult]++
|
||||
}
|
||||
writeJSON(w, http.StatusOK, statusResponse{TotalIPs: len(f.ips), IPsByState: byState, TotalValidators: len(f.validators)})
|
||||
}
|
||||
writeJSON(w, http.StatusOK, statusResponse{TotalIPs: len(f.ips), IPsByState: byState, TotalValidators: len(f.validators), ResultsByOverall: results})
|
||||
})
|
||||
|
||||
mux.HandleFunc("GET /api/v1/admin/ips", func(w http.ResponseWriter, r *http.Request) {
|
||||
f.mu.Lock()
|
||||
defer f.mu.Unlock()
|
||||
f.ipsQueries = append(f.ipsQueries, r.URL.RawQuery)
|
||||
qv := r.URL.Query()
|
||||
if qv.Get("limit") == "" {
|
||||
// Legacy shape: the whole queue as a bare array.
|
||||
f.bareIPsCalls++
|
||||
writeJSON(w, http.StatusOK, f.ips)
|
||||
return
|
||||
}
|
||||
limit, err := strconv.Atoi(qv.Get("limit"))
|
||||
if err != nil || limit < 1 || limit > 1000 {
|
||||
writeAPIErr(w, http.StatusBadRequest, "limit must be 1..1000")
|
||||
return
|
||||
}
|
||||
offset, _ := strconv.Atoi(qv.Get("offset"))
|
||||
var states map[string]bool
|
||||
if st := qv.Get("state"); st != "" {
|
||||
states = map[string]bool{}
|
||||
for _, x := range strings.Split(st, ",") {
|
||||
states[x] = true
|
||||
}
|
||||
}
|
||||
var matched []ipQueueItem
|
||||
for _, ip := range f.ips {
|
||||
if states != nil && !states[ip.State] {
|
||||
continue
|
||||
}
|
||||
if q := qv.Get("q"); q != "" && !strings.Contains(strings.ToLower(ip.IPAddress), strings.ToLower(q)) {
|
||||
continue
|
||||
}
|
||||
if res := qv.Get("result"); res != "" && ip.OverallResult != res {
|
||||
continue
|
||||
}
|
||||
matched = append(matched, ip)
|
||||
}
|
||||
if qv.Get("order") == "aggregated_at_desc" {
|
||||
sort.SliceStable(matched, func(i, j int) bool {
|
||||
a, b := matched[i].AggregatedAt, matched[j].AggregatedAt
|
||||
if a == nil || b == nil {
|
||||
return a != nil && b == nil
|
||||
}
|
||||
return a.After(*b)
|
||||
})
|
||||
} else {
|
||||
sort.SliceStable(matched, func(i, j int) bool { return matched[i].Sequence < matched[j].Sequence })
|
||||
}
|
||||
page := []ipQueueItem{}
|
||||
if offset < len(matched) {
|
||||
page = matched[offset:]
|
||||
if len(page) > limit {
|
||||
page = page[:limit]
|
||||
}
|
||||
}
|
||||
writeJSON(w, http.StatusOK, ipsPage{Items: page, Total: len(matched), Limit: limit, Offset: offset})
|
||||
})
|
||||
|
||||
mux.HandleFunc("GET /api/v1/admin/ips/{ip}", func(w http.ResponseWriter, r *http.Request) {
|
||||
@@ -122,6 +201,7 @@ func (f *fakeControlAPI) handler() http.Handler {
|
||||
writeAPIErr(w, http.StatusBadRequest, "addresses must not be empty")
|
||||
return
|
||||
}
|
||||
f.submitChunks = append(f.submitChunks, len(req.Addresses))
|
||||
resp := submitIPsResponse{}
|
||||
for _, addr := range req.Addresses {
|
||||
idx := f.findIP(addr)
|
||||
@@ -190,6 +270,7 @@ func (f *fakeControlAPI) handler() http.Handler {
|
||||
writeAPIErr(w, http.StatusBadRequest, "addresses must not be empty")
|
||||
return
|
||||
}
|
||||
f.deleteChunks = append(f.deleteChunks, len(req.Addresses))
|
||||
resp := deleteIPsResponse{}
|
||||
for _, addr := range req.Addresses {
|
||||
idx := f.findIP(addr)
|
||||
@@ -210,6 +291,7 @@ func (f *fakeControlAPI) handler() http.Handler {
|
||||
for _, ip := range f.ips {
|
||||
resp.Deleted = append(resp.Deleted, ip.IPAddress)
|
||||
}
|
||||
resp.Count = len(resp.Deleted)
|
||||
f.ips = nil
|
||||
writeJSON(w, http.StatusOK, resp)
|
||||
})
|
||||
@@ -246,18 +328,32 @@ func (f *fakeControlAPI) handler() http.Handler {
|
||||
mux.HandleFunc("POST /api/v1/admin/ips/scan", func(w http.ResponseWriter, r *http.Request) {
|
||||
f.mu.Lock()
|
||||
defer f.mu.Unlock()
|
||||
resp := scanIPsResponse{ScannedFree: len(f.scanFreeAddresses)}
|
||||
for _, addr := range f.scanFreeAddresses {
|
||||
idx := f.findIP(addr)
|
||||
if idx < 0 {
|
||||
if f.scanStartStatus != 0 {
|
||||
writeAPIErr(w, f.scanStartStatus, "scan refused")
|
||||
return
|
||||
}
|
||||
if !f.scan.Running {
|
||||
now := time.Now()
|
||||
f.ips = append(f.ips, ipQueueItem{IPAddress: addr, State: "queued", CreatedAt: now, UpdatedAt: now})
|
||||
resp.Added = append(resp.Added, addr)
|
||||
continue
|
||||
f.scan = scanStatusDTO{State: "listing", Running: true, DryRun: r.URL.Query().Get("dry_run") == "true", StartedAt: &now}
|
||||
f.scanPollsLeft = f.scanRunPolls
|
||||
if f.scanPollsLeft == 0 {
|
||||
f.finishScan()
|
||||
}
|
||||
resp.Reordered = append(resp.Reordered, addr)
|
||||
}
|
||||
writeJSON(w, http.StatusOK, resp)
|
||||
writeJSON(w, http.StatusAccepted, f.scan)
|
||||
})
|
||||
mux.HandleFunc("GET /api/v1/admin/ips/scan", func(w http.ResponseWriter, r *http.Request) {
|
||||
f.mu.Lock()
|
||||
defer f.mu.Unlock()
|
||||
if f.scan.Running && f.scanPollsLeft > 0 {
|
||||
f.scanPollsLeft--
|
||||
if f.scanPollsLeft == 0 {
|
||||
f.finishScan()
|
||||
} else {
|
||||
f.scan.Pages++
|
||||
}
|
||||
}
|
||||
writeJSON(w, http.StatusOK, f.scan)
|
||||
})
|
||||
|
||||
mux.HandleFunc("GET /api/v1/admin/auto-cycle", func(w http.ResponseWriter, r *http.Request) {
|
||||
@@ -329,11 +425,37 @@ func (f *fakeControlAPI) handler() http.Handler {
|
||||
mux.HandleFunc("GET /api/v1/admin/registry", func(w http.ResponseWriter, r *http.Request) {
|
||||
f.mu.Lock()
|
||||
defer f.mu.Unlock()
|
||||
out := make([]registryItem, 0, len(f.registry))
|
||||
qv := r.URL.Query()
|
||||
f.registryQueries = append(f.registryQueries, r.URL.RawQuery)
|
||||
matched := make([]registryItem, 0, len(f.registry))
|
||||
for _, item := range f.registry {
|
||||
out = append(out, item)
|
||||
if q := qv.Get("q"); q != "" && !strings.Contains(strings.ToLower(item.IPAddress), strings.ToLower(q)) {
|
||||
continue
|
||||
}
|
||||
writeJSON(w, http.StatusOK, out)
|
||||
if lr := qv.Get("last_result"); lr != "" && item.LastResult != lr {
|
||||
continue
|
||||
}
|
||||
matched = append(matched, item)
|
||||
}
|
||||
sort.Slice(matched, func(i, j int) bool { return matched[i].IPAddress < matched[j].IPAddress })
|
||||
if qv.Get("limit") == "" {
|
||||
writeJSON(w, http.StatusOK, matched)
|
||||
return
|
||||
}
|
||||
limit, err := strconv.Atoi(qv.Get("limit"))
|
||||
if err != nil || limit < 1 || limit > 1000 {
|
||||
writeAPIErr(w, http.StatusBadRequest, "limit must be 1..1000")
|
||||
return
|
||||
}
|
||||
offset, _ := strconv.Atoi(qv.Get("offset"))
|
||||
page := []registryItem{}
|
||||
if offset < len(matched) {
|
||||
page = matched[offset:]
|
||||
if len(page) > limit {
|
||||
page = page[:limit]
|
||||
}
|
||||
}
|
||||
writeJSON(w, http.StatusOK, registryPage{Items: page, Total: len(matched), Limit: limit, Offset: offset})
|
||||
})
|
||||
mux.HandleFunc("GET /api/v1/admin/registry/{ip}", func(w http.ResponseWriter, r *http.Request) {
|
||||
f.mu.Lock()
|
||||
@@ -544,6 +666,51 @@ func (f *fakeControlAPI) handler() http.Handler {
|
||||
})
|
||||
}
|
||||
|
||||
// finishScan ends the running scan job (caller holds f.mu): the discovered
|
||||
// free addresses are queued unless it was a dry run, or the job fails with
|
||||
// scanFinalError.
|
||||
func (f *fakeControlAPI) finishScan() {
|
||||
now := time.Now()
|
||||
f.scan.Running = false
|
||||
f.scan.FinishedAt = &now
|
||||
f.scan.Discovered = len(f.scanFreeAddresses)
|
||||
f.scan.Free = len(f.scanFreeAddresses)
|
||||
if f.scanFinalError != "" {
|
||||
f.scan.State = "error"
|
||||
f.scan.Error = f.scanFinalError
|
||||
return
|
||||
}
|
||||
f.scan.State = "done"
|
||||
if f.scan.DryRun {
|
||||
return
|
||||
}
|
||||
for _, addr := range f.scanFreeAddresses {
|
||||
if f.findIP(addr) < 0 {
|
||||
f.ips = append(f.ips, ipQueueItem{IPAddress: addr, State: "queued", Sequence: len(f.ips) + 1, CreatedAt: now, UpdatedAt: now})
|
||||
f.scan.Added++
|
||||
} else {
|
||||
f.scan.Reordered++
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// seedIPs appends n queued addresses 10.x.y.z (distinct, in sequence order)
|
||||
// and returns them. Use it for tests that need pages' worth of rows.
|
||||
func (f *fakeControlAPI) seedIPs(n int) []string {
|
||||
f.mu.Lock()
|
||||
defer f.mu.Unlock()
|
||||
now := time.Now()
|
||||
out := make([]string, 0, n)
|
||||
base := len(f.ips)
|
||||
for i := 0; i < n; i++ {
|
||||
k := base + i + 1
|
||||
addr := fmt.Sprintf("10.%d.%d.%d", k/65536, (k/256)%256, k%256)
|
||||
f.ips = append(f.ips, ipQueueItem{IPAddress: addr, State: "queued", Sequence: k, CreatedAt: now, UpdatedAt: now})
|
||||
out = append(out, addr)
|
||||
}
|
||||
return out
|
||||
}
|
||||
|
||||
func (f *fakeControlAPI) findIP(addr string) int {
|
||||
for i, ip := range f.ips {
|
||||
if ip.IPAddress == addr {
|
||||
@@ -596,3 +763,14 @@ func postForm(t *testing.T, ts *httptest.Server, method, path string, form url.V
|
||||
body, _ := io.ReadAll(resp.Body)
|
||||
return string(body)
|
||||
}
|
||||
|
||||
// newSlowAPI serves an empty JSON object for every request after delay.
|
||||
func newSlowAPI(t *testing.T, delay time.Duration) string {
|
||||
t.Helper()
|
||||
ts := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
|
||||
time.Sleep(delay)
|
||||
writeJSON(w, http.StatusOK, map[string]interface{}{})
|
||||
}))
|
||||
t.Cleanup(ts.Close)
|
||||
return ts.URL
|
||||
}
|
||||
+151
-6
@@ -1,6 +1,7 @@
|
||||
package dashboard
|
||||
|
||||
import (
|
||||
"fmt"
|
||||
"strconv"
|
||||
"time"
|
||||
)
|
||||
@@ -20,6 +21,27 @@ type statusResponse struct {
|
||||
TotalIPs int `json:"total_ips"`
|
||||
IPsByState map[string]int `json:"ips_by_state"`
|
||||
TotalValidators int `json:"total_validators"`
|
||||
// ResultsByOverall counts finished addresses by overall result
|
||||
// (pass/partial/fail/cancelled).
|
||||
ResultsByOverall map[string]int `json:"results_by_overall"`
|
||||
}
|
||||
|
||||
// Terminal queue states: the address needs no further processing. Shared by
|
||||
// the overview progress indicator; "occupied" counts as terminal too (the
|
||||
// check cycle never ran because the floating IP was already bound).
|
||||
var terminalStates = []string{"done", "failed", "occupied"}
|
||||
|
||||
// activeStates are the states of an address that is being worked on right
|
||||
// now (everything between "queued" and a terminal state).
|
||||
var activeStates = []string{"assigning_fip", "awaiting_self_check", "checking", "aggregating"}
|
||||
|
||||
// sumStates adds up the counts of the given states in a status breakdown.
|
||||
func sumStates(byState map[string]int, states []string) int {
|
||||
n := 0
|
||||
for _, s := range states {
|
||||
n += byState[s]
|
||||
}
|
||||
return n
|
||||
}
|
||||
|
||||
type ipQueueItem struct {
|
||||
@@ -99,14 +121,135 @@ type deleteIPsResponse struct {
|
||||
|
||||
type clearQueueResponse struct {
|
||||
Deleted []string `json:"deleted"`
|
||||
Count int `json:"count"`
|
||||
}
|
||||
|
||||
type scanIPsResponse struct {
|
||||
ScannedFree int `json:"scanned_free"`
|
||||
Added []string `json:"added"`
|
||||
Requeued []string `json:"requeued"`
|
||||
Reordered []string `json:"reordered"`
|
||||
SkippedInProgress []string `json:"skipped_in_progress"`
|
||||
// ipsPage is the paginated envelope of GET /admin/ips (sent when the request
|
||||
// carries `limit`).
|
||||
type ipsPage struct {
|
||||
Items []ipQueueItem `json:"items"`
|
||||
Total int `json:"total"`
|
||||
Limit int `json:"limit"`
|
||||
Offset int `json:"offset"`
|
||||
}
|
||||
|
||||
// registryPage is the paginated envelope of GET /admin/registry.
|
||||
type registryPage struct {
|
||||
Items []registryItem `json:"items"`
|
||||
Total int `json:"total"`
|
||||
Limit int `json:"limit"`
|
||||
Offset int `json:"offset"`
|
||||
}
|
||||
|
||||
// scanStatusDTO is the state of control-api's background floating-IP scan job
|
||||
// (POST/GET /api/v1/admin/ips/scan).
|
||||
type scanStatusDTO struct {
|
||||
State string `json:"state"`
|
||||
Running bool `json:"running"`
|
||||
DryRun bool `json:"dry_run"`
|
||||
Pages int `json:"pages"`
|
||||
Discovered int `json:"discovered"`
|
||||
Free int `json:"free"`
|
||||
Added int `json:"added"`
|
||||
Requeued int `json:"requeued"`
|
||||
Reordered int `json:"reordered"`
|
||||
SkippedInProgress int `json:"skipped_in_progress"`
|
||||
StartedAt *time.Time `json:"started_at"`
|
||||
FinishedAt *time.Time `json:"finished_at"`
|
||||
Error string `json:"error"`
|
||||
}
|
||||
|
||||
// Finished reports a job that has run to a terminal state (as opposed to
|
||||
// "idle" = never started, or still running).
|
||||
func (s scanStatusDTO) Finished() bool {
|
||||
if s.Running {
|
||||
return false
|
||||
}
|
||||
switch s.State {
|
||||
case "done", "error", "cancelled":
|
||||
return true
|
||||
}
|
||||
return false
|
||||
}
|
||||
|
||||
// StateLabel is the Russian description of the job's state.
|
||||
func (s scanStatusDTO) StateLabel() string {
|
||||
switch s.State {
|
||||
case "clearing":
|
||||
return "очистка"
|
||||
case "listing":
|
||||
return "читаются страницы"
|
||||
case "enqueuing":
|
||||
return "ставятся в очередь"
|
||||
case "done":
|
||||
return "готово"
|
||||
case "error":
|
||||
return "ошибка"
|
||||
case "cancelled":
|
||||
return "отменено"
|
||||
case "idle", "":
|
||||
return "нет активного сканирования"
|
||||
default:
|
||||
return s.State
|
||||
}
|
||||
}
|
||||
|
||||
// PillClass picks the pill style for the state.
|
||||
func (s scanStatusDTO) PillClass() string {
|
||||
switch s.State {
|
||||
case "done":
|
||||
return "pill-success"
|
||||
case "error":
|
||||
return "pill-danger"
|
||||
case "cancelled":
|
||||
return "pill-cancel"
|
||||
case "clearing", "listing", "enqueuing":
|
||||
return "pill-info"
|
||||
default:
|
||||
return "pill-neutral"
|
||||
}
|
||||
}
|
||||
|
||||
// Handled is how many of the free addresses the enqueuing phase has already
|
||||
// processed.
|
||||
func (s scanStatusDTO) Handled() int {
|
||||
return s.Added + s.Requeued + s.Reordered + s.SkippedInProgress
|
||||
}
|
||||
|
||||
// Indeterminate is true while the amount of work is not known yet.
|
||||
func (s scanStatusDTO) Indeterminate() bool {
|
||||
return s.State == "clearing" || s.State == "listing" || (s.State == "enqueuing" && s.Free <= 0)
|
||||
}
|
||||
|
||||
// Elapsed is the human-readable run time: until now while running, until
|
||||
// finished_at afterwards; empty when the job never started.
|
||||
func (s scanStatusDTO) Elapsed() string {
|
||||
if s.StartedAt == nil {
|
||||
return ""
|
||||
}
|
||||
end := time.Now()
|
||||
if !s.Running && s.FinishedAt != nil {
|
||||
end = *s.FinishedAt
|
||||
}
|
||||
return fmtDuration(end.Sub(*s.StartedAt))
|
||||
}
|
||||
|
||||
// fmtDuration renders a duration in Russian as "2 ч 05 мин", "3 мин 07 с" or
|
||||
// "42 с" — coarse on purpose (progress/ETA display).
|
||||
func fmtDuration(d time.Duration) string {
|
||||
if d < 0 {
|
||||
d = 0
|
||||
}
|
||||
sec := int(d.Round(time.Second) / time.Second)
|
||||
h, m, sc := sec/3600, (sec%3600)/60, sec%60
|
||||
switch {
|
||||
case h > 0:
|
||||
return fmt.Sprintf("%d ч %02d мин", h, m)
|
||||
case m > 0:
|
||||
return fmt.Sprintf("%d мин %02d с", m, sc)
|
||||
default:
|
||||
return fmt.Sprintf("%d с", sc)
|
||||
}
|
||||
}
|
||||
|
||||
// registryItem is one row of the durable per-address registry — see
|
||||
@@ -200,6 +343,8 @@ func (a autoCycleDTO) MaxRunMinutes() string { return secondsToMinutes(a.MaxRu
|
||||
// PhaseLabel is the Russian description of the current phase.
|
||||
func (a autoCycleDTO) PhaseLabel() string {
|
||||
switch a.Phase {
|
||||
case "scanning":
|
||||
return "сканирование Floating IP"
|
||||
case "running":
|
||||
return "идёт проверка"
|
||||
case "waiting":
|
||||
|
||||
@@ -1,14 +1,122 @@
|
||||
package dashboard
|
||||
|
||||
import (
|
||||
"context"
|
||||
"errors"
|
||||
"fmt"
|
||||
"net/http"
|
||||
"net/url"
|
||||
"strings"
|
||||
)
|
||||
|
||||
// bulkChunk is how many addresses go into one DeleteIPs/SubmitIPs call when
|
||||
// an operation spans a whole filter ("scope=all").
|
||||
const bulkChunk = 500
|
||||
|
||||
// ipStates are the queue states control-api accepts in the `state` filter.
|
||||
var ipStates = []string{"queued", "assigning_fip", "awaiting_self_check", "checking", "aggregating", "done", "failed", "occupied"}
|
||||
|
||||
var ipResults = []string{"pass", "partial", "fail", "cancelled"}
|
||||
|
||||
func containsStr(list []string, s string) bool {
|
||||
for _, v := range list {
|
||||
if v == s {
|
||||
return true
|
||||
}
|
||||
}
|
||||
return false
|
||||
}
|
||||
|
||||
// ipsFilter is the server-side filter of the /ips list. State is a filter
|
||||
// token: one of "", "queued", "active" (every in-progress state), "done",
|
||||
// "failed", "occupied", or a comma-separated list of raw queue states; the
|
||||
// result filter is separate (Result) and may be combined with it.
|
||||
type ipsFilter struct {
|
||||
Q string
|
||||
State string
|
||||
Result string
|
||||
}
|
||||
|
||||
// parseIPsFilter reads q/state/result from the request (query or form body).
|
||||
// The state <select> posts a single `state` value, so "result:pass" selects a
|
||||
// result filter; anything unrecognised is dropped rather than forwarded.
|
||||
func parseIPsFilter(r *http.Request) ipsFilter {
|
||||
f := ipsFilter{
|
||||
Q: strings.TrimSpace(r.FormValue("q")),
|
||||
State: strings.TrimSpace(r.FormValue("state")),
|
||||
Result: strings.TrimSpace(r.FormValue("result")),
|
||||
}
|
||||
if res, ok := strings.CutPrefix(f.State, "result:"); ok {
|
||||
f.State, f.Result = "", res
|
||||
}
|
||||
if !containsStr(ipResults, f.Result) {
|
||||
f.Result = ""
|
||||
}
|
||||
switch f.State {
|
||||
case "", "active":
|
||||
default:
|
||||
for _, st := range strings.Split(f.State, ",") {
|
||||
if !containsStr(ipStates, st) {
|
||||
f.State = ""
|
||||
break
|
||||
}
|
||||
}
|
||||
}
|
||||
return f
|
||||
}
|
||||
|
||||
// States expands the state token into the list sent to control-api.
|
||||
func (f ipsFilter) States() []string {
|
||||
switch f.State {
|
||||
case "":
|
||||
return nil
|
||||
case "active":
|
||||
return activeStates
|
||||
default:
|
||||
return strings.Split(f.State, ",")
|
||||
}
|
||||
}
|
||||
|
||||
// Active reports whether any filter is applied.
|
||||
func (f ipsFilter) Active() bool { return f.Q != "" || f.State != "" || f.Result != "" }
|
||||
|
||||
// Token is the value of the state <select> option matching the filter.
|
||||
func (f ipsFilter) Token() string {
|
||||
if f.State == "" && f.Result != "" {
|
||||
return "result:" + f.Result
|
||||
}
|
||||
return f.State
|
||||
}
|
||||
|
||||
func (f ipsFilter) values() url.Values {
|
||||
v := url.Values{}
|
||||
if f.Q != "" {
|
||||
v.Set("q", f.Q)
|
||||
}
|
||||
if f.State != "" {
|
||||
v.Set("state", f.State)
|
||||
}
|
||||
if f.Result != "" {
|
||||
v.Set("result", f.Result)
|
||||
}
|
||||
return v
|
||||
}
|
||||
|
||||
type ipsPageData struct {
|
||||
PageData
|
||||
Items []ipQueueItem
|
||||
FIPSettleSeconds int
|
||||
Filter ipsFilter
|
||||
Page, PerPage int
|
||||
// Total is the number of rows matching the filter; QueueTotal is the whole
|
||||
// queue (what "Очистить всё" would remove).
|
||||
Total, QueueTotal int
|
||||
Pager pagerData
|
||||
// SelfURL is this page's own URL (filter + page), re-requested to reload
|
||||
// the table when a scan finishes.
|
||||
SelfURL string
|
||||
PerPageOptions []int
|
||||
Scan scanProgressData
|
||||
}
|
||||
|
||||
type ipDetailData struct {
|
||||
@@ -16,13 +124,67 @@ type ipDetailData struct {
|
||||
Detail ipDetailResponse
|
||||
}
|
||||
|
||||
func (s *Server) handleIPsPage(w http.ResponseWriter, r *http.Request) {
|
||||
items, err := s.CA.ListIPs(r.Context())
|
||||
settings, settingsErr := s.CA.GetOrchestratorSettings(r.Context())
|
||||
// loadIPsData fetches the page of the queue selected by the request's
|
||||
// page/per_page/q/state/result params. A page that became empty (e.g. after
|
||||
// deleting its last rows) is clamped to the last non-empty page.
|
||||
func (s *Server) loadIPsData(r *http.Request) (ipsPageData, error) {
|
||||
ctx := r.Context()
|
||||
f := parseIPsFilter(r)
|
||||
perPage := parsePerPage(r.FormValue("per_page"))
|
||||
page := parsePage(r.FormValue("page"))
|
||||
q := ipsQuery{States: f.States(), Q: f.Q, Result: f.Result, Order: "sequence", Limit: perPage, Offset: (page - 1) * perPage}
|
||||
|
||||
res, err := s.CA.ListIPsPage(ctx, q)
|
||||
if err == nil {
|
||||
if clamped := clampPage(page, res.Total, perPage); clamped != page {
|
||||
page = clamped
|
||||
q.Offset = (page - 1) * perPage
|
||||
res, err = s.CA.ListIPsPage(ctx, q)
|
||||
}
|
||||
}
|
||||
data := ipsPageData{
|
||||
Items: res.Items,
|
||||
Filter: f,
|
||||
Page: page,
|
||||
PerPage: perPage,
|
||||
Total: res.Total,
|
||||
QueueTotal: res.Total,
|
||||
PerPageOptions: perPageOptions,
|
||||
}
|
||||
data.Pager = newPager("/ips", "ips-table-wrap", f.values(), page, perPage, res.Total)
|
||||
data.SelfURL = pageURL("/ips", f.values(), page, perPage)
|
||||
|
||||
if settings, settingsErr := s.CA.GetOrchestratorSettings(ctx); settingsErr != nil {
|
||||
if err == nil {
|
||||
err = settingsErr
|
||||
}
|
||||
data := ipsPageData{Items: items, FIPSettleSeconds: settings.FIPSettleSeconds}
|
||||
} else {
|
||||
data.FIPSettleSeconds = settings.FIPSettleSeconds
|
||||
}
|
||||
if err == nil && f.Active() {
|
||||
// "Очистить всё" ignores the filter: show the real queue size in its
|
||||
// confirmation. Non-fatal — the label just falls back to the filtered total.
|
||||
if st, stErr := s.CA.Status(ctx); stErr == nil {
|
||||
data.QueueTotal = st.TotalIPs
|
||||
}
|
||||
}
|
||||
return data, err
|
||||
}
|
||||
|
||||
func (s *Server) handleIPsPage(w http.ResponseWriter, r *http.Request) {
|
||||
data, err := s.loadIPsData(r)
|
||||
// A filter/pager request from htmx swaps only #ips-table-wrap (hx-select),
|
||||
// so there is no need to re-render the whole page (and re-query the scan
|
||||
// status). A history-restore fetch needs the full page.
|
||||
if r.Header.Get("HX-Request") == "true" && r.Header.Get("HX-History-Restore-Request") != "true" {
|
||||
s.renderFragment(w, "ips_table_wrap", data, err)
|
||||
return
|
||||
}
|
||||
if st, scanErr := s.CA.ScanStatus(r.Context()); scanErr != nil {
|
||||
s.Log.Warn("ips: scan status unavailable", "err", scanErr)
|
||||
} else {
|
||||
data.Scan = newScanProgress(st)
|
||||
}
|
||||
data.ActiveNav = "ips"
|
||||
data.Banner = bannerFor(err)
|
||||
s.renderPage(w, r, "ips_page", data)
|
||||
@@ -37,20 +199,18 @@ func (s *Server) handleIPDetail(w http.ResponseWriter, r *http.Request) {
|
||||
s.renderPage(w, r, "ip_detail_page", data)
|
||||
}
|
||||
|
||||
// renderIPsTable re-fetches the current queue and renders the ips_table
|
||||
// fragment, tagging actionErr (if any) on the shared error banner. Called
|
||||
// after every mutating /ips/* request so the table always reflects true
|
||||
// current state regardless of whether the mutation itself succeeded.
|
||||
// renderIPsTable re-fetches the current page of the queue (same page/filter as
|
||||
// the request, carried in hidden #ips-form inputs) and renders the
|
||||
// ips_table_wrap fragment, tagging actionErr (if any) on the shared error
|
||||
// banner. Called after every mutating /ips/* request so the table always
|
||||
// reflects true current state regardless of whether the mutation itself
|
||||
// succeeded.
|
||||
func (s *Server) renderIPsTable(w http.ResponseWriter, r *http.Request, actionErr error) {
|
||||
items, listErr := s.CA.ListIPs(r.Context())
|
||||
data, err := s.loadIPsData(r)
|
||||
if actionErr == nil {
|
||||
actionErr = listErr
|
||||
actionErr = err
|
||||
}
|
||||
settings, settingsErr := s.CA.GetOrchestratorSettings(r.Context())
|
||||
if actionErr == nil {
|
||||
actionErr = settingsErr
|
||||
}
|
||||
s.renderFragment(w, "ips_table", ipsPageData{Items: items, FIPSettleSeconds: settings.FIPSettleSeconds}, actionErr)
|
||||
s.renderFragment(w, "ips_table_wrap", data, actionErr)
|
||||
}
|
||||
|
||||
func (s *Server) handleIPsSubmit(w http.ResponseWriter, r *http.Request) {
|
||||
@@ -63,7 +223,7 @@ func (s *Server) handleIPsSubmit(w http.ResponseWriter, r *http.Request) {
|
||||
s.renderIPsTable(w, r, &apiErr{Status: http.StatusBadRequest, Message: "укажите хотя бы один адрес"})
|
||||
return
|
||||
}
|
||||
_, err := s.CA.SubmitIPs(r.Context(), addresses)
|
||||
err := s.submitChunked(r.Context(), addresses)
|
||||
s.renderIPsTable(w, r, err)
|
||||
}
|
||||
|
||||
@@ -85,31 +245,76 @@ func (s *Server) handleIPDelete(w http.ResponseWriter, r *http.Request) {
|
||||
s.renderIPsTable(w, r, err)
|
||||
}
|
||||
|
||||
func (s *Server) handleIPsDeleteSelected(w http.ResponseWriter, r *http.Request) {
|
||||
// submitChunked feeds addresses to SubmitIPs in chunks of bulkChunk so one
|
||||
// huge list never becomes one huge request/transaction.
|
||||
func (s *Server) submitChunked(ctx context.Context, addresses []string) error {
|
||||
for _, part := range chunk(addresses, bulkChunk) {
|
||||
if _, err := s.CA.SubmitIPs(ctx, part); err != nil {
|
||||
return err
|
||||
}
|
||||
}
|
||||
return nil
|
||||
}
|
||||
|
||||
// resolveFilterAddresses lists every address matching filter by paging
|
||||
// ListIPsPage (≤ maxPageLimit rows per call), for "select all N by filter".
|
||||
func (s *Server) resolveFilterAddresses(ctx context.Context, f ipsFilter) ([]string, error) {
|
||||
var out []string
|
||||
for offset := 0; ; {
|
||||
page, err := s.CA.ListIPsPage(ctx, ipsQuery{States: f.States(), Q: f.Q, Result: f.Result, Order: "sequence", Limit: maxPageLimit, Offset: offset})
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
for _, it := range page.Items {
|
||||
out = append(out, it.IPAddress)
|
||||
}
|
||||
offset += len(page.Items)
|
||||
if len(page.Items) == 0 || offset >= page.Total {
|
||||
return out, nil
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// bulkAddresses returns the addresses a bulk delete/recheck acts on: the
|
||||
// checked rows of the current page, or — with scope=all — everything that
|
||||
// matches the current filter, resolved server-side.
|
||||
func (s *Server) bulkAddresses(r *http.Request) ([]string, error) {
|
||||
if err := r.ParseForm(); err != nil {
|
||||
s.renderIPsTable(w, r, fmt.Errorf("invalid form: %w", err))
|
||||
return
|
||||
return nil, fmt.Errorf("invalid form: %w", err)
|
||||
}
|
||||
var addresses []string
|
||||
if r.FormValue("scope") == "all" {
|
||||
var err error
|
||||
addresses, err = s.resolveFilterAddresses(r.Context(), parseIPsFilter(r))
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
} else {
|
||||
addresses = r.Form["addresses"]
|
||||
}
|
||||
addresses := r.Form["addresses"]
|
||||
if len(addresses) == 0 {
|
||||
s.renderIPsTable(w, r, &apiErr{Status: http.StatusBadRequest, Message: "ничего не выбрано"})
|
||||
return
|
||||
return nil, &apiErr{Status: http.StatusBadRequest, Message: "ничего не выбрано"}
|
||||
}
|
||||
return addresses, nil
|
||||
}
|
||||
|
||||
func (s *Server) handleIPsDeleteSelected(w http.ResponseWriter, r *http.Request) {
|
||||
addresses, err := s.bulkAddresses(r)
|
||||
if err == nil {
|
||||
for _, part := range chunk(addresses, bulkChunk) {
|
||||
if _, err = s.CA.DeleteIPs(r.Context(), part); err != nil {
|
||||
break
|
||||
}
|
||||
}
|
||||
}
|
||||
_, err := s.CA.DeleteIPs(r.Context(), addresses)
|
||||
s.renderIPsTable(w, r, err)
|
||||
}
|
||||
|
||||
func (s *Server) handleIPsRecheckSelected(w http.ResponseWriter, r *http.Request) {
|
||||
if err := r.ParseForm(); err != nil {
|
||||
s.renderIPsTable(w, r, fmt.Errorf("invalid form: %w", err))
|
||||
return
|
||||
addresses, err := s.bulkAddresses(r)
|
||||
if err == nil {
|
||||
err = s.submitChunked(r.Context(), addresses)
|
||||
}
|
||||
addresses := r.Form["addresses"]
|
||||
if len(addresses) == 0 {
|
||||
s.renderIPsTable(w, r, &apiErr{Status: http.StatusBadRequest, Message: "ничего не выбрано"})
|
||||
return
|
||||
}
|
||||
_, err := s.CA.SubmitIPs(r.Context(), addresses)
|
||||
s.renderIPsTable(w, r, err)
|
||||
}
|
||||
|
||||
@@ -118,10 +323,46 @@ func (s *Server) handleIPsClear(w http.ResponseWriter, r *http.Request) {
|
||||
s.renderIPsTable(w, r, err)
|
||||
}
|
||||
|
||||
// handleIPsScan lists the OpenStack project's free (unassociated) floating
|
||||
// IPs and submits them to the check queue in one step — see
|
||||
// client.ScanFloatingIPs.
|
||||
// scanProgressData drives the scan_progress partial. Poll keeps the partial's
|
||||
// own hx-trigger="every 2s" alive: while the job runs, or while control-api is
|
||||
// transiently unreachable (so one failed poll doesn't freeze the panel).
|
||||
type scanProgressData struct {
|
||||
Status scanStatusDTO
|
||||
Poll bool
|
||||
}
|
||||
|
||||
func newScanProgress(st scanStatusDTO) scanProgressData {
|
||||
return scanProgressData{Status: st, Poll: st.Running}
|
||||
}
|
||||
|
||||
// renderScanProgress renders the scan panel fragment. A finished job also
|
||||
// sends `HX-Trigger: scan-finished`, which makes #ips-table-wrap reload.
|
||||
// Errors go to the shared banner; only transient ones (transport/5xx) keep
|
||||
// polling when keepPolling is set.
|
||||
func (s *Server) renderScanProgress(w http.ResponseWriter, st scanStatusDTO, err error, keepPolling bool) {
|
||||
data := newScanProgress(st)
|
||||
if err != nil {
|
||||
data = scanProgressData{}
|
||||
var ae *apiErr
|
||||
if keepPolling && errors.As(err, &ae) && (ae.Status == 0 || ae.Status >= 500) {
|
||||
data.Poll = true
|
||||
}
|
||||
} else if st.Finished() {
|
||||
w.Header().Set("HX-Trigger", "scan-finished")
|
||||
}
|
||||
s.renderFragment(w, "scan_progress", data, err)
|
||||
}
|
||||
|
||||
// handleIPsScan starts the background floating-IP scan (or joins the running
|
||||
// one) and returns the progress panel at once — the job itself runs in
|
||||
// control-api, so this never waits for OpenStack. ?dry_run=true only counts.
|
||||
func (s *Server) handleIPsScan(w http.ResponseWriter, r *http.Request) {
|
||||
_, err := s.CA.ScanFloatingIPs(r.Context())
|
||||
s.renderIPsTable(w, r, err)
|
||||
st, err := s.CA.StartScan(r.Context(), r.URL.Query().Get("dry_run") == "true")
|
||||
s.renderScanProgress(w, st, err, false)
|
||||
}
|
||||
|
||||
// handleIPsScanStatus is the progress panel's poll target.
|
||||
func (s *Server) handleIPsScanStatus(w http.ResponseWriter, r *http.Request) {
|
||||
st, err := s.CA.ScanStatus(r.Context())
|
||||
s.renderScanProgress(w, st, err, true)
|
||||
}
|
||||
@@ -2,23 +2,51 @@ package dashboard
|
||||
|
||||
import (
|
||||
"net/http"
|
||||
"sort"
|
||||
"strings"
|
||||
"time"
|
||||
)
|
||||
|
||||
const (
|
||||
// overviewActiveLimit caps the "В работе" table; overviewQueueLimit caps
|
||||
// the compact list of the next queued addresses. The queue itself can hold
|
||||
// thousands of rows, so the overview never lists more than these.
|
||||
overviewActiveLimit = 100
|
||||
overviewQueueLimit = 10
|
||||
// etaMinSamples is how many completed rows (with AggregatedAt) the ETA
|
||||
// needs to derive a rate from.
|
||||
etaMinSamples = 5
|
||||
)
|
||||
|
||||
// overviewProgress is the "Готово D из T" indicator of the stats block.
|
||||
type overviewProgress struct {
|
||||
Total, Done, Active, Queued, Percent int
|
||||
// ETA is a rough time-to-finish estimate ("" when not enough data).
|
||||
ETA string
|
||||
}
|
||||
|
||||
type overviewData struct {
|
||||
PageData
|
||||
Status statusResponse
|
||||
CurrentItems []ipQueueItem
|
||||
// ActiveItems are the addresses being checked right now (≤ overviewActiveLimit
|
||||
// of ActiveTotal); QueuedItems are the next few waiting ones (QueuedTotal
|
||||
// in total). Both stay empty under a result filter: an address without a
|
||||
// verdict can't match one.
|
||||
ActiveItems []ipQueueItem
|
||||
ActiveTotal int
|
||||
QueuedItems []ipQueueItem
|
||||
QueuedTotal int
|
||||
LastCompleted []ipQueueItem
|
||||
Breakdown map[string]int
|
||||
LastN int
|
||||
PollSeconds int
|
||||
Query string
|
||||
StatusFilter string
|
||||
Progress overviewProgress
|
||||
// AutoCycle is nil when the auto-cycle status could not be fetched; the
|
||||
// indicator is then simply hidden (a non-fatal failure).
|
||||
// indicator is then simply hidden (a non-fatal failure). Scan is likewise
|
||||
// nil when the scan status is unavailable.
|
||||
AutoCycle *autoCycleDTO
|
||||
Scan *scanStatusDTO
|
||||
}
|
||||
|
||||
func (s *Server) loadOverview(r *http.Request) (overviewData, error) {
|
||||
@@ -27,30 +55,58 @@ func (s *Server) loadOverview(r *http.Request) (overviewData, error) {
|
||||
if err != nil {
|
||||
return overviewData{}, err
|
||||
}
|
||||
ips, err := s.CA.ListIPs(ctx)
|
||||
if err != nil {
|
||||
return overviewData{}, err
|
||||
}
|
||||
var autoCycle *autoCycleDTO
|
||||
if ac, acErr := s.CA.GetAutoCycle(ctx); acErr != nil {
|
||||
s.Log.Warn("overview: auto-cycle status unavailable", "err", acErr)
|
||||
} else {
|
||||
autoCycle = &ac
|
||||
}
|
||||
q := strings.TrimSpace(r.URL.Query().Get("q"))
|
||||
resultFilter := r.URL.Query().Get("status")
|
||||
last := lastCompleted(ips, s.Cfg.LastCompletedCount)
|
||||
return overviewData{
|
||||
if !containsStr(ipResults, resultFilter) {
|
||||
resultFilter = ""
|
||||
}
|
||||
lastN := s.Cfg.LastCompletedCount
|
||||
if lastN < 1 {
|
||||
lastN = 20
|
||||
}
|
||||
|
||||
data := overviewData{
|
||||
Status: status,
|
||||
CurrentItems: filterQueueItems(currentlyChecking(ips), q, resultFilter),
|
||||
LastCompleted: filterQueueItems(last, q, resultFilter),
|
||||
Breakdown: resultBreakdown(last),
|
||||
LastN: s.Cfg.LastCompletedCount,
|
||||
LastN: lastN,
|
||||
PollSeconds: s.Cfg.OverviewPollIntervalS,
|
||||
Query: q,
|
||||
StatusFilter: resultFilter,
|
||||
AutoCycle: autoCycle,
|
||||
}, nil
|
||||
}
|
||||
|
||||
if resultFilter == "" {
|
||||
active, err := s.CA.ListIPsPage(ctx, ipsQuery{States: activeStates, Q: q, Order: "sequence", Limit: overviewActiveLimit})
|
||||
if err != nil {
|
||||
return overviewData{}, err
|
||||
}
|
||||
queued, err := s.CA.ListIPsPage(ctx, ipsQuery{States: []string{"queued"}, Q: q, Order: "sequence", Limit: overviewQueueLimit})
|
||||
if err != nil {
|
||||
return overviewData{}, err
|
||||
}
|
||||
data.ActiveItems, data.ActiveTotal = active.Items, active.Total
|
||||
data.QueuedItems, data.QueuedTotal = queued.Items, queued.Total
|
||||
}
|
||||
completed, err := s.CA.ListIPsPage(ctx, ipsQuery{States: []string{"done", "failed"}, Q: q, Result: resultFilter, Order: "aggregated_at_desc", Limit: lastN})
|
||||
if err != nil {
|
||||
return overviewData{}, err
|
||||
}
|
||||
data.LastCompleted = completed.Items
|
||||
data.Breakdown = resultBreakdown(completed.Items)
|
||||
|
||||
// ETA from the unfiltered window only: a filtered list is not a sample of
|
||||
// the checks' real throughput.
|
||||
data.Progress = computeProgress(status, completed.Items, q == "" && resultFilter == "")
|
||||
|
||||
if ac, acErr := s.CA.GetAutoCycle(ctx); acErr != nil {
|
||||
s.Log.Warn("overview: auto-cycle status unavailable", "err", acErr)
|
||||
} else {
|
||||
data.AutoCycle = &ac
|
||||
}
|
||||
if sc, scErr := s.CA.ScanStatus(ctx); scErr != nil {
|
||||
s.Log.Warn("overview: scan status unavailable", "err", scErr)
|
||||
} else if sc.Running {
|
||||
data.Scan = &sc
|
||||
}
|
||||
return data, nil
|
||||
}
|
||||
|
||||
func (s *Server) handleOverview(w http.ResponseWriter, r *http.Request) {
|
||||
@@ -75,40 +131,45 @@ func (s *Server) handleOverviewFragment(w http.ResponseWriter, r *http.Request)
|
||||
}
|
||||
}
|
||||
|
||||
// currentlyChecking is every IP not yet in a terminal state, ordered by
|
||||
// queue position — the "текущая проверка" live snapshot. No backend
|
||||
// concept of a "run" exists; this is computed fresh on every request.
|
||||
func currentlyChecking(ips []ipQueueItem) []ipQueueItem {
|
||||
var out []ipQueueItem
|
||||
for _, ip := range ips {
|
||||
if ip.State != "done" && ip.State != "failed" {
|
||||
out = append(out, ip)
|
||||
// computeProgress derives the done/active/queued counters from the status
|
||||
// breakdown (done = every terminal state, "occupied" included) and, when
|
||||
// useETA is set and the newest-first list of completed rows has at least
|
||||
// etaMinSamples with AggregatedAt, a rough ETA from their completion rate.
|
||||
func computeProgress(st statusResponse, completed []ipQueueItem, useETA bool) overviewProgress {
|
||||
p := overviewProgress{
|
||||
Total: st.TotalIPs,
|
||||
Done: sumStates(st.IPsByState, terminalStates),
|
||||
Active: sumStates(st.IPsByState, activeStates),
|
||||
Queued: st.IPsByState["queued"],
|
||||
}
|
||||
if p.Total > 0 {
|
||||
p.Percent = p.Done * 100 / p.Total
|
||||
}
|
||||
remaining := p.Active + p.Queued
|
||||
if !useETA || remaining == 0 {
|
||||
return p
|
||||
}
|
||||
var stamps []time.Time
|
||||
for _, ip := range completed {
|
||||
if ip.AggregatedAt != nil {
|
||||
stamps = append(stamps, *ip.AggregatedAt)
|
||||
}
|
||||
}
|
||||
sort.Slice(out, func(i, j int) bool { return out[i].Sequence < out[j].Sequence })
|
||||
return out
|
||||
}
|
||||
|
||||
// lastCompleted returns the n most recently completed (done/failed) IPs by
|
||||
// AggregatedAt descending — the "последняя завершённая проверка" summary
|
||||
// window. This is an operational definition, not a real "batch": resubmit
|
||||
// n if the window size needs tuning (overview.last_completed_count).
|
||||
func lastCompleted(ips []ipQueueItem, n int) []ipQueueItem {
|
||||
var done []ipQueueItem
|
||||
for _, ip := range ips {
|
||||
if (ip.State == "done" || ip.State == "failed") && ip.AggregatedAt != nil {
|
||||
done = append(done, ip)
|
||||
if len(stamps) < etaMinSamples {
|
||||
return p
|
||||
}
|
||||
// completed is newest-first: stamps[0] is the newest, the last the oldest.
|
||||
span := stamps[0].Sub(stamps[len(stamps)-1])
|
||||
if span <= 0 {
|
||||
return p
|
||||
}
|
||||
sort.Slice(done, func(i, j int) bool { return done[i].AggregatedAt.After(*done[j].AggregatedAt) })
|
||||
if len(done) > n {
|
||||
done = done[:n]
|
||||
}
|
||||
return done
|
||||
perItem := span / time.Duration(len(stamps)-1)
|
||||
p.ETA = fmtDuration(perItem * time.Duration(remaining))
|
||||
return p
|
||||
}
|
||||
|
||||
// resultBreakdown counts OverallResult values across exactly the given
|
||||
// items (normally the output of lastCompleted) — pass/partial/fail/cancelled.
|
||||
// items (the "последние N завершённых" window) — pass/partial/fail/cancelled.
|
||||
func resultBreakdown(items []ipQueueItem) map[string]int {
|
||||
out := map[string]int{"pass": 0, "partial": 0, "fail": 0, "cancelled": 0}
|
||||
for _, ip := range items {
|
||||
@@ -116,26 +177,3 @@ func resultBreakdown(items []ipQueueItem) map[string]int {
|
||||
}
|
||||
return out
|
||||
}
|
||||
|
||||
// filterQueueItems narrows items to those whose address contains q
|
||||
// (case-insensitive substring) and, if status is set, whose OverallResult
|
||||
// matches it exactly. A still-in-progress item always has an empty
|
||||
// OverallResult, so picking any specific status hides it — the intended
|
||||
// behavior for "Текущая проверка", which has no verdict yet.
|
||||
func filterQueueItems(items []ipQueueItem, q, status string) []ipQueueItem {
|
||||
if q == "" && status == "" {
|
||||
return items
|
||||
}
|
||||
q = strings.ToLower(q)
|
||||
out := make([]ipQueueItem, 0, len(items))
|
||||
for _, ip := range items {
|
||||
if q != "" && !strings.Contains(strings.ToLower(ip.IPAddress), q) {
|
||||
continue
|
||||
}
|
||||
if status != "" && ip.OverallResult != status {
|
||||
continue
|
||||
}
|
||||
out = append(out, ip)
|
||||
}
|
||||
return out
|
||||
}
|
||||
@@ -2,6 +2,7 @@ package dashboard
|
||||
|
||||
import (
|
||||
"net/http"
|
||||
"net/url"
|
||||
"strings"
|
||||
)
|
||||
|
||||
@@ -10,6 +11,10 @@ type registryPageData struct {
|
||||
Items []registryItem
|
||||
Query string
|
||||
StatusFilter string
|
||||
Page, PerPage int
|
||||
Total int
|
||||
Pager pagerData
|
||||
PerPageOptions []int
|
||||
}
|
||||
|
||||
type registryDetailData struct {
|
||||
@@ -22,43 +27,48 @@ type registryDetailData struct {
|
||||
// record that survives an address being deleted from /ips and later
|
||||
// re-added. See internal/db/migrations/0007_ip_registry.sql. Optional
|
||||
// ?q=&status= query params narrow the list by address substring and by
|
||||
// LastResult — see filterRegistryItems.
|
||||
// LastResult, and ?page=&per_page= select a page — all applied server-side
|
||||
// (control-api's ListRegistryPage), so only the visible rows are transferred.
|
||||
func (s *Server) handleRegistryPage(w http.ResponseWriter, r *http.Request) {
|
||||
items, err := s.CA.ListRegistry(r.Context())
|
||||
q := strings.TrimSpace(r.URL.Query().Get("q"))
|
||||
status := r.URL.Query().Get("status")
|
||||
if !containsStr(ipResults, status) {
|
||||
status = ""
|
||||
}
|
||||
perPage := parsePerPage(r.URL.Query().Get("per_page"))
|
||||
page := parsePage(r.URL.Query().Get("page"))
|
||||
query := registryQuery{Q: q, LastResult: status, Limit: perPage, Offset: (page - 1) * perPage}
|
||||
|
||||
res, err := s.CA.ListRegistryPage(r.Context(), query)
|
||||
if err == nil {
|
||||
if clamped := clampPage(page, res.Total, perPage); clamped != page {
|
||||
page = clamped
|
||||
query.Offset = (page - 1) * perPage
|
||||
res, err = s.CA.ListRegistryPage(r.Context(), query)
|
||||
}
|
||||
}
|
||||
params := url.Values{}
|
||||
if q != "" {
|
||||
params.Set("q", q)
|
||||
}
|
||||
if status != "" {
|
||||
params.Set("status", status)
|
||||
}
|
||||
data := registryPageData{
|
||||
Items: filterRegistryItems(items, q, status),
|
||||
Items: res.Items,
|
||||
Query: q,
|
||||
StatusFilter: status,
|
||||
Page: page,
|
||||
PerPage: perPage,
|
||||
Total: res.Total,
|
||||
Pager: newPager("/registry", "registry-table-wrap", params, page, perPage, res.Total),
|
||||
PerPageOptions: perPageOptions,
|
||||
}
|
||||
data.ActiveNav = "registry"
|
||||
data.Banner = bannerFor(err)
|
||||
s.renderPage(w, r, "registry_page", data)
|
||||
}
|
||||
|
||||
// filterRegistryItems narrows items to those whose address contains q
|
||||
// (case-insensitive substring) and, if status is set, whose LastResult
|
||||
// matches it exactly — the registry list's search-by-IP and
|
||||
// filter-by-status, mirroring filterQueueItems in handlers_overview.go.
|
||||
func filterRegistryItems(items []registryItem, q, status string) []registryItem {
|
||||
if q == "" && status == "" {
|
||||
return items
|
||||
}
|
||||
q = strings.ToLower(q)
|
||||
out := make([]registryItem, 0, len(items))
|
||||
for _, it := range items {
|
||||
if q != "" && !strings.Contains(strings.ToLower(it.IPAddress), q) {
|
||||
continue
|
||||
}
|
||||
if status != "" && it.LastResult != status {
|
||||
continue
|
||||
}
|
||||
out = append(out, it)
|
||||
}
|
||||
return out
|
||||
}
|
||||
|
||||
// handleRegistryDetail shows one address's full retained check history
|
||||
// across every cycle it has ever run, not just the current attempt — see
|
||||
// ip_detail_content in ip_detail.html for the attempt-scoped equivalent.
|
||||
|
||||
@@ -558,20 +558,6 @@ func TestTargetsAndCheckTypesRoundTrip(t *testing.T) {
|
||||
}
|
||||
}
|
||||
|
||||
func TestIPsScan(t *testing.T) {
|
||||
fake, caURL := newFakeControlAPI(t)
|
||||
fake.scanFreeAddresses = []string{"5.5.5.5"}
|
||||
ts := newTestServer(t, caURL)
|
||||
|
||||
body := postForm(t, ts, "POST", "/ips/scan", nil)
|
||||
if !strings.Contains(body, "5.5.5.5") {
|
||||
t.Fatalf("expected scanned address in re-rendered table, got:\n%s", body)
|
||||
}
|
||||
if len(fake.ips) != 1 || fake.ips[0].IPAddress != "5.5.5.5" {
|
||||
t.Fatalf("expected fake control-api queue to contain the scanned address, got %+v", fake.ips)
|
||||
}
|
||||
}
|
||||
|
||||
func TestRegistryPageAndDetail(t *testing.T) {
|
||||
fake, caURL := newFakeControlAPI(t)
|
||||
now := time.Now()
|
||||
@@ -667,64 +653,6 @@ func TestControlAPIUnreachable(t *testing.T) {
|
||||
}
|
||||
}
|
||||
|
||||
func TestFilterQueueItems(t *testing.T) {
|
||||
items := []ipQueueItem{
|
||||
{IPAddress: "1.1.1.1", OverallResult: "pass"},
|
||||
{IPAddress: "1.1.1.2", OverallResult: "fail"},
|
||||
{IPAddress: "2.2.2.2", OverallResult: "pass"},
|
||||
}
|
||||
|
||||
if got := filterQueueItems(items, "", ""); len(got) != 3 {
|
||||
t.Fatalf("expected no-op with empty q/status, got %+v", got)
|
||||
}
|
||||
if got := filterQueueItems(items, "1.1.1", ""); len(got) != 2 {
|
||||
t.Fatalf("expected 2 matches for q=1.1.1, got %+v", got)
|
||||
}
|
||||
if got := filterQueueItems(items, "1.1.1.1", ""); len(got) != 1 || got[0].IPAddress != "1.1.1.1" {
|
||||
t.Fatalf("expected exact-substring match, got %+v", got)
|
||||
}
|
||||
if got := filterQueueItems(items, "1.1.1.1", ""); len(got) != 1 {
|
||||
t.Fatalf("expected search to be case/substring based, got %+v", got)
|
||||
}
|
||||
if got := filterQueueItems(items, "", "pass"); len(got) != 2 {
|
||||
t.Fatalf("expected 2 matches for status=pass, got %+v", got)
|
||||
}
|
||||
if got := filterQueueItems(items, "1.1.1", "pass"); len(got) != 1 || got[0].IPAddress != "1.1.1.1" {
|
||||
t.Fatalf("expected q+status combined with AND, got %+v", got)
|
||||
}
|
||||
if got := filterQueueItems(items, "9.9.9.9", ""); len(got) != 0 {
|
||||
t.Fatalf("expected no matches, got %+v", got)
|
||||
}
|
||||
|
||||
// Case-insensitivity, via a query with mixed-case letters (IP octets
|
||||
// are numeric, so exercise it through IPv6-shaped input instead).
|
||||
mixed := []ipQueueItem{{IPAddress: "fe80::AbCd"}}
|
||||
if got := filterQueueItems(mixed, "abcd", ""); len(got) != 1 {
|
||||
t.Fatalf("expected case-insensitive search to match, got %+v", got)
|
||||
}
|
||||
}
|
||||
|
||||
func TestFilterRegistryItems(t *testing.T) {
|
||||
items := []registryItem{
|
||||
{IPAddress: "1.1.1.1", LastResult: "pass"},
|
||||
{IPAddress: "1.1.1.2", LastResult: "partial"},
|
||||
{IPAddress: "2.2.2.2", LastResult: ""},
|
||||
}
|
||||
|
||||
if got := filterRegistryItems(items, "", ""); len(got) != 3 {
|
||||
t.Fatalf("expected no-op with empty q/status, got %+v", got)
|
||||
}
|
||||
if got := filterRegistryItems(items, "1.1.1", ""); len(got) != 2 {
|
||||
t.Fatalf("expected 2 matches for q=1.1.1, got %+v", got)
|
||||
}
|
||||
if got := filterRegistryItems(items, "", "partial"); len(got) != 1 || got[0].IPAddress != "1.1.1.2" {
|
||||
t.Fatalf("expected exactly the partial-result address, got %+v", got)
|
||||
}
|
||||
if got := filterRegistryItems(items, "2.2.2", "partial"); len(got) != 0 {
|
||||
t.Fatalf("expected q+status combined with AND to exclude non-matching, got %+v", got)
|
||||
}
|
||||
}
|
||||
|
||||
func TestAutoCyclePanelRenders(t *testing.T) {
|
||||
fake, caURL := newFakeControlAPI(t)
|
||||
finished := time.Now().Add(-time.Hour)
|
||||
|
||||
@@ -0,0 +1,118 @@
|
||||
package dashboard
|
||||
|
||||
import (
|
||||
"net/url"
|
||||
"strconv"
|
||||
"strings"
|
||||
)
|
||||
|
||||
// Server-side pagination shared by /ips and /registry: `page` (1-based,
|
||||
// clamped to the last page) and `per_page` (one of perPageOptions, else
|
||||
// defaultPerPage).
|
||||
|
||||
const defaultPerPage = 50
|
||||
|
||||
var perPageOptions = []int{25, 50, 100, 200}
|
||||
|
||||
// parsePerPage accepts only the whitelisted page sizes.
|
||||
func parsePerPage(raw string) int {
|
||||
n, err := strconv.Atoi(strings.TrimSpace(raw))
|
||||
if err != nil {
|
||||
return defaultPerPage
|
||||
}
|
||||
for _, o := range perPageOptions {
|
||||
if n == o {
|
||||
return n
|
||||
}
|
||||
}
|
||||
return defaultPerPage
|
||||
}
|
||||
|
||||
// parsePage returns the 1-based page number; anything invalid is page 1.
|
||||
func parsePage(raw string) int {
|
||||
n, err := strconv.Atoi(strings.TrimSpace(raw))
|
||||
if err != nil || n < 1 {
|
||||
return 1
|
||||
}
|
||||
return n
|
||||
}
|
||||
|
||||
func pageCount(total, perPage int) int {
|
||||
if total <= 0 || perPage <= 0 {
|
||||
return 1
|
||||
}
|
||||
return (total + perPage - 1) / perPage
|
||||
}
|
||||
|
||||
// clampPage limits page to the last non-empty page for the given total.
|
||||
func clampPage(page, total, perPage int) int {
|
||||
if last := pageCount(total, perPage); page > last {
|
||||
page = last
|
||||
}
|
||||
if page < 1 {
|
||||
page = 1
|
||||
}
|
||||
return page
|
||||
}
|
||||
|
||||
// pagerData drives the shared "pager" template partial.
|
||||
type pagerData struct {
|
||||
// Wrap is the id of the swap-target element holding the table (the
|
||||
// ‹ › links replace it via hx-select + hx-target).
|
||||
Wrap string
|
||||
Page, PerPage, Total, Pages, From, To int
|
||||
PrevURL, NextURL string
|
||||
}
|
||||
|
||||
// newPager builds the pager for page (already clamped) of total rows. base is
|
||||
// the page path; params are the filters to preserve in the links (without
|
||||
// page/per_page, which the pager sets itself).
|
||||
func newPager(base, wrap string, params url.Values, page, perPage, total int) pagerData {
|
||||
p := pagerData{Wrap: wrap, Page: page, PerPage: perPage, Total: total, Pages: pageCount(total, perPage)}
|
||||
if total > 0 {
|
||||
p.From = (page-1)*perPage + 1
|
||||
p.To = page * perPage
|
||||
if p.To > total {
|
||||
p.To = total
|
||||
}
|
||||
}
|
||||
if page > 1 {
|
||||
p.PrevURL = pageURL(base, params, page-1, perPage)
|
||||
}
|
||||
if page < p.Pages {
|
||||
p.NextURL = pageURL(base, params, page+1, perPage)
|
||||
}
|
||||
return p
|
||||
}
|
||||
|
||||
// pageURL builds base?...&page=N&per_page=M with every value escaped by
|
||||
// url.Values (so a search string like "a&b" can't smuggle in a parameter).
|
||||
// Page 1 omits `page`.
|
||||
func pageURL(base string, params url.Values, page, perPage int) string {
|
||||
v := url.Values{}
|
||||
for k, vals := range params {
|
||||
for _, val := range vals {
|
||||
if val != "" {
|
||||
v.Add(k, val)
|
||||
}
|
||||
}
|
||||
}
|
||||
v.Set("per_page", strconv.Itoa(perPage))
|
||||
if page > 1 {
|
||||
v.Set("page", strconv.Itoa(page))
|
||||
}
|
||||
return base + "?" + v.Encode()
|
||||
}
|
||||
|
||||
// chunk splits list into slices of at most size elements.
|
||||
func chunk(list []string, size int) [][]string {
|
||||
var out [][]string
|
||||
for len(list) > size {
|
||||
out = append(out, list[:size])
|
||||
list = list[size:]
|
||||
}
|
||||
if len(list) > 0 {
|
||||
out = append(out, list)
|
||||
}
|
||||
return out
|
||||
}
|
||||
@@ -18,6 +18,7 @@ func (s *Server) routes(mux *http.ServeMux) {
|
||||
mux.HandleFunc("GET /ips/{ip}", s.handleIPDetail)
|
||||
mux.HandleFunc("POST /ips", s.handleIPsSubmit)
|
||||
mux.HandleFunc("POST /ips/scan", s.handleIPsScan)
|
||||
mux.HandleFunc("GET /ips/scan/status", s.handleIPsScanStatus)
|
||||
mux.HandleFunc("POST /ips/{ip}/recheck", s.handleIPRecheck)
|
||||
mux.HandleFunc("POST /ips/{ip}/cancel", s.handleIPCancel)
|
||||
mux.HandleFunc("DELETE /ips/{ip}", s.handleIPDelete)
|
||||
|
||||
@@ -0,0 +1,685 @@
|
||||
package dashboard
|
||||
|
||||
import (
|
||||
"fmt"
|
||||
"html"
|
||||
"net/http"
|
||||
"net/url"
|
||||
"regexp"
|
||||
"strconv"
|
||||
"strings"
|
||||
"testing"
|
||||
"time"
|
||||
)
|
||||
|
||||
// Tests for the scale features: background scan panel, server-side
|
||||
// pagination/filters on /ips and /registry, bulk selection by filter, and the
|
||||
// bounded overview.
|
||||
|
||||
func ipRowLink(ip string) string { return `href="/ips/` + ip + `"` }
|
||||
|
||||
func countRows(body string) int { return strings.Count(body, `name="addresses" value="`) }
|
||||
|
||||
func hasHXTriggerEvery(body string) bool { return strings.Contains(body, `hx-trigger="every 2s"`) }
|
||||
|
||||
var hrefRe = regexp.MustCompile(`href="([^"]*)"[^>]*rel="(prev|next)"`)
|
||||
|
||||
// pagerLink extracts the (unescaped, parsed) prev/next link of the pager.
|
||||
func pagerLink(t *testing.T, body, rel string) *url.URL {
|
||||
t.Helper()
|
||||
for _, m := range hrefRe.FindAllStringSubmatch(body, -1) {
|
||||
if m[2] == rel {
|
||||
u, err := url.Parse(html.UnescapeString(m[1]))
|
||||
if err != nil {
|
||||
t.Fatalf("parse pager link %q: %v", m[1], err)
|
||||
}
|
||||
return u
|
||||
}
|
||||
}
|
||||
return nil
|
||||
}
|
||||
|
||||
func TestIPsScan(t *testing.T) {
|
||||
fake, caURL := newFakeControlAPI(t)
|
||||
fake.scanFreeAddresses = []string{"5.5.5.5", "5.5.5.6"}
|
||||
fake.scanRunPolls = 2
|
||||
ts := newTestServer(t, caURL)
|
||||
|
||||
// Start: answers at once with the panel; it polls itself, the scan
|
||||
// buttons are disabled and nothing is queued yet.
|
||||
resp, body := doReq(t, ts, reqOpts{method: http.MethodPost, path: "/ips/scan"})
|
||||
if resp.StatusCode != http.StatusOK {
|
||||
t.Fatalf("status = %d, want 200", resp.StatusCode)
|
||||
}
|
||||
for _, want := range []string{
|
||||
`id="scan-progress"`, `hx-get="/ips/scan/status"`, `hx-trigger="every 2s"`, "читаются страницы",
|
||||
`<progress aria-label="Сканирование"></progress>`, ` disabled`,
|
||||
} {
|
||||
if !strings.Contains(body, want) {
|
||||
t.Fatalf("expected %q in the running panel, got:\n%s", want, body)
|
||||
}
|
||||
}
|
||||
if resp.Header.Get("HX-Trigger") != "" {
|
||||
t.Fatalf("a running scan must not fire scan-finished, got %q", resp.Header.Get("HX-Trigger"))
|
||||
}
|
||||
if len(fake.ips) != 0 {
|
||||
t.Fatalf("nothing should be queued before the job finishes, got %+v", fake.ips)
|
||||
}
|
||||
|
||||
// First poll: still running, keeps polling.
|
||||
resp, body = doReq(t, ts, reqOpts{path: "/ips/scan/status"})
|
||||
if !hasHXTriggerEvery(body) || resp.Header.Get("HX-Trigger") != "" {
|
||||
t.Fatalf("expected a still-polling panel without HX-Trigger, got %q:\n%s", resp.Header.Get("HX-Trigger"), body)
|
||||
}
|
||||
|
||||
// Second poll: finished — no polling trigger, HX-Trigger tells the table to reload.
|
||||
resp, body = doReq(t, ts, reqOpts{path: "/ips/scan/status"})
|
||||
if hasHXTriggerEvery(body) || strings.Contains(body, `hx-get="/ips/scan/status"`) {
|
||||
t.Fatalf("a finished panel must stop polling, got:\n%s", body)
|
||||
}
|
||||
if got := resp.Header.Get("HX-Trigger"); got != "scan-finished" {
|
||||
t.Fatalf("HX-Trigger = %q, want scan-finished", got)
|
||||
}
|
||||
for _, want := range []string{"готово", "<dt>добавлено</dt><dd>2</dd>", `<progress max="1" value="1"`} {
|
||||
if !strings.Contains(body, want) {
|
||||
t.Fatalf("expected %q in the finished panel, got:\n%s", want, body)
|
||||
}
|
||||
}
|
||||
if strings.Contains(body, " disabled") {
|
||||
t.Fatalf("buttons must be enabled again once finished, got:\n%s", body)
|
||||
}
|
||||
if len(fake.ips) != 2 {
|
||||
t.Fatalf("expected both scanned addresses queued, got %+v", fake.ips)
|
||||
}
|
||||
|
||||
// The table reload target: #ips-table-wrap listens for scan-finished and
|
||||
// re-fetches the current /ips URL.
|
||||
page := get(t, ts, "/ips")
|
||||
if !strings.Contains(page, `hx-trigger="scan-finished from:body"`) || !strings.Contains(page, `hx-select="#ips-table-wrap"`) {
|
||||
t.Fatalf("expected #ips-table-wrap to reload on scan-finished, got:\n%s", page)
|
||||
}
|
||||
if !strings.Contains(page, ipRowLink("5.5.5.5")) {
|
||||
t.Fatalf("expected the scanned address in the table, got:\n%s", page)
|
||||
}
|
||||
}
|
||||
|
||||
func TestIPsScanErrorAndDryRunAndRefusal(t *testing.T) {
|
||||
fake, caURL := newFakeControlAPI(t)
|
||||
fake.scanFreeAddresses = []string{"5.5.5.5"}
|
||||
fake.scanFinalError = "openstack: list floating ips: boom"
|
||||
ts := newTestServer(t, caURL)
|
||||
|
||||
// A job that fails: the panel shows the error and stops polling.
|
||||
resp, body := doReq(t, ts, reqOpts{method: http.MethodPost, path: "/ips/scan"})
|
||||
for _, want := range []string{"ошибка", "openstack: list floating ips: boom", `class="scan-error"`} {
|
||||
if !strings.Contains(body, want) {
|
||||
t.Fatalf("expected %q in the error panel, got:\n%s", want, body)
|
||||
}
|
||||
}
|
||||
if hasHXTriggerEvery(body) || resp.Header.Get("HX-Trigger") != "scan-finished" {
|
||||
t.Fatalf("an errored job must stop polling and signal scan-finished, got %q:\n%s", resp.Header.Get("HX-Trigger"), body)
|
||||
}
|
||||
if len(fake.ips) != 0 {
|
||||
t.Fatalf("a failed scan must not queue anything, got %+v", fake.ips)
|
||||
}
|
||||
|
||||
// Dry run: the flag reaches control-api, the queue stays untouched.
|
||||
fake.mu.Lock()
|
||||
fake.scanFinalError = ""
|
||||
fake.scan = scanStatusDTO{}
|
||||
fake.scanRunPolls = 1
|
||||
fake.mu.Unlock()
|
||||
_, body = doReq(t, ts, reqOpts{method: http.MethodPost, path: "/ips/scan?dry_run=true"})
|
||||
if !strings.Contains(body, "пробный запуск") {
|
||||
t.Fatalf("expected the dry-run marker, got:\n%s", body)
|
||||
}
|
||||
_, body = doReq(t, ts, reqOpts{path: "/ips/scan/status"})
|
||||
if !strings.Contains(body, "готово") || len(fake.ips) != 0 {
|
||||
t.Fatalf("dry run must finish without queueing, queue=%+v body:\n%s", fake.ips, body)
|
||||
}
|
||||
|
||||
// control-api refuses to start: banner, no polling.
|
||||
fake.mu.Lock()
|
||||
fake.scanStartStatus = http.StatusConflict
|
||||
fake.mu.Unlock()
|
||||
_, body = doReq(t, ts, reqOpts{method: http.MethodPost, path: "/ips/scan"})
|
||||
if !strings.Contains(body, "alert-warning") || hasHXTriggerEvery(body) {
|
||||
t.Fatalf("expected a client-error banner and no polling, got:\n%s", body)
|
||||
}
|
||||
}
|
||||
|
||||
func TestIPsPageRendersRunningScanPanelOutsideForm(t *testing.T) {
|
||||
fake, caURL := newFakeControlAPI(t)
|
||||
now := time.Now()
|
||||
fake.scan = scanStatusDTO{State: "enqueuing", Running: true, Pages: 33, Discovered: 6440, Free: 6440, Added: 1200, StartedAt: &now}
|
||||
ts := newTestServer(t, caURL)
|
||||
|
||||
page := get(t, ts, "/ips")
|
||||
for _, want := range []string{`hx-get="/ips/scan/status"`, "ставятся в очередь", `<progress max="6440" value="1200"`, "<dt>прочитано страниц</dt><dd>33</dd>"} {
|
||||
if !strings.Contains(page, want) {
|
||||
t.Fatalf("expected %q in the page, got:\n%s", want, page)
|
||||
}
|
||||
}
|
||||
if strings.Index(page, `id="scan-progress"`) > strings.Index(page, `<form id="ips-form"`) {
|
||||
t.Fatalf("the scan panel must sit above (outside) #ips-form")
|
||||
}
|
||||
}
|
||||
|
||||
func TestPageURLEscapesQuery(t *testing.T) {
|
||||
got := pageURL("/ips", url.Values{"q": {"a b&c=d"}, "state": {"queued"}}, 3, 100)
|
||||
u, err := url.Parse(got)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
q := u.Query()
|
||||
if u.Path != "/ips" || q.Get("q") != "a b&c=d" || q.Get("state") != "queued" || q.Get("page") != "3" || q.Get("per_page") != "100" || len(q) != 4 {
|
||||
t.Fatalf("unexpected URL %q (query %v)", got, q)
|
||||
}
|
||||
if first := pageURL("/ips", nil, 1, 50); strings.Contains(first, "page=") && !strings.Contains(first, "per_page=") {
|
||||
t.Fatalf("page 1 must not carry page=, got %q", first)
|
||||
}
|
||||
}
|
||||
|
||||
func TestIPsPagination(t *testing.T) {
|
||||
fake, caURL := newFakeControlAPI(t)
|
||||
fake.seedIPs(120)
|
||||
ts := newTestServer(t, caURL)
|
||||
|
||||
body := get(t, ts, "/ips")
|
||||
if !strings.Contains(body, "Показано 1–50 из 120") || countRows(body) != 50 {
|
||||
t.Fatalf("page 1: want «Показано 1–50 из 120» and 50 rows, got %d rows:\n%s", countRows(body), body)
|
||||
}
|
||||
if next := pagerLink(t, body, "next"); next == nil || next.Query().Get("page") != "2" {
|
||||
t.Fatalf("expected a next link to page 2, got %v", next)
|
||||
}
|
||||
if pagerLink(t, body, "prev") != nil {
|
||||
t.Fatalf("page 1 must have no prev link")
|
||||
}
|
||||
|
||||
body = get(t, ts, "/ips?page=2")
|
||||
if !strings.Contains(body, "Показано 51–100 из 120") || countRows(body) != 50 {
|
||||
t.Fatalf("page 2: got:\n%s", body)
|
||||
}
|
||||
if prev := pagerLink(t, body, "prev"); prev == nil || prev.Query().Get("page") != "" {
|
||||
t.Fatalf("page 2 must link back to page 1 (no page param), got %v", prev)
|
||||
}
|
||||
|
||||
// Last page, and a page past the end clamps to it; junk is page 1.
|
||||
for _, p := range []string{"3", "99"} {
|
||||
body = get(t, ts, "/ips?page="+p)
|
||||
if !strings.Contains(body, "Показано 101–120 из 120") || countRows(body) != 20 || pagerLink(t, body, "next") != nil {
|
||||
t.Fatalf("page=%s: want the last page, got:\n%s", p, body)
|
||||
}
|
||||
}
|
||||
for _, p := range []string{"0", "-4", "abc"} {
|
||||
if body = get(t, ts, "/ips?page="+p); !strings.Contains(body, "Показано 1–50 из 120") {
|
||||
t.Fatalf("page=%s: want page 1, got:\n%s", p, body)
|
||||
}
|
||||
}
|
||||
|
||||
// per_page: only 25/50/100/200 are accepted, anything else is the default.
|
||||
for per, want := range map[string]int{"25": 25, "100": 100, "200": 120, "77": 50, "": 50, "x": 50} {
|
||||
if body = get(t, ts, "/ips?per_page="+per); countRows(body) != want {
|
||||
t.Fatalf("per_page=%q: want %d rows, got %d", per, want, countRows(body))
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
func TestIPsPaginationKeepsFilterInLinks(t *testing.T) {
|
||||
fake, caURL := newFakeControlAPI(t)
|
||||
fake.seedIPs(120)
|
||||
ts := newTestServer(t, caURL)
|
||||
|
||||
// 10.0.0.1 matches .1, .10-.19, .100-.120: 32 rows.
|
||||
body := get(t, ts, "/ips?q=10.0.0.1&state=queued&per_page=25")
|
||||
if !strings.Contains(body, "Показано 1–25 из 32") {
|
||||
t.Fatalf("expected the filtered total, got:\n%s", body)
|
||||
}
|
||||
next := pagerLink(t, body, "next")
|
||||
if next == nil {
|
||||
t.Fatalf("expected a next link")
|
||||
}
|
||||
q := next.Query()
|
||||
if q.Get("q") != "10.0.0.1" || q.Get("state") != "queued" || q.Get("per_page") != "25" || q.Get("page") != "2" {
|
||||
t.Fatalf("next link lost the filter: %v", next)
|
||||
}
|
||||
// The same filter is carried by the hidden inputs of #ips-form and by the
|
||||
// wrapper's own reload URL.
|
||||
for _, want := range []string{
|
||||
`name="q" value="10.0.0.1"`, `name="state" value="queued"`, `name="per_page" value="25"`, `name="page" value="1"`,
|
||||
`hx-get="/ips?per_page=25&q=10.0.0.1&state=queued"`,
|
||||
} {
|
||||
if !strings.Contains(body, want) {
|
||||
t.Fatalf("expected %q in the page, got:\n%s", want, body)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
func TestIPsFilters(t *testing.T) {
|
||||
fake, caURL := newFakeControlAPI(t)
|
||||
now := time.Now()
|
||||
mk := func(ip, state, result string, seq int) ipQueueItem {
|
||||
return ipQueueItem{IPAddress: ip, State: state, OverallResult: result, Sequence: seq, AggregatedAt: &now, CreatedAt: now, UpdatedAt: now}
|
||||
}
|
||||
fake.ips = []ipQueueItem{
|
||||
mk("1.1.1.1", "queued", "", 1), mk("1.1.1.2", "checking", "", 2), mk("1.1.1.3", "awaiting_self_check", "", 3),
|
||||
mk("2.2.2.1", "done", "pass", 4), mk("2.2.2.2", "failed", "fail", 5), mk("2.2.2.3", "failed", "cancelled", 6),
|
||||
mk("3.3.3.3", "occupied", "", 7),
|
||||
}
|
||||
ts := newTestServer(t, caURL)
|
||||
|
||||
cases := []struct {
|
||||
query string
|
||||
want []string
|
||||
}{
|
||||
{"", []string{"1.1.1.1", "1.1.1.2", "1.1.1.3", "2.2.2.1", "2.2.2.2", "2.2.2.3", "3.3.3.3"}},
|
||||
{"state=queued", []string{"1.1.1.1"}},
|
||||
{"state=active", []string{"1.1.1.2", "1.1.1.3"}},
|
||||
{"state=done", []string{"2.2.2.1"}},
|
||||
{"state=failed", []string{"2.2.2.2", "2.2.2.3"}},
|
||||
{"state=occupied", []string{"3.3.3.3"}},
|
||||
{"state=result:fail", []string{"2.2.2.2"}},
|
||||
{"result=cancelled", []string{"2.2.2.3"}},
|
||||
{"state=failed&result=cancelled", []string{"2.2.2.3"}},
|
||||
{"q=1.1.1&state=active", []string{"1.1.1.2", "1.1.1.3"}},
|
||||
{"state=bogus", []string{"1.1.1.1", "1.1.1.2", "1.1.1.3", "2.2.2.1", "2.2.2.2", "2.2.2.3", "3.3.3.3"}},
|
||||
}
|
||||
for _, c := range cases {
|
||||
body := get(t, ts, "/ips?"+c.query)
|
||||
got := map[string]bool{}
|
||||
for _, ip := range []string{"1.1.1.1", "1.1.1.2", "1.1.1.3", "2.2.2.1", "2.2.2.2", "2.2.2.3", "3.3.3.3"} {
|
||||
got[ip] = strings.Contains(body, ipRowLink(ip))
|
||||
}
|
||||
for ip, in := range got {
|
||||
if want := containsStr(c.want, ip); in != want {
|
||||
t.Fatalf("?%s: row %s present=%v, want %v", c.query, ip, in, want)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// The select echoes the active filter; an empty filtered result says so.
|
||||
if body := get(t, ts, "/ips?state=result:fail"); !strings.Contains(body, `value="result:fail" selected`) {
|
||||
t.Fatalf("expected the result option to be selected, got:\n%s", body)
|
||||
}
|
||||
if body := get(t, ts, "/ips?q=zzz"); !strings.Contains(body, "Ничего не найдено по текущему фильтру") {
|
||||
t.Fatalf("expected the filtered-empty message, got:\n%s", body)
|
||||
}
|
||||
|
||||
// An htmx filter request gets just the swappable wrapper, not the page.
|
||||
_, frag := doReq(t, ts, reqOpts{path: "/ips?state=queued", headers: map[string]string{"HX-Request": "true"}})
|
||||
if strings.Contains(frag, "<html") || !strings.Contains(frag, `id="ips-table-wrap"`) || !strings.Contains(frag, ipRowLink("1.1.1.1")) {
|
||||
t.Fatalf("expected the table wrapper fragment, got:\n%s", frag)
|
||||
}
|
||||
|
||||
// Filters are applied server-side: control-api got them as parameters.
|
||||
fake.mu.Lock()
|
||||
defer fake.mu.Unlock()
|
||||
var sawState bool
|
||||
for _, q := range fake.ipsQueries {
|
||||
if strings.Contains(q, "state=queued") && strings.Contains(q, "limit=50") {
|
||||
sawState = true
|
||||
}
|
||||
}
|
||||
if !sawState || fake.bareIPsCalls != 0 {
|
||||
t.Fatalf("expected paginated filtered requests only (bare=%d), got %v", fake.bareIPsCalls, fake.ipsQueries)
|
||||
}
|
||||
}
|
||||
|
||||
func TestIPsMutationsKeepPageAndFilter(t *testing.T) {
|
||||
fake, caURL := newFakeControlAPI(t)
|
||||
addrs := fake.seedIPs(120)
|
||||
ts := newTestServer(t, caURL)
|
||||
|
||||
// Per-row delete on page 2 (the URL carries only the context params).
|
||||
_, body := doReq(t, ts, reqOpts{method: http.MethodDelete, path: "/ips/" + addrs[60] + "?page=2&per_page=50&state=queued"})
|
||||
if !strings.Contains(body, "Показано 51–100 из 119") || strings.Contains(body, ipRowLink(addrs[60])) {
|
||||
t.Fatalf("delete must re-render page 2 of the same filter, got:\n%s", body)
|
||||
}
|
||||
for _, want := range []string{`name="page" value="2"`, `name="state" value="queued"`} {
|
||||
if !strings.Contains(body, want) {
|
||||
t.Fatalf("expected %q in the re-rendered form, got:\n%s", want, body)
|
||||
}
|
||||
}
|
||||
|
||||
// Per-row recheck keeps the page too.
|
||||
body = postForm(t, ts, "POST", "/ips/"+addrs[70]+"/recheck", url.Values{"page": {"2"}, "per_page": {"50"}})
|
||||
if !strings.Contains(body, "Показано 51–100 из 119") {
|
||||
t.Fatalf("recheck must stay on page 2, got:\n%s", body)
|
||||
}
|
||||
|
||||
// Deleting every row of the last page clamps to the last non-empty page.
|
||||
fake.mu.Lock()
|
||||
var last []string
|
||||
for _, ip := range fake.ips[100:] {
|
||||
last = append(last, ip.IPAddress)
|
||||
}
|
||||
fake.mu.Unlock()
|
||||
body = postForm(t, ts, "POST", "/ips/delete", url.Values{"page": {"3"}, "per_page": {"50"}, "addresses": last})
|
||||
if !strings.Contains(body, "Показано 51–100 из 100") || !strings.Contains(body, `name="page" value="2"`) {
|
||||
t.Fatalf("expected a clamp to page 2 after the last page emptied, got:\n%s", body)
|
||||
}
|
||||
|
||||
// Adding addresses through the top form keeps the context hidden inputs' page too.
|
||||
body = postForm(t, ts, "POST", "/ips", url.Values{"addresses": {"9.9.9.9"}, "page": {"2"}, "per_page": {"50"}})
|
||||
if !strings.Contains(body, "Показано 51–100 из 101") {
|
||||
t.Fatalf("expected page 2 of 101 after adding, got:\n%s", body)
|
||||
}
|
||||
}
|
||||
|
||||
func TestIPsBulkScopeAllResolvesByFilterInChunks(t *testing.T) {
|
||||
fake, caURL := newFakeControlAPI(t)
|
||||
fake.seedIPs(1300)
|
||||
now := time.Now()
|
||||
fake.mu.Lock()
|
||||
for _, a := range []string{"2.2.2.1", "2.2.2.2", "2.2.2.3"} {
|
||||
fake.ips = append(fake.ips, ipQueueItem{IPAddress: a, State: "done", OverallResult: "pass", AggregatedAt: &now, CreatedAt: now, UpdatedAt: now})
|
||||
}
|
||||
fake.mu.Unlock()
|
||||
ts := newTestServer(t, caURL)
|
||||
|
||||
// Recheck everything in state=done: three addresses, one chunk.
|
||||
postForm(t, ts, "POST", "/ips/recheck", url.Values{"scope": {"all"}, "state": {"done"}})
|
||||
if len(fake.submitChunks) != 1 || fake.submitChunks[0] != 3 {
|
||||
t.Fatalf("submit chunks = %v, want [3]", fake.submitChunks)
|
||||
}
|
||||
|
||||
// Delete all 1300 queued+rechecked rows matching q=10. (the checked
|
||||
// boxes of the page are ignored under scope=all): chunks of ≤500.
|
||||
body := postForm(t, ts, "POST", "/ips/delete", url.Values{
|
||||
"scope": {"all"}, "q": {"10."}, "state": {"queued"}, "per_page": {"50"}, "addresses": {"2.2.2.1"},
|
||||
})
|
||||
if got := fmt.Sprint(fake.deleteChunks); got != "[500 500 300]" {
|
||||
t.Fatalf("delete chunks = %s, want [500 500 300]", got)
|
||||
}
|
||||
fake.mu.Lock()
|
||||
left := len(fake.ips)
|
||||
var big bool
|
||||
for _, q := range fake.ipsQueries {
|
||||
v, _ := url.ParseQuery(q)
|
||||
if n, _ := strconv.Atoi(v.Get("limit")); n > 1000 || n == 0 {
|
||||
big = true
|
||||
}
|
||||
}
|
||||
bare := fake.bareIPsCalls
|
||||
fake.mu.Unlock()
|
||||
if left != 3 {
|
||||
t.Fatalf("expected only the 3 non-matching rows left, got %d", left)
|
||||
}
|
||||
if big || bare != 0 {
|
||||
t.Fatalf("every list call must be paginated with limit ≤ 1000, got %v (bare=%d)", fake.ipsQueries, bare)
|
||||
}
|
||||
if !strings.Contains(body, "Ничего не найдено по текущему фильтру") {
|
||||
t.Fatalf("expected the emptied filter view, got:\n%s", body)
|
||||
}
|
||||
|
||||
// scope=all with nothing matching is a client error, not a silent no-op.
|
||||
body = postForm(t, ts, "POST", "/ips/delete", url.Values{"scope": {"all"}, "q": {"nothing"}})
|
||||
if !strings.Contains(body, "alert-warning") {
|
||||
t.Fatalf("expected a banner for an empty scope=all, got:\n%s", body)
|
||||
}
|
||||
}
|
||||
|
||||
func TestIPsTableControls(t *testing.T) {
|
||||
fake, caURL := newFakeControlAPI(t)
|
||||
fake.seedIPs(120)
|
||||
now := time.Now()
|
||||
fake.mu.Lock()
|
||||
fake.ips[0].State = "checking"
|
||||
fake.ips[1].State = "done"
|
||||
fake.ips[1].OverallResult = "pass"
|
||||
fake.ips[1].AggregatedAt = &now
|
||||
fake.mu.Unlock()
|
||||
ts := newTestServer(t, caURL)
|
||||
|
||||
body := get(t, ts, "/ips")
|
||||
const params = `hx-params="page,per_page,q,state,result"`
|
||||
// Every per-row button and the clear/scan buttons restrict the request
|
||||
// params (htmx would otherwise put every checked checkbox into the URL).
|
||||
var rowButtons int
|
||||
for _, b := range strings.Split(body, "<button")[1:] {
|
||||
tag := b[:strings.Index(b, ">")]
|
||||
perRow := strings.Contains(tag, `hx-delete="/ips/`) ||
|
||||
(strings.Contains(tag, `hx-post="/ips/1`) && (strings.Contains(tag, "/recheck") || strings.Contains(tag, "/cancel")))
|
||||
if perRow {
|
||||
rowButtons++
|
||||
}
|
||||
if perRow || strings.Contains(tag, `hx-post="/ips/clear"`) || strings.Contains(tag, `hx-post="/ips/scan`) {
|
||||
if !strings.Contains(tag, params) {
|
||||
t.Fatalf("button without hx-params: <button%s>", tag)
|
||||
}
|
||||
}
|
||||
}
|
||||
if rowButtons != 100 { // delete + recheck/cancel on each of the 50 rows
|
||||
t.Fatalf("expected 100 per-row buttons, found %d", rowButtons)
|
||||
}
|
||||
// Bulk delete/recheck submit the whole form (checked boxes + scope), so no whitelist there.
|
||||
for _, b := range strings.Split(body, "<button")[1:] {
|
||||
tag := b[:strings.Index(b, ">")]
|
||||
if (strings.Contains(tag, `hx-post="/ips/delete"`) || strings.Contains(tag, `hx-post="/ips/recheck"`)) && strings.Contains(tag, "hx-params") {
|
||||
t.Fatalf("bulk buttons must submit the whole form: <button%s>", tag)
|
||||
}
|
||||
}
|
||||
|
||||
// Bulk selection UI with real counts.
|
||||
for _, want := range []string{
|
||||
"Выбрано на странице: <b data-sel-count>0</b> из 50",
|
||||
"Выбрать все 120 по фильтру",
|
||||
"Удалить ВСЕ 120 адресов, включая идущие проверки? Действие необратимо.",
|
||||
"Удалить ВСЕ 120 адресов из очереди, включая идущие проверки? Действие необратимо.",
|
||||
`<input type="hidden" name="scope" value="">`,
|
||||
} {
|
||||
if !strings.Contains(body, want) {
|
||||
t.Fatalf("expected %q in the page, got:\n%s", want, body)
|
||||
}
|
||||
}
|
||||
|
||||
// Under a filter the clear confirmation shows the whole queue, the
|
||||
// delete-all one the filtered count.
|
||||
body = get(t, ts, "/ips?q=10.0.0.1")
|
||||
for _, want := range []string{"Удалить ВСЕ 32 адреса по текущему фильтру, включая", "Удалить ВСЕ 120 адресов из очереди"} {
|
||||
if !strings.Contains(body, want) {
|
||||
t.Fatalf("expected %q under a filter, got:\n%s", want, body)
|
||||
}
|
||||
}
|
||||
|
||||
// Everything on one page: no "select all by filter" offer.
|
||||
body = get(t, ts, "/ips?q=10.0.0.119")
|
||||
if strings.Contains(body, `class="select-all-link"`) {
|
||||
t.Fatalf("no select-all link when the filter fits on the page, got:\n%s", body)
|
||||
}
|
||||
}
|
||||
|
||||
func TestOverviewBoundedWithThousandsQueued(t *testing.T) {
|
||||
fake, caURL := newFakeControlAPI(t)
|
||||
fake.seedIPs(5000)
|
||||
now := time.Now()
|
||||
fake.mu.Lock()
|
||||
for i := 0; i < 3; i++ {
|
||||
fake.ips[i].State = "checking"
|
||||
}
|
||||
for i := 0; i < 30; i++ {
|
||||
at := now.Add(-time.Duration(i) * time.Minute)
|
||||
fake.ips = append(fake.ips, ipQueueItem{IPAddress: fmt.Sprintf("2.2.2.%d", i), State: "done", OverallResult: "pass", AggregatedAt: &at, CreatedAt: now, UpdatedAt: now})
|
||||
}
|
||||
fake.mu.Unlock()
|
||||
ts := newTestServer(t, caURL)
|
||||
|
||||
for _, path := range []string{"/overview", "/overview/fragment"} {
|
||||
body := get(t, ts, path)
|
||||
// 3 active + the next 10 queued + the last 20 completed.
|
||||
if n := strings.Count(body, `href="/ips/`); n != 33 {
|
||||
t.Fatalf("%s: expected 33 address rows (3+10+20), got %d", path, n)
|
||||
}
|
||||
for _, want := range []string{"В очереди: <b>4997</b>", `href="/ips?state=queued"`, "Последние 20 завершённых", "Ближайшие в очереди"} {
|
||||
if !strings.Contains(body, want) {
|
||||
t.Fatalf("%s: expected %q, got:\n%s", path, want, body)
|
||||
}
|
||||
}
|
||||
if !strings.Contains(body, ipRowLink("2.2.2.0")) || strings.Contains(body, ipRowLink("2.2.2.25")) {
|
||||
t.Fatalf("%s: expected only the 20 newest completed rows", path)
|
||||
}
|
||||
}
|
||||
|
||||
fake.mu.Lock()
|
||||
defer fake.mu.Unlock()
|
||||
if fake.bareIPsCalls != 0 {
|
||||
t.Fatalf("the overview must never load the whole queue, bare calls = %d", fake.bareIPsCalls)
|
||||
}
|
||||
if len(fake.ipsQueries) == 0 {
|
||||
t.Fatalf("expected paginated list requests")
|
||||
}
|
||||
for _, q := range fake.ipsQueries {
|
||||
v, _ := url.ParseQuery(q)
|
||||
if n, _ := strconv.Atoi(v.Get("limit")); n < 1 || n > 100 {
|
||||
t.Fatalf("overview list request without a small limit: %q", q)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
func TestOverviewProgressIndicatorAndETA(t *testing.T) {
|
||||
fake, caURL := newFakeControlAPI(t)
|
||||
fake.seedIPs(100)
|
||||
now := time.Now()
|
||||
fake.mu.Lock()
|
||||
// 6 completed rows, one every 10 s (span 50 s): 10 s per address.
|
||||
for i := 0; i < 6; i++ {
|
||||
at := now.Add(-time.Duration(i) * 10 * time.Second)
|
||||
fake.ips = append(fake.ips, ipQueueItem{IPAddress: fmt.Sprintf("2.2.2.%d", i), State: "done", OverallResult: "pass", AggregatedAt: &at, CreatedAt: now, UpdatedAt: now})
|
||||
}
|
||||
// "occupied" counts as finished work too.
|
||||
for i := 0; i < 2; i++ {
|
||||
fake.ips = append(fake.ips, ipQueueItem{IPAddress: fmt.Sprintf("3.3.3.%d", i), State: "occupied", CreatedAt: now, UpdatedAt: now})
|
||||
}
|
||||
fake.mu.Unlock()
|
||||
ts := newTestServer(t, caURL)
|
||||
|
||||
body := get(t, ts, "/overview")
|
||||
for _, want := range []string{
|
||||
"Готово 8 из 108 (7%) · в работе 0 · в очереди 100",
|
||||
`<progress max="108" value="8"`,
|
||||
"осталось ≈ 16 мин 40 с", // 100 queued × 10 s
|
||||
"Итоги проверок: pass 6",
|
||||
} {
|
||||
if !strings.Contains(body, want) {
|
||||
t.Fatalf("expected %q in the stats block, got:\n%s", want, body)
|
||||
}
|
||||
}
|
||||
// The ETA is not derived from a filtered window.
|
||||
if body = get(t, ts, "/overview?q=2.2.2"); strings.Contains(body, "осталось") {
|
||||
t.Fatalf("no ETA under a filter, got:\n%s", body)
|
||||
}
|
||||
// Too few samples: no ETA.
|
||||
fake.mu.Lock()
|
||||
fake.ips = append(fake.ips[:100], fake.ips[100:103]...)
|
||||
fake.mu.Unlock()
|
||||
if body = get(t, ts, "/overview"); strings.Contains(body, "осталось") || !strings.Contains(body, "Готово 3 из 103") {
|
||||
t.Fatalf("expected progress without ETA for <5 samples, got:\n%s", body)
|
||||
}
|
||||
}
|
||||
|
||||
func TestOverviewScanLineAndScanningPhase(t *testing.T) {
|
||||
fake, caURL := newFakeControlAPI(t)
|
||||
fake.seedIPs(2)
|
||||
now := time.Now()
|
||||
fake.autoCycle = autoCycleDTO{Enabled: true, IntervalSeconds: 3600, Phase: "scanning", NextRunAt: &now}
|
||||
ts := newTestServer(t, caURL)
|
||||
|
||||
// No running scan: no scan line.
|
||||
if body := get(t, ts, "/overview"); strings.Contains(body, "Сканирование:") {
|
||||
t.Fatalf("no scan line while idle, got:\n%s", body)
|
||||
}
|
||||
|
||||
fake.mu.Lock()
|
||||
fake.scan = scanStatusDTO{State: "enqueuing", Running: true, Pages: 33, Discovered: 6440, Free: 6439, Added: 500, StartedAt: &now}
|
||||
fake.mu.Unlock()
|
||||
body := get(t, ts, "/overview")
|
||||
for _, want := range []string{"Сканирование: ставятся в очередь", "прочитано страниц 33", "найдено адресов 6440", "свободных 6439", "сканирование Floating IP", "Автоцикл активен"} {
|
||||
if !strings.Contains(body, want) {
|
||||
t.Fatalf("expected %q, got:\n%s", want, body)
|
||||
}
|
||||
}
|
||||
// The polled fragment refreshes the line via the OOB stats block.
|
||||
frag := get(t, ts, "/overview/fragment")
|
||||
oob := strings.Index(frag, `id="overview-stats" hx-swap-oob="true"`)
|
||||
if oob < 0 || !strings.Contains(frag[oob:], "Сканирование: ставятся в очередь") {
|
||||
t.Fatalf("expected the scan line inside the OOB stats block, got:\n%s", frag)
|
||||
}
|
||||
}
|
||||
|
||||
func TestRegistryPaginationAndServerSideFilter(t *testing.T) {
|
||||
fake, caURL := newFakeControlAPI(t)
|
||||
now := time.Now()
|
||||
results := []string{"pass", "fail", "partial"}
|
||||
for i := 0; i < 120; i++ {
|
||||
ip := fmt.Sprintf("10.0.0.%d", i+1)
|
||||
fake.registry[ip] = registryItem{IPAddress: ip, FirstSeenAt: now, LastSeenAt: now, TotalCycles: 1, LastResult: results[i%3]}
|
||||
}
|
||||
ts := newTestServer(t, caURL)
|
||||
|
||||
body := get(t, ts, "/registry?page=2")
|
||||
if !strings.Contains(body, "Показано 51–100 из 120") || strings.Count(body, `href="/registry/10.`) != 50 {
|
||||
t.Fatalf("expected page 2 with 50 rows, got:\n%s", body)
|
||||
}
|
||||
if body = get(t, ts, "/registry?page=99"); !strings.Contains(body, "Показано 101–120 из 120") {
|
||||
t.Fatalf("expected a clamp to the last page, got:\n%s", body)
|
||||
}
|
||||
|
||||
// 40 of the 120 are "fail"; q and status survive in the pager links and
|
||||
// reach control-api as parameters (no client-side filtering).
|
||||
body = get(t, ts, "/registry?status=fail&q=10.0.0&per_page=25")
|
||||
if !strings.Contains(body, "Показано 1–25 из 40") {
|
||||
t.Fatalf("expected the server-side filtered total, got:\n%s", body)
|
||||
}
|
||||
next := pagerLink(t, body, "next")
|
||||
if next == nil || next.Path != "/registry" || next.Query().Get("status") != "fail" || next.Query().Get("q") != "10.0.0" ||
|
||||
next.Query().Get("per_page") != "25" || next.Query().Get("page") != "2" {
|
||||
t.Fatalf("pager link lost the filter: %v", next)
|
||||
}
|
||||
fake.mu.Lock()
|
||||
defer fake.mu.Unlock()
|
||||
last := fake.registryQueries[len(fake.registryQueries)-1]
|
||||
for _, want := range []string{"last_result=fail", "q=10.0.0", "limit=25"} {
|
||||
if !strings.Contains(last, want) {
|
||||
t.Fatalf("control-api request %q lacks %s", last, want)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
func TestLongTimeoutForClearAndBulkCalls(t *testing.T) {
|
||||
// A short per-call timeout (50 ms) breaks plain reads of a slow control-api
|
||||
// but not clear/bulk operations, which use the long client.
|
||||
slow := newSlowAPI(t, 200*time.Millisecond)
|
||||
c := newClient(slow, 50*time.Millisecond)
|
||||
if _, err := c.Status(t.Context()); err == nil {
|
||||
t.Fatalf("expected the short timeout to fail a plain read")
|
||||
}
|
||||
if _, err := c.ClearQueue(t.Context()); err != nil {
|
||||
t.Fatalf("ClearQueue must use the long timeout: %v", err)
|
||||
}
|
||||
if _, err := c.DeleteIPs(t.Context(), []string{"1.1.1.1"}); err != nil {
|
||||
t.Fatalf("DeleteIPs must use the long timeout: %v", err)
|
||||
}
|
||||
if _, err := c.SubmitIPs(t.Context(), []string{"1.1.1.1"}); err != nil {
|
||||
t.Fatalf("SubmitIPs must use the long timeout: %v", err)
|
||||
}
|
||||
}
|
||||
|
||||
func TestScanRoutesPassOriginCheckWhenAuthEnabled(t *testing.T) {
|
||||
fake, caURL := newFakeControlAPI(t)
|
||||
fake.scanRunPolls = 1
|
||||
_, ts := newAuthTestServer(t, caURL, nil)
|
||||
cookie := login(t, ts)
|
||||
|
||||
resp, body := doReq(t, ts, reqOpts{method: http.MethodPost, path: "/ips/scan", cookie: cookie})
|
||||
if resp.StatusCode != http.StatusForbidden {
|
||||
t.Fatalf("scan without Origin: status=%d, want 403", resp.StatusCode)
|
||||
}
|
||||
resp, body = doReq(t, ts, reqOpts{method: http.MethodPost, path: "/ips/scan", cookie: cookie, headers: map[string]string{"Origin": ts.URL}})
|
||||
if resp.StatusCode != http.StatusOK || !strings.Contains(body, `id="scan-progress"`) {
|
||||
t.Fatalf("scan with Origin: status=%d body:\n%s", resp.StatusCode, body)
|
||||
}
|
||||
resp, body = doReq(t, ts, reqOpts{path: "/ips/scan/status", cookie: cookie})
|
||||
if resp.StatusCode != http.StatusOK || resp.Header.Get("HX-Trigger") != "scan-finished" {
|
||||
t.Fatalf("status poll: status=%d trigger=%q body:\n%s", resp.StatusCode, resp.Header.Get("HX-Trigger"), body)
|
||||
}
|
||||
}
|
||||
@@ -559,3 +559,53 @@ textarea.autosize { overflow-y: hidden; resize: none; }
|
||||
}
|
||||
.sidebar-user-name { min-width: 0; overflow: hidden; text-overflow: ellipsis; white-space: nowrap; }
|
||||
.sidebar-user + .sidebar-foot { margin-top: 0; }
|
||||
|
||||
/* ---------- scan progress, progress bars, pager, bulk selection ---------- */
|
||||
progress {
|
||||
appearance: none; -webkit-appearance: none;
|
||||
display: block; width: 100%; height: 8px;
|
||||
border: 0; border-radius: var(--radius-sm);
|
||||
background: var(--surface-alt); color: var(--accent);
|
||||
overflow: hidden;
|
||||
}
|
||||
progress::-webkit-progress-bar { background: var(--surface-alt); border-radius: var(--radius-sm); }
|
||||
progress::-webkit-progress-value { background: var(--accent); border-radius: var(--radius-sm); }
|
||||
progress::-moz-progress-bar { background: var(--accent); border-radius: var(--radius-sm); }
|
||||
/* Indeterminate (no value attribute): a sliding stripe, Firefox/Chromium. */
|
||||
progress:indeterminate { background: linear-gradient(90deg, var(--surface-alt) 0, var(--accent-soft) 40%, var(--surface-alt) 80%); background-size: 200% 100%; animation: progress-slide 1.4s linear infinite; }
|
||||
progress:indeterminate::-webkit-progress-value { background: transparent; }
|
||||
progress:indeterminate::-moz-progress-bar { background: transparent; }
|
||||
@keyframes progress-slide { from { background-position: 200% 0; } to { background-position: -200% 0; } }
|
||||
@media (prefers-reduced-motion: reduce) { progress:indeterminate { animation: none; } }
|
||||
|
||||
.overall-progress { margin-top: 12px; font-size: 12.5px; color: var(--text-muted); }
|
||||
.overall-progress p { margin-bottom: 6px; }
|
||||
|
||||
.scan-progress { margin-top: 16px; }
|
||||
.scan-actions { display: flex; gap: 8px; align-items: center; flex-wrap: wrap; margin-bottom: 10px; }
|
||||
.scan-progress progress { margin-bottom: 10px; }
|
||||
.scan-counters { display: flex; flex-wrap: wrap; gap: 6px 22px; font-size: 12.5px; color: var(--text-muted); }
|
||||
.scan-counters dt { display: inline; }
|
||||
.scan-counters dd { display: inline; margin: 0 0 0 4px; color: var(--text); font-family: var(--font-mono); font-variant-numeric: tabular-nums; }
|
||||
.scan-error { margin-top: 8px; padding: 8px 10px; border: 1px solid var(--danger-border); border-radius: var(--radius-xs); background: var(--danger-soft); color: var(--danger); font-size: 12.5px; }
|
||||
.btn[disabled] { opacity: .5; cursor: not-allowed; }
|
||||
|
||||
.pager { display: flex; align-items: center; justify-content: space-between; gap: 12px; flex-wrap: wrap; margin-top: 12px; font-size: 12.5px; color: var(--text-muted); }
|
||||
.pager-nav { display: inline-flex; align-items: center; gap: 8px; }
|
||||
.pager-page { font-family: var(--font-mono); font-variant-numeric: tabular-nums; }
|
||||
.btn.is-disabled { opacity: .4; pointer-events: none; }
|
||||
.panel > .pager { margin: 0; padding: 10px 16px; border-top: 1px solid var(--border-soft); }
|
||||
|
||||
.bulk-bar { display: flex; gap: 8px; align-items: center; flex-wrap: wrap; }
|
||||
.bulk-hint { padding: 8px 16px; font-size: 12.5px; color: var(--text-muted); border-top: 1px solid var(--border-soft); border-bottom: 1px solid var(--border-soft); background: var(--surface-alt); }
|
||||
.bulk-hint b { color: var(--text); font-family: var(--font-mono); font-weight: 500; }
|
||||
.select-all-link, .scope-all-msg, .btn-del-all { display: none; }
|
||||
#ips-form.sel-all .select-all-link { display: inline; margin-left: 10px; }
|
||||
#ips-form.scope-all .select-all-link { display: none; }
|
||||
#ips-form.scope-all .scope-all-msg { display: inline; margin-left: 10px; color: var(--warning); }
|
||||
#ips-form.scope-all .btn-del-all { display: inline-flex; }
|
||||
#ips-form.scope-all .btn-del-sel { display: none; }
|
||||
|
||||
.queue-line { margin: 14px 0 8px; font-size: 13px; color: var(--text-muted); }
|
||||
.queue-line b { color: var(--text); font-family: var(--font-mono); font-weight: 500; }
|
||||
table.compact td, table.compact th { padding-top: 6px; padding-bottom: 6px; }
|
||||
@@ -24,7 +24,7 @@
|
||||
|
||||
<div class="panel">
|
||||
<div class="panel-body">
|
||||
<form hx-post="/ips" hx-target="#ips-table-wrap" hx-swap="innerHTML" hx-sync="#ips-table-wrap:queue last" hx-on::after-request="this.reset()">
|
||||
<form hx-post="/ips" hx-target="#ips-table-wrap" hx-swap="outerHTML" hx-include=".ips-ctx" hx-sync="#ips-table-wrap:queue last" hx-on::after-request="this.reset()">
|
||||
<div class="field-row">
|
||||
<div class="field" style="flex: 1 1 420px;">
|
||||
<label for="addresses">Адреса (по одному на строку или через запятую)</label>
|
||||
@@ -36,28 +36,127 @@
|
||||
<p class="muted" style="margin-top:8px">Новый адрес встаёт в очередь; уже завершённый (done/failed) запускается заново
|
||||
(тот же вызов); тот, что сейчас проверяется, не трогается. Кнопка «Сканировать Floating IP» ниже делает то же самое
|
||||
автоматически: находит в проекте OpenStack все свободные (не привязанные к порту) Floating IP и сразу передаёт их
|
||||
в очередь на проверку. Полная история проверок по каждому адресу — включая уже удалённые из очереди — доступна в
|
||||
<a href="/registry">реестре</a>.</p>
|
||||
в очередь на проверку — сканирование идёт в фоне, прогресс виден в панели под кнопкой. Полная история проверок по
|
||||
каждому адресу — включая уже удалённые из очереди — доступна в <a href="/registry">реестре</a>.</p>
|
||||
</div>
|
||||
</div>
|
||||
|
||||
<div id="ips-table-wrap">
|
||||
{{template "scan_progress" .Scan}}
|
||||
|
||||
<form id="ips-filter" class="panel" onsubmit="return false" style="margin-top:16px">
|
||||
<div class="panel-body field-row">
|
||||
<div class="field" style="flex:1 1 260px">
|
||||
<label for="ips-q">Поиск по IP</label>
|
||||
<input type="search" id="ips-q" name="q" value="{{.Filter.Q}}" placeholder="203.0.113.10"
|
||||
hx-get="/ips" hx-select="#ips-table-wrap" hx-target="#ips-table-wrap" hx-swap="outerHTML"
|
||||
hx-include="#ips-filter" hx-trigger="input changed delay:300ms"
|
||||
hx-replace-url="true" hx-sync="#ips-table-wrap:queue last">
|
||||
</div>
|
||||
<div class="field">
|
||||
<label for="ips-state">Состояние</label>
|
||||
<select id="ips-state" name="state"
|
||||
hx-get="/ips" hx-select="#ips-table-wrap" hx-target="#ips-table-wrap" hx-swap="outerHTML"
|
||||
hx-include="#ips-filter" hx-trigger="change"
|
||||
hx-replace-url="true" hx-sync="#ips-table-wrap:queue last">
|
||||
{{$tok := .Filter.Token}}
|
||||
<option value="">Все</option>
|
||||
<option value="queued" {{if eq $tok "queued"}}selected{{end}}>в очереди (queued)</option>
|
||||
<option value="active" {{if eq $tok "active"}}selected{{end}}>в работе (assigning_fip, awaiting_self_check, checking, aggregating)</option>
|
||||
<option value="done" {{if eq $tok "done"}}selected{{end}}>done</option>
|
||||
<option value="failed" {{if eq $tok "failed"}}selected{{end}}>failed</option>
|
||||
<option value="occupied" {{if eq $tok "occupied"}}selected{{end}}>occupied</option>
|
||||
<option value="result:pass" {{if eq $tok "result:pass"}}selected{{end}}>результат: pass</option>
|
||||
<option value="result:partial" {{if eq $tok "result:partial"}}selected{{end}}>результат: partial</option>
|
||||
<option value="result:fail" {{if eq $tok "result:fail"}}selected{{end}}>результат: fail</option>
|
||||
<option value="result:cancelled" {{if eq $tok "result:cancelled"}}selected{{end}}>результат: cancelled</option>
|
||||
</select>
|
||||
</div>
|
||||
<div class="field" style="flex:0 0 120px">
|
||||
<label for="ips-per-page">На странице</label>
|
||||
<select id="ips-per-page" name="per_page"
|
||||
hx-get="/ips" hx-select="#ips-table-wrap" hx-target="#ips-table-wrap" hx-swap="outerHTML"
|
||||
hx-include="#ips-filter" hx-trigger="change"
|
||||
hx-replace-url="true" hx-sync="#ips-table-wrap:queue last">
|
||||
{{$pp := .PerPage}}
|
||||
{{range .PerPageOptions}}<option value="{{.}}" {{if eq . $pp}}selected{{end}}>{{.}}</option>
|
||||
{{end}}
|
||||
</select>
|
||||
</div>
|
||||
</div>
|
||||
</form>
|
||||
|
||||
{{template "ips_table_wrap" .}}
|
||||
{{end}}
|
||||
|
||||
{{/* #ips-table-wrap is the swap target of every table mutation, the filter form
|
||||
and the pager (all hx-swap="outerHTML", so the wrapper's own hx-get always
|
||||
points at the page/filter it currently shows). It listens for the
|
||||
scan-finished event (HX-Trigger from the scan progress panel) and reloads
|
||||
itself. hx-disinherit="*": its attributes must not leak into the buttons
|
||||
inside. */}}
|
||||
{{define "ips_table_wrap"}}
|
||||
<div id="ips-table-wrap" hx-get="{{.SelfURL}}" hx-trigger="scan-finished from:body" hx-select="#ips-table-wrap" hx-target="this" hx-swap="outerHTML" hx-disinherit="*">
|
||||
{{template "ips_table" .}}
|
||||
</div>
|
||||
{{end}}
|
||||
|
||||
{{/* Scan progress panel. Lives outside #ips-form (no checkbox payload, survives
|
||||
table swaps). While the job runs it polls itself; the response that finds
|
||||
it finished has no hx-trigger and carries HX-Trigger: scan-finished. */}}
|
||||
{{define "scan_progress"}}
|
||||
{{$s := .Status}}
|
||||
<div id="scan-progress" class="panel scan-progress"{{if .Poll}} hx-get="/ips/scan/status" hx-trigger="every 2s" hx-swap="outerHTML"{{end}}>
|
||||
<div class="panel-body">
|
||||
<div class="scan-actions">
|
||||
<button type="button" class="btn btn-ghost btn-sm" hx-post="/ips/scan" hx-target="#scan-progress" hx-swap="outerHTML" hx-params="page,per_page,q,state,result"{{if $s.Running}} disabled{{end}}>Сканировать Floating IP</button>
|
||||
<button type="button" class="btn btn-ghost btn-sm" hx-post="/ips/scan?dry_run=true" hx-target="#scan-progress" hx-swap="outerHTML" hx-params="page,per_page,q,state,result"{{if $s.Running}} disabled{{end}} title="Только найти и посчитать свободные адреса, очередь не меняется">Пробное сканирование</button>
|
||||
{{if and $s.State (ne $s.State "idle")}}<span class="pill {{$s.PillClass}}">{{$s.StateLabel}}</span>{{if $s.DryRun}} <span class="muted">пробный запуск: очередь не меняется</span>{{end}}{{else}}<span class="muted">Найти все свободные Floating IP проекта и поставить их в очередь.</span>{{end}}
|
||||
</div>
|
||||
{{if and $s.State (ne $s.State "idle")}}
|
||||
{{if or $s.Running (eq $s.State "done")}}
|
||||
{{if $s.Indeterminate}}<progress aria-label="Сканирование"></progress>
|
||||
{{else if eq $s.State "done"}}<progress max="1" value="1" aria-label="Сканирование"></progress>
|
||||
{{else}}<progress max="{{$s.Free}}" value="{{$s.Handled}}" aria-label="Сканирование"></progress>{{end}}
|
||||
{{end}}
|
||||
<dl class="scan-counters">
|
||||
<div><dt>прочитано страниц</dt><dd>{{$s.Pages}}</dd></div>
|
||||
<div><dt>найдено адресов</dt><dd>{{$s.Discovered}}</dd></div>
|
||||
<div><dt>свободных</dt><dd>{{$s.Free}}</dd></div>
|
||||
<div><dt>добавлено</dt><dd>{{$s.Added}}</dd></div>
|
||||
<div><dt>повторно</dt><dd>{{$s.Requeued}}</dd></div>
|
||||
<div><dt>переупорядочено</dt><dd>{{$s.Reordered}}</dd></div>
|
||||
{{if $s.SkippedInProgress}}<div><dt>уже в работе</dt><dd>{{$s.SkippedInProgress}}</dd></div>{{end}}
|
||||
{{with $s.Elapsed}}<div><dt>время</dt><dd>{{.}}</dd></div>{{end}}
|
||||
</dl>
|
||||
{{if $s.Error}}<p class="scan-error">{{$s.Error}}</p>{{end}}
|
||||
{{end}}
|
||||
</div>
|
||||
</div>
|
||||
{{end}}
|
||||
|
||||
{{define "ips_table"}}
|
||||
<form id="ips-form">
|
||||
<form id="ips-form" onchange="var f=this,k=f.querySelectorAll('input[name=addresses]:checked').length,n=f.querySelectorAll('input[name=addresses]').length;f.querySelector('[data-sel-count]').textContent=k;f.classList.toggle('sel-all',n>0&&k===n);if(k!==n){f.elements.scope.value='';f.classList.remove('scope-all')}">
|
||||
<input type="hidden" class="ips-ctx" name="page" value="{{.Page}}">
|
||||
<input type="hidden" class="ips-ctx" name="per_page" value="{{.PerPage}}">
|
||||
<input type="hidden" class="ips-ctx" name="q" value="{{.Filter.Q}}">
|
||||
<input type="hidden" class="ips-ctx" name="state" value="{{.Filter.State}}">
|
||||
<input type="hidden" class="ips-ctx" name="result" value="{{.Filter.Result}}">
|
||||
<input type="hidden" name="scope" value="">
|
||||
<div class="panel">
|
||||
<div class="panel-body" style="display:flex; gap:8px; align-items:center;">
|
||||
<button type="button" class="btn btn-ghost btn-sm" hx-post="/ips/scan" hx-target="#ips-table-wrap" hx-swap="innerHTML" hx-sync="#ips-table-wrap:queue last">Сканировать Floating IP</button>
|
||||
<button type="button" class="btn btn-ghost btn-sm" hx-post="/ips/recheck" hx-target="#ips-table-wrap" hx-swap="innerHTML" hx-sync="#ips-table-wrap:queue last">Перепроверить выбранные</button>
|
||||
<button type="button" class="btn btn-danger-ghost btn-sm" hx-post="/ips/delete" hx-target="#ips-table-wrap" hx-swap="innerHTML" hx-sync="#ips-table-wrap:queue last" hx-confirm="Удалить выбранные адреса без возможности восстановления?">Удалить выбранные</button>
|
||||
<button type="button" class="btn btn-danger-ghost btn-sm" hx-post="/ips/clear" hx-target="#ips-table-wrap" hx-swap="innerHTML" hx-sync="#ips-table-wrap:queue last" hx-confirm="Удалить ВСЕ адреса из очереди, включая те, что сейчас проверяются? Действие необратимо.">Очистить всё</button>
|
||||
<div class="panel-body bulk-bar">
|
||||
<button type="button" class="btn btn-ghost btn-sm" hx-post="/ips/recheck" hx-target="#ips-table-wrap" hx-swap="outerHTML" hx-sync="#ips-table-wrap:queue last">Перепроверить выбранные</button>
|
||||
<button type="button" class="btn btn-danger-ghost btn-sm btn-del-sel" hx-post="/ips/delete" hx-target="#ips-table-wrap" hx-swap="outerHTML" hx-sync="#ips-table-wrap:queue last" hx-confirm="Удалить выбранные адреса без возможности восстановления?">Удалить выбранные</button>
|
||||
<button type="button" class="btn btn-danger-ghost btn-sm btn-del-all" hx-post="/ips/delete" hx-target="#ips-table-wrap" hx-swap="outerHTML" hx-sync="#ips-table-wrap:queue last" hx-confirm="Удалить ВСЕ {{.Total}} {{pluralAddr .Total}}{{if .Filter.Active}} по текущему фильтру{{end}}, включая идущие проверки? Действие необратимо.">Удалить все {{.Total}} по фильтру</button>
|
||||
<button type="button" class="btn btn-danger-ghost btn-sm" hx-post="/ips/clear" hx-target="#ips-table-wrap" hx-swap="outerHTML" hx-params="page,per_page,q,state,result" hx-sync="#ips-table-wrap:queue last" hx-confirm="Удалить ВСЕ {{.QueueTotal}} {{pluralAddr .QueueTotal}} из очереди, включая идущие проверки? Действие необратимо.">Очистить всё</button>
|
||||
</div>
|
||||
<div class="bulk-hint">
|
||||
Выбрано на странице: <b data-sel-count>0</b> из {{len .Items}}
|
||||
{{if gt .Total (len .Items)}}<a href="#" class="select-all-link" onclick="var f=this.closest('form');f.elements.scope.value='all';f.classList.add('scope-all');return false">Выбрать все {{.Total}} по фильтру</a>
|
||||
<span class="scope-all-msg">Выбраны все {{.Total}} {{pluralAddr .Total}} по фильтру (не только на этой странице).</span>{{end}}
|
||||
</div>
|
||||
<div class="table-scroll">
|
||||
<table>
|
||||
<thead><tr><th><input type="checkbox" onclick="this.closest('table').querySelectorAll('input[name=addresses]').forEach(cb => cb.checked = this.checked)"></th><th>Адрес</th><th>Состояние</th><th>Валидатор</th><th>Попытка</th><th>Обновлено</th><th></th></tr></thead>
|
||||
<thead><tr><th><input type="checkbox" aria-label="Выбрать все на странице" onclick="this.closest('table').querySelectorAll('input[name=addresses]').forEach(cb => cb.checked = this.checked)"></th><th>Адрес</th><th>Состояние</th><th>Валидатор</th><th>Попытка</th><th>Обновлено</th><th></th></tr></thead>
|
||||
<tbody>
|
||||
{{range .Items}}
|
||||
{{$b := ipBadge .State .OverallResult .FIPAssociatedAt $.FIPSettleSeconds}}
|
||||
@@ -72,11 +171,11 @@
|
||||
<td data-label="">
|
||||
<div class="actions">
|
||||
{{if $terminal}}
|
||||
<button class="btn btn-ghost btn-sm" hx-post="/ips/{{.IPAddress}}/recheck" hx-target="#ips-table-wrap" hx-swap="innerHTML" hx-sync="#ips-table-wrap:queue last">Перепроверить</button>
|
||||
<button class="btn btn-ghost btn-sm" hx-post="/ips/{{.IPAddress}}/recheck" hx-target="#ips-table-wrap" hx-swap="outerHTML" hx-params="page,per_page,q,state,result" hx-sync="#ips-table-wrap:queue last">Перепроверить</button>
|
||||
{{else}}
|
||||
<button class="btn btn-danger-ghost btn-sm" hx-post="/ips/{{.IPAddress}}/cancel" hx-target="#ips-table-wrap" hx-swap="innerHTML" hx-sync="#ips-table-wrap:queue last" hx-confirm="Остановить проверку {{.IPAddress}}?">Отменить</button>
|
||||
<button class="btn btn-danger-ghost btn-sm" hx-post="/ips/{{.IPAddress}}/cancel" hx-target="#ips-table-wrap" hx-swap="outerHTML" hx-params="page,per_page,q,state,result" hx-sync="#ips-table-wrap:queue last" hx-confirm="Остановить проверку {{.IPAddress}}?">Отменить</button>
|
||||
{{end}}
|
||||
<button class="btn btn-danger-ghost btn-sm" hx-delete="/ips/{{.IPAddress}}" hx-target="#ips-table-wrap" hx-swap="innerHTML" hx-sync="#ips-table-wrap:queue last" hx-confirm="Удалить {{.IPAddress}} без возможности восстановления?">Удалить</button>
|
||||
<button class="btn btn-danger-ghost btn-sm" hx-delete="/ips/{{.IPAddress}}" hx-target="#ips-table-wrap" hx-swap="outerHTML" hx-params="page,per_page,q,state,result" hx-sync="#ips-table-wrap:queue last" hx-confirm="Удалить {{.IPAddress}} без возможности восстановления?">Удалить</button>
|
||||
</div>
|
||||
</td>
|
||||
</tr>
|
||||
@@ -85,6 +184,18 @@
|
||||
</table>
|
||||
</div>
|
||||
</div>
|
||||
{{template "pager" .Pager}}
|
||||
</form>
|
||||
{{if not .Items}}<p class="muted">Очередь пуста.</p>{{end}}
|
||||
{{if not .Items}}<p class="muted">{{if .Filter.Active}}Ничего не найдено по текущему фильтру.{{else}}Очередь пуста.{{end}}</p>{{end}}
|
||||
{{end}}
|
||||
|
||||
{{define "pager"}}{{if .Total}}
|
||||
<div class="pager">
|
||||
<span class="pager-info">Показано {{.From}}–{{.To}} из {{.Total}}</span>
|
||||
<span class="pager-nav">
|
||||
{{if .PrevURL}}<a class="btn btn-ghost btn-sm" href="{{.PrevURL}}" hx-get="{{.PrevURL}}" hx-select="#{{.Wrap}}" hx-target="#{{.Wrap}}" hx-swap="outerHTML" hx-push-url="true" rel="prev" aria-label="Предыдущая страница">‹</a>{{else}}<span class="btn btn-ghost btn-sm is-disabled" aria-disabled="true">‹</span>{{end}}
|
||||
<span class="pager-page">{{.Page}} / {{.Pages}}</span>
|
||||
{{if .NextURL}}<a class="btn btn-ghost btn-sm" href="{{.NextURL}}" hx-get="{{.NextURL}}" hx-select="#{{.Wrap}}" hx-target="#{{.Wrap}}" hx-swap="outerHTML" hx-push-url="true" rel="next" aria-label="Следующая страница">›</a>{{else}}<span class="btn btn-ghost btn-sm is-disabled" aria-disabled="true">›</span>{{end}}
|
||||
</span>
|
||||
</div>
|
||||
{{end}}{{end}}
|
||||
@@ -50,8 +50,8 @@
|
||||
</select>
|
||||
</div>
|
||||
</div>
|
||||
<p class="muted" style="padding:0 16px 12px">Действует на обе таблицы ниже. Фильтр по статусу — это фильтр по
|
||||
итоговому результату, поэтому при выборе конкретного статуса строки «Текущей проверки» (у неё ещё нет результата)
|
||||
<p class="muted" style="padding:0 16px 12px">Действует на все списки ниже. Фильтр по статусу — это фильтр по
|
||||
итоговому результату, поэтому при выборе конкретного статуса строки «В работе» и «В очереди» (у них ещё нет результата)
|
||||
не показываются.</p>
|
||||
</form>
|
||||
|
||||
|
||||
@@ -6,6 +6,18 @@
|
||||
<div class="stat-card{{if eq $state "failed"}} bad{{else if eq $state "checking"}} accented{{end}}"><span class="value">{{$count}}</span><span class="label">{{$state}}</span></div>
|
||||
{{end}}
|
||||
</div>
|
||||
{{if .Progress.Total}}
|
||||
<div class="overall-progress" id="overview-progress">
|
||||
<p>Готово {{.Progress.Done}} из {{.Progress.Total}} ({{.Progress.Percent}}%) · в работе {{.Progress.Active}} · в очереди {{.Progress.Queued}}{{if .Progress.ETA}} · осталось ≈ {{.Progress.ETA}}{{end}}</p>
|
||||
<progress max="{{.Progress.Total}}" value="{{.Progress.Done}}" aria-label="Готово"></progress>
|
||||
</div>
|
||||
{{end}}
|
||||
{{with .Status.ResultsByOverall}}{{if or (index . "pass") (index . "partial") (index . "fail") (index . "cancelled")}}
|
||||
<p class="muted" id="overview-results" style="margin-top:8px">Итоги проверок: pass {{index . "pass"}} · partial {{index . "partial"}} · fail {{index . "fail"}} · cancelled {{index . "cancelled"}}</p>
|
||||
{{end}}{{end}}
|
||||
{{with .Scan}}
|
||||
<p class="muted" id="overview-scan" style="margin-top:8px">Сканирование: {{.StateLabel}}{{if .DryRun}} (пробное){{end}} · прочитано страниц {{.Pages}}, найдено адресов {{.Discovered}}, свободных {{.Free}}{{if .Handled}}, обработано {{.Handled}}{{end}} · <a href="/ips">подробнее</a></p>
|
||||
{{end}}
|
||||
{{with .AutoCycle}}{{if .Enabled}}
|
||||
<p class="muted" id="overview-auto-cycle" style="margin-top:12px"><span class="pill pill-info">Автоцикл активен</span> · {{.PhaseLabel}} · следующий запуск: {{fmtTime .NextRunAt}}</p>
|
||||
{{end}}{{end}}
|
||||
@@ -20,14 +32,14 @@
|
||||
{{define "overview_stats_oob"}}<div id="overview-stats" hx-swap-oob="true">{{template "overview_stats" .}}</div>{{end}}
|
||||
|
||||
{{define "overview_tables"}}
|
||||
<h2 class="section-title">Текущая проверка</h2>
|
||||
{{if .CurrentItems}}
|
||||
<h2 class="section-title">В работе</h2>
|
||||
{{if .ActiveItems}}
|
||||
<div class="panel">
|
||||
<div class="table-scroll">
|
||||
<table>
|
||||
<thead><tr><th>Адрес</th><th>Состояние</th><th>Валидатор</th><th>Попытка</th><th>Назначено</th></tr></thead>
|
||||
<tbody>
|
||||
{{range .CurrentItems}}
|
||||
{{range .ActiveItems}}
|
||||
{{$b := ipBadge .State .OverallResult .FIPAssociatedAt 0}}
|
||||
<tr>
|
||||
<td class="addr" data-label="Адрес"><a href="/ips/{{.IPAddress}}">{{.IPAddress}}</a></td>
|
||||
@@ -41,11 +53,35 @@
|
||||
</table>
|
||||
</div>
|
||||
</div>
|
||||
{{if gt .ActiveTotal (len .ActiveItems)}}<p class="muted">Показано {{len .ActiveItems}} из {{.ActiveTotal}} · <a href="/ips?state=active{{if .Query}}&q={{.Query}}{{end}}">все в работе</a></p>{{end}}
|
||||
{{else}}
|
||||
{{if or .Query .StatusFilter}}<p class="muted">Ничего не найдено по текущему фильтру.</p>
|
||||
{{else}}<p class="muted">Сейчас нет адресов в обработке.</p>{{end}}
|
||||
{{end}}
|
||||
|
||||
{{if not .StatusFilter}}
|
||||
<p class="queue-line">В очереди: <b>{{.QueuedTotal}}</b> · <a href="/ips?state=queued{{if .Query}}&q={{.Query}}{{end}}">открыть весь список</a></p>
|
||||
{{if .QueuedItems}}
|
||||
<div class="panel">
|
||||
<div class="table-scroll">
|
||||
<table class="compact">
|
||||
<thead><tr><th>Ближайшие в очереди</th><th>Состояние</th></tr></thead>
|
||||
<tbody>
|
||||
{{range .QueuedItems}}
|
||||
{{$b := ipBadge .State .OverallResult .FIPAssociatedAt 0}}
|
||||
<tr>
|
||||
<td class="addr" data-label="Адрес"><a href="/ips/{{.IPAddress}}">{{.IPAddress}}</a></td>
|
||||
<td data-label="Состояние"><span class="pill {{$b.Class}}">{{$b.Label}}</span></td>
|
||||
</tr>
|
||||
{{end}}
|
||||
</tbody>
|
||||
</table>
|
||||
</div>
|
||||
</div>
|
||||
{{if gt .QueuedTotal (len .QueuedItems)}}<p class="muted">Показаны ближайшие {{len .QueuedItems}} из {{.QueuedTotal}}.</p>{{end}}
|
||||
{{end}}
|
||||
{{end}}
|
||||
|
||||
<h2 class="section-title">Последние {{.LastN}} завершённых</h2>
|
||||
<div class="breakdown">
|
||||
<span>pass <b>{{index .Breakdown "pass"}}</b></span>
|
||||
|
||||
@@ -47,6 +47,17 @@
|
||||
<option value="cancelled" {{if eq .StatusFilter "cancelled"}}selected{{end}}>cancelled</option>
|
||||
</select>
|
||||
</div>
|
||||
<div class="field" style="flex:0 0 120px">
|
||||
<label for="registry-per-page">На странице</label>
|
||||
<select id="registry-per-page" name="per_page"
|
||||
hx-get="/registry" hx-select="#registry-table-wrap" hx-target="#registry-table-wrap" hx-swap="outerHTML"
|
||||
hx-include="#registry-filter" hx-trigger="change"
|
||||
hx-replace-url="true" hx-sync="#registry-table-wrap:queue last">
|
||||
{{$pp := .PerPage}}
|
||||
{{range .PerPageOptions}}<option value="{{.}}" {{if eq . $pp}}selected{{end}}>{{.}}</option>
|
||||
{{end}}
|
||||
</select>
|
||||
</div>
|
||||
</div>
|
||||
</form>
|
||||
|
||||
@@ -83,6 +94,7 @@
|
||||
</tbody>
|
||||
</table>
|
||||
</div>
|
||||
{{template "pager" .Pager}}
|
||||
</div>
|
||||
{{else if or .Query .StatusFilter}}<p class="muted">Ничего не найдено по текущему фильтру.</p>
|
||||
{{else}}<p class="muted">Реестр пуст — ни один адрес ещё не ставился на проверку.</p>{{end}}
|
||||
|
||||
@@ -37,6 +37,9 @@ var ipRegistrySchema string
|
||||
//go:embed migrations/0008_auto_cycle.sql
|
||||
var autoCycleSchema string
|
||||
|
||||
//go:embed migrations/0009_scale_indexes.sql
|
||||
var scaleIndexesSchema string
|
||||
|
||||
// migrations is the ordered list of schema versions. Each entry's SQL is
|
||||
// applied, in order, for any version greater than the database's current
|
||||
// PRAGMA user_version — so a fresh database walks the whole list and an
|
||||
@@ -53,6 +56,7 @@ var migrations = []struct {
|
||||
{6, proberHeartbeatSchema},
|
||||
{7, ipRegistrySchema},
|
||||
{8, autoCycleSchema},
|
||||
{9, scaleIndexesSchema},
|
||||
}
|
||||
|
||||
type DB struct {
|
||||
|
||||
@@ -0,0 +1,11 @@
|
||||
-- Scale indexes (see docs/changes: FIP scan at scale).
|
||||
--
|
||||
-- idx_ip_queue_registry: ListRegistry/ListRegistryPage look up the live
|
||||
-- ip_queue row of every registry address by registry_id; without an index
|
||||
-- that was a full scan per address, i.e. O(n^2) with thousands of addresses.
|
||||
--
|
||||
-- idx_ip_queue_state_aggregated: paged "recently finished" listings
|
||||
-- (state filter, newest aggregated_at first).
|
||||
|
||||
CREATE INDEX idx_ip_queue_registry ON ip_queue(registry_id);
|
||||
CREATE INDEX idx_ip_queue_state_aggregated ON ip_queue(state, aggregated_at);
|
||||
@@ -39,6 +39,32 @@ const (
|
||||
SourceEgress = "egress"
|
||||
)
|
||||
|
||||
// IPStates lists every valid ip_queue state, in lifecycle order.
|
||||
var IPStates = []string{
|
||||
IPQueued, IPAssigningFIP, IPAwaitingSelfCheck, IPChecking, IPAggregating,
|
||||
IPDone, IPFailed, IPOccupied,
|
||||
}
|
||||
|
||||
// IsValidIPState reports whether s is one of IPStates.
|
||||
func IsValidIPState(s string) bool {
|
||||
for _, v := range IPStates {
|
||||
if v == s {
|
||||
return true
|
||||
}
|
||||
}
|
||||
return false
|
||||
}
|
||||
|
||||
// IsValidResult reports whether s is a valid overall result value
|
||||
// (pass | partial | fail | cancelled).
|
||||
func IsValidResult(s string) bool {
|
||||
switch s {
|
||||
case ResultPass, ResultPartial, ResultFail, ResultCancelled:
|
||||
return true
|
||||
}
|
||||
return false
|
||||
}
|
||||
|
||||
// InboundSource returns the checks.source value for the given prober site
|
||||
// index (1-based), e.g. InboundSource(1) == "inbound-site-1".
|
||||
func InboundSource(siteIndex int) string {
|
||||
@@ -232,6 +258,9 @@ const (
|
||||
AutoCyclePhaseIdle = "idle"
|
||||
AutoCyclePhaseRunning = "running"
|
||||
AutoCyclePhaseWaiting = "waiting"
|
||||
// AutoCyclePhaseScanning: the queue is being cleared and the floating IPs
|
||||
// are being (re)discovered and enqueued by the background scan job.
|
||||
AutoCyclePhaseScanning = "scanning"
|
||||
|
||||
AutoCycleOutcomeCompleted = "completed"
|
||||
AutoCycleOutcomeNoFreeIPs = "no_free_ips"
|
||||
|
||||
@@ -102,7 +102,7 @@ func (d *DB) SetAutoCycleEnabled(ctx context.Context, enabled bool, nextRunAt *t
|
||||
// touch the configuration or the enabled flag.
|
||||
func (d *DB) UpdateAutoCycleState(ctx context.Context, s AutoCycleState) error {
|
||||
switch s.Phase {
|
||||
case AutoCyclePhaseIdle, AutoCyclePhaseRunning, AutoCyclePhaseWaiting:
|
||||
case AutoCyclePhaseIdle, AutoCyclePhaseScanning, AutoCyclePhaseRunning, AutoCyclePhaseWaiting:
|
||||
default:
|
||||
return fmt.Errorf("invalid auto-cycle phase %q: %w", s.Phase, ErrValidation)
|
||||
}
|
||||
|
||||
@@ -4,6 +4,7 @@ import (
|
||||
"context"
|
||||
"database/sql"
|
||||
"fmt"
|
||||
"strings"
|
||||
"time"
|
||||
)
|
||||
|
||||
@@ -497,6 +498,263 @@ func deleteIPTx(ctx context.Context, tx *sql.Tx, ipID int64) error {
|
||||
return nil
|
||||
}
|
||||
|
||||
// ClearAllIPs deletes every ip_queue row in one short, set-based
|
||||
// transaction (five statements regardless of the queue size) and returns the
|
||||
// deleted addresses in queue order. It does exactly what deleteIPTx does per
|
||||
// row: frees every validator that still points at a doomed row, detaches
|
||||
// checks/events (their registry_id keeps the history), drops the ephemeral
|
||||
// ip_site_checks progress flags and finally the rows themselves.
|
||||
// Disassociating attached floating IPs is the caller's (orchestrator's) job.
|
||||
func (d *DB) ClearAllIPs(ctx context.Context) ([]string, error) {
|
||||
tx, err := d.BeginTx(ctx, nil)
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
defer tx.Rollback()
|
||||
|
||||
rows, err := tx.QueryContext(ctx, `SELECT ip_address FROM ip_queue ORDER BY sequence`)
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
deleted := []string{}
|
||||
for rows.Next() {
|
||||
var addr string
|
||||
if err := rows.Scan(&addr); err != nil {
|
||||
rows.Close()
|
||||
return nil, err
|
||||
}
|
||||
deleted = append(deleted, addr)
|
||||
}
|
||||
if err := rows.Err(); err != nil {
|
||||
rows.Close()
|
||||
return nil, err
|
||||
}
|
||||
rows.Close()
|
||||
|
||||
now := timeToDB(Now())
|
||||
if _, err := tx.ExecContext(ctx, `
|
||||
UPDATE validators SET state=?, current_ip_id=NULL, updated_at=?
|
||||
WHERE current_ip_id IS NOT NULL
|
||||
`, ValidatorIdle, now); err != nil {
|
||||
return nil, fmt.Errorf("free owning validators: %w", err)
|
||||
}
|
||||
if _, err := tx.ExecContext(ctx, `UPDATE checks SET ip_id=NULL WHERE ip_id IS NOT NULL`); err != nil {
|
||||
return nil, fmt.Errorf("detach checks: %w", err)
|
||||
}
|
||||
if _, err := tx.ExecContext(ctx, `UPDATE events SET ip_id=NULL WHERE ip_id IS NOT NULL`); err != nil {
|
||||
return nil, fmt.Errorf("detach events: %w", err)
|
||||
}
|
||||
if _, err := tx.ExecContext(ctx, `DELETE FROM ip_site_checks`); err != nil {
|
||||
return nil, fmt.Errorf("delete ip_site_checks: %w", err)
|
||||
}
|
||||
if _, err := tx.ExecContext(ctx, `DELETE FROM ip_queue`); err != nil {
|
||||
return nil, fmt.Errorf("delete ip_queue rows: %w", err)
|
||||
}
|
||||
if err := tx.Commit(); err != nil {
|
||||
return nil, err
|
||||
}
|
||||
return deleted, nil
|
||||
}
|
||||
|
||||
// FIPRef identifies a queue row that currently holds a Neutron floating-IP
|
||||
// association.
|
||||
type FIPRef struct {
|
||||
IPID int64
|
||||
IPAddress string
|
||||
FIPID string
|
||||
}
|
||||
|
||||
// ListFIPRefs returns every queue row with an attached floating IP (fip_id
|
||||
// set) — typically at most one per validator — so a bulk clear can
|
||||
// disassociate them without loading the whole queue.
|
||||
func (d *DB) ListFIPRefs(ctx context.Context) ([]FIPRef, error) {
|
||||
rows, err := d.QueryContext(ctx, `SELECT id, ip_address, fip_id FROM ip_queue WHERE fip_id<>'' ORDER BY id`)
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
defer rows.Close()
|
||||
var out []FIPRef
|
||||
for rows.Next() {
|
||||
var r FIPRef
|
||||
if err := rows.Scan(&r.IPID, &r.IPAddress, &r.FIPID); err != nil {
|
||||
return nil, err
|
||||
}
|
||||
out = append(out, r)
|
||||
}
|
||||
return out, rows.Err()
|
||||
}
|
||||
|
||||
// ListFIPRefsByAddresses is ListFIPRefs restricted to the given addresses
|
||||
// (unknown addresses and rows without an attached floating IP are simply
|
||||
// absent), using a handful of IN (...) queries instead of one lookup per
|
||||
// address.
|
||||
func (d *DB) ListFIPRefsByAddresses(ctx context.Context, addresses []string) ([]FIPRef, error) {
|
||||
const chunk = 500
|
||||
var out []FIPRef
|
||||
for start := 0; start < len(addresses); start += chunk {
|
||||
end := start + chunk
|
||||
if end > len(addresses) {
|
||||
end = len(addresses)
|
||||
}
|
||||
part := addresses[start:end]
|
||||
args := make([]any, len(part))
|
||||
for i, a := range part {
|
||||
args[i] = a
|
||||
}
|
||||
rows, err := d.QueryContext(ctx,
|
||||
`SELECT id, ip_address, fip_id FROM ip_queue WHERE fip_id<>'' AND ip_address IN (`+placeholders(len(part))+`)`, args...)
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
for rows.Next() {
|
||||
var r FIPRef
|
||||
if err := rows.Scan(&r.IPID, &r.IPAddress, &r.FIPID); err != nil {
|
||||
rows.Close()
|
||||
return nil, err
|
||||
}
|
||||
out = append(out, r)
|
||||
}
|
||||
if err := rows.Err(); err != nil {
|
||||
rows.Close()
|
||||
return nil, err
|
||||
}
|
||||
rows.Close()
|
||||
}
|
||||
return out, nil
|
||||
}
|
||||
|
||||
func placeholders(n int) string {
|
||||
if n <= 0 {
|
||||
return ""
|
||||
}
|
||||
return strings.TrimSuffix(strings.Repeat("?,", n), ",")
|
||||
}
|
||||
|
||||
// CountIPsByState returns how many ip_queue rows are in each state, plus the
|
||||
// grand total, via a single GROUP BY (no row loading).
|
||||
func (d *DB) CountIPsByState(ctx context.Context) (map[string]int, int, error) {
|
||||
rows, err := d.QueryContext(ctx, `SELECT state, COUNT(*) FROM ip_queue GROUP BY state`)
|
||||
if err != nil {
|
||||
return nil, 0, err
|
||||
}
|
||||
defer rows.Close()
|
||||
counts := map[string]int{}
|
||||
total := 0
|
||||
for rows.Next() {
|
||||
var state string
|
||||
var n int
|
||||
if err := rows.Scan(&state, &n); err != nil {
|
||||
return nil, 0, err
|
||||
}
|
||||
counts[state] = n
|
||||
total += n
|
||||
}
|
||||
return counts, total, rows.Err()
|
||||
}
|
||||
|
||||
// CountIPsByResult returns how many ip_queue rows carry each non-empty
|
||||
// overall_result (pass/partial/fail/cancelled), via a single GROUP BY.
|
||||
func (d *DB) CountIPsByResult(ctx context.Context) (map[string]int, error) {
|
||||
rows, err := d.QueryContext(ctx, `
|
||||
SELECT overall_result, COUNT(*) FROM ip_queue WHERE overall_result<>'' GROUP BY overall_result
|
||||
`)
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
defer rows.Close()
|
||||
counts := map[string]int{}
|
||||
for rows.Next() {
|
||||
var res string
|
||||
var n int
|
||||
if err := rows.Scan(&res, &n); err != nil {
|
||||
return nil, err
|
||||
}
|
||||
counts[res] = n
|
||||
}
|
||||
return counts, rows.Err()
|
||||
}
|
||||
|
||||
// AnyNonTerminalIP reports whether any ip_queue row is still unfinished (its
|
||||
// state is none of done/failed/occupied).
|
||||
func (d *DB) AnyNonTerminalIP(ctx context.Context) (bool, error) {
|
||||
var any bool
|
||||
err := d.QueryRowContext(ctx, `SELECT EXISTS(SELECT 1 FROM ip_queue WHERE state NOT IN (?, ?, ?))`,
|
||||
IPDone, IPFailed, IPOccupied).Scan(&any)
|
||||
return any, err
|
||||
}
|
||||
|
||||
// Order values for IPFilter.Order.
|
||||
const (
|
||||
IPOrderSequence = "sequence"
|
||||
IPOrderAggregatedAtDesc = "aggregated_at_desc"
|
||||
)
|
||||
|
||||
// IPFilter narrows ListIPsPage. The zero value matches everything, ordered by
|
||||
// queue sequence.
|
||||
type IPFilter struct {
|
||||
States []string // any of these states (empty = all)
|
||||
Query string // substring of ip_address
|
||||
Result string // overall_result equals this (pass|partial|fail|cancelled)
|
||||
Order string // IPOrderSequence (default) | IPOrderAggregatedAtDesc
|
||||
}
|
||||
|
||||
// ListIPsPage returns one page (limit/offset) of ip_queue rows matching f,
|
||||
// plus the total number of matching rows. limit <= 0 means no limit.
|
||||
func (d *DB) ListIPsPage(ctx context.Context, f IPFilter, limit, offset int) ([]IPQueueItem, int, error) {
|
||||
var conds []string
|
||||
var args []any
|
||||
if len(f.States) > 0 {
|
||||
conds = append(conds, "state IN ("+placeholders(len(f.States))+")")
|
||||
for _, s := range f.States {
|
||||
args = append(args, s)
|
||||
}
|
||||
}
|
||||
if f.Query != "" {
|
||||
conds = append(conds, "instr(ip_address, ?) > 0")
|
||||
args = append(args, f.Query)
|
||||
}
|
||||
if f.Result != "" {
|
||||
conds = append(conds, "overall_result = ?")
|
||||
args = append(args, f.Result)
|
||||
}
|
||||
where := ""
|
||||
if len(conds) > 0 {
|
||||
where = "WHERE " + strings.Join(conds, " AND ") + " "
|
||||
}
|
||||
|
||||
var total int
|
||||
if err := d.QueryRowContext(ctx, `SELECT COUNT(*) FROM ip_queue `+where, args...).Scan(&total); err != nil {
|
||||
return nil, 0, err
|
||||
}
|
||||
|
||||
order := "ORDER BY sequence"
|
||||
if f.Order == IPOrderAggregatedAtDesc {
|
||||
// Timestamps are stored as RFC3339Nano, which trims trailing zeros
|
||||
// and so does not sort correctly as text; strftime normalizes them to
|
||||
// a fixed millisecond layout. NULL aggregated_at sorts last in DESC.
|
||||
order = "ORDER BY strftime('%Y-%m-%d %H:%M:%f', aggregated_at) DESC, sequence"
|
||||
}
|
||||
q := ipQueueSelect + where + order
|
||||
qargs := append([]any(nil), args...)
|
||||
if limit > 0 {
|
||||
q += " LIMIT ? OFFSET ?"
|
||||
qargs = append(qargs, limit, offset)
|
||||
}
|
||||
rows, err := d.QueryContext(ctx, q, qargs...)
|
||||
if err != nil {
|
||||
return nil, 0, err
|
||||
}
|
||||
defer rows.Close()
|
||||
items, err := scanIPQueueItems(rows)
|
||||
if err != nil {
|
||||
return nil, 0, err
|
||||
}
|
||||
if items == nil {
|
||||
items = []IPQueueItem{}
|
||||
}
|
||||
return items, total, nil
|
||||
}
|
||||
|
||||
func (d *DB) SetEgressComplete(ctx context.Context, ipID int64) error {
|
||||
_, err := d.ExecContext(ctx, `UPDATE ip_queue SET egress_complete=1, updated_at=? WHERE id=?`, timeToDB(Now()), ipID)
|
||||
return err
|
||||
|
||||
@@ -4,6 +4,7 @@ import (
|
||||
"context"
|
||||
"database/sql"
|
||||
"fmt"
|
||||
"strings"
|
||||
"time"
|
||||
)
|
||||
|
||||
@@ -66,7 +67,7 @@ type RegistrySummary struct {
|
||||
func (d *DB) ListRegistry(ctx context.Context) ([]RegistrySummary, error) {
|
||||
rows, err := d.QueryContext(ctx, `
|
||||
SELECT id, ip_address, first_seen_at, last_seen_at, next_cycle, created_at, updated_at
|
||||
FROM ip_registry ORDER BY first_seen_at
|
||||
FROM ip_registry ORDER BY first_seen_at, id
|
||||
`)
|
||||
if err != nil {
|
||||
return nil, err
|
||||
@@ -89,6 +90,84 @@ func (d *DB) ListRegistry(ctx context.Context) ([]RegistrySummary, error) {
|
||||
return out, nil
|
||||
}
|
||||
|
||||
// RegistryFilter narrows ListRegistryPage. The zero value matches everything.
|
||||
type RegistryFilter struct {
|
||||
Query string // substring of ip_address
|
||||
LastResult string // pass|partial|fail|cancelled — same meaning as RegistrySummary.LastResult
|
||||
}
|
||||
|
||||
// lastResultCond is the SQL form of fillRegistrySummary's LastResult rule,
|
||||
// over `ip_registry r LEFT JOIN ip_queue q ON q.registry_id=r.id`: a live
|
||||
// queue row's overall_result is authoritative (and a live row without one has
|
||||
// no verdict); with no live row the latest recorded cycle is classified from
|
||||
// its checks (pass if all succeeded, fail if none, partial otherwise).
|
||||
const lastResultCond = `(
|
||||
(q.id IS NOT NULL AND q.overall_result = ?)
|
||||
OR (q.id IS NULL AND (
|
||||
SELECT CASE WHEN SUM(c.success) = 0 THEN 'fail'
|
||||
WHEN SUM(c.success) = COUNT(*) THEN 'pass'
|
||||
ELSE 'partial' END
|
||||
FROM checks c
|
||||
WHERE c.registry_id = r.id
|
||||
AND c.cycle_id = (SELECT MAX(c2.cycle_id) FROM checks c2 WHERE c2.registry_id = r.id)
|
||||
HAVING COUNT(*) > 0
|
||||
) = ?)
|
||||
)`
|
||||
|
||||
// ListRegistryPage returns one page (limit/offset) of the registry in the same
|
||||
// order as ListRegistry, filtered by f, plus the total number of matching
|
||||
// rows. LIMIT/OFFSET are applied in SQL before the per-row summary queries,
|
||||
// so only the rows of the page pay for them. limit <= 0 means no limit.
|
||||
func (d *DB) ListRegistryPage(ctx context.Context, f RegistryFilter, limit, offset int) ([]RegistrySummary, int, error) {
|
||||
var conds []string
|
||||
var args []any
|
||||
if f.Query != "" {
|
||||
conds = append(conds, "instr(r.ip_address, ?) > 0")
|
||||
args = append(args, f.Query)
|
||||
}
|
||||
if f.LastResult != "" {
|
||||
conds = append(conds, lastResultCond)
|
||||
args = append(args, f.LastResult, f.LastResult)
|
||||
}
|
||||
from := ` FROM ip_registry r LEFT JOIN ip_queue q ON q.registry_id = r.id `
|
||||
where := ""
|
||||
if len(conds) > 0 {
|
||||
where = "WHERE " + strings.Join(conds, " AND ") + " "
|
||||
}
|
||||
|
||||
var total int
|
||||
if err := d.QueryRowContext(ctx, `SELECT COUNT(*)`+from+where, args...).Scan(&total); err != nil {
|
||||
return nil, 0, err
|
||||
}
|
||||
|
||||
q := `SELECT r.id, r.ip_address, r.first_seen_at, r.last_seen_at, r.next_cycle, r.created_at, r.updated_at` +
|
||||
from + where + `ORDER BY r.first_seen_at, r.id`
|
||||
qargs := append([]any(nil), args...)
|
||||
if limit > 0 {
|
||||
q += " LIMIT ? OFFSET ?"
|
||||
qargs = append(qargs, limit, offset)
|
||||
}
|
||||
rows, err := d.QueryContext(ctx, q, qargs...)
|
||||
if err != nil {
|
||||
return nil, 0, err
|
||||
}
|
||||
items, err := scanRegistryItems(rows)
|
||||
rows.Close()
|
||||
if err != nil {
|
||||
return nil, 0, err
|
||||
}
|
||||
|
||||
out := make([]RegistrySummary, len(items))
|
||||
for i, item := range items {
|
||||
s := RegistrySummary{RegistryItem: item}
|
||||
if err := d.fillRegistrySummary(ctx, &s); err != nil {
|
||||
return nil, 0, err
|
||||
}
|
||||
out[i] = s
|
||||
}
|
||||
return out, total, nil
|
||||
}
|
||||
|
||||
// GetRegistryByAddress returns the registry row (with summary) for a single
|
||||
// address, or ErrNotFound if it has never been submitted.
|
||||
func (d *DB) GetRegistryByAddress(ctx context.Context, address string) (*RegistrySummary, error) {
|
||||
|
||||
@@ -0,0 +1,405 @@
|
||||
package db
|
||||
|
||||
import (
|
||||
"context"
|
||||
"fmt"
|
||||
"reflect"
|
||||
"sort"
|
||||
"testing"
|
||||
"time"
|
||||
)
|
||||
|
||||
// scaleAddrs returns n distinct addresses 10.<a>.<b>.<c> in ascending order.
|
||||
func scaleAddrs(n int) []string {
|
||||
out := make([]string, n)
|
||||
for i := 0; i < n; i++ {
|
||||
out[i] = fmt.Sprintf("10.%d.%d.%d", (i/65536)%256, (i/256)%256, i%256)
|
||||
}
|
||||
return out
|
||||
}
|
||||
|
||||
// finishWithChecks submits the address, records ok successes and bad
|
||||
// failures as the current cycle's checks and, when finish != "", finishes the
|
||||
// row with that overall result.
|
||||
func finishWithChecks(t *testing.T, d *DB, addr string, ok, bad int, finish string) {
|
||||
t.Helper()
|
||||
ctx := context.Background()
|
||||
ip, err := d.GetIPByAddress(ctx, addr)
|
||||
if err != nil {
|
||||
t.Fatalf("get %s: %v", addr, err)
|
||||
}
|
||||
for i := 0; i < ok+bad; i++ {
|
||||
if err := d.UpsertCheck(ctx, Check{
|
||||
IPID: ip.ID, IPAddress: addr, AttemptNumber: ip.AttemptNumber,
|
||||
Source: SourceEgress, CheckType: "https", Target: fmt.Sprintf("https://t%d.test", i),
|
||||
Success: i < ok, CheckedAt: Now(),
|
||||
}); err != nil {
|
||||
t.Fatalf("upsert check: %v", err)
|
||||
}
|
||||
}
|
||||
if finish != "" {
|
||||
if err := d.FinishIP(ctx, ip.ID, finish); err != nil {
|
||||
t.Fatalf("finish: %v", err)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
func TestListIPsPageFiltersTotalOrder(t *testing.T) {
|
||||
d, ctx := newTestDB(t)
|
||||
addrs := []string{"10.0.0.1", "10.0.0.2", "10.0.0.3", "10.0.1.1", "10.0.1.2", "192.168.0.10"}
|
||||
if _, err := d.SubmitIPs(ctx, addrs); err != nil {
|
||||
t.Fatalf("submit: %v", err)
|
||||
}
|
||||
// 10.0.0.1 pass, 10.0.0.2 fail, 10.0.0.3 checking, 10.0.1.1 occupied.
|
||||
finishWithChecks(t, d, "10.0.0.1", 1, 0, ResultPass)
|
||||
time.Sleep(3 * time.Millisecond)
|
||||
finishWithChecks(t, d, "10.0.0.2", 0, 1, ResultFail)
|
||||
ip3, _ := d.GetIPByAddress(ctx, "10.0.0.3")
|
||||
if err := d.SetChecking(ctx, ip3.ID, time.Minute); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
ip4, _ := d.GetIPByAddress(ctx, "10.0.1.1")
|
||||
if err := d.MarkFIPOccupied(ctx, ip4.ID, ""); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
|
||||
addrsOf := func(items []IPQueueItem) []string {
|
||||
out := []string{}
|
||||
for _, it := range items {
|
||||
out = append(out, it.IPAddress)
|
||||
}
|
||||
return out
|
||||
}
|
||||
|
||||
items, total, err := d.ListIPsPage(ctx, IPFilter{}, 4, 0)
|
||||
if err != nil || total != 6 || !reflect.DeepEqual(addrsOf(items), addrs[:4]) {
|
||||
t.Fatalf("page 1: total=%d items=%v err=%v", total, addrsOf(items), err)
|
||||
}
|
||||
items, total, _ = d.ListIPsPage(ctx, IPFilter{}, 4, 4)
|
||||
if total != 6 || !reflect.DeepEqual(addrsOf(items), addrs[4:]) {
|
||||
t.Fatalf("page 2: total=%d items=%v", total, addrsOf(items))
|
||||
}
|
||||
items, total, _ = d.ListIPsPage(ctx, IPFilter{}, 4, 100)
|
||||
if total != 6 || len(items) != 0 || items == nil {
|
||||
t.Fatalf("offset past end: total=%d items=%v", total, items)
|
||||
}
|
||||
|
||||
items, total, _ = d.ListIPsPage(ctx, IPFilter{States: []string{IPDone, IPFailed}}, 50, 0)
|
||||
if total != 2 || !reflect.DeepEqual(addrsOf(items), []string{"10.0.0.1", "10.0.0.2"}) {
|
||||
t.Fatalf("states filter: total=%d items=%v", total, addrsOf(items))
|
||||
}
|
||||
items, total, _ = d.ListIPsPage(ctx, IPFilter{States: []string{IPQueued}, Query: "10.0.1."}, 50, 0)
|
||||
if total != 1 || addrsOf(items)[0] != "10.0.1.2" {
|
||||
t.Fatalf("state+q filter: total=%d items=%v", total, addrsOf(items))
|
||||
}
|
||||
items, total, _ = d.ListIPsPage(ctx, IPFilter{Result: ResultFail}, 50, 0)
|
||||
if total != 1 || addrsOf(items)[0] != "10.0.0.2" {
|
||||
t.Fatalf("result filter: total=%d items=%v", total, addrsOf(items))
|
||||
}
|
||||
// q is a plain substring, not a LIKE pattern: % and _ match literally.
|
||||
if _, total, _ = d.ListIPsPage(ctx, IPFilter{Query: "%"}, 50, 0); total != 0 {
|
||||
t.Fatalf("expected literal substring match, total=%d", total)
|
||||
}
|
||||
|
||||
// Newest aggregated first; never-aggregated rows last.
|
||||
items, _, _ = d.ListIPsPage(ctx, IPFilter{Order: IPOrderAggregatedAtDesc}, 50, 0)
|
||||
got := addrsOf(items)
|
||||
if got[0] != "10.0.1.1" && got[0] != "10.0.0.2" {
|
||||
t.Fatalf("expected a finished row first, got %v", got)
|
||||
}
|
||||
if items[0].AggregatedAt == nil || items[len(items)-1].AggregatedAt != nil {
|
||||
t.Fatalf("expected aggregated rows first and unaggregated last: %v", got)
|
||||
}
|
||||
for i := 1; i < len(items); i++ {
|
||||
a, b := items[i-1].AggregatedAt, items[i].AggregatedAt
|
||||
if a != nil && b != nil && a.Before(*b) {
|
||||
t.Fatalf("not sorted by aggregated_at desc at %d: %v", i, got)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
func TestCountAndNonTerminal(t *testing.T) {
|
||||
d, ctx := newTestDB(t)
|
||||
|
||||
byState, total, err := d.CountIPsByState(ctx)
|
||||
if err != nil || total != 0 || len(byState) != 0 {
|
||||
t.Fatalf("empty: %v %d %v", byState, total, err)
|
||||
}
|
||||
if any, err := d.AnyNonTerminalIP(ctx); err != nil || any {
|
||||
t.Fatalf("empty queue must not report non-terminal: %v %v", any, err)
|
||||
}
|
||||
|
||||
if _, err := d.SubmitIPs(ctx, []string{"1.1.1.1", "1.1.1.2", "1.1.1.3", "1.1.1.4"}); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
finishWithChecks(t, d, "1.1.1.1", 1, 0, ResultPass)
|
||||
finishWithChecks(t, d, "1.1.1.2", 1, 1, ResultPartial)
|
||||
ip3, _ := d.GetIPByAddress(ctx, "1.1.1.3")
|
||||
if err := d.CancelIP(ctx, ip3.ID); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
|
||||
byState, total, err = d.CountIPsByState(ctx)
|
||||
if err != nil || total != 4 || byState[IPDone] != 2 || byState[IPFailed] != 1 || byState[IPQueued] != 1 {
|
||||
t.Fatalf("by state: %v total=%d err=%v", byState, total, err)
|
||||
}
|
||||
byResult, err := d.CountIPsByResult(ctx)
|
||||
if err != nil || len(byResult) != 3 || byResult[ResultPass] != 1 || byResult[ResultPartial] != 1 || byResult[ResultCancelled] != 1 {
|
||||
t.Fatalf("by result: %v err=%v", byResult, err)
|
||||
}
|
||||
if any, _ := d.AnyNonTerminalIP(ctx); !any {
|
||||
t.Fatalf("a queued row is non-terminal")
|
||||
}
|
||||
ip4, _ := d.GetIPByAddress(ctx, "1.1.1.4")
|
||||
if err := d.MarkFIPOccupied(ctx, ip4.ID, ""); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if any, _ := d.AnyNonTerminalIP(ctx); any {
|
||||
t.Fatalf("done/failed/occupied only: expected terminal")
|
||||
}
|
||||
}
|
||||
|
||||
// TestListRegistryPageMatchesListRegistry builds a mixed dataset (live rows
|
||||
// with every overall result, an in-progress live row, deleted rows whose
|
||||
// last cycle classifies as pass/partial/fail, and an address without checks)
|
||||
// and verifies that ListRegistryPage's SQL filter selects exactly the rows
|
||||
// ListRegistry+fillRegistrySummary labels with the same LastResult.
|
||||
func TestListRegistryPageMatchesListRegistry(t *testing.T) {
|
||||
d, ctx := newTestDB(t)
|
||||
addrs := scaleAddrs(14)
|
||||
if _, err := d.SubmitIPs(ctx, addrs); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
// live rows with an aggregated result
|
||||
finishWithChecks(t, d, addrs[0], 2, 0, ResultPass)
|
||||
finishWithChecks(t, d, addrs[1], 1, 1, ResultPartial)
|
||||
finishWithChecks(t, d, addrs[2], 0, 2, ResultFail)
|
||||
ip3, _ := d.GetIPByAddress(ctx, addrs[3])
|
||||
if err := d.CancelIP(ctx, ip3.ID); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
// live row, passing checks recorded, but cycle unfinished -> no verdict
|
||||
finishWithChecks(t, d, addrs[4], 2, 0, "")
|
||||
// rows to be deleted: result derived from the checks of the last cycle
|
||||
finishWithChecks(t, d, addrs[5], 2, 0, ResultPass) // -> pass
|
||||
finishWithChecks(t, d, addrs[6], 1, 2, ResultPartial) // -> partial
|
||||
finishWithChecks(t, d, addrs[7], 0, 2, ResultFail) // -> fail
|
||||
finishWithChecks(t, d, addrs[8], 0, 0, "") // deleted, no checks -> ""
|
||||
// a deleted row whose first cycle failed but whose last passed
|
||||
finishWithChecks(t, d, addrs[9], 0, 1, ResultFail)
|
||||
if _, err := d.DeleteIPs(ctx, []string{addrs[5], addrs[6], addrs[7], addrs[8], addrs[9]}); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if _, err := d.SubmitIPs(ctx, []string{addrs[9]}); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
finishWithChecks(t, d, addrs[9], 1, 0, ResultPass)
|
||||
if _, err := d.DeleteIPs(ctx, []string{addrs[9]}); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
|
||||
all, err := d.ListRegistry(ctx)
|
||||
if err != nil || len(all) != 14 {
|
||||
t.Fatalf("list registry: n=%d err=%v", len(all), err)
|
||||
}
|
||||
|
||||
// No filter: same rows in the same order as ListRegistry, paged.
|
||||
var paged []RegistrySummary
|
||||
for off := 0; ; off += 5 {
|
||||
page, total, err := d.ListRegistryPage(ctx, RegistryFilter{}, 5, off)
|
||||
if err != nil || total != 14 {
|
||||
t.Fatalf("page off=%d total=%d err=%v", off, total, err)
|
||||
}
|
||||
if len(page) == 0 {
|
||||
break
|
||||
}
|
||||
paged = append(paged, page...)
|
||||
}
|
||||
if len(paged) != len(all) {
|
||||
t.Fatalf("paged %d rows, want %d", len(paged), len(all))
|
||||
}
|
||||
for i := range all {
|
||||
if all[i].IPAddress != paged[i].IPAddress || all[i].LastResult != paged[i].LastResult {
|
||||
t.Fatalf("row %d differs: %+v vs %+v", i, all[i], paged[i])
|
||||
}
|
||||
}
|
||||
|
||||
for _, res := range []string{ResultPass, ResultPartial, ResultFail, ResultCancelled} {
|
||||
var want []string
|
||||
for _, s := range all {
|
||||
if s.LastResult == res {
|
||||
want = append(want, s.IPAddress)
|
||||
}
|
||||
}
|
||||
page, total, err := d.ListRegistryPage(ctx, RegistryFilter{LastResult: res}, 100, 0)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
var got []string
|
||||
for _, s := range page {
|
||||
got = append(got, s.IPAddress)
|
||||
if s.LastResult != res {
|
||||
t.Fatalf("%s: row %s has LastResult %q", res, s.IPAddress, s.LastResult)
|
||||
}
|
||||
}
|
||||
sort.Strings(want)
|
||||
sort.Strings(got)
|
||||
if total != len(want) || !reflect.DeepEqual(got, want) || len(want) == 0 {
|
||||
t.Fatalf("last_result=%s: total=%d got=%v want=%v", res, total, got, want)
|
||||
}
|
||||
}
|
||||
|
||||
// pass: addrs[0] (live) + addrs[5] + addrs[9] (deleted, latest cycle).
|
||||
if _, total, _ := d.ListRegistryPage(ctx, RegistryFilter{LastResult: ResultPass}, 100, 0); total != 3 {
|
||||
t.Fatalf("expected 3 pass rows, got %d", total)
|
||||
}
|
||||
|
||||
// q filter + LIMIT applies after the filter, total is the filtered count.
|
||||
page, total, err := d.ListRegistryPage(ctx, RegistryFilter{Query: "10.0.0.1"}, 2, 0)
|
||||
if err != nil || total != 5 || len(page) != 2 { // 10.0.0.1, .10-.13
|
||||
t.Fatalf("q filter: total=%d len=%d err=%v", total, len(page), err)
|
||||
}
|
||||
page, total, _ = d.ListRegistryPage(ctx, RegistryFilter{Query: "10.0.0.1", LastResult: ResultPass}, 10, 0)
|
||||
if total != 0 || len(page) != 0 {
|
||||
// 10.0.0.1 is partial; 10.0.0.10-13 have no verdict or fail.
|
||||
t.Fatalf("q+last_result: total=%d page=%+v", total, page)
|
||||
}
|
||||
}
|
||||
|
||||
func TestClearAllIPsKeepsHistoryAndFreesValidators(t *testing.T) {
|
||||
d, ctx := newTestDB(t)
|
||||
if err := d.AdminCreateValidator(ctx, "validator-1", "port-1"); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
addrs := []string{"5.5.5.1", "5.5.5.2", "5.5.5.3"}
|
||||
if _, err := d.SubmitIPs(ctx, addrs); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
finishWithChecks(t, d, "5.5.5.2", 1, 1, ResultPartial)
|
||||
claimed, err := d.ClaimNextQueued(ctx, "validator-1", time.Minute)
|
||||
if err != nil || claimed == nil {
|
||||
t.Fatalf("claim: %v %v", claimed, err)
|
||||
}
|
||||
if err := d.SetFIPAssociated(ctx, claimed.ID, "fip-9", time.Minute); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if err := d.SetEgressComplete(ctx, claimed.ID); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if err := d.SetSiteComplete(ctx, claimed.ID, 1); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
|
||||
refs, err := d.ListFIPRefs(ctx)
|
||||
if err != nil || len(refs) != 1 || refs[0].FIPID != "fip-9" || refs[0].IPAddress != claimed.IPAddress {
|
||||
t.Fatalf("fip refs: %+v err=%v", refs, err)
|
||||
}
|
||||
refs, err = d.ListFIPRefsByAddresses(ctx, []string{claimed.IPAddress, "nope"})
|
||||
if err != nil || len(refs) != 1 {
|
||||
t.Fatalf("fip refs by address: %+v err=%v", refs, err)
|
||||
}
|
||||
if refs, _ := d.ListFIPRefsByAddresses(ctx, []string{"5.5.5.3"}); len(refs) != 0 {
|
||||
t.Fatalf("row without fip must not be listed: %+v", refs)
|
||||
}
|
||||
|
||||
deleted, err := d.ClearAllIPs(ctx)
|
||||
if err != nil {
|
||||
t.Fatalf("clear: %v", err)
|
||||
}
|
||||
sort.Strings(deleted)
|
||||
if !reflect.DeepEqual(deleted, addrs) {
|
||||
t.Fatalf("deleted %v, want %v", deleted, addrs)
|
||||
}
|
||||
if items, _ := d.ListIPs(ctx); len(items) != 0 {
|
||||
t.Fatalf("queue not empty: %+v", items)
|
||||
}
|
||||
v, err := d.GetValidator(ctx, "validator-1")
|
||||
if err != nil || v.State != ValidatorIdle || v.CurrentIPID != nil {
|
||||
t.Fatalf("validator not freed: %+v err=%v", v, err)
|
||||
}
|
||||
// History and registry rows survive, detached from the queue.
|
||||
reg, err := d.GetRegistryByAddress(ctx, "5.5.5.2")
|
||||
if err != nil || reg.TotalCycles != 1 || reg.LastResult != ResultPartial || reg.InQueue {
|
||||
t.Fatalf("registry after clear: %+v err=%v", reg, err)
|
||||
}
|
||||
checks, err := d.ListChecksForRegistry(ctx, reg.ID, nil)
|
||||
if err != nil || len(checks) != 2 || checks[0].IPID != 0 {
|
||||
t.Fatalf("checks after clear: %+v err=%v", checks, err)
|
||||
}
|
||||
// Clearing an empty queue is fine and returns an empty (non-nil) list.
|
||||
if deleted, err := d.ClearAllIPs(ctx); err != nil || deleted == nil || len(deleted) != 0 {
|
||||
t.Fatalf("second clear: %v %v", deleted, err)
|
||||
}
|
||||
// The same addresses can be re-submitted afterwards.
|
||||
if res, err := d.SubmitIPs(ctx, addrs); err != nil || len(res.Added) != 3 {
|
||||
t.Fatalf("resubmit: %+v err=%v", res, err)
|
||||
}
|
||||
}
|
||||
|
||||
func TestMigration0009Indexes(t *testing.T) {
|
||||
d, ctx := newTestDB(t)
|
||||
for _, name := range []string{"idx_ip_queue_registry", "idx_ip_queue_state_aggregated"} {
|
||||
var n int
|
||||
if err := d.QueryRowContext(ctx, `SELECT COUNT(*) FROM sqlite_master WHERE type='index' AND name=?`, name).Scan(&n); err != nil || n != 1 {
|
||||
t.Fatalf("index %s missing (n=%d err=%v)", name, n, err)
|
||||
}
|
||||
}
|
||||
var ver int
|
||||
if err := d.QueryRowContext(ctx, `PRAGMA user_version`).Scan(&ver); err != nil || ver < 9 {
|
||||
t.Fatalf("user_version=%d err=%v", ver, err)
|
||||
}
|
||||
}
|
||||
|
||||
// TestScaleSmoke6440 pushes a realistic project size through the hot paths
|
||||
// with a loose time bound: the point is the absence of O(n^2) / N+1 work, not
|
||||
// a benchmark.
|
||||
func TestScaleSmoke6440(t *testing.T) {
|
||||
if testing.Short() {
|
||||
t.Skip("scale smoke test skipped in -short mode")
|
||||
}
|
||||
d, ctx := newTestDB(t)
|
||||
addrs := scaleAddrs(6440)
|
||||
|
||||
start := time.Now()
|
||||
for off := 0; off < len(addrs); off += 500 {
|
||||
end := min(off+500, len(addrs))
|
||||
if _, err := d.SubmitIPs(ctx, addrs[off:end]); err != nil {
|
||||
t.Fatalf("submit chunk: %v", err)
|
||||
}
|
||||
}
|
||||
submitDur := time.Since(start)
|
||||
|
||||
start = time.Now()
|
||||
page, total, err := d.ListRegistryPage(ctx, RegistryFilter{}, 100, 3000)
|
||||
if err != nil || total != 6440 || len(page) != 100 {
|
||||
t.Fatalf("registry page: total=%d len=%d err=%v", total, len(page), err)
|
||||
}
|
||||
if _, total, err = d.ListRegistryPage(ctx, RegistryFilter{LastResult: ResultPass}, 100, 0); err != nil || total != 0 {
|
||||
t.Fatalf("registry last_result filter: total=%d err=%v", total, err)
|
||||
}
|
||||
registryDur := time.Since(start)
|
||||
|
||||
start = time.Now()
|
||||
if items, total, err := d.ListIPsPage(ctx, IPFilter{States: []string{IPQueued}, Query: "10.0.1."}, 50, 0); err != nil || total != 256 || len(items) != 50 {
|
||||
t.Fatalf("ips page: total=%d len=%d err=%v", total, len(items), err)
|
||||
}
|
||||
if by, total, err := d.CountIPsByState(ctx); err != nil || total != 6440 || by[IPQueued] != 6440 {
|
||||
t.Fatalf("count: %v %d %v", by, total, err)
|
||||
}
|
||||
if any, err := d.AnyNonTerminalIP(ctx); err != nil || !any {
|
||||
t.Fatalf("any non terminal: %v %v", any, err)
|
||||
}
|
||||
queryDur := time.Since(start)
|
||||
|
||||
start = time.Now()
|
||||
deleted, err := d.ClearAllIPs(ctx)
|
||||
if err != nil || len(deleted) != 6440 {
|
||||
t.Fatalf("clear: n=%d err=%v", len(deleted), err)
|
||||
}
|
||||
clearDur := time.Since(start)
|
||||
|
||||
t.Logf("submit=%v registry=%v queries=%v clear=%v", submitDur, registryDur, queryDur, clearDur)
|
||||
if registryDur > 10*time.Second || queryDur > 5*time.Second || clearDur > 10*time.Second {
|
||||
t.Fatalf("too slow: registry=%v queries=%v clear=%v", registryDur, queryDur, clearDur)
|
||||
}
|
||||
}
|
||||
@@ -40,7 +40,9 @@ func newAuthTestServer(t *testing.T, admin, agent string) (*Server, *httptest.Se
|
||||
t.Fatalf("bootstrap: %v", err)
|
||||
}
|
||||
log := slog.New(slog.NewTextHandler(os.Stderr, &slog.HandlerOptions{Level: slog.LevelError}))
|
||||
srv := New(d, orchestrator.New(d, openstack.NewMockClient(), cfg, log), log).WithAuth(admin, agent)
|
||||
orch := orchestrator.New(d, openstack.NewMockClient(), cfg, log)
|
||||
t.Cleanup(func() { orch.CancelScan() })
|
||||
srv := New(d, orch, log).WithAuth(admin, agent)
|
||||
ts := httptest.NewServer(srv.Handler())
|
||||
t.Cleanup(ts.Close)
|
||||
return srv, ts
|
||||
@@ -95,8 +97,8 @@ func TestRouteTableIsClassified(t *testing.T) {
|
||||
t.Fatalf("admin route %q is %s, want admin", rt.Pattern, rt.Access)
|
||||
}
|
||||
}
|
||||
if counts["admin"] != 33 || counts["agent"] != 5 || counts["open"] != 7 {
|
||||
t.Fatalf("access counts = %v, want admin=33 agent=5 open=7", counts)
|
||||
if counts["admin"] != 34 || counts["agent"] != 5 || counts["open"] != 7 {
|
||||
t.Fatalf("access counts = %v, want admin=34 agent=5 open=7", counts)
|
||||
}
|
||||
}
|
||||
|
||||
|
||||
@@ -1,6 +1,10 @@
|
||||
package httpapi
|
||||
|
||||
import "time"
|
||||
import (
|
||||
"time"
|
||||
|
||||
"cloudipvalidator/internal/db"
|
||||
)
|
||||
|
||||
// DTOs for the admin queue-management and dynamic-config endpoints
|
||||
// (/api/v1/admin/ips, /api/v1/admin/config/*). Unlike the older read-only
|
||||
@@ -30,6 +34,7 @@ type deleteIPsResponse struct {
|
||||
|
||||
type clearQueueResponse struct {
|
||||
Deleted []string `json:"deleted"`
|
||||
Count int `json:"count"`
|
||||
}
|
||||
|
||||
type scanIPsResponse struct {
|
||||
@@ -40,6 +45,43 @@ type scanIPsResponse struct {
|
||||
SkippedInProgress []string `json:"skipped_in_progress"`
|
||||
}
|
||||
|
||||
// scanStatusDTO is the progress/status object of the background floating-IP
|
||||
// scan job (POST/GET /api/v1/admin/ips/scan). state is one of
|
||||
// idle|clearing|listing|enqueuing|done|error|cancelled; "idle" means no scan
|
||||
// has run in this control-api process yet.
|
||||
type scanStatusDTO struct {
|
||||
State string `json:"state"`
|
||||
Running bool `json:"running"`
|
||||
DryRun bool `json:"dry_run"`
|
||||
Pages int `json:"pages"`
|
||||
Discovered int `json:"discovered"`
|
||||
Free int `json:"free"`
|
||||
Added int `json:"added"`
|
||||
Requeued int `json:"requeued"`
|
||||
Reordered int `json:"reordered"`
|
||||
SkippedInProgress int `json:"skipped_in_progress"`
|
||||
StartedAt *time.Time `json:"started_at"`
|
||||
FinishedAt *time.Time `json:"finished_at"`
|
||||
Error string `json:"error"`
|
||||
}
|
||||
|
||||
// ipsPageResponse / registryPageResponse are the paginated envelopes returned
|
||||
// by GET /admin/ips and GET /admin/registry when `limit` is given; total is
|
||||
// the number of rows matching the filters (before limit/offset).
|
||||
type ipsPageResponse struct {
|
||||
Items []db.IPQueueItem `json:"items"`
|
||||
Total int `json:"total"`
|
||||
Limit int `json:"limit"`
|
||||
Offset int `json:"offset"`
|
||||
}
|
||||
|
||||
type registryPageResponse struct {
|
||||
Items []registryDTO `json:"items"`
|
||||
Total int `json:"total"`
|
||||
Limit int `json:"limit"`
|
||||
Offset int `json:"offset"`
|
||||
}
|
||||
|
||||
// registryDTO is one row of the durable per-address registry — see
|
||||
// db.RegistrySummary.
|
||||
type registryDTO struct {
|
||||
|
||||
@@ -1,9 +1,15 @@
|
||||
package httpapi
|
||||
|
||||
import (
|
||||
"errors"
|
||||
"fmt"
|
||||
"net/http"
|
||||
"net/url"
|
||||
"strconv"
|
||||
"strings"
|
||||
|
||||
"cloudipvalidator/internal/db"
|
||||
"cloudipvalidator/internal/orchestrator"
|
||||
)
|
||||
|
||||
func (s *Server) handleHealthz(w http.ResponseWriter, r *http.Request) {
|
||||
@@ -11,7 +17,13 @@ func (s *Server) handleHealthz(w http.ResponseWriter, r *http.Request) {
|
||||
}
|
||||
|
||||
func (s *Server) handleAdminStatus(w http.ResponseWriter, r *http.Request) {
|
||||
ips, err := s.DB.ListIPs(r.Context())
|
||||
// GROUP BY counts instead of loading every row: the dashboard polls this.
|
||||
byState, total, err := s.DB.CountIPsByState(r.Context())
|
||||
if err != nil {
|
||||
writeError(w, http.StatusInternalServerError, err.Error())
|
||||
return
|
||||
}
|
||||
byResult, err := s.DB.CountIPsByResult(r.Context())
|
||||
if err != nil {
|
||||
writeError(w, http.StatusInternalServerError, err.Error())
|
||||
return
|
||||
@@ -21,24 +33,117 @@ func (s *Server) handleAdminStatus(w http.ResponseWriter, r *http.Request) {
|
||||
writeError(w, http.StatusInternalServerError, err.Error())
|
||||
return
|
||||
}
|
||||
byState := map[string]int{}
|
||||
for _, ip := range ips {
|
||||
byState[ip.State]++
|
||||
overall := map[string]int{
|
||||
db.ResultPass: 0, db.ResultPartial: 0, db.ResultFail: 0, db.ResultCancelled: 0,
|
||||
}
|
||||
for res, n := range byResult {
|
||||
overall[res] = n
|
||||
}
|
||||
writeJSON(w, http.StatusOK, map[string]interface{}{
|
||||
"total_ips": len(ips),
|
||||
"total_ips": total,
|
||||
"ips_by_state": byState,
|
||||
"total_validators": len(validators),
|
||||
"results_by_overall": overall,
|
||||
})
|
||||
}
|
||||
|
||||
// handleAdminIPs lists the queue. Without `limit` it returns the bare array of
|
||||
// every row (the original contract); with `limit` (1..1000) it returns the
|
||||
// envelope {items,total,limit,offset}. Filters: state (csv of valid ip
|
||||
// states), q (substring of the address), result (pass|partial|fail|
|
||||
// cancelled), order (sequence|aggregated_at_desc); offset >= 0 (needs limit).
|
||||
func (s *Server) handleAdminIPs(w http.ResponseWriter, r *http.Request) {
|
||||
q := r.URL.Query()
|
||||
limit, offset, paged, err := parsePaging(q)
|
||||
if err != nil {
|
||||
writeError(w, http.StatusBadRequest, err.Error())
|
||||
return
|
||||
}
|
||||
filter := db.IPFilter{Query: strings.TrimSpace(q.Get("q"))}
|
||||
for _, st := range strings.Split(q.Get("state"), ",") {
|
||||
st = strings.TrimSpace(st)
|
||||
if st == "" {
|
||||
continue
|
||||
}
|
||||
if !db.IsValidIPState(st) {
|
||||
writeError(w, http.StatusBadRequest, "invalid state "+strconv.Quote(st)+" (valid: "+strings.Join(db.IPStates, ", ")+")")
|
||||
return
|
||||
}
|
||||
filter.States = append(filter.States, st)
|
||||
}
|
||||
if res := q.Get("result"); res != "" {
|
||||
if !db.IsValidResult(res) {
|
||||
writeError(w, http.StatusBadRequest, "invalid result "+strconv.Quote(res)+" (valid: pass, partial, fail, cancelled)")
|
||||
return
|
||||
}
|
||||
filter.Result = res
|
||||
}
|
||||
switch order := q.Get("order"); order {
|
||||
case "", db.IPOrderSequence:
|
||||
filter.Order = db.IPOrderSequence
|
||||
case db.IPOrderAggregatedAtDesc:
|
||||
filter.Order = order
|
||||
default:
|
||||
writeError(w, http.StatusBadRequest, "invalid order "+strconv.Quote(order)+" (valid: sequence, aggregated_at_desc)")
|
||||
return
|
||||
}
|
||||
|
||||
if len(q) == 0 {
|
||||
ips, err := s.DB.ListIPs(r.Context())
|
||||
if err != nil {
|
||||
writeError(w, http.StatusInternalServerError, err.Error())
|
||||
return
|
||||
}
|
||||
writeJSON(w, http.StatusOK, ips)
|
||||
return
|
||||
}
|
||||
items, total, err := s.DB.ListIPsPage(r.Context(), filter, limit, offset)
|
||||
if err != nil {
|
||||
writeError(w, http.StatusInternalServerError, err.Error())
|
||||
return
|
||||
}
|
||||
if !paged {
|
||||
writeJSON(w, http.StatusOK, items)
|
||||
return
|
||||
}
|
||||
writeJSON(w, http.StatusOK, ipsPageResponse{Items: items, Total: total, Limit: limit, Offset: offset})
|
||||
}
|
||||
|
||||
// maxPageLimit is the largest allowed `limit` of the paginated list endpoints.
|
||||
const maxPageLimit = 1000
|
||||
|
||||
// parsePaging reads limit/offset. paged is true when `limit` was given (the
|
||||
// envelope form); `offset` without `limit` is rejected.
|
||||
func parsePaging(q url.Values) (limit, offset int, paged bool, err error) {
|
||||
if v := q.Get("limit"); v != "" {
|
||||
limit, err = strconv.Atoi(v)
|
||||
if err != nil || limit < 1 || limit > maxPageLimit {
|
||||
return 0, 0, false, fmt.Errorf("limit must be an integer in 1..%d", maxPageLimit)
|
||||
}
|
||||
paged = true
|
||||
}
|
||||
if v := q.Get("offset"); v != "" {
|
||||
if !paged {
|
||||
return 0, 0, false, errors.New("offset requires limit")
|
||||
}
|
||||
offset, err = strconv.Atoi(v)
|
||||
if err != nil || offset < 0 {
|
||||
return 0, 0, false, errors.New("offset must be an integer >= 0")
|
||||
}
|
||||
}
|
||||
return limit, offset, paged, nil
|
||||
}
|
||||
|
||||
// parseBoolParam reads an optional boolean query parameter.
|
||||
func parseBoolParam(q url.Values, name string) (bool, error) {
|
||||
v := strings.ToLower(q.Get(name))
|
||||
switch v {
|
||||
case "", "0", "false":
|
||||
return false, nil
|
||||
case "1", "true":
|
||||
return true, nil
|
||||
}
|
||||
return false, fmt.Errorf("%s must be true or false", name)
|
||||
}
|
||||
|
||||
func (s *Server) handleAdminIPDetail(w http.ResponseWriter, r *http.Request) {
|
||||
@@ -99,13 +204,34 @@ func (s *Server) handleAdminSubmitIPs(w http.ResponseWriter, r *http.Request) {
|
||||
})
|
||||
}
|
||||
|
||||
// handleAdminScanFloatingIPs lists every floating IP in the configured
|
||||
// OpenStack project, filters to the free (unassociated) pool, and submits
|
||||
// that address list to the check queue — see orchestrator.ScanFloatingIPs.
|
||||
// Takes no body; POST is used (rather than GET) because it mutates the
|
||||
// queue, matching handleAdminSubmitIPs.
|
||||
// handleAdminScanFloatingIPs starts the background floating-IP scan (see
|
||||
// orchestrator.StartScan) and answers 202 with its status at once; a scan that
|
||||
// is already running is joined (202 with the running job's status, no second
|
||||
// job). Query: dry_run=true only discovers and counts, leaving the queue
|
||||
// untouched; wait=true blocks until the job finishes and answers 200 with the
|
||||
// classic synchronous body {scanned_free, added, requeued, reordered,
|
||||
// skipped_in_progress} (502 if the scan failed). Takes no body; POST is used
|
||||
// (rather than GET) because it mutates the queue, matching handleAdminSubmitIPs.
|
||||
func (s *Server) handleAdminScanFloatingIPs(w http.ResponseWriter, r *http.Request) {
|
||||
result, scannedFree, err := s.Orch.ScanFloatingIPs(r.Context())
|
||||
q := r.URL.Query()
|
||||
dryRun, err := parseBoolParam(q, "dry_run")
|
||||
if err != nil {
|
||||
writeError(w, http.StatusBadRequest, err.Error())
|
||||
return
|
||||
}
|
||||
wait, err := parseBoolParam(q, "wait")
|
||||
if err != nil {
|
||||
writeError(w, http.StatusBadRequest, err.Error())
|
||||
return
|
||||
}
|
||||
opts := orchestrator.ScanOptions{DryRun: dryRun}
|
||||
|
||||
if !wait {
|
||||
st, _ := s.Orch.StartScan(opts)
|
||||
writeJSON(w, http.StatusAccepted, toScanStatusDTO(st))
|
||||
return
|
||||
}
|
||||
result, scannedFree, err := s.Orch.ScanAndWait(r.Context(), opts)
|
||||
if err != nil {
|
||||
writeError(w, http.StatusBadGateway, err.Error())
|
||||
return
|
||||
@@ -119,6 +245,29 @@ func (s *Server) handleAdminScanFloatingIPs(w http.ResponseWriter, r *http.Reque
|
||||
})
|
||||
}
|
||||
|
||||
// handleAdminScanStatus returns the current status/progress of the scan job.
|
||||
func (s *Server) handleAdminScanStatus(w http.ResponseWriter, r *http.Request) {
|
||||
writeJSON(w, http.StatusOK, toScanStatusDTO(s.Orch.ScanStatus()))
|
||||
}
|
||||
|
||||
func toScanStatusDTO(st orchestrator.ScanStatus) scanStatusDTO {
|
||||
return scanStatusDTO{
|
||||
State: string(st.State),
|
||||
Running: st.Running,
|
||||
DryRun: st.DryRun,
|
||||
Pages: st.Pages,
|
||||
Discovered: st.Discovered,
|
||||
Free: st.Free,
|
||||
Added: st.Added,
|
||||
Requeued: st.Requeued,
|
||||
Reordered: st.Reordered,
|
||||
SkippedInProgress: st.SkippedInProgress,
|
||||
StartedAt: st.StartedAt,
|
||||
FinishedAt: st.FinishedAt,
|
||||
Error: st.Error,
|
||||
}
|
||||
}
|
||||
|
||||
// handleAdminCancelIP force-stops a check in progress (or still-queued) for
|
||||
// the given address. Requires the orchestrator, since a floating IP may
|
||||
// need to be disassociated in OpenStack.
|
||||
@@ -176,7 +325,7 @@ func (s *Server) handleAdminClearQueue(w http.ResponseWriter, r *http.Request) {
|
||||
writeDBError(w, err)
|
||||
return
|
||||
}
|
||||
writeJSON(w, http.StatusOK, clearQueueResponse{Deleted: emptyIfNil(result.Deleted)})
|
||||
writeJSON(w, http.StatusOK, clearQueueResponse{Deleted: emptyIfNil(result.Deleted), Count: len(result.Deleted)})
|
||||
}
|
||||
|
||||
// emptyIfNil turns a nil slice into an empty one so these fields always
|
||||
|
||||
@@ -5,8 +5,35 @@ import (
|
||||
"net/http"
|
||||
"strings"
|
||||
"testing"
|
||||
"time"
|
||||
|
||||
"cloudipvalidator/internal/orchestrator"
|
||||
)
|
||||
|
||||
// stepAutoCycleThroughScan drives one cycle start: the first step launches the
|
||||
// background scan (phase scanning), then it waits for the job and runs the
|
||||
// step that consumes its result.
|
||||
func stepAutoCycleThroughScan(t *testing.T, orch *orchestrator.Orchestrator) {
|
||||
t.Helper()
|
||||
orch.AutoCycleStep(t.Context())
|
||||
waitForScan(t, orch)
|
||||
orch.AutoCycleStep(t.Context())
|
||||
}
|
||||
|
||||
func waitForScan(t *testing.T, orch *orchestrator.Orchestrator) orchestrator.ScanStatus {
|
||||
t.Helper()
|
||||
deadline := time.Now().Add(30 * time.Second)
|
||||
for {
|
||||
if st := orch.ScanStatus(); !st.Running {
|
||||
return st
|
||||
}
|
||||
if time.Now().After(deadline) {
|
||||
t.Fatalf("scan did not finish: %+v", orch.ScanStatus())
|
||||
}
|
||||
time.Sleep(2 * time.Millisecond)
|
||||
}
|
||||
}
|
||||
|
||||
func decodeAutoCycle(t *testing.T, body []byte) autoCycleDTO {
|
||||
t.Helper()
|
||||
var dto autoCycleDTO
|
||||
@@ -122,6 +149,12 @@ func TestAutoCycleStartStop(t *testing.T) {
|
||||
// The engine picks it up on the next step and the API reflects it.
|
||||
orch.AutoCycleStep(t.Context())
|
||||
resp, body = fc.do(http.MethodGet, "/api/v1/admin/auto-cycle", nil)
|
||||
if dto := decodeAutoCycle(t, body); dto.Phase != "scanning" && dto.Phase != "running" {
|
||||
t.Fatalf("expected scanning right after the first step, got %+v", dto)
|
||||
}
|
||||
waitForScan(t, orch)
|
||||
orch.AutoCycleStep(t.Context())
|
||||
resp, body = fc.do(http.MethodGet, "/api/v1/admin/auto-cycle", nil)
|
||||
if resp.StatusCode != http.StatusOK {
|
||||
t.Fatalf("get: status=%d body=%s", resp.StatusCode, body)
|
||||
}
|
||||
@@ -147,7 +180,7 @@ func TestAutoCycleStartIsIdempotent(t *testing.T) {
|
||||
if resp, body := fc.do(http.MethodPost, "/api/v1/admin/auto-cycle/start", nil); resp.StatusCode != http.StatusOK {
|
||||
t.Fatalf("start: status=%d body=%s", resp.StatusCode, body)
|
||||
}
|
||||
orch.AutoCycleStep(t.Context())
|
||||
stepAutoCycleThroughScan(t, orch)
|
||||
|
||||
resp, body := fc.do(http.MethodPost, "/api/v1/admin/auto-cycle/start", nil)
|
||||
if resp.StatusCode != http.StatusOK {
|
||||
|
||||
@@ -46,6 +46,8 @@ func newConfigTestHarness(t *testing.T) (*fakeClient, *db.DB, *orchestrator.Orch
|
||||
|
||||
log := slog.New(slog.NewTextHandler(os.Stderr, &slog.HandlerOptions{Level: slog.LevelError}))
|
||||
orch := orchestrator.New(d, mock, cfg, log)
|
||||
// A background scan must never outlive the database it writes to.
|
||||
t.Cleanup(func() { orch.CancelScan() })
|
||||
srv := New(d, orch, log)
|
||||
ts := httptest.NewServer(srv.Handler())
|
||||
t.Cleanup(ts.Close)
|
||||
|
||||
@@ -2,6 +2,8 @@ package httpapi
|
||||
|
||||
import (
|
||||
"net/http"
|
||||
"strconv"
|
||||
"strings"
|
||||
|
||||
"cloudipvalidator/internal/db"
|
||||
)
|
||||
@@ -9,9 +11,30 @@ import (
|
||||
// handleAdminRegistry lists every address ever submitted to the check
|
||||
// queue, each with a summary of its accumulated check history — the
|
||||
// durable record that survives an address being deleted from ip_queue and
|
||||
// later re-added. See migrations/0007_ip_registry.sql.
|
||||
// later re-added. See migrations/0007_ip_registry.sql. Without `limit` it is
|
||||
// the bare array of every row; with `limit` (1..1000) it returns the envelope
|
||||
// {items,total,limit,offset} (filters: offset, q = substring of the address,
|
||||
// last_result = pass|partial|fail|cancelled).
|
||||
func (s *Server) handleAdminRegistry(w http.ResponseWriter, r *http.Request) {
|
||||
items, err := s.DB.ListRegistry(r.Context())
|
||||
q := r.URL.Query()
|
||||
limit, offset, paged, err := parsePaging(q)
|
||||
if err != nil {
|
||||
writeError(w, http.StatusBadRequest, err.Error())
|
||||
return
|
||||
}
|
||||
filter := db.RegistryFilter{Query: strings.TrimSpace(q.Get("q")), LastResult: q.Get("last_result")}
|
||||
if filter.LastResult != "" && !db.IsValidResult(filter.LastResult) {
|
||||
writeError(w, http.StatusBadRequest, "invalid last_result "+strconv.Quote(filter.LastResult)+" (valid: pass, partial, fail, cancelled)")
|
||||
return
|
||||
}
|
||||
|
||||
var items []db.RegistrySummary
|
||||
var total int
|
||||
if len(q) == 0 {
|
||||
items, err = s.DB.ListRegistry(r.Context())
|
||||
} else {
|
||||
items, total, err = s.DB.ListRegistryPage(r.Context(), filter, limit, offset)
|
||||
}
|
||||
if err != nil {
|
||||
writeError(w, http.StatusInternalServerError, err.Error())
|
||||
return
|
||||
@@ -20,7 +43,11 @@ func (s *Server) handleAdminRegistry(w http.ResponseWriter, r *http.Request) {
|
||||
for i, it := range items {
|
||||
out[i] = registrySummaryToDTO(it)
|
||||
}
|
||||
if !paged {
|
||||
writeJSON(w, http.StatusOK, out)
|
||||
return
|
||||
}
|
||||
writeJSON(w, http.StatusOK, registryPageResponse{Items: out, Total: total, Limit: limit, Offset: offset})
|
||||
}
|
||||
|
||||
// handleAdminRegistryHistory returns one address's registry record plus its
|
||||
|
||||
@@ -10,14 +10,15 @@ import (
|
||||
"cloudipvalidator/internal/db"
|
||||
)
|
||||
|
||||
// TestScanFloatingIPsEndpoint proves POST /api/v1/admin/ips/scan only
|
||||
// queues floating IPs that are currently unassociated in OpenStack.
|
||||
// TestScanFloatingIPsEndpoint proves POST /api/v1/admin/ips/scan?wait=true
|
||||
// (the synchronous form) only queues floating IPs that are currently
|
||||
// unassociated in OpenStack and answers with the classic counters body.
|
||||
func TestScanFloatingIPsEndpoint(t *testing.T) {
|
||||
fc, _, _, mock := newConfigTestHarness(t)
|
||||
mock.Seed("fip-free", "5.5.5.5", "svc-project")
|
||||
mock.SeedWithPort("fip-occupied", "6.6.6.6", "svc-project", "some-port")
|
||||
|
||||
resp, body := fc.do(http.MethodPost, "/api/v1/admin/ips/scan", nil)
|
||||
resp, body := fc.do(http.MethodPost, "/api/v1/admin/ips/scan?wait=true", nil)
|
||||
if resp.StatusCode != http.StatusOK {
|
||||
t.Fatalf("scan: status=%d body=%s", resp.StatusCode, body)
|
||||
}
|
||||
|
||||
@@ -0,0 +1,364 @@
|
||||
package httpapi
|
||||
|
||||
import (
|
||||
"context"
|
||||
"encoding/json"
|
||||
"errors"
|
||||
"fmt"
|
||||
"net/http"
|
||||
"strings"
|
||||
"testing"
|
||||
"time"
|
||||
|
||||
"cloudipvalidator/internal/db"
|
||||
)
|
||||
|
||||
func getScanStatus(t *testing.T, fc *fakeClient) scanStatusDTO {
|
||||
t.Helper()
|
||||
resp, body := fc.do(http.MethodGet, "/api/v1/admin/ips/scan", nil)
|
||||
if resp.StatusCode != http.StatusOK {
|
||||
t.Fatalf("scan status: %d %s", resp.StatusCode, body)
|
||||
}
|
||||
var st scanStatusDTO
|
||||
if err := json.Unmarshal(body, &st); err != nil {
|
||||
t.Fatalf("unmarshal scan status %q: %v", body, err)
|
||||
}
|
||||
return st
|
||||
}
|
||||
|
||||
func waitScanDone(t *testing.T, fc *fakeClient) scanStatusDTO {
|
||||
t.Helper()
|
||||
deadline := time.Now().Add(30 * time.Second)
|
||||
for {
|
||||
st := getScanStatus(t, fc)
|
||||
if !st.Running {
|
||||
return st
|
||||
}
|
||||
if time.Now().After(deadline) {
|
||||
t.Fatalf("scan still running: %+v", st)
|
||||
}
|
||||
time.Sleep(5 * time.Millisecond)
|
||||
}
|
||||
}
|
||||
|
||||
func TestScanStartReturns202AndStatusIsPollable(t *testing.T) {
|
||||
fc, _, _, mock := newConfigTestHarness(t)
|
||||
|
||||
// Before any scan the status is idle, with explicit null times.
|
||||
resp, body := fc.do(http.MethodGet, "/api/v1/admin/ips/scan", nil)
|
||||
if resp.StatusCode != http.StatusOK {
|
||||
t.Fatalf("status: %d %s", resp.StatusCode, body)
|
||||
}
|
||||
var raw map[string]json.RawMessage
|
||||
if err := json.Unmarshal(body, &raw); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
for _, k := range []string{"state", "running", "dry_run", "pages", "discovered", "free", "added", "requeued",
|
||||
"reordered", "skipped_in_progress", "started_at", "finished_at", "error"} {
|
||||
if _, ok := raw[k]; !ok {
|
||||
t.Fatalf("missing key %q in %s", k, body)
|
||||
}
|
||||
}
|
||||
if string(raw["state"]) != `"idle"` || string(raw["started_at"]) != "null" || string(raw["finished_at"]) != "null" {
|
||||
t.Fatalf("expected idle with null times, got %s", body)
|
||||
}
|
||||
|
||||
mock.SeedMany("fip", 250)
|
||||
mock.SeedWithPort("busy", "203.0.113.5", "svc", "port-x")
|
||||
mock.PageSize = 100
|
||||
mock.PageDelay = 60 * time.Millisecond
|
||||
|
||||
resp, body = fc.do(http.MethodPost, "/api/v1/admin/ips/scan", nil)
|
||||
if resp.StatusCode != http.StatusAccepted {
|
||||
t.Fatalf("start: expected 202, got %d %s", resp.StatusCode, body)
|
||||
}
|
||||
var st scanStatusDTO
|
||||
if err := json.Unmarshal(body, &st); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if !st.Running || st.State != "listing" || st.StartedAt == nil {
|
||||
t.Fatalf("expected a running listing job, got %+v", st)
|
||||
}
|
||||
|
||||
// Starting again while it runs joins the same job: still 202, running.
|
||||
resp, body = fc.do(http.MethodPost, "/api/v1/admin/ips/scan?dry_run=true", nil)
|
||||
if resp.StatusCode != http.StatusAccepted {
|
||||
t.Fatalf("join: expected 202, got %d %s", resp.StatusCode, body)
|
||||
}
|
||||
var joined scanStatusDTO
|
||||
if err := json.Unmarshal(body, &joined); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if !joined.Running || joined.DryRun || joined.StartedAt == nil || !joined.StartedAt.Equal(*st.StartedAt) {
|
||||
t.Fatalf("expected the running (non-dry) job's status, got %+v", joined)
|
||||
}
|
||||
|
||||
fin := waitScanDone(t, fc)
|
||||
if fin.State != "done" || fin.Pages != 3 || fin.Discovered != 251 || fin.Free != 250 || fin.Added != 250 ||
|
||||
fin.FinishedAt == nil || fin.Error != "" {
|
||||
t.Fatalf("final status: %+v", fin)
|
||||
}
|
||||
if _, body := fc.do(http.MethodGet, "/api/v1/admin/status", nil); !strings.Contains(string(body), `"total_ips":250`) {
|
||||
t.Fatalf("expected 250 queued, got %s", body)
|
||||
}
|
||||
}
|
||||
|
||||
func TestScanDryRunDoesNotTouchQueue(t *testing.T) {
|
||||
fc, d, _, mock := newConfigTestHarness(t)
|
||||
mock.SeedMany("fip", 30)
|
||||
if err := d.SeedQueue(context.Background(), []string{"9.9.9.9"}); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
|
||||
resp, body := fc.do(http.MethodPost, "/api/v1/admin/ips/scan?dry_run=true", nil)
|
||||
if resp.StatusCode != http.StatusAccepted {
|
||||
t.Fatalf("dry run start: %d %s", resp.StatusCode, body)
|
||||
}
|
||||
fin := waitScanDone(t, fc)
|
||||
if fin.State != "done" || !fin.DryRun || fin.Free != 30 || fin.Added != 0 {
|
||||
t.Fatalf("dry run status: %+v", fin)
|
||||
}
|
||||
_, body = fc.do(http.MethodGet, "/api/v1/admin/ips", nil)
|
||||
var ips []db.IPQueueItem
|
||||
if err := json.Unmarshal(body, &ips); err != nil || len(ips) != 1 || ips[0].IPAddress != "9.9.9.9" {
|
||||
t.Fatalf("dry run must not touch the queue: %s err=%v", body, err)
|
||||
}
|
||||
|
||||
// dry_run + wait: counters in the classic body, still no queue change.
|
||||
resp, body = fc.do(http.MethodPost, "/api/v1/admin/ips/scan?dry_run=true&wait=true", nil)
|
||||
var sr scanIPsResponse
|
||||
if resp.StatusCode != http.StatusOK || json.Unmarshal(body, &sr) != nil || sr.ScannedFree != 30 || len(sr.Added) != 0 {
|
||||
t.Fatalf("dry run wait: %d %s", resp.StatusCode, body)
|
||||
}
|
||||
}
|
||||
|
||||
func TestScanWaitErrorIs502AndStatusShowsError(t *testing.T) {
|
||||
fc, d, _, mock := newConfigTestHarness(t)
|
||||
mock.SeedMany("fip", 5)
|
||||
mock.ListFailure = errors.New("neutron is down")
|
||||
if err := d.SeedQueue(context.Background(), []string{"9.9.9.9"}); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
|
||||
resp, body := fc.do(http.MethodPost, "/api/v1/admin/ips/scan?wait=true", nil)
|
||||
if resp.StatusCode != http.StatusBadGateway || !strings.Contains(string(body), "neutron is down") {
|
||||
t.Fatalf("expected 502 with the cause, got %d %s", resp.StatusCode, body)
|
||||
}
|
||||
st := getScanStatus(t, fc)
|
||||
if st.State != "error" || st.Running || !strings.Contains(st.Error, "neutron is down") {
|
||||
t.Fatalf("status after failure: %+v", st)
|
||||
}
|
||||
_, body = fc.do(http.MethodGet, "/api/v1/admin/ips", nil)
|
||||
var ips []db.IPQueueItem
|
||||
if err := json.Unmarshal(body, &ips); err != nil || len(ips) != 1 {
|
||||
t.Fatalf("queue must be untouched after a failed scan: %s", body)
|
||||
}
|
||||
|
||||
for _, q := range []string{"wait=maybe", "dry_run=2"} {
|
||||
if resp, body := fc.do(http.MethodPost, "/api/v1/admin/ips/scan?"+q, nil); resp.StatusCode != http.StatusBadRequest {
|
||||
t.Fatalf("%s: expected 400, got %d %s", q, resp.StatusCode, body)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// seedMixedQueue creates addresses 10.0.0.1..10.0.0.n (all queued), then moves
|
||||
// some to done/failed with results.
|
||||
func seedMixedQueue(t *testing.T, d *db.DB, n int) []string {
|
||||
t.Helper()
|
||||
ctx := context.Background()
|
||||
var addrs []string
|
||||
for i := 1; i <= n; i++ {
|
||||
addrs = append(addrs, fmt.Sprintf("10.0.0.%d", i))
|
||||
}
|
||||
if _, err := d.SubmitIPs(ctx, addrs); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
finish := func(addr, result string, ok, bad int) {
|
||||
ip, err := d.GetIPByAddress(ctx, addr)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
for i := 0; i < ok+bad; i++ {
|
||||
if err := d.UpsertCheck(ctx, db.Check{
|
||||
IPID: ip.ID, IPAddress: addr, AttemptNumber: ip.AttemptNumber, Source: db.SourceEgress,
|
||||
CheckType: "https", Target: fmt.Sprintf("t%d", i), Success: i < ok, CheckedAt: db.Now(),
|
||||
}); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
}
|
||||
if err := d.FinishIP(ctx, ip.ID, result); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
}
|
||||
finish("10.0.0.1", db.ResultPass, 1, 0)
|
||||
finish("10.0.0.2", db.ResultPass, 1, 0)
|
||||
finish("10.0.0.3", db.ResultFail, 0, 1)
|
||||
finish("10.0.0.4", db.ResultPartial, 1, 1)
|
||||
return addrs
|
||||
}
|
||||
|
||||
func TestAdminIPsPaginationFiltersAndCompat(t *testing.T) {
|
||||
fc, d, _, _ := newConfigTestHarness(t)
|
||||
seedMixedQueue(t, d, 12)
|
||||
|
||||
// No params: exactly the old bare array.
|
||||
resp, body := fc.do(http.MethodGet, "/api/v1/admin/ips", nil)
|
||||
if resp.StatusCode != http.StatusOK || !strings.HasPrefix(strings.TrimSpace(string(body)), "[") {
|
||||
t.Fatalf("expected a bare array, got %d %.60s", resp.StatusCode, body)
|
||||
}
|
||||
var all []db.IPQueueItem
|
||||
if err := json.Unmarshal(body, &all); err != nil || len(all) != 12 {
|
||||
t.Fatalf("bare array: n=%d err=%v", len(all), err)
|
||||
}
|
||||
|
||||
page := func(query string) ipsPageResponse {
|
||||
t.Helper()
|
||||
resp, body := fc.do(http.MethodGet, "/api/v1/admin/ips?"+query, nil)
|
||||
if resp.StatusCode != http.StatusOK {
|
||||
t.Fatalf("%s: %d %s", query, resp.StatusCode, body)
|
||||
}
|
||||
var p ipsPageResponse
|
||||
if err := json.Unmarshal(body, &p); err != nil {
|
||||
t.Fatalf("%s: unmarshal: %v (%s)", query, err, body)
|
||||
}
|
||||
return p
|
||||
}
|
||||
|
||||
p := page("limit=5")
|
||||
if len(p.Items) != 5 || p.Total != 12 || p.Limit != 5 || p.Offset != 0 || p.Items[0].IPAddress != "10.0.0.1" {
|
||||
t.Fatalf("limit=5: %+v", p)
|
||||
}
|
||||
p = page("limit=5&offset=10")
|
||||
if len(p.Items) != 2 || p.Total != 12 || p.Offset != 10 {
|
||||
t.Fatalf("offset=10: len=%d total=%d", len(p.Items), p.Total)
|
||||
}
|
||||
p = page("limit=50&state=done,failed")
|
||||
if p.Total != 4 || len(p.Items) != 4 {
|
||||
t.Fatalf("state csv: %+v", p)
|
||||
}
|
||||
p = page("limit=50&state=queued&q=0.0.1")
|
||||
if p.Total != 3 { // 10.0.0.10, .11, .12
|
||||
t.Fatalf("state+q: total=%d", p.Total)
|
||||
}
|
||||
p = page("limit=50&result=fail")
|
||||
if p.Total != 1 || p.Items[0].IPAddress != "10.0.0.3" {
|
||||
t.Fatalf("result: %+v", p)
|
||||
}
|
||||
p = page("limit=2&state=done,failed&order=aggregated_at_desc")
|
||||
if p.Total != 4 || len(p.Items) != 2 || p.Items[0].AggregatedAt == nil {
|
||||
t.Fatalf("order: %+v", p)
|
||||
}
|
||||
if p = page("limit=5&state=checking"); p.Total != 0 || p.Items == nil || len(p.Items) != 0 {
|
||||
t.Fatalf("empty page must have items: [] , got %+v", p)
|
||||
}
|
||||
|
||||
for _, bad := range []string{
|
||||
"limit=0", "limit=1001", "limit=abc", "limit=5&offset=-1", "limit=5&offset=x", "offset=5",
|
||||
"limit=5&state=bogus", "limit=5&state=done,nope", "limit=5&result=weird", "limit=5&order=random",
|
||||
} {
|
||||
if resp, body := fc.do(http.MethodGet, "/api/v1/admin/ips?"+bad, nil); resp.StatusCode != http.StatusBadRequest {
|
||||
t.Fatalf("%s: expected 400, got %d %s", bad, resp.StatusCode, body)
|
||||
}
|
||||
}
|
||||
if resp, _ := fc.do(http.MethodGet, "/api/v1/admin/ips?limit=1000", nil); resp.StatusCode != http.StatusOK {
|
||||
t.Fatalf("limit=1000 must be allowed")
|
||||
}
|
||||
}
|
||||
|
||||
func TestAdminRegistryPaginationFiltersAndCompat(t *testing.T) {
|
||||
fc, d, _, _ := newConfigTestHarness(t)
|
||||
seedMixedQueue(t, d, 12)
|
||||
|
||||
resp, body := fc.do(http.MethodGet, "/api/v1/admin/registry", nil)
|
||||
var bare []registryDTO
|
||||
if resp.StatusCode != http.StatusOK || json.Unmarshal(body, &bare) != nil || len(bare) != 12 {
|
||||
t.Fatalf("bare array: %d %.80s", resp.StatusCode, body)
|
||||
}
|
||||
|
||||
page := func(query string) registryPageResponse {
|
||||
t.Helper()
|
||||
resp, body := fc.do(http.MethodGet, "/api/v1/admin/registry?"+query, nil)
|
||||
if resp.StatusCode != http.StatusOK {
|
||||
t.Fatalf("%s: %d %s", query, resp.StatusCode, body)
|
||||
}
|
||||
var p registryPageResponse
|
||||
if err := json.Unmarshal(body, &p); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
return p
|
||||
}
|
||||
p := page("limit=5&offset=5")
|
||||
if len(p.Items) != 5 || p.Total != 12 || p.Limit != 5 || p.Offset != 5 || p.Items[0].IPAddress != bare[5].IPAddress {
|
||||
t.Fatalf("paging: %+v", p)
|
||||
}
|
||||
if p = page("limit=50&last_result=pass"); p.Total != 2 || len(p.Items) != 2 || p.Items[0].LastResult != "pass" {
|
||||
t.Fatalf("last_result=pass: %+v", p)
|
||||
}
|
||||
if p = page("limit=50&last_result=partial"); p.Total != 1 || p.Items[0].IPAddress != "10.0.0.4" {
|
||||
t.Fatalf("last_result=partial: %+v", p)
|
||||
}
|
||||
if p = page("limit=50&last_result=cancelled"); p.Total != 0 || p.Items == nil {
|
||||
t.Fatalf("last_result=cancelled: %+v", p)
|
||||
}
|
||||
if p = page("limit=50&q=0.0.1"); p.Total != 4 { // 10.0.0.1, .10, .11, .12
|
||||
t.Fatalf("q: %+v", p)
|
||||
}
|
||||
for _, bad := range []string{"limit=0", "limit=1001", "offset=1", "limit=5&last_result=bogus", "limit=5&offset=-3"} {
|
||||
if resp, body := fc.do(http.MethodGet, "/api/v1/admin/registry?"+bad, nil); resp.StatusCode != http.StatusBadRequest {
|
||||
t.Fatalf("%s: expected 400, got %d %s", bad, resp.StatusCode, body)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
func TestAdminStatusResultsByOverall(t *testing.T) {
|
||||
fc, d, _, _ := newConfigTestHarness(t)
|
||||
|
||||
// Empty: all four keys present with zeros.
|
||||
_, body := fc.do(http.MethodGet, "/api/v1/admin/status", nil)
|
||||
var st struct {
|
||||
TotalIPs int `json:"total_ips"`
|
||||
IPsByState map[string]int `json:"ips_by_state"`
|
||||
TotalValidators int `json:"total_validators"`
|
||||
ResultsByOverall map[string]int `json:"results_by_overall"`
|
||||
}
|
||||
if err := json.Unmarshal(body, &st); err != nil {
|
||||
t.Fatalf("%s: %v", body, err)
|
||||
}
|
||||
if st.TotalIPs != 0 || len(st.ResultsByOverall) != 4 || st.ResultsByOverall["pass"] != 0 {
|
||||
t.Fatalf("empty status: %s", body)
|
||||
}
|
||||
|
||||
seedMixedQueue(t, d, 6)
|
||||
_, body = fc.do(http.MethodGet, "/api/v1/admin/status", nil)
|
||||
if err := json.Unmarshal(body, &st); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
r := st.ResultsByOverall
|
||||
if st.TotalIPs != 6 || st.IPsByState["queued"] != 2 || st.IPsByState["done"] != 3 || st.IPsByState["failed"] != 1 {
|
||||
t.Fatalf("by state: %s", body)
|
||||
}
|
||||
if r["pass"] != 2 || r["partial"] != 1 || r["fail"] != 1 || r["cancelled"] != 0 {
|
||||
t.Fatalf("results_by_overall: %v", r)
|
||||
}
|
||||
}
|
||||
|
||||
func TestAdminClearReportsCount(t *testing.T) {
|
||||
fc, d, _, _ := newConfigTestHarness(t)
|
||||
seedMixedQueue(t, d, 7)
|
||||
|
||||
resp, body := fc.do(http.MethodPost, "/api/v1/admin/ips/clear", nil)
|
||||
if resp.StatusCode != http.StatusOK {
|
||||
t.Fatalf("clear: %d %s", resp.StatusCode, body)
|
||||
}
|
||||
var cr struct {
|
||||
Deleted []string `json:"deleted"`
|
||||
Count int `json:"count"`
|
||||
}
|
||||
if err := json.Unmarshal(body, &cr); err != nil || cr.Count != 7 || len(cr.Deleted) != 7 {
|
||||
t.Fatalf("clear body: %s err=%v", body, err)
|
||||
}
|
||||
resp, body = fc.do(http.MethodPost, "/api/v1/admin/ips/clear", nil)
|
||||
if err := json.Unmarshal(body, &cr); err != nil || cr.Count != 0 || cr.Deleted == nil || len(cr.Deleted) != 0 {
|
||||
t.Fatalf("empty clear: %s", body)
|
||||
}
|
||||
}
|
||||
@@ -51,6 +51,7 @@ func (s *Server) routeTable() []route {
|
||||
{"GET /api/v1/admin/ips", s.handleAdminIPs, accessAdmin},
|
||||
{"POST /api/v1/admin/ips", s.handleAdminSubmitIPs, accessAdmin},
|
||||
{"POST /api/v1/admin/ips/scan", s.handleAdminScanFloatingIPs, accessAdmin},
|
||||
{"GET /api/v1/admin/ips/scan", s.handleAdminScanStatus, accessAdmin},
|
||||
{"GET /api/v1/admin/ips/{ip}", s.handleAdminIPDetail, accessAdmin},
|
||||
{"POST /api/v1/admin/ips/{ip}/cancel", s.handleAdminCancelIP, accessAdmin},
|
||||
{"DELETE /api/v1/admin/ips/{ip}", s.handleAdminDeleteIP, accessAdmin},
|
||||
|
||||
@@ -3,10 +3,13 @@ package openstack
|
||||
import (
|
||||
"context"
|
||||
"fmt"
|
||||
"net/url"
|
||||
"time"
|
||||
|
||||
"github.com/gophercloud/gophercloud/v2"
|
||||
osauth "github.com/gophercloud/gophercloud/v2/openstack"
|
||||
"github.com/gophercloud/gophercloud/v2/openstack/networking/v2/extensions/layer3/floatingips"
|
||||
"github.com/gophercloud/gophercloud/v2/pagination"
|
||||
)
|
||||
|
||||
// AuthMethod selects how the client obtains the token it uses for Neutron
|
||||
@@ -51,10 +54,20 @@ type ClientConfig struct {
|
||||
ProjectID string
|
||||
Region string
|
||||
Interface string // "public" | "internal" | "admin"; "" defaults to "public"
|
||||
|
||||
// RequestTimeout bounds every single HTTP request to Keystone/Neutron
|
||||
// (provider.HTTPClient.Timeout). Zero means no timeout (not recommended:
|
||||
// a hung Neutron call would otherwise block the caller forever).
|
||||
RequestTimeout time.Duration
|
||||
// ListPageRetries is how many times one failed page of the floating-IP
|
||||
// listing is retried (exponential backoff 1s,2s,4s,...). Zero disables
|
||||
// retries; negative values mean "use DefaultListPageRetries".
|
||||
ListPageRetries int
|
||||
}
|
||||
|
||||
type Client struct {
|
||||
networking *gophercloud.ServiceClient
|
||||
retry pageRetry
|
||||
}
|
||||
|
||||
// buildAuthOptions translates ClientConfig into gophercloud.AuthOptions. It
|
||||
@@ -119,6 +132,10 @@ func NewClient(ctx context.Context, cfg ClientConfig) (*Client, error) {
|
||||
return nil, fmt.Errorf("openstack: authenticate: %w", err)
|
||||
}
|
||||
|
||||
if cfg.RequestTimeout > 0 {
|
||||
provider.HTTPClient.Timeout = cfg.RequestTimeout
|
||||
}
|
||||
|
||||
iface := cfg.Interface
|
||||
if iface == "" {
|
||||
iface = string(gophercloud.AvailabilityPublic)
|
||||
@@ -131,7 +148,11 @@ func NewClient(ctx context.Context, cfg ClientConfig) (*Client, error) {
|
||||
return nil, fmt.Errorf("openstack: networking client: %w", err)
|
||||
}
|
||||
|
||||
return &Client{networking: networking}, nil
|
||||
retries := cfg.ListPageRetries
|
||||
if retries < 0 {
|
||||
retries = DefaultListPageRetries
|
||||
}
|
||||
return &Client{networking: networking, retry: pageRetry{Retries: retries}}, nil
|
||||
}
|
||||
|
||||
func (c *Client) GetFloatingIPByAddress(ctx context.Context, address string) (*FloatingIP, error) {
|
||||
@@ -150,22 +171,78 @@ func (c *Client) GetFloatingIPByAddress(ctx context.Context, address string) (*F
|
||||
return &FloatingIP{ID: f.ID, Address: f.FloatingIP, PortID: f.PortID, ProjectID: f.TenantID}, nil
|
||||
}
|
||||
|
||||
func (c *Client) ListFloatingIPs(ctx context.Context) ([]FloatingIP, error) {
|
||||
pages, err := floatingips.List(c.networking, floatingips.ListOpts{}).AllPages(ctx)
|
||||
// fipListFields is the set of attributes requested from Neutron when listing:
|
||||
// everything FloatingIP needs and nothing more (the full resource is several
|
||||
// times larger, which matters with thousands of floating IPs).
|
||||
var fipListFields = []string{"id", "floating_ip_address", "port_id", "project_id"}
|
||||
|
||||
// pagedListOpts wraps floatingips.ListOpts to add the `fields` query
|
||||
// parameter, which gophercloud's ListOpts does not expose.
|
||||
type pagedListOpts struct {
|
||||
floatingips.ListOpts
|
||||
fields []string
|
||||
}
|
||||
|
||||
func (o pagedListOpts) ToFloatingIPListQuery() (string, error) {
|
||||
q, err := o.ListOpts.ToFloatingIPListQuery()
|
||||
if err != nil {
|
||||
return "", err
|
||||
}
|
||||
if len(o.fields) == 0 {
|
||||
return q, nil
|
||||
}
|
||||
v := url.Values{}
|
||||
for _, f := range o.fields {
|
||||
v.Add("fields", f)
|
||||
}
|
||||
if q == "" {
|
||||
return "?" + v.Encode(), nil
|
||||
}
|
||||
return q + "&" + v.Encode(), nil
|
||||
}
|
||||
|
||||
// fetchPage requests exactly one page (limit entries after marker). It uses
|
||||
// marker pagination driven by us rather than gophercloud's `next` link: behind
|
||||
// a proxy that link may point at an internal host.
|
||||
func (c *Client) fetchPage(ctx context.Context, marker string, limit int) ([]FloatingIP, error) {
|
||||
opts := pagedListOpts{
|
||||
ListOpts: floatingips.ListOpts{Limit: limit, Marker: marker},
|
||||
fields: fipListFields,
|
||||
}
|
||||
var out []FloatingIP
|
||||
err := floatingips.List(c.networking, opts).EachPage(ctx, func(_ context.Context, page pagination.Page) (bool, error) {
|
||||
list, err := floatingips.ExtractFloatingIPs(page)
|
||||
if err != nil {
|
||||
return false, fmt.Errorf("openstack: extract floating ips: %w", err)
|
||||
}
|
||||
out = make([]FloatingIP, 0, len(list))
|
||||
for _, f := range list {
|
||||
proj := f.TenantID
|
||||
if proj == "" {
|
||||
proj = f.ProjectID // Neutron may return only project_id when `fields` is used
|
||||
}
|
||||
out = append(out, FloatingIP{ID: f.ID, Address: f.FloatingIP, PortID: f.PortID, ProjectID: proj})
|
||||
}
|
||||
return false, nil // one page per request
|
||||
})
|
||||
if err != nil {
|
||||
return nil, fmt.Errorf("openstack: list floating ips: %w", err)
|
||||
}
|
||||
list, err := floatingips.ExtractFloatingIPs(pages)
|
||||
if err != nil {
|
||||
return nil, fmt.Errorf("openstack: extract floating ips: %w", err)
|
||||
}
|
||||
out := make([]FloatingIP, 0, len(list))
|
||||
for _, f := range list {
|
||||
out = append(out, FloatingIP{ID: f.ID, Address: f.FloatingIP, PortID: f.PortID, ProjectID: f.TenantID})
|
||||
}
|
||||
return out, nil
|
||||
}
|
||||
|
||||
// ListFreeFloatingIPs pages through every floating IP of the project (page by
|
||||
// page, retrying transient failures per page) and hands each page to onPage in
|
||||
// server order. All entries of a page are passed — free and associated; the
|
||||
// caller filters. It returns the number of non-empty pages read.
|
||||
func (c *Client) ListFreeFloatingIPs(ctx context.Context, pageSize int, onPage func([]FloatingIP) error) (int, error) {
|
||||
return paginate(ctx, pageSize, c.retry, c.fetchPage, onPage)
|
||||
}
|
||||
|
||||
func (c *Client) ListFloatingIPs(ctx context.Context) ([]FloatingIP, error) {
|
||||
return listAll(ctx, c)
|
||||
}
|
||||
|
||||
func (c *Client) AssociateFloatingIP(ctx context.Context, fipID, portID string) error {
|
||||
_, err := floatingips.Update(ctx, c.networking, fipID, floatingips.UpdateOpts{
|
||||
PortID: &portID,
|
||||
|
||||
@@ -0,0 +1,86 @@
|
||||
package openstack
|
||||
|
||||
import (
|
||||
"context"
|
||||
"encoding/json"
|
||||
"fmt"
|
||||
"net/http"
|
||||
"net/http/httptest"
|
||||
"testing"
|
||||
"time"
|
||||
|
||||
"github.com/gophercloud/gophercloud/v2"
|
||||
)
|
||||
|
||||
// TestClientListFreeFloatingIPsAgainstFakeNeutron drives the real Client
|
||||
// against an httptest server: marker pagination, the `fields` projection, a
|
||||
// transient 503 that is retried and an empty port_id meaning "free".
|
||||
func TestClientListFreeFloatingIPsAgainstFakeNeutron(t *testing.T) {
|
||||
const total = 5
|
||||
var requests int
|
||||
failedOnce := false
|
||||
srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
|
||||
requests++
|
||||
q := r.URL.Query()
|
||||
if got := q["fields"]; len(got) != 4 {
|
||||
t.Errorf("expected 4 fields params, got %v", got)
|
||||
}
|
||||
if q.Get("limit") != "2" {
|
||||
t.Errorf("expected limit=2, got %q", q.Get("limit"))
|
||||
}
|
||||
if q.Get("marker") == "id-1" && !failedOnce {
|
||||
failedOnce = true
|
||||
http.Error(w, "busy", http.StatusServiceUnavailable)
|
||||
return
|
||||
}
|
||||
start := 0
|
||||
if m := q.Get("marker"); m != "" {
|
||||
fmt.Sscanf(m, "id-%d", &start)
|
||||
start++
|
||||
}
|
||||
type fip struct {
|
||||
ID string `json:"id"`
|
||||
Addr string `json:"floating_ip_address"`
|
||||
PortID *string `json:"port_id"`
|
||||
TenantID string `json:"tenant_id"`
|
||||
}
|
||||
var items []fip
|
||||
for i := start; i < total && len(items) < 2; i++ {
|
||||
f := fip{ID: fmt.Sprintf("id-%d", i), Addr: fmt.Sprintf("203.0.113.%d", i+1), TenantID: "p"}
|
||||
if i == 1 {
|
||||
p := "port-1"
|
||||
f.PortID = &p
|
||||
}
|
||||
items = append(items, f)
|
||||
}
|
||||
w.Header().Set("Content-Type", "application/json")
|
||||
_ = json.NewEncoder(w).Encode(map[string]any{"floatingips": items})
|
||||
}))
|
||||
defer srv.Close()
|
||||
|
||||
c := &Client{
|
||||
networking: &gophercloud.ServiceClient{
|
||||
ProviderClient: &gophercloud.ProviderClient{HTTPClient: *srv.Client()},
|
||||
Endpoint: srv.URL + "/",
|
||||
ResourceBase: srv.URL + "/v2.0/",
|
||||
},
|
||||
retry: pageRetry{Retries: 3, Sleep: func(context.Context, time.Duration) error { return nil }},
|
||||
}
|
||||
var all []FloatingIP
|
||||
pages, err := c.ListFreeFloatingIPs(context.Background(), 2, func(p []FloatingIP) error {
|
||||
all = append(all, p...)
|
||||
return nil
|
||||
})
|
||||
if err != nil {
|
||||
t.Fatalf("list: %v", err)
|
||||
}
|
||||
if pages != 3 || len(all) != total {
|
||||
t.Fatalf("expected 3 pages / 5 fips, got %d / %d", pages, len(all))
|
||||
}
|
||||
if all[1].PortID != "port-1" || all[0].PortID != "" || all[4].Address != "203.0.113.5" {
|
||||
t.Fatalf("unexpected mapping: %+v", all)
|
||||
}
|
||||
if requests != 4 { // 3 pages + 1 retried 503
|
||||
t.Fatalf("expected 4 requests, got %d", requests)
|
||||
}
|
||||
}
|
||||
@@ -30,6 +30,15 @@ type FloatingIPClient interface {
|
||||
// the check queue without an admin having to enumerate them by hand.
|
||||
ListFloatingIPs(ctx context.Context) ([]FloatingIP, error)
|
||||
|
||||
// ListFreeFloatingIPs reads the project's floating IPs page by page
|
||||
// (pageSize entries per request; <=0 means DefaultListPageSize), calling
|
||||
// onPage with every page in server order — transient per-page failures
|
||||
// are retried inside, so a long listing survives a flaky Neutron. Every
|
||||
// entry of the page is passed, associated or free: the caller filters on
|
||||
// PortID. It returns the number of non-empty pages read. An error from
|
||||
// onPage stops the walk and is returned as-is.
|
||||
ListFreeFloatingIPs(ctx context.Context, pageSize int, onPage func(page []FloatingIP) error) (pages int, err error)
|
||||
|
||||
// AssociateFloatingIP attaches the floating IP to the given Neutron
|
||||
// port (the validator's primary NIC port).
|
||||
AssociateFloatingIP(ctx context.Context, fipID, portID string) error
|
||||
@@ -53,3 +62,16 @@ func IsNotFound(err error) bool {
|
||||
_, ok := err.(*notFoundError)
|
||||
return ok
|
||||
}
|
||||
|
||||
// listAll collects every page of c.ListFreeFloatingIPs into one slice; it is
|
||||
// the shared implementation of ListFloatingIPs on top of the paged method.
|
||||
func listAll(ctx context.Context, c FloatingIPClient) ([]FloatingIP, error) {
|
||||
var all []FloatingIP
|
||||
if _, err := c.ListFreeFloatingIPs(ctx, DefaultListPageSize, func(page []FloatingIP) error {
|
||||
all = append(all, page...)
|
||||
return nil
|
||||
}); err != nil {
|
||||
return nil, err
|
||||
}
|
||||
return all, nil
|
||||
}
|
||||
@@ -3,7 +3,10 @@ package openstack
|
||||
import (
|
||||
"context"
|
||||
"fmt"
|
||||
"net/netip"
|
||||
"sort"
|
||||
"sync"
|
||||
"time"
|
||||
)
|
||||
|
||||
// MockClient is an in-memory FloatingIPClient used by unit tests and the
|
||||
@@ -22,6 +25,24 @@ type MockClient struct {
|
||||
// ListFailure, when non-nil, is returned by every ListFloatingIPs call
|
||||
// until the test resets it to nil — to exercise the scan-error path.
|
||||
ListFailure error
|
||||
|
||||
// ListFailures is a queue of errors consumed one per page request (each
|
||||
// request, including retries, pops the head; an empty queue means no
|
||||
// failure), on top of the sticky ListFailure. Exercises per-page retry.
|
||||
ListFailures []error
|
||||
|
||||
// PageSize, when > 0, overrides the page size requested by the caller.
|
||||
PageSize int
|
||||
// PageDelay is an artificial latency added to every page request
|
||||
// (honours ctx cancellation) — to simulate a slow Neutron.
|
||||
PageDelay time.Duration
|
||||
// PageRetries / Sleep configure per-page retry like the real client;
|
||||
// zero retries by default so a failure surfaces immediately.
|
||||
PageRetries int
|
||||
Sleep func(ctx context.Context, d time.Duration) error
|
||||
|
||||
// PageCalls counts page requests served (successful or failed).
|
||||
PageCalls int
|
||||
}
|
||||
|
||||
func NewMockClient() *MockClient {
|
||||
@@ -60,19 +81,84 @@ func (m *MockClient) GetFloatingIPByAddress(ctx context.Context, address string)
|
||||
return &f, nil
|
||||
}
|
||||
|
||||
func (m *MockClient) ListFloatingIPs(ctx context.Context) ([]FloatingIP, error) {
|
||||
// SeedMany registers n free floating IPs with IDs "<prefix>-00000"… and
|
||||
// ascending addresses starting at 198.18.0.1 (TEST-NET benchmarking range),
|
||||
// returning the addresses in seed order.
|
||||
func (m *MockClient) SeedMany(prefix string, n int) []string {
|
||||
m.mu.Lock()
|
||||
defer m.mu.Unlock()
|
||||
base := netip.MustParseAddr("198.18.0.1")
|
||||
addrs := make([]string, 0, n)
|
||||
a := base
|
||||
for i := 0; i < n; i++ {
|
||||
id := fmt.Sprintf("%s-%05d", prefix, i)
|
||||
addr := a.String()
|
||||
m.fips[id] = &FloatingIP{ID: id, Address: addr, ProjectID: "mock-project"}
|
||||
m.byIP[addr] = id
|
||||
addrs = append(addrs, addr)
|
||||
a = a.Next()
|
||||
}
|
||||
return addrs
|
||||
}
|
||||
|
||||
func (m *MockClient) fetchPage(ctx context.Context, marker string, limit int) ([]FloatingIP, error) {
|
||||
m.mu.Lock()
|
||||
delay := m.PageDelay
|
||||
m.mu.Unlock()
|
||||
if delay > 0 {
|
||||
if err := sleepCtx(ctx, delay); err != nil {
|
||||
return nil, err
|
||||
}
|
||||
}
|
||||
m.mu.Lock()
|
||||
defer m.mu.Unlock()
|
||||
m.PageCalls++
|
||||
if m.ListFailure != nil {
|
||||
return nil, m.ListFailure
|
||||
}
|
||||
out := make([]FloatingIP, 0, len(m.fips))
|
||||
for _, f := range m.fips {
|
||||
out = append(out, *f)
|
||||
if len(m.ListFailures) > 0 {
|
||||
err := m.ListFailures[0]
|
||||
m.ListFailures = m.ListFailures[1:]
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
}
|
||||
if m.PageSize > 0 {
|
||||
limit = m.PageSize
|
||||
}
|
||||
ids := make([]string, 0, len(m.fips))
|
||||
for id := range m.fips {
|
||||
if id > marker {
|
||||
ids = append(ids, id)
|
||||
}
|
||||
}
|
||||
sort.Strings(ids)
|
||||
if len(ids) > limit {
|
||||
ids = ids[:limit]
|
||||
}
|
||||
out := make([]FloatingIP, 0, len(ids))
|
||||
for _, id := range ids {
|
||||
out = append(out, *m.fips[id])
|
||||
}
|
||||
return out, nil
|
||||
}
|
||||
|
||||
// ListFreeFloatingIPs pages through the seeded floating IPs sorted by ID, like
|
||||
// the real client does against Neutron (see FloatingIPClient).
|
||||
func (m *MockClient) ListFreeFloatingIPs(ctx context.Context, pageSize int, onPage func([]FloatingIP) error) (int, error) {
|
||||
m.mu.Lock()
|
||||
rp := pageRetry{Retries: m.PageRetries, Sleep: m.Sleep}
|
||||
if m.PageSize > 0 {
|
||||
pageSize = m.PageSize
|
||||
}
|
||||
m.mu.Unlock()
|
||||
return paginate(ctx, pageSize, rp, m.fetchPage, onPage)
|
||||
}
|
||||
|
||||
func (m *MockClient) ListFloatingIPs(ctx context.Context) ([]FloatingIP, error) {
|
||||
return listAll(ctx, m)
|
||||
}
|
||||
|
||||
func (m *MockClient) AssociateFloatingIP(ctx context.Context, fipID, portID string) error {
|
||||
m.mu.Lock()
|
||||
defer m.mu.Unlock()
|
||||
|
||||
@@ -0,0 +1,148 @@
|
||||
package openstack
|
||||
|
||||
import (
|
||||
"context"
|
||||
"errors"
|
||||
"io"
|
||||
"net"
|
||||
"strings"
|
||||
"time"
|
||||
)
|
||||
|
||||
// DefaultListPageSize is the page size used when the caller passes a
|
||||
// non-positive one.
|
||||
const DefaultListPageSize = 200
|
||||
|
||||
// DefaultListPageRetries is how many times a single failed page request is
|
||||
// retried (so up to DefaultListPageRetries+1 attempts in total).
|
||||
const DefaultListPageRetries = 5
|
||||
|
||||
// pageFetcher fetches one page of floating IPs: at most limit entries that
|
||||
// follow the entry with ID marker ("" = from the start), in a stable server
|
||||
// order. A page shorter than limit is the last one.
|
||||
type pageFetcher func(ctx context.Context, marker string, limit int) ([]FloatingIP, error)
|
||||
|
||||
// pageRetry configures per-page retries. The zero value retries nothing.
|
||||
type pageRetry struct {
|
||||
// Retries is the number of retries after the first failed attempt.
|
||||
Retries int
|
||||
// Sleep waits for d or until ctx is done; nil uses a real timer. Tests
|
||||
// inject a fake to avoid real backoff delays.
|
||||
Sleep func(ctx context.Context, d time.Duration) error
|
||||
}
|
||||
|
||||
// retryBackoff is the delay before retry number attempt (1-based):
|
||||
// 1s, 2s, 4s, 8s, 16s, then capped at 30s.
|
||||
func retryBackoff(attempt int) time.Duration {
|
||||
if attempt < 1 {
|
||||
attempt = 1
|
||||
}
|
||||
d := time.Second << uint(attempt-1)
|
||||
if d > 30*time.Second || d <= 0 {
|
||||
d = 30 * time.Second
|
||||
}
|
||||
return d
|
||||
}
|
||||
|
||||
func sleepCtx(ctx context.Context, d time.Duration) error {
|
||||
t := time.NewTimer(d)
|
||||
defer t.Stop()
|
||||
select {
|
||||
case <-ctx.Done():
|
||||
return ctx.Err()
|
||||
case <-t.C:
|
||||
return nil
|
||||
}
|
||||
}
|
||||
|
||||
// statusCoder is implemented by gophercloud.ErrUnexpectedResponseCode (and
|
||||
// anything wrapping it): it carries the HTTP status of a failed request.
|
||||
type statusCoder interface{ GetStatusCode() int }
|
||||
|
||||
// IsRetryableListError reports whether a failed page request is worth
|
||||
// repeating: transport-level failures (timeouts, resets, EOF /
|
||||
// RemoteDisconnected), HTTP 429 and HTTP 5xx. Any other 4xx, context
|
||||
// cancellation and unknown errors are not retried.
|
||||
func IsRetryableListError(err error) bool {
|
||||
if err == nil {
|
||||
return false
|
||||
}
|
||||
if errors.Is(err, context.Canceled) {
|
||||
return false
|
||||
}
|
||||
var sc statusCoder
|
||||
if errors.As(err, &sc) {
|
||||
code := sc.GetStatusCode()
|
||||
return code == 429 || code >= 500
|
||||
}
|
||||
if errors.Is(err, io.EOF) || errors.Is(err, io.ErrUnexpectedEOF) {
|
||||
return true
|
||||
}
|
||||
// A deadline is retryable only when it is a per-request network timeout
|
||||
// (net.Error), which is checked below; the caller's own ctx deadline is
|
||||
// filtered out by the paginator via ctx.Err().
|
||||
var ne net.Error
|
||||
if errors.As(err, &ne) {
|
||||
return true
|
||||
}
|
||||
msg := strings.ToLower(err.Error())
|
||||
for _, s := range []string{
|
||||
"remotedisconnected", "remote end closed connection",
|
||||
"connection reset", "connection refused", "broken pipe",
|
||||
"unexpected eof", "eof", "timeout", "tls handshake",
|
||||
} {
|
||||
if strings.Contains(msg, s) {
|
||||
return true
|
||||
}
|
||||
}
|
||||
return false
|
||||
}
|
||||
|
||||
// paginate walks all pages with marker-based pagination, retrying each failed
|
||||
// page per rp, and hands every non-empty page to onPage in server order. It
|
||||
// returns the number of non-empty pages delivered. An error from onPage, a
|
||||
// non-retryable error, exhausted retries or ctx cancellation stop the walk.
|
||||
func paginate(ctx context.Context, pageSize int, rp pageRetry, fetch pageFetcher, onPage func([]FloatingIP) error) (int, error) {
|
||||
if pageSize <= 0 {
|
||||
pageSize = DefaultListPageSize
|
||||
}
|
||||
sleep := rp.Sleep
|
||||
if sleep == nil {
|
||||
sleep = sleepCtx
|
||||
}
|
||||
marker := ""
|
||||
pages := 0
|
||||
for {
|
||||
if err := ctx.Err(); err != nil {
|
||||
return pages, err
|
||||
}
|
||||
var page []FloatingIP
|
||||
var err error
|
||||
for attempt := 0; ; attempt++ {
|
||||
page, err = fetch(ctx, marker, pageSize)
|
||||
if err == nil {
|
||||
break
|
||||
}
|
||||
if ctx.Err() != nil {
|
||||
return pages, ctx.Err()
|
||||
}
|
||||
if attempt >= rp.Retries || !IsRetryableListError(err) {
|
||||
return pages, err
|
||||
}
|
||||
if serr := sleep(ctx, retryBackoff(attempt+1)); serr != nil {
|
||||
return pages, serr
|
||||
}
|
||||
}
|
||||
if len(page) == 0 {
|
||||
return pages, nil
|
||||
}
|
||||
pages++
|
||||
if err := onPage(page); err != nil {
|
||||
return pages, err
|
||||
}
|
||||
if len(page) < pageSize {
|
||||
return pages, nil
|
||||
}
|
||||
marker = page[len(page)-1].ID
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,227 @@
|
||||
package openstack
|
||||
|
||||
import (
|
||||
"context"
|
||||
"errors"
|
||||
"fmt"
|
||||
"io"
|
||||
"net"
|
||||
"net/http"
|
||||
"net/url"
|
||||
"testing"
|
||||
"time"
|
||||
|
||||
"github.com/gophercloud/gophercloud/v2"
|
||||
"github.com/gophercloud/gophercloud/v2/openstack/networking/v2/extensions/layer3/floatingips"
|
||||
)
|
||||
|
||||
func TestMockPaginationReturnsAllPages(t *testing.T) {
|
||||
m := NewMockClient()
|
||||
m.PageSize = 200
|
||||
m.SeedMany("fip", 2500)
|
||||
|
||||
seen := map[string]bool{}
|
||||
var sizes []int
|
||||
pages, err := m.ListFreeFloatingIPs(context.Background(), 200, func(page []FloatingIP) error {
|
||||
sizes = append(sizes, len(page))
|
||||
for _, f := range page {
|
||||
if seen[f.ID] {
|
||||
t.Fatalf("duplicate id %s", f.ID)
|
||||
}
|
||||
seen[f.ID] = true
|
||||
}
|
||||
return nil
|
||||
})
|
||||
if err != nil {
|
||||
t.Fatalf("list: %v", err)
|
||||
}
|
||||
if len(seen) != 2500 {
|
||||
t.Fatalf("expected 2500 unique fips, got %d", len(seen))
|
||||
}
|
||||
if pages != 13 || len(sizes) != 13 || sizes[12] != 100 {
|
||||
t.Fatalf("expected 13 pages with a short last one, got pages=%d sizes=%v", pages, sizes)
|
||||
}
|
||||
if m.PageCalls != 13 {
|
||||
t.Fatalf("expected 13 page requests, got %d", m.PageCalls)
|
||||
}
|
||||
|
||||
// ListFloatingIPs stays available on top of the paged method.
|
||||
all, err := m.ListFloatingIPs(context.Background())
|
||||
if err != nil || len(all) != 2500 {
|
||||
t.Fatalf("ListFloatingIPs: n=%d err=%v", len(all), err)
|
||||
}
|
||||
}
|
||||
|
||||
func TestMockPaginationExactMultipleOfPageSize(t *testing.T) {
|
||||
m := NewMockClient()
|
||||
m.SeedMany("fip", 400)
|
||||
pages, err := m.ListFreeFloatingIPs(context.Background(), 200, func([]FloatingIP) error { return nil })
|
||||
if err != nil || pages != 2 {
|
||||
t.Fatalf("expected 2 non-empty pages, got pages=%d err=%v", pages, err)
|
||||
}
|
||||
}
|
||||
|
||||
func TestPageRetryOnTransientErrors(t *testing.T) {
|
||||
m := NewMockClient()
|
||||
m.PageSize = 100
|
||||
m.SeedMany("fip", 250)
|
||||
m.PageRetries = 5
|
||||
var slept []time.Duration
|
||||
m.Sleep = func(ctx context.Context, d time.Duration) error {
|
||||
slept = append(slept, d)
|
||||
return nil
|
||||
}
|
||||
// Fail the first request twice, then the 2nd page once.
|
||||
m.ListFailures = []error{io.ErrUnexpectedEOF, errors.New("RemoteDisconnected('Remote end closed connection')"), nil, io.EOF}
|
||||
|
||||
n := 0
|
||||
pages, err := m.ListFreeFloatingIPs(context.Background(), 100, func(p []FloatingIP) error {
|
||||
n += len(p)
|
||||
return nil
|
||||
})
|
||||
if err != nil {
|
||||
t.Fatalf("expected retries to succeed, got %v", err)
|
||||
}
|
||||
if n != 250 || pages != 3 {
|
||||
t.Fatalf("expected 250 fips in 3 pages, got %d in %d", n, pages)
|
||||
}
|
||||
want := []time.Duration{time.Second, 2 * time.Second, time.Second}
|
||||
if len(slept) != len(want) {
|
||||
t.Fatalf("expected %d backoffs, got %v", len(want), slept)
|
||||
}
|
||||
for i := range want {
|
||||
if slept[i] != want[i] {
|
||||
t.Fatalf("backoff[%d]=%v want %v (all %v)", i, slept[i], want[i], slept)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
func TestPageRetryGivesUpAfterRetries(t *testing.T) {
|
||||
m := NewMockClient()
|
||||
m.SeedMany("fip", 10)
|
||||
m.PageRetries = 2
|
||||
m.Sleep = func(context.Context, time.Duration) error { return nil }
|
||||
m.ListFailure = io.ErrUnexpectedEOF // sticky
|
||||
_, err := m.ListFreeFloatingIPs(context.Background(), 5, func([]FloatingIP) error { return nil })
|
||||
if !errors.Is(err, io.ErrUnexpectedEOF) {
|
||||
t.Fatalf("expected the page error after retries, got %v", err)
|
||||
}
|
||||
if m.PageCalls != 3 {
|
||||
t.Fatalf("expected 1+2 attempts, got %d", m.PageCalls)
|
||||
}
|
||||
}
|
||||
|
||||
func TestPageNonRetryableAborts(t *testing.T) {
|
||||
m := NewMockClient()
|
||||
m.SeedMany("fip", 10)
|
||||
m.PageRetries = 5
|
||||
m.Sleep = func(context.Context, time.Duration) error {
|
||||
t.Fatal("must not back off for a non-retryable error")
|
||||
return nil
|
||||
}
|
||||
boom := errors.New("403 forbidden")
|
||||
m.ListFailures = []error{boom}
|
||||
called := false
|
||||
_, err := m.ListFreeFloatingIPs(context.Background(), 5, func([]FloatingIP) error { called = true; return nil })
|
||||
if !errors.Is(err, boom) || called || m.PageCalls != 1 {
|
||||
t.Fatalf("expected immediate abort: err=%v called=%v calls=%d", err, called, m.PageCalls)
|
||||
}
|
||||
}
|
||||
|
||||
func TestPageOnPageErrorStops(t *testing.T) {
|
||||
m := NewMockClient()
|
||||
m.SeedMany("fip", 50)
|
||||
stop := errors.New("stop")
|
||||
pages, err := m.ListFreeFloatingIPs(context.Background(), 10, func([]FloatingIP) error { return stop })
|
||||
if !errors.Is(err, stop) || pages != 1 || m.PageCalls != 1 {
|
||||
t.Fatalf("expected stop after first page: pages=%d err=%v calls=%d", pages, err, m.PageCalls)
|
||||
}
|
||||
}
|
||||
|
||||
func TestPageContextCancel(t *testing.T) {
|
||||
m := NewMockClient()
|
||||
m.SeedMany("fip", 100)
|
||||
m.PageDelay = 5 * time.Second
|
||||
ctx, cancel := context.WithTimeout(context.Background(), 50*time.Millisecond)
|
||||
defer cancel()
|
||||
start := time.Now()
|
||||
_, err := m.ListFreeFloatingIPs(ctx, 10, func([]FloatingIP) error { return nil })
|
||||
if !errors.Is(err, context.DeadlineExceeded) {
|
||||
t.Fatalf("expected deadline error, got %v", err)
|
||||
}
|
||||
if time.Since(start) > 2*time.Second {
|
||||
t.Fatalf("cancel was not honoured promptly")
|
||||
}
|
||||
|
||||
// Cancellation during the backoff sleep is honoured too.
|
||||
m2 := NewMockClient()
|
||||
m2.SeedMany("fip", 10)
|
||||
m2.PageRetries = 3
|
||||
m2.ListFailure = io.EOF
|
||||
ctx2, cancel2 := context.WithCancel(context.Background())
|
||||
m2.Sleep = func(ctx context.Context, d time.Duration) error { cancel2(); return ctx.Err() }
|
||||
_, err = m2.ListFreeFloatingIPs(ctx2, 5, func([]FloatingIP) error { return nil })
|
||||
if !errors.Is(err, context.Canceled) {
|
||||
t.Fatalf("expected canceled, got %v", err)
|
||||
}
|
||||
}
|
||||
|
||||
func TestIsRetryableListError(t *testing.T) {
|
||||
status := func(code int) error {
|
||||
return gophercloud.ErrUnexpectedResponseCode{Actual: code, Expected: []int{200}}
|
||||
}
|
||||
cases := []struct {
|
||||
name string
|
||||
err error
|
||||
want bool
|
||||
}{
|
||||
{"nil", nil, false},
|
||||
{"eof", io.EOF, true},
|
||||
{"unexpected eof", io.ErrUnexpectedEOF, true},
|
||||
{"wrapped eof", fmt.Errorf("list: %w", io.EOF), true},
|
||||
{"remote disconnected text", errors.New("Get x: RemoteDisconnected('Remote end closed connection without response')"), true},
|
||||
{"connection reset", errors.New("read tcp: connection reset by peer"), true},
|
||||
{"connection refused", errors.New("dial tcp: connection refused"), true},
|
||||
{"net timeout", &url.Error{Op: "Get", URL: "http://x", Err: &net.DNSError{IsTimeout: true}}, true},
|
||||
{"500", status(http.StatusInternalServerError), true},
|
||||
{"503 wrapped", fmt.Errorf("openstack: list floating ips: %w", status(503)), true},
|
||||
{"429", status(http.StatusTooManyRequests), true},
|
||||
{"400", status(http.StatusBadRequest), false},
|
||||
{"401", status(http.StatusUnauthorized), false},
|
||||
{"404", status(http.StatusNotFound), false},
|
||||
{"canceled", context.Canceled, false},
|
||||
{"plain", errors.New("boom"), false},
|
||||
}
|
||||
for _, c := range cases {
|
||||
if got := IsRetryableListError(c.err); got != c.want {
|
||||
t.Errorf("%s: IsRetryableListError=%v want %v", c.name, got, c.want)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
func TestPagedListOptsQuery(t *testing.T) {
|
||||
o := pagedListOpts{ListOpts: floatingipsListOpts("m1", 200), fields: fipListFields}
|
||||
q, err := o.ToFloatingIPListQuery()
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
v, err := url.ParseQuery(q[1:])
|
||||
if err != nil {
|
||||
t.Fatalf("parse %q: %v", q, err)
|
||||
}
|
||||
if v.Get("limit") != "200" || v.Get("marker") != "m1" {
|
||||
t.Fatalf("limit/marker missing in %q", q)
|
||||
}
|
||||
if got := v["fields"]; len(got) != 4 || got[0] != "id" || got[1] != "floating_ip_address" || got[2] != "port_id" || got[3] != "project_id" {
|
||||
t.Fatalf("unexpected fields %v in %q", got, q)
|
||||
}
|
||||
// Without any other option the query must still start with '?'.
|
||||
q, _ = pagedListOpts{fields: []string{"id"}}.ToFloatingIPListQuery()
|
||||
if q != "?fields=id" {
|
||||
t.Fatalf("got %q", q)
|
||||
}
|
||||
}
|
||||
|
||||
func floatingipsListOpts(marker string, limit int) floatingips.ListOpts {
|
||||
return floatingips.ListOpts{Marker: marker, Limit: limit}
|
||||
}
|
||||
@@ -11,11 +11,14 @@ import (
|
||||
|
||||
// This file implements the automatic check cycle: an optional, repeating
|
||||
// "clear queue -> scan floating IPs -> wait until every queued address has
|
||||
// reached a terminal state -> wait interval" scenario. Steps 1-2 reuse
|
||||
// ClearQueue/ScanFloatingIPs verbatim; step 3 needs no code at all because
|
||||
// Tick already picks up `queued` addresses. All state lives in the database
|
||||
// (db.AutoCycle), so the cycle survives a control-api restart and the
|
||||
// interval/limits can be changed at runtime.
|
||||
// reached a terminal state -> wait interval" scenario. Steps 1-2 run as ONE
|
||||
// background scan job (StartScan{ClearFirst:true}) so the control loop and
|
||||
// autoCycleMu are never held across OpenStack/DB-heavy work: the cycle sits in
|
||||
// phase `scanning` while the job runs and each step merely polls it. Step 3
|
||||
// needs no code at all because Tick already picks up `queued` addresses. All
|
||||
// state lives in the database (db.AutoCycle), so the cycle survives a
|
||||
// control-api restart (a restart in phase `scanning` simply starts the scan
|
||||
// again) and the interval/limits can be changed at runtime.
|
||||
|
||||
// GetAutoCycle returns the current auto-cycle configuration and state.
|
||||
func (o *Orchestrator) GetAutoCycle(ctx context.Context) (db.AutoCycle, error) {
|
||||
@@ -65,12 +68,18 @@ func (o *Orchestrator) StopAutoCycle(ctx context.Context) error {
|
||||
// "stopped" describes an interrupted run. Stopping during the pause
|
||||
// between cycles must not overwrite the result of the last finished one.
|
||||
outcome := ""
|
||||
if ac.Phase == db.AutoCyclePhaseRunning {
|
||||
if ac.Phase == db.AutoCyclePhaseRunning || ac.Phase == db.AutoCyclePhaseScanning {
|
||||
outcome = db.AutoCycleOutcomeStopped
|
||||
}
|
||||
if err := o.DB.SetAutoCycleEnabled(ctx, false, nil, outcome); err != nil {
|
||||
return fmt.Errorf("disable auto cycle: %w", err)
|
||||
}
|
||||
if ac.Phase == db.AutoCyclePhaseScanning {
|
||||
// Persisted first, so a step racing in right after cannot restart the
|
||||
// scan; then abort the background job (queue chunks already enqueued
|
||||
// stay, like in-flight checks).
|
||||
o.CancelScan()
|
||||
}
|
||||
o.event(ctx, "control-api", "", nil, "auto_cycle_stopped", autoCyclePayload(map[string]any{
|
||||
"phase": ac.Phase,
|
||||
}))
|
||||
@@ -106,7 +115,11 @@ func (o *Orchestrator) autoCycleStep(ctx context.Context, now time.Time) {
|
||||
return
|
||||
}
|
||||
|
||||
if ac.Phase == db.AutoCyclePhaseRunning {
|
||||
switch ac.Phase {
|
||||
case db.AutoCyclePhaseScanning:
|
||||
o.autoCycleCheckScan(ctx, ac, now)
|
||||
return
|
||||
case db.AutoCyclePhaseRunning:
|
||||
o.autoCycleCheckRun(ctx, ac, now)
|
||||
return
|
||||
}
|
||||
@@ -118,49 +131,93 @@ func (o *Orchestrator) autoCycleStep(ctx context.Context, now time.Time) {
|
||||
o.autoCycleStartRun(ctx, ac, now)
|
||||
}
|
||||
|
||||
// autoCycleStartRun performs steps 1-2 of the scenario (clear the queue,
|
||||
// scan floating IPs) and moves the state to running, or straight to waiting
|
||||
// if there is nothing to wait for.
|
||||
// autoCycleStartRun begins a cycle: it starts the background scan job (clear
|
||||
// the queue, then discover and enqueue the free floating IPs) and moves to
|
||||
// phase `scanning` right away. Nothing slow happens here, so autoCycleMu is
|
||||
// released immediately and the control loop keeps ticking.
|
||||
func (o *Orchestrator) autoCycleStartRun(ctx context.Context, ac db.AutoCycle, now time.Time) {
|
||||
interval := time.Duration(ac.IntervalSeconds) * time.Second
|
||||
next := now.Add(interval)
|
||||
if _, started := o.StartScan(ScanOptions{ClearFirst: true}); !started {
|
||||
// A scan started by someone else (an operator's manual or dry-run scan,
|
||||
// the periodic scan) is in flight. It is not this cycle's scan: it may
|
||||
// not clear the queue, or may not enqueue anything at all (dry run), so
|
||||
// adopting it would end in a "completed" cycle over an untouched or
|
||||
// empty queue. Change nothing and try again on the next step, once it
|
||||
// has finished (it is bounded by fip_scan_timeout_seconds).
|
||||
o.Log.Info("auto cycle: another floating ip scan is running, waiting for it to finish")
|
||||
return
|
||||
}
|
||||
|
||||
st := autoCycleStateOf(ac)
|
||||
st.LastRunStartedAt = &now
|
||||
st.RunStartedAt = nil
|
||||
st.NextRunAt = nil
|
||||
st.Phase = db.AutoCyclePhaseScanning
|
||||
st.LastError = ""
|
||||
if err := o.DB.UpdateAutoCycleState(ctx, st); err != nil {
|
||||
o.Log.Error("auto cycle: save state", "err", err)
|
||||
}
|
||||
}
|
||||
|
||||
fail := func(step string, err error) {
|
||||
o.Log.Error("auto cycle: step failed", "step", step, "err", err)
|
||||
// autoCycleCheckScan handles the scanning phase by polling the scan job.
|
||||
func (o *Orchestrator) autoCycleCheckScan(ctx context.Context, ac db.AutoCycle, now time.Time) {
|
||||
scan := o.ScanStatus()
|
||||
switch {
|
||||
case scan.Running:
|
||||
return
|
||||
case scan.State == ScanIdle:
|
||||
// Phase `scanning` but no job in this process: control-api restarted
|
||||
// mid-scan. Starting again is idempotent (clear + scan).
|
||||
o.Log.Info("auto cycle: restarting floating ip scan after restart")
|
||||
o.StartScan(ScanOptions{ClearFirst: true})
|
||||
return
|
||||
}
|
||||
|
||||
interval := time.Duration(ac.IntervalSeconds) * time.Second
|
||||
next := now.Add(interval)
|
||||
st := autoCycleStateOf(ac)
|
||||
|
||||
switch scan.State {
|
||||
case ScanError:
|
||||
o.Log.Error("auto cycle: step failed", "step", "scan floating ips", "err", scan.Error)
|
||||
st.Phase = db.AutoCyclePhaseWaiting
|
||||
st.RunStartedAt = nil
|
||||
st.NextRunAt = &next
|
||||
st.LastRunFinishedAt = &now
|
||||
st.LastOutcome = db.AutoCycleOutcomeError
|
||||
st.LastError = fmt.Sprintf("%s: %v", step, err)
|
||||
if uerr := o.DB.UpdateAutoCycleState(ctx, st); uerr != nil {
|
||||
o.Log.Error("auto cycle: save state", "err", uerr)
|
||||
st.LastError = "scan floating ips: " + scan.Error
|
||||
if err := o.DB.UpdateAutoCycleState(ctx, st); err != nil {
|
||||
o.Log.Error("auto cycle: save state", "err", err)
|
||||
return
|
||||
}
|
||||
o.event(ctx, "control-api", "", nil, "auto_cycle_error", autoCyclePayload(map[string]any{
|
||||
"step": step,
|
||||
"error": err.Error(),
|
||||
"step": "scan floating ips",
|
||||
"error": scan.Error,
|
||||
}))
|
||||
|
||||
case ScanCancelled:
|
||||
if life := o.lifetimeErr(); life != nil {
|
||||
// The process is shutting down: leave the phase as is so the next
|
||||
// start resumes the cycle (scanning without a job restarts it).
|
||||
return
|
||||
}
|
||||
// Cancelled by an operator (CancelScan): treat as an interrupted run.
|
||||
st.Phase = db.AutoCyclePhaseWaiting
|
||||
st.RunStartedAt = nil
|
||||
st.NextRunAt = &next
|
||||
st.LastRunFinishedAt = &now
|
||||
st.LastOutcome = db.AutoCycleOutcomeStopped
|
||||
if err := o.DB.UpdateAutoCycleState(ctx, st); err != nil {
|
||||
o.Log.Error("auto cycle: save state", "err", err)
|
||||
}
|
||||
|
||||
if _, err := o.ClearQueue(ctx); err != nil {
|
||||
fail("clear queue", err)
|
||||
return
|
||||
}
|
||||
_, scanned, err := o.ScanFloatingIPs(ctx)
|
||||
if err != nil {
|
||||
fail("scan floating ips", err)
|
||||
return
|
||||
}
|
||||
st.LastScannedFree = scanned
|
||||
default: // ScanDone
|
||||
st.LastScannedFree = scan.Free
|
||||
st.LastError = ""
|
||||
|
||||
if scanned == 0 {
|
||||
if scan.Free == 0 {
|
||||
// Nothing was queued; waiting for completion would never end.
|
||||
o.Log.Info("auto cycle: no free floating ips, waiting for next interval")
|
||||
st.Phase = db.AutoCyclePhaseWaiting
|
||||
st.RunStartedAt = nil
|
||||
st.NextRunAt = &next
|
||||
st.LastRunFinishedAt = &now
|
||||
st.LastOutcome = db.AutoCycleOutcomeNoFreeIPs
|
||||
@@ -169,7 +226,7 @@ func (o *Orchestrator) autoCycleStartRun(ctx context.Context, ac db.AutoCycle, n
|
||||
}
|
||||
return
|
||||
}
|
||||
|
||||
// max_run_seconds counts from the end of the scan.
|
||||
st.Phase = db.AutoCyclePhaseRunning
|
||||
st.RunStartedAt = &now
|
||||
st.NextRunAt = nil
|
||||
@@ -179,34 +236,40 @@ func (o *Orchestrator) autoCycleStartRun(ctx context.Context, ac db.AutoCycle, n
|
||||
}
|
||||
o.event(ctx, "control-api", "", nil, "auto_cycle_started", autoCyclePayload(map[string]any{
|
||||
"reason": "cycle",
|
||||
"scanned_free": scanned,
|
||||
"scanned_free": scan.Free,
|
||||
}))
|
||||
}
|
||||
}
|
||||
|
||||
func (o *Orchestrator) lifetimeErr() error {
|
||||
o.scan.mu.Lock()
|
||||
defer o.scan.mu.Unlock()
|
||||
return o.scan.lifetime().Err()
|
||||
}
|
||||
|
||||
// autoCycleCheckRun handles the running phase: finish the cycle once every
|
||||
// queued address is terminal, or give up after max_run_seconds.
|
||||
func (o *Orchestrator) autoCycleCheckRun(ctx context.Context, ac db.AutoCycle, now time.Time) {
|
||||
items, err := o.DB.ListIPs(ctx)
|
||||
// An empty queue counts as finished: right after the scan it cannot be
|
||||
// empty (scanned > 0), so it only happens when an operator cleared or
|
||||
// deleted every address mid-cycle — and then there is nothing to wait for
|
||||
// (with max_run_seconds=0 the cycle would otherwise hang forever).
|
||||
// EXISTS is cheap enough to run on every tick, however long the queue.
|
||||
pending, err := o.DB.AnyNonTerminalIP(ctx)
|
||||
if err != nil {
|
||||
o.Log.Error("auto cycle: list ips", "err", err)
|
||||
o.Log.Error("auto cycle: check pending ips", "err", err)
|
||||
return
|
||||
}
|
||||
|
||||
interval := time.Duration(ac.IntervalSeconds) * time.Second
|
||||
next := now.Add(interval)
|
||||
|
||||
// An empty queue counts as finished: right after the scan it cannot be
|
||||
// empty (scanned > 0), so it only happens when an operator cleared or
|
||||
// deleted every address mid-cycle — and then there is nothing to wait for
|
||||
// (with max_run_seconds=0 the cycle would otherwise hang forever).
|
||||
allTerminal := true
|
||||
for _, it := range items {
|
||||
if !isTerminalIPState(it.State) {
|
||||
allTerminal = false
|
||||
break
|
||||
if !pending {
|
||||
_, addresses, err := o.DB.CountIPsByState(ctx)
|
||||
if err != nil {
|
||||
o.Log.Error("auto cycle: count ips", "err", err)
|
||||
return
|
||||
}
|
||||
}
|
||||
if allTerminal {
|
||||
st := autoCycleStateOf(ac)
|
||||
st.Phase = db.AutoCyclePhaseWaiting
|
||||
st.RunStartedAt = nil
|
||||
@@ -220,7 +283,7 @@ func (o *Orchestrator) autoCycleCheckRun(ctx context.Context, ac db.AutoCycle, n
|
||||
return
|
||||
}
|
||||
o.event(ctx, "control-api", "", nil, "auto_cycle_completed", autoCyclePayload(map[string]any{
|
||||
"addresses": len(items),
|
||||
"addresses": addresses,
|
||||
"runs": st.RunsTotal,
|
||||
}))
|
||||
return
|
||||
@@ -228,6 +291,11 @@ func (o *Orchestrator) autoCycleCheckRun(ctx context.Context, ac db.AutoCycle, n
|
||||
|
||||
if ac.MaxRunSeconds > 0 && ac.RunStartedAt != nil &&
|
||||
now.Sub(*ac.RunStartedAt) > time.Duration(ac.MaxRunSeconds)*time.Second {
|
||||
_, addresses, err := o.DB.CountIPsByState(ctx)
|
||||
if err != nil {
|
||||
o.Log.Error("auto cycle: count ips", "err", err)
|
||||
return
|
||||
}
|
||||
// The queue is left untouched: the next cycle clears it anyway, and
|
||||
// an operator can still inspect what got stuck.
|
||||
st := autoCycleStateOf(ac)
|
||||
@@ -243,7 +311,7 @@ func (o *Orchestrator) autoCycleCheckRun(ctx context.Context, ac db.AutoCycle, n
|
||||
}
|
||||
o.event(ctx, "control-api", "", nil, "auto_cycle_timeout", autoCyclePayload(map[string]any{
|
||||
"max_run_seconds": ac.MaxRunSeconds,
|
||||
"addresses": len(items),
|
||||
"addresses": addresses,
|
||||
}))
|
||||
}
|
||||
}
|
||||
|
||||
@@ -58,6 +58,37 @@ func countEvents(t *testing.T, d *db.DB, eventType string) int {
|
||||
return n
|
||||
}
|
||||
|
||||
// waitScan blocks until the background scan job is no longer running.
|
||||
func waitScan(t *testing.T, o *Orchestrator) ScanStatus {
|
||||
t.Helper()
|
||||
return waitScanFor(t, o, 30*time.Second)
|
||||
}
|
||||
|
||||
func waitScanFor(t *testing.T, o *Orchestrator, limit time.Duration) ScanStatus {
|
||||
t.Helper()
|
||||
deadline := time.Now().Add(limit)
|
||||
for {
|
||||
if st := o.ScanStatus(); !st.Running {
|
||||
return st
|
||||
}
|
||||
if time.Now().After(deadline) {
|
||||
t.Fatalf("scan job did not finish in time: %+v", o.ScanStatus())
|
||||
}
|
||||
time.Sleep(2 * time.Millisecond)
|
||||
}
|
||||
}
|
||||
|
||||
// stepStartCycle drives one cycle start: the first step launches the
|
||||
// background scan (phase scanning), then we wait for the job and the second
|
||||
// step, at the same virtual time, consumes its result.
|
||||
func stepStartCycle(t *testing.T, o *Orchestrator, now time.Time) {
|
||||
t.Helper()
|
||||
ctx := context.Background()
|
||||
o.autoCycleStep(ctx, now)
|
||||
waitScan(t, o)
|
||||
o.autoCycleStep(ctx, now)
|
||||
}
|
||||
|
||||
func TestAutoCycleDisabledIsNoOp(t *testing.T) {
|
||||
ctx := context.Background()
|
||||
o, d, mock := newTestOrchestrator(t, 180)
|
||||
@@ -111,7 +142,7 @@ func TestAutoCycleStartClearsQueueAndScans(t *testing.T) {
|
||||
}
|
||||
|
||||
now := db.Now()
|
||||
o.autoCycleStep(ctx, now)
|
||||
stepStartCycle(t, o, now)
|
||||
|
||||
got := queuedAddresses(t, d)
|
||||
if _, stale := got["9.9.9.9"]; stale {
|
||||
@@ -154,7 +185,7 @@ func TestAutoCycleWaitsWhileChecksInProgressAndTickPicksUpQueue(t *testing.T) {
|
||||
}
|
||||
|
||||
t0 := db.Now()
|
||||
o.autoCycleStep(ctx, t0)
|
||||
stepStartCycle(t, o, t0)
|
||||
|
||||
// Existing Tick logic starts the checks on its own.
|
||||
o.Tick(ctx)
|
||||
@@ -198,7 +229,7 @@ func TestAutoCycleCompletesAndRepeatsAfterInterval(t *testing.T) {
|
||||
}
|
||||
|
||||
t0 := db.Now()
|
||||
o.autoCycleStep(ctx, t0)
|
||||
stepStartCycle(t, o, t0)
|
||||
|
||||
// done, failed and occupied are all terminal.
|
||||
if _, err := d.ExecContext(ctx, `UPDATE ip_queue SET state=? WHERE ip_address=?`, db.IPDone, "1.1.1.1"); err != nil {
|
||||
@@ -241,7 +272,7 @@ func TestAutoCycleCompletesAndRepeatsAfterInterval(t *testing.T) {
|
||||
}
|
||||
|
||||
// At next_run_at: new cycle starts, queue is rebuilt from scratch.
|
||||
o.autoCycleStep(ctx, wantNext)
|
||||
stepStartCycle(t, o, wantNext)
|
||||
ac = getAutoCycle(t, d)
|
||||
if ac.Phase != db.AutoCyclePhaseRunning {
|
||||
t.Fatalf("expected running again, got %+v", ac)
|
||||
@@ -267,7 +298,7 @@ func TestAutoCycleTimeout(t *testing.T) {
|
||||
}
|
||||
|
||||
t0 := db.Now()
|
||||
o.autoCycleStep(ctx, t0)
|
||||
stepStartCycle(t, o, t0)
|
||||
|
||||
o.autoCycleStep(ctx, t0.Add(300*time.Second))
|
||||
if ac := getAutoCycle(t, d); ac.Phase != db.AutoCyclePhaseRunning {
|
||||
@@ -302,7 +333,7 @@ func TestAutoCycleNoLimitNeverTimesOut(t *testing.T) {
|
||||
t.Fatalf("start: %v", err)
|
||||
}
|
||||
t0 := db.Now()
|
||||
o.autoCycleStep(ctx, t0)
|
||||
stepStartCycle(t, o, t0)
|
||||
o.autoCycleStep(ctx, t0.Add(1000*time.Hour))
|
||||
if ac := getAutoCycle(t, d); ac.Phase != db.AutoCyclePhaseRunning {
|
||||
t.Fatalf("max_run_seconds=0 means no limit, got %+v", ac)
|
||||
@@ -320,7 +351,7 @@ func TestAutoCycleManualClearMidCycleCompletes(t *testing.T) {
|
||||
t.Fatalf("start: %v", err)
|
||||
}
|
||||
t0 := db.Now()
|
||||
o.autoCycleStep(ctx, t0)
|
||||
stepStartCycle(t, o, t0)
|
||||
if ac := getAutoCycle(t, d); ac.Phase != db.AutoCyclePhaseRunning {
|
||||
t.Fatalf("expected running after start, got %+v", ac)
|
||||
}
|
||||
@@ -350,7 +381,7 @@ func TestAutoCycleNoFreeIPs(t *testing.T) {
|
||||
}
|
||||
|
||||
now := db.Now()
|
||||
o.autoCycleStep(ctx, now)
|
||||
stepStartCycle(t, o, now)
|
||||
|
||||
ac := getAutoCycle(t, d)
|
||||
if ac.Phase != db.AutoCyclePhaseWaiting || ac.LastOutcome != db.AutoCycleOutcomeNoFreeIPs {
|
||||
@@ -378,7 +409,7 @@ func TestAutoCycleOpenStackErrorRetriesNextInterval(t *testing.T) {
|
||||
}
|
||||
|
||||
t0 := db.Now()
|
||||
o.autoCycleStep(ctx, t0)
|
||||
stepStartCycle(t, o, t0)
|
||||
|
||||
ac := getAutoCycle(t, d)
|
||||
if ac.Phase != db.AutoCyclePhaseWaiting || ac.LastOutcome != db.AutoCycleOutcomeError {
|
||||
@@ -400,7 +431,7 @@ func TestAutoCycleOpenStackErrorRetriesNextInterval(t *testing.T) {
|
||||
// OpenStack recovers: the next interval starts a normal cycle and the
|
||||
// stale error is cleared.
|
||||
mock.ListFailure = nil
|
||||
o.autoCycleStep(ctx, t0.Add(60*time.Second))
|
||||
stepStartCycle(t, o, t0.Add(60*time.Second))
|
||||
ac = getAutoCycle(t, d)
|
||||
if ac.Phase != db.AutoCyclePhaseRunning || ac.LastError != "" {
|
||||
t.Fatalf("expected running with cleared error, got %+v", ac)
|
||||
@@ -415,7 +446,7 @@ func TestAutoCycleStartIsIdempotent(t *testing.T) {
|
||||
t.Fatalf("start: %v", err)
|
||||
}
|
||||
t0 := db.Now()
|
||||
o.autoCycleStep(ctx, t0)
|
||||
stepStartCycle(t, o, t0)
|
||||
|
||||
if err := o.StartAutoCycle(ctx); err != nil {
|
||||
t.Fatalf("second start: %v", err)
|
||||
@@ -440,7 +471,7 @@ func TestAutoCycleStopMidCycle(t *testing.T) {
|
||||
t.Fatalf("start: %v", err)
|
||||
}
|
||||
t0 := db.Now()
|
||||
o.autoCycleStep(ctx, t0)
|
||||
stepStartCycle(t, o, t0)
|
||||
o.Tick(ctx) // check in flight
|
||||
|
||||
if err := o.StopAutoCycle(ctx); err != nil {
|
||||
@@ -484,7 +515,7 @@ func TestAutoCycleStopWhileWaitingKeepsLastOutcome(t *testing.T) {
|
||||
t.Fatalf("start: %v", err)
|
||||
}
|
||||
t0 := db.Now()
|
||||
o.autoCycleStep(ctx, t0)
|
||||
stepStartCycle(t, o, t0)
|
||||
finishAllIPs(t, d, db.IPDone)
|
||||
o.autoCycleStep(ctx, t0.Add(5*time.Second))
|
||||
if ac := getAutoCycle(t, d); ac.Phase != db.AutoCyclePhaseWaiting || ac.LastOutcome != db.AutoCycleOutcomeCompleted {
|
||||
@@ -512,7 +543,7 @@ func TestAutoCycleSurvivesRestart(t *testing.T) {
|
||||
t.Fatalf("start: %v", err)
|
||||
}
|
||||
t0 := db.Now()
|
||||
o.autoCycleStep(ctx, t0)
|
||||
stepStartCycle(t, o, t0)
|
||||
|
||||
// A fresh Orchestrator on the same database (a control-api restart)
|
||||
// continues the running phase instead of starting over.
|
||||
@@ -537,8 +568,256 @@ func TestAutoCycleSurvivesRestart(t *testing.T) {
|
||||
if ac := getAutoCycle(t, d); ac.Phase != db.AutoCyclePhaseWaiting {
|
||||
t.Fatalf("expected waiting until next_run_at, got %+v", ac)
|
||||
}
|
||||
o3.autoCycleStep(ctx, t1.Add(300*time.Second))
|
||||
stepStartCycle(t, o3, t1.Add(300*time.Second))
|
||||
if ac := getAutoCycle(t, d); ac.Phase != db.AutoCyclePhaseRunning {
|
||||
t.Fatalf("expected a new cycle at next_run_at, got %+v", ac)
|
||||
}
|
||||
}
|
||||
|
||||
func TestAutoCycleScanningPhaseThenRunning(t *testing.T) {
|
||||
ctx := context.Background()
|
||||
o, d, mock := newTestOrchestrator(t, 180)
|
||||
mock.SeedMany("fip", 30)
|
||||
mock.PageSize = 10
|
||||
mock.PageDelay = 100 * time.Millisecond
|
||||
setAutoCycleParams(t, d, 60, 300)
|
||||
if err := d.SeedQueue(ctx, []string{"9.9.9.9"}); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if err := o.StartAutoCycle(ctx); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
|
||||
t0 := db.Now()
|
||||
o.autoCycleStep(ctx, t0)
|
||||
ac := getAutoCycle(t, d)
|
||||
if ac.Phase != db.AutoCyclePhaseScanning || ac.RunStartedAt != nil || ac.NextRunAt != nil {
|
||||
t.Fatalf("expected scanning without run_started_at, got %+v", ac)
|
||||
}
|
||||
if ac.LastRunStartedAt == nil || !ac.LastRunStartedAt.Equal(t0) {
|
||||
t.Fatalf("expected last_run_started_at=%v, got %v", t0, ac.LastRunStartedAt)
|
||||
}
|
||||
if !o.ScanStatus().Running {
|
||||
t.Fatalf("expected the scan job to be running")
|
||||
}
|
||||
|
||||
// While the job runs, further steps leave the phase alone, even far past
|
||||
// max_run_seconds (it only counts from the end of the scan).
|
||||
o.autoCycleStep(ctx, t0.Add(time.Hour))
|
||||
if ac := getAutoCycle(t, d); ac.Phase != db.AutoCyclePhaseScanning || ac.LastOutcome != "" {
|
||||
t.Fatalf("expected still scanning, got %+v", ac)
|
||||
}
|
||||
|
||||
waitScan(t, o)
|
||||
t1 := t0.Add(2 * time.Hour)
|
||||
o.autoCycleStep(ctx, t1)
|
||||
ac = getAutoCycle(t, d)
|
||||
if ac.Phase != db.AutoCyclePhaseRunning || ac.LastScannedFree != 30 {
|
||||
t.Fatalf("expected running with last_scanned_free=30, got %+v", ac)
|
||||
}
|
||||
if ac.RunStartedAt == nil || !ac.RunStartedAt.Equal(t1) {
|
||||
t.Fatalf("run_started_at must be the scan end %v, got %v", t1, ac.RunStartedAt)
|
||||
}
|
||||
if got := queuedAddresses(t, d); len(got) != 30 {
|
||||
t.Fatalf("expected the 30 scanned addresses only, got %d", len(got))
|
||||
}
|
||||
// max_run_seconds is measured from t1.
|
||||
o.autoCycleStep(ctx, t1.Add(300*time.Second))
|
||||
if ac := getAutoCycle(t, d); ac.Phase != db.AutoCyclePhaseRunning {
|
||||
t.Fatalf("expected still running at the limit, got %+v", ac)
|
||||
}
|
||||
o.autoCycleStep(ctx, t1.Add(301*time.Second))
|
||||
if ac := getAutoCycle(t, d); ac.LastOutcome != db.AutoCycleOutcomeTimeout {
|
||||
t.Fatalf("expected timeout, got %+v", ac)
|
||||
}
|
||||
}
|
||||
|
||||
func TestAutoCycleScanErrorSurfacesAndRetries(t *testing.T) {
|
||||
ctx := context.Background()
|
||||
o, d, mock := newTestOrchestrator(t, 180)
|
||||
mock.SeedMany("fip", 5)
|
||||
mock.ListFailure = errors.New("neutron is down")
|
||||
setAutoCycleParams(t, d, 60, 0)
|
||||
if err := o.StartAutoCycle(ctx); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
t0 := db.Now()
|
||||
o.autoCycleStep(ctx, t0)
|
||||
if ac := getAutoCycle(t, d); ac.Phase != db.AutoCyclePhaseScanning {
|
||||
t.Fatalf("expected scanning first, got %+v", ac)
|
||||
}
|
||||
waitScan(t, o)
|
||||
t1 := t0.Add(time.Second)
|
||||
o.autoCycleStep(ctx, t1)
|
||||
ac := getAutoCycle(t, d)
|
||||
if ac.Phase != db.AutoCyclePhaseWaiting || ac.LastOutcome != db.AutoCycleOutcomeError ||
|
||||
!strings.Contains(ac.LastError, "neutron is down") || ac.NextRunAt == nil || !ac.NextRunAt.Equal(t1.Add(time.Minute)) {
|
||||
t.Fatalf("expected waiting/error with next_run_at=t1+interval, got %+v", ac)
|
||||
}
|
||||
if countEvents(t, d, "auto_cycle_error") != 1 {
|
||||
t.Fatalf("expected one auto_cycle_error event")
|
||||
}
|
||||
}
|
||||
|
||||
func TestAutoCycleRestartInScanningPhaseStartsScanAgain(t *testing.T) {
|
||||
ctx := context.Background()
|
||||
o, d, mock := newTestOrchestrator(t, 180)
|
||||
mock.SeedMany("fip", 12)
|
||||
if err := d.SeedQueue(ctx, []string{"9.9.9.9"}); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
setAutoCycleParams(t, d, 60, 0)
|
||||
if err := o.StartAutoCycle(ctx); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
t0 := db.Now()
|
||||
// Persisted state of a process that died mid-scan.
|
||||
if err := d.UpdateAutoCycleState(ctx, db.AutoCycleState{Phase: db.AutoCyclePhaseScanning, LastRunStartedAt: &t0}); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
|
||||
o2 := &Orchestrator{DB: d, OS: mock, Cfg: o.Cfg, Agg: o.Agg, Log: o.Log}
|
||||
t.Cleanup(func() { o2.CancelScan() })
|
||||
if st := o2.ScanStatus(); st.State != ScanIdle {
|
||||
t.Fatalf("fresh process must have no scan job, got %+v", st)
|
||||
}
|
||||
o2.autoCycleStep(ctx, t0.Add(time.Second))
|
||||
if ac := getAutoCycle(t, d); ac.Phase != db.AutoCyclePhaseScanning {
|
||||
t.Fatalf("expected to stay in scanning while the new job runs, got %+v", ac)
|
||||
}
|
||||
if st := o2.ScanStatus(); st.State == ScanIdle {
|
||||
t.Fatalf("expected a new scan job to be started")
|
||||
}
|
||||
waitScan(t, o2)
|
||||
t1 := t0.Add(2 * time.Second)
|
||||
o2.autoCycleStep(ctx, t1)
|
||||
ac := getAutoCycle(t, d)
|
||||
if ac.Phase != db.AutoCyclePhaseRunning || ac.LastScannedFree != 12 || ac.RunStartedAt == nil || !ac.RunStartedAt.Equal(t1) {
|
||||
t.Fatalf("expected running after the restarted scan, got %+v", ac)
|
||||
}
|
||||
got := queuedAddresses(t, d)
|
||||
if _, stale := got["9.9.9.9"]; stale || len(got) != 12 {
|
||||
t.Fatalf("restarted scan must clear and re-scan, got %d rows (stale=%v)", len(got), stale)
|
||||
}
|
||||
}
|
||||
|
||||
func TestAutoCycleStopCancelsScan(t *testing.T) {
|
||||
ctx := context.Background()
|
||||
o, d, mock := newTestOrchestrator(t, 180)
|
||||
mock.SeedMany("fip", 30)
|
||||
mock.PageSize = 10
|
||||
mock.PageDelay = 5 * time.Second
|
||||
if err := o.StartAutoCycle(ctx); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
t0 := db.Now()
|
||||
o.autoCycleStep(ctx, t0)
|
||||
if ac := getAutoCycle(t, d); ac.Phase != db.AutoCyclePhaseScanning || !o.ScanStatus().Running {
|
||||
t.Fatalf("precondition: scanning with a live job, got %+v", ac)
|
||||
}
|
||||
|
||||
if err := o.StopAutoCycle(ctx); err != nil {
|
||||
t.Fatalf("stop: %v", err)
|
||||
}
|
||||
ac := getAutoCycle(t, d)
|
||||
if ac.Enabled || ac.Phase != db.AutoCyclePhaseIdle || ac.LastOutcome != db.AutoCycleOutcomeStopped {
|
||||
t.Fatalf("expected disabled/idle/stopped, got %+v", ac)
|
||||
}
|
||||
if st := o.ScanStatus(); st.State != ScanCancelled || st.Running {
|
||||
t.Fatalf("expected the scan to be cancelled, got %+v", st)
|
||||
}
|
||||
// Later steps do nothing.
|
||||
o.autoCycleStep(ctx, t0.Add(time.Hour))
|
||||
if ac := getAutoCycle(t, d); ac.Enabled || ac.Phase != db.AutoCyclePhaseIdle {
|
||||
t.Fatalf("expected no activity after stop, got %+v", ac)
|
||||
}
|
||||
}
|
||||
|
||||
// A slow scan must neither block the control loop (Tick, auto-cycle steps)
|
||||
// nor Start/Stop/Get on the auto-cycle: autoCycleMu is never held across it.
|
||||
func TestSlowScanDoesNotBlockTickOrAutoCycleCalls(t *testing.T) {
|
||||
ctx := context.Background()
|
||||
o, d, mock := newTestOrchestrator(t, 180)
|
||||
mock.SeedMany("fip", 20)
|
||||
mock.PageSize = 10
|
||||
mock.PageDelay = 1500 * time.Millisecond
|
||||
if err := d.RegisterValidator(ctx, "validator-1", "host-1", "port-1", "v0.1"); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
|
||||
within := func(what string, limit time.Duration, f func()) {
|
||||
t.Helper()
|
||||
start := time.Now()
|
||||
f()
|
||||
if el := time.Since(start); el > limit {
|
||||
t.Fatalf("%s took %v while a scan was running (limit %v)", what, el, limit)
|
||||
}
|
||||
}
|
||||
|
||||
within("StartAutoCycle", 500*time.Millisecond, func() {
|
||||
if err := o.StartAutoCycle(ctx); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
})
|
||||
t0 := db.Now()
|
||||
within("first auto-cycle step", 500*time.Millisecond, func() { o.autoCycleStep(ctx, t0) })
|
||||
if !o.ScanStatus().Running {
|
||||
t.Fatalf("scan should be in flight")
|
||||
}
|
||||
within("Tick", 500*time.Millisecond, func() { o.Tick(ctx) })
|
||||
within("polling auto-cycle step", 500*time.Millisecond, func() { o.autoCycleStep(ctx, t0.Add(time.Second)) })
|
||||
within("GetAutoCycle", 500*time.Millisecond, func() {
|
||||
if _, err := o.GetAutoCycle(ctx); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
})
|
||||
within("StopAutoCycle", 3*time.Second, func() {
|
||||
if err := o.StopAutoCycle(ctx); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
})
|
||||
}
|
||||
|
||||
// A scan that is already running (an operator's dry run, a manual or periodic
|
||||
// scan) is NOT a cycle's scan: it may not clear the queue or enqueue anything.
|
||||
// The cycle must wait for it and then run its own clear+scan; following the
|
||||
// foreign job would end in "completed" over an untouched/empty queue.
|
||||
func TestAutoCycleWaitsForForeignScanInsteadOfFollowingIt(t *testing.T) {
|
||||
ctx := context.Background()
|
||||
o, d, mock := newTestOrchestrator(t, 180)
|
||||
mock.Seed("fip-1", "1.1.1.1", "svc")
|
||||
mock.PageSize = 1
|
||||
mock.PageDelay = 80 * time.Millisecond // keeps the dry run in flight for a while
|
||||
if err := d.SeedQueue(ctx, []string{"9.9.9.9"}); err != nil {
|
||||
t.Fatalf("seed queue: %v", err)
|
||||
}
|
||||
setAutoCycleParams(t, d, 60, 0)
|
||||
if err := o.StartAutoCycle(ctx); err != nil {
|
||||
t.Fatalf("start: %v", err)
|
||||
}
|
||||
|
||||
if _, started := o.StartScan(ScanOptions{DryRun: true}); !started {
|
||||
t.Fatalf("precondition: the dry run must start")
|
||||
}
|
||||
t0 := db.Now()
|
||||
o.autoCycleStep(ctx, t0)
|
||||
ac := getAutoCycle(t, d)
|
||||
if ac.Phase == db.AutoCyclePhaseScanning || ac.Phase == db.AutoCyclePhaseRunning {
|
||||
t.Fatalf("cycle must not adopt the foreign scan, got phase %q", ac.Phase)
|
||||
}
|
||||
if got := queuedAddresses(t, d); got["9.9.9.9"] != db.IPQueued {
|
||||
t.Fatalf("queue must be untouched while the foreign scan runs, got %v", got)
|
||||
}
|
||||
|
||||
waitScan(t, o) // the dry run ends
|
||||
mock.PageDelay = 0
|
||||
stepStartCycle(t, o, t0.Add(time.Second))
|
||||
ac = getAutoCycle(t, d)
|
||||
if ac.Phase != db.AutoCyclePhaseRunning || ac.LastScannedFree != 1 {
|
||||
t.Fatalf("expected the cycle's own scan to complete (running, 1 free), got %+v", ac)
|
||||
}
|
||||
got := queuedAddresses(t, d)
|
||||
if _, stale := got["9.9.9.9"]; stale || got["1.1.1.1"] != db.IPQueued {
|
||||
t.Fatalf("the cycle's own scan must clear the old queue and enqueue 1.1.1.1, got %v", got)
|
||||
}
|
||||
}
|
||||
@@ -41,6 +41,14 @@ type Orchestrator struct {
|
||||
// autoCycleMu serializes AutoCycleStep with StartAutoCycle/StopAutoCycle
|
||||
// so an API call can never interleave with a half-finished step.
|
||||
autoCycleMu sync.Mutex
|
||||
|
||||
// ScanPageSize is the Neutron page size used by the floating-IP scan
|
||||
// (config openstack.list_page_size); 0 means openstack.DefaultListPageSize.
|
||||
ScanPageSize int
|
||||
|
||||
// scan is the background floating-IP scan job (see scanjob.go); its zero
|
||||
// value is ready to use.
|
||||
scan scanJob
|
||||
}
|
||||
|
||||
// New constructs an Orchestrator. Egress check types/targets, prober sites,
|
||||
@@ -58,6 +66,8 @@ func New(d *db.DB, osClient openstack.FloatingIPClient, cfg *config.ControlAPI,
|
||||
Cfg: cfg.Orchestrator,
|
||||
Agg: cfg.Aggregation,
|
||||
Log: log,
|
||||
|
||||
ScanPageSize: cfg.OpenStack.ListPageSize,
|
||||
}
|
||||
}
|
||||
|
||||
@@ -349,17 +359,16 @@ func (o *Orchestrator) DeleteIP(ctx context.Context, ipAddress string) error {
|
||||
// in the list that has one attached, then deletes the whole list in one
|
||||
// DB.DeleteIPs call. Addresses not currently in the queue are simply
|
||||
// omitted from the disassociation pass and reported back in NotFound by
|
||||
// DB.DeleteIPs — not an error.
|
||||
// DB.DeleteIPs — not an error. The rows holding a floating IP are found with
|
||||
// a few IN (...) queries rather than one lookup per address.
|
||||
func (o *Orchestrator) DeleteIPs(ctx context.Context, addresses []string) (db.DeleteIPsResult, error) {
|
||||
for _, addr := range addresses {
|
||||
item, err := o.DB.GetIPByAddress(ctx, addr)
|
||||
refs, err := o.DB.ListFIPRefsByAddresses(ctx, addresses)
|
||||
if err != nil {
|
||||
continue // unknown address — DB.DeleteIPs will report it in NotFound
|
||||
}
|
||||
if item.FIPID != "" {
|
||||
if err := o.OS.DisassociateFloatingIP(ctx, item.FIPID); err != nil {
|
||||
o.Log.Error("disassociate fip on delete", "ip_id", item.ID, "fip_id", item.FIPID, "err", err)
|
||||
o.Log.Error("list attached fips before delete", "err", err)
|
||||
}
|
||||
for _, ref := range refs {
|
||||
if err := o.OS.DisassociateFloatingIP(ctx, ref.FIPID); err != nil {
|
||||
o.Log.Error("disassociate fip on delete", "ip_id", ref.IPID, "fip_id", ref.FIPID, "err", err)
|
||||
}
|
||||
}
|
||||
|
||||
@@ -372,79 +381,51 @@ func (o *Orchestrator) DeleteIPs(ctx context.Context, addresses []string) (db.De
|
||||
}
|
||||
|
||||
// ClearQueue deletes every address currently in the queue, regardless of
|
||||
// state — the "delete everything" operation, implemented as DeleteIPs over
|
||||
// the full current address list rather than a separate DB code path.
|
||||
// state — the "delete everything" operation. It is set-based (see
|
||||
// db.ClearAllIPs): O(1) statements however many rows there are. Floating IPs
|
||||
// attached to rows are disassociated first (best-effort), and only for rows
|
||||
// that actually hold one.
|
||||
func (o *Orchestrator) ClearQueue(ctx context.Context) (db.DeleteIPsResult, error) {
|
||||
items, err := o.DB.ListIPs(ctx)
|
||||
refs, err := o.DB.ListFIPRefs(ctx)
|
||||
if err != nil {
|
||||
return db.DeleteIPsResult{}, fmt.Errorf("list ips: %w", err)
|
||||
}
|
||||
addresses := make([]string, len(items))
|
||||
for i, item := range items {
|
||||
addresses[i] = item.IPAddress
|
||||
}
|
||||
|
||||
for _, item := range items {
|
||||
if item.FIPID != "" {
|
||||
if err := o.OS.DisassociateFloatingIP(ctx, item.FIPID); err != nil {
|
||||
o.Log.Error("disassociate fip on clear queue", "ip_id", item.ID, "fip_id", item.FIPID, "err", err)
|
||||
return db.DeleteIPsResult{}, fmt.Errorf("list attached fips: %w", err)
|
||||
}
|
||||
for _, ref := range refs {
|
||||
if err := o.OS.DisassociateFloatingIP(ctx, ref.FIPID); err != nil {
|
||||
o.Log.Error("disassociate fip on clear queue", "ip_id", ref.IPID, "fip_id", ref.FIPID, "err", err)
|
||||
}
|
||||
}
|
||||
|
||||
result, err := o.DB.DeleteIPs(ctx, addresses)
|
||||
deleted, err := o.DB.ClearAllIPs(ctx)
|
||||
if err != nil {
|
||||
return result, err
|
||||
return db.DeleteIPsResult{}, err
|
||||
}
|
||||
o.event(ctx, "control-api", "", nil, "queue_cleared", deletedAddressesPayload(result.Deleted))
|
||||
return result, nil
|
||||
o.event(ctx, "control-api", "", nil, "queue_cleared", deletedAddressesPayload(deleted))
|
||||
return db.DeleteIPsResult{Deleted: deleted}, nil
|
||||
}
|
||||
|
||||
// ScanFloatingIPs lists every floating IP in the OpenStack project, filters
|
||||
// to the ones not currently associated to any port (the free pool awaiting
|
||||
// validation before reissue), and submits that address list to the check
|
||||
// queue via db.SubmitIPs — the same entry point the admin API's "add
|
||||
// addresses" call uses, so add/requeue/reorder semantics are identical
|
||||
// whether the address list came from an operator or from this scan. Returns
|
||||
// the SubmitIPs outcome plus how many free floating IPs were found in total
|
||||
// (which can be larger than the sum of the SubmitIPsResult slices, since
|
||||
// addresses already mid-check are silently skipped — see db.SubmitIPs).
|
||||
func (o *Orchestrator) ScanFloatingIPs(ctx context.Context) (db.SubmitIPsResult, int, error) {
|
||||
fips, err := o.OS.ListFloatingIPs(ctx)
|
||||
if err != nil {
|
||||
return db.SubmitIPsResult{}, 0, fmt.Errorf("list floating ips: %w", err)
|
||||
}
|
||||
|
||||
var free []string
|
||||
for _, f := range fips {
|
||||
if f.PortID == "" {
|
||||
free = append(free, f.Address)
|
||||
}
|
||||
}
|
||||
|
||||
if len(free) == 0 {
|
||||
o.event(ctx, "control-api", "", nil, "fip_scan", `{"scanned_free":0}`)
|
||||
return db.SubmitIPsResult{}, 0, nil
|
||||
}
|
||||
|
||||
result, err := o.DB.SubmitIPs(ctx, free)
|
||||
if err != nil {
|
||||
return result, len(free), fmt.Errorf("submit scanned ips: %w", err)
|
||||
}
|
||||
o.event(ctx, "control-api", "", nil, "fip_scan", fmt.Sprintf(
|
||||
`{"scanned_free":%d,"added":%d,"requeued":%d,"reordered":%d,"skipped_in_progress":%d}`,
|
||||
len(free), len(result.Added), len(result.Requeued), len(result.Reordered), len(result.SkippedInProgress)))
|
||||
return result, len(free), nil
|
||||
}
|
||||
// maxEventAddresses caps how many addresses a batch-delete event payload
|
||||
// lists: clearing 6440 addresses must not write a 100 KB event row.
|
||||
const maxEventAddresses = 50
|
||||
|
||||
// deletedAddressesPayload builds the event payload for the batch delete
|
||||
// operations — a proper JSON array via encoding/json rather than fmt's %q
|
||||
// slice formatting (which produces space-separated quoted strings, not
|
||||
// valid JSON).
|
||||
// operations — a proper JSON object via encoding/json: the total count plus
|
||||
// at most the first maxEventAddresses addresses (and "truncated":true when
|
||||
// the list was cut).
|
||||
func deletedAddressesPayload(addresses []string) string {
|
||||
b, err := json.Marshal(struct {
|
||||
p := struct {
|
||||
Count int `json:"count"`
|
||||
Addresses []string `json:"addresses"`
|
||||
}{addresses})
|
||||
Truncated bool `json:"truncated,omitempty"`
|
||||
}{Count: len(addresses), Addresses: addresses}
|
||||
if p.Addresses == nil {
|
||||
p.Addresses = []string{}
|
||||
}
|
||||
if len(p.Addresses) > maxEventAddresses {
|
||||
p.Addresses = p.Addresses[:maxEventAddresses]
|
||||
p.Truncated = true
|
||||
}
|
||||
b, err := json.Marshal(p)
|
||||
if err != nil {
|
||||
return "{}"
|
||||
}
|
||||
|
||||
@@ -64,7 +64,10 @@ func newTestOrchestratorWithSites(t *testing.T, leaseTTLSeconds int, sites []con
|
||||
}
|
||||
|
||||
log := slog.New(slog.NewTextHandler(os.Stderr, &slog.HandlerOptions{Level: slog.LevelError}))
|
||||
return New(d, mock, cfg, log), d, mock
|
||||
o := New(d, mock, cfg, log)
|
||||
// A background scan must never outlive the database it writes to.
|
||||
t.Cleanup(func() { o.CancelScan() })
|
||||
return o, d, mock
|
||||
}
|
||||
|
||||
func TestHappyPath(t *testing.T) {
|
||||
|
||||
@@ -0,0 +1,370 @@
|
||||
package orchestrator
|
||||
|
||||
import (
|
||||
"context"
|
||||
"errors"
|
||||
"fmt"
|
||||
"net/netip"
|
||||
"sort"
|
||||
"sync"
|
||||
"time"
|
||||
|
||||
"cloudipvalidator/internal/db"
|
||||
"cloudipvalidator/internal/openstack"
|
||||
)
|
||||
|
||||
// This file implements the background floating-IP scan job. With thousands of
|
||||
// floating IPs a scan takes tens of seconds (Neutron is read page by page)
|
||||
// and then enqueues thousands of rows, so it can neither run inside an HTTP
|
||||
// request nor under the auto-cycle mutex. StartScan launches it on the
|
||||
// process-lifetime context and returns immediately; the job publishes its
|
||||
// progress as a ScanStatus that anyone can poll.
|
||||
|
||||
// ScanState is the phase of the scan job.
|
||||
type ScanState string
|
||||
|
||||
const (
|
||||
ScanIdle ScanState = "idle" // no scan has run in this process yet
|
||||
ScanClearing ScanState = "clearing"
|
||||
ScanListing ScanState = "listing"
|
||||
ScanEnqueuing ScanState = "enqueuing"
|
||||
ScanDone ScanState = "done"
|
||||
ScanError ScanState = "error"
|
||||
ScanCancelled ScanState = "cancelled"
|
||||
)
|
||||
|
||||
const (
|
||||
// scanChunkSize is how many addresses go into one db.SubmitIPs call, i.e.
|
||||
// one short transaction; the single DB connection is released in between
|
||||
// so the orchestrator tick and the API stay responsive.
|
||||
scanChunkSize = 500
|
||||
// defaultScanTimeout is used when Cfg.FIPScanTimeoutSeconds is zero.
|
||||
defaultScanTimeout = 1800 * time.Second
|
||||
// cancelWait bounds how long CancelScan waits for the job to wind down.
|
||||
cancelWait = 5 * time.Second
|
||||
)
|
||||
|
||||
// ScanOptions selects the variant of a scan.
|
||||
type ScanOptions struct {
|
||||
// ClearFirst clears the whole queue before scanning (auto-cycle step 1).
|
||||
// Ignored together with DryRun: a dry run never touches the queue.
|
||||
ClearFirst bool
|
||||
// DryRun only discovers and counts the free floating IPs; the queue is
|
||||
// left untouched.
|
||||
DryRun bool
|
||||
}
|
||||
|
||||
// ScanStatus is a snapshot of the scan job's progress.
|
||||
type ScanStatus struct {
|
||||
State ScanState
|
||||
Running bool
|
||||
DryRun bool
|
||||
Pages int // Neutron pages read so far
|
||||
Discovered int // floating IPs seen (free and associated)
|
||||
Free int // of those, free (no port) — the addresses to enqueue
|
||||
Added int
|
||||
Requeued int
|
||||
Reordered int
|
||||
SkippedInProgress int
|
||||
StartedAt *time.Time
|
||||
FinishedAt *time.Time
|
||||
Error string
|
||||
}
|
||||
|
||||
// scanResult is what a finished job leaves for the synchronous wrapper.
|
||||
type scanResult struct {
|
||||
submit db.SubmitIPsResult
|
||||
free int
|
||||
err error
|
||||
}
|
||||
|
||||
// scanRun is one execution of the job.
|
||||
type scanRun struct {
|
||||
done chan struct{} // closed when the job has finished (result is set)
|
||||
cancel context.CancelFunc
|
||||
result scanResult
|
||||
}
|
||||
|
||||
// scanJob is the zero-value-usable scan state embedded in Orchestrator
|
||||
// (tests build Orchestrator as a literal, so there is no constructor-only
|
||||
// initialization).
|
||||
type scanJob struct {
|
||||
mu sync.Mutex
|
||||
lifeCtx context.Context // process lifetime; nil = context.Background()
|
||||
status ScanStatus // zero value reads as idle (see snapshot)
|
||||
run *scanRun // current/last run
|
||||
cancelRequested bool
|
||||
}
|
||||
|
||||
// SetContext sets the lifetime context background jobs (the scan) run on.
|
||||
// Call once at startup; until then context.Background() is used. Cancelling
|
||||
// it cancels a running scan.
|
||||
func (o *Orchestrator) SetContext(ctx context.Context) {
|
||||
o.scan.mu.Lock()
|
||||
defer o.scan.mu.Unlock()
|
||||
o.scan.lifeCtx = ctx
|
||||
}
|
||||
|
||||
func (j *scanJob) lifetime() context.Context {
|
||||
if j.lifeCtx != nil {
|
||||
return j.lifeCtx
|
||||
}
|
||||
return context.Background()
|
||||
}
|
||||
|
||||
// snapshot returns a copy of the status; caller holds j.mu.
|
||||
func (j *scanJob) snapshot() ScanStatus {
|
||||
st := j.status
|
||||
if st.State == "" {
|
||||
st.State = ScanIdle
|
||||
}
|
||||
return st
|
||||
}
|
||||
|
||||
// ScanStatus returns the current scan status (state "idle" if no scan has run
|
||||
// in this process yet).
|
||||
func (o *Orchestrator) ScanStatus() ScanStatus {
|
||||
o.scan.mu.Lock()
|
||||
defer o.scan.mu.Unlock()
|
||||
return o.scan.snapshot()
|
||||
}
|
||||
|
||||
// StartScan starts the background scan job, or — single-flight — joins the one
|
||||
// already running: then started is false and the returned status is the
|
||||
// running job's. It never blocks on OpenStack or the database.
|
||||
func (o *Orchestrator) StartScan(opts ScanOptions) (ScanStatus, bool) {
|
||||
st, _, started := o.startScan(opts)
|
||||
return st, started
|
||||
}
|
||||
|
||||
func (o *Orchestrator) startScan(opts ScanOptions) (ScanStatus, *scanRun, bool) {
|
||||
j := &o.scan
|
||||
j.mu.Lock()
|
||||
defer j.mu.Unlock()
|
||||
if j.status.Running && j.run != nil {
|
||||
return j.snapshot(), j.run, false
|
||||
}
|
||||
|
||||
timeout := time.Duration(o.Cfg.FIPScanTimeoutSeconds) * time.Second
|
||||
if timeout <= 0 {
|
||||
timeout = defaultScanTimeout
|
||||
}
|
||||
ctx, cancel := context.WithTimeout(j.lifetime(), timeout)
|
||||
run := &scanRun{done: make(chan struct{}), cancel: cancel}
|
||||
if opts.DryRun {
|
||||
opts.ClearFirst = false
|
||||
}
|
||||
now := db.Now()
|
||||
state := ScanListing
|
||||
if opts.ClearFirst {
|
||||
state = ScanClearing
|
||||
}
|
||||
j.run = run
|
||||
j.cancelRequested = false
|
||||
j.status = ScanStatus{State: state, Running: true, DryRun: opts.DryRun, StartedAt: &now}
|
||||
go o.runScan(ctx, run, opts)
|
||||
return j.snapshot(), run, true
|
||||
}
|
||||
|
||||
// CancelScan cancels a running scan and waits (briefly) for it to wind down.
|
||||
// It reports whether a running scan was cancelled. Already-enqueued chunks
|
||||
// stay in the queue.
|
||||
func (o *Orchestrator) CancelScan() bool {
|
||||
j := &o.scan
|
||||
j.mu.Lock()
|
||||
run := j.run
|
||||
if !j.status.Running || run == nil {
|
||||
j.mu.Unlock()
|
||||
return false
|
||||
}
|
||||
j.cancelRequested = true
|
||||
j.mu.Unlock()
|
||||
|
||||
run.cancel()
|
||||
select {
|
||||
case <-run.done:
|
||||
case <-time.After(cancelWait):
|
||||
}
|
||||
return true
|
||||
}
|
||||
|
||||
// update mutates the status under the lock.
|
||||
func (j *scanJob) update(f func(*ScanStatus)) {
|
||||
j.mu.Lock()
|
||||
defer j.mu.Unlock()
|
||||
f(&j.status)
|
||||
}
|
||||
|
||||
func (o *Orchestrator) runScan(ctx context.Context, run *scanRun, opts ScanOptions) {
|
||||
j := &o.scan
|
||||
defer run.cancel()
|
||||
|
||||
res, err := o.doScan(ctx, opts)
|
||||
run.result = res
|
||||
|
||||
j.mu.Lock()
|
||||
finished := db.Now()
|
||||
switch {
|
||||
case err == nil:
|
||||
j.status.State = ScanDone
|
||||
case j.cancelRequested || errors.Is(j.lifetime().Err(), context.Canceled):
|
||||
j.status.State = ScanCancelled
|
||||
err = fmt.Errorf("scan cancelled: %w", context.Canceled)
|
||||
j.status.Error = "cancelled"
|
||||
case errors.Is(err, context.DeadlineExceeded):
|
||||
j.status.State = ScanError
|
||||
err = fmt.Errorf("scan timed out: %w", err)
|
||||
j.status.Error = err.Error()
|
||||
default:
|
||||
j.status.State = ScanError
|
||||
j.status.Error = err.Error()
|
||||
}
|
||||
j.status.Running = false
|
||||
j.status.FinishedAt = &finished
|
||||
state := j.status.State
|
||||
j.mu.Unlock()
|
||||
run.result.err = err
|
||||
|
||||
if state == ScanError || state == ScanCancelled {
|
||||
o.Log.Warn("floating ip scan did not complete", "state", state, "err", err)
|
||||
}
|
||||
close(run.done)
|
||||
}
|
||||
|
||||
// doScan is the scan algorithm: optional clear -> read every page into memory
|
||||
// -> sort -> (unless dry run) enqueue in chunks -> one fip_scan event.
|
||||
func (o *Orchestrator) doScan(ctx context.Context, opts ScanOptions) (scanResult, error) {
|
||||
j := &o.scan
|
||||
var res scanResult
|
||||
|
||||
if opts.ClearFirst {
|
||||
if _, err := o.ClearQueue(ctx); err != nil {
|
||||
return res, fmt.Errorf("clear queue: %w", err)
|
||||
}
|
||||
j.update(func(s *ScanStatus) { s.State = ScanListing })
|
||||
}
|
||||
|
||||
// Read everything first: a read error after retries must leave the queue
|
||||
// untouched, so nothing is enqueued until discovery is complete.
|
||||
var free []string
|
||||
seen := map[string]struct{}{}
|
||||
pages, err := o.OS.ListFreeFloatingIPs(ctx, o.ScanPageSize, func(page []openstack.FloatingIP) error {
|
||||
nFree := 0
|
||||
for _, f := range page {
|
||||
if f.PortID != "" || f.Address == "" {
|
||||
continue
|
||||
}
|
||||
if _, dup := seen[f.Address]; dup {
|
||||
continue
|
||||
}
|
||||
seen[f.Address] = struct{}{}
|
||||
free = append(free, f.Address)
|
||||
nFree++
|
||||
}
|
||||
j.update(func(s *ScanStatus) {
|
||||
s.Pages++
|
||||
s.Discovered += len(page)
|
||||
s.Free += nFree
|
||||
})
|
||||
return ctx.Err()
|
||||
})
|
||||
if err != nil {
|
||||
return res, fmt.Errorf("list floating ips: %w", err)
|
||||
}
|
||||
j.update(func(s *ScanStatus) { s.Pages = pages })
|
||||
res.free = len(free)
|
||||
sortAddressesAscending(free)
|
||||
|
||||
if !opts.DryRun && len(free) > 0 {
|
||||
j.update(func(s *ScanStatus) { s.State = ScanEnqueuing })
|
||||
for off := 0; off < len(free); off += scanChunkSize {
|
||||
if err := ctx.Err(); err != nil {
|
||||
return res, err
|
||||
}
|
||||
chunk := free[off:min(off+scanChunkSize, len(free))]
|
||||
r, err := o.DB.SubmitIPs(ctx, chunk)
|
||||
res.submit.Added = append(res.submit.Added, r.Added...)
|
||||
res.submit.Requeued = append(res.submit.Requeued, r.Requeued...)
|
||||
res.submit.Reordered = append(res.submit.Reordered, r.Reordered...)
|
||||
res.submit.SkippedInProgress = append(res.submit.SkippedInProgress, r.SkippedInProgress...)
|
||||
j.update(func(s *ScanStatus) {
|
||||
s.Added += len(r.Added)
|
||||
s.Requeued += len(r.Requeued)
|
||||
s.Reordered += len(r.Reordered)
|
||||
s.SkippedInProgress += len(r.SkippedInProgress)
|
||||
})
|
||||
if err != nil {
|
||||
return res, fmt.Errorf("submit scanned ips: %w", err)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
if !opts.DryRun {
|
||||
o.event(ctx, "control-api", "", nil, "fip_scan", fmt.Sprintf(
|
||||
`{"scanned_free":%d,"pages":%d,"added":%d,"requeued":%d,"reordered":%d,"skipped_in_progress":%d}`,
|
||||
len(free), pages, len(res.submit.Added), len(res.submit.Requeued),
|
||||
len(res.submit.Reordered), len(res.submit.SkippedInProgress)))
|
||||
}
|
||||
return res, nil
|
||||
}
|
||||
|
||||
// sortAddressesAscending orders addresses numerically (10.0.0.2 before
|
||||
// 10.0.0.10) so the queue order is deterministic; anything that does not parse
|
||||
// as an IP goes last, in string order.
|
||||
func sortAddressesAscending(addrs []string) {
|
||||
type keyed struct {
|
||||
s string
|
||||
ip netip.Addr
|
||||
ok bool
|
||||
}
|
||||
ks := make([]keyed, len(addrs))
|
||||
for i, a := range addrs {
|
||||
ip, err := netip.ParseAddr(a)
|
||||
ks[i] = keyed{s: a, ip: ip, ok: err == nil}
|
||||
}
|
||||
sort.SliceStable(ks, func(i, j int) bool {
|
||||
a, b := ks[i], ks[j]
|
||||
switch {
|
||||
case a.ok && b.ok:
|
||||
return a.ip.Compare(b.ip) < 0
|
||||
case a.ok != b.ok:
|
||||
return a.ok
|
||||
default:
|
||||
return a.s < b.s
|
||||
}
|
||||
})
|
||||
for i := range ks {
|
||||
addrs[i] = ks[i].s
|
||||
}
|
||||
}
|
||||
|
||||
// ScanFloatingIPs lists every floating IP in the OpenStack project, filters
|
||||
// to the ones not currently associated to any port (the free pool awaiting
|
||||
// validation before reissue), and submits that address list to the check
|
||||
// queue via db.SubmitIPs — the same entry point the admin API's "add
|
||||
// addresses" call uses, so add/requeue/reorder semantics are identical
|
||||
// whether the address list came from an operator or from this scan. It is the
|
||||
// synchronous "start (or join) the background scan and wait for it" wrapper:
|
||||
// returns the aggregated SubmitIPs outcome plus how many free floating IPs
|
||||
// were found in total (which can be larger than the sum of the
|
||||
// SubmitIPsResult slices, since addresses already mid-check are silently
|
||||
// skipped — see db.SubmitIPs). If ctx is cancelled first it returns
|
||||
// ctx.Err(); the background job keeps running.
|
||||
func (o *Orchestrator) ScanFloatingIPs(ctx context.Context) (db.SubmitIPsResult, int, error) {
|
||||
return o.ScanAndWait(ctx, ScanOptions{})
|
||||
}
|
||||
|
||||
// ScanAndWait starts (or joins) the background scan with opts and waits for it
|
||||
// to finish, returning the aggregated SubmitIPs outcome and the number of free
|
||||
// floating IPs found. Joining a scan that is already running returns that
|
||||
// scan's result regardless of opts. If ctx ends first it returns ctx.Err()
|
||||
// and the job keeps running.
|
||||
func (o *Orchestrator) ScanAndWait(ctx context.Context, opts ScanOptions) (db.SubmitIPsResult, int, error) {
|
||||
_, run, _ := o.startScan(opts)
|
||||
select {
|
||||
case <-run.done:
|
||||
case <-ctx.Done():
|
||||
return db.SubmitIPsResult{}, 0, ctx.Err()
|
||||
}
|
||||
return run.result.submit, run.result.free, run.result.err
|
||||
}
|
||||
@@ -0,0 +1,398 @@
|
||||
package orchestrator
|
||||
|
||||
import (
|
||||
"context"
|
||||
"encoding/json"
|
||||
"errors"
|
||||
"net/netip"
|
||||
"strconv"
|
||||
"strings"
|
||||
"testing"
|
||||
"time"
|
||||
|
||||
"cloudipvalidator/internal/db"
|
||||
)
|
||||
|
||||
func eventPayload(t *testing.T, d *db.DB, eventType string) string {
|
||||
t.Helper()
|
||||
var p string
|
||||
if err := d.QueryRowContext(context.Background(),
|
||||
`SELECT payload FROM events WHERE event_type=? ORDER BY id DESC LIMIT 1`, eventType).Scan(&p); err != nil {
|
||||
t.Fatalf("read %s event: %v", eventType, err)
|
||||
}
|
||||
return p
|
||||
}
|
||||
|
||||
func TestScanStatusIdleBeforeAnyScan(t *testing.T) {
|
||||
// Zero-value Orchestrator literal (as several tests build it) must work.
|
||||
o := &Orchestrator{}
|
||||
st := o.ScanStatus()
|
||||
if st.State != ScanIdle || st.Running || st.StartedAt != nil {
|
||||
t.Fatalf("expected idle, got %+v", st)
|
||||
}
|
||||
if o.CancelScan() {
|
||||
t.Fatalf("nothing to cancel")
|
||||
}
|
||||
}
|
||||
|
||||
func TestScanJobSingleFlightAndProgress(t *testing.T) {
|
||||
o, d, mock := newTestOrchestrator(t, 180)
|
||||
mock.SeedMany("fip", 600)
|
||||
mock.SeedWithPort("busy", "203.0.113.9", "svc", "port-x")
|
||||
mock.PageSize = 100
|
||||
mock.PageDelay = 40 * time.Millisecond
|
||||
|
||||
st, started := o.StartScan(ScanOptions{})
|
||||
if !started || !st.Running || st.State != ScanListing || st.StartedAt == nil {
|
||||
t.Fatalf("expected a started listing job, got started=%v %+v", started, st)
|
||||
}
|
||||
st2, started2 := o.StartScan(ScanOptions{DryRun: true})
|
||||
if started2 || !st2.Running || st2.DryRun {
|
||||
t.Fatalf("second start must join the running job (not dry-run), got started=%v %+v", started2, st2)
|
||||
}
|
||||
|
||||
// The synchronous wrapper joins the same job instead of scanning twice.
|
||||
type out struct {
|
||||
res db.SubmitIPsResult
|
||||
scanned int
|
||||
err error
|
||||
}
|
||||
ch := make(chan out, 1)
|
||||
go func() {
|
||||
r, n, err := o.ScanFloatingIPs(context.Background())
|
||||
ch <- out{r, n, err}
|
||||
}()
|
||||
|
||||
// Progress is visible while the job runs.
|
||||
sawProgress := false
|
||||
for i := 0; i < 200 && o.ScanStatus().Running; i++ {
|
||||
if s := o.ScanStatus(); s.Pages > 0 && s.Pages < 7 {
|
||||
sawProgress = true
|
||||
}
|
||||
time.Sleep(5 * time.Millisecond)
|
||||
}
|
||||
got := <-ch
|
||||
if got.err != nil || got.scanned != 600 || len(got.res.Added) != 600 {
|
||||
t.Fatalf("wrapper result: %+v", got)
|
||||
}
|
||||
if !sawProgress {
|
||||
t.Fatalf("expected to observe intermediate progress")
|
||||
}
|
||||
|
||||
fin := waitScan(t, o)
|
||||
if fin.State != ScanDone || fin.Running || fin.FinishedAt == nil || fin.Error != "" {
|
||||
t.Fatalf("expected done, got %+v", fin)
|
||||
}
|
||||
if fin.Pages != 7 || fin.Discovered != 601 || fin.Free != 600 || fin.Added != 600 {
|
||||
t.Fatalf("unexpected counters: %+v", fin)
|
||||
}
|
||||
if countEvents(t, d, "fip_scan") != 1 {
|
||||
t.Fatalf("expected exactly one fip_scan event, got %d", countEvents(t, d, "fip_scan"))
|
||||
}
|
||||
var p struct {
|
||||
ScannedFree int `json:"scanned_free"`
|
||||
Pages int `json:"pages"`
|
||||
Added int `json:"added"`
|
||||
}
|
||||
if err := json.Unmarshal([]byte(eventPayload(t, d, "fip_scan")), &p); err != nil || p.ScannedFree != 600 || p.Added != 600 || p.Pages != 7 {
|
||||
t.Fatalf("fip_scan payload: %+v err=%v", p, err)
|
||||
}
|
||||
|
||||
// A new scan after completion starts a fresh job; everything is requeued
|
||||
// or reordered (idempotent), nothing added.
|
||||
mock.PageDelay = 0
|
||||
st3, started3 := o.StartScan(ScanOptions{})
|
||||
if !started3 {
|
||||
t.Fatalf("a finished job must not block a new one: %+v", st3)
|
||||
}
|
||||
fin = waitScan(t, o)
|
||||
if fin.State != ScanDone || fin.Added != 0 || fin.Reordered != 600 {
|
||||
t.Fatalf("rescan: %+v", fin)
|
||||
}
|
||||
}
|
||||
|
||||
func TestScanJobDryRunLeavesQueueUntouched(t *testing.T) {
|
||||
ctx := context.Background()
|
||||
o, d, mock := newTestOrchestrator(t, 180)
|
||||
mock.SeedMany("fip", 450)
|
||||
mock.PageSize = 200
|
||||
if err := d.SeedQueue(ctx, []string{"9.9.9.9"}); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
|
||||
o.StartScan(ScanOptions{DryRun: true, ClearFirst: true}) // ClearFirst is ignored for dry runs
|
||||
st := waitScan(t, o)
|
||||
if st.State != ScanDone || !st.DryRun || st.Free != 450 || st.Pages != 3 || st.Added != 0 {
|
||||
t.Fatalf("dry run status: %+v", st)
|
||||
}
|
||||
if got := queuedAddresses(t, d); len(got) != 1 || got["9.9.9.9"] != db.IPQueued {
|
||||
t.Fatalf("dry run must not touch the queue, got %d rows", len(got))
|
||||
}
|
||||
if countEvents(t, d, "fip_scan") != 0 || countEvents(t, d, "queue_cleared") != 0 {
|
||||
t.Fatalf("dry run must not emit scan/clear events")
|
||||
}
|
||||
}
|
||||
|
||||
func TestScanJobReadErrorLeavesQueueUntouched(t *testing.T) {
|
||||
ctx := context.Background()
|
||||
o, d, mock := newTestOrchestrator(t, 180)
|
||||
mock.SeedMany("fip", 450)
|
||||
mock.PageSize = 200
|
||||
if err := d.SeedQueue(ctx, []string{"9.9.9.9"}); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
// Page 1 succeeds, page 2 fails: the 200 already read must NOT be queued.
|
||||
mock.ListFailures = []error{nil, errors.New("neutron exploded")}
|
||||
|
||||
o.StartScan(ScanOptions{})
|
||||
st := waitScan(t, o)
|
||||
if st.State != ScanError || !strings.Contains(st.Error, "neutron exploded") || st.FinishedAt == nil {
|
||||
t.Fatalf("expected error state, got %+v", st)
|
||||
}
|
||||
if st.Pages != 1 || st.Added != 0 {
|
||||
t.Fatalf("expected 1 page read and nothing added, got %+v", st)
|
||||
}
|
||||
if got := queuedAddresses(t, d); len(got) != 1 || got["9.9.9.9"] != db.IPQueued {
|
||||
t.Fatalf("queue must be untouched after a read error, got %d rows", len(got))
|
||||
}
|
||||
if countEvents(t, d, "fip_scan") != 0 {
|
||||
t.Fatalf("no fip_scan event on failure")
|
||||
}
|
||||
|
||||
// The wrapper reports the same error.
|
||||
mock.ListFailure = errors.New("still down")
|
||||
if _, _, err := o.ScanFloatingIPs(ctx); err == nil || !strings.Contains(err.Error(), "still down") {
|
||||
t.Fatalf("expected wrapped list error, got %v", err)
|
||||
}
|
||||
}
|
||||
|
||||
func TestScanJobClearFirst(t *testing.T) {
|
||||
ctx := context.Background()
|
||||
o, d, mock := newTestOrchestrator(t, 180)
|
||||
mock.Seed("fip-1", "1.1.1.1", "svc")
|
||||
if err := d.SeedQueue(ctx, []string{"9.9.9.9", "8.8.8.8"}); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
|
||||
st, _ := o.StartScan(ScanOptions{ClearFirst: true})
|
||||
if st.State != ScanClearing {
|
||||
t.Fatalf("expected clearing as first state, got %s", st.State)
|
||||
}
|
||||
fin := waitScan(t, o)
|
||||
if fin.State != ScanDone || fin.Added != 1 {
|
||||
t.Fatalf("clear-first scan: %+v", fin)
|
||||
}
|
||||
if got := queuedAddresses(t, d); len(got) != 1 || got["1.1.1.1"] != db.IPQueued {
|
||||
t.Fatalf("expected only the scanned address, got %v", got)
|
||||
}
|
||||
var payload struct {
|
||||
Count int `json:"count"`
|
||||
Addresses []string `json:"addresses"`
|
||||
}
|
||||
if err := json.Unmarshal([]byte(eventPayload(t, d, "queue_cleared")), &payload); err != nil || payload.Count != 2 || len(payload.Addresses) != 2 {
|
||||
t.Fatalf("queue_cleared payload: %+v err=%v", payload, err)
|
||||
}
|
||||
}
|
||||
|
||||
// 6440 free addresses (the real stand's size) plus a few out-of-order ones:
|
||||
// everything is queued in ascending numeric IP order within a sane time.
|
||||
func TestScanJobEnqueuesAllInAscendingOrderAtScale(t *testing.T) {
|
||||
if testing.Short() {
|
||||
t.Skip("scale test skipped in -short mode")
|
||||
}
|
||||
ctx := context.Background()
|
||||
o, d, mock := newTestOrchestrator(t, 180)
|
||||
o.ScanPageSize = 200
|
||||
mock.SeedMany("fip", 6440)
|
||||
// IDs sort before "fip-*", addresses sort numerically before 198.18.*;
|
||||
// 10.0.0.9 < 10.0.0.200 numerically but not as strings.
|
||||
mock.Seed("aaa-2", "10.0.0.200", "svc")
|
||||
mock.Seed("aaa-1", "10.0.0.9", "svc")
|
||||
mock.SeedWithPort("occupied", "203.0.113.1", "svc", "port-x")
|
||||
|
||||
start := time.Now()
|
||||
o.StartScan(ScanOptions{})
|
||||
st := waitScanFor(t, o, 5*time.Minute) // -race is an order of magnitude slower
|
||||
elapsed := time.Since(start)
|
||||
|
||||
if st.State != ScanDone || st.Free != 6442 || st.Added != 6442 || st.Discovered != 6443 || st.Pages != 33 {
|
||||
t.Fatalf("scan status: %+v", st)
|
||||
}
|
||||
if elapsed > 3*time.Minute {
|
||||
t.Fatalf("scanning 6442 addresses took %v", elapsed)
|
||||
}
|
||||
t.Logf("scan+enqueue of %d addresses took %v", st.Added, elapsed)
|
||||
|
||||
items, err := d.ListIPs(ctx)
|
||||
if err != nil || len(items) != 6442 {
|
||||
t.Fatalf("queue: n=%d err=%v", len(items), err)
|
||||
}
|
||||
var prev netip.Addr
|
||||
for i, it := range items {
|
||||
ip := netip.MustParseAddr(it.IPAddress)
|
||||
if i > 0 && ip.Compare(prev) <= 0 {
|
||||
t.Fatalf("queue not ascending at %d: %s after %s", i, it.IPAddress, prev)
|
||||
}
|
||||
if it.Sequence != i {
|
||||
t.Fatalf("expected contiguous sequences, row %d has %d", i, it.Sequence)
|
||||
}
|
||||
prev = ip
|
||||
}
|
||||
if items[0].IPAddress != "10.0.0.9" || items[1].IPAddress != "10.0.0.200" {
|
||||
t.Fatalf("numeric ordering broken: %s, %s", items[0].IPAddress, items[1].IPAddress)
|
||||
}
|
||||
if _, stale := queuedAddresses(t, d)["203.0.113.1"]; stale {
|
||||
t.Fatalf("occupied address must not be queued")
|
||||
}
|
||||
}
|
||||
|
||||
func TestScanJobCancelAndLifetimeContext(t *testing.T) {
|
||||
ctx := context.Background()
|
||||
o, d, mock := newTestOrchestrator(t, 180)
|
||||
mock.SeedMany("fip", 100)
|
||||
mock.PageSize = 10
|
||||
mock.PageDelay = 5 * time.Second
|
||||
if err := d.SeedQueue(ctx, []string{"9.9.9.9"}); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
|
||||
o.StartScan(ScanOptions{})
|
||||
time.Sleep(20 * time.Millisecond)
|
||||
start := time.Now()
|
||||
if !o.CancelScan() {
|
||||
t.Fatalf("expected a running scan to be cancelled")
|
||||
}
|
||||
if time.Since(start) > 3*time.Second {
|
||||
t.Fatalf("cancel took too long")
|
||||
}
|
||||
st := o.ScanStatus()
|
||||
if st.State != ScanCancelled || st.Running || st.FinishedAt == nil {
|
||||
t.Fatalf("expected cancelled, got %+v", st)
|
||||
}
|
||||
if got := queuedAddresses(t, d); len(got) != 1 {
|
||||
t.Fatalf("cancelled scan must not enqueue, got %d rows", len(got))
|
||||
}
|
||||
|
||||
// Cancelling the process-lifetime context also stops a running scan.
|
||||
life, cancelLife := context.WithCancel(context.Background())
|
||||
o.SetContext(life)
|
||||
o.StartScan(ScanOptions{})
|
||||
time.Sleep(20 * time.Millisecond)
|
||||
cancelLife()
|
||||
if st := waitScan(t, o); st.State != ScanCancelled {
|
||||
t.Fatalf("expected cancelled by lifetime ctx, got %+v", st)
|
||||
}
|
||||
}
|
||||
|
||||
func TestScanJobDeadline(t *testing.T) {
|
||||
o, _, mock := newTestOrchestrator(t, 180)
|
||||
o.Cfg.FIPScanTimeoutSeconds = 1
|
||||
mock.SeedMany("fip", 20)
|
||||
mock.PageSize = 10
|
||||
mock.PageDelay = 10 * time.Second
|
||||
|
||||
o.StartScan(ScanOptions{})
|
||||
st := waitScan(t, o)
|
||||
if st.State != ScanError || !strings.Contains(st.Error, "timed out") {
|
||||
t.Fatalf("expected timeout error, got %+v", st)
|
||||
}
|
||||
}
|
||||
|
||||
func TestScanFloatingIPsWrapperHonoursCallerContext(t *testing.T) {
|
||||
o, _, mock := newTestOrchestrator(t, 180)
|
||||
mock.SeedMany("fip", 20)
|
||||
mock.PageSize = 10
|
||||
mock.PageDelay = 5 * time.Second
|
||||
|
||||
ctx, cancel := context.WithTimeout(context.Background(), 50*time.Millisecond)
|
||||
defer cancel()
|
||||
if _, _, err := o.ScanFloatingIPs(ctx); !errors.Is(err, context.DeadlineExceeded) {
|
||||
t.Fatalf("expected the caller's ctx error, got %v", err)
|
||||
}
|
||||
// The job itself keeps going until cancelled (cleanup cancels it).
|
||||
if !o.ScanStatus().Running {
|
||||
t.Fatalf("background job should still be running")
|
||||
}
|
||||
}
|
||||
|
||||
func TestSortAddressesAscending(t *testing.T) {
|
||||
in := []string{"10.0.0.10", "zzz", "10.0.0.9", "2001:db8::1", "9.255.255.255", "aaa", "10.0.0.2"}
|
||||
sortAddressesAscending(in)
|
||||
want := []string{"9.255.255.255", "10.0.0.2", "10.0.0.9", "10.0.0.10", "2001:db8::1", "aaa", "zzz"}
|
||||
for i := range want {
|
||||
if in[i] != want[i] {
|
||||
t.Fatalf("got %v want %v", in, want)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
func TestClearQueueEventPayloadIsTruncated(t *testing.T) {
|
||||
ctx := context.Background()
|
||||
o, d, _ := newTestOrchestrator(t, 180)
|
||||
var addrs []string
|
||||
for i := 0; i < 120; i++ {
|
||||
addrs = append(addrs, "10.1.0."+strconv.Itoa(i))
|
||||
}
|
||||
if err := d.SeedQueue(ctx, addrs); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
res, err := o.ClearQueue(ctx)
|
||||
if err != nil || len(res.Deleted) != 120 {
|
||||
t.Fatalf("clear: %+v err=%v", res, err)
|
||||
}
|
||||
var p struct {
|
||||
Count int `json:"count"`
|
||||
Addresses []string `json:"addresses"`
|
||||
Truncated bool `json:"truncated"`
|
||||
}
|
||||
if err := json.Unmarshal([]byte(eventPayload(t, d, "queue_cleared")), &p); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if p.Count != 120 || len(p.Addresses) != 50 || !p.Truncated || p.Addresses[0] != "10.1.0.0" {
|
||||
t.Fatalf("payload: count=%d addrs=%d truncated=%v", p.Count, len(p.Addresses), p.Truncated)
|
||||
}
|
||||
}
|
||||
|
||||
func TestDeleteIPsDisassociatesOnlyAttachedFIPs(t *testing.T) {
|
||||
ctx := context.Background()
|
||||
o, d, mock := newTestOrchestrator(t, 180)
|
||||
mock.Seed("fip-a", "1.1.1.1", "svc")
|
||||
mock.Seed("fip-b", "2.2.2.2", "svc")
|
||||
if _, err := d.SubmitIPs(ctx, []string{"1.1.1.1", "2.2.2.2", "3.3.3.3"}); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if err := d.RegisterValidator(ctx, "validator-1", "host-1", "port-1", "v0.1"); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
claimed, err := d.ClaimNextQueued(ctx, "validator-1", time.Minute)
|
||||
if err != nil || claimed == nil || claimed.IPAddress != "1.1.1.1" {
|
||||
t.Fatalf("claim: %+v err=%v", claimed, err)
|
||||
}
|
||||
if err := mock.AssociateFloatingIP(ctx, "fip-a", "port-1"); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if err := d.SetFIPAssociated(ctx, claimed.ID, "fip-a", time.Minute); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
// An unrelated, externally associated FIP must stay as it is.
|
||||
if err := mock.AssociateFloatingIP(ctx, "fip-b", "other-port"); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
|
||||
res, err := o.DeleteIPs(ctx, []string{"1.1.1.1", "2.2.2.2", "nope"})
|
||||
if err != nil || len(res.Deleted) != 2 || len(res.NotFound) != 1 {
|
||||
t.Fatalf("delete: %+v err=%v", res, err)
|
||||
}
|
||||
a, _ := mock.GetFloatingIPByAddress(ctx, "1.1.1.1")
|
||||
b, _ := mock.GetFloatingIPByAddress(ctx, "2.2.2.2")
|
||||
if a.PortID != "" {
|
||||
t.Fatalf("attached fip must be disassociated, still on %q", a.PortID)
|
||||
}
|
||||
if b.PortID != "other-port" {
|
||||
t.Fatalf("row without recorded fip_id must not trigger a disassociate, got %q", b.PortID)
|
||||
}
|
||||
v, _ := d.GetValidator(ctx, "validator-1")
|
||||
if v.State != db.ValidatorIdle {
|
||||
t.Fatalf("validator must be freed, got %s", v.State)
|
||||
}
|
||||
}
|
||||
@@ -37,6 +37,20 @@ openstack:
|
||||
user_domain_name_env: "OS_USER_DOMAIN_NAME"
|
||||
password_env: "OS_PASSWORD"
|
||||
|
||||
# Постраничное чтение Floating IP (скан при тысячах адресов): сколько
|
||||
# адресов запрашивать у Neutron за один запрос. Default 200.
|
||||
# Page size of the paged floating-IP listing. Default 200.
|
||||
list_page_size: 200
|
||||
# Таймаут каждого HTTP-запроса к Keystone/Neutron, секунд. Default 60.
|
||||
# Per-request HTTP timeout (also protects the orchestrator tick from a
|
||||
# hung Neutron call). Default 60.
|
||||
request_timeout_seconds: 60
|
||||
# Сколько раз повторять неудавшуюся страницу (сетевая ошибка, EOF/
|
||||
# RemoteDisconnected, 5xx, 429) с паузами 1,2,4,8,16 с. Default 5;
|
||||
# отрицательное значение отключает повторы.
|
||||
# Retries per failed listing page. Default 5; negative disables retries.
|
||||
list_page_retries: 5
|
||||
|
||||
# Аутентификация API: здесь только ИМЕНА переменных окружения, значения
|
||||
# (статические bearer-токены) задаются окружением процесса — см.
|
||||
# deploy/systemd/control-api.service (EnvironmentFile=). Генерация:
|
||||
@@ -76,6 +90,10 @@ orchestrator:
|
||||
# scan on demand via POST /api/v1/admin/ips/scan or the dashboard's
|
||||
# "Scan Floating IPs" button.
|
||||
fip_scan_interval_seconds: 0
|
||||
# Общий таймаут одного фонового скана Floating IP (очистка + чтение всех
|
||||
# страниц + постановка в очередь), секунд. Default 1800.
|
||||
# Overall deadline of one background floating-IP scan. Default 1800.
|
||||
fip_scan_timeout_seconds: 1800
|
||||
|
||||
aggregation:
|
||||
missing_counts_as_fail: true
|
||||
|
||||
Reference in new issue
Block a user