Files
OpenVPN-Monitoring-Simple/DOCS/Operations/REBOOT-TEST.md
T
iclaoudezinandClaude Sonnet 5.5 3836049230 Docs: add operations records (SUMMARY, Hysteria chain manifest, reboot test, plans)
- DOCS/Operations: current-state SUMMARY of the ENTRY deployment (access,
  install, egress via Hysteria2, security findings and fixes, risk
  assessment, commits/backups, open items), the Hysteria chain manifest
  with an OpenVPN Monitor section, reboot test results, and the plan and
  rollout records for settings validation and privilege separation.
- Link them from README and DOCS/General/Index.md.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-09-30 12:59:55 +00:00

6.8 KiB

Reboot test: ENTRY (Timeweb) and EXIT (Hel)

Date: 2026-09-30. Goal: confirm the Hysteria 2 chain recovers by itself after a reboot of each node, with no manual steps.

Nodes: ENTRY 213.226.125.13 (Alpine, OpenRC), EXIT 85.234.86.140 (Debian 13, systemd).

Methodology

Order: ENTRY first, then EXIT. Each node is rebooted separately, so the other side is always up and its reconnect behavior is observable.

1. Pre-flight (before each reboot)

Check ENTRY EXIT
Services enabled rc-update show default (hysteria, hysteria-route, hysteria-edge) systemctl is-enabled hysteria-server
Kernel module for TUN persists lsmod | grep tun, grep -x tun /etc/modules not needed
sysctl persisted ls /etc/sysctl.d/ (99-hytun.conf, 99-quic.conf) ls /etc/sysctl.d/99-quic.conf
Temporary test settings removed n/a grep -c speedTest /etc/hysteria/config.yaml = 0

Any gap found is fixed before rebooting.

2. Reboot

sync, then a delayed reboot / systemctl reboot (detached, so the SSH command returns). Wait 45 s, then poll SSH every 6 s (up to 20 tries) until uptime answers.

3. Post-boot verification

ENTRY:

  1. rc-service hysteria|hysteria-route|hysteria-edge status → started.
  2. ls -l /dev/net/tun, ip -br a show hytun.
  3. ip rule | grep -E '^10[01]' and ip route show table 100 → both rules, default via hytun, blackhole fallback.
  4. sysctl net.ipv4.conf.all.rp_filter net.core.rmem_max → 2 and 7500000.
  5. ss -lunp | grep :443 → hysteria-edge listening.
  6. Tunnel TCP: su -s /bin/sh hyedge -c 'curl -4 -s https://ifconfig.me' → must equal the EXIT IP.
  7. Tunnel UDP: su -s /bin/sh hyedge -c 'nslookup example.com 1.1.1.1' → address returned.
  8. Logs: /var/log/hysteria.log, /var/log/hysteria-edge.log; a real client (mobile) reconnecting is a bonus signal.

EXIT:

  1. systemctl is-active hysteria-server, ss -lunp | grep :443.
  2. sudo /usr/sbin/sysctl net.core.rmem_max net.core.wmem_max (sysctl is not in a non-root user's PATH).
  3. journalctl -u hysteria-server -b → "server up and running" and "client connected" from ENTRY.
  4. From ENTRY: client log shows a new connected to server (count incremented) and the hyedge curl test (step 6 above) returns the EXIT IP.

Results

ENTRY reboot (Timeweb)

Check Result
Pre-flight Gap found: module tun loaded by hand but absent from /etc/modules → after reboot /dev/net/tun could be missing and the tunnel would not start. Fixed (echo tun >> /etc/modules) before rebooting.
SSH back Yes, about 45-60 s after the reboot command
Services hysteria, hysteria-route, hysteria-edge all started
TUN / routing /dev/net/tun present, hytun up (100.100.100.101/30), rules 100 and 101 present, table 100: default dev hytun + blackhole
sysctl rp_filter=2, rmem_max=7500000 applied
Listener hysteria-edge on UDP 443
Tunnel TCP 85.234.86.140 (EXIT IP)
Tunnel UDP DNS answer received
Client Hiddify (mobile) reconnected on its own right after boot
Noise in logs A few dial timeouts to individual remote hosts (unreachable destinations, not a tunnel fault)

Verdict: PASS (after the /etc/modules fix).

EXIT reboot (Hel)

Check Result
Pre-flight hysteria-server enabled, 99-quic.conf present, speedTest absent from config (0)
SSH back Yes, about 45-50 s after the reboot command
Service hysteria-server active, listening on UDP 443
sysctl rmem_max and wmem_max = 7500000
Journal "server up and running" at boot, then "client connected" from 213.226.125.13 about 30 s later
ENTRY side Client reconnected automatically (connected to server, count: 2); brief TUN UDP error ... EOF while EXIT was down; hyedge curl returned 85.234.86.140

Verdict: PASS.

Conclusions

  • Both nodes recover unattended; the ENTRY hysteria client (supervise-daemon, lazy: true) reconnects without restart when EXIT comes back.
  • During an EXIT outage the anti-leak blackhole route in table 100 keeps hyedge traffic from leaving via ENTRY's own uplink (leak test done separately: removing the tun default route makes the test curl fail instead of exposing the ENTRY IP).
  • Not covered: simultaneous reboot of both nodes, power-loss/unclean shutdown, and IKE/strongSwan (separate task, charon is not enabled at boot).

ENTRY reboot after the OpenVPN Monitor deployment (2026-09-30)

Goal: confirm the whole ENTRY stack (Hysteria chain + OpenVPN Monitor, OpenVPN, egress rules, hardening) recovers by itself. Only ENTRY was rebooted; EXIT stayed up.

Pre-flight

  • rc-update show default: hysteria, hysteria-route, hysteria-edge, ovpn-hytun, openvpn, ovpmon-api, ovpmon-gatherer, ovpmon-profiler, nginx, fail2ban, sshd, chronyd enabled.
  • tun in /etc/modules; sysctl files present (ip_forward=1, rp_filter=2); sshd -t and doas -C /etc/doas.d/ovpmon.conf valid.
  • Baseline snapshot /root/pre-reboot-snapshot.txt: service states, ip rule 100-102, table 100, nat/FORWARD rules, listeners, process owners.

Reboot and result

sync, detached reboot; SSH answered again after ~45 s. Post-boot snapshot /root/post-reboot-snapshot.txt is identical to the baseline.

Check Result
All services in default started by themselves
API processes gunicorn/uvicorn/gatherer run as ovpmon
hytun, tun0, rp_filter=2, ip_forward=1 present
ip rule 100/101/102, table 100 (hytun + blackhole), MASQUERADE to hytun, FORWARD DROP guard restored by hysteria-route / ovpn-hytun
Tunnel TCP as hyedge 85.234.86.140 (EXIT)
Hysteria client reconnected to EXIT
Panel https://…:8088/ 200, profiler without token 401, HTTP→HTTPS 301, login tstark works
Helper/doas service status and process stats OK; gatherer reads openvpn-status.log (root:ovpmon 640 persisted)
/etc/openvpn/crl.pem, /var/lib/ovpmon ownership intact
fail2ban 3 jails active
SSH passwordauthentication no, permitrootlogin prohibit-password
Errors in ovpmon-* logs since boot none (older entries predate the boot)
Real VPN client reconnected by itself ~3.5 min after boot (UDP client waits for its keepalive 10 120 timeout); ~5 MB downloaded through the tunnel, MASQUERADE counters rose, leak guard FORWARD DROP counter stayed 0

Notes

  • Expected client gap after a server reboot is up to ~2-4 minutes for UDP OpenVPN clients (ping-restart), not a fault of the server.
  • openvpn-status.log keeps root:ovpmon 640 because the file persists on disk; if it is ever deleted, OpenVPN recreates it as 600 and the gatherer cannot read it until the next ovpmon-helper service start|restart (which fixes the mode). Not triggered by a reboot.