mvm-s3 is a separate VK Cloud project whose admin pre-created two private networks/subnets with a known IP per router. Unify project-managed (private_network_cidrs, IPAM-assigned) and externally-owned (router_networks, fixed-IP) private interfaces into one local.router_interfaces so both share the existing port/dynamic-network mechanism instead of duplicating it. Switch from implicit *.auto.tfvars loading to explicit -var-file per environment (now two share this terraform/ directory) plus a dedicated Terraform workspace for mvm-s3, so PROD's state and credentials are never touched by mvm-s3 applies. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GHfG9FgpMrGdrvC1QUewTw
Quick Start
This is a conceptual Demo Scenario that will help you to bring highly available and secured connectivity with dynamic routing between Cloud and On-Prem using regular Internet circuits.
Key idea of this scenario based on limitations coming from On-Prem side which has two Internet circuits (Main and Backup where Backup is the
Radio Bridge)
See docs/QUICKSTART.md for a condensed step-by-step self-service deployment guide.
The materials from this repository will help you quickly build from the scratch the following network topology:
To prepare your admin workstation (desktop, laptop or maybe something else) follow these steps:
- Prepare your VK Cloud project (enable CLI and API access): URL
- Create and upload your SSH key into the cloud admin account: URL
- Install Terraform components depending on your OS: URL
- Install Ansible components depending on your OS: URL
- Install GIT components and copy this repo onto your admin workstation
Additional Steps:
- Use you private SSH key within Terraform and Ansible
- Use proper account credentials within Terraform
Under the Hood
Main part of this scenario related to the routers (a set of IaaS Virtual Machines (
IaaS Routers) converted into traditional routers with advanced functionality)
Adaptive router VM count and interfaces
The number of IaaS Router VMs is controlled by the Terraform variable router_count (default 2, tested with 3-4). Each router VM's private interfaces are driven by private_network_cidrs - a list of CIDR prefixes the admin must supply explicitly (no default, no auto-carving): one entry = one shared private network = one private interface per router, in addition to the single public (WAN) interface that stays fixed at 1. So a router VM ends up with 1 + length(private_network_cidrs) network interfaces total.
Each entry in private_network_cidrs is one network shared by all routers - every router gets its own port inside every listed network, similar to how the repo's original single shared LAN network worked, just generalized to an arbitrary number of networks and routers. Each port's address is auto-assigned by Neutron's IPAM (no explicit ip_address) rather than hand-computed, since VKCS auto-creates its own service ports on each network (observed: a network:dns port) that can otherwise collide with a manually-picked address - IPAM guarantees no double-booking. There is no VRRP between router VMs in this design, a change from the previous 2-NIC/VRRP model. The prefixes must not overlap each other and should be sized /28 (not /29 - too tight once platform-reserved addresses are accounted for) - both checked by terraform/tests/.
This deployment provisions only the router VMs - the repo's earlier priv_srv_01/priv_srv_02/priv_srv_03 demo instances (and the shared LAN network/security group that only they used) have been removed as out of scope.
New Terraform variables: router_count, private_network_cidrs, router_availability_zones (see terraform/variables.tf). The post-install script is a Terraform template (terraform/scripts/network-init.sh.tpl) rendered per-router via templatefile(), matching each private interface to its expected subnet deterministically instead of guessing - it already handles any interface count, no hardcoded assumption of 2. terraform/versions.tf now pins the provider source (vk-cs/vkcs, ~> 0.17), which was previously undeclared.
Provider authentication
terraform/versions.tf now has a provider "vkcs" { ... } block wiring auth_url, username, password, project_id, region, user_domain_name (previously declared in variables.tf but never actually connected to anything). auth_url/user_domain_name/region have sane defaults for a regular account; username/password/project_id still have none, same as before.
terraform.tfvars stays a committed, anonymized template - real credentials (e.g. from an openrc.sh for a service account) go in a separate, gitignored *.secrets.tfvars file per environment instead (see terraform/prod.secrets.tfvars.example / terraform/mvm-s3.secrets.tfvars.example for the field mapping from openrc.sh's OS_* variables), passed explicitly via -var-file on every plan/apply - see "External fixed-IP networks and multi-environment workspaces" below for why this is no longer auto-loaded. Never put real credentials in a committed *.tfvars file.
Default security group
Every VK Cloud project auto-creates a default security group with a UUID unique to that project. Rather than hardcoding one project's UUID, it's resolved dynamically via data.vkcs_networking_secgroup (matched by name = "default") and attached to every VM through local.default_security_group_id. default_security_group_id is a Terraform variable for the rare case that lookup doesn't fit a given project (non-standard name/SDN) - treat setting it explicitly (via terraform.tfvars or TF_VAR_default_security_group_id) as a last resort, not the normal path.
Horizontal scaling via environment variables
Terraform's variable precedence means a terraform.tfvars value always beats a TF_VAR_<name> environment variable, never the other way round - TF_VAR_* only takes effect for a variable terraform.tfvars leaves unset. This repo's terraform.tfvars currently pins real values for its actual deployment (router_count = 3, a specific private_network_cidrs), so TF_VAR_router_count/TF_VAR_private_network_cidrs have no effect while those stay set - to change scale, edit terraform.tfvars directly, or override on the command line with -var/-var-file (which does beat a tfvars file):
terraform apply -var-file=terraform.tfvars -var-file=prod.secrets.tfvars \
-var="router_count=4" -var='private_network_cidrs=["10.90.0.0/28","10.90.0.16/28","10.90.0.32/28"]'
If you instead comment router_count/private_network_cidrs back out of terraform.tfvars (e.g. for a fresh, non-PROD deployment), TF_VAR_router_count/TF_VAR_private_network_cidrs (JSON-encoded list) start working again as described above.
External fixed-IP networks and multi-environment workspaces
Not every private interface has to be a network Terraform creates itself: router_networks (default {}) is a map of pre-existing networks - typically owned by a different VK Cloud project, referenced by UUID only - that each router gets a fixed-IP interface into. Map key = role/interface name (e.g. "primary"/"backup"); ip_addresses[i] is the address for router(i+1), pre-agreed by that other project's admin (Terraform never invents or discovers it). locals.router_interfaces in main.tf merges both sources (private_network_cidrs and router_networks) into one role → definition map, so the rest of the pipeline - port creation, the instance's dynamic "network" blocks, network-init.sh.tpl's CIDR-based interface matching - stays a single generalized mechanism regardless of which source a role came from. A role's ip_address is left null (Neutron IPAM auto-assigns, same collision-avoidance rationale as before) when it comes from private_network_cidrs, and set explicitly when it comes from router_networks.
The first real deployment on this mechanism is mvm-s3 (see docs/changes/2026-09-09-mvm-s3-external-networks-*.md and terraform/mvm-s3.tfvars): a separate VK Cloud project where both private networks already exist, so private_network_cidrs = [] there (no project-managed networks at all) and both interfaces come from router_networks.
Since this repo's single terraform/ directory now serves more than one environment (PROD and mvm-s3), each with different router_count/networks/credentials, two conventions changed to keep them from interfering with each other:
- Separate Terraform workspaces per environment (
terraform workspace new mvm-s3), so each has its own state and applying one never touches the other's resources. - Explicit
-var-filefor everything environment-specific, including credentials -*.auto.tfvarsauto-loading was fine for exactly one environment, but with two*.auto.tfvarsfiles present at once Terraform would load both and silently merge them. Real credentials now live in a*.secrets.tfvarsper environment (gitignored,.gitignorepattern*.secrets.tfvars), never auto-loaded, always passed explicitly - seedocs/QUICKSTART.md§4/§4а for the exact commands.
Local delivery integrity tests
terraform/tests/ contains a local pytest suite that checks the delivery is internally consistent - required files present, terraform fmt clean, HCL parses, router_count/private_network_cidrs actually drive the resource/NIC count instead of being hardcoded, the example CIDRs in terraform.tfvars don't overlap and have room for router_count routers, and the post-install script template renders to valid bash. It also runs:
- a real
terraform init+terraform validateagainst the actualvkcsprovider schema (at severalrouter_count/private_network_cidrs/router_networksshapes, including anmvm-s3-shaped one), against a project-local filesystem-mirror copy of the provider - no cloud API is ever contacted and no credentials are needed; - a real
terraform planagainst an isolated, provider-free copy of justvariables.tf, to prove thevalidation { ... }blocks onrouter_count,private_network_cidrs(valid CIDR syntax, uniqueness) androuter_networks(valid CIDR/IPv4 syntax, each address falling inside its own role'scidr) are actually enforced -terraform validatealone does not enforce custom variable validations for externally-supplied values, onlyplan/applydo.
terraform/tests/setup-local-terraform.sh # one-time: provisions venv/ with terraform + the vkcs provider
venv/bin/pytest terraform/tests -v
setup-local-terraform.sh builds everything inside the git-ignored venv/ directory:
- the Python packages from
terraform/tests/requirements.txt(pytest,python-hcl2,checkov); - the
terraformCLI, downloaded (withSHA256SUMSverification) from a region-unrestricted HashiCorp releases mirror; - the
vk-cs/vkcsprovider binary, downloaded (withSHA256SUMSverification) directly from its GitHub releases - this bypassesregistry.terraform.io, which blocks some regions outright, and is what makes a realterraform validatepossible at all here; - a project-local CLI config (
venv/terraform.d/cli-config.tfrc) that pointsterraform initat that local provider copy via afilesystem_mirrorblock, instead of the network registry.
Note: the diagrams below (
ports.svg,topology.svg) and the Ansible layer (ansible/inventory.ini, rolesbase/frr_router/keepalived) still describe/assume the previous 2-router, 2-NIC, VRRP-based design and have not been updated for the new N-NIC/N-router topology yet - that's a separate follow-up.
IPv4 addressing plan for the project
Here is a card to assist with configuration planning. The card is filled out using the IP addressing from the Demo Scenario and the inventory.ini file, which will be used when running the Ansible playbook.
Terraform
Provisions router_count IaaS Routers (default 2), each with 1 public and length(private_network_cidrs) private ports (2 in the shipped example). Includes supplimentary Shell script template (which is a part of Terraform manifest) to maintain configuration across reboots.
Ansible
Configure IaaS Routers using role-based playbooks controlled via the Inventory File
Additional Software Used:
- strongSwan (to manage IPsec)
- FRR (to manage BGP)
- Keepalived (VRRP)
Each IaaS Router will use two secured connections to On-Prem environment through the Internet:
- IPsec Site-to-Site in Transport Mode (to protect GRE Tunnels)
- GRE Tunnel (to transfer a data)
GRE Tunnels topology clearly explained in the following diagram:
High Availability Design
BGP peering eliminates single points of failure on the Cloud side through:
- Bidirectional eBGP sessions from each
IaaS Routerto On-Premises - Optimized route metrics reflecting circuit priority (Primary/Backup)
- Automatic failover during circuit failures (including Cloud Availability Zone failures)
- Asymmetric routing prevention via MED and Local Preference configuration