Confirmed working against a real VK Cloud PROD deployment (3 routers, 19 resources, apply succeeded end to end). Fixes found along the way: - provider "vkcs" was never configured (versions.tf) - username/password/ project_id/region were declared but wired to nothing; added auth_url and user_domain_name to complete it. - Nova keypairs are per-user, not per-project - added an optional vkcs_compute_keypair resource (var.ssh_public_key) so Terraform can register a keypair under the deploying service account itself. - router_priv_port used a hand-computed fixed_ip offset that collided with VKCS's own auto-created service ports on each network (observed: a "network:dns" port) - now left unset so Neutron's IPAM auto-assigns, which is collision-free by construction. - vkcs_compute_instance set image_id at the top level while also booting from a volume via block_device - the provider docs say not to do this; Nova echoes back a sentinel string for image_id on a volume-booted server, which Terraform read as drift on a ForceNew attribute and wanted to destroy+recreate every already-created instance on every subsequent plan. - private_network_cidrs bumped from /29 to /28 - too tight once the platform's own reserved ports are accounted for. Also removed the priv_srv_01/02/03 demo instances and the LAN network/ security group only they used - this deployment provisions router VMs only, confirmed with the user. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011hXR2ftXZZhJ4Y3XuSoR8r
124 lines
10 KiB
Markdown
124 lines
10 KiB
Markdown
# Quick Start
|
|
|
|
This is a conceptual Demo Scenario that will help you to bring highly available and secured connectivity with dynamic routing between Cloud and On-Prem using regular Internet circuits.
|
|
|
|
>Key idea of this scenario based on limitations coming from On-Prem side which has two Internet circuits (Main and Backup where Backup is the `Radio Bridge`)
|
|
|
|
See [docs/QUICKSTART.md](docs/QUICKSTART.md) for a condensed step-by-step self-service deployment guide.
|
|
|
|
The materials from this repository will help you quickly build from the scratch the following network topology:
|
|
|
|

|
|
|
|
To prepare your admin workstation (desktop, laptop or maybe something else) follow these steps:
|
|
|
|
1. Prepare your VK Cloud project (enable CLI and API access): [URL](https://cloud.vk.com/docs/en/tools-for-using-services/api/rest-api/enable-api)
|
|
2. Create and upload your SSH key into the cloud admin account: [URL](https://cloud.vk.com/docs/tools-for-using-services/vk-cloud-account/instructions/account-manage/keypairs#importing_existing_key)
|
|
3. Install Terraform components depending on your OS: [URL](https://cloud.vk.com/docs/en/tools-for-using-services/terraform/quick-start)
|
|
4. Install Ansible components depending on your OS: [URL](https://docs.ansible.com/projects/ansible/latest/installation_guide/intro_installation.html#pipx-install)
|
|
5. Install GIT components and copy this repo onto your admin workstation
|
|
|
|
Additional Steps:
|
|
|
|
- Use you private SSH key within Terraform and Ansible
|
|
- Use proper account credentials within Terraform
|
|
|
|
# Under the Hood
|
|
|
|
>Main part of this scenario related to the routers (a set of IaaS Virtual Machines (`IaaS Routers`) converted into traditional routers with advanced functionality)
|
|
|
|
**Adaptive router VM count and interfaces**
|
|
|
|
The number of `IaaS Router` VMs is controlled by the Terraform variable `router_count` (default `2`, tested with 3-4). Each router VM's private interfaces are driven by `private_network_cidrs` - a list of CIDR prefixes the admin must supply explicitly (no default, no auto-carving): **one entry = one shared private network = one private interface per router**, in addition to the single public (WAN) interface that stays fixed at 1. So a router VM ends up with `1 + length(private_network_cidrs)` network interfaces total.
|
|
|
|
Each entry in `private_network_cidrs` is one network shared by *all* routers - every router gets its own port inside every listed network, similar to how the repo's original single shared LAN network worked, just generalized to an arbitrary number of networks and routers. Each port's address is auto-assigned by Neutron's IPAM (no explicit `ip_address`) rather than hand-computed, since VKCS auto-creates its own service ports on each network (observed: a `network:dns` port) that can otherwise collide with a manually-picked address - IPAM guarantees no double-booking. There is no VRRP between router VMs in this design, a change from the previous 2-NIC/VRRP model. The prefixes must not overlap each other and should be sized `/28` (not `/29` - too tight once platform-reserved addresses are accounted for) - both checked by `terraform/tests/`.
|
|
|
|
This deployment provisions **only the router VMs** - the repo's earlier `priv_srv_01`/`priv_srv_02`/`priv_srv_03` demo instances (and the shared LAN network/security group that only they used) have been removed as out of scope.
|
|
|
|
New Terraform variables: `router_count`, `private_network_cidrs`, `router_availability_zones` (see `terraform/variables.tf`). The post-install script is a Terraform template (`terraform/scripts/network-init.sh.tpl`) rendered per-router via `templatefile()`, matching each private interface to its expected subnet deterministically instead of guessing - it already handles any interface count, no hardcoded assumption of 2. `terraform/versions.tf` now pins the provider source (`vk-cs/vkcs`, `~> 0.17`), which was previously undeclared.
|
|
|
|
**Provider authentication**
|
|
|
|
`terraform/versions.tf` now has a `provider "vkcs" { ... }` block wiring `auth_url`, `username`, `password`, `project_id`, `region`, `user_domain_name` (previously declared in `variables.tf` but never actually connected to anything). `auth_url`/`user_domain_name`/`region` have sane defaults for a regular account; `username`/`password`/`project_id` still have none, same as before.
|
|
|
|
`terraform.tfvars` stays a committed, anonymized template - real credentials (e.g. from an `openrc.sh` for a service account) go in a separate, gitignored `*.auto.tfvars` file instead, which Terraform loads automatically on top of `terraform.tfvars`. See `terraform/prod.auto.tfvars.example` for the field mapping from `openrc.sh`'s `OS_*` variables. Never put real credentials in `terraform.tfvars` itself.
|
|
|
|
**Default security group**
|
|
|
|
Every VK Cloud project auto-creates a `default` security group with a UUID unique to that project. Rather than hardcoding one project's UUID, it's resolved dynamically via `data.vkcs_networking_secgroup` (matched by `name = "default"`) and attached to every VM through `local.default_security_group_id`. `default_security_group_id` is a Terraform variable for the rare case that lookup doesn't fit a given project (non-standard name/SDN) - treat setting it explicitly (via `terraform.tfvars` or `TF_VAR_default_security_group_id`) as a last resort, not the normal path.
|
|
|
|
**Horizontal scaling via environment variables**
|
|
|
|
Terraform's variable precedence means a `terraform.tfvars` value always beats a `TF_VAR_<name>` environment variable, never the other way round - `TF_VAR_*` only takes effect for a variable `terraform.tfvars` leaves unset. This repo's `terraform.tfvars` currently pins real values for its actual deployment (`router_count = 3`, a specific `private_network_cidrs`), so `TF_VAR_router_count`/`TF_VAR_private_network_cidrs` have **no effect** while those stay set - to change scale, edit `terraform.tfvars` directly, or override on the command line with `-var`/`-var-file` (which does beat a tfvars file):
|
|
|
|
```bash
|
|
terraform apply -var="router_count=4" -var='private_network_cidrs=["10.90.0.0/28","10.90.0.16/28","10.90.0.32/28"]'
|
|
```
|
|
|
|
If you instead comment `router_count`/`private_network_cidrs` back out of `terraform.tfvars` (e.g. for a fresh, non-PROD deployment), `TF_VAR_router_count`/`TF_VAR_private_network_cidrs` (JSON-encoded list) start working again as described above.
|
|
|
|
**Local delivery integrity tests**
|
|
|
|
`terraform/tests/` contains a local pytest suite that checks the delivery is internally consistent - required files present, `terraform fmt` clean, HCL parses, `router_count`/`private_network_cidrs` actually drive the resource/NIC count instead of being hardcoded, the example CIDRs in `terraform.tfvars` don't overlap and have room for `router_count` routers, and the post-install script template renders to valid bash. It also runs:
|
|
|
|
- a real `terraform init` + `terraform validate` against the actual `vkcs` provider schema (at several `router_count`/`private_network_cidrs` values), against a project-local filesystem-mirror copy of the provider - no cloud API is ever contacted and no credentials are needed;
|
|
- a real `terraform plan` against an isolated, provider-free copy of just `variables.tf`, to prove the `validation { ... }` blocks on `router_count` and `private_network_cidrs` (non-empty, valid CIDR syntax, uniqueness) are actually enforced - `terraform validate` alone does **not** enforce custom variable validations for externally-supplied values, only `plan`/`apply` do.
|
|
|
|
```bash
|
|
terraform/tests/setup-local-terraform.sh # one-time: provisions venv/ with terraform + the vkcs provider
|
|
venv/bin/pytest terraform/tests -v
|
|
```
|
|
|
|
`setup-local-terraform.sh` builds everything inside the git-ignored `venv/` directory:
|
|
|
|
- the Python packages from `terraform/tests/requirements.txt` (`pytest`, `python-hcl2`, `checkov`);
|
|
- the `terraform` CLI, downloaded (with `SHA256SUMS` verification) from a region-unrestricted HashiCorp releases mirror;
|
|
- the `vk-cs/vkcs` provider binary, downloaded (with `SHA256SUMS` verification) directly from its [GitHub releases](https://github.com/vk-cs/terraform-provider-vkcs/releases) - this bypasses `registry.terraform.io`, which blocks some regions outright, and is what makes a real `terraform validate` possible at all here;
|
|
- a project-local CLI config (`venv/terraform.d/cli-config.tfrc`) that points `terraform init` at that local provider copy via a `filesystem_mirror` block, instead of the network registry.
|
|
|
|
>Note: the diagrams below (`ports.svg`, `topology.svg`) and the Ansible layer (`ansible/inventory.ini`, roles `base`/`frr_router`/`keepalived`) still describe/assume the previous 2-router, 2-NIC, VRRP-based design and have **not** been updated for the new N-NIC/N-router topology yet - that's a separate follow-up.
|
|
|
|
**IPv4 addressing plan for the project**
|
|
|
|
Here is a card to assist with configuration planning. The card is filled out using the IP addressing from the Demo Scenario and the `inventory.ini` file, which will be used when running the Ansible playbook.
|
|
|
|

|
|
|
|
**Terraform**
|
|
|
|
Provisions `router_count` `IaaS Routers` (default 2), each with 1 public and `length(private_network_cidrs)` private ports (2 in the shipped example). Includes supplimentary Shell script template (which is a part of Terraform manifest) to maintain configuration across reboots.
|
|
|
|
**Ansible**
|
|
|
|
Configure `IaaS Routers` using role-based playbooks controlled via the [Inventory File](ansible/inventory.ini)
|
|
|
|
**Additional Software Used:**
|
|
|
|
- strongSwan (to manage IPsec)
|
|
- FRR (to manage BGP)
|
|
- Keepalived (VRRP)
|
|
|
|

|
|
|
|
Each IaaS Router will use two secured connections to On-Prem environment through the Internet:
|
|
|
|
- IPsec Site-to-Site in Transport Mode (to protect GRE Tunnels)
|
|
- GRE Tunnel (to transfer a data)
|
|
|
|

|
|
|
|
GRE Tunnels topology clearly explained in the following diagram:
|
|
|
|

|
|
|
|
**High Availability Design**
|
|
|
|
BGP peering eliminates single points of failure on the Cloud side through:
|
|
|
|
- Bidirectional eBGP sessions from each `IaaS Router` to On-Premises
|
|
- Optimized route metrics reflecting circuit priority (Primary/Backup)
|
|
- Automatic failover during circuit failures (including Cloud Availability Zone failures)
|
|
- Asymmetric routing prevention via MED and Local Preference configuration
|
|
|
|

|