Skip to content

feat(atlas): Move Atlas management traffic to private WireGuard - #311

Merged
tanmoysrt merged 16 commits into
developfrom
feat/private-network
Oct 2, 2026
Merged

tanmoysrt merged 16 commits into
developfrom
feat/private-network

Conversation

@tanmoysrt

@tanmoysrt tanmoysrt commented Oct 1, 2026 •

Copy link
Copy Markdown
Member

Summary

Atlas now manages Metal hosts and service VMs only through WireGuard on the provider private network. Hosts accept no management traffic on their public address, and fewer public IPv4 addresses are needed.

Why

Service VMs used a public IPv4 address only so that Atlas could SSH to them, and Metal and host SSH were open on the public address. A private path closes those ports and saves public IPv4 addresses.

What changed

  • Atlas reaches the Metal API, host SSH and guest SSH only through private WireGuard.
  • A host firewall rejects management traffic on the public address, and guests cannot reach the provider private network.
  • Cargo and the IPv6 router need no public IPv4 address. Only the HTTP proxy keeps one.
  • Import Server adds a Scaleway or AWS server that Atlas did not create.
  • A development gateway lets an Atlas outside the provider network reach hosts.
  • The use_public_ip_for_metald setting is removed.

Validation

  • Tested on AWS and Scaleway: host setup, Metal API, guest SSH, service VMs without IPv4, and the development gateway.
  • The Redis-dependent tests still need a run in CI.

Atlas is a wg0 peer on each host with a tenant-0 address. The VM hook
passes packets to that address to Linux, which routes them through wg0.
Other tenants stay blocked by the tenant check that runs first.

configure accepts --controller and stores the address in a map. Without
the flag, the stored address stays.
metald.toml names the Atlas mesh address as wg_mesh.controller_address.
Metal passes it to atlas-wg-mesh configure on every host setup.
Atlas gets one WireGuard identity for each region, with the last
tenant-0 VM address. Atlas Settings stores the key, so a site backup
keeps it. atlas-wireguard creates it once and writes atlas0.conf with one
peer for each host and its tenant-0 VMs. A scheduler job and the new
wireguard-link setup step write the file again.

install-atlas-wireguard.sh adds a root timer that applies the file every
10 seconds and an input filter on atlas0. configure-wireguard.sh adds the
Atlas peer to the host wg0, and metald.toml names the controller address.
Host SSH uses the wg0 address after the WireGuard setup.
metald binds its control API to the host wg0 address, and Atlas calls it
there. The use_public_ip_for_metald setting and the provider listen
address are gone.

A new host input firewall admits SSH and port 9000 only from the Atlas
address on wg0, ports 9001 and 9002 only from hosts, and WireGuard only
from the private network. Metal drops guest packets to the private
network CIDR. Provider SSH calls use the server ssh_host.
A VM SSH connection now runs through its current host. The proxy command
SSHes to the host wg0 address and runs nc to the guest inside the Metal
namespace, so a guest needs no public address for Atlas SSH. Atlas reads
the host for each connection, so a migrated VM stays reachable.

SSH Task, the Cargo, IPv6 router and proxy setup, and Cargo admin
commands use the proxy. install-metald.sh installs nc.
Cargo and the IPv6 router no longer take a public IPv4 allocation, and
the router setup no longer waits for one. Only the HTTP proxy keeps a
public IPv4 address. Its guest firewall admits TCP 80, 443 and ICMP from
any address and all traffic from tenant-0 mesh addresses.

Proxy peers replicate over the mesh. The apply command writes each peer
name with its mesh address into /etc/hosts.
The site config key atlas_internal_url names an Atlas listener on its
mesh address. Hosts and tenant-0 VMs use it for Atlas files, Cargo for
the Atlas API, and Cargo and the proxies for the JWKS. Without it, they
use atlas_base_url, then the site URL.
Import Server on the Metal Server list, and the import-metal-server
command, add a Scaleway or AWS server by its provider ID. Atlas matches
the size and image to the catalog, adds a missing Scaleway private
network option, and runs the normal setup. An AWS instance outside the
Atlas subnet is refused. Importing the same server again continues its
setup.

An import can name its own storage pool device. atlas-vm can import its
own host after setup.
The Firecracker CI kernel has no TUN device, so wireguard-go cannot run
in the guest. atlas-vm now downloads the Ubuntu 24.04 vmlinuz, checks its
SHA-256, and extracts the ELF kernel with zstd.

The unit removes a stale API socket before Firecracker starts, so a
restart no longer fails, and the unit is enabled for boot. The FORWARD
rules name the filter table explicitly.
atlas-vm setup has a new stage: it runs atlas-wireguard and installs the
root timer for atlas0.conf, so a new Atlas VM reaches hosts through wg0
with no manual step. It installs wireguard-tools and wireguard-go,
because the guest kernel has no WireGuard module.
A development Atlas runs outside the provider network. atlas-vm gateway
mode runs a WireGuard relay VM on one host, and atlas-dev-gateway deploys
it over SSH and writes atlas-gateway.conf on the developer machine. The
atlas0 sessions pass through it end to end, so hosts expose no WireGuard
port publicly.
The Atlas WireGuard address and keys now have their own Wireguard Config
section on the Networking tab, next to the private network settings.
@tanmoysrt tanmoysrt changed the title feat(atlas): Reach hosts and service VMs through a private WireGuard network feat(atlas): Move Atlas management traffic to private WireGuard Oct 1, 2026
atlas-wireguard is now configure-atlas-wireguard, and atlas-dev-gateway
is now deploy-dev-gateway, like the other Atlas commands. The import
rollback carries a bare nosemgrep, and the host commands drop the
set_user call that frappe.connect already does.
@tanmoysrt
tanmoysrt force-pushed the feat/private-network branch from 3bedbb5 to 13a4f08 Compare October 1, 2026 20:45
The Ubuntu kernel that atlas-vm boots has no modules in the guest root
file system, so nft could not load nf_tables and the atlas0 filter
failed to install. Setup and gateway mode now install the modules
package for the running kernel. It also gives the guest kernel
WireGuard, so wireguard-go is no longer needed.
The image catalog keeps only the newest AMI of each version, so an
instance launched from an older AMI could not be imported. The import
now reads the instance AMI and matches its version, such as
Ubuntu_24.04.
A new provider host has no root key pair, so atlas-vm create stopped
and asked the operator to run ssh-keygen. atlas-vm now creates the root
key, and the gateway deploy no longer needs to.
@tanmoysrt
tanmoysrt added this pull request to stack #313 October 2, 2026 08:23
@tanmoysrt
tanmoysrt merged commit 6dd7b41 into develop Oct 2, 2026
7 checks passed
@tanmoysrt
tanmoysrt deleted the feat/private-network branch October 2, 2026 17:44
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant