Skip to content

Repository files navigation

Terraform AWS Tailscale Module

This module deploys a Tailscale subnet router on AWS using an Auto Scaling Group and launch template. Instances authenticate directly with Tailscale using an OAuth client secret.

Tailscale Configuration

⚠️ Do this before deploying.

Please complete the following steps before running terraform apply for the first time:

  1. Configure the ACL
  2. Create an OAuth client

Otherwise the OAuth client can't be scoped to the right tag, and advertised routes get stuck pending instead of auto-propagating (see step 1 for why).

1. Configure the ACL

Do this step first, before creating the OAuth client. Two things depend on it:

  • An OAuth client scoped to auth_keys (used in step 2) can only be assigned an existing tag — a tag that doesn't exist yet cannot be selected.
  • autoApprovers only applies to route advertisements Tailscale receives after it's configured. Updating it later does not retroactively approve routes that are already pending — if the ACL isn't in place before the instance's first terraform apply, its advertised route will get stuck pending and must be manually removed and re-advertised.
  1. Open the Access Controls page in the Tailscale Admin Console.
  2. In the JSON editor, merge the following keys into the existing policy file — most tailnets already have other rules in place, so add to it rather than replacing the whole file:
{
  "tagOwners": {
    "tag:<environment>": []
  },
  "autoApprovers": {
    "routes": {
      "10.0.0.0/16": ["tag:<environment>"]
    }
  }
}
  1. Make sure the grants rules actually permit traffic to/from this tag, e.g.:
{
  "grants": [
    {
      "src": ["*"],
      "dst": ["*"],
      "ip": ["*"]
    }
  ]
}
  1. Save.

With autoApprovers configured, advertised routes are approved automatically — no manual approval in the admin console is needed.

The tag must be added to the ACL to disable automatic key expiration.

More info on tags: Tailscale ACL Tags.

2. Create an OAuth client

In the Tailscale Admin Console: Settings (top nav) → Trust credentials (left sidebar, under Tailnet Settings) → + Credential → OAuth.

Set the scope to Auth Keys — Write and select the tag created in step 1 (e.g. tag:<environment>).

Save the Client Secret — this is the only credential the module needs. See Storing the OAuth Secret further down for where to put it.

Usage

module "tailscale" {
  source  = "registry.terraform.io/hazelops/tailscale/aws"
  version = "~> 3.1"

  env                           = "prod"
  vpc_id                        = "vpc-0000000"
  subnets                       = ["subnet-aaa0001", "subnet-bbb0002"] # Two AZs for HA
  allowed_cidr_blocks           = ["10.0.0.0/16"]
  ec2_key_pair_name             = "default-key"
  tailscale_oauth_client_secret = "ts-xxxx1234567890"

  asg = {
    min_size = 2
    max_size = 2
  }
}

⚠️ Warning: Never hardcode secrets in Terraform code. Please fetch tailscale_oauth_client_secret at runtime from Secrets Manager or SSM Parameter Store.

More examples can be found in the examples directory.

Instance Sizing

The default instance_type is t4g.micro (1 GiB RAM) — don't go below this. Installing tailscale also pulls in iptables and its supporting libraries, and on t4g.nano (0.5 GiB RAM) this can exceed available memory on first boot and get killed by the OOM killer, silently leaving the instance without tailscale installed (cloud-init does not retry a failed module).

High Availability

Running two instances (e.g. asg = { min_size = 2, max_size = 2 }) in subnets across two Availability Zones provides automatic failover. Both instances advertise the same routes, and Tailscale handles failover transparently with ~15 seconds of downtime.

Each instance's Tailscale hostname includes a suffix from its EC2 instance ID (e.g. prod-tailscale-router-a1b2), so multiple concurrent instances are distinguishable in the Tailscale admin console instead of showing up with the same name.

Prerequisite: autoApprovers must be configured in the Tailscale ACL (see Configure the ACL). Without it, routes require manual approval and failover will not work automatically.

Recovery cycle: When one instance fails, Tailscale switches traffic to the surviving instance. The ASG detects the failure and launches a replacement. The new instance automatically joins the Tailnet, advertises the same routes, and becomes the standby node — restoring the HA pair without manual intervention.

Updates: The module adjusts rolling update behaviour based on asg.min_size. With a single instance (min_size = 1), the replacement is launched first and the old instance is terminated only after the new one is healthy — this temporarily runs two instances to avoid downtime. With two instances (min_size = 2), they are replaced one at a time so at least one is always active.

Exit Node

⚠️ Configure autoApprovers.exitNode before setting exit_node_enabled = true. Just like autoApprovers.routes (see step 1), this is not retroactive — it only applies to exit node advertisements Tailscale receives after it's configured. If the instance advertises itself as an exit node first, it stays pending even after the ACL is updated; toggling exit_node_enabled off and back on (or tailscale down/up on the instance) re-sends the advertisement so it can be auto-approved. Add this to the same policy file from step 1:

{
  "autoApprovers": {
    "exitNode": ["tag:<environment>"]
  }
}

Set exit_node_enabled = true to advertise this instance as a Tailscale exit node, routing all tailnet traffic (not just the advertised subnet routes) through it:

module "tailscale" {
  # ... other variables ...
  exit_node_enabled = true
}

Exit nodes must be approved before tailnet clients can select them. With autoApprovers.exitNode configured this happens automatically; if it isn't set up, approve the node manually in the Tailscale Admin Console.

Datadog Monitoring (Optional)

Enable Datadog Agent with a custom Tailscale health check:

module "tailscale" {
  # ... other variables ...
  datadog_enabled = true
  datadog_api_key = var.datadog_api_key
}

This installs Datadog Agent 7 and a custom check that emits:

  • tailscale.self.online (gauge, 0/1)
  • tailscale.peers.count (gauge)
  • tailscale.peers.online / tailscale.peers.direct / tailscale.peers.relayed (gauges) — tailnet-wide aggregates
  • tailscale.peer.online / tailscale.peer.rx_bytes / tailscale.peer.tx_bytes / tailscale.peer.direct / tailscale.peer.last_handshake_age_seconds (gauges, tagged peer:<hostname>)
  • tailscale.routes.primary_count (gauge) — subnets this node serves as primary subnet router (approved and active)
  • tailscale.health.issues (gauge) — total count of Tailscale health messages
  • tailscale.up (service check)

Service check states

The tailscale.up service check distinguishes real failures from benign health warnings, so it does not page on issues that are normal for a subnet router:

Node state Health messages tailscale.up
Online none OK
Online only benign (e.g. "advertising routes but --accept-routes is false") OK
Online real issue present (e.g. DERP region unreachable) WARNING (message lists the real issues only)
Offline any CRITICAL

Benign messages are filtered out before evaluation, so a subnet router that advertises routes without accepting others' does not trigger a false alert. tailscale.health.issues still reports the total count of health messages (including benign ones) for dashboard visibility.

Migrating from v2.x

  1. Remove api_token from the module call and replace with tailscale_oauth_client_secret.
  2. Drop the orphaned auth key from state (the Tailscale provider is no longer part of the module):
    terraform state rm module.tailscale.tailscale_tailnet_key.this
  3. Revoke the old auth key in the Tailscale Admin Console.
  4. Run terraform init -upgrade then terraform plan — expect only a launch template update followed by an instance refresh.

Security

IMDSv2

The launch template enforces IMDSv2 (http_tokens = required). This prevents unauthenticated access to the instance metadata endpoint (http://169.254.169.254), where EC2 user-data is served. Without IMDSv2, any process on the instance can retrieve user-data — including the Tailscale OAuth secret and Datadog API key — with a plain HTTP GET. IMDSv2 requires a session token obtained via a PUT request first, which blocks the most common SSRF-based metadata theft patterns.

EBS Encryption

The root EBS volume is encrypted by default using the AWS-managed key (alias/aws/ebs) — no extra setup required. To disable:

module "tailscale" {
  # ... other variables ...
  ebs_encrypted = false
}

There is no customer-supplied-KMS-key option: a customer-managed key would require granting the AWSServiceRoleForAutoScaling service-linked role explicit KMS permissions, which this module does not manage.

If using a custom ami_id: it must be an Amazon Linux 2023 (or AL2023-derived) AMI. The module's cloud-init content assumes a yum/dnf-based OS, and root-volume encryption assumes the AMI's root device is /dev/xvda — both are AL2023 defaults but are not validated by Terraform for arbitrary custom AMIs.

Storing the OAuth Secret

tailscale_oauth_client_secret accepts a plain string — store the Client Secret from step 2 above wherever fits the deployment. Two options:

SSM Parameter Store

aws ssm put-parameter \
  --name "/<env>/global/tailscale_oauth_client_secret" \
  --type "SecureString" \
  --value "<client-secret>"
data "aws_ssm_parameter" "tailscale_oauth_client_secret" {
  name = "/${var.env}/global/tailscale_oauth_client_secret"
}

module "tailscale" {
  # ... other variables ...
  tailscale_oauth_client_secret = data.aws_ssm_parameter.tailscale_oauth_client_secret.value
}

Secrets Manager

aws secretsmanager create-secret \
  --name "/<env>/global/tailscale_oauth_client_secret" \
  --secret-string "<client-secret>"
data "aws_secretsmanager_secret_version" "tailscale_oauth_client_secret" {
  secret_id = "/${var.env}/global/tailscale_oauth_client_secret"
}

module "tailscale" {
  # ... other variables ...
  tailscale_oauth_client_secret = data.aws_secretsmanager_secret_version.tailscale_oauth_client_secret.secret_string
}

Datadog Dashboard

A ready-made dashboard definition is available at datadog/datadog_dashboard.json, covering the metrics emitted when datadog_enabled = true: tailscale.up status, primary route ownership, health issues, peer connectivity, traffic by peer, and last handshake age.

To use it, click + New Dashboard in the Datadog UI, then open the settings (gear icon) menu on the new dashboard and select Import dashboard JSON to paste or upload the file — or create it via the Dashboards API.

Tailscale Router Datadog dashboard

Requirements

Name Version
terraform >=1.2.0
aws >=4.30.0

Providers

Name Version
aws >=4.30.0

Modules

No modules.

Resources

Name Type
aws_autoscaling_group.this resource
aws_iam_instance_profile.this resource
aws_iam_role.this resource
aws_iam_role_policy_attachment.this resource
aws_launch_template.this resource
aws_security_group.this resource
aws_ami.this data source
aws_iam_policy_document.this data source

Inputs

Name Description Type Default Required
allowed_cidr_blocks List of network subnets that are allowed. According to PCI-DSS, CIS AWS and SOC2 providing a default wide-open CIDR is not secure. list(string) n/a yes
ami_id Optional AMI ID for Tailscale instance. Otherwise the latest Amazon Linux 2023 AMI will be used. One might want to lock this down to avoid unexpected upgrades. Must be an Amazon Linux 2023 (or AL2023-derived) AMI: the module's cloud-init content assumes a yum/dnf-based OS, and root-volume encryption (see ebs_encrypted) assumes a /dev/xvda root device. string "" no
asg Scaling settings of an Auto Scaling Group
object({
min_size = number
max_size = number
})
{
"max_size": 1,
"min_size": 1
}
no
datadog_api_key Datadog API key (required if datadog_enabled is true) string "" no
datadog_enabled Whether to enable Datadog Agent monitoring on the Tailscale instance bool false no
ebs_encrypted Whether to encrypt the root EBS volume using the AWS-managed EBS key (alias/aws/ebs) bool true no
ec2_key_pair_name EC2 key pair name to use for Tailscale instance string n/a yes
env Environment name (typically dev/prod) string n/a yes
exit_node_enabled Whether to advertise this instance as a Tailscale exit node (--advertise-exit-node), allowing tailnet traffic to be routed through it bool false no
ext_security_groups External security groups to add to the Tailscale instance list(any) [] no
instance_type Type of Tailscale instance. Defaults to t4g.micro (1 GiB RAM) rather than t4g.nano (0.5 GiB): dnf install tailscale pulls in iptables-nft, libnetfilter_conntrack, and related dependencies, and on t4g.nano this can exceed available memory (RAM + swap) during first boot and get killed by the OOM killer, silently leaving the instance without tailscale installed (cloud-init does not retry). t4g.micro has enough headroom to avoid this reliably. string "t4g.micro" no
monitoring_enabled Whether to enable monitoring for the Auto Scaling Group bool true no
name Name for Tailscale instance string "tailscale-router" no
public_ip_enabled Whether to enable a public IP for Tailscale instance bool false no
ssm_role_arn SSM policy ARN to attach to the Tailscale instance IAM role string "arn:aws:iam::aws:policy/AmazonSSMManagedInstanceCore" no
subnets Subnets where the Tailscale instance will be placed. It is recommended to use a private subnet for better security. list(string) n/a yes
tags AWS tags for the Tailscale instance map(string) {} no
tailscale_oauth_client_secret Tailscale OAuth client secret string n/a yes
tailscale_tags List of Tailscale tags for the Tailnet device. It would be automatically tagged when it is authenticated with this key list(string) [] no
vpc_id VPC ID where the Tailscale instance will be placed string n/a yes

Outputs

Name Description
autoscaling_group_id n/a
name n/a
security_group_id n/a

About

No description, website, or topics provided.

Resources

Stars

2 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages