This project updates Debian-family Linux hosts through Ansible. It discovers tagged Ubuntu QEMU VMs from the Proxmox API, can also include manually listed non-Proxmox Debian/Ubuntu servers, creates pre-update snapshots for Proxmox VMs, applies safe apt upgrades, and reboots each host when the OS reports that a reboot is required.
The update playbook runs with serial: 1, so only one VM is being changed at a
time.
-
Create the control-node environment:
make bootstrap
-
Create your local environment file:
cp .env.example .env
-
Edit
.envwith your Proxmox API endpoint, token, and SSH user. -
In Proxmox, tag each VM that may be updated:
ansible-updateTo update a VM but suppress automatic reboots, also add:
no-reboot -
Make sure each target has:
- Debian or Ubuntu installed.
- QEMU guest agent installed and running for Proxmox VMs.
- SSH reachable from this control node.
- The configured SSH user allowed to run passwordless sudo.
Add external Debian or Ubuntu servers to inventory/static_hosts.yml. They use
the same ANSIBLE_SSH_USER and ANSIBLE_PRIVATE_KEY_FILE values from .env.
Example:
all:
children:
linux_update_targets:
children:
external_debian:
hosts:
debian-fileserver:
ansible_host: debian-fileserver.domain.com
no_reboot: trueValidate this file with:
make static-inventory-checkNon-Proxmox hosts are updated with apt like the Proxmox VMs, but snapshot and
snapshot-pruning tasks are skipped because they do not have Proxmox metadata.
Set no_reboot: true on a static host to suppress automatic reboots.
Create a dedicated SSH key for Ansible:
ssh-keygen -t ed25519 -f ~/.ssh/ansible_ed25519 -C ansible-homelabThen use your existing SSH account to create the ansible user on each VM:
scripts/create_ansible_user.sh --admin-user <your-current-user> --public-key ~/.ssh/ansible_ed25519.pub vm1 vm2 vm3If you have many hosts, put one hostname or IP per line in a file:
scripts/create_ansible_user.sh --admin-user <your-current-user> --hosts-file hosts.txtIf your current SSH access uses a non-default private key:
scripts/create_ansible_user.sh --admin-user <your-current-user> --identity-file ~/.ssh/current_admin_key --hosts-file hosts.txtThe script is idempotent. It creates or updates the user, installs the public
key, adds the user to sudo, and writes /etc/sudoers.d/90-ansible with
passwordless sudo. It also tests ssh -i ~/.ssh/ansible_ed25519 ansible@host sudo -n true after each host unless you pass --skip-verify. Afterward, set
these values in .env:
ANSIBLE_SSH_USER=ansible
ANSIBLE_PRIVATE_KEY_FILE=/Users/you/.ssh/ansible_ed25519
Test one VM directly:
ssh -i ~/.ssh/ansible_ed25519 ansible@vm1 sudo -n trueIf SSH connects but hangs during key exchange from this runner, leave
SSH_COMPATIBILITY_MODE=auto in .env. Update runs first check connectivity
with normal SSH options, then retry the preflight with:
-o ControlMaster=no -o KexAlgorithms=curve25519-sha256 -o HostKeyAlgorithms=ssh-ed25519 -o IPQoS=none
When the fallback succeeds, the playbook runs once with those compatibility options. To force the workaround for every scheduled run, set:
SSH_COMPATIBILITY_MODE=always
The SSH preflight has its own timeouts so a key-exchange stall does not leave the run looking frozen:
SSH_PREFLIGHT_CONNECT_TIMEOUT=10
SSH_PREFLIGHT_WALL_TIMEOUT=45
A host that fails the SSH preflight no longer cancels the run. The hosts that answered the preflight are updated, and each host that did not is logged as:
Skipping unreachable host: vm1 (SSH preflight [default] could not connect)
The run email then reports [PARTIAL] and names every skipped host along with
the SSH error that caused the skip, so a newly built VM that has not been
through make create-user yet does not block updates for the rest of the fleet.
The run still fails outright when no host is reachable. To restore the old
all-or-nothing behavior, set:
SSH_PREFLIGHT_REQUIRE_ALL=true
Every run refreshes known_hosts from the inventory before the SSH preflight
(scripts/refresh_known_hosts.py, also run by make known-hosts). New hosts
are added. A host whose key no longer matches, which is what happens when a VM
is rebuilt, is handled like this:
-
For a Proxmox VM the script reads the guest's own
/etc/ssh/ssh_host_*_key.pubthrough the QEMU guest agent. That path does not go over the SSH connection, so it cannot be spoofed by a man-in-the-middle. If the key inside the VM matches the key seen on the network, the oldknown_hostsentry is replaced and the run continues:Host key changed: vm1 (10.0.0.9): verified through the Proxmox guest agent; known_hosts updated. -
If the key inside the VM does not match, or the key cannot be verified (the API token lacks
VM.GuestAgent.FileRead, or the host is not a Proxmox VM),known_hostsis left alone. The preflight then skips the host and the run email explains why, with the new fingerprint and the exact commands to accept it by hand once you have confirmed it on the console:Host key changed: vm1 (10.0.0.9): NOT updated - could not be verified: ...
Set KNOWN_HOSTS_TRUST_CHANGED_KEYS=true in .env to accept unverifiable
changed keys automatically. That trades man-in-the-middle protection for
convenience, so prefer granting the guest-agent privilege instead. A key that
the guest itself disowns is never accepted, even with that flag.
Use a Proxmox API token with the narrowest permissions that work for your
cluster. For Proxmox VE 8, the practical role for /vms is usually:
VM.Audit VM.Snapshot VM.Monitor
For Proxmox VE 9, guest-agent privileges are more granular; use the guest-agent
audit privilege instead of VM.Monitor if your cluster rejects VM.Monitor.
Automatic rollback also needs VM.Snapshot.Rollback, and the default
post-rollback start behavior needs VM.PowerMgmt. Verifying a rebuilt VM's
new SSH host key through the guest agent needs VM.GuestAgent.FileRead (see
"Changed Host Keys"); without it a changed key is reported but not accepted.
On Proxmox VE 9 the full role therefore looks like:
pveum role modify Ansible -privs "Sys.Audit,VM.Audit,VM.GuestAgent.Audit,VM.GuestAgent.FileRead,VM.PowerMgmt,VM.Snapshot,VM.Snapshot.Rollback"If the token lacks guest-agent privileges on a particular VM, the inventory
script now logs an Inventory warning for that VM, drops it from the run, and
keeps every other host. Grant VM.GuestAgent.Audit on that VM to bring it back
into the update rotation.
The inventory script accepts token auth through:
PROXMOX_USER
PROXMOX_TOKEN_ID
PROXMOX_TOKEN_SECRET
For PROXMOX_TOKEN_ID, prefer only the token name, such as ansible. If you
paste a full user@realm!token value, the inventory script will use the part
after !.
List discovered hosts:
make inventoryCheck SSH connectivity:
make pingSeed SSH host keys for all discovered VMs:
make known-hostsCreate the dedicated VM user with the helper script:
make create-user EXTRA_ARGS="--admin-user <your-current-user> --hosts-file hosts.txt"Run a dry run:
make dry-runDry-run alert emails include [DRY RUN] in the subject and body so check-mode
failures are easy to distinguish from real update runs.
Run updates:
make updateApt tasks run noninteractively and have a one-hour timeout by default. Change
apt_task_timeout in group_vars/all.yml if a slow host needs more or less
time. Apt lock waits use apt_lock_timeout, which defaults to ten minutes. The
playbook runs apt-get update as a separate task so repository failures show
the underlying apt stdout/stderr in the run log.
Before upgrades, the playbook runs apt-get clean, checks free space under
/var/cache/apt/archives, and fails early if less than
apt_min_archive_free_mb is available. The default is 512 MB and can be
overridden globally in group_vars/all.yml or per host in inventory.
Limit to one host:
make update LIMIT=my-vm-nameOr pass raw Ansible arguments:
make update EXTRA_ARGS="--limit my-vm-name"
scripts/run_updates.sh --limit my-vm-nameInstall the update cron entry:
make install-cronThe installer prompts for daily, weekly, monthly, weekdays, or a custom cron expression. Press Enter through the prompts to keep the default weekly Sunday 03:00 schedule. For noninteractive setup:
make install-cron EXTRA_ARGS="--preset daily --time 04:15"
make install-cron EXTRA_ARGS="--preset weekly --weekday mon --time 02:30"
make install-cron EXTRA_ARGS="--preset custom --cron-expression '*/20 * * * *'"Logs are written to logs/update-YYYYMMDD-HHMMSS.log. Each scheduled run
re-runs the dynamic inventory and refreshes SSH known_hosts before applying
updates, so newly tagged Proxmox VMs are picked up automatically.
scripts/run_updates.sh sends an email after each run, including cron runs. It
uses the SMTP relay settings from .env. The body leads with the plain-English
per-host summary, then run metadata, then only the failing tasks and any
inventory warnings. The full Ansible log is attached rather than pasted inline
(gzipped when it is larger than 256 KB).
For your relay:
ALERT_EMAIL_ENABLED=true
ALERT_EMAIL_TO=you@example.com
ALERT_EMAIL_FROM=ansible-updates@domain.com
SMTP_RELAY_HOST=smtp.domain.com
SMTP_RELAY_PORT=587
SMTP_RELAY_STARTTLS=false
If ALERT_EMAIL_TO is unset, the run still completes and email is skipped. If
email sending fails, the script prints Email alert failed but still exits with
the original Ansible status.
When Debian or Ubuntu creates /var/run/reboot-required, the playbook reboots
that host automatically after updates. If a Proxmox VM has the Proxmox tag
no-reboot, or a static host has no_reboot: true, Ansible does not reboot it
and writes a message to the run log instead.
Each run creates snapshots named like:
ansible-pre-update-YYYYMMDD-HHMMSS
Snapshots do not include RAM (vmstate: false). The playbook keeps the newest
three snapshots created by this automation and prunes only snapshots whose names
start with ansible-pre-update-.
After package updates and any required reboot, the playbook verifies that SSH
recovers, passwordless sudo still works, apt-get check succeeds, and
dpkg --audit reports no interrupted package operations. If any update or
verification task fails on a Proxmox-discovered VM, Ansible rolls that VM back
to the pre-update snapshot created for that host and starts it afterward by
default.
Static non-Proxmox hosts run the same verification checks, but they cannot be rolled back automatically because they have no Proxmox snapshot metadata.
Rollback settings live in group_vars/all.yml:
post_update_health_timeout: 300
proxmox_rollback_start: true
proxmox_rollback_timeout: 300- This is for package updates, not Ubuntu release upgrades.
- Host key checking is enabled. If a VM is new, run
make known-hostsbeforemake pingormake update. A rebuilt VM's changed key is verified through the Proxmox guest agent; see "Changed Host Keys". - The cron job runs from this project directory and uses
.env, so keep that file readable only by the account running cron.