Node Management

Day-to-day operations for managing Talos nodes in the cluster. Lifecycle changes go through topf via just talos pitower <recipe> (recipes in talos/talos.justfile and talos/pitower/justfile); talosctl is for diagnostics. Run recipes from the repository root so mise.toml sets SOPS_AGE_KEY_FILE and KUBECONFIG.

Applying Configuration Changes

Through CI

The Talos Apply GitHub Actions workflow runs topf apply for the whole cluster on every merge to main that touches talos/. It runs on the self-hosted home-ops runners, which are pinned to worker-07: while worker-07 is drained or down the job queues, so apply locally instead. A change that reboots worker-07 also kills the job mid-run; finish it with a local apply.

Locally

bash
just talos pitower diff                     # pending changes (topf apply --dry-run)
just talos pitower apply                    # apply to all nodes
just talos pitower apply 'worker-0[12]'     # apply to a regex subset
just talos pitower apply-workers            # worker-04..10 and worker-ai-01

topf applies to control planes one at a time and waits for each node to stabilize before moving on. The justfile sets TOPF_CONFIRM=false, so recipes do not prompt; running topf directly from talos/pitower prompts y/n per node unless you pass --confirm=false.

Rebooting Nodes

bash
just talos pitower reboot-controlplanes   # 10.20.10.1-3, one at a time with --wait
just talos pitower reboot-workers         # 10.20.10.4-11, one at a time with --wait

These call talosctl reboot --wait with the talosconfig from just talos pitower talosconfig. To reboot one node:

bash
talosctl --talosconfig talos/pitower/output/talosconfig -n 10.20.10.7 reboot --wait

Resetting Nodes

To wipe and reset a node (for reprovisioning or troubleshooting):

bash
just talos pitower reset worker-05

This runs topf reset --nodes-filter '^worker-05$' --full=false, which wipes the STATE and EPHEMERAL partitions on the install disk and reboots the node into maintenance mode. topf's default graceful reset cordons and drains the node and leaves etcd first (for control planes). Data on other disks, such as worker-07's ZFS pools or worker-ai-01's models volume, is not touched.

Upgrading Nodes

Upgrades swap the Talos OS image atomically. topf upgrades each node to the talosVersion in topf.yaml with that node's schematic, cordoning and draining it first and uncordoning once it is Ready again.

Upgrade Flow

flowchart LR
    A[Bump talosVersion<br/>in topf.yaml] --> B[Submit any changed<br/>schematics]
    B --> C[upgrade-check]
    C --> D[Upgrade control planes<br/>one at a time]
    D --> E[Upgrade workers]
    E --> F[Verify cluster<br/>health]
bash
just talos pitower upgrade-check           # preview (exit 2 = upgrades due)
just talos pitower upgrade-controlplanes   # worker-01..03
just talos pitower upgrade-workers         # worker-04..10, worker-ai-01
just talos pitower upgrade                 # everything, or pass a regex filter

Image Management

bash
just talos pitower image-list 10.20.10.1   # images on a node, sorted by size
just talos pitower image-usage             # image count and containerd disk usage per node

Image garbage collection is configured in all/01-general.yaml: usage thresholds of 60% / 50% plus imageMaximumGCAge: 168h, so unused images are evicted after a week regardless of disk size.

Node-Specific Patches

Each node has a directory under talos/pitower/node/:

talos/pitower/node/
  worker-01/01-install.yaml      # AMD control plane
  worker-02/01-install.yaml      # AMD control plane
  worker-03/01-install.yaml      # AMD control plane
  worker-04/01-install.yaml      # Intel, eMMC, media-home taint
  worker-05/01-install.yaml      # Intel
  worker-06/01-install.yaml      # Intel
  worker-07/01-install.yaml      # Dell R630: RAID1 boot, ZFS
  worker-08/01-install.yaml      # Raspberry Pi 4
  worker-09/01-install.yaml      # Raspberry Pi 4
  worker-10/01-install.yaml      # Raspberry Pi 4
  worker-ai-01/01-install.yaml   # GPU node: NVIDIA/VFIO modules, gpu taint
  worker-ai-01/02-models.yaml    # models user volume
  worker-ai-01/03-network.yaml   # ignore the Bazzite NIC

The hostname, MAC address, VLAN 20 interface, and VIP come from topf.yaml through the all/*.tpl templates, so node patches only hold what is unique to that machine. See Talos Linux for what each one sets.

Adding or Renumbering a Node

  1. Add a DHCP reservation in terraform/unifi/reservations.tf.
  2. Add the node to nodes: in talos/pitower/topf.yaml (and a node/<host>/ directory if it needs one).
  3. Add its IP to the BGP neighbor list in terraform/unifi/frr-bgp.conf and push it with mise run unifi:bgp-upload from terraform/.
  4. Add the IP to nodes in talos/pitower/justfile (used by the diagnostics recipes).
  5. Boot the node into maintenance mode and run just talos pitower apply '^<host>$'.

Quick Reference

OperationCommand
Node statejust talos pitower status
Render configsjust talos pitower render
Preview changesjust talos pitower diff
Apply to all / subsetjust talos pitower apply [regex]
Apply to workersjust talos pitower apply-workers
Reboot control planesjust talos pitower reboot-controlplanes
Reboot workersjust talos pitower reboot-workers
Reset a nodejust talos pitower reset <name>
Preview upgradesjust talos pitower upgrade-check
Upgrade all / subsetjust talos pitower upgrade [regex]
Upgrade control planesjust talos pitower upgrade-controlplanes
Upgrade workersjust talos pitower upgrade-workers
Admin kubeconfig (12h)just talos pitower kubeconfig
Generate talosconfigjust talos pitower talosconfig
Schematic IDsjust talos pitower schematic-ids
Bootstrap addonsjust talos pitower addons
Images on a nodejust talos pitower image-list <node-ip>
Image usage summaryjust talos pitower image-usage
Node uptimejust talos pitower uptime
Cluster healthjust talos pitower health