Operations

Day-to-day operations for the pitower Kubernetes cluster running on Talos Linux.


Overview

Cluster lifecycle is managed declaratively with topf. talos/pitower/ contains a topf.yaml (cluster identity, node inventory, Talos/Kubernetes versions, schematic references) plus layered machine-config patches. Secrets stay SOPS-encrypted on disk; topf decrypts secrets.sops.yaml transparently with the age key that the root mise.toml points SOPS_AGE_KEY_FILE at.

Run topf from talos/pitower as mise exec -- topf ..., or through the justfile recipes that wrap it. The Talos Apply workflow also runs topf apply for the whole cluster on every merge to main that touches talos/.

talosctl is still used for read-only diagnostics (logs, services, dashboards) and Kubernetes minor upgrades (talosctl upgrade-k8s).

flowchart TD
    Operator((Operator))
    CI[Talos Apply workflow]
    Just[justfile recipes]
    Topf[topf]
    Talosctl[talosctl]
    Kubectl[kubectl]
    ArgoCD[ArgoCD]

    Operator -->|mise exec -- topf| Topf
    Operator --> Just
    CI -->|on merge to main| Topf
    Just -->|apply, upgrade, reset, render| Topf
    Just -->|diagnostics| Talosctl
    Just -->|addons| Kubectl
    ArgoCD -->|sync| Kubectl

    Topf -->|sops + age| Secrets[(secrets.sops.yaml)]
    Topf --> Nodes[Talos Nodes]
    Talosctl --> Nodes
    Kubectl --> API[Kubernetes API]

Quick Reference

Run from talos/pitower. apply and upgrade prompt y/n per node; review --dry-run first and add the global --confirm=false only when running non-interactively.

TaskCommandDetails
Show node statusmise exec -- topf nodesStage, readiness, schematic, version
Preview config changesmise exec -- topf apply --dry-runExit 2 = changes pending
Apply configmise exec -- topf applyAll nodes, or --nodes-filter 'worker-0[12]' (regex)
Render configsmise exec -- topf renderWrite merged machine configs to output/
Upgrade Talosmise exec -- topf upgradeTo talosVersion/schematicId from topf.yaml
Check pending upgradesmise exec -- topf upgrade --dry-runExit 2 = upgrades due
Reset a nodejust talos pitower reset <name>Wipes STATE+EPHEMERAL, back to maintenance mode
Admin kubeconfigmise exec -- topf kubeconfig > output/kubeconfigShort-lived (12h), printed to stdout
Cluster healthjust talos pitower healthtalosctl health

Sections

PageDescription
Justfile RecipesReference of the justfile modules and Talos recipes
Talos CommandsCommon talosctl commands for health checks, logs, and debugging
TroubleshootingCommon issues and their resolutions
UpgradesProcedures for upgrading Talos, Kubernetes, and applications

Node Layout (pitower)

All nodes sit on VLAN 20. worker-01..03 are the control planes behind the API VIP 10.20.10.0.

IP AddressHostnameRoleHardware (schematic)
10.20.10.1worker-01Control PlaneAMD Ryzen mini PC (amd), Ceph OSD
10.20.10.2worker-02Control PlaneAMD Ryzen mini PC (amd), Ceph OSD
10.20.10.3worker-03Control PlaneAMD Ryzen mini PC (amd), Ceph OSD
10.20.10.4worker-04WorkerIntel (intel), tainted dedicated=media-home
10.20.10.5worker-05WorkerIntel (intel), also on the untagged LAN
10.20.10.6worker-06WorkerIntel (intel), also on the untagged LAN
10.20.10.7worker-07WorkerDell R630 (r630), ZFS pools, Garage, CI runners
10.20.10.8worker-08WorkerRaspberry Pi 4 (rpi-poe)
10.20.10.9worker-09WorkerRaspberry Pi 4 (rpi-poe)
10.20.10.10worker-10WorkerRaspberry Pi 4 (rpi-poe)
10.20.10.11worker-ai-01WorkerGPU workstation (nvidia), tainted dedicated=gpu