Troubleshooting
Common issues and their resolutions for the cluster.
Node Not Joining Cluster
Symptoms: Node shows as not ready, or does not appear in kubectl get nodes.
Check Talos Health
talosctl health --nodes <node-ip>Look for failures in etcd, kubelet, or API server connectivity.
Check etcd Membership
talosctl etcd members --nodes 10.20.10.1If the node was previously part of the cluster and was reset, its stale etcd member entry may need to be removed:
talosctl etcd remove-member <member-id> --nodes 10.20.10.1Verify Machine Config
Ensure the node has the correct machine config applied (from talos/pitower; exit 2 = changes pending):
mise exec -- topf apply --dry-run --nodes-filter '^<hostname>$'Check kubelet-csr-approver
New nodes need their CSRs approved. Verify the kubelet-csr-approver is running:
kubectl get pods -n kube-system -l app.kubernetes.io/name=kubelet-csr-approver
kubectl get csrPod Stuck in Pending State
Symptoms: Pod stays in Pending status and never gets scheduled.
Check Node Resources
kubectl describe node <node-name> | grep -A10 "Allocated resources"
kubectl top nodesCheck Storage
If the pod requires a PVC, verify the storage class and available capacity:
kubectl get pvc -n <namespace>
kubectl describe pvc <pvc-name> -n <namespace>For Rook Ceph:
kubectl -n rook-ceph exec deploy/rook-ceph-tools -- ceph status
kubectl -n rook-ceph exec deploy/rook-ceph-tools -- ceph osd dfFor OpenEBS (local PV), check the provisioner and that the claim's node is allowed by the StorageClass (openebs-hostpath-fast, -runners and -media only exist on worker-07, -models on worker-ai-01):
kubectl -n openebs logs deploy/openebs-localpv-provisioner --tail=50
kubectl get sc <class> -o yaml | yq '.allowedTopologies'Check Pod Events
kubectl describe pod <pod-name> -n <namespace>
kubectl get events -n <namespace> --sort-by='.lastTimestamp' | tail -20Check Node Taints
kubectl get nodes -o custom-columns=NAME:.metadata.name,TAINTS:.spec.taintsDNS Not Resolving
Symptoms: Services cannot resolve DNS names, or external DNS records are not created.
UniFi DNS
Verify with DoH (DNS over HTTPS)
To check actual Cloudflare DNS records, bypass the router's interception using DoH:
# Using curl to query Cloudflare DoH
curl -sH 'accept: application/dns-json' \
'https://cloudflare-dns.com/dns-query?name=<app>.wibrow.dev&type=A' | jq
# Using dig with DoH (if supported)
dig @1.1.1.1 <app>.wibrow.dev +httpsCheck CoreDNS
kubectl get pods -n kube-system -l k8s-app=kube-dns
kubectl logs -n kube-system -l k8s-app=kube-dns --tail=50Check external-dns
Two instances run in networking: external-dns (Cloudflare) and external-dns-unifi (the UniFi gateway).
kubectl -n networking logs deploy/external-dns --tail=50
kubectl -n networking logs deploy/external-dns-unifi --tail=50Verify external-dns is watching the correct gateways:
kubectl get gateways -A -l external-dns.alpha.kubernetes.io/enabled=trueCheck HTTPRoute and Gateway
kubectl get httproutes -A
kubectl get gateways -ACertificate Issues
Symptoms: TLS errors, expired certificates, or certificates not being issued.
Check cert-manager
kubectl get pods -n cert-manager
kubectl logs -n cert-manager deploy/cert-manager --tail=50Check Certificate Status
kubectl get certificates -A
kubectl get certificaterequests -A
kubectl get orders.acme.cert-manager.io -A
kubectl get challenges.acme.cert-manager.io -ACheck ClusterIssuer
kubectl get clusterissuers
kubectl describe clusterissuer letsencrypt-productionForce Certificate Renewal
Delete the certificate to trigger re-issuance:
kubectl delete certificate <cert-name> -n <namespace>Service Not Accessible
Symptoms: Cannot reach a service via its URL, connection timeouts, or 404 errors.
Check Gateway Status
kubectl get gateways -n networking
kubectl describe gateway envoy-external -n networking
kubectl describe gateway envoy-internal -n networkingCheck HTTPRoute
kubectl get httproutes -A
kubectl describe httproute <route-name> -n <namespace>Verify the route's parentRefs point to the correct gateway:
- envoy-external (
10.20.10.239): public services, reached from the internet through the towonel tunnel - envoy-internal (
10.20.10.238): LAN and Tailscale only
Check LoadBalancer Announcements
LoadBalancer IPs (10.20.10.128-255) are announced by Cilium over L2 (ARP) and BGP to the UniFi gateway:
kubectl get svc -A | grep LoadBalancer
kubectl get ciliuml2announcementpolicies,ciliumbgpclusterconfigs,ciliumbgppeerconfigs
kubectl -n kube-system exec ds/cilium -c cilium-agent -- cilium-dbg shell -- bgp/peersCheck the Tunnel
For externally exposed services:
kubectl get pods -n networking -l app.kubernetes.io/name=towonel-agent
kubectl logs -n networking -l app.kubernetes.io/name=towonel-agent --tail=50See Towonel Tunnel for the hub side on ovh-vps.
End-to-End Request Flow
flowchart LR
Client -->|"DNS: Cloudflare, unproxied CNAME"| Hub[towonel hub<br/>ovh-vps]
Hub --> Tunnel[towonel-agent]
Tunnel --> EE[envoy-external<br/>10.20.10.239]
EE --> Route[HTTPRoute]
Route --> Svc[Service]
Svc --> Pod[Pod]
Verify each hop in the chain to isolate where the failure occurs.
ArgoCD Sync Failed
Symptoms: Application shows OutOfSync, Degraded, or Unknown in ArgoCD.
Check Application Status
kubectl get applications -n argocd
kubectl describe application <app-name> -n argocdArgoCD CLI
argocd app list
argocd app get <app-name>
argocd app diff <app-name>Common Sync Failures
Resource Already Exists
If a resource was manually created, ArgoCD may fail to adopt it:
argocd app sync <app-name> --forceSchema Validation Errors
CRDs may not be installed yet when the app tries to sync:
# Check if CRDs exist
kubectl get crds | grep <crd-name>
# Sync CRDs first if needed
argocd app sync <crd-app-name>Health Check Failures
Check pod health and events:
kubectl get pods -n <namespace> -l app.kubernetes.io/name=<app>
kubectl describe pod <pod-name> -n <namespace>
kubectl get events -n <namespace> --sort-by='.lastTimestamp'Helm Template Errors
For apps using Helm, test rendering locally:
cd kubernetes/apps/pitower/<category>/<app>
kustomize build . --enable-helmStorage Issues
Rook Ceph Degraded
kubectl -n rook-ceph exec deploy/rook-ceph-tools -- ceph status
kubectl -n rook-ceph exec deploy/rook-ceph-tools -- ceph health detail
kubectl -n rook-ceph exec deploy/rook-ceph-tools -- ceph osd treePVC Stuck in Pending
kubectl get pvc -A | grep Pending
kubectl describe pvc <pvc-name> -n <namespace>
kubectl get sckopiur Backup Failures
kubectl get snapshotpolicies.kopiur.home-operations.com -A
kubectl -n <namespace> get snapshots.kopiur.home-operations.com
kubectl -n <namespace> describe snapshot.kopiur.home-operations.com <name>See Backup & Restore. A mover that cannot read the app's files usually needs spec.mover.securityContext set to the app's uid/gid.
Network Issues
Pod-to-Pod Communication
# Test from a debug pod
kubectl run -it --rm debug --image=busybox -- sh
# Inside the pod:
wget -qO- http://<service>.<namespace>.svc.cluster.local:<port>Cilium Connectivity
cilium connectivity test
cilium status --verboseEnvoy Gateway Logs
kubectl logs -n networking deploy/envoy-gateway --tail=50