Monitoring

The observability stack runs in the monitoring namespace, each component its own ArgoCD application under kubernetes/apps/pitower/monitoring/. Prometheus scrapes metrics and remote-writes them to VictoriaMetrics, Fluent Bit ships container logs to VictoriaLogs, Tempo stores traces, and Grafana (run by the grafana-operator) queries all of them. Alertmanager routes alerts to Discord and ntfy.

Overview

flowchart LR
    subgraph Sources
        Apps[Applications\n/metrics]
        Logs[Container logs]
        Ext[External hosts\nOTLP]
        BB[Blackbox / smartctl /\nNUT / UniFi exporters]
    end

    subgraph Metrics
        Prom[Prometheus\n10d]
        VM[VictoriaMetrics\n2y]
    end

    subgraph Logs and Traces
        FB[Fluent Bit]
        VL[VictoriaLogs\n14d]
        OC[otel-collector]
        Tempo[Tempo\n72h]
    end

    AM[Alertmanager]
    Grafana[Grafana]
    Notify[Discord / ntfy]

    Apps & BB -->|ServiceMonitor / PodMonitor / Probe| Prom
    Prom -->|remote_write| VM
    Logs --> FB --> VL
    Ext -->|otlp.wibrow.dev| OC
    OC -->|traces| Tempo
    OC -->|logs| VL
    OC -->|metrics :8889| Prom
    Prom -->|alerts| AM --> Notify
    Prom & VM & VL & Tempo --> Grafana

Observability Strategy

  • Metrics: applications expose /metrics; ServiceMonitors, PodMonitors, Probes and ScrapeConfigs tell Prometheus what to scrape. Prometheus keeps 10 days locally and remote-writes everything to VictoriaMetrics, which keeps 2 years.
  • Logs: Fluent Bit runs on every node, tails /var/log/containers/, and pushes to VictoriaLogs (14 days), queried with LogsQL.
  • Traces: apps send OTLP to the OpenTelemetry Collector, which forwards traces to Tempo. The OpenTelemetry Operator can auto-instrument pods via the monitoring/otel-instrumentation Instrumentation.
  • Dashboards: GrafanaDashboard CRs, in any namespace, select the Grafana instance by label.
  • Alerting: Alertmanager sends everything to Discord; alerts labelled notify="ntfy" also go to ntfy as push notifications.
  • Uptime: Gatus checks every HTTPRoute it discovers; the public status page is a Cloudflare Worker.

Prometheus, VictoriaMetrics, VictoriaLogs and Tempo run on worker-07 (the only wibrow.dev/compute=true node) on openebs-hostpath-fast, the fast/extra ZFS dataset.

Components

ComponentChart / imageVersionPurpose
kube-prometheus-stackprometheus-community/kube-prometheus-stack91.9.0Prometheus, Alertmanager, operator, node-exporter, kube-state-metrics, rules
VictoriaMetricsvictoriametrics/victoria-metrics-single0.48.0Long-term metrics (2y), vm.wibrow.dev
Grafanagrafana-operator5.25.0Grafana instance, datasources and dashboards as CRs
VictoriaLogsvictoriametrics/victoria-logs-single0.13.10Log storage (14d)
Fluent Bitfluent/fluent-bit0.58.3Log collection from all nodes
OpenTelemetryopen-telemetry/opentelemetry-operator0.124.1Operator + otel collector, external OTLP ingest
Tempografana/tempo1.24.4Trace storage (72h, 20Gi)
Gatusapp-template, ghcr.io/twin/gatusv5.37.0Endpoint checks from HTTPRoute annotations, up.wibrow.dev
ntfyapp-template, binwiederhier/ntfyv2.28.0Push notifications, ntfy.wibrow.dev
ntfy-alertmanagerapp-template, xenrox/ntfy-alertmanagerlatestAlertmanager webhook to ntfy bridge
blackbox-exporterprometheus-community/prometheus-blackbox-exporter11.19.1ICMP/TCP/TLS/DNS/HTTP probes
smartctl-exporterprometheus-community/prometheus-smartctl-exporter0.17.1SMART data on every amd64 node, disk-health alerts
nut-exporterapp-template, druggeri/nut_exporter3.3.0UPS metrics from NUT
unpollerapp-template, unpoller/unpollerv5.5.0UniFi metrics
dozzleapp-template, amir20/dozzlev11.3.0Live container log viewer, dozzle.wibrow.dev
infra-healthPrometheusRule--infra:* recording rules behind the status page
garageScrapeConfig + PrometheusRule--Garage S3 metrics and alerts

Namespace Configuration

The monitoring namespace (created by the kube-prometheus-stack app) runs with privileged Pod Security Standards for node-exporter, Fluent Bit and smartctl-exporter, which need host access:

yaml
apiVersion: v1
kind: Namespace
metadata:
  name: monitoring
  labels:
    pod-security.kubernetes.io/audit: privileged
    pod-security.kubernetes.io/enforce: privileged
    pod-security.kubernetes.io/warn: privileged

Key Endpoints

ServiceURLGateway
Grafanahttps://grafana.wibrow.devenvoy-external
Prometheushttps://prometheus.wibrow.devenvoy-internal
Alertmanagerhttps://alertmanager.wibrow.devenvoy-internal
VictoriaMetricshttps://vm.wibrow.devenvoy-internal
Dozzlehttps://dozzle.wibrow.devenvoy-internal
Gatushttps://up.wibrow.devenvoy-external
ntfyhttps://ntfy.wibrow.devenvoy-external
OTLP ingesthttps://otlp.wibrow.devenvoy-external (API key)
Status pagehttps://status.wibrow.devCloudflare Worker (no origin)

Alert Silences

Before a cluster upgrade or other disruptive work (draining or rebooting nodes, storage/ZFS changes, gateway or CNI changes), silence Alertmanager so ntfy is not flooded, and expire the silence once the work is verified. See Prometheus Stack → Silences.

Key Design Decisions

  • VictoriaMetrics for history -- Prometheus stays small (10d) while VictoriaMetrics holds 2 years of everything Prometheus scrapes.
  • VictoriaLogs over Loki -- single binary, LogsQL, and fed directly by Fluent Bit's HTTP output.
  • Operator-managed Grafana -- datasources and dashboards are CRs, so any app can ship its own dashboard with a GrafanaDashboard.
  • Fluent Bit over Promtail -- low resource footprint and a flexible filtering pipeline.
  • Status page outside the cluster -- a Cloudflare Worker can still report when the cluster is down.