Rook Ceph
Rook Ceph provides distributed block storage for the cluster. Three SATA SSDs, one in each control-plane node (worker-01/02/03), form a Ceph cluster managed by the Rook operator. ceph-block is the cluster's default StorageClass.
Architecture
flowchart TD
subgraph Operators
OP[rook-ceph-operator\nManages Ceph lifecycle]
CSIOP[ceph-csi-operator\nManages CSI drivers]
end
subgraph Ceph Cluster
MON1[MON\nworker-01]
MON2[MON\nworker-02]
MON3[MON\nworker-03]
MGR[MGR\nCluster manager]
OSD1[OSD\nworker-01\nSATA SSD]
OSD2[OSD\nworker-02\nSATA SSD]
OSD3[OSD\nworker-03\nSATA SSD]
end
subgraph Resources
BP[CephBlockPool\nceph-blockpool\nReplicated x3]
SC[StorageClass\nceph-block]
VSC[VolumeSnapshotClass\ncsi-ceph-blockpool]
DASH[Dashboard\nrook.wibrow.dev]
end
OP --> MON1 & MON2 & MON3
OP --> MGR
OP --> OSD1 & OSD2 & OSD3
BP --> OSD1 & OSD2 & OSD3
SC --> BP
VSC --> BP
CSIOP -->|rbd Driver| SC
MGR --> DASH
Repository Layout
The deployment is split into four directories under kubernetes/apps/pitower/rook-ceph/, each its own ArgoCD application:
kubernetes/apps/pitower/rook-ceph/
├── operator/ # rook-ceph chart: operator, CRDs, ceph-csi-operator
├── csi-drivers/ # ceph-csi-drivers chart: OperatorConfig + rbd Driver CR
├── cluster/ # rook-ceph-cluster chart: CephCluster, pool, StorageClass, dashboard HTTPRoute
└── add-ons/ # Ceph Grafana dashboards (GrafanaDashboard CRs)Operator
The Rook operator is deployed via the rook-ceph Helm chart (v1.20.8) into the rook-ceph namespace:
image:
repository: ghcr.io/rook/ceph
crds:
enabled: true
monitoring:
enabled: true
nodeSelector: &cephControllerNodes
node-role.kubernetes.io/control-plane: ""
ceph-csi-operator:
controllerManager:
nodeSelector: *cephControllerNodesKey decisions:
- CRDs managed by the chart:
crds.enabled: trueinstalls and upgrades the CRDs with the operator. monitoring.enabled: true: grants the operator RBAC to create therook-ceph-mgr/rook-ceph-exporterServiceMonitors and PrometheusRules. Without it noceph_*metrics are scraped.- Controllers stay on the OSD hosts: the chart only exposes
nodeSelector, so the operator and ceph-csi-operator ride the control-plane role label (today exactly worker-01/02/03).
CSI Drivers
Since Rook v1.20 the operator chart no longer configures the CSI drivers; the ceph-csi-operator does. csi-drivers/ deploys the ceph-csi-drivers chart (1.1.0), which renders the OperatorConfig and the rbd Driver CR:
- Only RBD is enabled; CephFS, NFS and NVMe-oF drivers are disabled.
- The RBD node plugin tolerates every
dedicated=*taint, so volumes can mount wherever pods land. - The controller plugin (provisioner, attacher, snapshotter, resizer) runs 2 replicas pinned to worker-01/02/03.
CephConnection/ClientProfileCRs are left to Rook, which owns the liverook-cephones.
Cluster Configuration
The rook-ceph-cluster Helm chart (v1.20.8) deploys the CephCluster custom resource.
Monitors and Managers
| Component | Count | Purpose |
|---|---|---|
| MON | 3 | Maintain cluster map consensus (one per node) |
| MGR | 1 | Cluster management, dashboard, metrics |
| OSD | 3 | One per SSD, stores actual data |
Placement
placement.all pins every Ceph daemon to worker-01/02/03 and tolerates the node-issue and dedicated=storage taints. The toolbox is a chart-owned Deployment, so it carries the same node affinity separately.
Storage Nodes
Each OSD is pinned to a specific device by disk ID:
cephClusterSpec:
storage:
useAllNodes: false
useAllDevices: false
config:
osdsPerDevice: "1"
nodes:
- name: "worker-01"
devices:
- name: "/dev/disk/by-id/ata-SanDisk_SD7SB2Q512G1001_153141400183"
- name: "worker-02"
devices:
- name: "/dev/disk/by-id/ata-Samsung_SSD_850_EVO_500GB_S21JNSAG160704R"
- name: "worker-03"
devices:
- name: "/dev/disk/by-id/ata-KINGSTON_SA400S37480G_50026B7384393346"Network
cephClusterSpec:
network:
provider: hostHost networking is used for Ceph daemons to maximize throughput and minimize latency between OSDs and monitors.
cephx Key Rotation
Daemon keys are rotated with security.cephx.daemon.keyRotationPolicy: KeyGeneration; bump keyGeneration to rotate. Rook rolls the daemons one at a time, so the pool must be active+clean first. The CSI client keys deliberately stay on aes until every node runs a Linux kernel >= 7.0; the reasoning and the muted health checks are documented in cluster/values.yaml.
Resource Limits
| Daemon | CPU Request | Memory Request | Memory Limit |
|---|---|---|---|
| MGR | 73m | 283Mi | 2Gi |
| MON | 30m | 512Mi | 1Gi |
| OSD | 48m | 1Gi | 6Gi |
| MGR Sidecar | 49m | 128Mi | 256Mi |
| Crash Collector | 15m | 7Mi | 64Mi |
| Log Collector | 25m | 100Mi | 1Gi |
Storage Resources
CephBlockPool
The chart's default ceph-blockpool pool replicates three ways with host as the failure domain, so each OSD node holds one copy.
StorageClass
The chart creates the ceph-block StorageClass (the cluster default) that provisions RBD volumes from the pool:
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
name: my-app-data
spec:
accessModes:
- ReadWriteOnce
storageClassName: ceph-block
resources:
requests:
storage: 10GiDashboard
The Ceph dashboard is enabled and exposed via the internal gateway:
apiVersion: gateway.networking.k8s.io/v1
kind: HTTPRoute
metadata:
name: rook-ceph-dashboard
namespace: rook-ceph
spec:
hostnames:
- rook.wibrow.dev
parentRefs:
- name: envoy-internal
namespace: networking
sectionName: https
rules:
- backendRefs:
- name: rook-ceph-mgr-dashboard
port: 7000Access the dashboard at https://rook.wibrow.dev from the LAN or via Tailscale.
Monitoring
The cluster chart sets monitoring.enabled: true and createPrometheusRules: true, so Prometheus scrapes the mgr and exporter and the stock Ceph alert rules are installed.
Grafana Dashboards
add-ons/dashboard/ ships three dashboard JSON files as ConfigMaps, each referenced by a GrafanaDashboard CR in the Storage folder:
| Dashboard | Grafana ID | Purpose |
|---|---|---|
| Ceph Cluster | 2842 | Overall cluster health, IOPS, throughput |
| Ceph OSD | 5336 | Per-OSD performance and utilization |
| Ceph Pools | 5342 | Pool-level statistics and capacity |
Toolbox
The Rook toolbox pod is enabled (toolbox.enabled: true) for interactive Ceph CLI troubleshooting:
kubectl -n rook-ceph exec -it deploy/rook-ceph-tools -- bash
# Inside the toolbox
ceph status
ceph osd status
ceph df
rados dfHealth Checks
Common commands to verify Ceph cluster health:
Quick Status
kubectl -n rook-ceph exec deploy/rook-ceph-tools -- ceph statusOSD Tree
kubectl -n rook-ceph exec deploy/rook-ceph-tools -- ceph osd df treePool Usage
kubectl -n rook-ceph exec deploy/rook-ceph-tools -- ceph dfPG Status
kubectl -n rook-ceph exec deploy/rook-ceph-tools -- ceph pg stat