CephFS mounts unavailable

Minor incident Zone LPG 2 Storage
2026-09-01 12:03 CEST · 4 hours, 9 minutes

Updates

Resolved

We’ve not observed any further problems with CephFS mounts. We will continue to monitor the situation.

September 1, 2026 · 16:11 CEST
Monitoring

We’ve now rebooted all worker nodes to clear stale CephFS mounts and are currently monitoring the cluster.

The root cause for the incident was that we applied the last step (rotating the CSI driver CephX keys) for mitigating Ceph CVE-2025-30156 before all worker nodes were rebooted during the maintenance. Because of that, existing CephFS mounts broke because their CephX keys were dropped in the mitigation step.

Timeline:
10:17 Maintenance engineer applies Rook-Ceph config that rotates the CSI driver CephX keys without retaining any previous keys.
11:00 Internal notification that a VSHN application is unavailable.
11:40 Internal notification that the VSHN application has been unavailable for longer than usual.
11:45 Engineer investigates the VSHN application and restarts pods.
11:49 First customer reports about applications issues.
~12:00 VSHN engineer realizes that all applications that exhibit issues have CephFS volumes as the common factor.
12:02 VSHN engineer realizes that CSI CephX keys got rotated without retaining any previous keys.
12:05 VSHN engineer restarts all CSI driver pods to ensure that CSI drivers use latest keys for new mounts. Starting from this point, new pods that use CephFS mounts work again unless they’re scheduled on a node where the target CephFS volume is already mounted with an invalidated key.
12:10 - 14:30: VSHN engineers observe maintenance and ensure that all worker nodes are rebooted as quickly as possible to clear out any stale CephFS mounts.
15:00: VSHN engineers are monitoring the cluster to ensure that all CephFS mounts work normally again.

September 1, 2026 · 15:12 CEST
Update

We’ve identified that the quickest way to clean up stale CephFS mounts is to reboot nodes. We’re currently actively monitoring the scheduled maintenance to ensure that node reboots happen as quickly as possible.

September 1, 2026 · 12:58 CEST
Update

We’ve restarted the CephFS CSI drivers, new CephFS mounts should work again. Pods with old CephFS mounts may get stuck in Terminating or Completed states.

We are currently determining what’s necessary to correctly clean up stuck pods.

September 1, 2026 · 12:12 CEST
Investigating

We’re investigating an issue with CephFS mounts on APPUiO LPG 2.

September 1, 2026 · 12:03 CEST

← Back