NFS Gateway Node Crash
Problem
One NFS gateway (out of a 4-node cluster) crashed and dropped out of the cluster. The NFS mount remained reachable, but I/O operations stopped completing because the cluster was stuck waiting for the missing gateway to rejoin during the grace period.
Symptoms
- NFS mount is accessible from clients.
- NFS operations (reads and writes) hang or do not go through.
Cause
NFS-Ganesha's grace period was not lifted. The cluster was waiting for the crashed gateway to come back online before allowing normal operations to resume.
Resolution
Unregister the crashed host from the RADOS grace period object so the remaining gateways can proceed without waiting for it.
Option 1 — via ganesha-rados-grace:
cephadm shell -- ganesha-rados-grace --pool .nfs --ns <nfs_cluster_namespace> lift <missing_gateway_id>
Option 2 — direct RADOS OMap key removal:
rados -p .nfs -N <nfs_cluster_name> rmomapkey grace <old_node_id>
Notes
- This clears the stale entry so the grace period can be lifted without waiting indefinitely for a node that is not coming back.
- Confirm that the crashed host is genuinely down or removed before running these commands, to avoid clearing grace state for a node that is still expected to rejoin.