Skip to main content

NFS Gateway Node Crash

Problem​

One NFS gateway (out of a 4-node cluster) crashed and dropped out of the cluster. The NFS mount remained reachable, but I/O operations stopped completing because the cluster was stuck waiting for the missing gateway to rejoin during the grace period.

Symptoms​

  • NFS mount is accessible from clients.
  • NFS operations (reads and writes) hang or do not go through.

Cause​

NFS-Ganesha's grace period was not lifted. The cluster was waiting for the crashed gateway to come back online before allowing normal operations to resume.

Resolution​

Unregister the crashed host from the RADOS grace period object so the remaining gateways can proceed without waiting for it.

Option 1 — via ganesha-rados-grace:

cephadm shell -- ganesha-rados-grace --pool .nfs --ns <nfs_cluster_namespace> lift <missing_gateway_id>

Option 2 — direct RADOS OMap key removal:

rados -p .nfs -N <nfs_cluster_name> rmomapkey grace <old_node_id>

Notes​

  • This clears the stale entry so the grace period can be lifted without waiting indefinitely for a node that is not coming back.
  • Confirm that the crashed host is genuinely down or removed before running these commands, to avoid clearing grace state for a node that is still expected to rejoin.