Skip to main content

Ceph Tentacle (v20)

This document lists known critical bugs affecting Ceph Tentacle (v20) releases.

Cephadm OSD Key Rotation Bug Causes Unnecessary OSD Restarts​

Severity: Critical Affected Versions: v19.2.6, v20.2.4 Bug Report: https://tracker.ceph.com/issues/80440

Description​

When a user or cephadm issues an ceph orch daemon reconfig osd.X command on an OSD, it triggers an unnecessary restart of that OSD. The issue is triggered automatically if the manager communicates a new monmap to the OSDs. An example of this is adding a monitor, which triggers a reconfig operation on an OSD. Such a reconfig operation is normal behavior in all versions of Ceph, but added logic in this (the affected) version causes the OSD always to restart.

On an active cluster, this can lead to multiple OSDs starting randomly, and that can cause PGs to become inactive. Sometimes this occurs during a MGR failover. This is always triggered when adding or removing a monitor.

This can be reproduced manually by running the following command:

ceph orch daemon reconfig osd.2 # restarts osd.2, which it shouldn't do

If this has happened in your cluster, it will show up in the cephadm.log of the host's OSD, or in logs that are returned by the command ceph -W cephadm. The following excerpt shows the text that confirms that your OSD key has been rotated:

[DBG] err: Reconfig daemon osd.2 ...
Stopping osd.2 to update osd_key bluestore label
Rotating osd.2 key with ceph-bluestore-tool
Successfully rotated osd.2 keyring

Recommendation​

  • Be aware that mgr failover, monitor add/remove operations, and post-upgrade key rotation can trigger unexpected OSD restarts — plan maintenance windows accordingly.
  • Watch the manager logs for the "Reconfig daemon" / "Rotating ... key" sequence shown above to confirm whether an OSD is being affected by this bug.

OSD Crash When Enabling EC Optimizations on CephFS​

Severity: High Affected Versions: 20.2.0 Bug Report: https://tracker.ceph.com/issues/71642

Description​

OSDs crash when allow_ec_optimizations is enabled on an existing CephFS Erasure Coded (EC) data pool that does not have allow_ec_overwrites explicitly enabled. The crash occurs in ECTransaction::WritePlanObj when accessing a non-existent transaction key.

Recommendation​

New or Recreated OSD Missing DB Device with Hybrid Spec​

Severity: Medium Affected Versions: 20.2.0, 20.2.1
Bug Report: https://tracker.ceph.com/issues/72696

Description​

When using a hybrid OSD spec (HDDs for data, SSDs/NVMe for DB devices), newly created or recreated OSDs are deployed without a DB device. The ceph-volume lvm batch command issued by the orchestrator omits the --db-device argument.

The root cause is a regression introduced by the fix for tracker #68576: the ceph_device attribute in ceph-volume inventory JSON output was renamed to ceph_device_lvm. This causes the selector to treat all existing RocksDB volumes as unavailable, so it cannot assign DB devices when recreating OSDs.

Note: This bug also affects Ceph Squid (v19). See the Squid known bugs page for details.

Recommendation​

  • A patch is available in PR #65986
  • Follow the bug tracker for fix and backport updates

FastEC Scrub Errors After Recovery​

Severity: Medium Affected Versions: 20.2.0 Bug Report: https://tracker.ceph.com/issues/73184

Description​

After recovering from the EC optimization crash, clusters may experience excessive scrub errors with messages like candidate size X info size Y mismatch. This is a secondary issue related to the FastEC code path.

Recommendation​