Skip to main content

Stretch Mode PGs Inactive with Device-Class CRUSH Rules

Known issue. Affects every release with stretch mode, 16.2.0 through 19.2.6 and 20.2.4 at the time of writing.

Problem​

A stretch-mode cluster loses one site, or the link between sites. Pools whose CRUSH rule selects a device class become unavailable: their PGs stay undersized+peered instead of active. Pools on rules without a device class keep serving I/O.

When the site comes back, those PGs stay inactive (clean+peered). ceph osd dump keeps reporting degraded_stretch_mode 1 and recovering_stretch_mode 1, and the cluster never returns to healthy stretch mode by itself.

While both sites are up, the cluster reports HEALTH_OK and gives no sign of the problem.

Am I affected?​

You are affected if stretch mode is enabled and any pool uses a rule with a step take <bucket> class <class>. This lists those rules and the pools that use them:

ceph osd crush rule dump -f json | jq -r '.[]
| select(any(.steps[]; .op == "take" and (.item_name | contains("~"))))
| "\(.rule_id) \(.rule_name)"' | while read id name; do
echo "rule $name: $(ceph osd pool ls detail -f json |
jq -r --argjson id $id '[.[] | select(.crush_rule == $id) | .pool_name] | join(" ")')"
done

No output means you are not affected.

Cause​

A device-class rule selects OSDs from a shadow CRUSH tree, where site bs1 appears as bs1~ssd. Stretch-mode peering resolves each OSD to its site through the pool's rule, so it sees bs1~ssd. After a site loss, the Monitors require every PG to include the surviving site, and they identify that site by its real bucket, bs1. The shadow bucket never matches, so no PG of an affected pool can go active. Leaving recovery requires every PG to be active first, so the cluster waits indefinitely.

Workaround​

While a site is down, affected pools are unavailable and there is no safe workaround. Plan site maintenance and network changes with this in mind.

Once the lost site is back, the network is stable, and all OSDs are up, force the cluster back to healthy stretch mode:

ceph osd force_healthy_stretch_mode --yes-i-really-mean-it

The affected PGs peer within seconds and recovery proceeds normally. Do not run this while the site is still down or the link is still flapping.

Solution​

  • Fix: the patch in the tracker above makes stretch peering map shadow buckets back to the real site. With it, device-class rules behave the same as rules without a class, both during a site loss and after it. Clyso can provide a hotfix build ahead of the upstream release; contact support.
  • Without a patched build: move affected pools to a stretch rule without a device class, such as step take bs1 / step chooseleaf firstn 2 type host per site. Changing a pool's rule moves most of its data. It also mixes device classes, so it only suits clusters that don't depend on class separation.