ODF 4.15 upgrades can break Ceph’s integrated Prometheus exporter
A Python exception can push the storage cluster into warning health during an in-stream 4.15 upgrade, making before-and-after monitoring checks part of the change plan.
Red Hat has verified an OpenShift Data Foundation 4.15 upgrade failure in the Ceph manager’s integrated Prometheus module. During an upgrade within the 4.15 stream, the storage cluster can enter a warning state while the manager reports Module 'prometheus' has failed: AttributeError("'NoneType' object has no attribute 'fileno'"), according to the Customer Portal solution updated Sept. 19.
How monitoring fails
The exception is inside Ceph’s Prometheus exporter subsystem. That means operators can see two related symptoms: Ceph health reports a module failure, and the metrics path used by Prometheus can stop representing the storage system correctly. A green application dashboard is therefore not enough evidence that the storage upgrade is healthy; the Ceph manager module itself must also be checked.
Red Hat identifies ODF 4.15.x with Red Hat Ceph Storage 6.x as the affected environment. The public advisory does not narrow the issue to a particular 4.15 z-stream pair, so any in-stream 4.15 upgrade using that combination deserves the same preflight until Red Hat publishes a narrower boundary.
Before the upgrade
Capture the Ceph health output and manager-module state, and confirm that the Prometheus module is enabled and currently serving metrics. Save a short baseline of storage alerts and a handful of metrics that exercise capacity, health and daemon status. Also confirm who will evaluate Ceph health during the maintenance window; a generic cluster-upgrade checklist can miss a module-specific warning.
Because Red Hat’s corrective procedure is subscriber-only, do not improvise by repeatedly disabling or removing the module during an active storage upgrade. Have the full solution or a support case available before the change if the environment matches the stated versions.
After the move
Check the upgrade status, then inspect Ceph health and manager logs for the exact NoneType/fileno exception. Verify that the exporter endpoint is producing fresh samples rather than merely remaining reachable, and confirm that expected storage alerts can still be evaluated.
A practical acceptance sequence is:
- Compare Ceph health with the pre-upgrade baseline.
- Confirm the active manager and Prometheus module state.
- Search manager logs for the documented exception.
- Check timestamps on representative Ceph metrics.
- Validate storage dashboards and alert rules with current data.
- If the module has failed, pause further rollout and use Red Hat’s supported remediation.
The warning is operationally important because it can turn monitoring blind at the same moment operators most need reliable storage telemetry.
sources
comments · 0