Three thermal derates on the same compute ring over a short review interval have triggered a ring-level capacity assessment after operators concluded the events could no longer be treated as unrelated local adjustments.
None of the derates reached catastrophic levels, and service continuity was largely maintained through controlled load shedding. The deeper concern is structural. Shared assumptions about cooling margin, maintenance timing, and fallback distribution appear to have been more correlated than the operating model admitted.
This is exactly the kind of event insurers and customers watch closely. Single derates are manageable. Clusters suggest that graceful degradation may be less independent, and therefore less abundant, than the commercial story implied.
The review now underway is as much about institutional credibility as engineering. Operators need to show that the ring can still be trusted as a network rather than a collection of near-misses.