In high-density AI data centres, a single cooling failure can trigger thermal runaway within seconds. As GPU clusters outpace the limits of air cooling, the question is how fast organisations can adapt and rethink liquid cooling maintenance.
The AI data centre liquid cooling market is projected to grow from $3,2 billion in 2025 to $15,3 billion by 2035, representing a compound annual growth rate of 16,9%, according to Precedence Research. With 75% of new data centre projects targeting AI workloads, and 53% of operators expecting liquid cooling to dominate future high-density deployments, the infrastructure landscape is changing rapidly.
This trend is increasingly relevant in South Africa where growing investment in hyperscale facilities, cloud services and AI capabilities is placing greater emphasis on data centre resilience, energy efficiency, and infrastructure reliability.
The new reality of AI data centre cooling
What is often overlooked is the operational complexity that accompanies liquid cooling. Coolant distribution units, heat exchangers, precision pumps and manifold assemblies, whether in direct-to-chip, immersion or hybrid configurations, introduce entirely new failure modes compared to conventional HVAC systems. These assets cannot be effectively maintained using calendar-based maintenance schedules designed for a simpler era.
Schneider Electric’s whitepaper on direct liquid cooling system challenges in data centres highlights material compatibility challenges, corrosion pathways and the importance of engineered system integration.

Why preventive maintenance falls short
Traditional preventive maintenance was designed for environments with lower power densities, simpler architectures and limited operational data. Today’s AI facilities generate continuous streams of information from electrical power monitoring systems and building management systems. Relying solely on fixed inspection intervals while ignoring this data means missing valuable indicators of asset health. The consequences are significant:
• Developing faults that emerge between scheduled inspections often remain undetected until they become critical.
• In AI environments, even a minor cooling disruption can result in downtime, performance degradation, and loss of computational capacity.
• For South African operators where maintaining uptime is complicated by inconsistent power supply, reducing avoidable failures becomes even more important.
Schneider Electric’s recent white paper on AI-driven systemic condition-based maintenance, ‘Rethinking Data Centre Service with an AI-Driven Systemic Asset Management Strategy’, demonstrates why reactive and calendar-based approaches are increasingly ineffective in high-density environments. As liquid cooling becomes central to AI infrastructure, predictive and condition-based maintenance are emerging as critical enablers of thermal stability and operational resilience.
The predictive maintenance advantage
Condition-based and predictive maintenance strategies use real-time monitoring, advanced analytics, and AI-driven algorithms to transform maintenance from a reactive cost centre into a strategic advantage. Research shows that the majority of organisations implementing predictive maintenance report positive returns, while a considerable number achieve full payback within the first year.
Effective predictive maintenance depends on more than sensors alone. Success comes from combining monitoring technologies with advanced analytics and actionable insights. Key monitoring capabilities include:
• Vibration monitoring on pumps and cooling distribution units to detect mechanical degradation before failure occurs.
• Thermal imaging on heat exchangers and coolant distribution manifolds to identify temperature anomalies caused by fouling, flow restrictions, or thermal interface degradation.
• Continuous analytics that compare real-time operational data against known failure patterns, helping identify trends and anomalies before they affect performance.
Schneider Electric’s next generation of EcoCare for cooling services is designed to provide remote monitoring and data-driven insights that forecast maintenance requirements with sufficient lead time to schedule interventions without disrupting operations. Adopting predictive maintenance for liquid cooling systems is an operational strategy; the infrastructure is already generating the necessary data. The question is whether organisations are equipped to use it.
As South Africa’s digital economy continues to expand and AI adoption accelerates, predictive maintenance will play an important role in ensuring that next-generation data centres remain resilient, efficient and capable of supporting the workloads of the future.
© Technews Publishing (Pty) Ltd | All Rights Reserved