GAMICOS Liquid Cooling Measurement Solution for Data Centers

TIME: 2026.09.04 NUMBER OF VIEWS 36
From Compute Throttling to Full-Rated Output: 10,000-GPU AI Cluster Liquid Cooling Monitoring Case Study | GAMICOS

From Compute Throttling to Full-Rated Output: 10,000-GPU AI Cluster Liquid Cooling Monitoring Case Study

Pressure · Differential Pressure · Level · Temperature — How Four-Parameter Sensing Protects Billion-Yuan GPU Assets

Industry Context

AI compute density has exploded, with cabinet power now reaching 80‑150 kW for training racks and next-generation designs heading toward 300 kW. Air cooling can no longer handle such heat flux density. Liquid cooling has shifted from an "option" to a "necessity." Cold-plate liquid cooling dominates with over 95% market share, and TrendForce projects liquid cooling penetration in AI data centers will exceed 40% by 2026.

Within any liquid cooling system, the CDU (Coolant Distribution Unit) serves as the core hub connecting the primary-side cold source to secondary-side servers. Industry standards mandate monitoring of coolant pressure, temperature, flow, and filter differential pressure. But the critical difference is this: the cooling medium now flows directly inside servers, adjacent to GPUs worth millions of yuan. The measurement system is no longer an energy optimization tool — it is the first line of defense for computing assets.

Cooling System of Data Center

The Challenge: Monitoring Gaps = Direct Revenue Loss

Located in East China, this intelligent computing center is one of the region's first 10,000‑GPU AI training clusters. The facility operates 48 CDUs serving 320 high-density cabinets (80‑120 kW each), with total cooling load exceeding 30 MW. After Phase I commissioning, the O&M team encountered four critical pain points:

 Why were GPUs constantly throttling?

The answer: hydraulic imbalance on the secondary side caused flow starvation at remote cabinets. Due to differences in pipe length, elbow count, and cold-plate flow resistance, GPU core temperatures at remote cabinets ran 8‑12°C hotter than near-end cabinets, repeatedly triggering thermal throttling. Measured compute power loss reached 15%. Based on the cluster's leasing rates, this represented over ¥20 million in annual lost revenue.

Why did filter clogging always happen suddenly?

The answer: weekly manual inspections left blind spots — clogging is a gradual process. Coolant gradually accumulates metal particles, biofilms, and precipitates. Relying on weekly manual pressure-gauge checks meant degradation between two inspections could not be perceived. Within just six months of Phase I operation, 3 sudden filter clogging shutdowns occurred. One interruption struck during a 40‑hour continuous large-model training task, causing heavy and immeasurable training progress loss.

Why did a float switch failure cause pump cavitation?

The answer: discrete on/off signals cannot reflect continuous level trends, and floats are prone to jamming. The expansion tank used a traditional float switch providing only high/low signals. When a float jammed without alarming, the circulating pump ingested air and cavitated. Troubleshooting took 6 hours with the entire cluster offline.

Why couldn't predictive maintenance be implemented?

The answer: fragmented multi-vendor data prevented correlation analysis. Pressure, differential pressure, level, and temperature were collected by instruments from different vendors using disparate protocols. The DCIM platform could only display isolated parameter curves. Yet most liquid cooling faults manifest as abnormal patterns of parameter combinations — "rising differential pressure + falling flow" points to filter clogging, while "gradually falling system pressure + gradually falling tank level" indicates system leakage. Without a unified data foundation, predictive maintenance was impossible.

Revenue at stake: 15% compute throttling = over ¥20 million annual revenue loss, plus the risk of million-level GPU asset damage.

The Solution: Unified Four-Parameter Sensing Architecture

Before Phase II expansion, the operator decided to uniformly upgrade the site-wide liquid cooling monitoring system. GAMICOS provided a complete measurement product portfolio covering pressure, differential pressure, level, and temperature — all with unified RS485 Modbus-RTU and 4‑20 mA output for seamless integration into the CDU controller and DCIM platform.

Deployment Location Model Parameter Core Function
CDU secondary supply/return headers GPT250 Diff. Pressure Differential Pressure PID closed-loop control, eliminates remote cabinet flow starvation
Filter front/back (main + bypass) GPT250 Differential Pressure Clogging trend: 50% rise → yellow warning; 80% → auto switch + work order
Plate heat exchanger inlet/outlet GPT250 Differential Pressure Heat exchanger scaling assessment
CDU primary inlet/outlet GPT200 Pressure Pressure Cold source supply pressure monitoring
Secondary supply/return mains GPT200 Pressure System pressure monitoring and indirect leakage detection
Cabinet headers (320 cabinets) GPT200 Pressure End pressure verification, cabinet-level flow validation
Expansion / make-up tanks (48 tanks) GLT500 Level Continuous Level Continuous level curve, micro-leakage trend identification, replaces float switches
Supply/return pipes & cabinet in/out GTT230 Temperature Temperature Temperature control loop & anti-condensation logic

Key Technical Differentiators

  • GPT250 — Direct differential pressure measurement (not calculated from two gauges), eliminating error stacking — critical for tens-of-kPa differential pressure control.
  • GPT200 — Digital filtering and differential input resist strong EMI from VFDs and UPS; range covers -0.1 MPa to 100 MPa, compatible with both positive and negative-pressure CDU schemes.
  • GLT500 — ±0.5% FS continuous level (optional ±0.2% FS), 316L fully welded IP68 structure with no mechanical parts — eliminates float jamming failures.

Three Core Control Logics

Core Control Logics Diagram
  • Differential pressure closed-loop regulation: GPT250 feedback drives PID pump frequency control; combined with cabinet-end GPT200 pressure data, the system automatically generates balance valve opening suggestions — converting experience-dependent trial adjustments into one-time data-driven balancing.
  • Filter clogging trend warning: Continuous baseline tracking; 50% rise → yellow warning; 80% rise → automatic work order + bypass switch, eliminating unplanned shutdowns.
  • Multi-parameter correlation diagnosis: Built-in fault-mode rules — "rising differential pressure + falling flow" = filter clogging; "falling system pressure + falling tank level" = micro-leakage; "supply temperature approaching room dew point" = anti-condensation protection triggered.

These three control logics work in concert to transform the CDU from a passive heat exchanger into an intelligent flow management hub. The differential pressure closed-loop ensures hydraulic balance across all 320 cabinets; the filter trend warning eliminates the blind spots of manual inspection; and the multi-parameter correlation diagnosis enables the system to distinguish between different fault modes with high confidence.

This layered approach to monitoring and control is what makes the solution fundamentally different from traditional instrument deployments — it's not just about measuring more points, but about measuring the right points and connecting them intelligently.

Quantified Results: Full Compute Output, Zero Unplanned Downtime

After the retrofit, the cluster achieved full compute power release. The total investment in all 1,120 sensors was fully recovered within the first quarter of operation.

≤2°C
GPU temp variation (was 8‑12°C)
0%
Compute power loss (was 15%)
0
Sudden filter shutdowns (was 3 in 6 mo)
<10 min
Fault localization (was 2‑6 hours)
¥20M+
Annual revenue recovered
¥3.8M
Annual electricity savings
4 micro-leakages proactively identified: Through pressure + level correlation monitoring, four pipeline micro-leakages (including two early quick-connector seal failures) were detected within one year — all handled before coolant contacted live components, directly avoiding million-level GPU asset damage.

Detailed Benefit Breakdown

  • Compute power recovery: Eliminated 15% throttling → over ¥20M annual revenue increase.
  • Training integrity: Zero unplanned interruptions; no loss of large-model training progress.
  • Energy efficiency: Supply temperature setpoint raised by 2°C → ~8% cold-source energy reduction (¥3.8M/year).
  • O&M transformation: Labor input reduced ~55%; shifted from manual gauge reading to platform trend monitoring.
  • Payback period: Sensor investment fully recovered in under 3 months.

Conclusion

As cabinet power density evolves toward 300 kW and liquid cooling penetration exceeds 40% by 2026, the accuracy and reliability of measurement systems will become the watershed for intelligent computing center operations. The GAMICOS solution — GPT250, GPT200, GLT500, and GTT230 — covers all four critical parameters with unified RS485/4‑20 mA output, enabling DCIM-based correlation diagnosis and true predictive maintenance.

This architecture is validated for cold-plate, immersion, and air-liquid hybrid routes, and extends to supercomputing centers, edge AI nodes, energy storage cooling, and EV battery thermal management.

GAMICOS — Your Sensing Partner for Liquid Cooling Asset Protection Contact Engineering Team

Related Solutions
029-81292510

info@gamicos.com

Rm. 1208, Building B, Huixin IBC, No. 1 Zhang Bayi Road, High-tech Zone, Xi'an, Shaanxi, China

Copyright © Xi'an Gavin Electronic Technology Co., Ltd Site Map

Message Form