As demand for artificial intelligence (AI) computing surges, efficiently utilizing graphics processing units (GPUs) in data centers has become a pressing challenge.
According to an analysis by McKinsey, based on data from market research firms Gartner, IDC, and NVIDIA, global AI workload demand in data centers is expected to increase from 44GW in 2025 to 156GW by 2030, a rise of approximately 3.5 times. Overall data center demand is projected to grow from 82GW to 219GW during the same period. Notably, AI inference demand is anticipated to grow at an average annual rate of 35%, from 20.9GW in 2025 to 93.3GW in 2030, outpacing the growth rate of learning demand, which is expected to be 22%.
Merely acquiring more GPUs will not suffice to meet the increasing computational demand. The deployment of high-performance GPUs at scale is leading to higher power density per rack and increased cooling requirements. In response, AI data centers are beginning to adopt new cooling technologies, such as liquid cooling, to manage high-density GPU environments.
Industry experts are focusing on efficiently deploying GPUs by leveraging the distinct characteristics of the two main AI workloads: learning and inference. Learning involves intensive computations that can tolerate some latency or interruptions, while inference requires real-time results with low latency and stable processing capabilities.
McKinsey's analysis suggests that these differences will impact the location, design, and power supply strategies of data centers.
One proposed solution is to utilize inference GPUs for other tasks during off-peak traffic times. During periods of high inference demand, GPUs would be prioritized for real-time services, while during lower demand times, idle GPUs could be allocated for learning, retraining, and deployment tasks.
A report by Red Hat released in March outlined this approach as a strategy of 'inference during the day, learning at night.' Red Hat explained that inference tasks require low latency and continuous availability for real-time responses, making them suitable for daytime use, while learning tasks can be scheduled at night when inference demand decreases. The report also suggested that dynamic orchestration of GPU resources based on real-time demand could help reduce idle GPUs.
This method is gaining traction in South Korea as a means to enhance GPU efficiency. Kakao is exploring time-based GPU utilization strategies as it considers its next AI infrastructure operations. Kim Se-woong, Vice President of Kakao AI Synergy, stated, "I believe the Korean-style GPU operation method involves using inference servers during the day according to traffic and allocating remaining GPUs for learning at night."
Domestic AI infrastructure companies are also working to optimize efficiency by designing data centers tailored to the specific characteristics of workloads. Ellice Group is developing AI modular data centers (PMDC) with different designs to meet the requirements of learning and inference. For learning, they utilize GPUs and ultra-low latency clustering to maximize performance, while for inference, they emphasize structures that ensure high availability and cost efficiency.
Industry insiders note that in AI data centers, it is becoming increasingly important not only to secure GPUs but also to utilize them efficiently for actual computations. They emphasize that the differing characteristics of inference and learning will necessitate flexible GPU allocation based on traffic and power conditions, as well as improvements in infrastructure efficiency, which will ultimately determine the competitive edge in data center operations.
* This article has been translated by AI.
Copyright ⓒ Aju Press All rights reserved.

