To survive in the competitive landscape against the U.S. and China, the K-AI strategy must shift from focusing on computational power to refinement and optimization, overcoming limitations in power and memory, said Yuminsoo, Chief Research Officer of A2Sys and a professor at KAIST.
Speaking at the 18th Good Growth, Good Jobs Forum (2026 GGGF) held on September 2 at The Plaza Hotel in Seoul, Yuminsoo emphasized the need for this transition.
Yuminsoo, a former NVIDIA expert in computer systems and GPU architecture, recently founded the AI infrastructure startup A2Sys. He presented quantitative research findings on the systemic burdens and power crises arising from the paradigm shift from traditional chatbots to autonomous AI agents.
He noted that while global tech giants like OpenAI and Meta are building massive AI data centers with capacities in the tens of gigawatts, the current infrastructure and power grid have clear sustainability limits.
According to Yuminsoo's research team, the response latency of AI agents has increased by an average of over 5.2 times compared to traditional large language model (LLM) chatbots. Unlike LLMs, which provide immediate outputs upon receiving input, AI agents engage with external tools like web browsers, continuously thinking and acting until their goals are achieved.
Yuminsoo explained that AI agents must retain previous conversation histories and search results in memory, quickly filling up storage space (KV cache). This leads to a breakdown of existing data center memory design assumptions, resulting in significantly slower processing speeds and exacerbated memory shortages.
As AI spends more time deliberating and recalculating to improve accuracy (test-time scaling), energy consumption has surged. Yuminsoo highlighted that processing a single agent query can increase GPU power and energy consumption by as much as 136.5 to 140 times compared to traditional chatbots.
The research team estimated that if AI agents were to scale to traffic levels comparable to Google Search (approximately 13.7 billion queries per day), the required power demand for data centers could reach up to 198.9 GW, nearly half of the total capacity of the U.S. power grid (approximately 476.9 GW).
Yuminsoo stated, "We have reached a point where securing GPU quantities is no longer sufficient; there is a shortage of commercial sites for power supply and data center installation. The random computation approach of pouring resources without strategy has clear limitations in terms of cost efficiency and pricing."
He proposed a 'refinement and optimization-focused K-AI strategy' as a solution to overcome the capital and infrastructure gaps in the U.S.-China power competition. He cautioned that simply increasing computing resources to boost performance leads to diminishing returns in accuracy relative to costs, resulting in a 'cost reduction' phenomenon. He urged that specialized systems and software optimization technologies for agentic AI workloads are South Korea's strong assets, and national efforts should focus on technological innovations that maximize infrastructure efficiency.
* This article has been translated by AI.
Copyright ⓒ Aju Press All rights reserved.

