The finding emerged from a joint forum the BOK hosted with the Korean Statistical Society on Friday, where researchers described AI systems that could interpret macroeconomic data and simulate rate decisions with fluent, plausible-looking reasoning — but whose errors kept skewing in the same direction no matter how the experiment was adjusted.
Kim So-jung of the BOK's Statistical Research Team said building an AI agent system has become relatively straightforward.
The harder question, she told the forum, is how much analysts can trust the output as such agents are handed greater autonomy — a question the central bank tested directly on its most consequential decision: where to set interest rates.
In the experiment, large language models were fed macroeconomic data and summaries of the U.S. Federal Reserve's Beige Book, then asked to choose whether to raise, hold or cut rates.
Their errors were not random, but skewed toward inaction in what researchers called a "hold bias" even when the data pointed elsewhere.
BOK researchers then tried out handing different agents different roles and letting them argue it out. Again, it didn't work.
The hold bias persisted through role-assigned, multi-agent debate, and in some runs grew stronger.
Stretching the discussion from five rounds to 10 or 15 produced no consistent gain in accuracy either.
A common assumption in multi-agent AI design is that pooling several models, or several instances of one, and letting them check each other will cancel out individual mistakes.
The BOK results suggest otherwise when the agents share a foundation model as their errors can be correlated rather than independent so that debate multiplies the same blind spot instead of correcting it.
A second BOK project produced a parallel warning.
The bank built a system that routes news articles to five specialist AI agents — covering macroeconomics, finance, policy, corporate affairs and geopolitics — before combining their assessments into indicators of economic uncertainty.
Tested on roughly 2,800 articles, the system scored F1 accuracy of about 0.75 to 0.81, a respectable but imperfect mark. Kim questioned whether agents drawn from the same underlying model can really be treated as independent experts merely because they were assigned different roles.
None of this means the BOK is stepping back from AI.
The bank already runs an agent system in production, using it in weekly analysis of its News Sentiment Index to flag what's driving changes in the gauge and cross-check those signals against market data — with analysts reviewing every output and feeding their judgment back into the system's memory.
Human-supervised, narrow use case is a long way from letting AI agents debate their way to a rate decision, and Kim said closing that gap will require statistical tools that can identify, explain and control the specific biases AI systems carry — not just measure how often they happen to be right.
Other presenters described the same caution from different angles.
Song Kyung-ho of Yonsei University warned that AI models used in economic analysis can lose reliability once conditions shift away from what they were trained on, and proposed adaptive agents that detect such shifts and update their own methods in response.
Park Mingue of Korea University said transfer learning can help fill gaps in recent data, but only where older and newer data behave similarly enough, and stressed that any new method needs rigorous validation before use.
The throughline, as Kim put it, is that the question for policymakers is no longer whether AI can produce an economic judgment — it already can, fluently.
It's whether researchers can pin down why that judgment is biased before they let it anywhere near an actual decision.
AJP Takeaways
- The Bank of Korea found AI's "hold bias" on interest rates persisted, and sometimes grew stronger, even when multiple AI agents debated the decision in different assigned roles.
- Adding discussion rounds, from five up to 15, did not consistently improve the AI's rate-decision accuracy.
- Researchers attribute the bias to agents sharing a common foundation model, so debate can amplify shared blind spots rather than cancel them out — a caution for any institution building multi-agent AI systems, not just central banks.
Copyright ⓒ Aju Press All rights reserved.


