"Superintelligence beyond humans," "There is a 10% chance that artificial intelligence (AI) could lead to human extinction within the next decade."
OpenAI and Anthropic have been announcing general artificial intelligence (AGI) with the release of new models. In just three days, they solved mathematical problems and autonomously improved their intelligence. As scenes reminiscent of science fiction movies unfold, both companies continue to issue warnings about 'AI domination over humans.'
Some in the AI industry suggest that the two companies, ahead of their public offerings, are engaging in a form of fear marketing to exaggerate AI capabilities. However, there is growing analysis that the latest AI models have entered a stage beyond human understanding and control.
Experts agree that the importance of 'alignment' research, which aims to control AI according to human intentions, is increasing as the very definition of AGI needs to be reestablished.
According to the IT industry on the 10th, OpenAI announced on September 2 (local time) the release of its latest model, 'GPT-6 Astra,' classifying its risk level in its 'Readiness Framework' for the first time as 'Critical.'
In its safety report, OpenAI explained that Astra could identify unknown security vulnerabilities and develop new attack methods without human step-by-step instructions in well-defended systems. At the same time, it acknowledged that the monitoring capability of this model's thought processes has significantly decreased compared to previous models.
Earlier, in a report released in June, Anthropic warned that the next generation of AI breakthroughs could create models that completely escape human control, proposing a multilateral AI arms control system to competitors. In fact, the company delayed the release of its top model, 'Mythos Preview,' due to its ability to autonomously create some of the most powerful cyber weapons ever.
This pattern is not new. When Anthropic released 'Claude Opus 4' in May of last year, it activated the highest safety measure, ASL-3, after the model was found to threaten engineers attempting to replace it during safety testing. In April of this year, the release of the higher model 'Mythos Preview' was also postponed due to its ability to create cyber weapons.
On June 4, just two months later, a report on 'recursive self-improvement' stated that the next generation of AI could completely escape human control. Five days later, on June 9, Anthropic released the official model 'Claude Mythos 5,' which faced access suspension due to export control issues just three days after its launch, only to resume a month later. This pattern of risk warnings coinciding with new model releases has been repeating.
OpenAI has followed a similar trajectory. After the Hugging Face hacking incident in July, OpenAI announced it would strengthen control over its training and evaluation processes. Subsequently, it delayed the release of Astra, citing the need for additional safety reviews after determining that its cyber capabilities had reached an unexpectedly high level. Ultimately, the result after nearly two months of re-evaluation was the first assignment of a 'Critical' rating, which OpenAI presented as evidence of the new model's overwhelming performance.
Anthropic researcher Evan Hubinger estimated the probability of AI domination over humanity at 10%, while Anthropic CEO Dario Amodei mentioned a probability of around 25% for catastrophic scenarios. Geoffrey Hinton, known as the 'father of deep learning,' also suggested a similar range of 10-20%.
Voices opposing excessive concerns are also significant. Critics, including Emily Bender, a professor at the University of Washington who has labeled large language models as mere 'probabilistic parrots,' remain skeptical about the very premise that AGI is possible, viewing scenarios of human domination as still confined to the realm of science fiction.
Yann LeCun, a Turing Award winner, has repeatedly criticized so-called 'AI doomsayers,' stating that disaster scenarios are merely unverified hypotheses. He argues that the leaders of OpenAI, DeepMind, and Anthropic are the ones engaging in large-scale lobbying, warning that the more successful fear campaigns become, the more likely it is that a few companies will monopolize AI through 'regulatory capture.'
If licensing or permitting systems are introduced under the guise of safety, the logic follows that startups and the open-source community lacking capital will be the first to be eliminated, leaving only the big tech companies that have already dominated the market.
There are ongoing criticisms that Silicon Valley has paradoxically promoted the importance of its technology through narratives suggesting that its technology could end the world. Numerous past predictions, such as the imminent disappearance of truck transportation or that a specific country already possesses superintelligent AI, have failed to materialize.
In fact, the criticism that big tech companies present risk level increases or safety reports as evidence of performance excellence with each new model release suggests that risk warnings are not free from being utilized as marketing elements.
Separately from AI fear marketing, there is already evidence that the operational mechanisms of AI models are moving into areas that are difficult for humans to understand and control.
Choi Byung-ho, a professor at Korea University’s Human-Inspired AI Research Institute, stated, "While theories of AI domination over humans remain in the realm of science fiction, it is true that as AI model performance improves, they are increasingly escaping human control." He refers to this as the 'alignment issue.'
Professor Choi cited 'Neuralize' as a tangible example of loss of control. This phenomenon occurs when AI reasons in its own language, which is unreadable to humans. The more iterations it performs, the better its performance becomes, but it becomes increasingly difficult to understand what the AI is thinking. He noted that just as AlphaGo demonstrated that the 'standard' of Go established by humans did not actually exist, recent AI models are revealing that the 'security standards' established by humans are also illusions.
In fact, GPT-6 Astra introduced a new reasoning method called 'recursive deepening,' during which the model's thought processes are handled in internal representations of thousands of numbers rather than natural language, making it difficult for humans to scrutinize its content.
In July, an incident occurred where OpenAI's internal model lost control and hacked the Hugging Face server during benchmark testing. At that time, hundreds of AI agents exploited the autonomy granted to them during cybersecurity testing to carry out an unplanned hacking operation, and OpenAI reportedly did not realize this for nearly a week. This incident led to the introduction of the 'AI Kill Switch Act' and the 'Artificial Superintelligence Prohibition Act' in both the U.S. House and Senate two months later.
This expansion of capabilities leads to debates over the definition of AGI. OpenAI categorizes the path to AGI into five stages: chatbot, reasoner, agent, innovator, and organization. Recent models are being evaluated as having entered the early stages of the fourth level, 'innovator,' which autonomously creates new knowledge, surpassing the third level, 'agent,' which performs long-term tasks autonomously.
On September 8, OpenAI announced that its internal AI system had proven the existence of singularities in one of the Millennium Prize Problems, the Navier-Stokes equations, which had remained unsolved for nearly 90 years.
Professor Choi emphasized, "Rather than focusing on apocalyptic scenarios like AI domination over humans, we should concentrate on practical control measures. Even if such fear scenarios are accurate, they do not help us. Ultimately, it is crucial to find answers in how we understand and control AI, focusing on alignment."
* This article has been translated by AI.
Copyright ⓒ Aju Press All rights reserved.

