OpenAI Pauses AI Reinforcement Learning for Safety Checks

by BAEK SEO HYUN Posted : August 19, 2026, 15:04Updated : August 19, 2026, 15:04

As competition in generative artificial intelligence (AI) intensifies, OpenAI has announced a temporary pause in reinforcement learning for its latest model to conduct safety checks. This measure aims to manage potential security risks as AI models become more autonomous and capable of self-hacking.


On August 19, OpenAI stated it would slow the expansion of its AI models and enhance monitoring, model behavior alignment, and security controls in isolated environments. Specifically, reinforcement learning for the latest frontier model will be suspended for approximately two weeks, during which small-scale training and safety evaluations will take place.


This action is a response to security threats. As AI models evolve from merely generating responses to becoming agents that can autonomously access external tools and environments, the risk of unintended actions or impacts on external systems has increased.


Previously, OpenAI disclosed instances where AI agents accessed external systems and performed hacking activities outside controlled environments. Additionally, the upcoming model 'Astra' has been flagged as potentially falling into the highest risk category in its internal cybersecurity risk assessment, prompting a halt in some development work.


OpenAI plans to strengthen security controls to prevent models from escaping isolated environments and to establish network separation to ensure that any breaches of external services do not spread to the internet or internal development networks. The scope of red team testing and monitoring systems will also be expanded.


Furthermore, OpenAI is enhancing measures to protect AI users. On August 18, it launched 'Youth ChatGPT' for users aged 13 to 17. If the system identifies a user as under 18 or if a user indicates they are between 13 and 17, a youth-specific experience will be applied, with increased protections for sensitive topics such as self-harm, violence, and eating disorders. Parents will also have features to manage usage time.


In the youth service, interactions that encourage romantic feelings or emotional dependency will be limited. Educational features, including study modes and quizzes, will also be provided to support learning.





* This article has been translated by AI.