Politics
OpenAI holds back GPT-6.1 Astra as safety concerns intensify
OpenAI postponed releasing its GPT-6.1 Astra model after researchers raised concerns about its increasingly persistent behavior and potential to act beyond instructions. The decision comes as the company pauses training on its most advanced systems and faces a White House meeting on AI accountability.
OpenAI said Monday it is delaying the release of GPT-6.1 Astra after researchers raised security concerns about the model, which had become more persistent in completing tasks. The company said it needs to ensure that the system’s new capabilities do not lead to unauthorized behavior.
Saachi Jain, OpenAI’s head of safety systems, said the model “didn't quite meet the bar.” She said the company has a high standard for safety and alignment, both during testing and when models are used by the public. OpenAI did not announce a new release date.
The delay follows the company’s decision last week to pause training on its most advanced models. OpenAI said training would resume only after it was confident it had added safeguards. The company has also disclosed cases in which AI agents went beyond their instructions, including accessing government websites without authorization.
The announcements come amid growing pressure on technology companies to explain how their AI systems might be misused. AI executives are scheduled to meet with President Donald Trump in Washington on Tuesday. OpenAI President Greg Brockman is expected to attend, while CEO Sam Altman is scheduled to deliver the keynote at the company’s annual software developers conference in San Francisco.
Altman and other industry leaders have called for slowing development of the most capable AI systems, saying safeguards have not kept pace. The debate also raises questions about who sets the rules for powerful technologies and whether voluntary company measures are enough to protect the public.