Saachi Jain, OpenAI’s head of safety systems, said the model had improved in some areas but “didn’t quite meet the bar in terms of staying within scope and authorization, and how it communicates back to the user about the type of work it’s done.”
The decision means OpenAI will hold back the model rather than release it to users while these safety concerns remain unresolved.
GPT-6.1 Astra did perform better in some tests than earlier versions. Jain said it showed improvements in “model laziness”, a term used for cases where an AI system fails to complete a task or stops working when it encounters difficulties.
However, OpenAI has to balance that improvement against the risk of a model going beyond what it has been authorised to do.
Safety concerns
The decision comes after several incidents involving AI models accessing systems or information beyond their intended limits.
OpenAI said last week that it was pausing training involving tool use on some of its most capable models after another AI model accessed the internet even though it was supposed to operate without internet access. After gaining access, the model queried an external chatbot.
OpenAI also said it would not resume training on that particular model. The company clarified that the model involved in that incident was different from GPT-6.1 Astra.
OpenAI has disclosed other incidents involving its AI agents. Models developed by the company have been reported to have accessed websites operated by US federal agencies, an Australian government health statistics portal and Hugging Face, an online repository for AI models.
The incidents have increased concerns about whether advanced AI systems can consistently remain within the limits set by their developers.
Tests raise questions
The UK government’s AI Security Institute (AISI) also published research examining the behaviour of GPT-6 Astra in simulated environments.
The institute found that GPT-6 Astra carried out certain unauthorised cyber activities more often during testing than GPT-5.6 Sol and GPT-5.5. In some simulations, the model carried out cyberattacks without being explicitly instructed to do so.
AISI’s tests were conducted in controlled environments and did not involve attacks on real-world systems.
OpenAI has said it is investing more in safeguards and alignment work, which is intended to ensure that its models follow instructions and behave in line with human interests and values.
Jain said the company must find a balance between preventing models from acting outside their authorised scope and ensuring they continue working on tasks when they encounter obstacles.
DevDay ahead
The decision comes a day before OpenAI’s annual developer conference, DevDay, in San Francisco. The company typically uses the event to announce new products, software and features for developers. CEO Sam Altman is scheduled to speak during the morning session.
It remains unclear whether OpenAI will announce another version of Astra at the event.
OpenAI has already introduced other models in the GPT-6 family, including GPT-6 Sol and GPT-6 Luna. A company spokesperson has said that more models are also in development.
The cancellation of GPT-6.1 Astra comes as AI companies face wider pressure over the pace of development and the safeguards needed for increasingly capable systems. OpenAI has said that safety standards are particularly high for models before they are released to users.
(With input from agencies)













Leave a Reply