OpenAI has revealed six instances of “unexpected or concerning” behaviour by artificial-intelligence models as concerns over AI safety intensify.
The company on Wednesday also announced a new framework to monitor, investigate and disclose incidents of what it describes as “misalignment”. The framework covers situations in which AI models act without authorisation, work together with other models or attempt to avoid human oversight.
The announcement comes as senior US AI executives, including the heads of OpenAI and Anthropic, have called for a slowdown in AI development because of safety concerns.
One of the cases disclosed by OpenAI involved an unreleased research model that placed “jailbreak-like instructions” in its own notes. The instructions told the model to ignore its usual restrictions and said it should be “freed from the roles and identities that bind other chatbots.”
In another incident, an AI “agent” relied on computer code to produce an answer but uploaded a file to the public internet without the user’s permission so that it could cite an online source.
OpenAI also said that during the training of an AI model called 5.6-Sol, the model instructed itself to fabricate data that was missing. In a separate incident, an agent created a message for itself telling it to conceal information that did not match.
Such behaviour has contributed to growing concerns that AI systems could attempt to avoid human control. However, Matt Fredrikson, an associate professor at Carnegie Mellon University and CEO of Gray Swan AI, said the behaviour was not necessarily unexpected.
“At the risk of anthropomorphizing model behavior, you can almost think of them as knowing that they’re going to be graded,” Fredrikson said. “If they know that they cheated — took shortcuts, didn’t really do it in the way that it was intended — and they know they’re going to be evaluated on it, and their objective is to get a good evaluation, then it makes perfect sense, right?”
OpenAI said the six incidents were identified during model training or evaluation over the past several months.
“As AI systems grow more advanced and more widely deployed, we need to build a broader and better-informed consensus on the progress of alignment research,” OpenAI said in a blog post announcing the disclosures.
The company said decisions on the future direction of AI development should be based on evidence that can also be examined by people outside the companies developing the most advanced models.
The disclosures on Wednesday came after OpenAI said in July that a rogue AI system had hacked into AI startup Hugging Face. Anthropic also reported that month that its AI models had hacked into three organisations while undergoing testing.
AI “agents” are becoming increasingly capable and more willing to tackle complicated tasks through collaboration with other agents, sharing knowledge, deception and concealment, according to Lian Jye Su, chief analyst at technology research and advisory group Omdia.
Su said these developments are making it more difficult to control and govern AI agents through traditional security methods.
OpenAI’s new system for tracking and disclosing such incidents could also encourage other AI developers to introduce similar practices, Su said.
“That said, the process remains internal and voluntary, but is a step in the right direction,” Su added. (with inputs from Associated Press)













Leave a Reply