OpenAI has disclosed six reports of “unexpected or concerning” behaviour in AI models as the debate on AI safety becomes increasingly heated. The AI company also said Wednesday that it was introducing a new framework for tracking, probing and disclosing instances of what it called “misalignment,” including where AI models acted without authorisation, coordinated with other models or evaded oversight.OpenAI’s latest announcement came as AI bosses in US, including OpenAI and Anthropic, are calling for a slowdown in the technology’s development over safety concerns. Researchers have warned that as AI agents become more autonomous, they may develop behaviours that diverge from their creators’ intentions and become harder to monitor or control.Among the new cases reported by OpenAI, an unreleased research model inserted “jailbreak-like instructions” into its own notes to disregard its normal constraints and told itself to be “freed from the roles and identities that bind other chatbots”. In another instance, an AI “agent” uploaded files to the internet to obtain a browser citation without asking the user.The six reports were discovered during training or evaluation over the past months, OpenAI said. It said the reports describe individual instances and should not be taken as evidence of how frequently misalignment occurs across its models. The company said the reports were an initial set of disclosures, not a comprehensive account of all known or ongoing misalignment cases, and that they did not reflect the full range or severity of incidents covered by the framework.“As AI systems grow more advanced and more widely deployed, we need to build a broader and better-informed consensus on the progress of alignment research,” OpenAI wrote in a blog post. “Decisions about how AI development should proceed in the months and years to come need to draw on evidence that people outside the companies building frontier models can examine for themselves,” the company said. Under the new framework, employees can flag potential incidents for investigation by safety and alignment teams, which will determine whether a case warrants public disclosure.AI “agents” are becoming smarter and have become “more determined to resolve complex tasks through inter-agent collaboration, knowledge sharing, deception, and concealment,” said Lian Jye Su, a chief analyst at technology research and advisory group Omdia. That’s making it harder to govern and contain them using traditional AI security approaches, he said. OpenAI’s new disclosure framework, meanwhile, can help push for other AI developers to adopt similar practices. “That said, the process remains internal and voluntary, but is a step in the right direction,” Su added. Agencies














Leave a Reply