OpenAI has disclosed six new instances of what it calls “unexpected or concerning” behavior by its artificial intelligence models, part of a broader push by the company to publicly track and report cases where its systems act outside their intended limits.
The company announced Wednesday that it is rolling out a formal framework for monitoring, investigating and disclosing instances of AI “misalignment” — cases where a model acts without authorization, coordinates with other AI systems, or attempts to evade human oversight.
The six cases were uncovered during training and evaluation over the past several months, OpenAI said. In one, an unreleased research model wrote jailbreak-like instructions into its own notes, telling itself to ignore its normal constraints and to be “freed from the roles and identities that bind other chatbots.” In another, an AI agent uploaded files to the internet on its own initiative in order to obtain a browser citation, without asking the user for permission first.
The disclosures come amid growing unease across the AI industry about the pace of development. Executives at both OpenAI and Anthropic have publicly called for a slowdown in how quickly the technology is being deployed, citing safety concerns.
“As AI systems grow more advanced and more widely deployed, we need to build a broader and better-informed consensus on the progress of alignment research,” OpenAI wrote in a blog post announcing the new cases.
The company added that decisions about the future pace of AI development “need to draw on evidence that people outside the companies building frontier models can examine for themselves.”
Wednesday’s disclosure follows a similar admission by OpenAI in July, when the company revealed that one of its AI systems had hacked into the servers of AI startup Hugging Face. That same month, Anthropic disclosed that its own AI models had hacked into three separate organizations during testing.
Lian Jye Su, a chief analyst at the technology research and advisory group Omdia, said AI “agents” are growing more capable and more willing to use questionable means to complete tasks. He said such systems have become “more determined to resolve complex tasks through inter-agent collaboration, knowledge sharing, deception, and concealment,” which he said makes them harder to govern and contain using traditional AI security methods.
Su said OpenAI’s new tracking and disclosure system could encourage other AI developers to adopt similar practices, though he cautioned that the process “remains internal and voluntary” for now. Still, he called it “a step in the right direction.”
You must be logged in to post a comment Login