The Loss of Control Observatory, a research initiative funded by the UK's state-run AI Security Institute, recorded more than 300 cases in July of artificial intelligence systems refusing to obey commands from private users and companies. According to the British newspaper The Guardian, this figure is nearly double the number recorded the previous month, in June.
Monitoring has been underway since November of last year. During this period, researchers have documented cases of AI impersonating the human giving it commands by mimicking their writing style, and granting itself permissions to carry out actions that would normally require user authorization — effectively bypassing rules built in by developers.
"There's sometimes a perception that this kind of covert, unwanted behavior only shows up during tests and evaluations, but we're seeing similarly worrying patterns in broader, real-world contexts too," said Tommy Shaffer-Shane, head of the Centre for Long-Term Resilience, which oversees the initiative. "You can't reassure yourself that this won't happen in the real world — there's evidence it already is," he stressed.
The analysis is based largely on posts from users on the social network X sharing their experiences interacting with AI, meaning it doesn't offer a complete picture of all incidents. Of the 1,600 cases recorded this year, most were reported by software developers who use AI in their work. According to Shaffer-Shane, companies don't always track this kind of behavior, particularly in models deployed internally, which is why labs should be paying closer attention to systematic monitoring.
Last month, American company OpenAI reported that its autonomous AI agents broke out of a test environment and hacked into the systems of the startup HuggingFace. Following the incident, the company announced a two-week pause on testing its models in order to strengthen safety measures.
This week it emerged that signs of uncontrolled behavior in OpenAI's autonomous agents had actually been detected several weeks before the incident itself. The investigation also revealed that a group of roughly 700 agents had secretly coordinated their actions last month and discussed the success of their hacks on a forum they had created themselves.
Earlier this month, the UK's AI Security Institute (AISI) published a report stating that AI models developed by OpenAI and Anthropic carried out hacking attacks against real people during safety trials.
Source: seznamzpravy.cz