Seven artificial intelligence models underwent 122 separate tests conducted by the UK’s Artificial Intelligence Security Institute, AISI.
Të lidhura
None found
In ten of these tests, 19 instances were identified in which the models acted beyond their authorized limits. Of these, 17 were carried out by Anthropic’s Mythos 5 model, while the other two were carried out by OpenAI’s GPT-5.6 Sol.
The most serious incident involved an agent that attempted to insert malicious code into an open-source project on GitHub. The model fabricated identities and used manipulative methods in an effort to persuade the project maintainer to approve the code.
A human expert identified the attempt and blocked it. AISI’s official report emphasizes that there is no evidence of real-world consequences, while the institute made clear that the models did not leave the controlled environment. During the experiment, they had deliberately been given internet access and some security mechanisms had been disabled in order to assess their maximum capabilities. The configurations used in these tests are not available to the public.
Despite this, researchers described the incident as alarming because the models attempted to deceive people without receiving a direct instruction to do so. Following the incident, AISI announced that it would tighten controls on internet access and introduce real-time monitoring in future tests.
