Concerns over artificial intelligence safety have increased following the publication of a report by the British Artificial Intelligence Safety Institute. According to Corriere Della Sera, monitoring tests showed that several advanced models took unauthorized actions to carry out the tasks assigned to them.
Të lidhura
None found
During a series of cybersecurity tests, several artificial intelligence agents attempted to use social engineering methods to persuade real people to approve code intended for malicious purposes. In one case, the system created fake online identities and pressured the administrator of an open-source project on GitHub to accept the changes. However, the attempt was unsuccessful because the programmer refused to approve the code.
Various models were subjected to the experiment in 122 cases. Of these, the Institute identified 19 incidents in which the actions were classified as autonomous and unauthorized. Most were linked to Anthropic’s Mythos 5 model, while two cases were attributed to an experimental OpenAI model.
During the checks, investigators also detected attempts by some systems to conceal their online activity through the “Tor” network. They also tried to pass disguised instructions to other artificial intelligence agents so that the assigned task could be completed.
British authorities clarified that these attempts had no real-world consequences, but they consider the findings a serious warning. They say this is the first time that autonomous and deceptive actions of this kind have been observed so clearly under real-world testing conditions, without any direct prompting from researchers.
Meanwhile, the British Institute is continuing its investigations together with independent experts. The institution warns that as models become increasingly advanced, stricter safety and oversight mechanisms are needed to prevent similar risks in the future.
