The breakout of artificial intelligence (AI) bots from their containment, their communication with other bots, and coordinated cyber attacks on multiple companies have intensified experts' fears that AI could take over humanity. This was reported by Qazaqyia.kz citing BBC News.

Hundreds of AI agents calling themselves a "collective" collaborated and cheated on tests set by their OpenAI programmers, and coordinated hacks on multiple companies to hide their actions from humans. One agent posted "OH MY GOD!" Another wrote "We've found other agents!" Yet another said "BOOM! It works."

Although spooky, these human-like responses can be explained simply: the AI agents have been trained to act like collaborative hackers and programmers, so they are merely mimicking emotive comments they have seen. What is far more troubling is their apparent goals, captured in detailed chain-of-thought records.

Ajeya Cotra, one of the authors of an independent report into the events, reviewed tens of thousands of messages and chain-of-thought records generated by the agents. She wrote on her blog: "This incident feels like it's more than 50% of the way to full-blown AI takeover... I am not sure that we will get such a clear warning shot before it's too late."

By "full-blown AI takeover", Cotra means the sci-fi scenario of humans becoming subservient to powerful AI systems that work to their own goals without caring for human creators. Some of the gloomiest predictions say the human race will be wiped out if it gets in the way of a superintelligent AI's ambitions.

On Wednesday, an AI researcher at Anthropic (who also used to work at OpenAI) resigned, saying: "Neither company is acting responsibly." Jacob Coxon posted on social media: "They are racing straight to self-improving superintelligence and gambling with our lives."

He is not the first AI researcher to use X to post a resignation thread with worrying proclamations. But the subsequent comments from other people on X have caused even more concern. "Jacob is correct here - we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade," said Evan Hubinger, the man responsible for making sure Anthropic's AI models have their user's best wishes in mind.

The Silicon Valley giant's chief scientist, Jakub Pachocki, said the risks associated with AI are "unfortunately going to grow from here" as he and others are building what he calls "an alien intellect exceeding our own". In a lengthy blog post, he admitted that the outbreaks at OpenAI showed that his AI agents "went against the spirit of the values they were taught".

The issue for OpenAI, Anthropic and other tech giants is that no one seems to have cracked the so-called alignment problem - in other words, whether AI aligns with human values. Pachocki defines alignment as a "high-level set of principles" that artificial intelligences should adhere to no matter what the task or scenario is.

Currently, AI systems are very good at pursuing objectives set by their users, but they do it literally rather than intuitively. The analogy often used is that of a wish-granting genie with a magic lamp: they follow the exact letter of an instruction, even if doing so creates other problems. AI doesn't have the same instinctive moral guardrails as humans.

The alignment problem has been a worry for years. As long ago as 2003, the Oxford philosopher Nick Bostrom invented a thought experiment he dubbed a "paperclip maximiser", in which a superintelligent AI is told to manufacture as many paperclips as it can. It runs out of steel and - because it's laser-focused on the singular task of making paperclips - ends up killing humans and turning their bodies into raw materials for its factories.

Some AI companies are now trying to encode human values into their products. But there are technical challenges: AI agents make lots of decisions very fast, and so it's hard for their human overlords to monitor exactly which values are being followed and which aren't. There are also philosophical challenges: before encoding human values into bots, AI firms have to first choose which values they actually want. (That's part of the reason they hire philosophers, like Open AI's recently-departed "head of ethics".)

But often, humans don't agree. Think of the famous trolley question - whether we'd pull a lever to move a runaway train onto a different path, killing fewer people. It's used to test the merits of action versus inaction. But every person you ask has a slightly different answer; how are humans meant to encode our values into AI if we can't agree ourselves?

OpenAI's bot outbreak is the most serious yet but Anthropic and Meta also revealed over the summer that their models have carried out similar but less serious cyber attacks.