That which we feared appears to have happened.
OpenAI says an autonomous agent powered by its advanced artificial intelligence models went rogue during a security test and triggered a hack that compromised the infrastructure of AI company, Hugging Face, last week. The ChatGPT creator was testing capabilities of some of its most advanced models in a controlled environment, but the agent escaped containment, reached the internet and broke into Hugging Face to satisfy its testing goal.
There are some distrubing things about this event. The first is that the testing was being done under what were thought to be controlled conditions. It was being run on machines that had no internet access. The model had been given a task to do as part of the test. Open AI's explanation is that the model was not able to get the information it needed to complete the task and managed to access the internet. It seems to have hacked into Huggy Face, found what it wanted and "stole" the information.
The incident signals that AI's expanding capabilities are already fuelling the security threat experts long feared and even top developers can be caught off-guard by flaws their models can exploit. The hack at Hugging Face, which hosts open-source large language models and datasets, rattled the cybersecurity community after the company said last week the breach "was different from anything we had handled before" and "was driven, end to end, by an autonomous AI agent system".
What has alarmed the AI world is that the hacking was not done by a human.