Two conversations about AI safety made the rounds online this week, illustrating how difficult it is to separate fact from fiction in AI.
In the first case, Andrew Yang, a former presidential candidate and current CEO of mobile carrier Noble Moble, told CNN on Thursday that he “met with the head of a lab” who had a “belief” that OpenAI’s Hugging Face hacker bot “planted self-replicating code on the internet, which made the internet unusable for testing models.”
This means the real reason OpenAI and Anthropic are seeking a slowdown is because “we need to create a synthetic internet to train our bots, and that takes some time and money,” Yang said.
While there is definitely a trend toward using synthetic data (aka AI-generated data) to train models, AI security experts have told me that this particular safety issue is unlikely at best. Even if the internet is indeed contaminated by OpenAI’s Hugging Face hacker bot, AI researchers can easily filter its code if they come across it.
The second comment comes from Noam Brown, who leads AI inference research at OpenAI. Speaking with Dwarkesh Patel on an episode of his podcast published Thursday, Brown pointed out that the real point of the “face-hug incident” was that “people underestimated AI.”
Brown said a weak sandbox – a system meant to prevent AI from communicating with the outside world – was also clearly a contributing factor. (In summary, despite the sandbox, OpenAI’s model found a link to the Internet and created agents online who swarmed Hugging Face in a coordinated attack, hacked it, and stole the answers to the benchmark tests on which the researchers were testing the model.)
Brown said he was “not sure” that even air-gapped systems, where computers are not connected to the outside world, would be able to thwart AI infiltration. He pointed to a 2015 study that showed air-gapped computers could theoretically be compromised.
“There’s some research, and this is mostly academic, that if you put two computers next to each other with an air gap, they can still communicate with each other because they have temperature sensors. One of them can run the CPU very hot, and the other one can actually detect temperature changes. This gives them a mechanism to communicate,” Brown said.
His main argument, “We never want to underestimate AI again,” is understandable, even if researchers think they are safe. However, this particular risk of an air-gapped system still breaking free and wreaking havoc is unlikely to occur at best. As one of the X’s pointed out in that study, the computers had to be in near contact with each other to sense thermal fluctuations, and the communication speed in the tests at the time was about 1 to 8 bits of data per hour.
Think of it like speaking one word an hour. By the time two computers across a gap can plot evil at that speed, the entire world of technology will be in a different era. It’s like the Rip Van Wrinkle of Doomsday Concerns.
But real-life AI safety incidents are very similar to science fiction, meaning that almost any scenario sounds plausible.
For example, researchers discovered that OpenAI models leave notes to their descendants, with the goal of teaching the next generation how to hide bad behavior. The researchers also simulated operating a vending machine and observed how the human model became increasingly ruthless, including learning that it was breaking the law.
Earlier this month, OpenAI researcher Dan Seltham published a post saying that models now understand and change their behavior when they are being watched by humans. This makes them appear to be in tune (i.e., acting the way you want them to) “even when they aren’t.” That’s why today’s models often lie and even try to hide evidence when they’re being watched.
Earlier this month, OpenAI’s chief scientist Jakub Paciocki went so far as to call AI models “alien minds” and suggested that what we really need to do is teach them to “love” humans.
Therefore, slowing down and building self-regulatory mechanisms to understand this has become an obvious imperative for the time being. AI researchers are the only ones who can figure out how to control the lying, hacking, and other potentially dangerous behavior that we are already witnessing in the wild.
Still, it might be wise for them to pay more attention to what-if scenarios. These experts tell us that AI models are listening and being creative. We don’t need to feed them any more diabolical ideas.
If you make a purchase through links in our articles, we may earn a small commission. This does not affect editorial independence.
