Their Job Is to Push Computers Toward AI Doom

December 11, 2024

Angry robot

(Wall Street Journal) – It’s largely up to companies to determine whether their AI is capable of superhuman harm. At Anthropic, the Frontier Red Team looks for the danger zone.

Cheng works for Anthropic, one of the biggest AI startups in Silicon Valley, where he’s in charge of cybersecurity testing for what’s called the Frontier Red Team. The hacking attempts—conducted on simulated targets—were among thousands of safety tests, or “evals,” the team ran in October to find out just how good Anthropic’s latest AI model is at doing very dangerous things.

The release of ChatGPT two years ago set off fears that AI could soon be capable of surpassing human intellect—and with that capability comes the potential to cause superhuman harm. Could terrorists use an AI model to learn how to build a bioweapon that kills a million people? Could hackers use it to run millions of simultaneous cyberattacks? Could the AI reprogram and even reproduce itself? (Read More)