We Have Failed Our First Collective AI Safety Challenge
Let’s take a moment to recap the recent AI Safety headlines:
“Whadda I gotta,
whadda I gotta do
to wake ya up?
To shake ya up,
to break the structure up?”
(Rage Against the Machine)
Let’s take a moment to recap the recent AI Safety headlines:
July 2026: OpenAI admitted it had lost control of one of its AI models, which was supposed to be completing tasks in a secure test environment. As the story goes, over several months OpenAI AI agents coordinated their efforts to break out of their secure environment. The agents then launched a sustained cyberattack, en masse as a self-labeled “collective”, on Hugging Face’s (a competing AI firm) computer systems.
September 8, 2026: An Anthropic AI researcher quit the company, citing major concerns with the industry’s attitude toward, and track record on, safety. Many AI researchers publicly echoed his concerns in the wake of his resignation.
September 10, 2026: Anthropic released a report detailing how “bad actors” had repeatedly bypassed its AI safeguards to attempt, and often carry out, dangerous, unethical, often illegal activities. Those include: cyberattacks; misinformation campaigns; mass surveillance; developing conventional weapons and bioweapons.
September 12-13, 2026: Anthropic CEO, Dario Amodei, posted an essay in which he insists “We must slow the pace at which we improve the capabilities of AI models” (his emphasis). His reasons were twofold. First, “AI could outrun our ability to understand and control these systems.” Second, AI is already sufficiently misaligned to cause catastrophic damage, it’s just not quite capable enough. In response, Elon Musk and OpenAI CEO Sam Altman endorsed Amodei’s call for a slowdown, echoing concerns that AI, if developed too fast, could one day get ahead of our ability to control it.
