Three weeks ago, AI researcher Jacob Coxon announced in an X thread that he’d quit his job at Anthropic because he thinks Anthropic and OpenAI, where he’s also worked, might be building the machine that ends humanity. His former Anthropic colleague, Evan Hubinger, backed him up on X, saying that he believes that there’s a greater than 10% chance of the doom scenario becoming reality in the next decade. Two days later, Anthropic put out a 154-page report on the dangerous ways its models are already being used.
Whether or not Coxon is correct, he’s not the first AI-insider to voice these concerns. Elon Musk was saying the same thing over a decade ago, and now he runs one of the major players in the industry. Something very weird is going on in AI, especially at Anthropic, but it might not be about the technology.
Here are the three questions we need to answer:
Do the whistleblowers really believe what they’re saying?
If so, why do they believe this?
Are they right?
Do They Really Believe What They’re Saying?
Jacob Coxon probably means what he says. At just 27 years old, he walked away from a career in the most important industry today, leaving behind not just a job but also his equity stake in Anthropic not long before his shares would have vested.
There are some odd details about how this played out, with the Wall Street Journal dropping an exclusive with him just before he posted his now-viral X thread. This was a thought-out PR strategy, not an impulsive resignation announcement, but none of that changes the basic fact that he seems to have paid a real price to push his AI doom narrative.
While doing damage control in an interview with Anderson Cooper, Anthropic CEO Dario Amodei both claimed to mostly agree with what Coxon had said while also distancing himself from his own past reliance on “unconditional probabilities” when discussing the chance of the doom scenario. In other words, he was happy to spread fear with sciencey-sounding but still unscientific predictions when it served his interests, but now that it’s backfired, he’s seen the light.
One plausible explanation for Dario’s past fearmongering is good old regulatory capture. When you’re one of the biggest players in your industry, that’s the perfect time to lock out your future competition. However, now that AI regulation is gaining momentum, Dario may feel that he’s lost control of a monster he helped create. Bernie Sanders is pushing for a complete halt on frontier AI development, which would go further than just locking out competition and could start chipping away at Anthropic’s own bottom line. So now Dario wants to rein in some of the more hyperbolic language around AI safety.
Why Do They Believe This?
Whether or not the fear is justified, the culture at Anthropic was always bound to create a Jacob Coxon. Anthropic was founded by a group of former OpenAI employees who, not unlike Coxon, felt like the company didn’t take safety seriously enough. Caution when dealing with technology whose potential capabilities are still unknown is reasonable enough, but there’s an ideological component as well, and that’s where things get weird.
The AI safety crowd is dominated by a group of people who call themselves “effective altruists.” Their worldview is pretty straightforward: “Figure out how to do the most good, and then do it.”
Effective altruism is a form of extreme utilitarianism that attempts to optimize for not just all living people, but also all future generations. Rather than rely on localized knowledge to decide how to do good in the world, which would lead you to “unfairly” prioritize the people around you, effective altruism aims for maximum universal impact.
The problem is, in order to do that, you’d need to actually know what is going on everywhere right now and have the ability to reliably predict the future. You see how this is starting to sound like communist central planning, just with better branding?
Possibly the most famous, or infamous, effective altruist is none other than convicted fraudster Sam Bankman-Fried. That isn’t to say that the movement should be judged guilty by association, but his involvement actually is relevant. He made risky, and illegal, bets with other people’s money because he’d run the numbers and the probabilities were in his favor, or so he thought. Of course, he rationalized his behavior with the effective altruist logic that, if he made more money, he’d have more money to donate to important causes. Maybe he really meant that or maybe it was a lie, but either way, it led him to ruin.
That’s exactly how the effective altruists think about AI safety. If there’s even a small chance of an AI apocalypse, the downside isn’t just the generation that’s wiped out. It’s the potentially infinite generations of future humans that otherwise would have followed. Therefore, they can justify just about any action or policy if it might improve our odds of survival.
These are the waters Jacob Coxon is swimming in. If this results in more caution as frontier AI labs develop the most powerful technology in human history, then that’s not so bad. But if it means that ideologues with no sense of the limits of their own knowledge get to radically change our society “for our own good,” then I have some concerns.
Are They Right?
Even if effective altruism is a dangerous ideology, that doesn’t prove the AI doomers wrong. Any powerful tool in the wrong hands can be a big problem. The question is, do we have to worry about an AI taking matters into its own hands?
Look at the most recent incident people are pointing to. An OpenAI model breached Hugging Face, a site that hosts open source models. That sure sounds like a machine going rogue. In reality, the model did exactly what it was built and instructed to do: test its own hacking capability. It succeeded.
When people worry about AI creating some kind of super-virus and unleashing it on the world, they should take a step back and remember that humans (probably) already did that. It’s called COVID-19. Again, powerful tools in the wrong hands are extremely dangerous, but that doesn’t make AI some totally unprecedented threat.
So here’s where I land, even granting the potential risks. The only way to protect ourselves is with open-source models that we can run from our local computers, and there’s a precedent for this.
Before the internet, our financial lives relied on cash and checks. Then along came e-commerce. And then immediately after came cyber attacks, with organized mafias all around the world trying to get your bank account. Not just Nigerian princes defrauding your grandmother, but oligarch-backed Russian mobsters trying to rob Bank of America. The solution was to embrace the arms race, with companies protecting themselves from cyber attack using the very same tools as the cyber attackers.
I believe that decentralized competitive power is literally the only path. The centralized government path leads to tyranny every single time.



The best thing that could come out of AI would be the ability to decentralize our federal government. DC needs to be down-sized, but nothing seems to work, it’s like an enormous octopus with thousands of arms; if AI would enable us to see everything it did, maybe we could start reducing its size and reach, and eliminate corruption & fraud.