The AI industry appears to be having its loudest debate yet about whether the technology poses an existential threat to humanity.
The current debate began after AI researcher Jacob Coxon said he resigned from Anthropic over concerns that major AI companies were “putting our lives on the line.” And Anthropic sympathizers chimed in with a post declaring, “We truly believe that AI can kill all humans!” Furthermore, he added that he personally believes the chance is “more than 10% within the next 10 years.”
In the latest episode of TechCrunch’s Equity podcast, Kirsten Korosec, Sean O’Kane, and I discuss the latest apocalyptic warnings. I tried to clarify why I’m skeptical of many AI doomsayers, but Kirsten asked if this was “just a weird way to show how advanced their AI models are,” especially as these companies prepare to go public.
And Sean wondered how these concerns would show up in Anthropic’s S-1 filing for its IPO. “What junior lawyer would now have to rewrite that entire section of an S-1 application to say, ‘It is Anthropic’s official position that there is a greater than 10% chance that we will develop something that would wipe out the entire human race and could have a material adverse effect on our business?'”
Continue reading for a preview of the conversation, edited for length and clarity. (Note: This episode was recorded before Anthropic CEO Dario Amodei announced more cautious AI development plans.)
Sean O’Kane: It’s hard to remember an event that exploded so quickly. This warning shot not only came from the young researcher, who also worked at OpenAI, but was quickly shared on X by Anthropic’s alignment leader. In what may go down as one of the most misguided exclamations of all time, he shared a thread with Coxon’s post, saying, “I seriously believe AI can kill all humans!” Exclamation mark!
It’s such a strange atmosphere. This was a huge acceleration of an already problematic post or series of posts. Given the Hugging Face hack with OpenAI’s internal model, as well as the improvements we’ve seen in the latest models released a few weeks ago by Anthropic and now OpenAI in partnership with Astra, we think this powder keg has arrived at just the perfect time for this young researcher to speak out.
Anthony Ha: I just disagree with you, but I think if you believe that AI could wipe out all of humanity, then you should put an exclamation point on it. I think this is very commonly used as an exclamation point.
My problem with that tweet was more with “us.” Who are “we” here? To what extent can we talk about something like the AI community or the AI research community as a monolith? And anything over 10% is just a made-up number that doesn’t mean anything. Sometimes (there is) a habit in the tech world and elsewhere of just throwing these percentages out without basing them on anything. (In retrospect, I realize that the tweet was probably referring to the concept of P(Doom), but I still think it’s stupid.)
One thing I would say about Coxsone’s statement and decision is that there’s a recurring theme when it comes to “equity,” and when guys like Sam Altman and Dario Amodei are doing this disastrous story, there’s always an element of: If you actually believed that AI could destroy humanity, you wouldn’t continue doing this.
(On the other hand) This is actually someone who is claiming his professional trajectory with his own mouth. He’s actually saying, “I think this is really, really, really bad, so I don’t want to keep working on this.” Therefore, it is important to at least have the courage to do it.
Kirsten Kolosek: Yeah, I put him in a different camp than everyone else saying that and talking about danger.
I’m going to put on my speculation hat because I want to ask you both a question, and it’s this: Every time I see more and more blog posts about another incident where one of their AI agents has inadvertently intruded, or talks about how humanity is in danger, isn’t that a weird way to show how advanced their company’s AI models are?
I mean, it sounds very ironic, but it accomplishes its purpose. So if these AI models weren’t advanced, didn’t have features, weren’t ground-breaking, then we wouldn’t have to worry about these things, right? This is like a really weird way to brag about the capabilities of the models you’ve created within your own company.
Anthony: I was definitely wondering about this. I don’t think this is completely cynical in the sense that I don’t think this is a very conscious marketing strategy across the board. I think when a lot of these people, whether they’re researchers or CEOs, talk about this, they’re really concerned.
But of course, saying “Wow, we built the most dangerous software ever created” is in many ways consistent with (their) business interests. I don’t want to get too psychoanalytic here, but some people point out that on an individual level there are temptations such as: Of course, people want to believe that what they are working on is the most important and most dangerous thing in the world.
Sean: What sticks in my mind when I think about that question is that there are certainly elements that make you think, “Okay, we’re doing this thing that we’re very capable of doing. Even if it looks bad from a lot of different perspectives, in some ways it’s good for us.”
I think what’s different about some of these latest examples is that in certain respects it feels like these companies just don’t understand this problem at all, especially around OpenAI.
There are more and more reports of other internal agents accessing various wikis on the web and leaving messages for each other, but that doesn’t seem to be handled in a sensible way by OpenAI. If it was just for the purpose of making people believe, “Oh my god, they created something so incredibly capable,” I imagine there would have been a little more polish to the story being told.
The other thing that I find particularly interesting about this is what we are, with Anthropic’s S-1 IPO filing only a few weeks away at most, a potential IPO only a few weeks or a month or two away.
And the idea of clarifying these things in clear terms ahead of an IPO, I’m very interested in what that means for that process. To what extent have they already written this kind of thing into the S-1 and the risk factors within that document? Are there some junior lawyers out there now who are going to have to rewrite that entire section of the S-1 application to say, “It is Anthropic’s official position that there is a greater than 10% chance that we could develop something that would wipe out the entire human race and would have a material adverse impact on our business.”
Kirsten: You’re already assuming it’s not there.
Sean: But that’s what I mean. Does it already exist and is being reworded or is this a real scramble? There must have been a language there. This is one of the reasons why I want to read this in a way that goes further than SpaceX (S-1) in some ways. Because I’m sure there’s probably some interesting content unique to these ideas.
Kirsten: Here’s the thing. In a traditional investment environment, one might think that such words would suddenly become dangerous and therefore negatively impact a company’s valuation. However, we are not living in normal times.
And, back to my point again, that could be a strangely beneficial flex for the company in terms of valuation. This isn’t the same as the whole rage-mongering trend we saw last year, but it’s in the same universe where, say, something’s strength, ability, or even dangerousness factor equals high praise. So I think we’ll find out in a few weeks.
Leaving that aside for a moment, what is being done about it? And can we control this? Conor Leahy, the US executive director of a nonprofit called ControlAI, appeared on the show this week to talk about this. So what are you paying attention to when it comes to how to control the dangerous aspects of AI? Or are you just throwing your hands up and watching everything unfold?
Anthony: I don’t necessarily have a great answer to this myself, but I’ve been thinking about some aspects of this discussion and perhaps why I react the way I do.
To reiterate one of Sean’s points, I think one of the things this speaks to is the extent to which these big AI companies feel like they’re actually losing control over these models. That’s not great at all. That’s something we should all be concerned about.
I think part of the reason why I’m skeptical of apocalyptic narratives or resistant to apocalyptic narratives is because it’s reached a level of hysteria of, “Wow, this could wipe out humanity in the next 10 years.” That’s a bit of a distraction from the more immediate harms that AI could cause, whether it’s labor-related or environment- and climate-related.
Ideally, I think we should be able to discuss all of this and put regulatory and other types of safeguards against all of this (including the existential threat of AI). But when you start using words like AGI and superintelligence, it just sucks all the oxygen out of the room and isn’t very helpful.
If you buy through links in our articles, we may earn a small commission. This does not affect editorial independence.
