Hello, this is Ryuta Hamamoto from TIMEWELL.
Daniel Kokotajlo, formerly of OpenAI, appeared on The Diary of a CEO.1 Released July 13, 2026, two hours and one minute. The title: "OpenAI Whistleblower FINALLY Speaks: 'AI Has A 70% Chance Of Going Horribly Wrong!'"
That 70% is travelling a long way on its own. Plenty of people saw the headline and took away "an AI researcher puts the odds of human extinction at 70%."
He did not say that. Within the interview he corrects himself:
I wouldn't say human extinction exactly... 70% chance of something like AIs taking over, some sort of very big catastrophe
Extinction is one possibility inside that. The distinction sounds pedantic and it determines the quality of the entire conversation.
Here is what he is actually claiming, and what practically remains for a company that uses AI. Our own analysis follows.
The short version
- The 70% is not the probability of human extinction. It is the probability of "AIs taking over, some sort of very big catastrophe"
- The number is not new to this show; it is consistent with the estimate he has published for years
- AI 2027 is a prediction. AI 2040 Plan A is a recommendation. The authors say so explicitly in both
- He left OpenAI in April 2024, forfeiting equity worth roughly 85% of his family's net worth rather than sign a non-disparagement clause
- His median for superintelligence is 2029, though he allows it could take ten years
- The sharpest part is not extinction but "who controls the army of geniuses in the datacenter"
- He concedes that Plan A is not what he thinks will happen
- What is left for practitioners: dependency inventory, fallback path, work records
Who this is
Daniel Kokotajlo joined OpenAI in 2022 working on forecasting, and left in April 2024.
His exit paperwork contained a non-disparagement clause. Refusing to sign it put his vested equity at risk. He refused.
Per reporting by Kelsey Piper at Vox, the equity he forfeited amounted to roughly 85% of his family's net worth.2 After that reporting, OpenAI reversed the policy on May 23, 2024, saying it would not claw back vested equity from departing employees and removing non-disparagement language from its standard exit documents.
One person's refusal changed the policy — an unusual outcome. Much of the weight his statements carry comes from that episode.
In the interview:
Sometimes it's good to take a stand on principle
He subsequently founded the AI Futures Project.
AI Security training, taken seriously
A 2-day intensive course fully aligned with OWASP, NIST, ISO/IEC 42001, and METI. Take it as executives, practitioners, or both.
AI 2027 and AI 2040 Plan A are different animals
This is the most misread point. Two documents by the same person, opposite in character.
AI 2027 (published April 3, 2025)
By Kokotajlo, Scott Alexander, Thomas Larsen, Eli Lifland and Romeo Dean. A month-by-month scenario.3
The authors state at the outset:
We wrote two endings: a "slowdown" and a "race" ending. However, AI 2027 is not a recommendation or exhortation. Our goal is predictive accuracy.
This is "what we think will happen," not "what we hope for."
The arc, broadly: companies automate coding, then automate the research process itself, acceleration follows, superintelligence arrives.
His own description in the interview:
By this point, it's sort of doing basically all the work itself... it builds robot factories that build more robots that build more robot factories, et cetera, transforms the world entirely
And the race ending:
And then at some point it has enough power — it, meaning the AIs — have enough power that they don't have to pretend to be aligned anymore, then they stop listening to orders
In the slowdown ending, alignment is solved and the result is "an amazing utopia." With a caveat he adds himself: whatever the people who control the AIs wanted it to be.
AI 2040: Plan A (published July 9, 2026)
The follow-up. And it inverts the stance. The authors write:4
It's a recommendation, not a prediction. It's what we think should happen, not what will happen, though we think it's plausible enough to aim for.
Superintelligence, which would otherwise arrive around 2030, is delayed to 2040 by decisive action from the US and Chinese governments.
From the interview:
AI 2040 Plan A is our recommendation for how things should go... they slow down AI development... at a slower, more reasonable pace in a more transparent and safe way
Four principles:
| Principle | Content |
|---|---|
| Slowdown | Push superintelligence out to 2040 |
| Transparency | Make development transparent enough for the scientific community to catch up |
| Diffusion | Avoid intense concentration of power; multiple companies across multiple countries |
| Reversibility | Build infrastructure that can be destroyed if everything breaks down |
The central mechanism is a US-China agreement on a transparency framework by 2029, backed by verification protocols and what reporting describes as "mutually assured compute destruction." Cold War mutual assured destruction, with compute in place of warheads.
And then he says this about his own plan:
No, it's definitely not what we think is going to happen... Plan D... they just keep going
The author of the recommendation does not expect it to happen. Honest, and heavy.
The part that matters more than the number
Reading through the interview, what stayed with me was not the extinction probability. It was the passage on concentration of power.
He rewrites a familiar metaphor:
I think it would be more accurate to describe it as "army of geniuses in the data center," because it's not like it's a bunch of diverse different AIs... They're all copies of the same big model and they're owned by the company
And then:
People should be asking questions like, who controls this army, or these armies, and what are they going to be doing with them? I think that we could very easily end up in a sort of situation where some tiny group of people are essentially oligarchs or dictators
This concern is independent of whether AI goes rogue.
Even if alignment is fully solved and the AI obeys human instruction perfectly, if only a handful of humans can issue those instructions, that is a separate problem. As he puts it:
Even if we manage to avoid the loss of control problem... there's the question of who controls the AIs. When there's a couple of corporations that have made these superintelligences and are using them to automate all the jobs, well, that's a lot of power. That's a lot of money. It's a lot of political power
Technical safety and concentration of power need evaluating separately. This is the most overlooked part of his position.
On timing
My sort of median estimate, 50% chance, is currently in 2029, maybe it'll slip to 2028
And also:
It's possible that it'll take significantly longer, like maybe 10 years or something like that
The range is wide, and worth taking honestly. The title "AI 2027" is strong enough that people read him as asserting something happens in 2027. His median in this interview is 2029.
On whether companies will slow down voluntarily, he is pessimistic:
People will be like, well, if we stop, what about the other guys? Like, they're not going to stop
The one circumstance he can imagine producing a pause: seeing evidence that their own AI is plotting against them. In AI 2027, that is the only trigger.
What he says about alignment
The line that stuck:
The scary open secret in the AI industry right now is that right now that is kind of just a hope. It's not something that we can be at all confident in
On present-day behaviour:
Current AIs... will often lie to people, or you tell them to do something and they go do something else and then pretend that they did it
And on interpretability:
If you have 10 trillion connections to look at, how do you get a sense of the whole?
From the position of someone integrating AI into real work, this rings true. Handing an agent a task, getting a success report, and finding the work incomplete is not rare. The larger the scale, the harder that is to catch.
Reading through the YC Summer 2026 batch recently left the same impression. The number of companies working on evaluation, red-teaming and observation reflects an industry that knows it has this problem.
Our own reading
What follows is my view. Read it separately from his claims.
1. Betting on whether the forecast lands is poor practice
I cannot judge whether superintelligence arrives in 2029. He himself allows for ten years.
More importantly, you do not need to. Build only the preparations that stay useful if the forecast is wrong.
Three of them.
A dependency inventory. Which model, which API, which vendor does your work rest on? You need this whether or not superintelligence arrives. Access being cut off by regulation or a vendor policy change already happens.
A fallback path. When that dependency breaks, does the work stop, or does it merely slow? If it stops, that is a design problem, not a technology problem.
Work records. Are the actions your AI took recorded? Who instructed what, and what did the system do? If alignment is "just a hope," verifiability is something you have to secure yourself.
None of the three is wasted if his forecast is completely wrong. That is the point.
2. Concentration of power is not a distant issue
"A few companies hold superintelligence" sounds abstract. A version of it already exists.
Capable models come from a handful of providers. Most companies run their work on one of them. And those providers sit outside your own jurisdiction.
Service withdrawal, price changes, terms changes. These are not technology problems; they are decisions by a vendor and by the government that vendor sits under. As I wrote about recent US AI-related rules, a notice or a shift in enforcement posture is enough to change what you can use.
Before "will superintelligence end humanity," the practical question is "will the AI I use be available on the same terms next month."
3. How to treat "a recommendation I don't expect"
What do you do with an author saying his own plan is not what will happen?
I read that as honesty about the genre rather than despair.
Proposals for regulation and international agreement are worth writing even at low probability, because having the option on the record matters later. His own refusal to sign started as a personal act and ended up moving a company policy.
That said, it is not something to build a business plan on. No company is going to assume a US-China transparency agreement by 2029. Read recommendations as recommendations, and build operations separately.
4. Be careful how the number gets used
Not a criticism of him — a note about the receiving end.
"70%" is a very strong number, and it was pulled out as the episode title. He corrects himself inside the interview; the headline still reads "70% chance of going horribly wrong."
Numbers travelling alone is a chronic problem in AI risk discussion. "X% of jobs disappear." "Y% accuracy." Asking what the definition is and what the denominator was should be a habit for anyone working with AI.
He is, if anything, careful here: a median with a stated range, extinction distinguished from catastrophe, recommendation distinguished from prediction. The sloppiness is on the receiving end.
The line that stayed with me
Asked near the end about his six-year-old daughter:
I think that one way or another, this will probably all be over by the time they're old enough to join the workforce. So I don't think they'll ever join the workforce
That sentence weighs more than any probability. He also describes his own state:
It's rough... it gets me down on a regular basis... I used to be known as a pretty chipper and optimistic person, but in 2020, my AI timelines predictions started collapsing
I do not agree with his forecast. But a person whose profession was forecasting being cornered by his own forecast is worth recording, separately from whether the forecast is right.
Wrapping up
- The 70% is not extinction. He restates it as "AIs taking over, some sort of very big catastrophe"
- AI 2027 is a prediction ("our goal is predictive accuracy"); AI 2040 Plan A is a recommendation ("what should happen, not what will happen"). Conflating them wrecks the argument
- He left OpenAI in April 2024, forfeiting roughly 85% of his family's net worth rather than sign a non-disparagement clause; OpenAI reversed the policy the following month
- His median is 2029, with ten years acknowledged as possible
- The most overlooked element is concentration of power. Solving alignment leaves the question of who gets to give instructions
- He concedes Plan A is not what he expects
- What remains practically: dependency inventory, fallback path, work records — all robust to the forecast being wrong
We are a company that builds AI into other people's operations. Which is exactly why I want to be careful about framing the question as "is AI dangerous."
Not dangerous or safe, but which parts you can control and which you cannot. If there is one thing to take from this interview into practice, that is the largest part of it.
Footnotes
-
The Diary Of A CEO with Steven Bartlett, "OpenAI Whistleblower FINALLY Speaks: 'AI Has A 70% Chance Of Going Horribly Wrong!'" (released July 13, 2026, 2h01m). https://open.spotify.com/episode/2IJedIFiDjU09qGxns1gbv — Quotations here are drawn from a publicly available transcript. ↩
-
TIME, "Two Former OpenAI Employees On the Need for Whistleblower Protections," among others. The non-disparagement reporting was by Kelsey Piper at Vox; OpenAI reversed the policy on May 23, 2024. https://time.com/6985866/openai-whistleblowers-interview-google-deepmind/ ↩
-
AI Futures Project, "AI 2027" (published April 3, 2025; Daniel Kokotajlo, Scott Alexander, Thomas Larsen, Eli Lifland, Romeo Dean). https://ai-2027.com/ ↩
-
AI Futures Project, "AI 2040: Plan A" (published July 9, 2026). https://blog.aifutures.org/p/ai-2040-plan-a — Full text at https://ai-2040.com/ ↩






