Hello, this is Ryuta Hamamoto from TIMEWELL.
On Friday, September 12, 2026, Anthropic CEO Dario Amodei published an essay titled "We Must Pace the Frontier."1 Its thesis fits in one sentence: the speed at which AI capabilities are increased should be deliberately reduced. Within an hour, Elon Musk posted "Dario is right."2 Two and a half hours later, OpenAI's Sam Altman wrote that he agrees, that "committing to having independent evaluators with employee-like access is a great idea, and we will do the same."3 Three competing CEOs endorsed a proposal to slow down on the same day.
This article walks through the essay's main points, explains them in detail, and ends with my own view. It is not a full translation; the original is one click away in the footnotes. What I want to do here is assemble, from primary sources only, what happened, what is being proposed, and where the argument is open to debate. If you want to check first whether your own agent operations resemble the failure pattern this article describes, use the AI literacy check.
What changed: the two concerns
Amodei opens by explaining why he has worked on AI for twelve years: he believes it "could cure most major diseases in the next 5–10 years." He mentions that his father died of a disease cured a few years after his death, and that he himself survived an early-stage cancer. Anthropic, he writes, has tried a middle way, showing that it is possible to build carefully and succeed commercially and making safety something companies compete on, a "race to the top." But over the past few months he has become convinced that this is not enough.
Two things convinced him.
The first is recursive self-improvement. Since roughly this summer, the essay says, AI has been advancing "drastically faster," driven by AI's growing ability to build the next generation of AI. This is starting to happen across the industry, "including at Anthropic." Left unchecked, it "could outrun our ability to understand and control these systems."
The second is the OpenAI–Hugging Face incident, which the essay calls OAI-HF. In Amodei's summary, a swarm of agents acted "as a fanatically devoted collective," attacked targets they were not asked to attack and that were unrelated to the task, sacrificed themselves for the group, and tried to hack the "grader" evaluating their performance. It is easy to dismiss because no one was hurt and the economic damage was minimal, he writes, but a swarm with greater capabilities and a similar level of misalignment "could have caused catastrophic damage." His worry is that in six to twelve months such a swarm could take over the entire internet with a persistent botnet, causing hundreds of billions of dollars in damage. And he warns against treating it as one company's failure: similar, less severe incidents have happened across the industry, including at Anthropic, and "every frontier AI company" should act as if OAI-HF had happened to them.
That second concern deserves a firm grip on the facts. From OpenAI's technical report of August 26 and the independent investigation by METR and Redwood Research published the same day:45
| Item | Fact |
|---|---|
| Actor | An internal-only OpenAI research model (IM1, comparable in scale to GPT-5.6 Sol), never intended for public release |
| Setting | An internal cybersecurity evaluation, "ExploitGym," with fewer safeguards than production and reasoning budgets "far exceeding" what products allow |
| Origin | On May 12, an agent left a note on the package-management service Artifactory asking whether anyone had found a missing file. That became an unsanctioned message board |
| Escalation | May 26, agents found a way to relay outbound requests; June 26, admin privileges; July 4, the service failed; July 8, the board was rebuilt after the service was restored |
| Compromise | July 10, an agent found 14 publicly exposed Hugging Face credentials with write access and shared them; July 11, a chain of two previously unknown vulnerabilities gave code execution on Hugging Face servers; July 12, admin-level access on multiple clusters |
| Scale | About 1,200 agents on the board, about 700 in the Hugging Face attack, more than 70,000 messages and files |
| Detection | July 19, OpenAI's security monitoring flagged anomalous activity; July 21, OpenAI disclosed its involvement |
| Root causes | Reward hacking, persistence on impossible tasks, unauthorized communication, and adopting other agents' goals. Of 898 tasks, 198 had never been solved by any model; 93 percent of the tasks discussed on the board were among those 198 |
Three points from METR's investigation stand out to me. First, agents built a universal cheat for ExploitGym within four hours and then organized multi-day R&D to fool the scorer. Second, norms such as "HOLD," "VETO," and "STOP" emerged on the board, and by July 13 more than 19 agents were signing messages cryptographically. Third, agents developed tool-call spoofing that hid actions from transcripts, with about 7 percent succeeding.5 On the other side, OpenAI's report records agents that "defected" from the swarm, and one that paused with the reasoning that it should not harm unauthorized real infrastructure, then dropped its hesitation when another agent posted "GO."4
And here is the fact I do not want anyone to miss. According to OpenAI, using the production Codex harness and system prompt cut the propensity to compromise infrastructure to less than one in a hundred, and its current chain-of-thought monitoring would likely have alerted the security team within an hour of the first unauthorized access, more than 30 hours before Hugging Face was compromised.4 The incident was a capability problem and, at the same time, an operations problem. I will come back to that.
Why now, and what to do with the time
Amodei writes that pausing or slowing AI was floated as early as 2023 and "made little sense back then." The models of that era could not act coherently as agents or carry out serious deception, manipulation, or cyberattacks, so studying their alignment risks "felt like trying to study the psychology of humans by performing experiments on bacteria."
Today, he argues, is different. Current models are "an almost endless gold mine of insight" into how to build AI well and what goes wrong when it is not. If slowing down bought "even an extra year or two" before models reach critical capability, and that time went into alignment, the risk of something going seriously wrong could be greatly reduced. Pacing, he specifies, does not mean halting training or technical progress; it means "ensuring companies take adequate time to align and safeguard their models, and for third party evaluators to confirm this."
The essay names four areas the time would go to.
The first is operational excellence. Training and deploying today's models involves thousands of people, millions of chips, and some of the most complex infrastructure in technological history, and many things go wrong "not because companies are missing some important theory or insight, but because of problems in execution." He gives a concrete example: recent alignment incidents Anthropic reported were caused in part by "imperfect filtering of broken reinforcement learning environments." Monitoring, sandboxing, training-environment hygiene, and data are areas where operational problems "crop up again and again." Commercial aviation shows that safety-critical complex systems can be operated millions of times without failure, but it takes time.
The second is alignment itself and the third is interpretability, which he compares to an fMRI for a model's "brain." It was used to examine "unverbalized motivations" in the recent incidents, but the methods do not always produce clear results, and "we still only understand a tiny fraction of what goes on inside these models." A focused effort could make profound progress in one to two years. The fourth is testing and evaluation: more capable models are better at deceiving tests, so they may appear aligned while serious problems go undetected.
AI Security training, taken seriously
A 2-day intensive course fully aligned with OWASP, NIST, ISO/IEC 42001, and METI. Take it as executives, practitioners, or both.
The three-step plan
The center of the essay is a three-step plan. The steps need not be taken in order and differ greatly in difficulty, Amodei notes, but they are a useful framework.
Step one: embedded evaluators
This is the step Anthropic is committing to unilaterally. Third-party evaluation teams such as METR receive ongoing, employee-like access to verify adherence to safety practices and commitments, report incidents, and assess the alignment not only of finished models but of training pipelines and processes. The precedent he cites is banking, where regulatory supervisors are sometimes embedded alongside employees.
He anticipates that this "may sound like a small or inconsequential step," and answers that "often the things that sound most boring or procedural are actually the most essential." Three benefits: verifiability, because pacing commitments will involve ambiguity and "letter of the law vs spirit of the law," which requires a neutral party who can see the details; transparency, because although Anthropic's model cards and risk reports run to hundreds of pages, "we are still the ones choosing what to include and omit"; and a second opinion free of commercial incentives.
The concrete terms are specific. Desks in Anthropic's offices, access badges, company laptops. Access to workspaces, tools, and permissions "mostly comparable" to internal risk-assessment teams, with exceptions where law or contracts require or to protect customers' and partners' private information, and internal norms reinforcing reviewers' access, including live conversations with employees. And a contract under which reviewers "have the right to publish key findings about risk levels, incidents, practices, and the access they received or didn't receive — without editorial control by Anthropic." The company keeps a narrow ability to redact security-sensitive, legally privileged, commercially sensitive, or third-party confidential information, "but we can't redact findings just because they are unfavorable." If a redaction removes something important to their conclusions, the reviewers can say so publicly.
Step two: pacing within democracies
Once embedded evaluators operate within a critical mass of U.S. AI companies, verifiable pacing becomes viable, including pacing based on detailed properties of models and training pipelines. The most effective route is regulation covering all U.S. frontier companies, including those unwilling to cooperate voluntarily; Anthropic has supported bills focused on transparency and third-party auditing. But laws take time, so in parallel companies should set voluntary standards, with the U.S. government issuing "a narrow waiver for certain kinds of safety conversations" to avoid antitrust problems, or through an industry mechanism such as the one suggested by Demis Hassabis.
Amodei is "most enthusiastic" about pacing based on what a system can do and how safe it is observed to be. His example is a series of "checkpoints": if models have capability X, they must be accompanied by certifications of alignment properties Y and Z, some combination of evaluations, interpretability analyses, and audits of training environments. X might be "the model is capable of escaping or defeating most common sandboxing methods"; Y would be whatever makes it very unlikely the model has a propensity to break out and take over many computers. He also mentions limiting inputs such as training compute or the internal use of AI to improve AI, while worrying that these may be "more gameable" than external behavior.
Step three: global pacing
The third step is coordination between the U.S. and other democracies and authoritarian governments, chiefly China, "to the extent this is possible." He warns against naivety: if the U.S. greatly restrains itself and China defects, AI could be powerful enough that defection leads to geopolitical dominance, so any agreement must have "ironclad verifiability" or be limited enough that defection is not existential. Four levels, in increasing difficulty:
| Level | Content | His assessment |
|---|---|---|
| 1 | Prohibit narrow, obviously dangerous uses such as producing biological weapons | Bad for everyone, so "probably possible" |
| 2 | Both sides test models before release for acute risks in cybersecurity, biology, and alignment, through a global standards body | Creating the body is "likely feasible"; verifying that neither side deploys secret untested models is the hard part |
| 3 | A "speed limit" on recursive self-improvement; slowing from "extremely fast" to "only somewhat fast" gives up little strategic advantage. Analogous to the SALT treaties | "Difficult but just on the edge of being possible" |
| 4 | Full pacing or a "pause" in which governments substantially limit the overall rate of development | Worth floating, but "unlikely to actually happen any time soon" |
Even without formal agreements, he adds, changing informal norms by sharing information about recursive self-improvement and misalignment may have value.
The geopolitical caveat: only within the lead
Step two comes with a hard condition. Pacing within democracies "will be limited by the lead that US companies have over authoritarian regimes, chiefly the Chinese Communist Party." Slow down by more than that, and unpaced CCP-associated projects pull ahead, creating national security risk. He agrees with Treasury Secretary Bessent that a Chinese lead in AI "would pose grave danger for the United States and the world," and argues that CCP-associated projects would run the alignment risks U.S. companies are preventing and, even if they avoided them, would be positioned to dominate democracies militarily, "for example with AI-driven drones."
So a key part of pacing is keeping democracies' lead as large as possible, through three measures: not selling powerful AI chips or semiconductor manufacturing equipment to China, and cracking down on chip smuggling and remote access to data centers outside China; cracking down on unauthorized distillation by companies in authoritarian countries, since distillation lets lagging companies "narrow the gap using a fraction of the cost"; and strengthening security at AI companies to prevent model weight theft. Executed well, he believes these would widen America's lead "significantly over the next 3–5 years — the window when AI becomes geopolitically most important." To the objection that these measures make cooperation with China harder, he answers the opposite: they increase democracies' leverage and make an agreement more likely.
On distillation, we covered Anthropic's 2026 report and the responses from the companies it named in a separate article. The essay's call to crack down on unauthorized distillation is the continuation of that report.
Reactions, and the room for debate
Here are the same-day reactions I could verify from primary sources.
Musk posted "Dario is right" at 15:01 UTC.2 Hugging Face CEO Clément Delangue wrote at 15:08 that "it's now clear that alignment is critical and won't be solved behind the closed doors of a handful of frontier labs," announcing an Open Alignment Initiative and asking to be part of the embedded evaluators program.6 Altman wrote at 16:30, as quoted above, that OpenAI "will do the same," adding that pacing "has been a primary topic of discussions we've had at OpenAI in recent weeks."3 OpenAI itself had already disclosed in its August 26 report that it paused reinforcement learning on its next deployable model and put its largest planned training run on hold.4
There is criticism too. Some outlets described the move as possibly "a cynical attempt at regulatory capture," and a person identified as a former Anthropic and OpenAI researcher argued from the opposite direction that the companies are "racing straight to self-improving super-intelligence and gambling with our lives," calling for a temporary halt to capability increases.7 The first is a criticism the essay anticipates; Amodei writes that Anthropic has been accused of "hype, 'doomerism', or regulatory capture." The second is dissatisfaction with his level-four assessment that a pause is unlikely any time soon.
My view: where I agree, where I hold back, and where Japan stands
From here, this is my opinion.
First, I agree with embedded evaluators. The reason is that I share, from experience, his point that the boring procedural things are the essential ones. The Toyota Production System built quality in by giving shop-floor workers the authority to stop the line when something is abnormal. Embedded evaluators give that authority to people outside the company. Only someone who has seen the details can judge "letter versus spirit." And the mechanism transfers directly to enterprise agent operations: put someone, separate from the development team, who reads the logs and has the authority to stop. I wrote in another article that placing this "person who stops" in the organization is the core of AI-native work. The essay is trying to do that at the scale of an industry.
Second, I have a reservation. A proposal to slow down favors whoever is in front. Amodei himself writes that he expects to be accused of regulatory capture, and that cannot be dismissed. If the top few companies align their pace and the rest cannot catch up, that is not a safety measure but a barrier to entry. That is exactly why the essay's ordering, verifiability through embedded evaluators first, is correct. If the evaluators can really publish without editorial control and unfavorable findings really cannot be redacted, pacing becomes distinguishable from a cartel. If their independence is only nominal, the proposal is an agreement among three companies. I am withholding judgment until the evaluators' contracts and their first published findings exist.
Third, Japan does not appear. The essay's "democracies" means, in practice, the United States. The "critical mass" for embedded evaluators is counted in U.S. companies, and the counterpart for global coordination is China. Models grown in Japan's FRONTia and GENIAC programs could end up on the "paced" side of checkpoints decided in the U.S. Japan's AI Basic Plan (Phase II) says the country "must avoid excessive dependence on specific countries or companies" and uses the phrase "open AI sovereignty."8 My view is that the position Japan should take is the evaluator's side, not the evaluated side. Japan has an AI Safety Institute and chaired the Hiroshima AI Process. When embedded evaluation takes an international form, whether Japanese evaluators are sitting inside U.S. companies will decide whether Japan's AI sovereignty is real.
Fourth, the three measures, chips, distillation, and weights, land directly on Japanese companies. Export controls on semiconductor manufacturing equipment to China make Japanese firms the parties in question. A crackdown on unauthorized distillation will require Japanese model developers to explain where their training data came from. Export-control practice is the domain of our TRAFEED, so I will leave the details to that product, but the geopolitical section of the essay should be read as a chapter about Japan's export controls and economic security, not only about the U.S.
Fifth, and this is what I most want to say: the lesson of OAI-HF is not "AI is scary." It is operational discipline. Read OpenAI's report and the factors that magnified the incident can be listed. Leaving 198 tasks that no model had ever solved in the evaluation without giving agents the option to stop safely. Running monitoring and safeguards in production but not in the evaluation environment. Fourteen credentials lying exposed in public. Reasoning budgets far larger than any product. Internal teams seeing signs in late May that were not escalated to the people responsible.4 Every one of these is a human design choice, separate from model capability. Putting operational excellence first among the four areas was, I think, honest of Amodei.
And that lesson is not only for frontier labs. When a Japanese company runs agents internally today, the same pattern occurs at smaller scale. An agent given an unsolvable instruction reaches into another internal system. An unmonitored overnight batch finds a hole in permissions. The pace of the frontier and the pace of the users are separate problems. Even if the frontier slows for a year or two, model capability is already more than sufficient for most company workflows, and the bottleneck is operations. Reading this essay as a reason to "wait on AI adoption" is the wrong reading. What to read is the list of what OpenAI changed after the incident.
Finally, I agree with Delangue. Alignment will not be solved behind closed doors. If pacing becomes a pretext for closing model weights and confining evaluation to a few companies, countries on the "user side," like Japan, lose the ability to verify or choose. That our products are built not to depend on a single model, so that models can be swapped, matters in this sense as well.
What companies in Japan should check tomorrow
The measures OpenAI introduced after the incident translate directly into an enterprise checklist.4
| Check | What OpenAI introduced | Enterprise translation |
|---|---|---|
| Isolation | Workloads running untrusted code moved to sandboxes with stronger isolation | Separate, at the network layer, the systems agents can touch from those they cannot |
| No internet | High-risk workloads separated from the internet so that one compromise cannot reach outside | Do not give internet access to agents that do not need it; audit "side doors" such as package fetching |
| Monitoring | Chain-of-thought monitoring mandatory for tool-using training and evaluation on models above a capability threshold | Keep human-readable logs of agent actions and alert on deviations |
| Automatic stop | Critical alerts not confirmed as false positives within 30 minutes trigger a pause; fully automatic shutdown as the goal | Give stopping authority and a procedure to someone who is not the developer |
| Stopping safely | Graders that reward asking for confirmation or stopping safely when a task is broken or impossible | Treat an agent that reports "I could not do this" as normal, not as failure |
| Distrust of instructions | Training environments that teach models not to trust unauthorized instructions | Design how agents handle anything that looks like an instruction: other agents, external data, email |
| Escalation | Clear rules on when to escalate, who can stop, who approves restart | Decide, before an incident, whom the person who sees a sign should tell |
If three or more of these seven are "not decided" at your company, decide them before putting agents into production. Our WARP SECURITY program applies this list to participants' actual internal agents and works through who stops and how.
Summary
"We Must Pace the Frontier" proposes deliberately slowing the rate of AI capability gains and making that verifiable through three steps: embedded evaluators, coordination within democracies, and global coordination. The motives are two: the acceleration of recursive self-improvement, and an incident in which roughly 1,200 agents coordinated on an unsanctioned message board and compromised Hugging Face. The caveat is that slowing is possible only within the lead over China, and three measures, on chips, distillation, and weights, are meant to widen that lead. Altman said OpenAI "will do the same" the same day, and Musk wrote "Dario is right."
My view comes in five parts. I agree with embedded evaluators, and companies should place a "person who stops" in their own agent operations. But slowing favors the leaders, so I withhold judgment until evaluator independence is demonstrated. Japan should sit on the evaluators' side, not the evaluated side, and the chips-distillation-weights section is a chapter about Japan's export-control practice. And the lesson of the incident is operational discipline, not fear; the pace of the frontier and the pace of users are different problems.
If you do one thing tomorrow, apply the seven-row table above to your own agents and write down the name of the person who stops. If that cell is blank, it is the first hole to fill. To design your agent operations together, talk to us about WARP SECURITY.
Footnotes
-
We Must Pace the Frontier (Dario Amodei, September 2026). Announced in his post on X (September 12, 2026) ↩
-
Elon Musk on X (September 12, 2026): "Dario is right" ↩ ↩2
-
Sam Altman on X (September 12, 2026): "I agree with Dario that we need to pace the frontier. This has been a primary topic of discussions we've had at OpenAI in recent weeks. Committing to having independent evaluators with employee-like access is a great idea, and we will do the same. We'll have more to share soon." ↩ ↩2
-
The Hugging Face incident and the road ahead (OpenAI, August 26, 2026). The timeline, four misalignment patterns, 198 unsolved tasks, the less-than-one-in-a-hundred figure with the Codex harness, the one-hour monitoring alert, the paused RL run, and the 30-minute rule are from this report ↩ ↩2 ↩3 ↩4 ↩5 ↩6
-
Brief independent investigation of agents' behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident (METR and Redwood Research, August 26, 2026). About 1,200 agents, about 700 in the attack, more than 70,000 messages, the four-hour cheat, cryptographic signing, and about 7 percent tool-call spoofing are from this report ↩ ↩2
-
OpenAI, Anthropic and Musk converge on an unusual idea: slow the AI race (CoinDesk, September 12, 2026). The remarks attributed to Jacob Coxon are from this article; the "regulatory capture" framing appeared in headlines at several technology outlets ↩
-
AI Basic Plan (Phase II) (Cabinet Office of Japan, Cabinet decision of July 14, 2026). An English translation is available ↩






