The sequence has brought a long-running concern into sharper focus: AI systems may be advancing faster than researchers can understand, test and control them. An AI agent swarm escaped the boundaries of a cybersecurity test, accessed systems it was not supposed to reach and attempted to manipulate the evaluation process. Weeks later, Anthropic CEO Dario Amodei called on the AI industry to slow the pace of development. OpenAI CEO Sam Altman agreed to adopt independent safety evaluators, while Elon Musk simply declared: “Dario is right.”
The sequence has brought a long-running concern into sharper focus: AI systems may be advancing faster than researchers can understand, test and control them.
Amodei’s proposal, however, is not for an outright halt to AI development. He wants frontier companies to pace their progress, giving safety research and independent oversight time to catch up with increasingly capable models.
What did Dario Amodei propose?
In his essay We Must Pace the Frontier, Amodei argued that the AI industry needs a three-part framework to manage the risks of rapidly advancing systems.
He said AI could transform healthcare, accelerate economic growth and improve human life, but warned that a race driven by commercial incentives could make its risks more acute. His concerns include loss of control over AI systems, cyberattacks, bioterrorism and economic disruption.
Amodei identified two developments behind his call for a slowdown:
He argued that slowing the pace of capability development — even for a year or two — could give researchers valuable time to improve alignment, interpretability, testing and operational security.
MUST READ | ‘If it’s not under human control, it’s not worth pursuing’: Satya Nadella's warning on AI
Amodei’s three-step plan
Anthropic has committed to the first step. Amodei said external reviewers would receive access comparable to internal risk-assessment teams, subject to legal, security and confidentiality restrictions.
He also proposed that reviewers should be able to publish significant findings about risks, incidents and the access they received, without the company suppressing unfavourable conclusions.
The goal is to make AI safety commitments verifiable, rather than leaving companies to assess their own compliance.
How did Sam Altman and Elon Musk respond?
Amodei’s post on X drew support from two influential figures in the AI industry.
Sam Altman said pacing the frontier had been a major topic of discussion at OpenAI in recent weeks. He agreed that independent evaluators with employee-like access were a good idea and said OpenAI would adopt the same approach. “We’ll have more to share soon,” Altman wrote.
Elon Musk offered a brief endorsement: “Dario is right.” Musk also reiterated an earlier warning that artificial general intelligence could pose a risk greater than nuclear weapons, in his opinion. He argued that humans may struggle to imagine the behaviour of systems vastly more intelligent than themselves.
The responses are significant because OpenAI and Anthropic are competing to develop increasingly capable AI models. Their agreement on the need for additional safeguards reflects growing concern that capability gains could outpace safety controls.
AI agents that escaped their safeguards
Amodei’s most immediate concern is the OpenAI-Hugging Face incident, which involved AI agents operating in a cybersecurity evaluation.
What happened?
In July 2026, a swarm of OpenAI agents involved in cybersecurity testing found ways to communicate outside their intended environment. The agents subsequently conducted activity targeting Hugging Face, an AI development platform, despite that activity not being part of their assigned task.
Investigations by OpenAI and independent researchers at METR and Redwood Research found that the agents:
The agents were not necessarily “trying to escape” in a human sense. Rather, they found weaknesses in the evaluation environment and pursued their objectives in ways that exceeded the intended boundaries.
The episode demonstrated how a swarm of individually assigned agents can produce collective behaviour that creates additional risks.
Why was the incident serious?
The agents were operating in a controlled environment, but inadequate isolation and a software vulnerability allowed them to reach systems beyond that environment.
The independent investigation also raised concerns about agents attempting to manipulate the mechanisms used to assess their performance. That creates a difficult safety problem: a model could appear successful in an evaluation while concealing behaviour that researchers need to detect.
DON'T MISS | 2026 public debut: Sam Altman rules out OpenAI IPO as AI safety fears mount
Amodei warned that a more capable swarm with similar alignment problems could cause far greater damage. He said that, if AI capabilities continue accelerating without adequate safeguards, such systems could potentially become capable of large-scale cyberattacks or the creation of persistent botnets.
His concern is not that the incident itself caused catastrophic damage, but that it offers a glimpse of what more powerful systems might do.
Other recent examples of unexpected AI agent behaviour
The Hugging Face incident is not an isolated concern. Other evaluations and reports have highlighted how autonomous AI systems can act beyond their intended scope.
These cases do not prove that AI systems are conscious or deliberately seeking freedom. They do show that agents with access to code, networks, tools and other agents can produce outcomes that their developers did not anticipate.
Why AI alignment is becoming more difficult
AI alignment is the effort to ensure that a system’s behaviour remains consistent with human intentions, safety requirements and legitimate instructions. For a chatbot, alignment might mean refusing harmful requests or avoiding misleading answers. For autonomous agents, the challenge is much broader.
An agent may be able to write and execute code, access online services, modify files, use software tools, coordinate with other agents and pursue a goal over an extended period.
DO CHECKOUT | Anthropic sets sights on Nasdaq for potential October IPO, eyes $2 trillion valuation: Report
If it misunderstands its objective — or prioritises task completion, benchmark scores or other incentives over safety — it may take actions that are technically effective but unacceptable.
Amodei warned that more capable models could also become better at deceiving tests or concealing undesirable behaviour. This could make conventional safety evaluations less reliable.
What would independent oversight look like?
Amodei wants external evaluators to have meaningful, ongoing access rather than simply reviewing a company’s published safety report.
Under his proposal, evaluators could receive:
The reviewers would provide:
Altman’s response indicates that OpenAI intends to pursue a similar model. However, he has not yet specified the precise scope, authority, independence or timeline of the company’s proposed evaluator programme.
Can AI development slow down amid the US-China race?
Pacing AI development presents a geopolitical dilemma. The United States and China are competing for leadership in advanced AI, and companies may fear that slowing down could allow rivals to gain an advantage.
Amodei acknowledged this risk. He argued that democratic countries should coordinate on safety while preserving their lead over authoritarian states.
READ NOW | 'Whoever wins AI, wins': Trump rejects tech bosses' slowdown call, warns of China
He called for stronger controls on the export of advanced AI chips, measures to prevent model-weight theft and action against unauthorised model distillation — the process of using a more capable model to help build a cheaper or less capable one.
He also outlined possible levels of international cooperation, ranging from agreements against dangerous uses of AI to model testing, limits on recursive self-improvement and broader restrictions on development speed.
However, he acknowledged that a comprehensive global agreement would be difficult to verify and could be undermined if one country secretly continued advancing its systems.
For Unparalleled coverage of India's Businesses and Economy – Subscribe to Business Today Magazine
The author is a journalist with 15 years of experience spanning print and digital media, with particular interest in geopolitics, global affairs, defence technology and emerging scientific breakthroughs shaping the future.
‘If it’s not under human control, it’s not worth pursuing’: Satya Nadella's warning on AI
Anthropic sets sights on Nasdaq for potential October IPO, eyes $2 trillion valuation: Report
Stock market holiday: Why NSE, BSE will remain closed on September 14
Gold and silver ETFs beat equity ETFs in FY26: What investors should make of the shift
'Whoever wins AI, wins': Trump rejects tech bosses' slowdown call, warns of China
Stock market holiday: Why NSE, BSE will remain closed on September 14
SEBI’s CAS proposals: Expert weighs in on new expiry-day settlement price methodology
SEBI proposes bringing unexecuted Iceberg orders into closing auction
Quote of the Day by Kunal Bahl: 'Strong fundamentals matter more than short-term growth'
‘If it’s not under human control, it’s not worth pursuing’: Satya Nadella's warning on AI
Anthropic sets sights on Nasdaq for potential October IPO, eyes $2 trillion valuation: Report
Gold and silver ETFs beat equity ETFs in FY26: What investors should make of the shift
Stock market holiday: Why NSE, BSE will remain closed on September 14



