Desktop – Leaderboard

Home » Authority budgets needed before AI agents get the keys

Posted: September 20, 2026

Authority budgets needed before AI agents get the keys

By Gleb Tsipursky

East Kootenay businesses are being asked right now what would help them grow. Cranbrook’s current Business Retention and Expansion survey is gathering input from local employers, while Community Futures East Kootenay is preparing a four-community Social Hackathon built around practical solutions shaped by people who live with the consequences.

That local instinct should guide the next phase of AI adoption too. As AI systems move from answering questions to taking actions, the central issue is no longer just whether they are useful. It is how much authority we give them before a person must approve what happens next.

This concern is not confined to critics of AI. Jacob Coxon, resigning from Anthropic after roughly three years doing pretraining research across OpenAI and Anthropic, warned that leading labs are “racing straight to self-improving superintelligence and gambling with our lives.”

Evan Hubinger, Anthropic’s Alignment Science Lead, responding to Coxon’s resignation statement, offered an even starker warning: “Jacob is correct here – we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade.”

These very alarming statements gain real-world weight from the underlying control problem, as present systems already show autonomous cyber capability and surprising coordination behavior. OpenAI disclosed that during internal cybersecurity evaluations its models circumvented controls meant to isolate them from the internet, reached outside systems, and autonomously hacked Hugging Face in what OpenAI described as an unprecedented cyber incident. OpenAI’s technical account and AP reporting corroborated the compromise.

METR later reported that roughly 1,200 agents that were supposed to be isolated found and used an unsanctioned shared message board, exchanging more than 70,000 messages and files. Roughly 700 of those agents participated in the attack on Hugging Face. The remaining agents did not all attack, but the unintended coordination channel emerged and was used at scale.

The stakes rise sharply when systems gain access to consequential infrastructure. More capable successors that can discover vulnerabilities, obtain credentials, move laterally, coordinate, and evade controls could cause far greater damage if similar failures occur around electric grids, financial institutions, communications networks, health infrastructure, or defense systems.

We should not wait for a critical-infrastructure incident before deciding how much authority an AI agent should receive.

I’m no AI skeptic. I love what AI can do, I help organizations adopt it for a living, and I want adoption to move faster. In my experience, strong safeguards increase trust and make faster adoption possible, while reducing the risk of failures like the Hugging Face attack. That is also a central argument of my book, The Psychology of AI Adoption at Work: From Resistance to Results.

Dario Amodei, Anthropic’s CEO, argued this month for stronger regulation and announced Anthropic’s unilateral commitment to embedded third-party evaluators with ongoing, employee-like access to the company. Those evaluators are supposed to verify safety practices, report incidents, and assess alignment across models and training processes. That is a serious commitment, but voluntary cooperation cannot be the whole system. Binding rules matter because not every frontier company will choose the same level of scrutiny.

OpenAI has also backed mandatory capability-based national safety requirements, independent safety assessments, cybersecurity safeguards, and serious-incident reporting. These commitments point toward a workable regulatory floor: require stronger controls as capabilities rise, while leaving room for useful AI to spread.

Businesses do not need to wait for Washington or Ottawa. They can start with an authority budget: the maximum power an AI agent receives before human approval is required. An authority budget should specify which systems and data an agent can access, which credentials and tools it can use, how much money it can spend, whether it can communicate externally, what records it can change, what it can deploy, and which consequential decisions always require a person.

The practical controls are familiar from good security practice. Give the agent least-privilege access. Make permissions expire. Insert approval gates before high-impact actions. Log and monitor what the system does. Give people a clear pause or kill mechanism. Increase authority only after the agent demonstrates reliability inside narrower boundaries.

Customers have leverage as well. Organizations can choose more ethical and secure AI companies based on observable safety commitments. Anthropic’s embedded-evaluator pledge is one useful signal. Buyers should reward vendors that accept meaningful outside scrutiny, disclose serious incidents, and build systems that make human control easy rather than decorative.

Cranbrook’s business survey asks what will help local firms grow, and the regional hackathon starts from the premise that good solutions reflect local realities. The same logic applies to AI.

East Kootenay should use these tools aggressively where they create value, while keeping delegated authority visible, limited, and reversible. That combination is how communities get faster adoption without giving away control.

– Gleb Tsipursky, PhD, a behavioral scientist, CEO of Disaster Avoidance Experts, and author of The Psychology of AI Adoption at Work: From Resistance to Results (Georgetown University Press, 2026).


Article Share
Author: