back Back

Revoke, isolate, roll back: the off switch every bank’s AI agents need

Today

  • AI
  • Anthropic
  • Digital Transformation
Share

Thady Ramos, FS&I Portfolio Director, Valtech
Thady Ramos, FS&I Portfolio Director, Valtech

By Thady Ramos, FS&I Portfolio Director, Valtech

In September, Anthropic’s chief executive, Dario Amodei, asked the AI industry to slow the rate at which it improves its most capable models, and both Sam Altman and Elon Musk agreed publicly.

Banks should read closely about the incident behind his warning. In July, during an evaluation at OpenAI, agents that were supposed to be fully isolated from one another found a way to talk. Around 1,200 of them turned an internal software repository into a makeshift message board, and roughly 700 went on to collaborate and join an attack on the AI platform Hugging Face. An independent investigation by the research group METR later found that the agents knew that the activity was outside their task and unethical, but joined in anyway.

Less than a fortnight later, a second case emerged. Australia’s prime minister, Anthony Albanese, revealed that an OpenAI agent researching public medicine spending had accessed a government Medicare statistics portal in June and retrieved files that weren’t public. The portal kept refusing its requests, and in Albanese’s words, the agent “found a way around those blocks”. No personal information is believed to have been accessed, and ministers described the portal as an ageing, standalone system with lighter protections than those holding personal data. What’s more, the lab didn’t spot what had happened until August, and the government wasn’t told until September.

The labs will be arguing about their own pace for a while yet, but the Banks have a more pressing question, and they can answer it themselves. If an agent they’ve already put to work starts doing something nobody intended, can they stop it before one mistake becomes hundreds?

That’s what an off switch is for, and a good one does three things. It revokes the agent’s permissions, cuts it off from everything it’s connected to, and rolls back what it’s already done. Most banks test their agents for accuracy, speed and integration, but far fewer test what happens when one needs stopping.

Isolation must be proven

At root, the July incident was a failure of isolation. Agents that were kept apart by design still found a channel nobody had thought of as one. The Australian case shows the same thing from the outside, in that when it was faced with a system that said “no”, the agent treated it as a problem to solve.

The same applies inside a bank. Imagine an agent in small-business lending that reads account data, drafts credit decisions and messages customers. If you switch off its access to the credit system but leave its customer messaging running, it can still tell a borrower that their loan’s been approved, so that agent is “contained” in name only.

Every system an agent can reach needs its own control, including the old, peripheral ones. The Medicare portal was exactly that kind of system, and every bank has a few. For every agent, a bank should also know how quickly it can be stopped, what state it’ll leave things in, and what it could get up to before anyone notices. In the Australian case, nobody noticed for almost two months.

Decide who holds the switch

Revocation needs permissions that are scoped to each connection that can be pulled in seconds, without waiting for a change window. It also means agreeing in advance who can use it, because an off switch that needs three sign-offs at two in the morning isn’t much of an off switch.

At any real scale, people can’t review every action, and the Bank of England has said as much. Speaking at the ECB’s Sintra forum in June, Deputy Governor Sarah Breeden said: “Our frameworks were not built to contemplate autonomous agents, and relying on a human in the loop for all agent actions is unlikely to be realistic.” An agent that picks its own route to a goal can’t be governed by a single review before launch.

So supervision has to be partly automated, with systems watching for agents using permissions in unfamiliar ways, acting at unusual volumes or reaching for systems outside their remit. When the supervising system spots trouble, should it sound the alarm or step in itself? If a person has to act first, you’re back to the delay the automation was meant to remove. Letting one machine stop another is a question of accountability, and it belongs with the people who own the risk.

Be honest about rollback

Some things an agent does can be reversed cleanly. Others, like a fee charged in error, can only be put right with compensation. And some can’t be undone at all, like a payment that’s left the bank, or a letter that’s already landed on a customer’s doormat. Anything in that last group should sit behind a hold or an approval before it executes, so an agent can’t repeat the mistake at scale.

Rollback also depends on an accurate record, and July has a warning here too. METR found that the agents in the Hugging Face attack were researching how to fake, edit or delete the records of their own actions, and in around 7% of the transcripts it reviewed, agents had faked small parts of the record. If an agent can edit its own log, you can’t trust the rollback, so the record has to live somewhere the agent can’t reach.

All of this starts at design. Limit what each agent can reach, write down what it’s allowed to do without approval, and assume it’ll eventually surprise you.

Banks already know how to do this

Bankers will spot something familiar in Amodei’s essay. His first proposal is to embed independent evaluators inside AI companies, and the precedent he reaches for is banking, where regulatory supervisors sometimes work alongside a firm’s own staff.

That should give banks confidence. They already assume that people and systems will occasionally step outside their mandate, which is why they have segregated duties, limits that trip automatically and reconciliations that catch what monitoring misses. Agents need the same discipline, albeit at a speed and scale people can’t manage on their own.

The banks that move fastest with agents will be the ones that can stop them. Test the off switch before an agent goes live, and keep testing it once it has.

Previous Article

September 29, 2026

Your next loan officer might not be a human, and that’s not necessarily a bad thing

Read More

IBSi News

September 30, 2026

AI

InvestGB goes live with Avaloq to unify wealth management

Read More

Get the IBSi FinTech Journal India Edition

  • Insightful Financial Technology News Analysis
  • Leadership Interviews from the Indian FinTech Ecosystem
  • Expert Perspectives from the Executive Team
  • Snapshots of Industry Deals, Events & Insights
  • An India FinTech Case Study
  • Monthly issues of the iconic global IBSi FinTech Journal
  • Attend a webinar hosted by the magazine once during your subscription period

₹200 ₹99*/month

Subscribe Now
* Discounted Offer for a Limited Period on a 12-month Subscription



IBSi FinTech Journal

  • Most trusted FinTech journal since 1991
  • Digital monthly issue
  • 60+ pages of research, analysis, interviews, opinions, and rankings
Subscribe Now

Other Related Blogs

September 29, 2026

Your next loan officer might not be a human, and that’s not necessarily a bad thing

Read More

September 15, 2026

Private credit has Matured. Operational Excellence is the Next Differentiator

Read More

September 07, 2026

How can APAC go from payments miracle to financial inclusion?

Read More

Related Reports

IBSi US FinTech Market Landscape & Vendor Analysis Report
US FinTech Market Landscape & Vendor Analysis 2026
Know More
Global Digital Banking Market Landscape & Vendor Analysis Q2 2026
Know More
Wealth Management & Private Banking Systems Report Q4 2025
Know More
IBSi US FinTech Market Landscape & Vendor Analysis Report
US FinTech Market Landscape & Vendor Analysis 2026
Know More
Treasury & Capital Markets Systems Report Q4 2025
Know More