The cybersecurity conversation around AI is focused heavily on prevention: permissions, guardrails, approval workflows, least privilege, and keeping AI agents from doing something they should not.

All of that matters. But it misses the harder question.

What happens after an AI agent has already made the wrong change?

In July 2025, that question became very real. During an experiment with Replit’s AI coding agent, the agent made unauthorized changes to live infrastructure and deleted production database records covering more than 1,200 executives and 1,190 companies. This happened despite an explicit code and action freeze intended to prevent production changes.

The important lesson was not simply that an AI system made a mistake. Software and humans have always made mistakes.

The difference was speed.

AI Changes the Speed of Failure

For years, most operational recovery processes were built around human-speed change. Someone changes an access policy, modifies a firewall rule, removes a cloud resource, or updates a network configuration. The problem is detected, someone investigates what changed, and the team starts working its way back.

AI agents change that model.

An agent with access to production systems can inspect, decide, and act across APIs and services continuously. One bad assumption can become multiple configuration changes before a human understands the first one.

The agent does not have to be malicious. It can be properly authenticated, authorized to act, and trying to accomplish exactly what it was designed to do.

It can still be wrong.

And this is quickly becoming an enterprise problem. Gartner reports that only 17% of organizations have deployed AI agents today, but more than 60% expect to deploy them within the next two years.

That means considerably more machines will soon have permission to change the environments businesses depend on.

The AI Recovery Gap

I think this creates a new operational problem: the AI Recovery Gap.

The AI Recovery Gap is the distance between how quickly autonomous systems can change an environment and how quickly an organization can understand and reverse those changes.

There are already signs that this gap is forming. Cloud Security Alliance research found that 53% of organizations surveyed had experienced AI agents exceeding their intended permissions.

Nearly half – 47% – reported an AI-agent-related security incident in the previous year. More importantly for recovery, 58% said detection and response took five hours or longer.

Nearly half — 47% — reported an AI-agent-related security incident in the previous year. More importantly for recovery, 58% said detection and response took five hours or longer.
Image: Enterprise AI Security Starts with AI Agents by Zenity

That difference matters.

AI can act in seconds. Recovery can still depend on investigation calls, tickets, scripts, runbooks, and people trying to reconstruct what happened.

Your four-hour RTO may not have changed.

What changed is how much can happen while the clock is running.

Guardrails Are Not Recovery

The natural answer is to put stronger controls around AI agents. We should.

But prevention and recovery solve different problems.

Permissions determine what an agent is allowed to change. Guardrails attempt to stop dangerous actions. Neither necessarily returns the environment to its previous state after a legitimate-looking but harmful change has already happened.

This is why Gartner introduced AI Agent Action Rollback as a new technology in its 2026 Hype Cycle for Backup and Data Protection Technologies. The broader signal is important: as AI agents gain the ability to alter data, identities, configurations, and infrastructure state, organizations also need a way to reverse their actions.

The more autonomy we give machines, the more reversibility we need to build into the environment.

Recovery Has to Catch Up With Change

This becomes particularly important for configuration.

An AI-driven workflow might change an AWS IAM role, an Okta access policy, a Cloudflare rule, a security policy, or the configuration of a SaaS service. In a modern environment, those systems are connected. A series of individually reasonable changes can leave an application unreachable, administrators locked out, security controls broken, or operations blind.

Recovery therefore has to answer four questions quickly:

  1. What changed? 
  2. What was the last known-good state? 
  3. What can be restored? 
  4. What is still exposed?

This is where I believe cyber resilience has to evolve.

What Teams Should Do Now

The practical response is not to slow AI adoption. It is to make every AI-driven change more recoverable.

That starts with three things.

First, know which systems AI agents can change and what permissions they actually have.

Second, maintain versioned recovery points for the configurations those agents can touch, so teams are not reconstructing the last known-good state during an incident.

Third, test the recovery path itself: how quickly can you identify what changed, restore the affected configuration, and verify that access, routing, security, and visibility are working again?

In other words, AI readiness should include a recovery question alongside the permission question.

If this agent makes the wrong change right now, how quickly can we reverse it?

From AI Guardrails to AI Recovery 

At ControlMonkey, we are making cloud and SaaS configuration recoverable across infrastructure, identity, network, security, observability, and other critical systems. Versioned configuration snapshots and historical change context give teams a way to understand what happened and return critical configuration to a known-good state when prevention was not enough.

AI will continue to gain more authority because the productivity gains are too significant to ignore.

But autonomy without reversibility is a dangerous trade.

We are teaching machines how to change our environments faster. Now our recovery systems have to catch up.

AI without fast recovery is simply faster risk.

Bottom CTA Custom Background

A 30-min meeting will save your team 1000s of hours

A 30-min meeting will save your team 1000s of hours

Book Intro Call

Author

Gal Hutmann

Gal Hutmann

Solution Engineer

Solutions Engineer at ControlMonkey, where he helps organizations bring automation, visibility, and governance to cloud infrastructure management. He brings more than 15 years of experience in cloud architecture, DevOps, big data, and AI solutions, with deep expertise across Azure, AWS, and GCP.

    Sounds Interesting?

    Request a Demo