TL;DR

  • Azure Backup protects data and workloads, but not the configuration required to operate – including networking, RBAC, policies, and identity.
  • RTO and RPO need to reflect the full business service, including the time required to recover configuration, not just restore data.
  • Recovery order matters: identity → networking → security → infrastructure → applications → data. Restoring resources without their dependencies can still leave the environment unusable.
  • ControlMonkey complements Azure Backup by making Azure configuration recoverable, with versioned snapshots and recovery from known-good states – helping restore the operating environment before data and workloads come back online.

At enterprise scale, an Azure outage or ransomware event rarely comes down to “can we restore the data?” It comes down to whether the environment the data lands in – the virtual networks, the identity roles, the policies, the dependencies between resources – still exists in a form the application can actually run in. Azure Backup is built primarily to protect and recover data and workloads. Recovering the broader configuration required to operate is a separate part of the recovery problem.

What Does Azure Cloud Backup Protect?

Azure Backup protects data and workloads, including Azure VM disks, SQL and SAP HANA databases, Azure Files, and Blob storage. It creates point-in-time recovery points stored in Recovery Services vaults or Backup vaults, while Azure Site Recovery supports VM failover for disaster recovery. However, protecting workloads and data does not necessarily protect the broader configuration required to operate – such as identity, networking, security policies, and other cloud configuration.

What none of this covers is the configuration layer that the workload actually depends on: which subnet a VM sits in, which network security group rules allow it to talk to anything, which Azure Policy definitions and RBAC assignments govern who can touch it, and which identity roles and Conditional Access policies gate who can even sign in once the environment is back up. A backup of a VM disk restores the disk. It doesn’t restore the resource graph around it.

Where ControlMonkey Fits Into Azure Backup and Recovery

ControlMonkey complements Azure Backup by protecting and recovering the configuration layer that Azure environments depend on to operate. While Azure Backup focuses on restoring data and workloads, ControlMonkey continuously discovers and snapshots configurations across identity, networking, security, infrastructure, and third-party systems, enabling teams to restore them from a known-good state before bringing applications and data back online.

Azure Backup and ControlMonkey aren’t solving the same problem, which is exactly why they need to be planned together instead of separately:

  • Azure Backup → data and workloads
  • ControlMonkey → cloud and SaaS configuration
  • Together → a more complete recovery of the operating environment, not just the files inside it

This distinction is the throughline for the rest of this article. Every practice below either strengthens the data-recovery side (Azure Backup’s job) or the configuration-recovery side (where ControlMonkey fits) – and the practices that most enterprise runbooks miss are almost always on the configuration side.

icon

Azure Backup protects your data. What protects the configuration required to operate?

Check how ControlMonkey extends disaster recovery beyond data backup by making critical cloud and SaaS configurations recoverable.

1. Define RTO and RPO Based on Business Criticality

Recovery time objective (RTO) and recovery point objective (RPO) are usually defined per resource or per workload – “this database gets a 15-minute RPO.” That’s a reasonable start for data recovery, but it doesn’t answer the question the business is actually asking, which is: how long until this service is back.

RTO and RPO should be set at the business-service level – order processing, customer authentication, the billing pipeline – and then decomposed into the RTOs of every dependency underneath it: the data, yes, but also the infrastructure, the network path, the identity and access configuration, and the third-party or SaaS dependencies the service relies on. A database that restores in 20 minutes doesn’t help if the application can’t authenticate against Entra ID for another two hours because nobody defined an RTO for identity configuration recovery at all. This is the starting point for most Azure backup disaster recovery best practices worth following: get the objective right at the service level before optimizing anything underneath it.

2. Define Azure Backup Frequency and Retention by Workload

There’s no universally “correct” Azure backup frequency – the right answer depends on how much data loss the business can actually tolerate for that specific workload, and there’s real tension between backing up often enough to hit an aggressive RPO and the storage cost and operational overhead of doing so.
Frequency and retention should be set per workload against a few concrete inputs:

  • Workload criticality – what breaks downstream if this data is unavailable or stale
  • RPO requirements – how much data loss between backups is acceptable
  • Short- vs. long-term retention – daily/weekly recovery points for operational recovery, separate from monthly/yearly retention kept for compliance or long-tail investigation
  • Regulatory requirements – retention windows that are set by a compliance framework, not by the team

Treating every workload with the same backup policy is a common source of both wasted spend (over-protecting low-criticality data) and unpleasant surprises (under-protecting something that turned out to matter). Among Azure backup policy best practices, workload-by-workload differentiation is the one that’s easiest to skip under time pressure and most expensive to skip in practice.

3. Apply Azure Backup Security Best Practices

Backup infrastructure is itself a target. Ransomware operators that get into an environment increasingly go after the backup vault before they detonate, on the assumption that a victim with no recovery point is far more likely to pay. Azure Backup has purpose-built defenses against exactly that scenario:

  • Backup isolation – keeping backup vaults and the roles that manage them separate from production management paths, so a compromised production credential doesn’t automatically have a path to the backup data too.
  • Immutable vaults – Azure Backup supports locking a vault’s immutability setting, which blocks any operation that would reduce retention or delete recovery points before their expiry – and once locked, that setting can’t be reversed, including by an attacker with admin access.
  • Ransomware resilience as a layered control – Azure Backup now enforces soft delete by default on new vaults, so deleted backup items are recoverable for a retention window rather than gone immediately. Locked immutability closes the gap soft delete leaves open (an attacker who can still shorten retention or purge soft-deleted items). Multi-user authorization, via a separate Resource Guard resource ideally owned by a different administrator, adds a second-approval requirement to critical vault operations – so no single compromised account can unilaterally delete backup data.

These three controls protect the backup vault and the recovery points inside it. They deliberately don’t cover who’s allowed to sign in with elevated access, what MFA policy applies to that sign-in, or how recovery operations get authorized at the identity layer – that’s access configuration, not backup data, and it’s covered in Sections 4 and 5 below, where it actually lives.

4. Protect Azure Configuration Alongside Your Backups

This is where most enterprise backup strategies quietly stop, and where ControlMonkey’s role starts. Recovering a workload also means recovering everything the workload needs to run inside Azure:

  • Virtual networks (VNets) and subnet layout
  • Network security groups (NSGs) and the rules that govern east-west and inbound traffic
  • Azure RBAC – including the least-privilege role assignments and separation-of-duties boundaries that keep any one identity from having more access than its job requires
  • Azure Policy definitions and assignments that enforce compliance and guardrails across subscriptions
  • Routing – route tables, peering, and the paths traffic actually takes
  • Resource configurations and dependencies more broadly – the settings and relationships that don’t show up in a VM disk snapshot at all

Why IaC Alone Is Not an Azure Configuration Backup

Terraform or another IaC tool is the right way to provision Azure infrastructure, but an IaC repository is not a backup of that infrastructure’s actual state. Manual changes made directly in the Azure portal during an incident, emergency firewall rule additions, temporary RBAC grants that never got cleaned up – all of this creates configuration drift, and drift means the code in your repository stops matching what’s actually running in production. If you recover from the IaC repo alone, you recover the environment as it was designed, not the environment as it existed the moment before the incident – including whatever undocumented changes were keeping things working.

How ControlMonkey Protects Azure Configuration

ControlMonkey continuously protects Azure configuration by capturing versioned recovery points of the actual resource state – not just the IaC that was supposed to produce it – and surfacing drift between the two as it happens, rather than during a post-incident audit. When configuration needs to be recovered, that means restoring to a known-good state that reflects what was genuinely running, including the manual changes an IaC repo would have missed entirely.

This is the same problem the Block case study addresses directly: recovering configuration across a complex multi-cloud footprint fast enough to matter, instead of manually reconstructing it resource by resource.

5. Include Microsoft Entra ID in Azure Backup and Recovery

Every recovery scenario above assumes someone can actually sign in to Azure once the infrastructure is back – and that assumption depends entirely on Entra ID configuration surviving the incident intact. This is also, not coincidentally, the layer most enterprise DR plans forget to plan for at all.

Entra ID recovery needs to account for:

  • Users and groups – including the group memberships that drive access and Conditional Access scoping
  • Roles – both built-in and custom directory roles, and specifically the privileged role assignments that are the most attractive target for an attacker and the most damaging to lose or leave misconfigured after a restore
  • Conditional Access policies – including MFA enforcement, which is a policy setting, not a credential – losing or misconfiguring it during recovery either locks legitimate users out or, worse, silently drops the MFA requirement altogether
  • App registrations and service principals – the identities that automation, CI/CD, and service-to-service auth depend on, which are easy to overlook because they don’t belong to a person

Microsoft’s own native Entra ID Backup and Recovery, currently in public preview, is a meaningful step here – it automatically takes daily snapshots of a defined set of directory objects and lets admins generate a difference report and run a targeted restore. It’s also, by design, narrow: retention is measured in days, the object set it covers is still expanding, and it’s explicitly built to complement existing identity resilience practices, not replace them. ControlMonkey’s role is to protect and version Entra ID configuration – roles, Conditional Access, app registrations – alongside Azure infrastructure configuration, so identity and infrastructure configuration can be protected as part of the same configuration-recovery strategy, rather than treated as separate DR workstreams.

6. Design Azure Backup and Recovery Across Multiple Subscriptions

Enterprise Azure estates aren’t one subscription – they’re tens or hundreds of them, organized under management groups, often spanning shared services subscriptions, multiple regions, and centralized networking and security policy that’s meant to apply consistently across all of it.

At that scale, backup and recovery planning has to account for subscription-level blast radius (what’s isolated vs. shared), where centrally-enforced policy actually lives and whether it’s recoverable on its own, and how a recovery in one subscription affects shared dependencies – like hub networking or centralized identity – that other subscriptions rely on. A backup strategy designed for a five-subscription environment usually doesn’t scale cleanly to fifty without deliberate redesign.

ControlMonkey provides recovery-readiness visibility across the Azure estate – helping teams understand what configuration is protected, what changed, what is recoverable, and where recovery gaps remain.

7. Map Dependencies and Automate the Azure Recovery Sequence

Recovery order matters as much as recovery completeness. The dependency chain in a typical Azure recovery runs roughly:

Identity → Networking → Security → Infrastructure → Applications → Data

Restoring everything at once, or restoring resources independently without regard to this order, doesn’t guarantee the application comes back – an application restored ahead of the network path it depends on, or ahead of the identity configuration it authenticates against, just fails differently than before. Automating the sequence – so infrastructure, security, and networking configuration recover in the right order before data and applications come online – is what makes an application-level recovery actually reliable, instead of a race between resources that happen to restore at different speeds.

ControlMonkey automates recovery of the configuration layers that need to be in place before applications and data can operate, including identity, networking, security, and infrastructure configuration. Its dependency-aware recovery helps restore configurations in the right order, reducing manual rebuilding and human error. This places ControlMonkey upstream of Azure Backup in the recovery sequence: ControlMonkey restores the configuration required for the environment to operate, while Azure Backup restores the data and workloads that run within it.

8. Test Complete Azure Recovery, Not Just Backups

The equation that actually matters for DR testing is:

Data + Identity + Configuration + Infrastructure + Dependencies → Business Service

ControlMonkey Complete Azure Recovery

Confirming that a specific backup can be restored is a necessary test, but it only validates one term in that equation. It doesn’t prove the business service comes back within its defined RTO, because it says nothing about whether the identity configuration, network path, and dependency chain around that data are also in a recoverable state.

Recovery testing at the enterprise level should validate the whole equation – which means including configuration recovery and Entra ID recovery in the same exercise as the backup restore, not as a separate, less-rehearsed workstream. ControlMonkey’s configuration recovery should be part of that same test, exercised on the same schedule as the Azure Backup restore test, so both sides of the recovery are validated together instead of one side being assumed to work because the other one does.

Azure Cloud Service Backup Across Multiple Subscriptions

How ControlMonkey Supports Multi-Subscription Azure Recovery

Across a complex, multi-subscription Azure estate, ControlMonkey supports recovery with:

  • Centralized configuration protection across Azure subscriptions and management groups, not managed subscription by subscription
  • Versioned configuration snapshots that provide known-good recovery points
  • Drift visibility that surfaces where live configuration has diverged from what’s documented or intended, before that gap becomes an incident
  • Recovery automation that applies known-good configuration back across affected subscriptions without manual, resource-by-resource rebuilding
  • Cross-platform configuration recovery – extending the same model beyond Azure to the other cloud and SaaS platforms an enterprise estate typically depends on

Azure Backup and Azure Site Recovery will keep doing what they’re built for: recovering data and workloads. The eight practices above only add up to a complete recovery when the configuration, identity, and dependency layers are planned for with the same rigor – which is the gap ControlMonkey is built to close.

Bottom CTA Custom Background

A 30-min meeting will save your team 1000s of hours

A 30-min meeting will save your team 1000s of hours

Book Intro Call

Author

Gal Hutmann

Gal Hutmann

Solution Engineer

Solutions Engineer at ControlMonkey, where he helps organizations bring automation, visibility, and governance to cloud infrastructure management. He brings more than 15 years of experience in cloud architecture, DevOps, big data, and AI solutions, with deep expertise across Azure, AWS, and GCP.

    Sounds Interesting?

    Request a Demo