Operational disruption is unavoidable. The objective of FCA operational resilience is therefore not to prevent every incident, but to ensure that UK financial firms can continue delivering their most important business services when disruption occurs.
not to prevent every incident, but to ensure that UK financial firms can continue delivering their most important business services when disruption occurs.
The Financial Conduct Authority’s rules require firms in scope to identify important business services, set limits on how much disruption those services can tolerate, map the resources behind them and test their ability to withstand severe but plausible scenarios.
Since the transition period ended on 31 March 2025, firms have been expected to demonstrate that they can remain within the impact tolerances established for each important business service – not simply that plans and backups exist.
TL;DR: FCA Operational Resilience
- FCA operational resilience rules require in-scope UK financial firms to identify the important business services whose disruption could cause intolerable harm to customers or markets.
- Firms must set an impact tolerance for each important business service, defining the maximum disruption that can occur before that harm becomes intolerable.
- Organisations must map the people, processes, technology, information, facilities and third parties behind each service, then test them against severe but plausible disruption scenarios.
- For cloud-dependent financial services, operational resilience should include the ability to recover critical identity, network, security, DNS, observability, SaaS and cloud infrastructure configurations—not only applications and data.
- ControlMonkey supports FCA operational resilience programmes by helping organisations discover, back up, compare and recover the cloud and SaaS configurations their important business services depend on.
What Is FCA Operational Resilience?
Operational resilience is the ability of a firm to prevent, respond to, recover from and learn from operational disruption.
The FCA framework focuses on the delivery of business services rather than the availability of individual systems. The objective is to ensure that disruption does not cause intolerable harm to consumers or threaten the integrity of the wider financial system.
This is an important distinction.
A firm can have working servers, available databases and successful backups while still being unable to deliver a service to its customers. An authentication failure, deleted DNS record, corrupted network route or unavailable third-party platform may be enough to prevent the service from operating.
Operational resilience therefore asks a broader question:
Can the organisation continue delivering the service, or restore it before the disruption causes intolerable harm?
Who Do the FCA Operational Resilience Rules Apply To?
FCA operational resilience is a UK financial-services requirement. It does not apply to every UK business or every FCA-authorised organisation.
The FCA lists firms in scope including:
- Banks and building societies
- PRA-designated investment firms
- Insurers
- Recognised Investment Exchanges
- Enhanced-scope Senior Managers and Certification Regime firms
- Certain payment and electronic-money institutions
- Consolidated tape providers
- Qualifying cryptoasset firms
The FCA rules and guidance came into force on 31 March 2022. The transition period ended on 31 March 2025, by which point firms were expected to have completed the mapping and testing necessary to remain within their impact tolerances.
Although the FCA framework is UK-specific, it reflects a broader international shift. Regulators increasingly expect financial organisations to demonstrate that critical services can withstand and recover from technology and cyber disruption—not merely maintain written continuity plans.
What Does the FCA Expect Firms to Do?
The FCA operational resilience framework can be understood as a continuous cycle:
- Identify important business services.
- Set an impact tolerance for each service.
- Map the resources and dependencies supporting it.
- Test the service against severe but plausible disruption.
- Identify and remediate vulnerabilities.
- Learn from incidents and continue improving.
The framework is not satisfied by completing these steps once. Important business services, tolerances and supporting resources should be reviewed at least annually and following material changes to the business or its operating environment.
“Firms need to expect the unexpected and be prepared to maintain their services in all severe but plausible scenarios to prevent intolerable harm.
The statement was published shortly after the transition period ended and captures the central purpose of the rules: preparation must translate into an ability to maintain services when a serious disruption actually occurs.

1. Identify Important Business Services
The FCA requires firms to identify the services whose disruption could cause intolerable harm to consumers or create risks for markets.
An important business service is considered from the perspective of the outcome delivered to an identifiable customer or market participant. It is not simply an internal department, application or technology platform.
Depending on the organisation, examples might include:
- Allowing customers to access their accounts
- Processing or receiving payments
- Executing financial transactions
- Processing insurance claims
- Providing customers with access to funds
- Supporting merchant payments
- Managing time-sensitive investment instructions
This service-led approach prevents operational resilience from becoming a list of individual systems marked “critical” without an understanding of the customer outcome they collectively support.
2. Set Impact Tolerances
For every important business service, the firm must set an impact tolerance.
An impact tolerance represents the maximum tolerable level of disruption to a service. It marks the point beyond which further disruption could cause intolerable harm to consumers, firms or markets.
Time will normally be an important measure, but it may not be sufficient on its own. Depending on the service, firms may also consider:
- The number of affected customers
- The nature of the affected transactions
- The value or volume of disrupted payments
- Financial loss
- Data integrity
- Market impact
- Customer vulnerability
- Reputational consequences
An impact tolerance is related to, but distinct from, a traditional recovery time objective.
An RTO typically measures the targeted time for recovering a system or technology component. An impact tolerance measures the maximum disruption the business service can sustain before the consequences become intolerable.
Technical recovery objectives may therefore need to sit comfortably inside the impact tolerance. The firm still needs time to validate recovered systems, reconnect dependencies, investigate failures and confirm that the service is usable by customers.
3. Map the Resources Behind the Service
Once an important business service has been identified, the firm must understand what is required to deliver it.
The FCA expects mapping to consider the people, processes, technology, facilities, information and third parties supporting each important business service.
For a cloud-dependent service, this map may include:
- Applications and workloads
- Databases and customer data
- Identity providers and access policies
- Cloud accounts and subscriptions
- Virtual networks and subnets
- DNS and traffic-management records
- Routing tables and load balancers
- Security groups and firewall rules
- Encryption and key-management settings
- Observability, dashboards and alerts
- SaaS platforms and third-party services
- Internal teams and escalation processes
- Manual workarounds and recovery procedures
The mapping exercise should reveal more than which vendors are used. It should show how dependencies connect, which ones are necessary to deliver the service and what would happen if one or more became unavailable.
This is especially important in modern cloud environments, where service configuration is distributed across cloud providers, SaaS platforms, identity systems, networking products and observability tools.
4. Test Severe but Plausible Scenarios
The FCA expects firms to maintain testing plans that demonstrate how they can remain within impact tolerances during severe but plausible disruption.
The scenarios should vary in nature, severity and duration and should reflect the organisation’s actual risks and vulnerabilities. Testing should provide evidence for senior management and the governing body, helping them approve and fund remediation plans.
The FCA has also encouraged firms to mature beyond judgement-based or desktop exercises and incorporate more empirical testing, including:
- Disaster recovery and failover testing
- Simulations
- Penetration testing
- Lessons from real incidents
- Testing involving material third parties
For cloud-dependent financial services, severe but plausible scenarios could include:
- A privileged identity or authentication configuration is deleted.
- Network routes or firewall rules are maliciously altered.
- A critical DNS configuration becomes unavailable.
- A cloud account is compromised.
- An infrastructure deployment corrupts production settings.
- A third-party SaaS platform becomes unavailable.
- Security or observability configurations are lost during an incident.
- A ransomware attack disrupts both production and the recovery path.
- Several dependencies fail at the same time.
- An important environment must be reconstructed from a known-good state.
The purpose is not merely to demonstrate that a backup job completed. It is to determine whether the complete business service can remain available or be restored within tolerance.
5. Identify and Remediate Vulnerabilities
Mapping and testing will often uncover dependencies or recovery assumptions that were not previously visible.
For example, a scenario test may reveal that:
- A critical configuration is not backed up.
- A recovery process depends on one engineer’s knowledge.
- The current recovery point is too old.
- A third party cannot meet the firm’s tolerance.
- Dependencies must be restored in a specific order.
- Recovery instructions do not reflect the live environment.
- Infrastructure created outside approved automation is absent from the plan.
- Restoring data does not restore access, routing or security.
- A service cannot be fully validated before its impact tolerance is breached.
The FCA expects identified vulnerabilities to be prioritised, funded, governed and addressed. Closure should be supported by repeat testing that demonstrates the vulnerability has actually been resolved.
This makes operational resilience an evidence-based improvement programme, not a documentation exercise.
Why Data Backup Is Not Enough for Operational Resilience
Data protection remains essential. Financial services cannot operate without trustworthy customer, transaction and business data.
But restoring data does not automatically restore the environment required to use it.
A recovered database may still be inaccessible if identity policies have been deleted. An application may remain unreachable if DNS records or routing tables are missing. Security teams may lack visibility if dashboards, monitors and alert configurations have been corrupted.
The business service depends on both:
- The data and workloads being protected
- The configuration required to access, route, secure, monitor and operate them
Traditional backup restores data. Operational resilience also depends on restoring the configuration required to operate.
That configuration can exist across:
- AWS, Microsoft Azure and Google Cloud
- Identity providers
- Network and security platforms
- Observability systems
- SaaS tools
- Version-control services
- Third-party operational platforms
Some of it may be represented in Infrastructure as Code. Other resources may have been created or modified through consoles, APIs, scripts, vendor interfaces or automation. An IaC repository alone may therefore not represent the complete state that existed before an incident.

Connect Impact Tolerances to Configuration Recovery
Impact tolerances are set at the level of the business service. Recovery, however, happens across the technology and third-party dependencies behind that service.
For cloud-based financial services, those dependencies may include identity policies, DNS records, network routes, firewall rules, load balancers, monitoring settings and SaaS configurations. If one of these cannot be restored, the service may remain unavailable even after its applications and data have been recovered.
This creates a practical question for resilience teams:
Can the configuration behind an important business service be restored within its impact tolerance?
Answering it requires more than confirming that backups exist. Firms need to know:
- Which configurations the service depends on
- Whether those configurations are protected
- Which previous state can be trusted
- How the configurations would be restored
- Whether recovery has been tested
The FCA does not prescribe a specific technology for this work. But its emphasis on mapping, testing and remaining within impact tolerances means firms need evidence that the full service—not only its data—can recover.

The Configuration Gap in Cloud Recovery
Most disaster recovery plans are built around applications, workloads and data.
That leaves a gap.
A payment platform may have a healthy database but remain inaccessible because its identity configuration was deleted. A customer portal may be restored but unreachable because its DNS or network policies were changed. A service may be running while security and operations teams have lost the dashboards and alerts needed to validate it.
These are not secondary technical details. They are part of the operating environment the business service depends on.
Traditional backup restores data. Operational resilience also requires the configuration needed to access, route, secure, monitor and operate that data.
What Good Configuration Recovery Looks Like
A resilient cloud recovery process should give teams four things:
| Visibility into what exists | Teams need an accurate view of the cloud and SaaS configurations behind important business services, including resources created outside approved automation. |
| Known-good recovery points | Configuration states should be captured over time so teams can understand what changed and select a trusted point for recovery. |
| A tested recovery path | The organisation should know how individual configurations or wider environments would be restored, including dependencies and recovery order. |
| Evidence of recovery readiness | Resilience and security leaders should be able to see what is protected, what can be restored and where gaps remain. |
This makes configuration recovery measurable. It also gives scenario testing something concrete to validate.
How ControlMonkey Supports FCA Operational Resilience
ControlMonkey is a Cyber Resilience Platform for Cloud Configuration Disaster Recovery.
It helps organisations discover, back up, compare and recover critical configurations across cloud infrastructure, SaaS, identity, network and observability.
ControlMonkey continuously discovers configurations across the operating environment, including resources managed and unmanaged by Infrastructure as Code. It captures versioned snapshots that help teams understand what changed and identify a previous known-good state.
When an incident occurs, teams can recover individual resources, configurations or wider environments from those trusted recovery points. They can also review recovery coverage and identify configurations that remain exposed.
This can support several parts of an FCA operational resilience programme:
- Mapping the technology behind important business services
- Identifying cloud and SaaS recovery gaps
- Maintaining known-good configuration states
- Testing configuration recovery
- Providing evidence of recovery readiness
ControlMonkey does not make a firm FCA-compliant on its own. Operational resilience also includes governance, people, processes, communications, facilities and third-party management.
Its role is more specific: helping firms make the cloud and SaaS configuration behind important business services visible, testable and recoverable.frastructure in hours, moving from the two-month cohort toward the three-day one. Infrastructure configuration determines RTO. That is the whole argument in one sentence.
