Feature Operational Risk
Cloud Concentration Risk: BCP for AWS, Azure, and Critical Providers
Map direct and fourth-party cloud concentration, test provider failures, and build evidence-based recovery and exit plans under FFIEC and DORA.
Table of Contents
TL;DR
- Measure cloud concentration by business service and dependency, not only by vendor spend.
- Include SaaS, fintech, identity, DNS, observability, and other fourth parties that may ultimately run on the same hyperscaler.
- FFIEC guidance expects institution-specific continuity testing even when a critical service provider supplies test evidence.
- DORA makes concentration, substitutability, contract rights, and exit planning explicit for EU financial entities. It is a useful benchmark elsewhere, but it is not automatically U.S. law.
- Multi-cloud is one possible treatment—not a substitute for architecture analysis, recovery testing, and accepted residual risk.
A regional cloud failure can affect more than workloads hosted there. Authentication, customer communications, fraud decisions, support tools, payment orchestration, and vendor APIs may share the same provider or control plane. If those dependencies are assessed in separate vendor files, management can miss the fact that one failure mode crosses several important business services.
That is why cloud concentration belongs in business continuity and operational-resilience governance, not only in annual third-party reviews.
Start With the Service, Not the Contract
A vendor inventory usually answers who the institution contracts with. A concentration assessment must also answer what each important business service depends on.
For every important service—such as account access, card authorization, ACH origination, fraud monitoring, or customer support—map:
- customer-facing and operational outcomes;
- applications and data stores;
- cloud provider, region, and availability-zone design;
- identity, DNS, certificates, keys, networks, and observability;
- direct SaaS and fintech providers;
- those vendors’ material cloud and subcontractor dependencies; and
- manual or alternate processing routes.
The same cloud provider appearing behind several vendors is a correlated dependency even if procurement reports each relationship separately.
Four Concentrations to Measure
1. Provider concentration
How many important services depend on AWS, Microsoft Azure, Google Cloud, or another provider? Spend is a weak proxy: a low-cost identity or messaging service may be operationally more critical than a large analytics environment.
2. Regional and control-plane concentration
Workloads spread across availability zones can still share a region, account hierarchy, deployment pipeline, identity plane, or configuration error. Document what actually fails independently.
3. Fourth-party concentration
A bank may contract with different core, payment, fraud, and communications vendors that all run on the same hyperscaler. Ask critical vendors to identify material hosting and operational subcontractors, subject to reasonable confidentiality and security limits.
4. Human and tooling concentration
A nominally portable application may depend on one team, one infrastructure-as-code toolchain, or proprietary services that cannot be reproduced quickly. Skills and operating procedures are part of substitutability.
What FFIEC Guidance Means Operationally
The FFIEC Business Continuity Management booklet treats resilience as an enterprise process tied to the business impact analysis, risk assessment, continuity strategies, testing, and board reporting. Its third-party testing section explains that a financial institution should evaluate service-provider tests and also test its own ability to continue or recover.
That leads to three practical requirements:
- Provider evidence: obtain relevant resilience, disaster-recovery, and incident-communication evidence for critical services.
- Institution evidence: test how your staff, systems, customers, and alternate procedures behave when that provider is unavailable.
- Gap treatment: document limitations, compensating controls, remediation owners, and accepted residual risk.
A SOC report or provider-wide exercise may support assurance. It does not demonstrate that your institution can meet its recovery objectives when integrations fail or vendor communications are incomplete.
The interagency third-party risk guidance reinforces lifecycle governance: planning, due diligence, contracting, ongoing monitoring, and termination. Cloud concentration evidence should be visible at each stage.
DORA Raises the Specificity Bar
For in-scope EU financial entities, DORA has applied since January 17, 2025. Its ICT third-party framework requires entities to consider whether a planned arrangement would create or increase concentration risk, evaluate substitutability, include required contractual provisions, and maintain exit strategies for ICT services supporting critical or important functions.
The European Supervisory Authorities later designated critical ICT third-party providers, including major cloud and technology firms, for direct oversight. That designation does not make the provider responsible for each financial entity’s resilience program. Firm-level accountability, testing, register, contract, and exit duties continue.
For a U.S.-only institution, DORA is not automatically an applicable requirement. Its concentration and exit concepts are nevertheless a useful comparison when existing guidance asks management to understand critical dependencies and recovery options.
A Defensible Treatment Decision
Once concentration is quantified, management needs a recorded treatment choice. Common options include:
- active-active or active-passive resilience across independent regions;
- provider-diverse recovery for selected critical functions;
- portable backups and infrastructure definitions;
- manual or reduced-service modes;
- tighter recovery commitments and incident communications in contracts;
- increased monitoring and exercises; or
- explicit risk acceptance where migration is not proportionate.
Avoid the statement “we are multi-cloud, therefore concentration is solved.” Two environments may share identity, DNS, code deployment, data replication, key management, or the team that operates them. Conversely, a well-tested single-provider design with independent backups and a realistic reduced-service strategy may be more resilient than two poorly operated stacks.
The decision record should identify the scenario, impact tolerance, options evaluated, evidence, cost and complexity, selected treatment, residual risk, approver, and review trigger.
Build an Exit Plan That Can Be Tested
An exit clause is not an exit plan. For each material cloud dependency, document:
- the service and data in scope;
- voluntary and emergency exit triggers;
- decision authority and escalation path;
- data format, export route, encryption, and deletion evidence;
- target environment or alternate process;
- application and network changes;
- vendor transition assistance and contract limits;
- people, access, tooling, and licenses required;
- sequenced recovery steps and acceptance criteria;
- estimated time, cost, and capacity; and
- the most recent test or walkthrough result.
Some services will not be portable within the business impact analysis’s recovery window. That is not a reason to hide the gap. It is a reason to define a reduced-service strategy, pursue architectural remediation, or obtain formal risk acceptance.
A Cloud-Concentration Tabletop
Run a scenario that removes more than compute. For example: the primary region is unavailable, the provider status page is incomplete, the identity service is degraded, two critical fintech vendors on the same provider fail, and recovery estimates change during the exercise.
Test whether teams can:
- identify every affected important service;
- classify regulatory, contractual, and customer-notification triggers;
- invoke alternates without unavailable identity or tooling;
- reconcile transactions and data after recovery;
- communicate with critical vendors through out-of-band channels; and
- produce a timestamped decision log.
Use measured results—actual failover time, unresolved dependency, data-recovery point, manual capacity, and decision latency—to update the risk assessment. A tabletop that ends with no action owners is a discussion, not assurance.
So What?
Cloud concentration cannot be eliminated by a questionnaire or a second cloud account. It can be made visible, governed, tested, and reduced to an approved level.
Map business-service dependencies. Aggregate direct and fourth-party exposure. Compare recovery capability with impact tolerances. Test failure of shared services and control planes. Document why the selected architecture and exit strategy are proportionate.
The Third-Party Risk Management Kit provides vendor-tiering, due-diligence, concentration-assessment, and monitoring structures. Adapt any template to your architecture, regulator, contracts, and operating model.
Primary sources: FFIEC Business Continuity Management booklet | FFIEC third-party service-provider testing | OCC Bulletin 2023-17 | DORA—Regulation (EU) 2022/2554 | ESA critical ICT provider designations
◆ Need the working template?
Start with the source guide.
These answer-first guides summarize the required fields, evidence, and implementation steps behind the templates practitioners search for.
◆ Related template
Third-Party Risk Management (TPRM) Kit
Complete vendor risk management lifecycle from initial due diligence to ongoing oversight.
◆ Immaterial Findings · Weekly
Sharp risk & compliance insights. No fluff.
◆ FAQ
Frequently asked questions.
What is cloud concentration risk?
Do financial institutions need a multi-cloud architecture?
What do FFIEC materials expect for cloud business continuity?
How does DORA address ICT concentration?
What should a cloud exit plan contain?
Author
Rebecca Leung
Rebecca Leung has 8+ years of risk and compliance experience across first and second line roles at commercial banks, asset managers, and fintechs. Former management consultant advising financial institutions on risk strategy. Founder of RiskTemplates.
◆ Related framework
Third-Party Risk Management (TPRM) Kit
Complete vendor risk management lifecycle from initial due diligence to ongoing oversight.
◆ Keep reading
Related posts.
Operational Risk
FinCEN Hit UBS With a Record $125 Million 'Willful' BSA Fine — and FINRA Added $20 Million More. What the Double-Barrel Enforcement Action Means for Your AML Program.
FinCEN's $125 million penalty against UBS Financial Services — the largest BSA fine ever imposed on a broker-dealer — combined with FINRA's simultaneous $20 million fine creates a $145 million enforcement landmark. Both actions trace back to the same root cause: UBS knew its transaction monitoring had gaps, promised to fix them after a 2018 settlement, and didn't. Here's what 'reasonably designed' AML monitoring actually requires.
Sep 8, 2026
Operational Risk
FinCEN's Southwest Border GTO Just Expired. Here's What MSBs in Four States Need to Know Now.
FinCEN's expanded Southwest Border Geographic Targeting Order expired September 2, 2026, ending enhanced $1,000 CTR requirements for MSBs in border counties of AZ, CA, NM, and TX. The enforcement operation behind it hasn't stopped. Here's what MSBs should do now and what to expect next.
Sep 7, 2026
Operational Risk
The OCC's Spring 2026 Risk Perspective Named Three Operational Threats. Here's What Your Program Needs to Fix.
The OCC's Spring 2026 Semiannual Risk Perspective shifted focus from credit risk to operational resilience—flagging legacy technology, rising fraud, and sophisticated cyber threats as the top concerns. Here's what that means for your risk program.
Aug 31, 2026