Cloud Outage Business Continuity: Lessons From AWS and Azure
AWS, Azure, and Google Cloud have each had a significant outage in the past thirteen months, and the service credit sitting in your cloud contract will not cover what a bad afternoon actually costs your business. Cloud outage business continuity is not about picking a more reliable provider. Every major provider has now had a headline outage. It is about having a plan for the day your provider, whichever one you use, has a bad day anyway.
Five Outages, Three Providers, Thirteen Months
Here is the pattern, in order:
Google Cloud, June 12, 2025. A major internal failure disrupted Google's own services and rippled outward to platforms that depend on Google Cloud infrastructure, including Spotify.
AWS, October 20, 2025. A network health monitoring issue took down a wide range of services, centered again on the US-East-1 region in Northern Virginia, a region with a long history of causing trouble.
Azure, October 29, 2025. A configuration change in Azure Front Door's global control plane generated more than 18,000 user reports at its peak and disrupted logins and web traffic worldwide.
Azure, February 2 to 3, 2026. This one ran over ten hours. A policy meant to tighten security accidentally blocked read access on the storage accounts that host virtual machine extension packages, which cascaded into a second failure in Azure's Managed Identities service. The combined disruption reached well beyond virtual machines into Azure Kubernetes Service, GitHub Actions, and Microsoft Copilot Studio.
AWS, May 2026. A thermal event and power loss at a Virginia data center impaired EC2 and EBS, the compute and storage layers a huge share of the internet quietly runs on.
None of this means one provider is worse than another. It means concentrated cloud infrastructure has concentrated failure points, and every provider has now proven it.
The Outage You Do Not See Coming
Most businesses in the 25-to-250-employee range do not think of themselves as AWS or Azure customers. They think of themselves as Microsoft 365 customers, or Google Workspace customers, or users of a dozen other SaaS tools. What gets missed is that Microsoft 365 runs on Azure. Google Workspace runs on Google's own cloud infrastructure. A large share of the SaaS tools in your stack, your CRM, your VoIP provider, the AI-powered chatbot answering routine customer questions, sit on top of AWS or Azure without ever putting either name on an invoice.
That February 2026 Azure outage is a useful example of why this matters now more than it used to. The Managed Identities failure reached Copilot Studio, the platform behind a growing number of business AI assistants. A business that added an AI tool to its stack in the last year did not just add a helpful feature. It added a new dependency on the same infrastructure that already carries email, file storage, and phone systems. When that infrastructure has a bad day, everything riding on it can go dark at once, for a reason that has nothing to do with anything your own team configured.
What an Hour of Downtime Actually Costs
Gartner's commonly cited baseline puts average downtime cost across all company sizes at roughly $5,600 a minute, or about $336,000 an hour. That figure is skewed heavily by large enterprises, so it is not the right number to hang a decision on if you run a 60-person firm.
A more useful exercise is to do the math for a business your own size. Take a 50-person company with an average salary of $80,000. The moment nobody can work, that is roughly $1,923 an hour in idle payroll alone, before a single dollar of lost revenue, missed client deadlines, or recovery cost gets added. Industry surveys of businesses in the 20-to-100-employee range put the fuller number, once lost revenue and recovery expenses are included, somewhere between $5,000 and $50,000 an hour depending on the industry and what systems went down.
Run that math for your own headcount and average pay. It tends to be a bigger number than most owners expect, and it is the number that actually matters, not the industry-wide average.
What Your Cloud SLA Actually Promises
Cloud contracts are clear about one thing: the compensation for an outage is a service credit, not a payment for what the outage cost you. AWS's own EC2 service terms describe the credit as a percentage of what you paid for the affected service that billing cycle, issued only against future charges, only if the credit exceeds one dollar, and only if you file a claim yourself with supporting evidence. The terms explicitly exclude anything caused by a force majeure event, an issue outside AWS's control, or an action on the customer's own side.
The scale of these credits is smaller than most businesses assume. Independent analysis from the Uptime Institute walks through AWS's own EC2 credit schedule: a virtual machine down for under about seven hours in a month qualifies for a 10 percent credit. Azure and Google Cloud use comparable structures, built to acknowledge a miss against a target, not to make a business whole.
What the credit never covers is the part that actually hurts. Cloud contracts routinely exclude lost sales, penalties you owe your own clients for a missed deadline, and reputational harm, and they typically cap direct damages at whatever you paid the provider the month before. An SLA describes the provider's target. It is not a business continuity plan, and treating it as one is where a lot of businesses get caught off guard.
What Actually Closes the Gap
None of this means a 60-person business should build its own multi-region failover architecture. That is not realistic for most growing businesses to run in-house. The useful lesson is narrower: know which critical systems depend on which provider, write down who does what when one goes down, and have someone actively watching provider status pages so your team is not the last to find out.
That last piece is the one businesses without a managed IT relationship tend to miss entirely. When AWS or Azure has a bad morning, the fastest question to answer is whether the problem is on your end or theirs. A business monitoring its own systems in isolation can burn an hour figuring that out. A managed IT partner who is already watching the relevant status pages, already has a documented plan for your specific stack, and already knows which of your tools sit on which provider closes that gap before it costs you the first hour, let alone the fourth.
For related reading, see our breakdown of cloud backup versus disaster recovery if your continuity plan does not yet distinguish between the two, and our post on AI tool outages and what they mean for your business if your bigger concern is the AI layer specifically. For a broader look at what a managed IT relationship actually covers, the managed IT services overview is a good starting point.
Frequently Asked Questions
What happens to my business if AWS or Azure goes down? It depends on which of your tools depend on that provider, which most businesses have not mapped out ahead of time. Microsoft 365 runs on Azure and Google Workspace runs on Google Cloud, so a major outage at either provider can affect email, file access, and any AI or automation tools layered on top, even if you have never directly used AWS or Azure yourself.
Does my Microsoft 365 or Google Workspace subscription protect me from a cloud outage? No. Your subscription includes an uptime target from the provider, not a business continuity plan for your company. If Microsoft or Google misses that target, you may be eligible for a small service credit, but that credit does not cover lost revenue, missed deadlines, or the time your team spends working around the disruption.
How often do major cloud provider outages actually happen? Google Cloud, AWS, and Microsoft Azure have each had at least one significant, widely reported outage within the past thirteen months. This is not a single provider's problem. It is a predictable outcome of how concentrated modern cloud infrastructure is.
Will my cloud provider's SLA credit cover my losses from an outage? Almost never in full. SLA credits are typically a percentage of what you paid for the affected service during that billing cycle, capped at a modest amount, and they explicitly exclude lost revenue, contractual penalties owed to your own clients, and reputational damage.
What should a growing business do to prepare for a cloud outage? Map which critical systems depend on which cloud provider, write down who does what when one of them goes down, and make sure someone is actively watching provider status pages rather than waiting for your team to notice something is wrong on its own.
Not sure which of your business tools depend on which cloud provider, or what happens the next time one of them has a bad day? That is worth mapping out before it matters. Get in touch here.