Every so often a large cloud provider has a bad morning and a surprising share of the internet stops working at once. Point-of-sale systems freeze. Airlines stop boarding. Somebody’s doorbell will not open. Then it gets fixed, everybody writes think pieces about resilience for two days, and it is forgotten until the next one.

This is a pattern, not an event. Treat it that way and planning gets simpler, because you stop asking “will this happen again” and start asking “what do we do for the four hours when it does.” Harrison’s line fits here: the cloud is just someone else’s computer. A very good computer, run by very good engineers, in a building with better power protection than yours. Still a computer, and computers have bad days. Your job is not to prevent that. It is to decide in advance what your business does during one.

Why One Provider’s Bad Day Takes So Much With It

Concentration is the whole story. A handful of providers host an enormous share of the world’s business software, and tools you buy from smaller vendors often run on those same platforms without saying so. Your project management tool, phone system, payment processor, and scheduling software may all sit in one data center region, and you would have no way of knowing.

The Uptime Institute noted in its Annual Outage Analysis 2026 that across the nine years it has tracked publicly reported outages, third-party IT and data center service providers, including cloud and internet companies, telecommunications firms, and colocation providers, accounted for about two thirds of them.

The mechanics are usually mundane, which is oddly reassuring. Amazon Web Services published a summary of a disruption affecting its Northern Virginia region on October 19 and 20, 2025, tracing it to a latent race condition in the DNS management system for one database service that produced an empty DNS record for that service’s regional endpoint. DNS is the internet’s address book. One bad entry, and services depending on that database could not find it. By the company’s own timeline, effects on some services persisted roughly fourteen hours. Nobody was hacked. A piece of automation did something unexpected at the wrong moment.

The same shape appears in the July 19, 2024 CrowdStrike incident, where a content configuration update to a Windows security sensor caused widespread crashes. CrowdStrike’s own account notes roughly 99 percent of Windows sensors were back online by 8:00pm Eastern on July 29, 2024, ten days later. Recovery from a global software problem is not instant even when the fix is known.

What You Can and Cannot Do During One

Honesty here separates a useful plan from a fantasy.

You cannot fix it. You cannot escalate it. Your provider already knows, and their status page is the one everyone else is refreshing. Calling your IT company about Microsoft’s authentication service is understandable and futile, and any provider who implies otherwise is selling. You also cannot usefully migrate mid-outage. Standing up an alternative platform while your team is idle and stressed turns a four hour problem into a four week one.

What you can do is keep working on the parts that do not depend on the broken thing, communicate clearly, and avoid making it worse. That is a smaller list than people want, which is exactly why preparation has to happen beforehand.

The Fallbacks Worth Building Before You Need Them

None of these are expensive. All of them have to exist before the outage, because you cannot download a contact list from a system that is offline.

  • Offline copies of the documents you cannot work without. Not the whole file server. The ten or fifteen items that stop the day: the price list, the standard contract, today’s schedule, the safety procedure. A folder synced to local storage on two or three machines is enough.
  • A second communication channel on a different platform. If your email and chat come from the same vendor, they go down together. Pick something independent, tell everyone what it is, and log into it once a quarter so people remember it exists.
  • A phone tree that lives on paper. Printed, in a drawer, updated twice a year. Mobile numbers for every employee and key vendor. It feels antiquated right up until your contact directory is the thing that is down.
  • Cached client contacts. An exported spreadsheet of clients and phone numbers, refreshed monthly, stored somewhere that does not depend on the platform it came from. Slightly stale contact data beats none by a wide margin.
  • A manual fallback for the transaction that pays you. If you take payments, know how to take one without the system. If you book appointments, know how to book on paper. Note what gets re-entered afterward.
  • One page that says who decides. Who declares fallbacks in effect, who posts the client message, who calls vendors. Ambiguity costs more time than the outage.

How to Tell Clients You Are Still Working

Clients rarely get upset about an outage. They get upset about silence. Reach out before they ask, and be specific about what is affected and what is not.

A good message has four parts and takes ten minutes to write in advance: what is happening in plain terms, what it does and does not affect for that client, how to reach you right now, and when you will update them next. Then send the next update, even if it is “still down, next note at 2pm.” Consistency reads as competence.

Two things to avoid. Do not overexplain the technical cause, because your client will read it as excuse-making. And do not promise a restoration time you do not control. “We expect to be back within the hour” is a promise your cloud provider has not made to you.

Draft the template now, while nothing is broken. Writing clearly is harder at 8am with fourteen people standing in your office. We covered more of that in the cloud can go down, and what that means for your business.

What Not to Do About This

Here is the contrarian part, and it matters because outages generate bad advice. For almost every small and mid-sized business, these are wrong answers.

  • Do not build a multi-cloud architecture. Running the same workload across two providers so it survives one failing is genuinely hard engineering. It doubles complexity, and complexity causes its own outages. That is an enterprise answer to an enterprise problem.
  • Do not move everything back on-premises. Your server room is not more reliable than a major provider. It is less reliable, and its failures are yours at 2am.
  • Do not switch vendors out of anger. Every provider has an outage history. Migrating after a bad week means paying real costs to acquire a different set of outages.
  • Do not skip the boring version. The printed phone tree and the offline price list do more in a real outage than any architecture diagram.

The Bottom Line

Cloud outages are a recurring feature of modern business technology, not a scandal and not a reason to retreat. The providers are still more reliable than what you would run yourself. Accepting the pattern means you stop treating each outage as a surprise and start treating it as a scheduled inconvenience of unknown date, which you can plan for in an afternoon.

Do the afternoon. Write the list of what stops, build the offline copies, print the phone tree, draft the client message, name who decides. Then set a reminder to refresh it twice a year. That is the entire program, and it beats any amount of architecture for a business your size.

Harrison Ward Technology helps small and mid-sized businesses across Denton County work out which systems genuinely stop the day, and build fallbacks simple enough to work under pressure. If you are weighing how much of this to handle in-house, we wrote about that in why more businesses are outsourcing IT and what to look for in a partner. For a hand mapping your dependencies before the next bad morning, we are glad to help. Contact us today


Sources:

Comments are closed

This website uses cookies and asks your personal data to enhance your browsing experience. We are committed to protecting your privacy and ensuring your data is handled in compliance with the General Data Protection Regulation (GDPR).