Everybody in a crisis is worse at their job than on a normal Tuesday. That is not a criticism of your people. It is how humans work under pressure. Sleep is short, information is incomplete, phones will not stop, and the person who normally makes the call is stuck in traffic. Judgment gets expensive exactly when you need it.
A disaster recovery runbook is how you spend that judgment ahead of time. It is a plain document, written on a calm afternoon, that answers the questions you will not want to think through at 3 a.m. Who decides this is a disaster. Who gets called, in what order. What comes back online first. Where the credentials live. What we tell clients. It turns a panic into a checklist, and a checklist is something a tired person can follow.
A Runbook Is Not the Same as a Plan
Plenty of businesses have a continuity plan somewhere. It describes strategy: acceptable downtime, priorities, maybe an insurance summary. Useful for planning, nearly useless at 3 a.m., because it says what should happen without telling anyone what to do.
The National Institute of Standards and Technology draws this line clearly in Special Publication 800-34, its contingency planning guide. NIST says the plan “should contain detailed guidance and procedures for restoring a damaged system unique to the system’s security impact level and recovery requirements.” Detailed guidance and procedures. Not principles.
NIST structures recovery into three phases, and the structure is worth stealing. Activation and Notification “describes the process of activating the plan based on outage impacts and notifying recovery personnel.” Recovery “details a suggested course of action for recovery teams to restore system operations.” Reconstitution “includes activities to test and validate system capability and functionality.” Decide, restore, verify. Write your runbook in that order and it will make sense to whoever opens it.
What Actually Goes In It
Keep it short enough that people will read it. Ten pages is plenty for most small businesses. Here is what earns space.
- A contact list with personal cell numbers. Work numbers ring on systems that may be down. Get personal numbers for every key employee, and get permission to publish them internally. CISA recommends building “a cybersecurity list of key people who may be needed during a crisis” as part of incident response planning.
- Vendor and carrier contacts. IT provider, internet provider, phone provider, payroll company, bank fraud desk, insurance carrier. Include policy and account numbers, because every one of them will ask.
- Who declares a disaster. One named person, with two named backups in order. Declaring is what unlocks spending, calling vendors, and taking systems offline. If nobody has that authority written down, the first hour goes to figuring out who does.
- Restore order by system. A numbered list of what comes back first, second, third, with dependencies noted. More on this below.
- Where credentials live and who can reach them. Not the passwords themselves. The location of the password manager, who has emergency access, and how that emergency access is triggered. If one person is the only path to your systems, that is a finding, not a plan.
- Communication templates. A short message for staff, a short message for clients, and a note on who is authorized to send them. Writing these while calm is dramatically easier than writing them while your phone is ringing.
On that last one, CISA’s incident response guidance assigns a communications role to one person who “will interact with reporters, post updates on social media, and may interact with external stakeholders.” Even in a 30-person company, name that person. Otherwise five people improvise five different stories.
It Has to Exist on Paper
This is the part owners push back on, and it is the part we are least willing to negotiate. A runbook stored only in your document system is a runbook you may not be able to open. The thing that is down might be the thing holding the instructions for fixing what is down.
CISA says it directly in its Incident Response Plan Basics guidance: “Print these documents and the associated contact list and give a copy to everyone you expect to play a role in an incident,” because “your internal email, chat, and document storage services may be down.” That is the entire argument. Print it.
Practically: a printed copy at the office, one at the owner’s home, and a copy on an encrypted USB drive outside your business systems. It feels dated. It is also why some businesses are back in a day and others are still calling around on day three. If your instinct is that cloud services never go down, see the cloud can go down and what that means for your business.
Restore Order Is Where Plans Fall Apart
Ask an owner what to restore first and the answer is usually “everything.” Under pressure, teams restore whatever is easiest or whoever is loudest. That is how a company spends six hours bringing back a file server nobody can log into, because the identity system it depends on is still down.
Build the list by asking what stops the business from taking money and serving customers. For most of our clients the order looks roughly like this: identity and authentication first, then network and internet access, then email, then the one line of business application that the company genuinely cannot operate without, then file storage, then everything else. Your order will differ. What matters is that it exists as a numbered list, with dependencies written next to each item.
Write an honest time estimate beside each system too. Not the vendor’s marketing number. The number you have measured. If nobody has ever measured it, write “unknown” and put it on the list of things to test this quarter. Unknown is uncomfortable, and that is exactly why it is more useful than a guess.
Keeping It Current, and Proving It Works
A runbook decays. People leave, phone numbers change, you migrate a system, and eighteen months later half of it is fiction. NIST Special Publication 800-34 sets the baseline: plans “must be reviewed annually and updated as necessary.” CISA suggests reviewing your incident response plan quarterly and treating it as a “living document.” Somewhere in that range is right for a small business.
- Update it whenever someone with a role leaves. Turnover is the number one cause of a broken runbook. Make it part of offboarding.
- Test one restore per year, end to end. Pick a real system, restore it somewhere safe, and time it. NIST is blunt about why: “Testing validates recovery capabilities, whereas training prepares recovery personnel for plan activation and exercising the plan identifies planning gaps.”
- Reprint after every change. An updated file and a stale printout is worse than no printout, because someone will trust the paper.
- Date every version. Put the revision date in the header of every page. Anyone picking it up should know instantly whether they are holding something current.
The Bottom Line
A DR runbook is not a compliance artifact. It is a letter to your future self, written by the version of you that has time to think, addressed to the version that does not. Most small businesses can draft a workable one in two hours: contacts, who declares, restore order, credential locations, two message templates. Print it, date it, hand it out, revisit it yearly. That is the difference between a bad week and a very bad quarter. It pairs with the broader security work we described in why cybersecurity is no longer optional for mid-sized businesses.
If writing one sounds like a project you will keep postponing, we can build it with you. We interview your team, document the restore order against your actual systems, verify the contacts, and hand you a printed runbook plus a schedule for keeping it honest. Contact us today.
Sources:
Comments are closed