Skip to content

CISSP Certification Guide / Chapter 5

Business Continuity and Disaster Recovery Planning

The business impact analysis that values recovery, the time metrics that set the targets, recovery strategies for data, systems, sites, and people, and the test ladder from checklist to full interruption.

The plan nobody reads until it is the only thing that matters

Chapter 4 gave you the machinery of risk: threats, vulnerabilities, likelihood, impact, and the response ladder that turns analysis into a decision with an owner. Business continuity and disaster recovery are what happen when the decision loop has to run at three in the morning with the building unreachable. Everything you learned about risk applies, but the pace changes and the stakes concentrate. A risk register entry that said “tolerate a low likelihood of extended outage” stops being an abstraction the moment the outage is real, and the only asset left is the thinking you did in advance.

That is the first and most important thing to understand about this material, and it is the reason the exam keeps returning to it. Continuity planning is not a technology discipline. It is a discipline of deciding in advance: deciding how long the organization can be down, deciding how much data it can afford to lose, deciding who gets to call the event a disaster, and deciding where the work will happen while the usual place is gone. The technology, the backup software, the alternate site, the failover scripts, is the last mile of a long chain of decisions that are fundamentally business decisions. A CISO who picks a recovery strategy before the business impact analysis exists has built the last mile of a road to nowhere.

The second thing to understand is that the exam tests the shape of the discipline more than its technology. It asks which step comes first, which metric measures what, which site matches which target, and which test proves what claim. The answers are rarely obscure. They are wrong when a candidate confuses the order, swaps two metrics, or buys a strategy that does not fit the target the BIA established. Get the shape right and the questions become almost mechanical.

Here is the shape. Analysis comes first: the business impact analysis finds out what the organization actually loses over time when a process stops, and it produces the targets. Then strategy: recovery options for data, systems, sites, and people are matched to the targets and the budget. Then the plans: the documents that record activation authority, procedures, and reconstitution. Then the program: testing, training, and maintenance, because a plan that has never been exercised is a document with good intentions. Everything in this chapter is one of those four layers, and the exam is testing whether you know which layer you are in and what belongs there.

The vocabulary of plans: BCP, DRP, COOP, OEP, and the program that holds them

The discipline answers to a family of names, and the names overlap enough that the exam can and will use them as discriminators. The reference that keeps them straight for the US federal world is NIST Special Publication 800-34 Revision 1, “Contingency Planning Guide for Federal Information Systems,” which organizes contingency planning in three tiers: business continuity planning for mission and business processes, continuity of operations planning for essential functions at an alternate site, and information technology contingency planning for the systems themselves. Those tiers are not competitors. They are the same concern at three altitudes.

The business continuity plan (BCP) is the broadest document. It addresses the processes that keep the organization in business through and after a disruption, not just the computers that support them. If the payments team cannot do its work because the office is closed, the BCP is the document that says where they work, how they are reached, what they need, and what they do first. It is written for the business, and it is owned by the business.

The disaster recovery plan (DRP) is the technical sibling. It restores information systems, applications, and data after a major failure, and it is the plan that names the alternate site, the restore procedures, the failover sequence, and the point to which data will be recovered. In most organizations the DRP is the thickest document in the set, because it is the one the engineers actually execute. In the plan family, think of the DRP as the child of the BCP: the business plan says the revenue process must run again within four hours, and the DRP is how the systems make that promise true.

Continuity of operations planning (COOP) sits between the two. It concentrates on essential functions, the small set of activities the organization cannot stop doing, and on running those functions from an alternate location. The term comes from the federal world, where agencies must keep essential functions going after an event that removes their primary site, but the concept is universal: a bank, a hospital, a manufacturer, and a government agency all have a handful of functions without which they cease to be what they claim to be, and COOP is the plan for those.

Below and beside these sits the occupant emergency plan (OEP), which is often forgotten and always first in sequence. The OEP addresses life safety within a facility: evacuation routes, assembly points, wardens, head counts, and the immediate actions of the first minutes. It is the only plan that is always executed, because it covers events too small to be disasters and the first moments of events that become disasters. The ordering that matters for the exam is the ordering that matters in reality: life safety comes before business resumption, which comes before system recovery. A plan that has people standing in a burning lobby while the servers are recovered is not a plan, it is a liability. The crisis communication plan rounds out the family: it says who learns what from whom, who is authorized to speak, and what the messages are, to employees, customers, regulators, and the public.

Hold the vocabulary as a ladder: OEP for the immediate minutes, BCP for the business process, DRP for the systems, COOP for the essential functions at an alternate site, crisis communication for the story. Underneath all of them sits the business continuity management program, the governance layer that funds them, schedules their exercises, and makes sure they stay current. A pile of plans without a program is a shelf; the program is what makes the shelf matter.

The BIA: the analysis that makes the plan honest

Every continuity document in the world is downstream of one analysis, and that analysis is the business impact analysis (BIA). The BIA answers the question the whole discipline exists to answer: what does this organization lose, and how quickly does the loss become unacceptable, when a given process stops? It is the first major step in building a continuity program, and the exam treats it as the hinge the rest of the plan turns on. If you remember one ordering rule from this chapter, make it this one: the BIA comes before the strategy, the strategy before the plan, the plan before the test.

NIST SP 800-34 Revision 1 distills the BIA process into three activities. The first is determining mission and business processes and their recovery criticality: identifying what the organization does, and ranking it. The second is identifying resource requirements: the people, systems, data, facilities, and third-party services each critical process needs to run. The third is identifying dependencies: which resources a process depends on, and which other processes depend on it. Do those three things honestly and the BIA writes itself.

Recovery criticality is the ranking exercise. The organization’s processes are not equal in the eyes of a disruption. Some must resume in minutes because the loss compounds immediately; a trading desk, an emergency room, a payment gateway. Some can wait hours, some a day, some a week. The BIA does not ask what the IT department can restore. It asks what the business needs, and the gap between the two is the work the rest of the program exists to close. This is the first place candidates lose their way, because the instinct is to start from the systems, what runs on what, and work outward. The BIA starts from the process, what the organization must keep doing, and works inward to the resources. Process first, systems later.

Dependencies are the second trap and the second source of BIA value. A process rarely fails alone. The order-to-cash process depends on the order database, which depends on the storage array, which depends on power, which depends on the facility, which depends on the utility grid. It also depends on people, on a payment processor, on a shipping carrier, and on a regulatory filing window. Mapping those dependencies is what turns a process ranking into an actual recovery design, because the recovery of any process is a chain of dependencies, and the chain is only as strong as its weakest, slowest, or most forgotten link. The BIA also captures the reverse dependencies: which other processes wait on this one, so that the plan recovers in an order that unblocks the most work the fastest.

The signature insight of the BIA is that impact is a function of time. A two-hour outage of the order entry system may cost nothing, because customers wait. A two-day outage may cost real revenue, because competitors get the orders. A two-week outage may end the business, because the customers did not just wait, they left. So the BIA does not rate processes by a single impact number. It rates impact over time, and the curve of that relationship, flat for a while and then steep, is what produces the targets. This is why the BIA must estimate both quantitative impact, lost revenue, regulatory fines, contractual penalties, and qualitative impact, reputation, customer confidence, employee morale, legal exposure. A process can be critical with a flat revenue curve because its failure breaks a regulatory promise or a customer’s trust, and a BIA that only counts money will rank it too low.

The BIA ends with a set of outputs, and those outputs are the currency the rest of the program spends: a ranked priority order for recovery, the resource requirements per process, and the time targets, the maximum tolerable downtime, the recovery time objective, and the recovery point objective, which are the subject of the next section. It produces no sites, no vendors, no budgets, no technology choices. Those come later, and they come from these numbers. A BIA that jumps straight to recommending a hot site has skipped its own job.

The time metrics: MTD, RTO, RPO, and the pieces inside the RTO

Three time metrics carry the exam’s continuity questions, and they are distinguished with lawyer-like precision because each one answers a different question. Get the three questions fixed in your head and the metric distinctions stop being a trap.

The maximum tolerable downtime (MTD) is the outer bound: the total time the organization’s leadership is willing to accept for a disruption of a process before the consequences become unacceptable. NIST SP 800-34 Revision 1 frames it as the amount of time an owner is willing to accept for an outage, and it is the bluntest of the three metrics, a management judgment about how much pain the organization can absorb. The business impact analysis surfaces the curve, and management draws the line. Nothing in the plan may exceed the MTD; it is the boundary the whole plan lives inside.

The recovery time objective (RTO) is the target: the maximum time that may elapse after a disruption before the process or system is restored to an acceptable level of operation. It is the number the engineers design to. If the BIA says the payment process cannot be down more than four hours total, the RTO is set inside that boundary, at two hours, say, to leave room for the things that always go wrong. The RTO is always set inside the MTD, strictly smaller, because the MTD is the full span management will tolerate and the RTO is a promise that must leave margin for error. It is the promised delivery time, and the MTD is the deadline the customer, management, will not tolerate missing.

The recovery point objective (RPO) is the data loss target: the maximum acceptable loss, measured in time, or equivalently the point in time to which data must be recovered. If the RPO is fifteen minutes, the recovery must put the data back to no earlier than fifteen minutes before the disruption. Every transaction older than that must exist; everything in the last fifteen minutes may be gone. The RPO is the answer to the question “how much history can we lose?”, and it is set by the business, not by the backup schedule. A business that cannot lose more than fifteen minutes of orders will demand an RPO of fifteen minutes, and the backup and replication design must deliver it.

These two numbers drive the two halves of the strategy. The RPO drives the data strategy: how often backups are taken, whether transaction logs stream offsite, whether data replicates in real time. The RTO drives the restoration design: how fast the environment can come back, which means the site, the procedures, and the people. The exam loves to swap them: a scenario states that the system must resume in two hours and may lose ten minutes of data, and the question asks which strategy fits. The answer follows from treating the two hours as the RTO and the ten minutes as the RPO, and never the other way around.

Two refinements complete the picture. Planners commonly decompose the RTO into the work processing time, the time to bring the alternate environment up, and the work recovery time, the time to validate the restored data and operations and resume normal service. The business does not care about the split, only about the sum, because the clock on the RTO runs from the moment of disruption and stops when the process actually works again. The second refinement is that the RTO and RPO are targets, not guarantees, which is exactly why testing exists: the only way to know a promised recovery time and data point are real is to attempt them and measure the gap.

Recovery strategy: data first, then systems, then people

With the targets set, the next layer is strategy: the recovery options matched to the targets and the money. This is where the discipline’s risk-management soul shows up, because every strategy is a purchase of time with capital. The BIA said what the time is worth. The strategy layer decides what the organization will pay to buy it back, and the exam wants you to make the same match: expensive strategies for short targets, economical strategies for long ones, and nothing that misses its own target.

Data strategy comes first, because data is the thing that, once lost, cannot be bought. Backup schemes are the foundation. A full backup copies everything. An incremental backup copies only what changed since the last backup of any kind, full or incremental. A differential backup copies only what changed since the last full backup. The tradeoffs run in the familiar direction: full backups are simple to restore from but expensive and slow to take; incrementals are quick to take and small, but restoring means laying down the full and then every incremental in order; differentials sit between, bigger than incrementals, and restore from just the last full plus the last differential. More recovery speed costs more storage and more backup load, and the RPO is the arbiter: the backup interval must be tight enough that the gap between backups never exceeds the RPO, and the restore path must be fast enough to meet the RTO.

Rotation and offsite storage complete the basic discipline. The grandfather-father-son scheme keeps daily, weekly, and monthly generations on a cycle, giving both recovery depth and a measure of protection against a corrupt or infected backup, since an older generation survives a new one’s failure. The widely used three-two-one practice, three copies of the data on two media types with one copy offsite, is not a formal standard but a sound rule of thumb, and the offsite copy is the point that actually matters for disaster recovery: a backup sitting in the same building as the data it protects is not a backup for the events that take the building. NIST SP 800-34 Revision 1 describes the escalation from offsite batch storage to electronic vaulting, the scheduled transfer of backup data to a remote site, and remote journaling, the near-real-time transfer of transaction logs, up to disk mirroring and replication, where the data never leaves a synchronized pair. Each step costs more and shrinks the RPO. Nothing in this ladder is inherently correct; the correct rung is the one whose RPO matches the target.

System strategy decides where the work runs while the primary site is gone, and the classic options form the price-and-speed axis the exam tests most. The cold site has the basic infrastructure, power, space, cooling, network, but no equipment, and takes days to weeks to make usable; it is the cheapest insurance that buys the longest recovery time. The warm site has some equipment and configuration in place, enough to cut the recovery time to hours or a day. The hot site is fully configured with systems, data, and communications ready to take over in hours; it is expensive, and it is the answer whenever the RTO is short. The mirrored site is the top of the ladder: full real-time replication, an RTO and RPO near zero, and the highest cost of all. A mobile site, a self-contained trailer or container that can be delivered and connected, rounds out the family, and is a favorite exam answer because candidates forget it exists.

Two external options deserve their own place. A mutual aid agreement, in which two or more organizations pledge to support each other in a disaster, is cheap and politically attractive, and it carries the structural weakness that the partner may be busy with its own disaster, or in the middle of its own workload, when the call comes; such agreements must be tested or they are promises. A service bureau, an outside firm that provides processing capacity under contract, moves the risk to a vendor and adds a procurement and contract management burden. Cloud services now function as alternate sites for many organizations, offering the elasticity to stand up capacity on demand, and they bring their own cautions: the provider’s outages are the customer’s outages, the contract must actually commit to the RTO and RPO the BIA demanded, data location and exit strategy need clauses, and a cloud failure is not automatically a failover event.

People strategy is the layer candidates skip, and the exam rarely skips it. Recovery plans are executed by people, and the plan must say where they work, how they connect, and who is authorized to decide what. Alternate work locations, remote access enablement, and delegated authority matter as much as any site. If the alternate site is fully configured but the crisis communication plan does not say how the team learns they are on duty, the configuration is theater. And key-person coverage, documented procedures that do not live in one person’s head, is the difference between a plan and a hostage situation.

Writing the plans that someone can actually execute

Strategy selects the means; the plan records the execution. A good continuity or recovery plan is two things at once: a decision record, showing who decided what and who authorized it, and an operating procedure, showing who does what, in what order, starting from the first minute. NIST SP 800-34 Revision 1 gives the contingency planning process seven steps, and the order is itself worth memorizing because it is the exam’s favorite sequencing question: develop the contingency planning policy; conduct the BIA; identify preventive controls; create recovery strategies; develop the contingency plan; plan testing, training, and exercises; and maintain the plan. The plan sits at step five, downstream of analysis and strategy, and upstream of the testing that keeps it honest.

A plan that someone can actually execute in a crisis contains the activation material before it contains any technology. First, the activation criteria and the declaration authority: the conditions that make an event a disaster for this plan, and the named role, not the anonymous committee, with the authority to declare it and invoke the plan. Second, notification: the contact rosters, the escalation sequence, and who calls whom, with multiple contact channels, because the office phone is down by definition. Third, the roles: the teams, the incident commander or crisis manager, the recovery teams, the communications lead, the liaison to vendors and regulators, each with named alternates. Fourth, the procedures: per process and per system, the steps to stand up the alternate environment, restore the data, validate it, and resume. Fifth, communication: internal and external messaging, who speaks, who approves. Sixth, the dependency notes: the third-party contracts, the vendors, the carriers, the utilities, and what each has promised in writing.

Two features separate a usable plan from a binder. The first is that it knows how it ends: reconstitution, the return to the primary site or the new normal, the validation that the primary environment is sound, the orderly migration back, and the deactivation of the emergency structure. Plans that stop at “we recovered” strand the organization in its alternate site with no path home, which is how recoveries that succeeded become operating budgets that never end. The second is that it names who decides. A plan that cannot be invoked by a named person with clear authority will be invoked by whoever is loudest at the moment of crisis, and that person may not be the one who signed the risk decisions. The disaster recovery plan must state the declaration authority and the criteria, and the organization must practice using them, because the cost of declaring early is cheap theater and the cost of declaring late is measured in the losses the plan existed to prevent.

Preventive controls deserve a paragraph of their own because they are the quiet step in the seven-step process. Redundant power, redundant network paths, surge protection, fire suppression, replication, and maintenance discipline all exist to make the plan unnecessary. NIST SP 800-34 Revision 1 places identifying preventive controls as step three, immediately after the BIA and before recovery strategy, because it is cheaper to keep a process running than to recover it. The exam rewards candidates who reach for prevention first in scenario questions, exactly as it did in the risk chapters: the best recovery is the one that never happens.

Plan maintenance is where plans actually die. Contact rosters rot in months. Organizations change systems, move facilities, change vendors, reorganize, and hire and lose the people whose names fill the roles. The discipline is to treat the plan as a living document: version-controlled, reviewed on a schedule, and updated whenever an exercise reveals a gap or a material change occurs. A plan’s age is a measure of its dishonesty. ISO 22301:2019, the international standard for business continuity management systems, builds this entire loop into its requirements: the business impact analysis and risk assessment at clause 8.2, the strategies and solutions at 8.3, the plans and procedures at 8.4, and the exercise programme at 8.5, all inside a management system that leadership reviews and audits. Its cousin, ISO/IEC 27001:2022, requires continuity discipline from the information security side: Annex A control 5.30 demands ICT readiness for business continuity, and 5.29 requires that security arrangements survive the disruption itself. That pairing matters. A recovery that brings the systems back without their security controls is not a recovery; it is a fresh incident running on restored infrastructure.

Testing: where the plan stops being fiction

Every continuity professional has a story about a plan that looked flawless on paper and collapsed in its first real exercise, usually in the first hour. The collapse is the point. A plan that has never been tested is fiction with a logo on it, and the exam wants you to know the ladder of tests, what each one proves, and the price each one charges, because the choice of test is a risk decision in miniature.

The checklist test is the cheapest rung: someone reads the plan against a checklist, verifying that sections exist, contact lists are current, and references resolve. It catches documentation gaps and nothing else, and it is better than nothing by exactly that margin. The tabletop exercise, also called a walkthrough, brings the team together to discuss a scenario, walking through the plan role by role with a facilitator; it costs no downtime, and it catches the coordination failures, the unclear authorities, the missing handoffs, that no amount of reading will surface. The simulation exercise, the functional exercise in the taxonomy of NIST SP 800-84, “Guide to Test, Training, and Exercise Programs for IT Plans and Capabilities,” moves participants into action: they execute parts of the plan, make decisions under scenario pressure, and sometimes perform real recovery steps, without a full cutover. This is where decision-making gets tested, and it is often the highest-value test per dollar spent.

The parallel test runs the alternate site in production-parallel: the workload is executed at both the primary and the recovery site, and the results are compared, without the primary being abandoned. It proves that the recovery environment actually processes real work, which is the hardest claim in the discipline to make true, and it risks no production availability; its price is the cost of operating both environments at once. At the top of the ladder sits the full interruption test: the organization actually shuts down the primary operation and runs from the recovery site, the only test that proves the whole claim, and the riskiest one, because if recovery fails, the failure is real. NIST SP 800-84 sets out the broader exercise program thinking behind all of these, with tabletop, functional, and full-scale exercises as the standard progression, and it is the reference to cite when a scenario asks what a given test can and cannot establish.

The practical guidance behind the ladder is annual rhythm plus change triggers. Test the plan at least annually, and after every material change: new systems, new sites, new vendors, new leadership, a merger, a migration. Test at multiple rungs, because the rungs prove different things: the tabletop keeps the people coordinated, the parallel test keeps the systems honest, and an occasional full interruption keeps everyone humble. And do the after-action review with the same seriousness as the test itself: capture what failed, what nearly failed, what surprised, fix the plan, and re-test the fix. A failed test is a cheap disaster bought at a discount. A failed recovery is the real thing at full price. The organizations that survive are not the ones with the best plans; they are the ones whose plans have failed in rehearsal often enough to be nearly true.

Training and testing are siblings, and the exam expects you to keep them straight. Training gives people the knowledge and skills to execute their roles: it builds the bench. Testing proves the plan: it measures the bench against the scenario. You train first, so that the test measures the plan and not the team’s ignorance, and you test repeatedly, so that the training does not decay into a binder on a shelf. The two together, plus maintenance, are the program, and the program is the thing that survives when every individual document has gone stale.

The continuity mindset the exam scores

Step back and the whole domain reduces to a few postures, and the exam questions are mostly tests of posture. First, order: BIA before strategy before plan before test, and policy, prevention, and maintenance woven around the sequence. Second, business first: the BIA starts from processes, not systems; the RTO and RPO are set by the business consequences, not by the backup software; and the plan is owned by the organization, not the IT department. Third, metrics as the vocabulary: MTD is the outer bound, RTO is the resumption target, RPO is the data loss target, and every strategy question is answered by matching the option to the numbers. Fourth, cost versus speed as the axis: hot and mirrored sites buy short recovery at high price, cold sites buy long recovery cheaply, and the correct answer is the one consistent with the stated targets and the stated budget. Fifth, prevention before recovery, always. Sixth, accountability: the plans name decision-makers, management signs the targets, and the declaration authority is a person.

Keep those postures and the scenario questions solve themselves. A question about a bank that cannot tolerate losing more than a few minutes of transactions points at a mirrored site or remote journaling, not a cold site and weekly tape rotation. A question about a small firm with a modest budget and a three-day RTO points at a warm or cold site, not a mirrored one. A question asking what the BIA produces does not get answered with “a recovery site”; it gets answered with priorities, resource requirements, and time targets. A question asking which test proves recovery capability answers “parallel” or “full interruption,” and a question asking which test risks production answers “full interruption.” The exam is not testing your knowledge of failover software. It is testing whether you can take a business statement, a budget, and a target, and make the plan-shaped decision a continuity manager would make.

Practice questions

  1. An organization is beginning its continuity program from scratch. Which sequence reflects correct practice?

    A. Select the alternate site, then write the plan, then run a tabletop exercise. B. Conduct the business impact analysis, then set recovery targets, then choose strategies, then write and test the plan. C. Write the disaster recovery plan first, then conduct the BIA to justify it. D. Purchase insurance, then draft the plan, then commission the BIA.

  2. A business analyst states that the order entry system must resume accepting orders within two hours of any disruption and may lose at most fifteen minutes of transactions. What are the recovery time objective and recovery point objective?

    A. RTO of fifteen minutes, RPO of two hours. B. RTO of two hours, RPO of fifteen minutes. C. RTO of two hours, RPO of two hours. D. RTO of fifteen minutes, RPO of fifteen minutes.

  3. The head of operations says the payroll process cannot be down for more than 24 hours, full stop, and that no recovery target may exceed that line. Which metric is she setting?

    A. The recovery point objective. B. The maximum tolerable downtime. C. The annualized loss expectancy. D. The recovery time objective.

  4. A trading operation cannot tolerate losing any material amount of transaction data and must resume processing within minutes of a disruption. Which recovery strategy best matches these targets?

    A. A cold site with weekly full backups. B. A warm site with daily differential backups. C. A mirrored site with real-time data replication. D. A mutual aid agreement with a peer firm.

  5. During the business impact analysis, the team documents that the warehouse shipping process depends on the inventory database, the carrier’s pickup schedule, and the dock staff, and that two other processes depend on shipping for their own recovery. What is the team documenting?

    A. The recovery time objectives of the processes. B. The resource requirements and dependencies of the process. C. The preventive controls currently in place. D. The list of threats in the organization’s threat model.

  6. A firm takes a full backup every Saturday night, and each weeknight backs up only files changed since the previous Saturday. After a failure on Thursday, what does restoration require?

    A. The Saturday full backup only. B. The Saturday full backup and the Monday, Tuesday, and Wednesday differentials. C. The Saturday full backup and every incremental backup taken since Saturday. D. The Thursday backup only.

  7. An administrator restores a database by loading the last full backup and then applying every backup taken after it, in chronological order. Which backup scheme is in use?

    A. Incremental backups. B. Differential backups. C. Grandfather-father-son rotation. D. Electronic vaulting.

  8. Management wants to prove the recovery environment can actually process real workloads without taking production down. Which test does this best?

    A. A checklist test. B. A tabletop exercise. C. A parallel test. D. A full interruption test.

  9. A facilitator presents a scenario and the team discusses, hour by hour, who would act, what they would do, and where the plan’s handoffs would occur. No systems are touched. What has the team conducted?

    A. A checklist test. B. A tabletop exercise. C. A parallel test. D. A full interruption test.

  10. A fire alarm sounds in the office tower. Which plan governs the first actions of the next fifteen minutes?

    A. The disaster recovery plan. B. The occupant emergency plan. C. The continuity of operations plan. D. The business continuity plan.

  11. A disaster recovery plan states that a disaster is declared when a disruption is expected to exceed four hours, and names the director of operations as the only person who may invoke it. What is the plan specifying?

    A. The recovery point objective of the affected systems. B. The activation criteria and declaration authority. C. The reconstitution procedures. D. The alternate site selection.

  12. After a full interruption test, the team finds that the alternate site’s restored database was missing two hours of transactions and that failover took twice the planned time. What is the correct next step?

    A. Abandon full interruption tests and rely on parallel tests, which carry no risk. B. Record the findings in the after-action review, fix the recovery procedures, and re-test the fix. C. Increase the RPO so the missing transactions are within tolerance. D. Declare the exercise a failure and remove the alternate site from the plan.

Answers and rationales

  1. B. The BIA precedes everything: it establishes priorities, resource requirements, and the time targets that strategy and planning must satisfy. Selecting a site before the BIA (option A) commits to a strategy before the requirements exist, and writing the plan before the analysis (option C) builds on guesses. Insurance is a transfer option that complements the program, but it does not replace the BIA as the first step (option D).

  2. B. The RTO is the resumption target, two hours here, and the RPO is the acceptable data loss, fifteen minutes here. Options A and D reverse or conflate the two, and option C ignores the stated data limit. The RPO is expressed as the maximum tolerable loss of history, not as a second resumption target.

  3. B. The outer boundary management will accept, beyond which consequences become unacceptable, is the maximum tolerable downtime. The RTO is a target set inside that boundary, and the RPO concerns data loss rather than downtime. ALE is a quantitative risk calculation from Chapter 4, unrelated to a time boundary.

  4. C. A mirrored site with real-time replication is the strategy whose RTO and RPO approach zero, which is what near-zero data loss and minute-scale resumption require. A cold site with weekly backups loses a week of data and takes days to come up (option A); a warm site with daily differentials loses up to a day and resumes in hours (option B); a mutual aid agreement is a promise, not a guarantee, and unlikely to meet minute-scale targets (option D).

  5. B. The BIA identifies resource requirements, the people, systems, data, and services a process needs, and dependencies, both the resources the process depends on and the processes that depend on it. Recovery priorities and targets are outputs, not this documentation; preventive controls and threat lists belong to other steps.

  6. B. A differential backs up changes since the last full backup, so after a Thursday failure the restore needs the Saturday full and the differentials taken since: Monday, Tuesday, and Wednesday. Option C describes the incremental scheme, which is not in use. Restoring the full alone (option A) loses three days of changes, and option D is meaningless on its own.

  7. A. Restoring from the last full plus every backup after it in order is the incremental restore path, because each incremental contains only the changes since the previous backup. A differential restore needs only the last full and the last differential. Grandfather-father-son is a rotation scheme, not a backup type, and electronic vaulting is a transfer method.

  8. C. A parallel test runs the workload at both sites and compares results, proving the recovery environment processes real work without abandoning production. The checklist and tabletop prove documentation and coordination, not processing capability (options A and B). The full interruption test also proves capability but it takes production down to do it (option D), which is more than the question requires.

  9. B. A discussion-based, no-impact walkthrough of the plan against a scenario is a tabletop exercise. The checklist is a document review, and the parallel and full interruption tests touch real systems.

  10. B. The occupant emergency plan governs life safety in the first minutes: evacuation, assembly, and head counts. Life safety precedes business resumption and system recovery, so the BCP, DRP, and COOP all wait until the immediate hazard is handled.

  11. B. Activation criteria, the conditions that make an event a declared disaster, and declaration authority, the named role empowered to invoke the plan, are the activation material every executable plan must carry. The RPO is a target metric, reconstitution is the return to normal operations, and site selection is strategy, not activation.

  12. B. The value of a test is the failure it reveals. The after-action review captures the findings, the plan is corrected, and the fix is re-tested, which is the maintenance loop the discipline requires. Abandoning the most realistic test (option A) hides the gap; inflating the RPO to fit the failure (option C) changes the business requirement to excuse a shortfall; and removing the alternate site (option D) eliminates the recovery option instead of repairing it.

Business continuity on one page

When the details blur, hold the shape. The discipline decides in advance: how long the organization can be down, how much data it can lose, who declares the disaster, and where the work happens. The analysis comes first. The business impact analysis ranks processes by recovery criticality, identifies their resource requirements and dependencies, and measures impact as a function of time. It produces the three metrics that carry every decision: MTD, the outer bound management will accept; RTO, the resumption target set inside it; and RPO, the data loss the business can tolerate, with the work processing and work recovery time inside the RTO.

Then the strategy, always matched to the targets: backups and replication to satisfy the RPO, full, incremental, and differential schemes with offsite storage and rotation; sites to satisfy the RTO, cold, warm, hot, mirrored, and mobile, with mutual aid, service bureaus, and cloud as alternates; and people, facilities, and communications to make it executable. Then the plans, which record who decides, who acts, and how the operation returns home, and then the program: checklist, tabletop, simulation, parallel, and full interruption tests, at least annually and after every material change, each test proving a larger claim at a larger price, with training first, after-action review always, and maintenance forever. Prevention sits ahead of all of it, because the best recovery is the one that never happens.

Chapter 6 takes the same decision discipline and turns it to the human layer of the program: the policy hierarchy, personnel security, and the training and awareness that make every other control land. Before you go, hold the one sentence that organizes this chapter and the exam’s treatment of it: the plan is not the deliverable, the capability is, and the capability is built by analysis, paid for by strategy, written into plans, and proven by tests.