Skip to content

Project Management Mastery / Chapter 23

Plan for Resilience and Alternative Futures

Resilience is not prediction; it is what holds when the future refuses to hold still. On day fifteen of Northstar's flood response, six small failures in one day — a broken drive shaft, a fuel reserve thirty percent below plan, a satellite link down six hours, a fuel authority with no backup — make the argument in one scene: the response that cannot name what it will do when its assumptions fail is running on luck. The chapter builds the four capabilities of resilience, the scenario set and the pre-mortem, the dependency map and its single points of failure, the five moves that change the shape of failure, the playbook that opens with a trigger and pre-delegates its authority, and the rehearsal that proves the design — plus the honest price, because resilience costs, and the discipline dies at the first budget cut if the trade was never decided.

Chapter 23: Plan for Resilience and Alternative Futures

The convoy was late and nobody had failed

Day fifteen of the flood response, and the incident log reads like a report of small things. Convoy 7 left the hub four hours after its slot because the fuel truck that should have met it at the midpoint broke a drive shaft on the northern road. The reserve at the midpoint depot was 30 percent below the level the plan assumed, because two top-ups in the past week had been logged as deferred. The satellite terminal at the hub lost its link for six hours in the afternoon, and the field team reached the convoy through a relay on a health post’s radio. And the one person with authority to approve a fuel purchase above 50,000 units was asleep through four of those six hours, because nobody had named a backup for that decision.

Everything recovered. The convoy arrived six hours late instead of never. The fuel held, narrowly. The radio relay worked. The response moved 41 tons that day against a plan of 44, closer to the target than it had been on the third day, when the gap was 16 tons. Nobody in the room had failed. That is precisely what Daniel Okafor, the relief response lead, says when he calls the meeting for the following morning.

“A near miss is the system telling us what it cannot absorb,” he says. “We got lucky six times in one day. I want to know how much luck we have left.”

The room is the same one from the third day, when the forty-truck offer forced the tailoring decision: Priya Raman on procurement, Farida Aziz on compliance and safeguarding, Mateo Herrera on field logistics, Ines Carvalho on the donor relationship. The fast-tracked northern route is now the spine of the response, carrying most of the tonnage. It is also unverified after the flood, served by carriers with thin safety records, and protected by scouts and checkpoints that Mateo’s team improvised rather than planned. The register carries the row: because the northern route is unverified and the carrier’s record is thin, a convoy may be delayed or fail mid-route, and Mateo owns the response. The response also depends on one satellite link, one fuel supplier, one generator in the cold-chain warehouse that has not been load-tested since the flood, and one person holding the fuel authority.

Mateo makes the case first. “The eastern route through the highlands is passable. I have scouted it twice. It is slower and rougher, and the last crossing will not take the big trucks at night, but it carries medicine and cold chain. If the northern road closes, we do not have a fallback, we have a reroute that has never been driven with a load.”

Ines asks the question the donor will ask. “What does it cost, and what are we giving up to run two routes?”

“The scouting is done. The signing of two more carriers is the same vetting we did on the fast track, with the same red-flag controls,” Priya says. “The cost is fuel and maintenance on a road that will not be the cheapest ton. The risk is that we spend capacity keeping a backup warm that we might never use.”

Farida adds the floor. “Whatever we sign, the safeguards hold. Unverified carriers do not carry medicines alone, and every convoy keeps the scout and the checkpoint discipline. That part is not negotiable.”

Daniel does not decide on the spot. He asks each of them to come back the next morning with the response’s dependencies written down: everything the promise depends on, and what happens if it fails. He wants the map before he spends the money. This chapter is what that map is for.

Prediction is a tool; resilience is a property

The risk register is where the project names what it can name: the cause, the event, the effect, the trigger, the owner. It is a decision instrument, and chapter 22 built it as one. But it has a boundary, and the boundary is the point of this chapter. The register can only carry the futures someone in the room can imagine, and some of the most damaging failures are not the ones the room fails to rate. They are the ones the room cannot see at all: not a risk, but a combination; not a scenario, but a condition the system was never designed to meet.

The sociologist Charles Perrow made this argument about high-technology systems in 1984, in a book that gave the idea its name, Normal Accidents. His finding: when a system is interactively complex — parts connect in ways the designers did not fully map — and tightly coupled, a failure in one part propagates to the next before anyone can intervene. The accident is then not a deviation from the design; it is a property of it. The system generates failures its own designers could not have imagined, and the imagination of any single person or workshop is the wrong instrument for the problem. The register is a telescope; the accident is the shape of the whole observatory.

This is not a counsel of despair. It is a shift in what the project optimizes. Prediction asks: what will happen, and what is the probability? Resilience asks: when our assumptions fail, how much do we lose, how fast do we recover, and what keeps working? Prediction builds a better model of the future. Resilience builds a system that holds when the model is wrong. They are complementary, and the complement is not optional: the more the project depends on assumptions that cannot be verified in advance, the more of its safety margin must live in recovery rather than in forecast.

The resilience engineering tradition, in the work of Erik Hollnagel, David Woods, and Nancy Leveson, gives the practitioner four capabilities to build. I describe them in my own words because they are the chapter’s spine. Respond: the system knows what to do when something happens. Monitor: the system knows what to look for, and looks. Anticipate: the system knows what to expect, and notices when expectation diverges. Learn: the system knows what has happened, and changes because of it. The near miss at Northstar tests all four at once. The response had a workaround, so it could respond, barely. It had no instrument watching the fuel reserve, so it could not monitor its own margin. It had not anticipated the combination, so the first signal arrived as a six-hour delay. And whether it learns is the question the room is now trying to answer.

The field signal that a project is running on prediction alone is easy to read once you look: the plan is a single path through a single set of assumptions; the schedule is a chain rather than a network, so one broken link stops the whole promise; the team can name its own single points of failure out loud and has never written them down; the near misses are logged and nothing changes. Any one of these is a sentence in an incident log. All four together are a design.

Four futures you can argue about

The first instrument of the resilience map is not a register row, it is a set of futures. Scenario planning, the practice Royal Dutch Shell made famous in the 1970s and Pierre Wack described publicly in the Harvard Business Review in 1985, does not try to forecast which future will happen. It builds a small set of coherent, alternative futures, each plausible enough to argue about, and then asks what each one would force the organization to do. Wack’s claim was that the value of a scenario is not the probability attached to it; it is the way it changes the mental models of the people who must decide. The decision maker who has lived inside four futures in a workshop has already rehearsed being wrong.

The minimum viable scenario set for a project is small: three or four futures, built on the two or three uncertainties that matter most. Not the dozens of risks in the register, the few forces that would change the shape of the whole undertaking. At Northstar, the forces are the condition of the roads, the condition of the funding, and the behavior of the environment, whether the rivers rise again. From those, the room builds four worlds. In the first, the single disruption: the northern route closes for a week and everything else holds, the world where the convoy’s bad day becomes a bad month. In the second, the creeping degradation: the roads worsen slowly, fuel prices double, the donor adds reporting, and the response bleeds capacity one percent at a time. In the third, the second disaster: a new flood on the southern catchment takes out the verified alternative and the eastern route must carry what the whole region needs. In the fourth, everything holds, the world where the current plan is right.

Each world needs three things before it is useful. An indicator: what would tell us we are in this world now, the river gauge at the crossing, the convoy delay trend, the fuel price at the depot, the donor’s latest reporting request. A decision: what this world would force us to do, close the northern route and shift to the eastern, buy fuel forward, pre-position stock at the midpoint. And an owner: the person who watches the indicator and would raise the decision. The scenario set is not a forecast and not a wish. It is a set of rehearsed choices, and the rehearsal is the point, because the decision maker who has already decided, in a workshop, what the second-disaster world means has removed the delay that real disasters punish.

The failure mode of scenario work is subtle, and it is the reason many scenario exercises die in binders. The first failure is the futures that are all the same future wearing different labels, four versions of a single optimistic assumption. The test is brutal: would the decision differ if this world arrived? If the answer is no, the world is decoration. The second failure is the missing indicator, a future nobody can recognize arriving. The third is the scenario that stops at the story and never reaches the decision, the narrative that makes the room feel wise and changes nothing. The discipline that prevents all three is the same: every scenario ends with a column headed “then we would,” and every row of that column has an owner and a trigger.

The pre-mortem and the stress test

Scenarios open the imagination forward. The pre-mortem opens it backward. The cognitive psychologist Gary Klein described the practice in the Harvard Business Review in 2007, and his design is worth keeping because the details are the method. You gather the people who know the work, and you tell them: it is a year from now, and the project has failed. It is not a possibility, it is a fact, the launch missed, the corridor did not open, the relief ran out, the merchant promise broke. Write, each of you, alone and in silence, the story of how it happened. Then share, cluster, and read the list.

The silence is not ceremony. In an open discussion the first speaker anchors the room, and the room converges on the loudest story, which is usually the most familiar one. The silent pre-mortem collects the stories the quiet people hold, and in a pre-mortem the quiet people are often the ones closest to the mechanism. The second detail is the backward stance. Imagining that the failure has already happened releases the room from the politeness of uncertainty, nobody has to argue a probability, they simply have to write a plausible obituary. The output is not a risk register. It is a set of mechanisms and the decisions they imply, and the chapter’s drill will run one later.

The stress test is the third instrument, and it is the one most often confused with prediction. A stress test does not ask what will happen. It pushes a parameter, or a combination of parameters, far beyond expectation, and observes what breaks. The purpose is not to believe the parameter; it is to learn where the structure has its breaking point. The financial system gave the practice its canonical form after the 2007 and 2008 crisis: the Dodd-Frank Wall Street Reform and Consumer Protection Act of 2010 made annual stress tests mandatory for the largest U.S. banks, and the Federal Reserve has run its Comprehensive Capital Analysis and Review since 2011. The regulators do not believe the severe scenario will occur. They push the balance sheet against it and read where capital breaks, because a bank that knows its breaking point in a simulation is a bank that has priced its own fragility. The project version is identical in spirit: take the plan, double the volume, halve the funding, close the road, lose the person, and watch the model, not for the number, but for the first promise that goes red.

The single-risk habit is the enemy of all three instruments, and it fails in a specific way. The risk register tests one thing at a time: the supplier fails, or the link fails, or the person leaves. Real failure is combinatorial: the supplier fails because the road closed because the flood came because the funding ran out. The two “independent” routes that both flood in the same storm are not independent, they are one correlated failure wearing two names. This is why the stress test that matters for resilience is the simultaneous one, and it is the chapter’s mastery drill: not the supplier failing, but the supplier, the technology, and the leadership failing in the same week. The point of the simultaneous assumption is not to be realistic about all three at once. It is to find the promise that breaks first, because that promise is the system’s true constraint, and the constraint is the design problem.

The map that ties the instruments together is a dependency network: the nodes the promise depends on, the links between them, and the annotation of what fails, what it takes down with it, how fast, and what can substitute. The map is this chapter’s primary visual, and it earns the description because it changes the questions the project asks. The register asks: what could go wrong with this node? The map asks: if this node goes, what else goes with it? The difference is the difference between a list and a circuit.

Two concepts structure the map. A single point of failure is a node whose loss takes down the outcome, no matter how healthy everything else is. Concentration is the quieter cousin: no single node is indispensable, but so much of the promise runs through one place, one supplier, one route, one team, that the place has become a de facto single point, protected only by the assumption that it will hold. The map reveals both. At Northstar, the northern route carries most of the tonnage, one carrier group moves the medicines, one satellite link reaches the field, one person holds the fuel authority, and one generator keeps the cold chain honest. Any one of these is a node. The response is a series circuit with five elements, and a series circuit is as strong as its weakest element.

Documented cases teach the cost of forgetting this. In March 2021, the container ship Ever Given ran aground in the Suez Canal and blocked it for six days, and the world discovered how much of its trade ran through one waterway and one strait; shipping rates and delivery delays rippled through global supply chains for months. In 2011, severe flooding in Thailand disrupted the factories that produced a large share of the world’s hard disk drives, and the global industry felt the shortage for the better part of a year. Neither event was a single point of failure in the strict sense; each was concentration, so much of the world’s capacity in one geography that a regional event became a global one. The lesson the map teaches is the same at every scale: the more promise runs through one place, the more the project should know exactly how much, because that number is the size of the bet it has made without noticing.

Building the map is a workshop, and the raw material already exists in the project’s own documents. The decomposition of chapter 14 gives the nodes. The schedule logic of chapter 17 gives the links, the critical path and the near-critical paths. The supplier map and the exit plans of chapter 20 give the external nodes. The knowledge map of chapter 19 gives the human ones, the person who holds the banking interface, the clinician who knows the escalation workflow, the one operator who can configure the cold chain. The overlay that makes it a resilience map rather than a diagram is the promise, from chapter 2: trace each success criterion backwards to the nodes it depends on, and any node that appears on the critical path of the same promise twice is a single point by definition. The map is then annotated with the interventions, and the interventions are the next section.

The failure mode of the map is that it is drawn and adored. The map that shows the technical network and not the human one is missing the most common failure of all, the knowledge, the authority, the password, that lives in one head. The map that is never updated is a museum. The test of the map is the same as the test of the scenario set: does the room act differently because it exists, and does the map change when the plan changes?

Five moves that change the shape of failure

Once the map shows where the failure propagates, the project has five moves available, and each is a design choice with conditions, not a checkbox. The moves are buffers, redundancy, substitutability, modularity, and graceful degradation. The temptation is to treat them as interchangeable, more of all of them, and the discipline is to treat each as a named answer to a named failure on the map.

A buffer is stock set aside against variability: time in the schedule, cash in the budget, fuel in the depot, capacity in the team. The buffer has value only if it has a policy, and the policy has four clauses: how much, what consumes it, who may consume it, and how it is replenished. At Northstar the fuel reserve is the working example. The plan assumes three days of consumption in the depot. If consumption runs at one and a half times plan, the reserve empties in two days, not three. The replenishment lead from the supplier is one day, and at one and a half times plan that lead consumes one and a half days of stock. The trigger, replenish when the reserve falls below two and a half days of stock, therefore leaves one day of margin, and that day is the difference between a hold and a stop. The policy makes it explicit: below the trigger, the reserve is a decision, not an asset, and the decision belongs to Mateo, who owns the route. The failure mode of the buffer is the comfort blanket: the reserve that is hoarded, never consumed even when the policy allows it, or consumed silently, in the small top-ups logged as deferred that drifted the midpoint reserve 30 percent below plan. A buffer consumed without a decision is not a buffer; it is early failure wearing a disguise, and the incident log is where the disguise shows.

Redundancy duplicates the capacity: a second generator, a second link, a second carrier for the same job. The condition that makes redundancy real is that the standby has been tested, and the first lesson of every rehearsal is that the standby does not stand up. A redundant unit that has never been exercised is cold redundancy, and cold redundancy fails on the day it is called, because the failure that called it is usually the failure that its configuration never met. The warm version is exercised on a cadence, rotated into real use, and its performance is logged. Redundancy also carries a second cost that the map reveals: the split brain. Two systems that can both do the job can also disagree about which one is doing it, two generators that both think they are primary, two records that both think they are current. The discipline that contains the split brain is the same as the discipline of chapter 39’s configuration control: one of them is the authority at any moment, and the handover is rehearsed, not improvised.

Substitutability is the move most projects already half-own without naming it, and naming it changes what they decide. Redundancy duplicates; substitutability diversifies. A different route, a different carrier, a different communications channel, a generalist who can hold the fuel authority when the specialist is asleep. The substitute does not match the primary’s performance, the eastern route is slower and rougher, the radio relay is thinner than the satellite link, and the design question is not which is better, it is what the project gives up for the margin. Substitution trades efficiency for resilience, and the trade must be decided and recorded, because the efficiency loss is permanent while the resilience gain is invisible until it is needed. The arithmetic of substitution is the chapter’s numbers that matter, and it is worth doing once. If each route passes independently on 9 of 10 days, a single route delivers on 9 of 10 days, 90 percent. A journey that needs two such segments in series delivers on 9 of 10 squared, about 81 percent, because both must work. Two substitutable routes, either of which carries the load, fail only when both fail: 10 percent squared, so the system delivers on 99 percent of days. Series multiplies failure; substitution multiplies reliability. And the assumption is in the word independent, because a storm that floods both routes makes the two substitutable routes one correlated failure, which is precisely the failure the simultaneous stress test exists to find. The arithmetic is honest only when the correlation is questioned, and questioning it is part of the discipline, not a footnote.

Modularity is the move that keeps failure local. The idea has a proper genealogy: Herbert Simon, in a 1962 paper on the architecture of complexity, observed that complex systems in nature and in human design are typically hierarchic and nearly decomposable, built of components that interact strongly inside and weakly across boundaries, and that this structure is what lets a complex system evolve, be repaired, and survive damage. The project version: interfaces that are owned and small, batches that are separable, teams that can be isolated, delivery segments that can proceed when a neighbor stops. BlueLine’s phased segment opening is modularity in infrastructure: the central and northern segments can open while the eastern segment is delayed, so the corridor delivers value without waiting for its weakest link. KijaniPay’s canary rollback, the small cohort that can be withdrawn without taking down the whole launch, is modularity in software: the module that can be rolled back is a module that can fail without becoming a catastrophe. Modularity has a cost: the interface contracts must be designed and maintained, which is chapter 34’s discipline at the multi-team scale. The failure mode is the module sealed so tightly it cannot integrate — the walled garden whose failure is local only because its value is local too.

Graceful degradation is the last move, and it is the one most projects never write down, which is why they discover its absence at the worst moment. When failure is unavoidable, something must give, and the question is what gives first, in what order, and what keeps working until the end. The degradation order is a values decision before it is a technical one. The guardrails of chapter 2 decide it: at KijaniPay, the settlement promise is the near-zero-appetite dimension, so the degradation order protects the promise to the last, and the first thing to degrade is the nice-to-have reporting, not the settlement. At Northstar, the cold chain and the insulin never degrade, the fuel reserve and the convoy schedule do. The order is written, owned, rehearsed, and reviewed. The rehearsal is not optional: a degradation order that nobody has practiced will be improvised under panic, and improvisation under panic selects for whoever is loudest in the room, not for the promise. The failure mode is the degradation theater: the plan that says fail safe and the authority that has not been delegated, so that the moment the order must be executed, the person who must say stop to one promise to save another is the person who has no standing to say it. The pre-delegation is the subject of the playbook.

The playbook that opens with a trigger

A contingency plan is a response that has been decided before the condition arrives. The distinction that matters is between the contingency plan and the improvisation, and the test is the trigger. A contingency plan opens with a trigger: an observable, objective condition, defined in advance, that activates the response. Improvisation is what happens when the condition arrives without a trigger, because nobody wrote one. Crisis management, chapter 42’s subject, is what happens when the improvisation must be coordinated under pressure. The project that does contingency work well is the project that keeps its improvisation for the futures it could not name and pre-decides the futures it could.

The contingency playbook is the minimum viable form, and the minimum is one page per named scenario, with five elements. The trigger, stated so that a monitor can see it: the river gauge at the crossing above a mark, the convoy delay trend above six hours, the fuel reserve below its two-and-a-half-day line. The evidence that confirms the trigger, because the first report is a hypothesis: the gauge reading logged, the delay confirmed by two checkpoints, the depot stock reconciled. The owner and the pre-delegated authority: who decides, and up to what threshold without calling anyone. The actions in order, written so that the first action is the safest and the cheapest. And the exit: what must be true to stand the response down, because a contingency that never stands down becomes a new operating mode nobody approved. At Northstar, the eastern route contingency reads, in its one page: when the northern route closes or the delay trend exceeds six hours, Mateo shifts the medicine and cold-chain schedule to the eastern route, using the two carriers signed under the fast-track vetting, and notifies Daniel and Ines; the first convoy carries only the load the route has proven, the medicines, not the bulk; and the response stands down when the northern route carries three consecutive convoys at the checkpoint times. The one page is the deliverable, and the binder is the failure mode, with a reliable signal: nobody in the room can tell you, without opening it, what the first action is when the trigger fires. A playbook that cannot be recalled from memory has not been rehearsed, and a playbook that has not been rehearsed is a document.

The trigger dashboard is the instrument that keeps the playbook honest. It is a short list of indicators, each with an owner, a threshold, and the trigger it maps to: the fuel reserve and the two-and-a-half-day line, the northern route delay trend and the six-hour line, the satellite link uptime and the six-hour line, the cold-chain stock and the level below which the generator becomes a decision, the donor’s reporting requests and the point at which Ines raises the creep scenario. The dashboard is watched on a cadence, and the cadence is the same rhythm as the risk burndown and the schedule review, not a monthly ritual: at Northstar it is daily, at the convoy report, because the fuel reserve moves daily. The dashboard’s discipline is the discipline chapter 22 taught about thresholds: the line that moves when the calendar presses is not a line, it is a mood. The two-and-a-half-day fuel trigger does not become a one-day trigger because the convoy is urgent; it becomes a decision, escalated to Daniel, with the ethical decision record of chapter 4 as the instrument that keeps the escalation honest.

The authorization ladder is the playbook’s spine, and pre-delegating it is the work most projects skip because it feels like paperwork. The ladder states the thresholds in advance: Mateo may commit up to 50,000 units to keep the response moving, fuel, repairs, the small charters. Daniel may commit up to 500,000. Above that, the decision goes to the donor through Ines, with the fact base and the recommendation attached. The ladder exists because the near miss named its own cost: four hours of a crisis passed while the one person with authority slept, not because anyone was careless, but because the authority had never been distributed. The pre-delegation is recorded, communicated to the people who must act on it, and rehearsed, and it carries the chapter 4 discipline with it: the person who exercises delegated authority in an emergency writes the ethical decision record afterward, the context, the options, who decided, what evidence existed, what was preserved, what would have changed the decision. The ladder is the organizational form of graceful degradation: when the leader is unavailable, the authority degrades to the next rung, and the response continues without a moment’s improvisation about who may act.

Roles you name before the crisis

Every plan that fails in the first hour fails for a reason that is visible in advance: nobody had named the roles. The crisis roles are few, and naming them is a half-hour of work that saves the first hours of every crisis. The decision maker, the single accountable person for the response window, Daniel at Northstar, with the alternates named, Priya if Daniel is unreachable, and the authority ladder as the rule for what each rung may do. The logistics lead, Mateo, who owns the operational response. The communications lead, the one authorized voice to the donor, the authorities, the field, and the public, Ines in her role, and the discipline that in a crisis only that voice speaks outward, because the moment two voices speak, the world hears a contradiction that becomes a fact. The records keeper, who owns the fact base, the incident log, the decision log, the evidence that the review after the crisis will need. And the safeguarding floor, Farida, who holds the obligations that do not degrade, the dignity floor of chapter 4, and who can stop the response’s own momentum when it would trade a person’s safety for a tonnage number.

The information discipline is the second half of the same work, and it has three rules that fit on one page. One channel for facts: the incident log, with every entry timestamped, sourced, and labeled with confidence, so that the fact base of chapter 42 exists before the crisis instead of being improvised during it. One channel for decisions: the decision log, with the who, the what, the authority exercised, and the record to return to. One voice outward: the communications lead, and the principle that in a crisis the first report is a hypothesis, not a truth, and is labeled as one. The near miss showed the cost of the missing discipline: by evening, three versions of the convoy’s delay were circulating, the mechanical failure, the fuel failure, the coordination failure, and the incident log reconciled them only because Mateo’s team had written down what they knew as they knew it. In a real crisis, the versions do not reconcile; they harden, and the plan that did not name its channels in advance will spend its first hours discovering who speaks.

Near misses are the cheapest lesson the project will ever receive, and the price of the lesson is the discipline of taking it. The aviation industry built the model: since 1976, NASA has operated the Aviation Safety Reporting System for the Federal Aviation Administration, collecting confidential reports of incidents and near misses from pilots and controllers. Its founding lesson is that reporting flourishes when it is protected from punishment. The reports are about the system, not the person, and the person who reports is treated as the system’s sensor, not its scapegoat. The project version is the near-miss register, and the discipline is the question that turns a near miss into design change: why not worse? The convoy arrived six hours late; why not never? Because the radio relay existed, because the checkpoint discipline held, because the fuel held by a margin nobody had measured. Each answer is a working control, and the register that records the near miss without asking why not worse is a graveyard. The test of the near-miss system is the same as the test of the resilience map: does the design change because the near miss was recorded, and is the change rehearsed? At Northstar, the near miss changes the design: the fuel reserve gets an owner and a trigger, the eastern route gets a signed carrier, the fuel authority gets a named alternate. The near miss earned its keep when the design changed, and it would have been free either way; the failure would be the near miss that is logged in the register and lives there forever, the incident that becomes a statistic instead of a repair.

The rehearsal that fails on purpose

The last instrument is the one that proves the others, and it is the one most projects skip because it feels expensive. A rehearsal is the plan executed on purpose, in a controlled setting, with the explicit expectation that it will fail, because the failure of the rehearsal is the point. The tabletop is the cheapest form: the room, the playbook, a facilitator, and a series of injects, the conditions arriving one at a time, the gauge rising, the link dropping, the donor calling, and the team deciding what it would do, in real time, with the real triggers and the real authority ladder. The simulation adds the environment, the field, the radios, the stock, the clock, and the cost. The live exercise is the most expensive and the most honest, and the discipline is the same at every level: the rehearsal is not a demonstration, it is a test, and the test is designed to fail somewhere, because the somewhere it fails is the lesson. The first rehearsal of the eastern route will find the crossing that will not take the big truck at night, and the find is the value of the rehearsal, because the alternative is the crossing that refuses at night, in the flood, with the medicines on board.

The recovery targets are the numbers the rehearsal proves, and they come from the continuity vocabulary. The recovery time objective, the RTO, is how long the function can be down before the harm becomes unacceptable: the cold chain cannot be down six hours, so its RTO is measured in hours, and the generator is the instrument that meets it. The recovery point objective, the RPO, is how much state or data may be lost when the function recovers: the registration data cannot lose a day, so the RPO is the last synchronized backup, and the record-keeping discipline exists to meet it. The vocabulary comes from business continuity practice, the discipline standardized in ISO 22301:2019, the international standard for business continuity management systems. It is worth adopting at project scale, because it converts the vague ambition — we must be able to recover — into a testable pair: the function returns within this time, and this much is lost. The rehearsal proves the pair, and the proof is the number the project can defend to the donor, the sponsor, and the regulator.

Learning after near misses completes the loop that the four capabilities opened. The system responds, and the response becomes the monitored indicator, and the indicator becomes the anticipation, and the anticipation becomes the learning that changes the design. The loop closes only when the change is rehearsed, and the review cadence is the clock that keeps the loop turning: the scenario set revisited when an assumption changes, the pre-mortem run at the major milestones, the stress test re-run when the plan changes shape, the near-miss register reviewed on the same rhythm as the risk burndown. The cadence is not a calendar artifact; it is the answer to the question the opening scene asked. Daniel asked how much luck the response had left. The answer is the design: the map, the scenarios, the playbook, the ladder, the rehearsal, and the loop. Luck is what the design buys when it is not needed, and the measure of the design is not how much luck remains, but how little the response needs.

Resilience has a price and a voice

Every move in this chapter costs something. The buffer ties up fuel and cash. The redundancy doubles the equipment and the maintenance. The substitution runs a slower road at a higher cost per ton. The modularity demands interface contracts that must be designed and maintained. The rehearsal takes the team out of production and spends its attention. The cost is real. The project that pretends otherwise will quietly drop the discipline at the first budget cut, and the quiet drop is the resilience plan’s most common death. The honest version is the trade decided in advance, against the risk appetite of chapter 22: the appetite statement says how much risk the project will retain per value dimension, and the resilience budget is the spending that converts retained risk into absorbed shock. The question is not whether resilience is worth it; it is which failures the project will absorb, which it will accept, and which it will design out, and the answer is a decision with an owner, not a sentiment.

Resilience also has diminishing returns, and the map is the instrument that shows where the returns flatten. The project that adds a second route, a second carrier, and a second link is buying the margin that protects the promise. The project that adds a fourth route and a sixth carrier is buying insurance against a future so specific that the premium has outrun the exposure. The discipline of the map is the discipline of the stress test: push until it breaks, find the first promise that goes red, and buy resilience at the point where the promise breaks first, not uniformly across the network. The uniform resilience program, every node gets a redundancy, every risk gets a buffer, is the resilience theater of the project world: it feels comprehensive, it is expensive, and it is aimed at a map that was never drawn.

The equity question is the one the techniques do not answer. The degradation order is a values decision, and the values belong to more people than the room that wrote them. At Northstar, the order that protects the insulin and the cold chain before the bulk tonnage is not a technical choice; it is a promise to the people in the shelters and the health posts, and those people have no seat at the table where the order was written. At BlueLine, the resilience that keeps the corridor open matters differently to the commuters, the traders, the residents, and the disability advocates of chapter 9. An order that protects aggregate travel time before local access is a choice about who bears the risk — made by people who do not bear it themselves. The stewardship discipline of chapter 4 applies to resilience as it applies to compliance: the project is the temporary custodian of other people’s reliance, and the people who depend on the system should be visible in the room when the degradation order is written. The practical form is the review question: for each line in the degradation order, who depends on this staying up, and were they consulted about the order in which it goes down?

The delivery approach changes which of the moves carries the weight, and the differences matter enough to name. A predictive project, BlueLine’s corridor, holds its resilience in the schedule and the resources: buffers at the seams between segments, redundancy in the critical equipment, modularity in the staged opening that lets the corridor deliver value while its weakest segment is still troubled, and a degradation order that names which segment opens first when everything cannot open. An adaptive project, KijaniPay’s platform, holds its resilience in the architecture: modularity in the services, the canary rollback, the degradation order written from the guardrails, and the trigger dashboard that watches the promise in production rather than in a plan. A crisis project, Northstar’s response, compresses everything: the scenario set is four worlds on one page, the playbook is one page per named condition, the rehearsal is the trial convoy rather than the full exercise, and the authority ladder is the difference between a response that continues and a response that stops to ask. The hybrid, Meridian’s clinics and platform, holds its resilience across both registers: the construction buffers in the schedule, the clinical workflow degradable by design, the digital rollout staged so a platform incident does not close a clinic. The constant across all four is not the method; it is the question. What fails first, what keeps working, who decides, and what did the last near miss teach?

The machine has a bounded place in this work, and the boundary is the same one the whole book draws. A language model can draft the scenario narratives from the project’s assumption inventory and risk register, generating the four futures and their indicators for the room to argue with; it can propose stress-test parameter sets, doubling the volume and halving the funding and closing the road, and it can cluster the near-miss reports into the patterns a tired team will miss. Each of these is a draft, a hypothesis, a provocation, and none of them is evidence or decision. The scenario choice, the trigger thresholds, the degradation order, the authority ladder, and the signature on the playbook stay human, because they encode values and accountability that no draft can carry. And the data boundary is absolute: the humanitarian registration data, the patient records, the merchant transactions, the regulator filings, none of it enters an unapproved system, whatever the tool promises. The machine that drafts the future is useful; the machine that decides the degradation order is not a machine, it is an abdication.

The map, the playbook, the rehearsal, and the luck

The discipline of this chapter reduces to a question the project can ask at any moment, on any plan, at any scale: if the future fails my assumptions, in the worst combination I can name, what is the first promise to break, what keeps working, what is the first action, and who has the authority to take it? The map answers the first two, the playbook and the ladder answer the last two, and the rehearsal proves all four. The scenario set, the pre-mortem, and the stress test are the instruments that find the questions; the near-miss register and the why-not-worse review are the loop that keeps the answers current. The register of chapter 22 names what the project can name; the resilience design absorbs what it cannot.

At Northstar, the morning after the near miss, the room has the map, and the decisions follow the map. The eastern route gets its signed carriers and its trial convoy, carrying medicines and cold chain, the load the route has proven. The fuel reserve gets its owner, its trigger, and its replenishment policy, and the midpoint depot’s next top-up is not deferred. The satellite link gets its alternates named, the radio relay tested, and the escalation for a drop beyond the six-hour line. The fuel authority gets its ladder, Mateo below 50,000 units, Daniel below 500,000, the donor call above, and the ethical decision record as the instrument that keeps the ladder honest. The trigger dashboard goes on the convoy report, read daily, and the eastern route trial becomes the rehearsal that will fail somewhere useful. None of it costs more than the fast-track contract saved, and all of it answers the question the near miss asked. The luck the response had left is now a design.

Practice

One. A scenario field drill: four futures for your own project. Take the work you lead or know best. (a) Name the two or three uncertainties that would change the shape of the whole undertaking, not the routine risks, the forces. (b) Build four worlds from them: the single disruption, the creeping degradation, the second shock, and the world where the current plan holds. (c) For each world, write the indicator that would tell you it had arrived, the decision it would force, and the owner who would raise it. (d) Apply the one-question test to every world: would the project do anything differently if this world arrived, and if the answer is no, why is it in the set?

The drill passes when every world has an observable indicator, a concrete decision, and an owner, and when at least two of the worlds would force different decisions. The most common failure is the four futures that are the same future wearing different labels, four versions of the optimistic assumption, and the repair is the one-question test. The second failure is the world with no indicator, a story nobody can recognize arriving, and the repair is the monitor: the gauge, the trend, the price, the request, something on a dashboard. Credit belongs to any set where the worlds disagree with each other, because disagreement is what rehearses the judgment.

Two. A numbers drill: the arithmetic of substitution. Reproduce the chapter’s arithmetic, then extend it. (a) Route A passes on 9 of 10 days; what is the delivery reliability of a journey that needs two such segments in series? (b) What is the reliability of a system with two substitutable routes, each 90 percent? (c) Recompute both if each route passes on 8 of 10 days. (d) The fuel reserve: three days at plan, consumption at 1.5 times plan, replenishment lead of one day, trigger at two and a half days of stock; how many days of margin does the policy leave? (e) What happens if the lead becomes two days, and what would the trigger have to be to preserve the margin? (f) The cold chain: 40,000 units a day move through the hub; a six-hour outage exposes one quarter of a day’s flow; how many units are at risk, and what does a recovery time objective of 4 hours versus 12 hours mean in units?

(a) 9 of 10 squared, 0.81, about 81 percent: two serial segments deliver only when both work. (b) 1 minus 0.1 squared, 0.99, about 99 percent: two substitutable routes fail only when both fail. (c) At 8 of 10, the serial pair is 0.64, about 64 percent, and the substitutable pair is 1 minus 0.2 squared, 0.96, about 96 percent: the gap between the two designs widens as the parts weaken, which is why the choice of design matters more when conditions are worse. (d) At 1.5 times plan, three days of reserve cover two days. The trigger at two and a half days of stock fires after one-third of a day, the one-day lead consumes one and a half days of stock, and the replenishment arrives with one day of stock remaining, a margin of exactly one day, the difference between a hold and a stop, assuming the supplier is reliable. (e) At a two-day lead, the lead consumes three days of stock at 1.5 times plan, the entire reserve: the policy fails, the depot reaches zero before the replenishment arrives, and no trigger can repair it, because a trigger cannot exceed the three days the depot holds. The reserve must grow, the lead must shorten, or a second fuel source must exist, which is the substitutability lesson of the drill: the arithmetic cannot be fixed with a trigger alone, and the trigger is a function of the lead time and the consumption rate, not a mood. (f) 40,000 units times 0.25, 10,000 units at risk in a six-hour outage: an RTO of 4 hours restores the chain before the batch is lost, an RTO of 12 hours loses the day’s flow, and the difference, 10,000 units, is the number the generator investment is priced against. The trap in every branch of this drill is the independence assumption: the arithmetic assumes the routes and the supplier fail independently, and the storm that floods both routes invalidates the 99 percent, which is why the simultaneous stress test is the companion to the arithmetic, not a footnote.

Three. A field drill: the resilience map. Take the plan you know best. (a) Draw the dependency network: the nodes the promise depends on, technical, human, and external, and the links between them. (b) Trace each success criterion from chapter 2 backwards to the nodes it depends on, and mark every node that appears on the critical path of the same promise twice: that node is a single point by definition. (c) Annotate each node: what fails, what it takes down with it, how fast, and what can substitute. (d) For each single point, choose the move, buffer, redundancy, substitutability, modularity, or a place in the degradation order, and name the cost. (e) Write the degradation order: what fails first, what keeps working, and who decided, and who depends on each line.

The drill succeeds when every promise can be traced to the nodes it depends on, when the human dependencies are on the map, the person who holds the knowledge, the authority, the password, and when each single point has a named move with a named cost. The most common failure is the map that stops at the technical network: the suppliers, the links, the equipment, and never reaches the person who holds the only key. The second failure is the uniform resilience program, a redundancy for every node and a buffer for every risk, and the repair is the stress test: push until the first promise breaks, and spend where it broke first. The test of the map is the question it answers without opening a binder: what is the first promise to break, and what keeps working?

Four. A decision room: the pre-mortem at the milestone. It is the week before the next major milestone of the work you lead, and the milestone is real: the launch, the opening, the go-live, the next tranche. Run the pre-mortem as Klein designed it. It is now three months later, and the milestone failed; each person writes, alone and in silence, the story of how it happened, with a mechanism, not a mood. Share the stories, cluster them, and choose the three most credible. For each of the three, name the leading indicator that would have warned the room, the decision the warning should have forced, and the owner who would have raised it. Then answer the question the pre-mortem cannot answer alone: which of the three would the simultaneous stress test have caught that the single-risk register would have missed?

The discipline is the silence. Written first, in silence, the pre-mortem collects the stories the quiet people hold, and in a pre-mortem the quiet people are often the ones closest to the mechanism; spoken first, the room converges on the loudest story and the loudest story is usually the most familiar one. The output is not a risk register, it is mechanisms and decisions, and the drill passes when every story has a mechanism and every mechanism has an indicator and an owner. The most common failure is the room that runs the pre-mortem as a brainstorm, the failure floating in the air, and the repair is the pencil. Credit belongs to the room that names the failure that embarrasses its own assumptions, because that is the one the plan was built to miss.

Five. The mastery drill: three failures at once. Stress-test your plan against simultaneous supplier, technology, and leadership failure. (a) Name your most critical supplier, the technology whose loss would hurt most, and the person whose absence would cost the project most. (b) Assume all three fail in the same week: the supplier stops, the technology dies, the person is gone. What is the first promise to break, and what keeps working? (c) What is the first action, and who has the authority to take it without calling anyone? (d) Reduce the assumptions one at a time: which single point matters most, the supplier, the technology, or the leadership, and what does that finding say about where the resilience budget belongs? (e) For the answer to (c) to be true, what must already exist, the signed substitute, the tested standby, the pre-delegated authority, and which of them does not exist today?

The drill passes when you can name the degradation order, the first action, and the pre-delegated authority, and when you can say which single point matters most. The most common failure is solving the three failures separately: the supplier plan, the technology plan, the succession plan, each perfect and each useless on the week when they arrive together. The simultaneous assumption is the point, because the register tests one thing at a time and real failure is combinatorial. The second failure is the answer to (c) that begins with a meeting: the first action in the same-week failure is execution, not consultation, and if the authority does not exist in advance, the first action is an improvisation. The third failure is the comfort of the redundancy that has never been tested: the standby that does not stand up, the substitute carrier that has never run the route, the alternate who has never held the authority. The drill’s final question is the one the chapter has been asking since the opening scene: how much of the response is a design, and how much is luck?

Six. The transfer question. On the project you lead, what is the one road, one carrier, one link, one person: the single point of failure the plan assumes will hold, and what would the map show if it were drawn today? When did you last run a rehearsal, and what did it fail to do? What would the near-miss register of your last six months say if it could speak, and which of its entries changed the design? If the most critical person in the project did not answer the phone for a week, what is the first decision that would have to wait, and who is named to make it? And if the worst combination you can name happened on the same day, what keeps working, and what does that tell you about what you are really depending on?

The durable principle: resilience is not prediction, it is the design that holds when the prediction fails, and the design is a map, a playbook, a ladder, and a rehearsal. The register names what the project can name; the scenario set, the pre-mortem, and the stress test open the futures it cannot; the map shows where one failure becomes the whole story; the five moves, the buffer, the redundancy, the substitute, the module, and the degradation order, change the shape of the failure rather than the forecast of it. Every plan will meet the future wrong, and the only question is how much it loses and how fast it recovers, which is why the trigger, the owner, the pre-delegated authority, and the rehearsal are the load-bearing parts of the playbook, and why the near miss is the cheapest lesson the project will ever be offered, if the design changes because of it. The most common next failure is the one the opening scene embodied: the near miss is logged, the meeting is held, the decisions are made, and then the response returns to the single route and the single person, because the eastern route costs more per ton and the fuel authority was never needed again, and the discipline that was built on a bad day quietly dissolves on a hundred good ones. The control is the cadence and the question, the scenario set revisited when an assumption changes, the stress test re-run when the plan changes shape, the near miss reviewed on the same rhythm as the risk burndown, and the rehearsal that fails on purpose before the failure happens by accident. And the discipline has a boundary the next chapter will walk: resilience absorbs the shock, and some obligations cannot be absorbed, the safety floor, the privacy floor, the regulatory evidence, the security of the data, the duty of care that does not degrade, and the assurance that must stay independent of the very pressure that asks it to yield, which is why the next chapter turns from the resilience the project designs to the obligations it cannot design away.

Notes

  • The composite cases remain author-created illustrative material. The Northstar near miss of day fifteen, the convoy that ran six hours late, the midpoint fuel reserve 30 percent below plan with the two deferred top-ups, the satellite link down six hours with the health post radio relay, the fuel authority held by one person with no named alternate, the 41 tons moved against a plan of 44, the eastern route through the highlands scouted but unsigned, and the decisions of the following morning, the signed carriers with the fast-track vetting, the fuel reserve trigger at two and a half days, the named alternates for the link and the authority, the ladder from Mateo at 50,000 units to Daniel at 500,000 units to the donor call through Ines, and the trigger dashboard on the daily convoy report, are all teaching constructions consistent with the facts established in earlier chapters: the flood response and the 16-ton gap from chapter 1 and chapter 4; the fast-tracked northern route, the red-flag controls, the scouts and checkpoints, and the ethical decision record from chapter 4; the framework agreements and the one-page exit plans for the emergency contracts from chapter 20; the northern-route convoy row, the unverified road, the thin safety records, the alternate routes, and Mateo’s ownership from chapter 22; and the cold chain and the medicines as the non-negotiable floors of the response from chapter 4. The cold-chain figures, 40,000 units a day with 10,000 units at risk in a six-hour outage, and the fuel arithmetic, three days at plan, one and a half times plan consumption, a two-and-a-half-day trigger, and a one-day replenishment lead, are simple teaching numbers introduced here and fully reproducible from the text; the route arithmetic, 0.9 squared for the serial pair and 1 minus 0.1 squared for the substitutable pair, and the 0.8 variants, reconcile to the reliabilities stated in the text and the exercise guidance.
  • The scenario practice follows Pierre Wack, “Scenarios: Uncharted Waters Ahead,” Harvard Business Review, September-October 1985, the article that made Shell’s scenario work public, and its argument that scenarios are valuable not for their probabilities but for changing the mental models of decision makers is described here in the author’s own words; the four-world structure, the indicators, the decisions, and the owners are this book’s method-neutral working form, not a reproduction of any proprietary scenario method. The pre-mortem follows Gary Klein, “Performing a Project Premortem,” Harvard Business Review, September 2007, whose design, imagine the failure as a fact and write the story in silence before sharing, is described in the author’s own words. The resilience engineering tradition is attributed to Erik Hollnagel, David D. Woods, and Nancy Leveson, editors of Resilience Engineering: Concepts and Precepts (Ashgate, 2006); the four capabilities, respond, monitor, anticipate, and learn, follow that tradition’s account of the potentials a resilient system must build, stated here in the author’s own words rather than reproduced from any chapter or list in that volume. The high-reliability organization literature, including Karl E. Weick and Kathleen M. Sutcliffe, Managing the Unexpected: Resilient Performance in an Age of Uncertainty (2nd ed., Jossey-Bass, 2007), informs the treatment of near misses and of preoccupation with failure, again described in the author’s own words. Charles Perrow’s argument that interactively complex and tightly coupled systems generate accidents their designers could not imagine comes from Normal Accidents: Living with High-Risk Technologies (Basic Books, 1984), summarized in the author’s own words. Herbert A. Simon’s observation that complex systems are hierarchic and nearly decomposable comes from “The Architecture of Complexity,” Proceedings of the American Philosophical Society 106(6), December 1962, pp. 467-482, paraphrased here. The continuity vocabulary, the recovery time objective and the recovery point objective, follows business continuity practice as standardized in ISO 22301:2019, Security and resilience: Business continuity management systems: Requirements, published by the International Organization for Standardization, with the terms defined in the author’s own words. The near-miss reporting model follows the NASA Aviation Safety Reporting System, operated by NASA for the Federal Aviation Administration since 1976, whose founding design, confidential and non-punitive reporting, is described in the author’s own words. The financial stress-test practice follows the Dodd-Frank Wall Street Reform and Consumer Protection Act of 2010, which made annual stress tests mandatory for the largest U.S. bank holding companies, and the Federal Reserve’s Comprehensive Capital Analysis and Review, run since 2011; the description of stress testing as a probe of structure rather than a prediction of events is the author’s own framing. The Suez Canal obstruction of March 2021, when the container ship Ever Given grounded and blocked the canal for about six days, and the severe flooding in Thailand in 2011, which disrupted global hard disk drive supply for the better part of a year, are used as documented cases at the level of widely reported public events, cited to illustrate concentration risk rather than to assert any disputed causal claim.
  • The chapter’s cross-references to chapters 1, 2, 3, 4, 9, 14, 17, 19, 20, 21, 22, 34, 39, and 42, and the preview of chapter 24, follow the book’s outline. The failure characters, the binder that cannot be recalled from memory, the cold standby that does not stand up, the near-miss register as graveyard, the uniform resilience program, and the degradation theater, are the author’s own constructions, consistent with the failure-aware teaching style established in chapters 17 through 22. The five moves, the buffer, the redundancy, the substitute, the module, and the degradation order, the one-page playbook, the authorization ladder, and the why-not-worse question are the author’s method-neutral working instruments and names. The Northstar data-sensitivity discipline, that humanitarian registration data, patient records, and field incident information do not enter unapproved systems, follows the data-boundary facts established in chapter 4 and the companion discipline established for KijaniPay in chapter 16 and carried in chapters 21 and 22. No proprietary certification manual, commercial text, or framework guide is reproduced or paraphrased here; PMBOK Guide, Scrum, PRINCE2, and similar named materials are not drawn upon for this chapter’s content.