Skip to content

Cybersecurity Engineering Handbook / Chapter 44

Incident Response Operating Model

Define incident roles, severity, evidence handling, communication cadence, and decision records before a crisis begins.

A release identity has deployed an artifact whose digest was never approved. The workload has read an export credential and begun scanning tenant records. Chapter 43 supplied enough evidence for a responder to suspend the release lane, but opening an incident creates a harder problem: several necessary actions now compete with one another.

The technical lead wants to isolate the workload. The security lead wants its volatile state preserved. Customer support needs to know whether any tenant data left the service. Legal and privacy need a defensible account of what is known, not a confident guess. An executive asks when production can reopen. If all five questions enter one undifferentiated chat stream, activity rises while control disappears.

An incident operating model makes that pressure governable. It gives each workstream an owner while keeping authority, facts, decisions, evidence, communication, and recovery in one shared account of the incident.

Incident bridge operating model with incident commander, technical lead, security lead, communications, legal privacy, customer support, scribe, executive sponsor, decision log, timeline, evidence repository, containment, customer updates, and executive updates.
The incident bridge works when roles feed one decision log and timeline while technical, evidence, customer, and executive streams stay coordinated.

Declare on credible harm, not perfect certainty

An event is an observable fact: an identity accepted an artifact, a workload read a credential, a policy changed. An alert is a signal selected for triage. Neither is automatically an incident. A security incident begins when credible unauthorized activity, abuse, exposure, or control failure requires coordinated action. The declaration is an operating decision, not a claim that the team already knows the cause.

Some incidents need additional lanes. A major incident threatens broad customer, operational, financial, safety, legal, or trust impact. A privacy incident may involve unauthorized access, disclosure, alteration, loss, or unavailability of personal data and therefore brings privacy and legal judgment into the response. A vulnerability emergency is a time-sensitive weakness that demands coordinated mitigation even when exploitation has not been observed.

These definitions should lower the cost of declaring. Waiting for proof can give an attacker time; calling every suspicious event a crisis exhausts the response system. The release case qualifies because a privileged production transition, an unapproved artifact, credential access, and tenant reads require several owners to make coupled decisions. The commander can later downgrade or close it without treating the original declaration as a mistake.

The declaration record is short:

  • incident identifier, title, declaration time, and declaring person;
  • current severity and the observations that justify it;
  • affected or potentially affected services, environments, identities, data, and customers;
  • commander, technical lead, security lead, scribe, and required specialist lanes;
  • bridge and evidence-repository locations, access restrictions, and next update time.

Put command beside investigation

Assign functions, not prestigious titles. The incident commander owns coordination: severity, objectives, cadence, decision routing, handoffs, and closure. The commander asks whether the team has the people, authority, and shared facts to act. They should not also become the deepest investigator; command vanishes when its owner disappears into a shell session.

The technical lead directs investigation, containment, eradication, and recovery across the affected system. The security lead maintains the threat interpretation, attacker-path questions, evidence needs, and residual-risk view. Their advice can conflict for good reasons. Isolating the workload may stop tenant reads but destroy a volatile trail or trigger attacker behavior elsewhere. The commander makes the trade-off explicit, finds the authorized decision owner, and records the result.

The scribe keeps the operational memory: timeline, facts, hypotheses, decisions, actions, evidence links, and unanswered questions. This is active response work, not meeting administration. A precise record allows the next shift to inherit the incident instead of reconstructing it.

Add lanes as consequences demand them. A communications lead coordinates internal, executive, customer, and public messages. A legal/privacy lead advises on personal data, contractual or regulatory duties, evidence concerns, and law-enforcement contact; engineering should surface facts and deadlines without improvising legal conclusions. A customer/support lead joins reports across customers and prevents unverified internal theories from leaking into replies. An executive sponsor removes organizational blockers and accepts business-level risk that exceeds the commander’s authority.

In a small incident, one person may hold several functions. Keep the function labels anyway. “Mara is both commander and communications lead” reveals a load and review risk that “Mara is handling it” conceals. Every handoff should state current severity and objectives, unresolved high-risk questions, active actions, pending decisions, the next update, and who now owns each role.

Run parallel lanes against one clock

Prepare, detect, analyze, contain, eradicate, recover, communicate, and learn are not a relay race. During the release incident, analysis continues while the team contains the release identity; communication starts before customer impact is settled; recovery design begins before every persistence mechanism is known. Treat the lifecycle as concurrent lanes joined by decisions.

The first command cycle should establish five things:

  1. Current harm and uncertainty. Are tenant reads continuing? Did data leave the trust boundary? Is the identity still usable? Which logs or systems may disappear with time?
  2. A bounded objective. For the next thirty minutes: stop further exports, preserve the workload and control-plane evidence, and determine whether any second workload accepted the artifact.
  3. Authority. Who may suspend the production identity, isolate the workload, disable an export route, notify customers, or accept continued operation?
  4. Parallel owners. Name the people investigating scope, executing containment, preserving evidence, assessing customer and data impact, and drafting the next update.
  5. A return time. Every action and question has an owner and an expected update. The bridge reconvenes for decisions, not a continuous recital of debugging.

Containment stops or limits harm. Prefer a narrow, reversible action when it is fast enough: suspend the release identity and export route, quarantine the artifact, isolate the workload, and protect the evidence sources. A broad shutdown may be right when scope is uncertain and harm is severe, but the service cost and evidence consequences belong in the decision record.

Eradication removes the attacker’s access and the vulnerable state: unauthorized artifacts, persistence, compromised credentials, malicious configuration, and the path that admitted them. Recovery restores a trustworthy service, not merely a running one. For this incident, that means an approved artifact from a trusted build path, reissued credentials, verified policy and logging, no unexplained tenant reads, monitoring for replay, and business validation of the export boundary. Reopening production is a new decision with named evidence and rollback conditions.

Severity is a promise about response

Choose severity from plausible impact and required urgency, then revise it as facts change. The number should determine authority, staffing, and communication cadence rather than decorate a ticket.

  • Sev 1 — organization-level crisis. Use for active or likely broad compromise, destructive activity, severe sensitive-data exposure, loss of a critical control plane, major customer harm, or an urgent legal or safety concern. Maintain active command and technical bridges. Set an explicit internal cadence—often thirty minutes or less—and an executive cadence appropriate to the decisions required. Staff communications, legal/privacy, and customer impact early.
  • Sev 2 — dedicated incident response. Use for confirmed compromise with limited or uncertain scope, material exposure contained to a bounded population, or a vulnerability emergency with a credible path into critical systems. Keep a dedicated team and bridge; update stakeholders at a declared interval, commonly hourly while conditions are changing.
  • Sev 3 — managed investigation. Use for contained suspicious activity with credible risk or a localized control failure that still needs coordinated ownership. Assign a commander or incident owner, schedule checkpoints, and set an escalation deadline rather than leaving an open-ended ticket.
  • Sev 4 — tracked security event. Use for low-impact activity needing investigation or corrective follow-up. A ticketed workflow is sufficient, with an owner and due date.

Read together, the four levels form the severity matrix: each one binds impact to an operating response. Local plans should replace illustrative cadences with values the organization can actually staff.

Escalate immediately for privileged or control-plane access, lateral movement, continuing sensitive-data access, destructive action, disabled evidence sources, credible customer reports, public disclosure, legal or privacy concern, or uncertain scope in a critical system. Do not wait for the next scheduled update. Downgrade only when the record shows why the higher-impact path is no longer plausible, not because the bridge has become quiet.

Maintain one record without flattening uncertainty

The incident record must distinguish what happened from what the team thinks happened. Give every entry a type:

  • a fact points to an observation and its source;
  • a hypothesis states an explanation, supporting and conflicting evidence, and how it will be tested;
  • a decision records an authorized choice;
  • an action has an owner, status, and result;
  • an open question has an owner and next check;
  • an evidence reference points to a preserved artifact without copying sensitive material into the bridge.

The timeline keeps three clocks when they differ: when activity occurred, when the organization observed it, and when the response acted. For example:

10:02 event     release identity accepted digest sha256:… (deployment record D-184)
10:06 event     workload requested tenant-export credential (audit event A-772)
10:11 observed  correlation case C-91 reached on-call
10:17 decision  suspend identity and export route; IC approved; preserve workload first
10:19 action    identity suspended by technical lead; control-plane record E-203

Avoid editing an old entry into the new truth. Append the correction and link it. That preserves how the team’s understanding changed, which is essential when evaluating decision quality later.

A decision entry uses a repeatable shape:

Time / decision / decision owner
Facts used and uncertainty that remains
Options considered
Expected reduction in harm
Operational and evidence risks
Action owner and deadline
Reversal condition or next review

Keep the bridge for coordination, the record for durable state, and the evidence repository for sensitive artifacts. If chat is the only durable record, access control, discovery, handoff, and chronology will all become harder at once.

Preserve before changing state

Containment often alters the system the team must understand. Before rotating credentials, terminating workloads, deleting resources, changing policy, wiping hosts, or restoring backups, ask what evidence the action will destroy and whether the delay required to preserve it creates unacceptable harm. Record the answer. Evidence preservation is a decision constraint, not an automatic veto on containment.

For the release incident, preserve the suspect artifact and digest, deployment and provenance records, identity and credential audit events, workload metadata and volatile state where feasible, tenant-read and egress logs, relevant configuration, alert payload, and source-health status. Extend retention on affected sources early; ordinary rollover must not erase the first hour while the incident is still expanding.

Store originals or defensible copies in an access-controlled location. Record collector, collection time, source, method, relevant system time, hash where meaningful, storage location, and every transformation. Limit access to people with an incident need. If litigation, law enforcement, insurance, regulation, or internal policy may require formal chain of custody, assign an evidence owner and obtain specialist guidance. Engineers do not need to imitate forensic examiners, but they must not make later examination impossible through undocumented handling.

Communicate the state, not the theory

Every update answers the same operational questions: What is confirmed? What is still being investigated? What customer or business effect is known? What has the team done? What decisions or help are needed? When is the next update?

An executive update can be copied from this structure:

Incident / severity / time
Confirmed impact: …
Potential impact still under investigation: …
Current containment and service state: …
Customer/data assessment: …
Decisions or resources needed by: …
Next update: …

Say “we have confirmed reads of 18 tenant records; egress remains under investigation,” not “no data was exfiltrated” merely because the egress query is unfinished. Do not announce root cause while alternatives remain live. Express recovery times as current estimates with assumptions, owner, and next revision point.

The internal bridge may contain restricted operational detail. Executive messages compress facts for risk decisions. Customer updates need service- and customer-specific accuracy plus the organization’s review path. Legal and privacy should be engaged promptly when personal data, contractual notice, regulated systems, insurers, or law enforcement may be involved. The public status page should describe customer-visible service effects and restoration work without disclosing defensive details that create further risk. These streams share facts, but they are not interchangeable audiences.

Close only after recovery earns trust

Stopping the alert is not closure. Before lowering command, the team should show that active harm has stopped; known attacker access and persistence are removed; affected identities, artifacts, policies, and data paths are accounted for; the restored service came from a trusted state; monitoring and evidence sources are healthy; customer and legal decisions have owners; and any temporary control has an expiry or a path to permanence.

The post-incident review reconstructs how the system and response behaved. Its working template is:

Scope and customer/data impact
Detection, declaration, containment, eradication, and recovery timeline
Root cause and contributing conditions
What limited harm; what increased it
Missed or delayed signals and evidence gaps
Command, handoff, and communication decisions
Recovery proof and residual risk
Actions: owner / due date / verification / closure approver

Root cause should explain a controllable system condition, not stop at “human error” or the compromised identity. In the release case, ask why an unapproved digest could become accepted production state, why the identity could reach the export credential, which boundary should have limited tenant reads, and which observation finally supported containment. Also examine the response: Did command begin soon enough? Did preserving the workload delay a necessary block? Could support determine customer impact from authoritative data?

A useful review is candid without becoming a blame document. It distinguishes choices that were reasonable with the facts available from choices weakened by missing preparation. Its actions alter code, architecture, permissions, telemetry, detection logic, runbooks, ownership, exercises, or recovery evidence. Each action needs an owner, deadline, verification method, and someone authorized to accept closure. “Share the lesson” is not a control improvement.

The plan to prepare before the bridge opens

The durable incident response plan need not predict every scenario. It must make the first decisions executable. Keep it where responders can reach it during identity, network, or collaboration-system failure. Name declaration paths, role authorities and alternates, severity commitments, protected contact methods, bridge and record locations, evidence rules and repositories, containment authorities, communication review paths, customer/data impact ownership, handoff expectations, recovery gates, post-incident review timing, and exercise cadence.

Test the plan with an inconvenient case. Declare the unapproved-release incident after hours. Remove the primary commander. Make one telemetry source late. Require a choice between rapid workload termination and volatile evidence. Ask support to identify affected tenants and an executive to decide whether a bounded export outage may continue. The exercise succeeds when the team can preserve a coherent record and make authorized decisions despite those constraints—not when everyone reaches the expected answer.

Chapter 45 turns this operating model into scenario-specific first-hour playbooks. Those playbooks can name what to inspect and contain. They work only because command, severity, evidence, communication, and recovery authority no longer have to be invented while harm is unfolding.