Skip to content

AI Systems Handbook / Chapter 9

Risk Triage Before Building

Classify AI use cases by consequence, sensitivity, autonomy, scale, affected rights, and recoverability so required reviews and controls precede prototype momentum.

The Prototype That Arrived Before the Question

An HR team asks an analyst whether AI could help managers prepare promotion reviews. Two weeks later, a prototype scores employee narratives against examples of previously promoted staff. Leaders like the ranked dashboard and ask when it can connect to the performance system.

The artifact makes the proposal feel more settled than it is. No one has decided whether past promotion outcomes are an appropriate target, what sensitive information the narratives expose, how a score would influence managers, whether an employee could contest an error, or which employment and labor rules apply. Stopping now looks like losing work, even though the work answered none of the consequential questions.

Risk triage belongs before the prototype. It does not predict every harm or confer legal approval. It makes an earlier decision: stop, narrow, investigate, or proceed under gates strong enough for the proposed action and its consequences.

Split the Proposal at the Action Boundary

“Help with promotion reviews” contains at least three systems.

An employee-facing writing aid could compare a self-authored narrative with a published rubric. The employee controls both input and output, and nothing enters the personnel record automatically. A manager-facing evidence organizer could retrieve approved records and group them under rubric categories for verification. A promotion score could rank employees against historical decisions and feed that rank into the promotion process.

These systems might use similar retrieval, language, or ranking capabilities. Their risk differs because of what happens after the output. The writing aid can shape how an employee presents a case. The organizer can omit or misattribute evidence seen by a manager. The ranking can influence access to pay and opportunity while hiding the assumptions learned from previous decisions.

Follow each output until a person, record, permission, payment, service, or physical condition changes. Then ask who bears a wrong result, whether they can see it, how quickly it can be corrected, and what happens while correction is pending. The useful unit of triage is this complete action path, not the model family or the product name.

That is why identical capabilities move between consequence classes. A summary kept as a private note differs from one entered into a medical record. A classifier that organizes personal photographs differs from one that routes public-benefit applications. A tool that drafts an access request differs from one that grants production access.

Find the First Decisive Escalation Signal

Triage begins with the intended purpose, direct users, and people affected by the eventual decision. It then follows the data through prompts, retrieval, logs, evaluation sets, human review, vendors, and deletion. Employment, health, financial, biometric, location, confidential, and other sensitive data can change both the severity of exposure and the reviewers needed before any sample is collected.

Autonomy changes the action path. An output that can send a message, move money, change a record, alter access, execute code, or repeat an action deserves separate classification from the suggestion that preceded it. A confirmation dialog is useful only when the reviewer can see the target, evidence, and effect and has time and authority to refuse.

Error becomes more serious as it becomes hidden, difficult to detect, or difficult to reverse. Lost opportunity, denied service, public disclosure, physical harm, or a cascading infrastructure change cannot be averaged away by many harmless cases. Scale and duration matter for the same reason: a small error rate across a large population, or a persistent feedback loop, can accumulate into substantial harm.

Rights and democratic values become engineering constraints at this point. Equal treatment requires evaluation that can reveal uneven burdens. Privacy limits what data enters the system and where it travels. The ability to contest a consequential outcome requires traceable evidence, a qualified human, a timely correction path, and a non-AI route when the system fails. If the organization cannot provide those things, it has found a design limit rather than a documentation gap.

Security asks who could influence the action path through prompts, retrieved content, uploaded files, tools, or feedback. Operations asks for the worst plausible event at launch scale, its first observable signal, and a person who can contain it. A proposed pilot reduces exposure only when its boundaries, stop signals, fallback, and cleanup have been proved.

Do not average these signals into a reassuring score. One decisive condition can control the route. A legally barred action remains barred even if the data is ordinary and the interface is excellent. An irreversible safety failure remains high consequence even if it is rare. A system with no accountable pause authority is not operationally governable.

Use Tiers to Route Work

An internal tier is a routing device. It answers which evidence, reviewers, controls, and decision authority must exist before work crosses the next boundary.

Tier 0 routes away from a deployed AI use. The team may choose a non-AI solution or confine exploration to synthetic, non-sensitive data with no operational effect. Standard software and research controls still apply, but the experiment must not quietly become a product.

Tier 1 covers assistive, limited-consequence use. Errors are visible and recoverable, and no material individual outcome follows without ordinary review. Proceeding requires a named owner, a bounded use-case record, representative task tests, usage guidance, data minimization, basic logging, a correction path, and monitoring appropriate to the workflow.

Tier 2 covers material but recoverable impact. Output can change service, money, access, workload, reputation, or user experience. Before a pilot, the team needs formal success and severe-error thresholds, segment and end-to-end evaluation, a threat model, tested permissions, measured reviewer workload, escalation and incident playbooks, documentation, and an accountable approval decision.

Tier 3 covers high impact. Safety, rights, sensitive data, consequential decisions, vulnerable populations, broad scale, significant autonomy, or difficult recovery call for an impact assessment and alternatives analysis, independent challenge, domain and affected-party evidence, specialist security and privacy gates, meaningful appeal and fallback, tightly staged exposure, and senior risk acceptance. Some proposals will remain unacceptable after every feasible control is added.

Tier 4 is prohibited. Law, contract, organizational policy or values, or the inability to control an intolerable consequence bars the proposed use or action boundary. Do not build or deploy it. A narrower use must return as a new proposal, not inherit approval from the prohibited one.

A risk-triage board shows six escalation signals and four tiers from assistive through material and high impact to prohibited, with increasingly strong evidence, security, privacy, oversight, and independent assurance gates.
Higher consequence requires stronger evidence and decision authority before development. Tiers route work to controls and reviewers; they are not badges of inherent model safety.

Tiering does not make controls uniform. A Tier 1 coding assistant that handles confidential source may need a strong security gate. A Tier 3 system may need continuous containment in one setting and staffed-hours operation with a safe shutdown in another. Assign controls from the particular failure path, then use the tier to prevent the team from understating the necessary assurance.

Route the Promotion Proposal

The employee writing aid could begin at Tier 1 only if the employee knowingly chooses the tool, controls what enters it, sees and can discard the output, and the result stays outside the employment record. Before even that bounded exploration, privacy and security owners must approve the data path; accessibility review must ensure the aid does not create a new participation barrier; and evaluation must test whether rubric guidance is accurate across the forms of evidence employees actually use. If the employer makes the tool effectively mandatory, stores its output, or allows managers to infer who used it, the original route no longer holds.

The manager evidence organizer begins at least at Tier 2. It touches approved employment records and can shape what a decision-maker notices. The team must prove access enforcement and provenance, test omission and misattribution by relevant group and evidence type, measure whether managers verify rather than rubber-stamp summaries, and preserve the employee’s correction and appeal path. Because the use sits inside an employment process, legal, privacy, labor or works-council, accessibility, domain, and security reviewers may each have a distinct question that engineering cannot answer for them.

The promotion score reaches Tier 3 before model evaluation begins. Historical promotion is not neutral ground truth; it records earlier criteria, discretion, opportunity, and possible inequity. The score would affect a consequential employment decision, sensitive proxies may enter indirectly, over-reliance may be difficult to observe, and lost opportunity is not easily restored after a cycle closes. Development stops while legal and policy classification, labor obligations, impact assessment, affected-party challenge, and the availability of a less harmful design are resolved. If law or organizational policy prohibits the use, the route is Tier 4. A polished prototype cannot lower it.

The action path, not the sophistication of the model, produced all three routes.

Name Reviewers by the Question They Own

“Send it to governance” hides responsibility. The product owner defines the permitted purpose and the next development boundary. Domain experts judge whether the task and evidence represent the real decision. Security examines attack paths and permissions. Privacy examines necessity, purpose, movement, retention, and deletion of data. Accessibility reviewers test whether the interface and fallback work for the people expected to use them. Legal specialists classify applicable obligations; labor representatives or works councils address workplace processes where relevant. Affected people can expose consequences and assumptions invisible to the delivery team.

Internal tiering and legal classification are separate records. They may use different definitions and reach different results. A low internal tier cannot waive a legal duty, and a system outside a named statutory category may still violate policy or create unacceptable harm. Record the jurisdiction, source, reviewer, and date rather than writing “legally approved.”

People also need enough AI literacy for the authority assigned to them. Builders need to understand limits and foreseeable misuse. Operators need to recognize failure and escalate it. Reviewers need evidence, time, and permission to reject an output. Affected users need an honest explanation of the system’s role and a usable route to human help. Training is therefore a control to test—through decisions people can actually make—not an attendance record. In jurisdictions that impose specific AI-literacy duties, legal review must translate those duties into roles and evidence.

Higher-tier review needs independence from the pressure to deliver. Product and engineering provide facts, but a reviewer who can challenge the framing, tests, controls, and residual risk must not depend on launch for success. Risk acceptance should name the remaining exposure, affected people, evidence and uncertainty, controls, maximum intended scale, review date, change triggers, and the authority for pause, rollback, and retirement. “Approved by governance” cannot reconstruct why the system was allowed to operate.

Leave With a Route, Not a Score

The initial triage record should let another reader reconstruct the decision. Identify the system version and owner; intended purpose and prohibited use; users and affected people; data and processors; output and action boundary; permissions, scale, severity, detectability, reversibility, and appeal; plausible misuse and failure; the decisive escalation signals; internal tier; separate legal or policy classification still required; reviewers and evidence gates; blocking assumptions; and the authority and expiry of the decision.

Finish with one of five routes: stop, redesign the action boundary, gather named evidence, proceed to a bounded experiment, or escalate to specialist review. Every route needs an owner and a condition for returning. “More review” without a question or decision-maker merely delays the same ambiguity.

Try the route on five proposals before reading anyone else’s answer: a private meeting summarizer, a clinical diagnosis assistant, a warehouse demand forecast, an employee performance scorer, and an internal document-search assistant. Split each proposal if its outputs lead to different actions. For the strongest escalation signal, name the reviewer, evidence, and stop condition it creates. If a tier changes when you add automatic action, sensitive data, tenfold scale, or removal of human review, the exercise has exposed the boundary that matters.

Triage is not permanent. Reopen it when advice becomes action, review is reduced, a population or region is added, data purpose or provider changes, a new tool or permission appears, scale expands, evaluation reveals a severe failure, or law, policy, threats, or organizational values change. A stable product name can conceal a materially different system.

The HR team now has three proposals instead of one impressive dashboard. One may proceed under narrow gates, one requires substantial evidence and specialist review, and one stops before prototype momentum becomes an argument. Business Case, Cost, and Value Realization asks whether any proposal that survives this route is worth its full operating commitment.

Source Notes

  • NIST AI Risk Management Framework 1.0, voluntary and use-case-agnostic framework for managing risks to individuals, organizations, and society; published 2023.
  • NIST AI RMF Core, continuous Govern, Map, Measure, and Manage functions; contextual go/no-go decisions; trained and accountable roles; independent assessment; and appeal, override, and change management; verified 2026-07-19.
  • NIST AI RMF: Generative Artificial Intelligence Profile, NIST AI 600-1, generative-AI risk actions spanning governance, content, privacy, security, and human configuration; published 2024.
  • Regulation (EU) 2024/1689, Article 4, AI-literacy obligations framed around provider and deployer personnel and others operating AI systems on their behalf, with regard to knowledge, experience, training, context, and affected people; verified 2026-07-19.