Skip to content

AI Systems Handbook

Appendix A: AI Appropriateness Scorecard

Score whether an AI proposal has a worthwhile problem, viable evidence, controllable error, and a credible path to responsible operation.

Use This Before Choosing a Model

The proposal on the review agenda is a support-reply assistant. The sponsor has a persuasive demo, an estimated saving, and a preferred vendor. Operations has not agreed who would own a harmful reply. Security has not tested whether one customer’s text can expose another customer’s data. Frontline reviewers expect the tool to save time, but nobody has measured the work of checking its answers.

Those omissions are not minor deductions from an otherwise attractive total. They determine what may happen next. The first question is: does this workflow contain a bounded, valuable, evaluable task for which probabilistic behavior is acceptable and governable?

Use the scorecard as a structured challenge, not an arithmetic permission slip. Clear the hard gates first. Then have reviewers score the remaining criteria from evidence, expose disagreements, and record the next decision. Do not calculate a total or average: strength in one row cannot compensate for a blocker or an unsupported assumption in another.

An AI appropriateness path with three hard gates—lawful data and use, tolerable recoverable error, and feasible evaluation—followed by scored value, baseline, oversight, security, and operating readiness.
Remember the order: clear hard blockers before scoring the remaining evidence. A promising business case cannot compensate for unlawful data, intolerable error, or an outcome the team cannot evaluate.

Clear the Hard Gates

Stop at the proposed use and action boundary when any of these is true:

  • the intended use is prohibited by applicable law or binding policy;
  • the team lacks a lawful, authorized, and supportable path for necessary data;
  • a reasonably foreseeable error can cause intolerable harm and cannot be prevented, detected, contained, corrected, or appealed;
  • success and material harm cannot be evaluated with credible evidence;
  • nobody has authority and resources to own the production system and its consequences.

“Stop” does not always mean abandoning the underlying problem. The team may remove an automated action, narrow the population, choose a non-AI route, or return with a materially different proposal. That new boundary needs its own review. Do not let it inherit the earlier proposal’s scores.

Escalate legal, security, privacy, safety, accessibility, labor, procurement, or domain questions to people qualified and authorized to decide them. The scorecard records those decisions; it does not replace specialist review.

Score the Evidence That Remains

Use four scores. A 0 means the proposal cannot proceed at its current boundary. A 1 marks a consequential assumption or missing evidence and requires redesign or investigation. A 2 means there is enough evidence for a bounded experiment with named conditions. A 3 means the design and evidence are strong enough to enter the next risk gate; it is still not deployment approval.

Ask reviewers to score independently before discussion and to attach the evidence they used. A bare number is not a review.

Criterion 0 — Blocked 1 — Weak 2 — Adequate 3 — Strong Evidence to attach
Problem and task No concrete problem or task Broad aspiration Defined task and user Task, workflow boundary, and affected outcomes defined Use case canvas, workflow map
Non-AI baseline None considered Asserted but unmeasured Baseline defined Simpler option tested and comparison recorded Baseline results
AI-specific fit Deterministic behavior is required AI adds novelty, not a clear capability Uncertainty is useful and controllable AI addresses a language, perception, pattern, or scale need demonstrably better Prototype evidence, task analysis
Data and rights Use or data is prohibited, unlawful, or unavailable Provenance, permission, or coverage has major gaps Usable data with bounded gaps Governed, sufficiently representative data with lineage and retention Data inventory or card, rights review
Error tolerance Plausible error creates intolerable, irrecoverable exposure Error costs are poorly understood Errors are detectable or recoverable with controls Error classes, owners, containment, and recovery tested Error analysis, control matrix
Evaluation feasibility No valid success or safety measure Proxy only or no representative sample Metrics, sample, and baseline defined Thresholds, segment tests, human review, and decision rule ready Evaluation plan
Human and process fit No accountable owner or usable oversight Human presence is symbolic Authority and escalation defined Workload, information, override, appeal, and quality assurance tested Oversight plan, service design
Security and privacy Uncontrolled exposure or forbidden access Major unknowns Feasible controls identified Threat and privacy controls designed and testable Threat model, privacy review
Value and lifecycle cost No material value Benefit asserted; lifecycle costs omitted Credible value and total-cost range Measured counterfactual, sensitivity analysis, and benefit owner Business case
Operating readiness No owner after an experiment Monitoring and change control vague Owner, monitoring, incident, and rollback approach defined Service levels, runbooks, change gates, and exit plan tested Operating plan

Decision Rule

After the hard-gate review, use the weakest scores to choose the work rather than the strongest scores to sell the proposal.

  1. Any 0 rejects the proposal at its current boundary. Record whether the response is to stop, choose a non-AI alternative, or submit a materially narrower design.
  2. Every 1 needs an owner, the missing evidence or redesign, a due date, and the decision that evidence will inform. A bounded experiment may be designed to resolve a 1 only when it does not expose people to the unresolved risk.
  3. A proposal with evidence of at least 2 in every applicable row may proceed to risk triage and bounded evaluation. It has not been approved for a pilot or deployment.
  4. Consequential or high-impact systems require independent challenge even when every score is strong.

Record each reviewer’s score and the range for each criterion. A spread of 1–3 is more useful than a median: it exposes a disagreement about evidence, system boundaries, or acceptable risk that the group must name. Do not settle it by voting.

Worked Example: Support Reply Drafting

Criterion Score Evidence or condition
Problem and task 3 agents spend measurable time drafting repetitive replies
Non-AI baseline 2 macros exist; comparison sample defined
AI-specific fit 2 language variation is useful; deterministic account actions excluded
Data and rights 2 authorized tickets; retention and redaction changes required
Error tolerance 2 human sends reply; high-risk intents routed out
Evaluation feasibility 3 correctness, grounding, tone, privacy, and handle-time rubric ready
Human and process fit 2 reviewer authority exists; workload test still needed
Security and privacy 2 tenant isolation and injection tests are launch conditions
Value and lifecycle cost 2 benefit plausible; inference sensitivity modeled
Operating readiness 1 incident ownership and model-change control unresolved

The sponsor calls this a pass because nine rows score at least 2. The scorecard says otherwise. Operating readiness is a 1, reviewer workload remains untested, and the security conditions are designs rather than results.

The next decision is redesign and gather evidence. The operations lead must accept incident ownership and define model-change gates. The evaluation owner must test whether review remains meaningful under representative queue pressure. Security must test tenant isolation and prompt-injection cases before real customer exposure. Only then can the group rescore the affected rows and decide whether to enter risk triage for a bounded experiment.

Decision Record

Proposal:
Review date and version:
Facilitator:
Reviewers and roles:

Hard-gate result: clear / blocked / specialist review required
Scores and evidence by criterion:
Score ranges and unresolved disagreements:

Decision: reject / redesign / bounded experiment / proceed to risk triage
Conditions and owners:
Evidence still required:
Review date:

Quality Check

  • Evidence supports every score; enthusiasm and vendor claims are not evidence.
  • Reviewers considered workflow redesign and a non-AI baseline.
  • Error costs include affected people, not only the buyer or operator.
  • Hard blockers were resolved before the remaining evidence was scored.
  • Disagreement was resolved through evidence or recorded for a named decision-maker; scores were not averaged.
  • Conditions have owners, due dates, and a verification method.
  • The recorded decision distinguishes experimentation, pilot, and deployment.

Continue with the AI Use Case Canvas, The AI Appropriateness Question, and Risk Triage Before Building.