Appendix A: AI Appropriateness Scorecard
Score whether an AI proposal has a worthwhile problem, viable evidence, controllable error, and a credible path to responsible operation.
Use This Before Choosing a Model
The proposal on the review agenda is a support-reply assistant. The sponsor has a persuasive demo, an estimated saving, and a preferred vendor. Operations has not agreed who would own a harmful reply. Security has not tested whether one customer’s text can expose another customer’s data. Frontline reviewers expect the tool to save time, but nobody has measured the work of checking its answers.
Those omissions are not minor deductions from an otherwise attractive total. They determine what may happen next. The first question is: does this workflow contain a bounded, valuable, evaluable task for which probabilistic behavior is acceptable and governable?
Use the scorecard as a structured challenge, not an arithmetic permission slip. Clear the hard gates first. Then have reviewers score the remaining criteria from evidence, expose disagreements, and record the next decision. Do not calculate a total or average: strength in one row cannot compensate for a blocker or an unsupported assumption in another.
Clear the Hard Gates
Stop at the proposed use and action boundary when any of these is true:
- the intended use is prohibited by applicable law or binding policy;
- the team lacks a lawful, authorized, and supportable path for necessary data;
- a reasonably foreseeable error can cause intolerable harm and cannot be prevented, detected, contained, corrected, or appealed;
- success and material harm cannot be evaluated with credible evidence;
- nobody has authority and resources to own the production system and its consequences.
“Stop” does not always mean abandoning the underlying problem. The team may remove an automated action, narrow the population, choose a non-AI route, or return with a materially different proposal. That new boundary needs its own review. Do not let it inherit the earlier proposal’s scores.
Escalate legal, security, privacy, safety, accessibility, labor, procurement, or domain questions to people qualified and authorized to decide them. The scorecard records those decisions; it does not replace specialist review.
Score the Evidence That Remains
Use four scores. A 0 means the proposal cannot proceed at its current boundary. A 1 marks a consequential assumption or missing evidence and requires redesign or investigation. A 2 means there is enough evidence for a bounded experiment with named conditions. A 3 means the design and evidence are strong enough to enter the next risk gate; it is still not deployment approval.
Ask reviewers to score independently before discussion and to attach the evidence they used. A bare number is not a review.
| Criterion | 0 — Blocked | 1 — Weak | 2 — Adequate | 3 — Strong | Evidence to attach |
|---|---|---|---|---|---|
| Problem and task | No concrete problem or task | Broad aspiration | Defined task and user | Task, workflow boundary, and affected outcomes defined | Use case canvas, workflow map |
| Non-AI baseline | None considered | Asserted but unmeasured | Baseline defined | Simpler option tested and comparison recorded | Baseline results |
| AI-specific fit | Deterministic behavior is required | AI adds novelty, not a clear capability | Uncertainty is useful and controllable | AI addresses a language, perception, pattern, or scale need demonstrably better | Prototype evidence, task analysis |
| Data and rights | Use or data is prohibited, unlawful, or unavailable | Provenance, permission, or coverage has major gaps | Usable data with bounded gaps | Governed, sufficiently representative data with lineage and retention | Data inventory or card, rights review |
| Error tolerance | Plausible error creates intolerable, irrecoverable exposure | Error costs are poorly understood | Errors are detectable or recoverable with controls | Error classes, owners, containment, and recovery tested | Error analysis, control matrix |
| Evaluation feasibility | No valid success or safety measure | Proxy only or no representative sample | Metrics, sample, and baseline defined | Thresholds, segment tests, human review, and decision rule ready | Evaluation plan |
| Human and process fit | No accountable owner or usable oversight | Human presence is symbolic | Authority and escalation defined | Workload, information, override, appeal, and quality assurance tested | Oversight plan, service design |
| Security and privacy | Uncontrolled exposure or forbidden access | Major unknowns | Feasible controls identified | Threat and privacy controls designed and testable | Threat model, privacy review |
| Value and lifecycle cost | No material value | Benefit asserted; lifecycle costs omitted | Credible value and total-cost range | Measured counterfactual, sensitivity analysis, and benefit owner | Business case |
| Operating readiness | No owner after an experiment | Monitoring and change control vague | Owner, monitoring, incident, and rollback approach defined | Service levels, runbooks, change gates, and exit plan tested | Operating plan |
Decision Rule
After the hard-gate review, use the weakest scores to choose the work rather than the strongest scores to sell the proposal.
- Any 0 rejects the proposal at its current boundary. Record whether the response is to stop, choose a non-AI alternative, or submit a materially narrower design.
- Every 1 needs an owner, the missing evidence or redesign, a due date, and the decision that evidence will inform. A bounded experiment may be designed to resolve a 1 only when it does not expose people to the unresolved risk.
- A proposal with evidence of at least 2 in every applicable row may proceed to risk triage and bounded evaluation. It has not been approved for a pilot or deployment.
- Consequential or high-impact systems require independent challenge even when every score is strong.
Record each reviewer’s score and the range for each criterion. A spread of 1–3 is more useful than a median: it exposes a disagreement about evidence, system boundaries, or acceptable risk that the group must name. Do not settle it by voting.
Worked Example: Support Reply Drafting
| Criterion | Score | Evidence or condition |
|---|---|---|
| Problem and task | 3 | agents spend measurable time drafting repetitive replies |
| Non-AI baseline | 2 | macros exist; comparison sample defined |
| AI-specific fit | 2 | language variation is useful; deterministic account actions excluded |
| Data and rights | 2 | authorized tickets; retention and redaction changes required |
| Error tolerance | 2 | human sends reply; high-risk intents routed out |
| Evaluation feasibility | 3 | correctness, grounding, tone, privacy, and handle-time rubric ready |
| Human and process fit | 2 | reviewer authority exists; workload test still needed |
| Security and privacy | 2 | tenant isolation and injection tests are launch conditions |
| Value and lifecycle cost | 2 | benefit plausible; inference sensitivity modeled |
| Operating readiness | 1 | incident ownership and model-change control unresolved |
The sponsor calls this a pass because nine rows score at least 2. The scorecard says otherwise. Operating readiness is a 1, reviewer workload remains untested, and the security conditions are designs rather than results.
The next decision is redesign and gather evidence. The operations lead must accept incident ownership and define model-change gates. The evaluation owner must test whether review remains meaningful under representative queue pressure. Security must test tenant isolation and prompt-injection cases before real customer exposure. Only then can the group rescore the affected rows and decide whether to enter risk triage for a bounded experiment.
Decision Record
Proposal:
Review date and version:
Facilitator:
Reviewers and roles:
Hard-gate result: clear / blocked / specialist review required
Scores and evidence by criterion:
Score ranges and unresolved disagreements:
Decision: reject / redesign / bounded experiment / proceed to risk triage
Conditions and owners:
Evidence still required:
Review date:
Quality Check
- Evidence supports every score; enthusiasm and vendor claims are not evidence.
- Reviewers considered workflow redesign and a non-AI baseline.
- Error costs include affected people, not only the buyer or operator.
- Hard blockers were resolved before the remaining evidence was scored.
- Disagreement was resolved through evidence or recorded for a named decision-maker; scores were not averaged.
- Conditions have owners, due dates, and a verification method.
- The recorded decision distinguishes experimentation, pilot, and deployment.
Continue with the AI Use Case Canvas, The AI Appropriateness Question, and Risk Triage Before Building.
Continue reading
Full table of contents