Appendix B: AI Use Case Canvas
Turn an AI idea into a testable system proposal with explicit workflow, affected parties, data, action path, evidence, controls, and stop conditions.
Turn the Pitch Into an Inspectable System
“Use AI to improve operations” is not a use case. A usable proposal names a person, task, workflow, input, output, action, uncertainty, consequence, owner, and stopping rule. Complete this canvas in a cross-functional session before choosing architecture or vendor.
The controlling question is: what changes in the real workflow when this output appears, and who bears the consequences if it is wrong?
Facilitation Method
Invite the workflow owner, a frontline user, technical and data owners, a risk or domain reviewer, and—where feasible—someone who represents affected people. Use actual examples rather than abstract claims. Mark unknowns explicitly. If participants cannot agree on the current workflow or action path, stop and investigate before prototyping.
Keep three boundaries visible throughout:
- System boundary: components, people, providers, data stores, interfaces, and environments included.
- Authority boundary: what the system may suggest, decide, or do; what stays with humans or deterministic policy.
- Evidence boundary: what evaluation can and cannot establish before and after launch.
Follow the Consequence Path
Consider an invoice-intake proposal. The first pitch—“use an LLM to automate invoices”—names a technology and an aspiration, but no system anyone can evaluate. The canvas becomes useful when the team follows one output into the work it changes.
Begin With the Work as It Exists
Accounts-payable staff now open submitted invoices, find the supplier and purchase order, transcribe line items, check totals, and investigate exceptions. New bank details receive a separate verification. The visible cost is handling time; the less visible costs are correction work, delayed suppliers, duplicate payment, and fraud exposure.
That account establishes both the problem and the comparison. Before adding AI, the team must ask what form validation, supplier portals, deterministic extraction, workflow repair, or additional staffing could achieve. “No change” is also a baseline. A proposal earns further work by outperforming a credible alternative, not by outperforming an inefficient process the team has chosen to preserve.
Record who performs each step, which systems and handoffs are involved, where work queues form, and which controls already prevent mistakes. Name the observable outcome that should improve and the people who experience it. If participants cannot agree on this account, the next action is workflow investigation, not prototyping.
Bound the New Authority
In the narrower proposal, the AI component extracts supplier, purchase-order reference, line items, tax, currency, and payment date. Deterministic rules validate totals and approved suppliers. Low-confidence fields, duplicates, mismatches, and new bank details go to accounts-payable review.
The exclusions are as important as the capability: the component cannot create a supplier, change bank details, approve payment, or release funds. Those statements locate the authority boundary. “A human is involved” would not: the canvas must say what the reviewer sees, how much time and skill the review requires, and whether that person can reject the output without penalty.
Now trace the output at least three steps forward. Who sees the extracted fields? Which policy checks run? What enters the payment workflow? What can be corrected after an error, and what cannot? Include suppliers and operational staff among the affected parties, not only the system’s buyer and direct user.
Make the Evidence Match the Consequences
List the input categories and their sources, owners, permissions, sensitivity, expected quality, coverage gaps, freshness, retention, and deletion paths. State what the system must never receive. For invoice intake, evaluation must include the document formats, languages, currencies, handwriting, scan quality, and supplier populations the service will actually encounter.
The value hypothesis is now testable: reduce handling time without increasing duplicate payment, correction, fraud, delayed-payment, or supplier-support burden. Field accuracy alone cannot establish that claim. The evidence plan needs the current workflow baseline, representative samples, segment results, service and safety floors, lifecycle cost, and a decision rule. Lifecycle cost includes integration, evaluation, inference, monitoring, human review, support, compliance, incidents, and exit—not only the model bill.
Error hypotheses should describe consequences rather than merely name risk categories. A missed duplicate can release a second payment. A wrong bank-detail extraction can send a reviewer toward a fraudulent change. Repeated low confidence for one language can shift delay and support work onto the same suppliers. For each credible failure, name who bears it and how the team will prevent, detect, contain, correct, and learn from it.
Give the Proposal an Ending
A deployable proposal names the product, technical, data, operations, domain, security, privacy, and risk owners that its context requires. It also names the person with authority to accept residual risk. “The governance committee” is not an owner unless a particular member has a defined decision.
Specify the environments, regions, user groups, channels, languages, volumes, dependencies, and rollout stages in scope. Changes to the model, prompt, data, threshold, tool, policy, or workflow may invalidate the evidence; name which changes trigger re-evaluation.
Finally, decide in advance what will block development, pause the pilot, roll back production, or retire the system. For invoice intake, those conditions might include an unresolved right to process a document source, failure of the duplicate-payment floor, inability to monitor a critical segment, a severe bank-detail incident, loss of the human review team, or lifecycle cost that erases the expected value. A stopping rule that cannot be observed or executed is only reassurance.
Capture the Use Case
Write compact, evidenced answers. Attach workflow maps, sample inventories, evaluation plans, and risk records rather than squeezing their contents into this page.
USE CASE:
Decision this canvas supports:
Problem and affected outcome:
Current workflow:
Non-AI baseline:
Bounded AI role:
Outside authority:
Users and affected parties:
Inputs, rights, and provenance:
Output and next actions:
Reversibility and correction:
Value hypothesis and lifecycle cost:
Measures, thresholds, and segments:
Error and harm hypotheses:
Human oversight and contestability:
Preventive, detective, and corrective controls:
Fallback and incident handoff:
Deployment context:
Change boundary and re-evaluation triggers:
Owners and required reviews:
Required evidence and artifacts:
Stop, rollback, and retirement conditions:
Open assumptions, owner, and due date:
Decide What the Canvas Supports
The completed canvas should lead to a decision: reject the proposal, redesign it, gather named evidence, run bounded discovery, pilot under explicit conditions, or prepare for deployment review. It does not itself authorize launch. Use these questions to test whether the decision rests on a system rather than a pitch:
- The problem can be understood without mentioning a model or vendor.
- The current workflow and non-AI baseline are evidence, not assumptions.
- The AI role is narrower than the system boundary and has explicit exclusions.
- Affected parties include people beyond the buyer and direct user.
- Every important output has an action path, consequence, owner, and correction route.
- Metrics include safety and service floors as well as value.
- Unknowns have owners and dates; they are not hidden by optimistic prose.
- Stop, rollback, and retirement conditions can be observed and executed.
Use the completed canvas with the AI Appropriateness Scorecard, From Vague Idea to Testable Use Case, and Evaluation Mindset.
Continue reading
Full table of contents