Skip to content

AI Systems Handbook

Appendix K: AI Launch Readiness Checklist

Make an evidence-based launch, limited-launch, or block decision across use case, evaluation, controls, operations, people, and rollback.

Make Launch a Decision, Not a Calendar Event

A claims-triage service passes its aggregate recall target, and the launch review is ready to approve nationwide release. Buried in the evaluation record, one low-volume urgent claim type misses its floor. The proposed safeguard is to route those claims to the existing manual queue, but nobody has measured that queue at launch volume. The provider name in the evaluation also points to a moving alias rather than the version that produced the results. Every team completed its document; the evidence still does not support the proposed scope.

Launch readiness is an argument built from linked evidence. The checklist does not award points for filled boxes. It asks whether the exact production version, population, authority, controls, staff, and recovery path support a launch decision—and whether a reviewer can limit or block release when they do not.

A launch-readiness chain passes through use case, evidence, controls, operations, and people gates before reaching launch, limit, or block decisions, with tested rollback under the entire route.
Readiness is a chain: approved purpose, sufficient evidence, effective controls, prepared operations, and ready people. Tested rollback supports every gate; one material break can limit or block launch.

Freeze the Candidate and the Scope

Run this review on an immutable release candidate: model, prompt, retrieval corpus, data pipeline, policy, tools, permissions, interface, and code. Record a digest or other release identifier that lets a reviewer match the candidate to the evaluation evidence and deployed system. Name the users, regions, languages, channels, traffic, and action authority requested. Evidence from a narrower version or population cannot silently transfer to a broader launch.

Assign each item an evidence link and owner. Mark N/A only with a reason and approver. A waiver must state the gap, compensating controls, expiry, and residual-risk owner. Reopen the review if any behavior-changing component or launch assumption changes.

AI Launch Readiness Record

RELEASE IDENTITY
System and immutable release-candidate version:
Release digest and evaluated artifact references:
Changes since the last evaluation or review:
Requested launch date and decision date:
Requested scope: users, regions, languages, channels, volume:
AI role and maximum action authority:
Product, technical, operations, and risk owners:
Decision forum and required approvers:

DECISIVE CLAIMS (repeat)
Claim that must remain true in production:
Acceptance floor and required segments:
Evidence for this candidate and scope:
Control or fallback if the claim fails:
Owner with authority to act:
What remains uncertain:

USE CASE AND GOVERNANCE
[ ] Use case canvas names workflow, users, affected people, and non-AI baseline
[ ] Appropriateness decision remains valid for this scope
[ ] Risk triage and impact assessment are current
[ ] Prohibited and out-of-scope uses are enforceable and communicated
[ ] Required legal, privacy, security, accessibility, labor, and domain reviews complete
[ ] Accountable owners and residual-risk authority are named
Evidence / owner / gap:

DATA, MODEL, AND SYSTEM DOCUMENTATION
[ ] Dataset records cover released snapshots, rights, gaps, and maintenance
[ ] Model and system cards identify versions, evidence, limits, and constraints
[ ] Prompt, policy, retrieval, tool, and code versions are traceable
[ ] Data retention, deletion, correction, and provenance paths work
[ ] Vendor terms, change notice, availability, and exit assumptions are reviewed
Evidence / owner / gap:

EVALUATION AND ASSURANCE
[ ] Evaluation plan was approved before acceptance testing
[ ] Non-AI and prior-system baselines were compared
[ ] Primary, segment, safety, security, privacy, fairness, and service floors pass
[ ] Human evaluation and domain review pass where judgment is required
[ ] Robustness, adversarial, abuse, accessibility, and edge-case tests pass
[ ] Failed cases became fixes, restrictions, or protected regression tests
[ ] Evidence supports this exact scope; uncertainty is visible
Evidence / owner / gap:

CONTROL EFFECTIVENESS
[ ] Identity, access, and least privilege are tested end to end
[ ] Input, retrieval, tool, output, and downstream controls are tested
[ ] Human oversight has information, time, authority, escalation, and fallback
[ ] Abstention, safe mode, and deterministic or human fallback work
[ ] Rate, time, cost, and action limits contain maximum credible blast radius
[ ] Privacy and security logging support investigation without excess collection
Evidence / owner / gap:

OPERATIONS AND RELIABILITY
[ ] Service-level objectives and operating envelope are approved
[ ] Dashboards and actionable alerts cover service, data, model, safety, user, and cost
[ ] Alert owners, rotations, runbooks, and escalation targets are staffed
[ ] Provider outage, degraded quality, stale data, and dependency failures rehearsed
[ ] Kill switch, rollback, traffic limit, and recovery verification are tested
[ ] Incident response and evidence preservation completed in a tabletop exercise
[ ] Change classification, re-evaluation, and re-approval rules are active
Evidence / owner / gap:

PEOPLE AND USER READINESS
[ ] Users understand capability, limitations, uncertainty, and prohibited use
[ ] Reviewers and operators completed role-specific practice and assessment
[ ] Notices, evidence, correction, appeal, and human-contact routes work
[ ] Support, accessibility, multilingual, and affected-party needs are covered
[ ] Workload, incentives, and staffing do not undermine control effectiveness
[ ] Feedback and complaint channels route to accountable owners
Evidence / owner / gap:

STAGED EXPOSURE PLAN
Stage 1 scope, duration, success and stop conditions:
Stage 2 scope, duration, success and stop conditions:
Canary / shadow / pilot mechanics:
Fallback capacity at expected and peak volume:
Who may expand, pause, roll back, or terminate:
Post-launch verification window and evidence:

OPEN CONDITIONS AND WAIVERS (repeat)
Gap:
Consequence:
Compensating control:
Owner and due date:
Expiry and re-review trigger:
Residual-risk approver:

DECISION
[ ] LAUNCH within requested scope
[ ] LIMITED LAUNCH with explicit restrictions and conditions
[ ] BLOCK pending evidence or remediation
[ ] RETIRE / do not proceed
Decision rationale and decisive evidence:
Approved scope and prohibited expansion:
Conditions, owners, dates, and expiry:
Approvers and timestamps:

Decision Rules

Do not average away hard gates. A critical security path, unavailable rollback, unlawful or unauthorized data use, missing accountable owner, failure on a non-tradeable safety or rights floor, or inability to evaluate a consequential claim blocks launch. A bounded segment failure may justify a limited launch only when that segment can be reliably excluded and the exclusion is monitored.

A conditional launch is not a polite approval. Conditions need owners, due dates, enforced restrictions, expiry, and a consequence when unmet. Launch authority must remain separate enough from delivery incentives to challenge the evidence. An approver who cannot narrow the scope, delay the date, or demand a new candidate is witnessing the launch, not governing it.

Work the Decision Until the Scope Is Honest

The claims model passes aggregate recall, but urgent property claims fall below their minimum floor. The system can identify that type deterministically, so the proposal routes those claims to the existing manual process. On paper, the weak segment is excluded.

The load rehearsal changes the decision. The manual queue can absorb twelve additional claims an hour before breaching its response target; the nationwide peak would send thirty-one. A fallback that collapses at credible volume is not a control. During the same review, the team discovers that its evaluation names a provider alias that advanced after testing. The forum pauses the decision until the evaluated version is pinned and the evidence is reproduced.

The revised record authorizes a four-week pilot for two trained teams, caps traffic below the measured manual capacity, routes every urgent property claim to that queue, and assigns a responder to the queue-depth stop condition. Expansion requires a new load test and segment results from the pinned candidate. The nationwide launch is blocked; it has not been approved with cautious wording.

Rehearse the Decision Forum Under Pressure

Before approval, rehearse the meeting with one broken link in the evidence chain and a reason to ignore it. Let the candidate change after evaluation, make the fallback too small for peak load, remove the only person authorized to stop traffic, or reveal a segment failure on the promised launch morning. Give the forum the real release date and delivery pressure. Ask for a decision, not a list of future fixes.

Follow the proposed restriction through the running system. Can the team identify the excluded population before the model acts? Can monitoring prove the exclusion is holding? Does the fallback have enough capacity, and does its owner have authority to pause expansion? Then make the forum record what it will launch, what it will not launch, who can reverse the decision, and which evidence would permit a wider scope. If a material uncertainty produces only a note to monitor closely, the gate has failed.

A reviewer should be able to choose any checked item and reach the exact evidence, candidate, scope, owner, and consequence behind it. The record is ready when those links support a launch, limited-launch, block, or retirement decision without hiding a hard failure inside completion language.

Use this record with Online Experiments, Pilots, Shadow Mode, and Launch Decisions, the AI Evaluation Plan, and the AI Monitoring Plan.