AI Systems Handbook / Chapter 10
Business Case, Cost, and Value Realization
Build an AI business case from measured workflow outcomes, full lifecycle costs, expected risk loss, a credible baseline, and explicit scale or stop decisions.
Preparing audio…
Audio edition
Business Case, Cost, and Value Realization
The Fraction-of-a-Cent Business Case
A support team demonstrates an assistant that drafts replies for a fraction of a cent per model call. Finance multiplies that price by monthly conversations, subtracts it from agent labor, and predicts substantial savings.
The arithmetic prices a model invocation. The proposed product has a larger job: find the right policy under the agent’s permissions, distinguish current guidance from an obsolete article, cite its evidence, abstain on exceptions, survive malicious customer text, fit the case-management system, and leave the agent enough time to verify the answer. Production therefore adds document cleanup, retrieval, evaluation cases, security review, integration, human verification, monitoring, vendor support, incident response, and re-evaluation after model changes.
Even the benefit has been misidentified. A draft does not save money merely by appearing quickly. It creates value only if the case reaches an acceptable outcome with less total effort or better service—and if the organization can use the released capacity.
The business case can be stated in one line:
net value = measured workflow benefit
- full lifecycle cost
- expected risk loss
relative to the best credible alternative
The hard work lies in making each term refer to the same workflow, population, period, and action boundary.
Start With the Case That Must Be Resolved
The support organization handles 20,000 cases a month. A resolved case takes a median of twelve agent minutes. Reopened cases, incorrect answers, customer effort, and escalations are measured separately. Those facts provide a baseline; “agent productivity” does not.
The team names a narrower value hypothesis:
For two high-volume, low-consequence policy categories, reduce agent minutes per resolved case without increasing unsupported claims, repeat contact, correction time, customer complaints, or unequal performance by language.
This statement identifies the beneficiary, unit, scope, and guardrails. It also exposes possible conflicts. Faster initial handling can create more repeat contact. A concise answer can omit an exception. A tool that helps experienced agents may slow new agents or screen-reader users. Average improvement can conceal a worse result for one language queue.
AI value can come from revenue, avoided cost, faster cycle time, improved quality, risk reduction, accessibility, personalization, resilience, or learning that enables a later decision. But every claimed lever needs a causal path. If drafts become shorter handling time, what must agents do differently? If shorter handling time becomes lower cost, will schedules, backlog, service hours, or case volume actually change? If no operational decision follows, the time saving is capacity, not yet cash.
Make the Alternative Strong Enough to Win
The relevant comparison is not assistant versus today’s neglected workflow. Before evaluating the assistant, the support team improves templates and search. That simpler option reduces median handling time by 1.8 minutes, requires little new review, and makes obsolete policy pages easier to find.
This is the counterfactual the AI proposal must beat.
Evidence strengthens as it moves from baseline observation through historical replay, usability and time-and-motion studies, shadow operation, human-reviewed pilots, controlled comparisons, and sustained production outcomes. The exact design depends on the workflow, but the principle does not: compare options under the same outcome and guardrail measures.
Do not multiply a laboratory time saving by every future task. Adoption, exceptions, correction, learning curves, queue constraints, demand changes, and selection into the pilot all weaken that projection. Record a range and the assumption responsible for each end of it.
Price the System Beneath the Demo
Follow the system from discovery to retirement. Early costs include workflow research, data cleanup and licensing, evaluation sets, application and interface work, identity and permission enforcement, security and privacy review, migration, and training. Operation adds inference, storage, network and latency capacity, human verification, quality sampling, escalation, support, monitoring, incident response, and vendor management. Material model, prompt, corpus, policy, or provider changes can trigger new evaluation and approval. Exit requires export, replacement, archival, deletion, contract termination, and user transition.
Classify these as fixed, variable, or step costs. Evaluation design may be fixed for one release but recur after each material change. Inference varies with traffic. Human review varies until queue volume requires another staffed shift. Vendor minimums, reserved capacity, regional hosting, and assurance obligations create steps that an average cost curve can hide.
The prototype-to-production cliff is steepest when the design needs permission-aware data, expert review of rare severe errors, grounded and cited output, consequential tool use, several regions or languages, low latency at peak traffic, appeal and correction, or continuous monitoring. These are not miscellaneous overhead around the model. They are part of the product being valued.
Calculate a Resolved Case, Not a Token
For the proposed copilot, the useful unit is a resolved case:
cost per resolved case
= model, retrieval, and infrastructure
+ verification and correction
+ quality sampling and escalation
+ allocated monitoring, support, and incident capacity
+ allocated fixed build, assurance, change, and exit cost
Suppose a planning model uses these explicitly provisional assumptions:
- 20,000 eligible cases per month;
- a loaded agent cost of $0.75 per minute;
- five minutes of raw drafting time removed, but 1.8 minutes added for evidence review and correction;
- $0.22 per case for model, retrieval, and infrastructure;
- $7,000 a month for monitoring, sampling, support, and escalation capacity;
- $10,000 a month as the planning-period allocation of fixed build and assurance cost.
The net 3.2 minutes released is worth at most $2.40 per case, or $48,000 a month, and only if the organization can turn that capacity into a real outcome. Variable model and infrastructure cost $4,400. The allocated operating and fixed costs add $17,000. The modeled monthly contribution is therefore $26,600 before risk loss.
Now compare the improved-template option on the same basis. Its 1.8-minute reduction is worth at most $27,000 a month. If ongoing search and content maintenance cost $1,000, its modeled contribution is $26,000 after launch. The AI option’s impressive raw time saving has become a $600 advantage inside assumptions far less certain than $600.
The calculation has done useful work: it has not approved the copilot. It has found the evidence that can change the decision. The pilot must measure actual verification time, repeat contact, unsupported claims, adoption, language-specific performance, and whether released capacity reduces backlog or cost. It must also test the expensive tail. Long contexts, repeated retrieval, retries, rare escalations, and peak traffic can dominate an average.
Tokens and calls remain useful engineering drivers. They are not the business unit.
Keep Risk in the Decision
The previous chapter routed the support copilot according to consequence. The business case now prices the residual exposure without pretending every harm is money.
For monetary planning, use ranges across ordinary, severe, and tail scenarios. Estimate plausible frequency, consequence exposure, and uncertainty about control effectiveness. Include remediation, fraud, downtime, rework, contract penalties, incident response, and opportunity loss where they belong. A precise-looking expected value built from weak inputs is less honest than a range with named assumptions.
Rights, dignity, safety, trust, and unequal burden may resist monetary conversion. Preserve them as guardrails and stop conditions. If the direct-send design can issue an unsupported eligibility answer before meaningful review, a large labor saving does not purchase permission to cross that boundary. Controls cost money, but they may be what makes a narrower use governable.
In the support comparison, direct automated response appears to remove the most labor. It is still rejected because the current evidence cannot bound severe errors or provide meaningful review before customer impact. The team keeps templates as fallback and considers the copilot only for two categories where errors are visible and recoverable.
Choose an Operating Commitment
Build, buy, and partner are not three prices for the same box.
Building can be justified when specialized workflow integration, control, data boundaries, or evaluation create genuine differentiation. The comparison must include the enduring staff and infrastructure needed to own those capabilities. Buying fits a commodity-like capability only when the vendor can meet the required evidence, security, data-use, change-notice, service, and exit conditions; integration and oversight remain the buyer’s work. A partnership may supply domain expertise, distribution, data, or operating capacity, but it must assign accountability, intellectual property, incident coordination, evaluation rights, and termination.
For any vendor option, examine model and policy changes, submitted-data use, retention and subprocessors, evaluation access, audit evidence, portability, rate limits, latency, regional availability, service commitments, incident notification, and exit costs. A low license fee paired with weak evaluation rights or expensive migration is not a low-cost option.
Declining or deferring belongs in the comparison. So does improving the baseline. Avoiding a negative-value project is value realization.
The support team therefore authorizes a twelve-week human-reviewed pilot, not a general launch. It retains improved templates and search as both comparator and fallback. Expansion requires sustained improvement in resolved-case time and first-contact resolution, severe-error rates below the predeclared threshold by language, review load within staffed capacity, and positive net value after operating and risk costs. Failure on a hard guardrail stops the pilot even if average handling time improves.
Turn the Forecast Into an Owned Decision
A business case decays unless someone compares forecast with observation. Keep a benefits record that names:
- the workflow problem, beneficiary, baseline, and best simpler option;
- the causal path from system output to benefit;
- primary outcome, guardrails, distributional measures, and evidence method;
- adoption and exposure denominators;
- fixed, variable, and step costs, including change and exit;
- risk scenarios and non-monetary constraints;
- owners for product, finance, data, technology, risk, and operations;
- pilot, expansion, pause, provider-switch, and retirement rules;
- observed value, uncertainty, and reasons for variance from forecast.
Review leading evidence such as adoption, task success, correction, latency, reviewer load, and unit cost beside lagging outcomes such as repeat contact, retention, loss avoided, complaints, incidents, and workforce effects. Provider pricing, traffic mix, model behavior, policy, and user workflow can turn a valuable system negative after launch. The benefits owner needs authority to narrow, pause, or retire it.
For a new proposal, write the business case in this order:
- State the current outcome and the strongest credible non-AI baseline.
- Name the workflow unit, beneficiary, value mechanism, and guardrails.
- Design the comparison that could disprove the benefit hypothesis.
- Estimate lifecycle cost by fixed, variable, and step behavior.
- Calculate unit economics across typical, peak, and exception cases.
- Add risk scenarios while preserving hard constraints.
- Compare build, buy, partner, baseline improvement, defer, and decline under the same requirements.
- Set pilot, scale, pause, and retirement rules, with an owner for each decision.
Write assumptions so they can be tested. “Agents will adopt it” becomes adoption and correction hypotheses. “Costs will fall with scale” becomes a curve that includes review load and capacity steps. “The model is cheap” becomes irrelevant unless the resolved case is cheaper or better.
The next feasibility gate is data. A use case can survive risk triage and produce promising economics yet still fail if its evidence is unavailable, unrepresentative, stale, or unusable under its rights and retention conditions.
Source Notes
- Google Rules of Machine Learning, guidance on metrics, simple baselines, observable objectives, and launch decisions across multiple product measures; verified 2026-07-10.
- NIST AI RMF Core, continuous mapping, measurement, prioritization, monitoring, and management of AI risk across the lifecycle; verified 2026-07-10.
Continue reading
Full table of contents