Skip to content

AI Systems Handbook / Chapter 48

AI Portfolio Management

Fund AI work in stages using risk-adjusted value, feasibility, learning, reuse, benefits evidence, and explicit stop criteria.

The Portfolio of Permanent Pilots

An executive dashboard shows 37 AI initiatives. Twenty-nine are “green” because each team reported a successful demonstration. Six have production users, two measure a business outcome, and none has a documented stop condition. Shared evaluation, retrieval, and monitoring work is buried inside project budgets, so every new team rebuilds it. Leaders can count activity but cannot say what the portfolio has learned or whether it deserves another dollar.

This is an idea queue wearing portfolio language.

AI portfolio management allocates scarce attention, money, data, platform capacity, assurance, and change effort to staged evidence—not to enthusiasm—and stops work when the evidence no longer justifies continuation.

AI proposals enter an intake gate and move through explore, prove, scale, or stop zones. Cards are judged on value, feasibility, risk, learning, and reuse; redesign and kill-criteria paths prevent automatic progression.
A portfolio is a flow of evidence. Work earns greater commitment by reducing important uncertainty, while redesign and stop decisions preserve capacity for stronger opportunities.

The Meeting Has Ten Proposals and Four Places

The next quarterly forum can support four substantial pieces of work. Ten proposals arrive with projected savings, executive sponsors, and polished demonstrations. If the forum treats the score as a verdict, the proposals with the boldest estimates will win. Their estimates are also the least tested.

The forum first asks what commitment each proposal is seeking. One team needs two weeks to inspect data rights. Another has already tested a bounded workflow and needs a production-like proof. A third is asking for a year of integration, training, monitoring, and support on the strength of a demo. These are not comparable funding requests merely because each carries the word project.

A score is useful here, but only as a way to expose the next decision. The forum looks at five dimensions. Value begins with a named user or operational outcome and a measured baseline. Feasibility covers the complete system: data, integration, skills, evaluation, operating load, and dependencies. Risk asks who can be affected, how severely, how failure is detected and recovered, and who may accept what remains. Learning names the important uncertainty that the next stage can actually resolve. Reuse asks whether the work creates or consumes a shared capability and who will maintain it.

Strategic alignment constrains all five; it cannot substitute for them. Calling a proposal strategic does not supply data rights, an operating owner, a credible evaluation, or a user willing to change the workflow.

Price the Next Question, Not the Imagined Finish

One proposal would summarize customer-service cases. Its sponsor forecasts thousands of saved hours by multiplying handling time by case volume. The portfolio record changes the view. Agents already skim many short cases, the difficult cases require source checking, and summaries add review work. The forecast might prove true, but it is not yet evidence.

The forum records three things separately. Expected contribution includes the size, reach, duration, strategic fit, and confidence of the outcome. Total commitment includes engineering, data, integration, evaluation, assurance, human review, training, operation, vendor, and eventual exit costs. Downside exposure includes affected people, severity, likelihood ranges, recoverability, concentration, and uncertainty.

These views should not be collapsed into a spurious monetary answer. A one-to-five rating can support comparison if the evidence and confidence remain visible. Some constraints are not tradable at all. Unlawful data access, intolerable harm, no accountable owner, or no credible way to evaluate the system cannot be averaged away by a large revenue estimate. Such a proposal must be redesigned, sent to the proper exception authority, or rejected.

For the summarizer, the cheapest useful commitment is a bounded proof. The team will compare assisted and unassisted handling on representative cases, including long and multilingual cases, and count source-checking, correction, and review time. It has six weeks, a named operational owner, and a decision date. The portfolio has bought an answer, not promised a product.

Make Each Stage Earn the Next One

An initiative in explore should clarify the problem, affected people, current workflow, non-AI alternatives, data rights, and largest uncertainties. Interviews, workflow observation, small data analysis, and paper prototypes may be enough. Exploration ends with a testable use case or a reason to stop.

In prove, the team tests value, feasibility, and controls under bounded but credible conditions. It compares against the baseline, uses representative data and failure cases, estimates operating load, and involves real users safely. A prototype that demonstrates model capability without testing workflow value has not earned a production commitment.

Scale funds a service, not a larger demo. Reliability, security, accessibility, change management, support, monitoring, assurance, benefits tracking, and exit belong in the commitment. Expansion should remain progressive and reversible where consequences require it.

Stop is also a stage decision. Work closes when its criterion is met, its opportunity cost no longer makes sense, or a non-AI change solves the problem better. The team preserves findings, components, data decisions, and user insight. Returning capacity before a weak idea becomes an operated dependency is a portfolio gain.

The summarizer therefore receives its stop conditions before anyone sees the new demo. It will not advance if it fails to improve accepted resolution time after review; if serious errors exceed the agreed threshold for any included language; if correction work erases the claimed benefit; or if agents cannot verify summaries within the real queue. A recorded decision may revise a criterion when new evidence changes the premise. Quietly moving it after disappointing results converts a test into theater.

Four Places Do Not Mean Four Projects

The portfolio view changes the meeting’s allocation. Two proposals cannot establish data rights or an accountable owner, so they leave with explicit reopening conditions. An employee-scoring proposal is redirected toward redesigning the underlying management workflow. Three uncertain but plausible uses receive small exploration commitments. Document triage earns a bounded proof because it has a credible baseline and evaluation route. The summarizer is held while its latest results are recalculated with rework included.

The fourth substantial allocation does not belong to any one use case. Four teams need the same evaluation infrastructure, and each has hidden a smaller version inside its estimate. The forum funds that shared capability directly, with a product owner, intended consumers, service expectations, maintenance cost, and adoption evidence.

This is why portfolio management cannot end with card ranking. Individually attractive projects can create a brittle whole. The forum inspects dependence on one provider, dataset, platform, or scarce reviewer group; the balance between early learning and expensive scaled services; the distribution of benefits, review labor, and harm; and whether shared security, data quality, observability, accessibility, and literacy are funded visibly. A portfolio of ten high-scoring cards can still exceed the organization’s capacity to assure or operate them.

Benefits Must Survive the Full Workflow

Six weeks later, the summarizer returns. Generated summaries look good, and raw drafting time fell. Once source checking and correction are counted, ordinary cases show little change. Long cases improve substantially. Multilingual cases improve on average but produce an unacceptable severe-error rate in one language.

The decision is neither “AI works” nor “the pilot failed.” The forum stops the broad proposal, preserves the evaluation set and review findings, and offers a narrower proof for long cases in the languages that met the threshold. It also sends the language failure to the shared evaluation owner because other initiatives rely on the same provider.

That judgment was possible because benefits tracking began before the pilot. The baseline included time, quality, cost, accessibility, risk, customer effort, and worker load. The result included review, rework, incidents, support, training, vendor expense, and displaced work. Output counts, licenses, prompt volume, and training attendance remained activity signals rather than benefits. Segment results prevented an average improvement from concealing who received worse service.

The dashboard now shows fewer permanent pilots. It also shows capacity returned, assumptions disproved, components reused, and decisions reopened. Those are signs of a portfolio learning how to invest.

AI Portfolio Scoring Matrix

For each proposal, record:

  • Use case and owner: user, affected parties, workflow, baseline, sponsor, product owner, operational owner.
  • Value: intended outcome, measure, magnitude range, time horizon, strategic fit, confidence.
  • Feasibility: data, model or non-AI pattern, integration, skills, operating load, dependencies, confidence.
  • Risk: impact tier, principal harms, recoverability, control maturity, residual-risk authority.
  • Learning: largest uncertainty, cheapest credible test, threshold, artifact, decision date.
  • Reuse: shared data, evaluation, platform, policy, or workflow asset; owner and consumers.
  • Commitment: stage, people, funding, capacity, duration, and opportunity cost.
  • Decision: explore, prove, scale, redesign, hold, stop; rationale and expiry.

Take ten proposals from a real or plausible organization and assume the next quarter can support only four substantial commitments. Do not rank the proposals first. For each one, identify the largest decision-relevant uncertainty and the cheapest credible evidence that could reduce it. Then allocate capacity across exploration, proof, scale, and shared capability. Name what you would stop or hold, what evidence could reopen it, which constraint cannot be averaged away, and where the resulting portfolio remains dangerously concentrated.

Source Notes