AI Systems Handbook / Chapter 47
Building an AI Operating Model
Design the roles, decision forums, shared capabilities, evidence flows, and escalation routes that carry AI work from intake through operation.
Preparing audio…
Audio edition
Building an AI Operating Model
Twelve Pilots, No Way Through
A company has twelve AI pilots. A central innovation team built four, business units bought five, and employees assembled three with general-purpose tools. Security reviews only systems that request production credentials. Legal learns about customer-facing uses during contract review. The platform team sees traffic but not business purpose. Every group is busy; nobody can list the portfolio or explain how a pilot becomes an operated product.
One of the twelve is a customer-service assistant. Three business units want it, each with different products and complaint rules. The innovation team owns the prototype. The service lines will receive its errors. The platform team will carry its traffic. When a unit asks who may launch it, the central AI council schedules a slide review for next month.
Another principles document will not resolve this confusion. The company needs an operating model: a repeatable allocation of outcome ownership, enabling capability, independent challenge, decision authority, evidence, and escalation across the AI lifecycle.
Replace the Org Chart With a Decision Route
The council begins by asking where AI expertise should report. That question is premature. Reporting lines cannot tell the service team who may accept the use case, approve customer data, set evaluation thresholds, authorize a limited launch, respond to a complaint, or retire the system.
Trace those decisions first. For each one, name a single accountable owner; the people who contribute; an independent challenger where the consequence requires one; the evidence needed; the owner’s authority limit; the route for escalation; and the record left behind. “The council provides oversight” closes none of these gaps.
The customer-service assistant reveals the first hard boundary. The innovation lead can prove that the assistant generates plausible answers, but cannot own the service outcome. Each service-line owner must define the customer problem, the existing baseline, affected people, acceptable errors, human workflow, adoption plan, and stop conditions. The owner may not delegate those consequences to the model team or the vendor.
This does not mean every business unit should build its own stack. A platform team can supply approved model access, identity, secrets, data connectors, retrieval, evaluation harnesses, observability, version records, deployment patterns, and cost controls. Its success is not the number of adopted components. It is whether teams can produce trustworthy evidence and operate systems with less reinvention.
Nor should “risk” become one universal approval queue. Security, privacy, legal, compliance, safety, accessibility, internal audit, and responsible-AI specialists ask different questions. Independent challenge tests the product owner’s assumptions, evidence, control design, and residual-risk reasoning; it does not quietly inherit the product decision. Executive governance sets risk appetite, restricted uses, materiality thresholds, investment priorities, and escalation authority. It intervenes when a decision exceeds delegated limits or when patterns across the portfolio demand a change—not to replay every feature review.
Choose Topology From the Tension
The company could centralize the entire program while expertise is scarce and the portfolio is small. That would make standards and capability easier to establish, but the center would soon become a queue and domain ownership would weaken. It could federate authority to strong business units. Decisions would sit closer to customers, but standards, platforms, and even the inventory could diverge.
The customer-service program needs both shared capability and local judgment, so the company chooses a hybrid. The center owns minimum controls, the inventory, reusable platform services, high-consequence assurance, cross-unit learning, and literacy resources. Each service line owns problem selection, workflow design, adoption, outcomes, routine operation, and retirement. The design works only because the boundary is explicit. “Hybrid” without a decision map is ambiguity with a modern label.
An AI center of excellence can help establish this arrangement, but its enduring job must be specific. It may incubate capabilities, teach teams, curate approved patterns, convene specialists, and make lessons portable. If it remains the owner of every product, it prevents the capability transfer that federation requires. If it becomes a ceremonial council, it adds a meeting without adding a route.
Let One Record Travel With the System
At intake, the service-line owner registers the assistant’s purpose, baseline, users, affected people, providers, initial risk tier, and accountable parties. That record becomes the spine for later work. Data and architecture choices attach to it. So do evaluation plans and results, security and privacy findings, the approved launch scope, accepted residual risk, versions, incidents, complaints, material changes, and eventual retirement evidence.
Specialists can have views suited to their work without asking the team to retype the same facts into unrelated forms. A provider-version change should notify every owner whose earlier decision depended on it. A complaint should link to the deployed version, known limitations, launch conditions, and the person able to restrict the service. An inventory refreshed only for an annual audit is an archive; this record must participate in operations.
The same spine prevents a common handoff failure. When the innovation team transfers the assistant, knowledge does not disappear into a presentation. The service line receives the evaluation design, control assumptions, unresolved findings, runbooks, fallback, and named experts. The platform team records the reusable pattern. Training uses the real decision route, not a general awareness deck. Searchable decisions, examples, and incident lessons allow the next team to reuse judgment as well as code.
Give Different Work Different Routes
The first service line wants a drafting assistant whose answer a trained agent must approve. It uses an established retrieval pattern and can fall back to search. That may fit a standard lane with known evidence and delegated launch authority.
The second wants to send answers directly to customers in several languages. External exposure, accessibility, harder-to-recover errors, and a new translation provider move it to an enhanced lane with additional evaluation and specialist challenge. The third proposes using complaint sentiment to change customer entitlements automatically. The consequence and unresolved legal questions put it on an exception route beyond the product owner’s authority.
Meanwhile, the innovation team may continue a bounded experiment, but the experiment has explicit data, users, permissions, duration, and an end date. Production credentials do not magically convert learning into permission. Moving from experiment to service requires a new decision.
These routes are not labels applied after the design is finished. Each specifies its entry conditions, evidence, reviewers, thresholds, decision owner, authority limit, expected response time, and exit. Fast review comes from prepared evidence and known routes. A single queue makes low-consequence work wait unnecessarily while specialists discover serious questions too late.
Make Every Forum Close Something
The monthly slide council becomes several smaller mechanisms. Intake decides whether a proposal should enter exploration and who owns the next evidence-producing step. A weekly design clinic lets domain, platform, and specialist teams resolve questions before commitments harden. The delegated launch authority closes decisions within each lane. Operational review examines incidents, drift, complaints, exceptions, and material changes. The executive forum sets the boundary and resolves only escalations or portfolio patterns above delegated authority.
Each forum publishes its required inputs, decision rights, quorum, response target, and durable output. The company watches time spent waiting, repeated findings, exceptions, late discoveries, and overturned decisions. Those signals show where authority is unclear, evidence is poor, or a shared capability is missing.
Two months after the first limited launch, customer complaints rise for one product line after a corpus update. The service owner restricts the assistant under delegated incident authority, the platform team identifies that its ingestion pattern did not preserve a required product boundary, and assurance reopens the affected evaluation. Executive approval is unnecessary for the immediate restriction. The executive forum later sees the pattern across two systems and funds a shared corpus-isolation control.
That episode completes the operating circuit. A local owner protects the outcome, shared capability makes correction reusable, independent challenge tests the repair, and portfolio authority acts on the systemic lesson.
AI Operating Model Blueprint
- Mandate: business outcomes, risk appetite, principles, covered systems, and executive sponsor.
- Topology: centralized, federated, or hybrid rationale; organizational boundaries and local variation.
- Decision map: lifecycle decision, accountable owner, authority, contributors, challenger, evidence, service target, record, escalation.
- Lifecycle lanes: standard, enhanced, exception, and experiment entry and exit criteria.
- Shared capabilities: platform services, approved components, evaluation, observability, documentation, support, and funding.
- Forums: purpose, inputs, cadence, quorum, decision rights, and performance measures.
- Evidence spine: inventory, identifiers, artifacts, versions, access, retention, and interoperability.
- Assurance: competence, independence, sampling, control tests, findings, closure, and audit route.
- Learning: incidents, user feedback, policy changes, reusable patterns, training, and model improvements.
Test the blueprint against a company with many business units. Choose one use case that three units want in different forms. Decide which lifecycle choices remain local, which capabilities and minimum controls are shared, and which facts must follow every version. Then introduce three disturbances: one unit wants to skip a control, a provider change invalidates an evaluation, and a production complaint requires immediate restriction. If the blueprint cannot identify who acts now, what evidence travels, and when authority escalates, it describes an organization but does not yet operate one.
Source Notes
- NIST AI RMF Govern describes policies, processes, roles, inventory, training, executive responsibility, ongoing monitoring, and lifecycle risk management; voluntary Framework 1.0 and its in-progress revision verified 2026-07-20.
- ISO/IEC 42001:2023 specifies requirements for establishing, implementing, maintaining, and continually improving an organizational AI management system; standard overview verified 2026-07-20.
- NIST AI RMF Playbook provides voluntary suggested actions for operationalizing the Framework; revision status verified 2026-07-20.
- See Governance as an Operating System, Change Management and Continuous Improvement, and AI Portfolio Management.
Continue reading
Full table of contents