AI Systems Handbook / Chapter 49
AI Literacy, Training, and Change Management
Build role-specific competence through realistic rehearsal, supported workflow change, and observation of safe behavior in practice.
Preparing audio…
Audio edition
AI Literacy, Training, and Change Management
Everyone Passed the Course
A customer-service team completes mandatory AI training on Monday. Everyone passes the quiz on hallucinations, privacy, and human oversight. On Tuesday the queue surges. Agents paste restricted customer details into an unapproved assistant because the approved one is slow. Reviewers accept fluent drafts without opening their sources. Supervisors reward handling time and never count correction.
The course described the risks correctly. The workplace taught a stronger lesson.
AI literacy is the competence to make the decisions a role actually owns: to recognize when AI is involved, work within its limits, verify what matters, recover from failure, and escalate without penalty. Training succeeds only when the surrounding workflow makes those actions possible.
Begin With the Changed Work
The service program is about to add an assistant that retrieves account guidance and drafts replies. A feature tour would begin with the prompt box. The training designer instead sits with agents before the workflow hardens.
The difficult work is not typing a reply. It is recognizing when account history is incomplete, when policy admits an exception, when a distressed customer needs a human judgment rather than a polished answer, and when the queue is too busy for careful source checking. Agents also point out a proposed telemetry field that could be used to score individuals even though it was introduced to study the tool. These observations change the pilot: sensitive cases remain outside scope, the source panel becomes easier to open, review time enters the staffing plan, and the telemetry purpose and access are narrowed.
Participation is useful because workers know where the process already bends. It is also a question of standing. An agent may be the assistant’s user, its reviewer, the source of domain knowledge, and the person whose pace is newly measured. Describing that agent only as a “human control” misses most of the change.
Close the participation loop before launch. Record what changed, what did not, why, who decided, and how a concern can be reopened. Consultation that disappears into a workshop summary trains people not to bother next time.
Teach a Decision, Not a Population
The same system creates different obligations. Agents decide when drafting is allowed, what data may enter, whether the cited guidance supports the reply, and when to abandon the draft. Reviewers must catch subtle defects without rejecting every unfamiliar answer. Supervisors set queue expectations, protect escalation, and decide when operating conditions require fallback. Builders reproduce failures, bind evaluations to versions, constrain permissions, and make controls operable. The service owner chooses thresholds, funds review capacity, watches benefits and harms, and can restrict the system. Assurance staff need enough independence and evidence access to challenge those choices. Executives must know which questions to ask and must fund the correction implied by an unwelcome answer.
Calling all of this “AI awareness” conceals the curriculum. For each role, name the consequential decisions, the knowledge needed to make them, the action and record expected, the boundary of authority, and the route out when the case exceeds that authority. A role is ready when it can perform those decisions in credible conditions—not when it can repeat a vocabulary.
This also keeps training durable as vendors and interfaces change. The learner should understand the complete system: the use case and non-AI baseline; what the model contributes; what retrieval, rules, software, and people contribute; which inputs and purposes are allowed; characteristic failures and operating limits; the required depth of verification; and the routes for correction, complaint, incident response, fallback, and pause. A button may move. The decision remains.
Let the First Rehearsal Fail Usefully
Agents begin with ordinary cases and real artifacts: source passages, draft replies, account context, review screens, and escalation forms. They first explain the system boundary and their authority. Then they compare plausible replies and locate missing evidence. Only after that do they complete a case from intake through record keeping.
The first failure is revealing. A draft cites an obsolete refund rule in fluent, reassuring language. Several agents notice that the answer feels wrong but cannot find a route between silently correcting it and filing a severe incident. The exercise has found an operating defect, not merely a gap in memory. The owner adds a lightweight content-report route, names who receives it, and explains when the system should be restricted.
Later rehearsals introduce missing evidence, a prompt-injection attempt inside a customer attachment, an outage, a harmful reply, and a mistaken action that must be repaired. Learners must choose a route, preserve the useful evidence, communicate urgency, and stop when authorized. Time pressure and representative queue load come last; applied too early, they mostly measure interface familiarity and reading speed.
Reviewer practice needs equally careful resistance. Seed errors that polished prose can hide: the wrong customer, an unsupported concession, an outdated source, a request to use a forbidden tool, and a case outside the pilot’s operating limit. Count serious misses, but also count needless rejection. A reviewer who blocks everything has not learned calibrated oversight.
Use relevant languages, accessible materials, and alternate demonstrations that preserve the actual job requirement. Competence should not be confused with one narrow way of showing it.
Tuesday Is Part of the Curriculum
The pilot opens to a small cohort with staffed fallback, office hours, and protected incident reporting. Then the queue spikes.
Source inspection falls sharply. The tempting conclusion is that agents ignored their training. Observation shows something else: the source panel requires an extra navigation step, service targets did not change, and corrections count against individual handling time. Supervisors have been taught to encourage verification while being rewarded for throughput.
The response is not a refresher course. The team makes sources visible beside the draft, reduces the cohort’s queue target, excludes good-faith overrides from performance penalties, and gives supervisors a threshold for returning to the non-AI route. Where restricted data is repeatedly pasted into another tool, it also makes the approved path usable and enforces the boundary. Training cannot compensate for excessive workload, unusable controls, unclear authority, missing fallback, or retaliation.
This is the work of change management: product design, staffing, incentives, support, communication, measurement, and participation moving with the technical release. The curriculum includes the environment because the environment is where the real examination occurs.
Observe Without Building a Surveillance Program
Attendance and quiz scores show that instruction occurred. They do not show whether the new workflow is safe or useful. The service owner compares rehearsal performance with production evidence: whether sources are opened when required, when agents abstain or override, the quality of escalation and recovery, severe misses and false alarms, privacy near misses, rework, queue load, customer outcomes, and worker feedback. Results are segmented where an average could conceal a language, case type, or customer group receiving worse service.
Raw adoption is a poor success measure. Refusing the assistant on an out-of-scope case is competent use. Returning to the old workflow during an outage is recovery, not resistance. A decline in escalations may mean that the system improved, or that reporting became costly. Measures need interpretation beside the workflow that produced them.
The program collects only what it needs for learning, safety, and the declared outcome. It states who can see the records, how long they remain, whether they affect employment decisions, and how a worker can challenge them. Coaching is separated from discipline where appropriate. Trust does not mean enthusiasm or agreement with leadership. It is an informed willingness to rely on the system within limits, strengthened when the institution responds competently to bad news.
Two weeks into the pilot, raw drafting time is down but correction and source-checking erase most of the gain. Complex cases improve; routine cases do not. Agents report that fallback works but the escalation queue is understaffed. The program narrows the use case, funds the queue, and repeats the affected rehearsals. Expansion waits.
Nobody failed a generic literacy test. The organization learned which competence, capacity, and authority this particular change requires.
AI Training Plan
For a specific AI-enabled change, record:
- The changed work: the workflow before and after, affected people, intended benefit, new burden and risk, operating limits, and non-AI route.
- Role decisions: what each role must recognize, decide, do, record, refuse, recover, and escalate; where its authority ends.
- Learning experience: explanation, real artifacts, worked examples, practice, failure rehearsal, feedback, accessibility, language, remediation, and refresh triggers.
- Supporting environment: approved tools, embedded guidance, staffing, incentives, review capacity, fallback, support, and protected reporting.
- Evidence of competence: representative scenarios, acceptance criteria, observation in practice, privacy limits, retention after time, and expiry after material change.
- Participation and communication: what remains uncertain, how telemetry will be used, which decisions remain open, how concerns receive answers, and who owns improvement.
Now design the plan for the service assistant. Give support agents, reviewers, supervisors, builders, the service owner, executives, and assurance staff different objectives. Include one ambiguous case and one operational disturbance for each role. Then assume the queue doubles, the model changes without notice, and workers learn that telemetry may affect performance reviews. Decide what must change in the workflow, incentives, communication, and competence evidence before the pilot may continue.
Source Notes
- NIST AI RMF Govern includes a risk-aware culture, clear roles, adequate resources, and training; voluntary framework verified 2026-07-20.
- OECD.AI, A Socio-technical Approach to AI Literacy describes understanding AI, critical evaluation, responsible use, and human oversight in organizational context; quick guide verified 2026-07-20.
- People + AI Guidebook provides human-centered guidance for user needs, mental models, explainability, feedback, and graceful failure; practitioner resource verified 2026-07-20.
- See Human Oversight, Escalation, and Accountability, Incident Response for AI Systems, and Building an AI Operating Model.
Continue reading
Full table of contents