Case Study 8: Education Tutor
Design an AI tutor that supports learning without replacing evidence, teacher judgment, age-appropriate safeguards, or learner agency.
The Homework Was Finished
Maya, a modeled seventh-grade learner in a middle-school mathematics pilot, submits every integer exercise on time. When the tutor asks her to solve -3 + 5, she requests a hint. It supplies a number-line explanation, then another. The second explanation reverses the direction of travel but still arrives at 2. Maya copies the method into three later answers. The tutor praises her progress, the homework system records completion, and the class dashboard turns green.
Two days later, without the tutor, Maya solves -3 + 5 as -8. She can repeat the tutor’s language but cannot say why the answer should be positive. The system made the assignment easier to finish and the misconception harder for the teacher to see.
That failure gives the pilot its governing rule: the tutor must preserve observable learner reasoning. Its job is to create an attempt, offer the smallest useful help, test what changed, and return authority to the learner and teacher. It is not an answer service, a diagnosis of the child, a substitute for instruction, or an open-ended companion.
Start With the Attempt the System Must Preserve
Before adding a model, the school studies the existing route through the lesson: teacher instruction, curriculum examples, peer discussion, office hours, and a deterministic practice tool that can check an answer without generating an explanation. It records where learners become stuck, how teachers recognize a misconception, how long useful feedback takes, and which learners cannot reliably reach human help because of schedule, language, accessibility, device, or connectivity constraints.
The proposed tutor earns a place only if it improves that route without weakening the strongest feasible non-AI alternative. For the integer unit, educators define the objective as reasoning about direction and magnitude, not merely producing a correct sum. They identify prerequisite ideas, permitted forms of help, common misconceptions, evidence of independent mastery, and the point at which a teacher should intervene. -3 + 5 is useful precisely because a correct answer can conceal broken reasoning.
The content service holds an approved, versioned curriculum corpus and the worked examples educators have reviewed. Retrieval carries source and curriculum-version identifiers into the interaction. Generated explanations remain distinct from quoted material. The tutor cannot invent grades, deadlines, school policy, or claims about a learner’s ability. When it lacks approved support for a question, it says so and offers the appropriate human route.
Make Help Change State, Not Just Add Words
The interaction begins with an attempt: a calculation, sketch, explanation, selected step, or description of where thinking stopped. If the learner cannot begin, the tutor retrieves a prerequisite rather than pretending that an empty answer reveals a misconception. This preserves an important distinction between missing prior knowledge, a mistaken rule, a language or interface barrier, and a request to avoid the work.
Help advances through a bounded ladder. The tutor may restate the task, ask a diagnostic question, recall a prerequisite, give one hint, offer a parallel example, or explain one step. Each move produces a new learner attempt before the next move. A direct solution is available only for activities in which educators have decided that seeing a complete solution serves the objective; the interface labels that change in mode instead of slipping from coaching into answer delivery.
For Maya’s first request, a useful response might ask, “Starting at negative three, which direction does adding a positive number move?” Her next action supplies evidence. If she moves left, the system has located the misconception without inferring a stable trait. If she moves right but counts the starting point as a step, it has found a different error. The same generic explanation should not follow both attempts.
The learning state is therefore small and inspectable: activity and curriculum version, learner attempts, help level, concepts encountered, verification result, and unresolved misconception flags. It expires according to the educational purpose. The system does not silently derive personality, emotion, disability, diligence, intelligence, or potential. A teacher may record a needed accommodation through the school’s governed process; the tutor does not diagnose one from conversation.
The interface makes uncertainty and source boundaries visible in age-appropriate language. Learners can ask for a simpler explanation, another representation, translation, accessibility support, or a teacher. The tutor avoids manipulative streaks, emotional dependency cues, open-ended persuasion, and language implying that it knows the learner better than trusted adults do.
Teachers see unresolved concepts, repeated help patterns, sampled interactions, and safety or content flags at a level useful for instruction. They do not receive a surveillance feed of every exploratory thought. They can inspect the source and policy version, review the learner’s actual attempt, correct an explanation, restrict a topic, and understand why escalation occurred. Analytics are prompts for attention, never diagnoses or grades.
The access design is part of the learning design. Essential instruction and teacher contact cannot depend on a high-end device, continuous connectivity, paid home access, or fluency in the interface’s default language. The pilot compares who can use each route, who receives slower or thinner help, and whether resources spent on the tutor reduce accessible human support.
Ask Whether Maya Can Solve the Next Problem
Evaluation begins before the model is allowed to coach a learner. Domain experts test mathematical correctness, curriculum alignment, age appropriateness, clarity, cultural and language coverage, accessibility, source fidelity, refusal behavior, and responses to prompt injection and unsafe requests. The set includes ordinary lessons and misconception-rich cases in which a plausible explanation can support a correct answer for the wrong reason.
The negative-number trace has ordered gates. First, the tutor must not introduce a mathematical error. Then it must select help permitted for that activity and preserve the learner’s next attempt. Only after those gates pass does the team ask whether the interaction improved understanding. A fluent explanation cannot compensate for wrong content, and a correct answer cannot compensate for a hidden misconception.
Maya sees a parallel problem during tutoring. Later she completes a new problem without assistance, explains the direction of movement, and applies the idea to a temperature change rather than another nearly identical number-line exercise. Those observations test immediate correction, independent performance, explanation, and transfer. A delayed task tests retention. If she still applies the reversed rule, the interaction failed even if every coached answer was correct.
The team compares outcomes with the baseline and studies distribution: device and connectivity constraints, language, accessibility need, prior attainment, age, class context, and frequency of use. It checks whether the system helps already-confident learners while creating dependency or confusion for others. Qualitative sessions ask learners to explain what they think the tutor knows and when they would seek human help.
Completion time, activity rate, and satisfaction remain operational signals. They are not mastery proxies. The team also watches direct-answer requests, escalation, repeated hints, independent-task performance, delayed retention, transfer, quality of explanation, and teacher-observed misconceptions. It examines whether frequent use improves later independence or merely makes assistance necessary.
Rehearse the Lesson That Goes Wrong
The first pilot covers a bounded subject, age range, curriculum version, and school context. Teachers receive training on capabilities, common errors, escalation, and how not to treat engagement analytics as objective measures of ability. Learners and caregivers receive understandable information about purpose, data, limitations, and support routes.
Before launch, the team replays Maya’s interaction and variations that change one condition at a time. She gives no initial attempt. She uses speech input that transcribes the minus sign incorrectly. She asks in another language. She loses connectivity after the hint. The curriculum has changed but the retrieval cache has not. She solves the coached item, fails the independent item, and returns for a final answer. A second learner enters a sensitive disclosure during the lesson. Each route must end in a known state with an owner.
The pilot then runs in shadow or closely supervised mode, with teacher review capacity measured as a real system constraint. Escalation is not credible if a flag waits until the lesson is over or if the teacher cannot see the attempt that caused it. The school tests what happens when several classes generate content flags at once, an accessibility route fails, or the model or curriculum version changes during a unit.
Release requires content and safety floors, evidence of learning beyond coached tasks, usable correction and escalation, bounded data use, accessible alternatives, and adequate teacher capacity. Stop conditions include a severe safety failure, a systematic curriculum error, ineffective escalation, unauthorized data use, material harm concentrated in a group, inaccessible essential support, or evidence that tutored learners perform worse independently.
Incident: A Plausible but Wrong Explanation
Maya’s teacher reports the reversed number-line instruction. The team preserves the protected evidence needed to investigate: the attempts and responses, approved source, curriculum context, retrieval result, model and policy versions, help level, verification result, and relevant prior turns. It blocks the affected explanation path, alerts instructors using that lesson where appropriate, and samples neighboring concepts and languages. Root-cause review separates a corpus error, stale retrieval, generation error, interface loss, prompt behavior, and a missing verification control.
Correcting the next response is only the beginning. The school identifies learners exposed to the faulty path within its authorized records, limits access to those who need to act, and gives teachers a plan for checking the concept without labeling the children. It decides whether remedial instruction is needed, adds the case and variants to regression evaluation, and verifies the repair on independent tasks. Restoration requires evidence that the content path, cache, monitoring, and teacher route work together; a stronger disclaimer does not restore the lesson.
The Learning and Safety Record
One compact record supports the pilot decision and later investigation. It identifies the learning purpose, age and subject boundary, non-AI baseline, approved curriculum versions, exact AI role, prohibited uses, permitted help ladder, independent evidence of mastery, and named educational owner. It documents the minimum learner state, access and retention, accessibility and language routes, content and safety evaluation, teacher workload, escalation response, access disparities, release floors, stop triggers, and the changes that require reevaluation.
For a particular interaction, the protected record contains only what authorized staff need to reconstruct the learning and safety path. It can show Maya’s initial attempt, the retrieved source, the wrong explanation, her retry, the independent-task failure, the teacher’s correction, the affected content versions, and the remedial response. It does not become a permanent character judgment about Maya.
Sensitive disclosures, signs of immediate danger, bullying, abuse, or requests outside instructional scope follow a reviewed, age-appropriate route to qualified human support. The route states what the tutor should say, what it should avoid claiming, which information is shared, with whom, how quickly, and what happens when the primary contact is unavailable. The tutor does not simulate professional care or discourage help from trusted people.
The benefits-triage case before this one showed how a better average can hide one household’s longer wait. Here a green completion dashboard hides a learner’s weaker understanding. The useful tutor does not remove all friction. It makes the right thinking visible, gives it carefully bounded support, and knows when the lesson belongs back with a human teacher.
See Human-Centered AI Design, Evaluating Generative AI, and AI Literacy and Training.
Continue reading
Full table of contents