TutorMoments: Do AI tutors know when to help and when to hold back?

2026-08-20 · Hugging Face

TutorMoments: Do AI tutors know when to help and when to hold back?

Introduction

The Allen Institute for AI today introduced a preview of TutorMoments, a framework designed to evaluate whether state-of-the-art large language models can successfully navigate one of education’s hardest trade-offs: when to step in and help a student, and when to hold back to let the student do more of the intellectual work themselves.

TutorMoments is a replay-based evaluation grounded in real one-on-one math tutoring sessions. Experienced math teachers reviewed transcripts collected from a U.S. tutoring program and flagged moments where the tutor faced a genuine choice between making the problem easier to help the student get started or pushing the student to perform more of the reasoning independently.

At each flagged decision point, the transcript is handed to an LLM, which assumes the role of tutor and continues the session for five turns with another LLM acting as the student. Researchers then analyze the model’s choices.

Key Findings

When told only to "tutor well," current models tend to over-help. They provide excessive support and rarely push students toward deeper thinking. Explicitly spelling out the help-versus-hold-back trade-off in the system prompt improves performance. However, it does not close the gap to human tutors, who consistently adapt to the specific moment. LLMs also differ widely in how reliably they make these pedagogical decisions.

In keeping with its commitment to open research, the team is releasing:

  • A dataset of 462 de-identified tutoring transcripts (TutorMoments-Preview)
  • The code for the replay evaluation pipeline
  • Model-generated tutor replays for all evaluated key moments

These resources aim to give educators, researchers, and AI tutor developers a sharper tool for assessing how models handle the nuanced decisions that matter most in teaching.

What Makes a Good Tutor?

A skilled math tutor rarely jumps straight to the solution. Instead, they often respond with a question such as "What do you know about what the problem is asking?" This is not unhelpfulness. Strong teaching requires diagnosing what a student already understands and providing the right level of support at the right moment.

Immediately volunteering full support can rob students of the productive struggle — the effortful, sometimes frustrating work that research consistently links to stronger, longer-lasting learning. Sometimes scaffolding is necessary; other times, the best move is to push the student to explain a correct answer themselves.

Language models, trained to be maximally helpful, tend to do the hard work for the user: explaining concepts, laying out steps, and guiding directly to the answer. In tutoring, this can short-circuit the very process that produces deep understanding.

Limitations of Existing Benchmarks

Most current benchmarks for LLM tutors fail to capture this tension. They typically reward fixed behaviors — such as never revealing the answer or always giving a hint — without considering whether that behavior was appropriate for the student’s actual state of understanding in that specific moment.

Good tutoring is not a uniform policy. It is a contextual judgment call: what does *this* student need *right now* on *this* problem?

How TutorMoments Works

The framework is built on authentic data. The released TutorMoments-Preview dataset contains 462 de-identified, text-only transcripts of real math tutoring sessions with U.S. students in grades 2–7. It includes more than 1,500 teacher-annotated key moments and thousands of free-text annotations from 27 experienced U.S. math teachers.

All transcripts came from a high-dosage tutoring program serving students who largely attend Title I schools. Data was shared under a research agreement and underwent rigorous de-identification.

Teachers identified decision points where the tutor had to weigh scaffolding (making the problem more accessible) against pushing for rigor (encouraging harder independent thinking). The evaluation pauses the transcript at these points and lets the LLM tutor for five turns with a simulated student. Each continuation is called a replay.

An LLM-based scoring pipeline evaluates each replay on three axes:

  • Did it scaffold when support was needed?
  • Did it push for rigor when the student was ready?
  • Did it avoid over-scaffolding?

Ground truth is determined by majority vote among multiple teacher annotators. A validated classifier then judges whether the model’s action matched the moment’s requirement.

Preliminary Results

Seven LLMs were tested with two prompts: a plain "tutor well" prompt and an evaluation-aware prompt that explicitly describes the help/hold-back trade-off.

Models consistently over-helped under minimal guidance. The explicit prompt improved results, but a substantial gap remains compared to human tutors, and reliability varies significantly across models.

The open release of data, code, and replays is intended to help the field build AI tutors that truly adapt to students rather than doing the work for them.

Source