← Back to Research

LLM Pedagogical Behavior in AI Tutoring Interactions

Lee, S., J. Baek, J. Park, & D. Shin
Findings of the Association for Computational Linguistics: EMNLP, 2026
Forthcoming

Abstract

Students increasingly use LLMs as tutors for coursework and problem solving. Little is known about the level of assistance LLMs provide when students use them as tutors in authentic learning interactions. This matters because tutoring responses can differ substantially in how directly they help students complete a task. We operationalize this dimension as scaffolding level and develop a five-level scale, validated against human annotations, that characterizes responses according to the degree of direct assistance they provide.

We apply the scale to 14,637 LLM responses from 203 students in a university AI course. Responses are overwhelmingly concentrated at high levels of assistance, with more than 95% classified as either Explaining or Solving. Scaffolding level is systematically associated with students' subsequent conversational behavior, but provides little additional predictive information about performance on three subsequent exams beyond prior achievement and dialogue behavior. These findings provide an empirical baseline for LLM assistance in tutoring interactions and a measurement framework for evaluating how alternative tutoring designs change that assistance.

Approach

A scale for how much help a response gives

  • Five levels ordered by how directly a response completes the task for the student
  • Validated against human annotations rather than assumed
  • Applied at the level of individual tutoring responses, not whole sessions

Authentic classroom data

  • 14,637 LLM responses drawn from real coursework interactions
  • 203 students in a university AI course
  • Three subsequent exams available as downstream outcomes

What We Find

📈

Assistance runs high

Over 95% of responses fall into the two most direct levels, Explaining or Solving.

💬

It shapes the dialogue

Scaffolding level is systematically associated with what students do next in the conversation.

📝

But not exam scores

Beyond prior achievement and dialogue behavior, it adds little in predicting later exam performance.

📏

A baseline to build on

The scale gives a way to evaluate whether alternative tutoring designs actually change assistance.

Keywords

AI Tutoring Scaffolding Human-AI Interaction LLM Measurement Learning Analytics