At University of Michigan, I built MathMentorDB from 5.4 million tutoring chat messages and released it on HuggingFace. I use it as a real-data baseline for evaluating AI tutoring conversations and synthetic data.
I tested whether AI-generated tutoring conversations could pass for real ones and found they still fall short: AI tutors over-scaffold, while AI students get unstuck faster than real students. I’m developing a four-part test for when synthetic data can replace human data.
In the Geometry for Teachers project, I measured students’ growth and linked their scores to those of practicing teachers using IRT. I’ve also modeled tutoring conversations, evaluated AI graders, and taught statistics and mathematics.

