Int-Bench
Int-Bench is a simulation-based evaluation for LLM intervention behavior in tutoring. It simulates a student solving problems across code debugging, mathematics, and brain teasers, with a teacher deciding whether, when, and how to intervene, measuring frequency, timing, and impact on task success and generalization.
- Released
- 2026-07-23
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
The evaluation gap is that AI assistants may over-assist, hindering learning. This benchmark aims to quantify intervention timing and content, providing a method to compare models on supportive versus answer-giving behavior, which is critical for educational AI.
Motivation
Large language models (LLMs) are increasingly used as tutors and thought partners, helping users reason through problems.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.