Benchmark Radar
AI BENCHMARK PROFILE

Int-Bench

General AIKnowledge & Reasoning

Int-Bench is a simulation-based evaluation for LLM intervention behavior in tutoring. It simulates a student solving problems across code debugging, mathematics, and brain teasers, with a teacher deciding whether, when, and how to intervene, measuring frequency, timing, and impact on task success and generalization.

Released
2026-07-23
Readiness
Paper only
Primary field
General AI

Why it matters

The evaluation gap is that AI assistants may over-assist, hindering learning. This benchmark aims to quantify intervention timing and content, providing a method to compare models on supportive versus answer-giving behavior, which is critical for educational AI.

Motivation

Large language models (LLMs) are increasingly used as tutors and thought partners, helping users reason through problems.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.