Benchmark Radar
AI BENCHMARK PROFILE

EduClaw-Bench

General AIKnowledge & Reasoning

EduClaw-Bench evaluates pedagogical LLM agents in a simulated 30-day tutoring relationship with a knowledge-tracing-based simulated learner, scoring learning gain, responsiveness, helpfulness, and curriculum design across 55 scenarios.

Released
2026-08-04
Readiness
Paper only
Primary field
General AI

Why it matters

Existing benchmarks focus on single-turn tasks, leaving long-horizon tutoring unmeasured. This benchmark provides a standardized way to assess sustained pedagogical interaction, helping developers and educators choose agents that maintain effective teaching over time.

Motivation

Large language models (LLMs) power educational applications from tutoring to essay scoring, but each is a point solution to a single task, and only recently have these point solutions been integrated into agents operating over a learning management system (LMS).

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.