AI BENCHMARK PROFILE
CSTutorBench
CSTutorBench evaluates small language models as tutors in VEX VR block-based programming, with 17 scenario-based questions scored via a rubric and LLM-as-judge.
- Released
- 2026-07-06
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
The benchmark addresses the gap in evaluating SLMs for block-based programming tutoring, where models often lack domain-specific training.
Motivation
Large language models are increasingly explored as AI tutors, yet deploying them in K-12 settings raises concerns around privacy, cost, and reliance on proprietary models.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.