Benchmark Radar
AI BENCHMARK PROFILE

CSTutorBench

General AICoding & Software Engineering

CSTutorBench evaluates small language models as tutors in VEX VR block-based programming, with 17 scenario-based questions scored via a rubric and LLM-as-judge.

Released
2026-07-06
Readiness
Paper only
Primary field
General AI

Why it matters

The benchmark addresses the gap in evaluating SLMs for block-based programming tutoring, where models often lack domain-specific training.

Motivation

Large language models are increasingly explored as AI tutors, yet deploying them in K-12 settings raises concerns around privacy, cost, and reliance on proprietary models.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.