Benchmark Radar
AI BENCHMARK PROFILE

HG-Bench

General AIMultimodal Perception

HG-Bench evaluates page-aware, two-level answer-region grounding: given multi-page handwritten homework images, models must localize complete answer regions and ordered step-level subregions, with question- and step-level boxes under a hierarchical constraint.

Released
2026-06-24
Readiness
Paper only
Primary field
General AI

Why it matters

Automated homework assessment needs both answer recognition and spatial grounding of reasoning steps. Prior benchmarks miss page-aware, multi-level localization; HG-Bench provides a reproducible protocol to measure this capability gap across models.

Motivation

Automated homework assessment depends not only on recognizing student answers, but also on accurately locating where each answer and each intermediate reasoning step appears in noisy, multi-page handwritten work.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.