HLL
HLL is a benchmark that evaluates multimodal agents on interactive CAPTCHA verification in a closed-loop GUI environment, covering diverse task types such as text transcription, image selection, sliders, jigsaw puzzles, and logic-based challenges. Scoring is based on task completion and trace-conditioned validation.
- Released
- 2026-06-01
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
HLL addresses the gap in measuring agent capability at automation-protected workflows, providing a testbed for comparing progress in human-like interaction and process consistency.
Motivation
Multimodal agents are increasingly expected to operate interfaces on behalf of users, raising a central deployment question: can they truly substitute for humans in workflows that services deliberately protect against automation?
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.