Benchmark Radar
AI BENCHMARK PROFILE

LabOSBench

General AIAgents

LabOSBench evaluates multimodal GUI agents on 96 subtasks across eight web-based scientific-instrument simulators, covering workflows from sample loading to result inspection. Agents operate via a browser, with execution-based evaluation on task completion.

Released
2026-06-15
Readiness
Paper only
Primary field
General AI

Why it matters

Existing computer-use benchmarks focus on software tasks, leaving a gap for scientific instrument control. LabOSBench provides a safe, reproducible, low-cost testbed to assess agents' feedback-driven and long-horizon capabilities in instrument operation, supporting practical adoption in laboratory automation.

Motivation

Current computer-use benchmarks primarily focus on software operation tasks in virtualized systems, whereas scientific instrumentation scenarios require coordinated control over complex interfaces, and feedback-driven parameter adjustment.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.