LabOSBench
LabOSBench evaluates multimodal GUI agents on 96 subtasks across eight web-based scientific-instrument simulators, covering workflows from sample loading to result inspection. Agents operate via a browser, with execution-based evaluation on task completion.
- Released
- 2026-06-15
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Existing computer-use benchmarks focus on software tasks, leaving a gap for scientific instrument control. LabOSBench provides a safe, reproducible, low-cost testbed to assess agents' feedback-driven and long-horizon capabilities in instrument operation, supporting practical adoption in laboratory automation.
Motivation
Current computer-use benchmarks primarily focus on software operation tasks in virtualized systems, whereas scientific instrumentation scenarios require coordinated control over complex interfaces, and feedback-driven parameter adjustment.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.