Benchmark Radar
AI BENCHMARK PROFILE

OmegaUse-OfficeVal

General AIKnowledge & ReasoningBaidu Frontier Research

Evaluates LLM agents on long-horizon office-suite tasks from 100 practitioner-derived scenarios, scoring deliverable quality through code-based verifiers and comparing performance against economic signals of human labor time and price.

Released
2026-07-29
Readiness
Runnable
Primary field
General AI

Why it matters

Existing agent benchmarks rarely consider cost-effectiveness for real office workflows. This benchmark provides a way to compare agent output quality relative to human labor costs, supporting value-weighted evaluation of office automation.

Motivation

Large language model (LLM) agents are increasingly expected to assist users in completing tasks.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.