Benchmark Radar
AI BENCHMARK PROFILE

WorkSurface-Bench

General AIKnowledge & Reasoninghaolpku

WorkSurface-Bench evaluates enterprise agents on knowledge routing across documents, tables, and graphs. It includes 1,151 atomic tasks with auditable reference answers and scoring for route, evidence, answer, and efficiency.

Released
2026-07-28
Readiness
Runnable
Primary field
General AI

Why it matters

It isolates surface routing from evidence acquisition and answer generation, showing that correct routing is necessary but insufficient. This helps diagnose why agents fail on multi-surface tasks and informs design of routing-aware systems.

Motivation

Enterprise agents often need to integrate heterogeneous knowledge sources: documents for narrative facts, tables for computation, and dependency graphs for file relationships.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.