AI BENCHMARK PROFILE
WorkSurface-Bench
WorkSurface-Bench evaluates enterprise agents on knowledge routing across documents, tables, and graphs. It includes 1,151 atomic tasks with auditable reference answers and scoring for route, evidence, answer, and efficiency.
- Released
- 2026-07-28
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
It isolates surface routing from evidence acquisition and answer generation, showing that correct routing is necessary but insufficient. This helps diagnose why agents fail on multi-surface tasks and informs design of routing-aware systems.
Motivation
Enterprise agents often need to integrate heterogeneous knowledge sources: documents for narrative facts, tables for computation, and dependency graphs for file relationships.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.