Benchmark Radar
AI BENCHMARK PROFILE

DataSpace

General AIKnowledge & ReasoningHKUST Dial

DataSpace evaluates data agents on analytical questions over heterogeneous workspaces containing CSV, JSON, SQLite, Markdown, PDF, and video artifacts. Agents must discover and integrate relevant evidence and return complete tabular results, scored by a deterministic evaluator using header-invariant column alignment and type-aware comparison.

Released
2026-08-04
Readiness
Runnable
Primary field
General AI

Why it matters

Existing benchmarks isolate structured querying, retrieval, or open-ended analysis. DataSpace unifies evidence discovery, tabular output completeness, and deterministic evaluation, addressing a gap in assessing data agents for real-world analytical tasks. It provides a practical decision value for comparing agent harnesses and multimodal backbones.

Motivation

Data agents enable natural-language analytics over organizational workspaces, where relevant evidence may be scattered across databases, structured files, long documents, and multimedia.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.