Benchmark Radar
AI BENCHMARK PROFILE

DRFLOW

General AIKnowledge & Reasoning

DRFLOW evaluates an agent's ability to predict personalized workflows, sequences of action-steps, from heterogeneous sources across five domains. It includes 100 tasks with reference workflow steps and multiple diagnostic metrics.

Released
2026-06-16
Readiness
Paper only
Primary field
General AI

Why it matters

Deep research systems are typically evaluated on report generation, but enterprise tasks often require actionable workflows. DRFLOW addresses this gap by assessing workflow prediction, offering metrics for factual grounding, step recovery, and personalization, which can guide development of more practical agents.

Motivation

Deep research (DR) systems are increasingly used for complex information-seeking tasks, but existing works mainly focus on generating reports and summaries.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.