Benchmark Radar
AI BENCHMARK PROFILE

LongWoF-Bench

General AIAgentsCoding & Software EngineeringMathematics & Formal Sciences

Evaluates verifiable long-workflow tasks across code generation, agent-environment synthesis, mathematical reasoning, and rule following, with machine-verifiable scoring.

Released
2026-08-24
Readiness
Paper only
Primary field
General AI

Why it matters

Provides a reusable benchmark for studying experience reuse in long-horizon LLM workflows, offering a standardized testbed for approaches like EvoMap.

Motivation

Large language models are increasingly expected to execute complex workflows whose success depends on maintaining interdependent constraints and producing artifacts that satisfy strict end-to-end verification.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.