AI BENCHMARK PROFILE
LongWoF-Bench
Evaluates verifiable long-workflow tasks across code generation, agent-environment synthesis, mathematical reasoning, and rule following, with machine-verifiable scoring.
- Released
- 2026-08-24
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Provides a reusable benchmark for studying experience reuse in long-horizon LLM workflows, offering a standardized testbed for approaches like EvoMap.
Motivation
Large language models are increasingly expected to execute complex workflows whose success depends on maintaining interdependent constraints and producing artifacts that satisfy strict end-to-end verification.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.