AI BENCHMARK PROFILE
VideoWeaver
VideoWeaver is an agent harness and benchmark for long video generation, with 16 task categories and 285 cases, evaluating agents via evidence-grounded judge on process and output.
- Released
- 2026-06-06
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
General-purpose agents are underevaluated on long-horizon multimodal tasks; VideoWeaver offers a reproducible benchmark to assess and evolve agent skills for video generation.
Motivation
Recent agent frameworks such as Claude Code, Codex, and OpenClaw are strong at tool use and orchestration, but whether they can handle long video generation, a long-horizon multimodal task, remains underexplored.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.