Benchmark Radar
AI BENCHMARK PROFILE

VideoWeaver

General AIMultimodal PerceptionJianhuiWei7

VideoWeaver is an agent harness and benchmark for long video generation, with 16 task categories and 285 cases, evaluating agents via evidence-grounded judge on process and output.

Released
2026-06-06
Readiness
Runnable
Primary field
General AI

Why it matters

General-purpose agents are underevaluated on long-horizon multimodal tasks; VideoWeaver offers a reproducible benchmark to assess and evolve agent skills for video generation.

Motivation

Recent agent frameworks such as Claude Code, Codex, and OpenClaw are strong at tool use and orchestration, but whether they can handle long video generation, a long-horizon multimodal task, remains underexplored.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.