MMBench-Live
MMBench-Live is a continuously evolving multimodal benchmark built by a multi-agent pipeline from MMBench. It contains 5.9K newly generated evaluation instances with executable reasoning, and evaluates vision-language models across question-answer generation and reasoning tasks. Updates cost about USD 30 and take 1-2 hours, with distribution-consistent strategy to maintain comparability.
- Released
- 2026-07-02
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
Static benchmarks suffer from contamination and staleness; this benchmark provides a scalable, low-cost paradigm for sustainable evaluation that preserves model rankings and reduces memorization signals.
Motivation
Evaluation benchmarks are essential for assessing vision-language models (VLMs), but most multimodal benchmarks are static, making them vulnerable to temporal staleness, data contamination, and costly maintenance.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.