CLBench-V
CLBench-V evaluates multimodal context learning across three dimensions: context grounding, new information application, and new knowledge learning. It includes 3,443 instances across 14 subdatasets spanning science, finance, long-document understanding, spatial reasoning, and web-based VQA.
- Released
- 2026-07-28
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
Existing context learning benchmarks focus on text, missing multimodal settings where context is in figures, tables, and maps. CLBench-V provides a structured evaluation to localize where context use breaks down, aiding progress in multimodal models for real-world tasks.
Motivation
Real-world tasks often require models to learn from task-specific context rather than relying only on pre-trained knowledge.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.