Benchmark Radar
AI BENCHMARK PROFILE

VQABench

General AIMultimodal Perception

Evaluates 12 image preprocessing techniques for cloud VLM-based visual question answering across 3 VQA datasets and 4 commercial models, measuring accuracy, cost, and latency.

Released
2026-08-08
Readiness
Paper only
Primary field
General AI

Why it matters

Assesses the impact of client-side preprocessing on cost-quality trade-offs for offloaded VQA inference, offering practical guidance for system design.

Motivation

Vision-language models (VLMs) are becoming a practical backend for mobile visual question answering (VQA) systems, enabling smartphones and smart glasses to answer users' questions about the physical world.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.