RandomBench
RandomBench evaluates whether multimodal LLMs maintain distributionally neutral behavior when selecting among equivalent options, providing metrics for entropy and distributional bias under explicit random instructions.
- Released
- 2026-06-04
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Logic-neutral scenarios are underexplored in MLLM evaluation. RandomBench introduces a way to quantify stochastic collapse, a bias toward non-uniform choices that affects repetitive behavior and coverage, aiding design of more robust models.
Motivation
Current evaluations for Multimodal Large Language Models (MLLMs) overwhelmingly focus on utility-driven objectives, leaving model behavior under logic-neutral scenarios largely underexplored.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.