Benchmark Radar
AI BENCHMARK PROFILE

RandomBench

General AISafety & Trustworthiness

RandomBench evaluates whether multimodal LLMs maintain distributionally neutral behavior when selecting among equivalent options, providing metrics for entropy and distributional bias under explicit random instructions.

Released
2026-06-04
Readiness
Paper only
Primary field
General AI

Why it matters

Logic-neutral scenarios are underexplored in MLLM evaluation. RandomBench introduces a way to quantify stochastic collapse, a bias toward non-uniform choices that affects repetitive behavior and coverage, aiding design of more robust models.

Motivation

Current evaluations for Multimodal Large Language Models (MLLMs) overwhelmingly focus on utility-driven objectives, leaving model behavior under logic-neutral scenarios largely underexplored.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.