Benchmark Radar
AI BENCHMARK PROFILE

AOR-Bench

General AIMultimodal PerceptionAOR-Bench Team

AOR-Bench is a benchmark of 3,000 pseudo-harmful audio samples across six categories, designed to evaluate over-refusal in large audio language models. It assesses whether models incorrectly reject benign audio queries.

Released
2026-06-19
Readiness
Paper only
Primary field
General AI

Why it matters

Over-refusal in audio models is understudied. AOR-Bench provides a standardized set to measure this behavior, helping align safety mechanisms with real-world audio context.

Motivation

Large Audio Language Models (LALMs) have demonstrated strong performance across a wide range of audio tasks.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.