AI BENCHMARK PROFILE
AOR-Bench
AOR-Bench is a benchmark of 3,000 pseudo-harmful audio samples across six categories, designed to evaluate over-refusal in large audio language models. It assesses whether models incorrectly reject benign audio queries.
- Released
- 2026-06-19
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Over-refusal in audio models is understudied. AOR-Bench provides a standardized set to measure this behavior, helping align safety mechanisms with real-world audio context.
Motivation
Large Audio Language Models (LALMs) have demonstrated strong performance across a wide range of audio tasks.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.