AI BENCHMARK PROFILE
MMAE
MMAE is a benchmark for instruction-based audio editing with 2,000 samples across 7 modalities, 6 complexity levels, and rubric-based evaluation with 17,741 criteria for instruction following and context consistency.
- Released
- 2026-06-05
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
It provides the first comprehensive evaluation testbed for general-purpose audio editing, enabling precise multi-dimensional assessment and identifying bottlenecks in current models.
Motivation
We introduce MMAE, a Massive Multitask Audio Editing benchmark, serving as the first comprehensive evaluation testbed designed for general-purpose instruction-based audio editing.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.