Benchmark Radar
AI BENCHMARK PROFILE

MMAE

General AIMultimodal Perception

MMAE is a benchmark for instruction-based audio editing with 2,000 samples across 7 modalities, 6 complexity levels, and rubric-based evaluation with 17,741 criteria for instruction following and context consistency.

Released
2026-06-05
Readiness
Runnable
Primary field
General AI

Why it matters

It provides the first comprehensive evaluation testbed for general-purpose audio editing, enabling precise multi-dimensional assessment and identifying bottlenecks in current models.

Motivation

We introduce MMAE, a Massive Multitask Audio Editing benchmark, serving as the first comprehensive evaluation testbed designed for general-purpose instruction-based audio editing.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.