Benchmark Radar
AI BENCHMARK PROFILE

MLUBench

General AIKnowledge & ReasoningMLUBench Project

MLUBench evaluates multimodal large language models on lifelong unlearning across 127 entities and 9 classes, providing QA pairs and images. It includes a protocol for sequential unlearning requests and evaluation of forgetting and retention.

Released
2026-06-11
Readiness
Runnable
Primary field
General AI

Why it matters

Lifelong unlearning is a practical challenge for MLLMs as data removal requests arrive over time. This benchmark enables systematic evaluation of unlearning methods and highlights the unique constraint of preserving multimodal alignment.

Motivation

Multimodal large language models (MLLMs) are trained on massive multimodal data, making data unlearning increasingly important as data owners may request the removal of specific content.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.