MLUBench
MLUBench evaluates multimodal large language models on lifelong unlearning across 127 entities and 9 classes, providing QA pairs and images. It includes a protocol for sequential unlearning requests and evaluation of forgetting and retention.
- Released
- 2026-06-11
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
Lifelong unlearning is a practical challenge for MLLMs as data removal requests arrive over time. This benchmark enables systematic evaluation of unlearning methods and highlights the unique constraint of preserving multimodal alignment.
Motivation
Multimodal large language models (MLLMs) are trained on massive multimodal data, making data unlearning increasingly important as data owners may request the removal of specific content.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.