Benchmark Radar
AI BENCHMARK PROFILE

PPE-Bench

General AIKnowledge & Reasoning

PPE-Bench evaluates machine unlearning in multimodal large language models under private-public entanglement, where images contain a target individual to forget and public elements to preserve.

Released
2026-07-03
Readiness
Paper only
Primary field
General AI

Why it matters

Addresses the lack of benchmarks that reflect real-world image complexity and entanglement of private and public information, supporting evaluation of unlearning methods that must preserve public context.

Motivation

Multimodal Large Language Models (MLLMs) have shown strong capabilities, but they may memorize private information from web data, raising privacy concerns.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.