Benchmark Radar
AI BENCHMARK PROFILE

PAL-Bench

General AICoding & Software Engineering

PAL-Bench evaluates evidence-grounded profile reconstruction from longitudinal personal albums, scoring agents on owner facts, identities, and relations with a seven-metric protocol.

Released
2026-06-15
Readiness
Paper only
Primary field
General AI

Why it matters

Existing benchmarks test sub-problems of multimodal understanding, but PAL-Bench addresses the gap in album-scale reconstruction with social identity binding and evidence citation, offering a controlled public-record contract for evaluating perceptual entity resolution and multimodal integration.

Motivation

Longitudinal personal albums are weak-schema multimodal databases: noisy perceptual records whose key facts require joins across faces, text, timestamps, locations, and repeated events.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.