Benchmark Radar
AI BENCHMARK PROFILE

SocialPersona

General AIMultimodal Perception

SocialPersona evaluates multimodal LLMs on recovering revealed preferences from longitudinal social-media timelines and using them in dialogue. It includes 171 user timelines, 2,597 human-verified preference tags across seven domains, and supports profile construction and response generation tasks.

Released
2026-06-25
Readiness
Paper only
Primary field
General AI

Why it matters

Existing personalization benchmarks rely on explicitly stated preferences; SocialPersona tests inference from natural multimodal traces, a harder capability. It provides a reusable benchmark for measuring progress on long-horizon user modeling and personalized dialogue.

Motivation

Personalized language-model assistants are often evaluated through a memory lens: can a model recall preferences users have explicitly stated in dialogue?

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.