AI BENCHMARK PROFILE
LUNAR
LUNAR is a benchmark for evaluating LLM personalization from longitudinal app interaction logs across domains like clothing, food, housing, and mobility. It uses a synthetic data pipeline and evaluates 19 LLMs.
- Released
- 2026-08-05
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
LUNAR could address the need for evaluating cross-domain personalization from behavioral logs, but without public artifacts or clear scoring details, its utility is unclear.
Motivation
Existing personalized LLM benchmarks primarily rely on textual personas or isolated behavioral signals, providing limited evaluation of cross-domain behavioral personalization, where responses must be grounded in heterogeneous daily-life activities.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.