Benchmark Radar
AI BENCHMARK PROFILE

LUNAR

General AIKnowledge & Reasoning

LUNAR is a benchmark for evaluating LLM personalization from longitudinal app interaction logs across domains like clothing, food, housing, and mobility. It uses a synthetic data pipeline and evaluates 19 LLMs.

Released
2026-08-05
Readiness
Paper only
Primary field
General AI

Why it matters

LUNAR could address the need for evaluating cross-domain personalization from behavioral logs, but without public artifacts or clear scoring details, its utility is unclear.

Motivation

Existing personalized LLM benchmarks primarily rely on textual personas or isolated behavioral signals, providing limited evaluation of cross-domain behavioral personalization, where responses must be grounded in heterogeneous daily-life activities.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.