VitaBench
VitaBench 2.0 evaluates personalized and proactive agent behavior in long-term, multi-session user interactions across food delivery, in-store consumption, and online travel domains. Tasks are per-user sequences of subtasks requiring agents to infer, utilize, and update user preferences from fragmented interaction history, with an extensible memory interface for controlled comparison across memory architectures.
- Released
- 2026-05-26
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
Existing agent benchmarks focus on reasoning and tool use, overlooking the challenges of inferring and leveraging user preferences over time. This benchmark isolates personalization and proactivity in long-horizon tasks, providing a measurement of practical readiness for life-serving applications where models must act on implicit and evolving user needs.
Motivation
Large language models (LLMs) have evolved into interactive agents that collaborate with users in real-world tasks.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.