VibeLifeBench
VibeLifeBench evaluates LLM agents on 200 long-horizon, multi-week tasks across 10 everyday-life domains in a simulated world of 22 mock services. Agents must manage silent world changes and implicit constraints, with fine-grained weighted checks on end state, timeliness, and constraint adherence.
- Released
- 2026-08-11
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
Existing benchmarks focus on short, static tasks, leaving a gap in measuring long-horizon proactive assistance. VibeLifeBench provides a credible public evaluation for agents that must operate over weeks with changing environments, offering practical decision value for deploying personal assistants.
Motivation
Large language model (LLM) agents are increasingly deployed as personal assistants.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.