Benchmark Radar
AI BENCHMARK PROFILE

VibeLifeBench

General AIKnowledge & ReasoningEvolvent AI

VibeLifeBench evaluates LLM agents on 200 long-horizon, multi-week tasks across 10 everyday-life domains in a simulated world of 22 mock services. Agents must manage silent world changes and implicit constraints, with fine-grained weighted checks on end state, timeliness, and constraint adherence.

Released
2026-08-11
Readiness
Runnable
Primary field
General AI

Why it matters

Existing benchmarks focus on short, static tasks, leaving a gap in measuring long-horizon proactive assistance. VibeLifeBench provides a credible public evaluation for agents that must operate over weeks with changing environments, offering practical decision value for deploying personal assistants.

Motivation

Large language model (LLM) agents are increasingly deployed as personal assistants.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.