AI BENCHMARK PROFILE
LifePlanner
Evaluates LLM agents on geo-spatial planning tasks using map data enriched with social media posts, measuring pass rate across four task categories and three difficulty levels.
- Released
- 2026-08-25
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Introduces realistic open-ended social signals into geo-spatial agent evaluation, revealing degradation in grounded planning beyond simple retrieval.
Motivation
Geo-spatial planning, like trip design, is a realistic testbed for LLM agents because it requires grounded tool use, noisy evidence retrieval, and multi-constraint reasoning.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.