Benchmark Radar
AI BENCHMARK PROFILE

LifePlanner

General AIKnowledge & Reasoning

Evaluates LLM agents on geo-spatial planning tasks using map data enriched with social media posts, measuring pass rate across four task categories and three difficulty levels.

Released
2026-08-25
Readiness
Paper only
Primary field
General AI

Why it matters

Introduces realistic open-ended social signals into geo-spatial agent evaluation, revealing degradation in grounded planning beyond simple retrieval.

Motivation

Geo-spatial planning, like trip design, is a realistic testbed for LLM agents because it requires grounded tool use, noisy evidence retrieval, and multi-constraint reasoning.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.