AI BENCHMARK PROFILE
WorldLines
WorldLines evaluates long-horizon stateful embodied agents in household environments through Memory QA and Embodied Task Planning, using temporally extended household traces with dialogues and state changes.
- Released
- 2026-06-17
- Readiness
- Paper only
- Primary field
- Robotics & Autonomous Systems
Why it matters
Existing benchmarks lack evaluation of long-term memory in dynamic embodied settings; WorldLines fills this gap by testing both memory retrieval and planning over extended interactions.
Motivation
To assist humans over extended periods in real homes, embodied agents must remember user routines, world states, and past interactions.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.