AI BENCHMARK PROFILE
DunphyBench
DunphyBench evaluates long-horizon embodied decision-making in housing environments, requiring agents to navigate and choose options aligned with multi-dimensional human preferences under partial observations.
- Released
- 2026-08-02
- Readiness
- Paper only
- Primary field
- Robotics & Autonomous Systems
Why it matters
Addresses the gap in evaluating agents on long-horizon, human-centered decisions beyond procedural tasks, providing a reference for progress in integrating multimodal evidence and preference reasoning.
Motivation
Agents are increasingly expected to act not only as task executors, but also as decision-makers on behalf of human users.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.