EgoBench
EgoBench is an interactive multimodal benchmark for tool-using agents. It comprises 1,045 egocentric-video-grounded tasks across four daily scenarios, with a user-agent-tool interactive environment. It assesses multimodal perception, tool invocation with multi-hop reasoning, and dynamic user interaction.
- Released
- 2026-05-27
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
AI agents in open environments need joint multimodal and tool-use capabilities. EgoBench provides a standardized interactive environment and deterministic joint validation to objectively measure these skills, addressing a lack of comparable benchmarks for dynamic tool-using agents.
Motivation
As AI agents increasingly operate in open, real-world environments, they require a deep synergy of multimodal perception, tool invocation with multi-hop reasoning, and dynamic interaction with users.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.