AI BENCHMARK PROFILE
PhysTool-Bench
PhysTool-Bench evaluates multimodal LLMs on physical tool use through two tasks: recognizing all tools in a scene and selecting and sequencing tools for a given task, using 2,510 queries over 2,678 real-world tools.
- Released
- 2026-06-09
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
Physical tool use is underexplored in MLLMs; this benchmark isolates recognition and planning deficits, supporting progress in embodied AI and practical human-robot collaboration.
Motivation
Multimodal Large Language Models (MLLMs) excel at utilizing digital APIs and increasingly serve as the "brain" of embodied AI, instructing robots to interact with the physical world.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.