Benchmark Radar
AI BENCHMARK PROFILE

PhysTool-Bench

General AIAgentsTool CallingModalityDance

PhysTool-Bench evaluates multimodal LLMs on physical tool use through two tasks: recognizing all tools in a scene and selecting and sequencing tools for a given task, using 2,510 queries over 2,678 real-world tools.

Released
2026-06-09
Readiness
Runnable
Primary field
General AI

Why it matters

Physical tool use is underexplored in MLLMs; this benchmark isolates recognition and planning deficits, supporting progress in embodied AI and practical human-robot collaboration.

Motivation

Multimodal Large Language Models (MLLMs) excel at utilizing digital APIs and increasingly serve as the "brain" of embodied AI, instructing robots to interact with the physical world.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.