MulRobBench
An offline protocol-conditioned benchmark for Vision-Language-Action UAV agents, evaluating operational context understanding, multimodal evidence arbitration, degradation-aware reasoning, and risk-aware action planning across 3,024 samples with semantic scoring and structural diagnostics.
- Released
- 2026-07-26
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Most UAV benchmarks focus on perception or navigation, leaving a gap in assessing coupled physical evidence, protocol constraints, and action risk. MulRobBench's diagnostic dimensions could inform safety-critical UAV deployment decisions.
Motivation
Smart-city airspace is transforming Uncrewed Aerial Vehicles (UAVs) from passive sensing platforms into cyber-physical decision makers that must follow operational rules under degraded observations and ambiguous language.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.