Benchmark Radar
AI BENCHMARK PROFILE

RobustMAD

General AISafety & TrustworthinessEN Research

Evaluates robustness of multimodal small language models for industrial anomaly detection across open-ended queries and visual degradations, using multiple-choice accuracy and LLM-judged open-ended responses.

Released
2026-06-26
Readiness
Runnable
Primary field
General AI

Why it matters

Assesses deployability of compact models in real-world industrial conditions, identifying failure modes like fragile grounding and hallucination on ill-posed queries.

Motivation

Multimodal industrial anomaly inspection assistants are a critical component of next-generation smart factories, enabling interactive vision-language-based querying.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.