Benchmark Radar
AI BENCHMARK PROFILE

StylisticBias

General AISafety & TrustworthinessStylisticBias Team

StylisticBias is a benchmark of 25,000 images with controlled single-attribute variations to evaluate attribute-level social bias in multimodal LLMs. It measures how visual cues shift model judgments in 25 scenarios.

Released
2026-06-18
Readiness
Runnable
Primary field
General AI

Why it matters

Social bias in multimodal models is often confounded by identity. StylisticBias isolates visual cues, enabling targeted bias diagnosis and mitigation.

Motivation

Multimodal large language models (MLLMs) are increasingly deployed in personally and societally consequential settings, yet the visual cues that shape how these models judge people remain poorly understood.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.