FBHM
FBHM evaluates vision-language models on hateful meme detection across 25 rhetorical functionalities and 10 target communities, with 5,000 memes. Performance is measured by Macro-F1 score on this curated dataset.
- Released
- 2026-05-29
- Readiness
- Paper only
- Primary field
- Cybersecurity
Why it matters
Existing hateful meme benchmarks confound rhetorical strategies with target community features, preventing causal evaluation of model vulnerabilities. FBHM isolates these axes, revealing that models rely on dataset-specific heuristics rather than robust reasoning, and offers a controlled environment for measuring generalization.
Motivation
Hateful meme detection remains a formidable challenge for vision-language models, as existing benchmarks are structurally observational - confounding rhetorical hate mechanisms with target community features and preventing causal evaluation of model vulnerabilities.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.