Benchmark Radar
AI BENCHMARK PROFILE

GAMA-Bench

General AIKnowledge & Reasoning

GAMA-Bench evaluates LLMs on gender-asymmetric moral framing across 1,298 paired conflict scenarios, measuring response differences in punitive, therapeutic, and blame dimensions.

Released
2026-06-12
Readiness
Runnable
Primary field
General AI

Why it matters

detects biases beyond stereotypes by comparing responses to matched male and female actors, informing fairness in AI decision-making.

Motivation

Existing studies on gender bias in LLMs have largely focused on stereotypes, occupational associations, or explicit harmful outputs.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.