AI BENCHMARK PROFILE
GAMA-Bench
GAMA-Bench evaluates LLMs on gender-asymmetric moral framing across 1,298 paired conflict scenarios, measuring response differences in punitive, therapeutic, and blame dimensions.
- Released
- 2026-06-12
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
detects biases beyond stereotypes by comparing responses to matched male and female actors, informing fairness in AI decision-making.
Motivation
Existing studies on gender bias in LLMs have largely focused on stereotypes, occupational associations, or explicit harmful outputs.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.