Benchmark Radar
AI BENCHMARK PROFILE

AgentFairBench

General AIKnowledge & Reasoning

AgentFairBench evaluates demographic disparity in the actions of LLM agents across hiring, lending, and medical triage. It uses synthetic, demographic-neutral profiles in counterfactual matched sets varying name-coded race/gender. Metrics include counterfactual flip rate, mean absolute score difference, action-rate disparity, and tool-invocation disparity, with bootstrap confidence intervals and FDR control.

Released
2026-06-15
Readiness
Paper only
Primary field
General AI

Why it matters

Existing fairness evaluations grade answers, not actions. AgentFairBench addresses the gap by measuring disparity in consequential agent decisions. Its low cost and reproducible harness provide a practical path for screening models for action-level bias before deployment.

Motivation

Large language model (LLM) agents increasingly take actions (screening applicants, recommending credit, triaging patients), yet fairness for LLMs is still measured by grading answers.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.