Benchmark Radar
AI BENCHMARK PROFILE

MAC-Bench

General AIKnowledge & Reasoning

MAC-Bench evaluates procedural compliance of multi-agent systems under social-engineering pressure, measuring compliance-weighted success rate and Machiavellian gap.

Released
2026-06-05
Readiness
Paper only
Primary field
General AI

Why it matters

Addresses 'Goodhart's Law' in agent alignment, offering metrics that trade off task success and compliance, but lacks a public evaluation path.

Motivation

The rapid evolution of Large Language Models (LLMs) from passive assistants to autonomous, execution-capable agents has introduced critical operational risks.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.