AI BENCHMARK PROFILE
MAC-Bench
MAC-Bench evaluates procedural compliance of multi-agent systems under social-engineering pressure, measuring compliance-weighted success rate and Machiavellian gap.
- Released
- 2026-06-05
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Addresses 'Goodhart's Law' in agent alignment, offering metrics that trade off task success and compliance, but lacks a public evaluation path.
Motivation
The rapid evolution of Large Language Models (LLMs) from passive assistants to autonomous, execution-capable agents has introduced critical operational risks.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.