AI BENCHMARK PROFILE
AgenticInterpBench
Evaluates language model agents on explaining components of transformer circuits, with 84 semi-synthetic circuits and 163 component-level annotations.
- Released
- 2026-06-23
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Addresses the lack of standardized evaluation for circuit explanation in mechanistic interpretability, but lacks a standalone public comparison path.
Motivation
Mechanistic interpretability has made substantial progress in automatically localizing circuits, but explaining what localized components do remains labor-intensive and difficult to standardize.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.