Benchmark Radar
AI BENCHMARK PROFILE

RepBench

General AIKnowledge & Reasoning

RepBench compiles 353 public benchmarks into 46,149 probe texts spanning 94 capabilities, with a taxonomy of 182 clusters in 13 families. It evaluates representation readout methods across 12 models under cross-benchmark transfer, providing a reusable closed-loop pipeline for capability-aligned probing.

Released
2026-07-30
Readiness
Paper only
Primary field
General AI

Why it matters

RepBench addresses the lack of comparable and reproducible evaluation for representation engineering by grounding probes in multiple public benchmarks, reducing single-source bias and enabling meaningful comparison of readout methods and aggregation criteria across models.

Motivation

Representation engineering reads and steers capability directions in large language models, yet methods are typically evaluated on paper-specific synthetic data.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.