Benchmark Radar
AI BENCHMARK PROFILE

SelectBench

General AIKnowledge & ReasoningSearch & Retrieval

SelectBench evaluates selective evidence adoption in retrieval-augmented language models, focusing on rejecting deceptive content. The benchmark includes a 325-example test set and rule- or judge-based reward scoring.

Released
2026-07-22
Readiness
Paper only
Primary field
General AI

Why it matters

Retrieval-augmented models often face mixed contexts with misleading content. A standardized evaluation for selective evidence adoption helps measure safety and reliability in real-world deployments.

Motivation

Retrieval-augmented large language models frequently face contexts that interleave useful evidence with misleading statements or instruction-like content.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.