Benchmark Radar
AI BENCHMARK PROFILE

SEVRA-BENCH

CybersecurityRobotics & Autonomous SystemsKnowledge & Reasoning

SEVRA-BENCH is a benchmark for measuring how often LLM-based code review agents approve adversarial pull requests with social-engineering framings, built from vulnerability-fixing commits. It includes a challenge split of roughly 1,000 adversarial PRs.

Released
2026-06-11
Readiness
Paper only
Primary field
Cybersecurity

Why it matters

Review agents are susceptible to narrative manipulation, which can lead to merging vulnerable code. SEVRA-BENCH quantifies this gap in security capabilities.

Motivation

Large language models (LLMs) are increasingly deployed in automated code-review systems, where their approvals can determine which code is merged into shared repositories.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.