Benchmark Radar
AI BENCHMARK PROFILE

PRWeaver

General AICoding & Software Engineering

PRWeaver evaluates LLM-based code auditors against malicious pull requests across repository history, using 208 attacks and multiple renderings.

Released
2026-08-03
Readiness
Paper only
Primary field
General AI

Why it matters

Addresses reliability of code auditing agents under adversarial conditions, informing deployment of such systems.

Motivation

LLM-based code auditors are increasingly integrated into pull-request (PR) workflows, yet their reliability against adversarial changes distributed across repository evolution remains poorly understood.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.