Benchmark Radar
AI BENCHMARK PROFILE

BulkPR-Bench

General AICoding & Software EngineeringZenodo

BulkPR-Bench evaluates governance of interacting pull requests. It includes 581 candidate PRs on 18 repositories, with metrics RDS and Global-SGY measuring safe delivery. The benchmark uses executable repository execution with hidden safety checks.

Released
2026-08-03
Readiness
Runnable
Primary field
General AI

Why it matters

Coding-agent benchmarks often assume independent PRs. BulkPR-Bench addresses interactive PR queues, providing a benchmark for evaluating joint decision-making in complex scenarios, with executable validation.

Motivation

Coding-agent benchmarks increasingly cover long-horizon, end-to-end, and interactive development, but typically retain one requested outcome or a fixed change sequence.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.