Benchmark Radar
AI BENCHMARK PROFILE

RSIBench-Data

General AICoding & Software Engineeringevolvent-ai

RSIBench-Data evaluates LLM agents as data-centric researchers, where agents iteratively revise training-data strategies for a fixed target model on six benchmarks, with real training and evaluation runs.

Released
2026-07-28
Readiness
Runnable
Primary field
General AI

Why it matters

It isolates research capability from engineering, showing that current agents can improve from feedback but inconsistently. This provides an auditable testbed for capabilities needed in recursive self-improvement.

Motivation

Recursive self-improvement requires turning evidence of model failures into better models.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.