AI BENCHMARK PROFILE
RSIBench-Data
RSIBench-Data evaluates LLM agents as data-centric researchers, where agents iteratively revise training-data strategies for a fixed target model on six benchmarks, with real training and evaluation runs.
- Released
- 2026-07-28
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
It isolates research capability from engineering, showing that current agents can improve from feedback but inconsistently. This provides an auditable testbed for capabilities needed in recursive self-improvement.
Motivation
Recursive self-improvement requires turning evidence of model failures into better models.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.