AI BENCHMARK PROFILE
AARRI-Bench
AARRI-Bench evaluates LLM agents on entry-level research intern tasks, measuring success rate in containerized environments with fixed tasks and scoring.
- Released
- 2026-06-05
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
Provides a reproducible benchmark for agentic research behavior, highlighting gaps in nuanced reasoning and offering a public comparison platform.
Motivation
As foundation models advance and agent scaffolding becomes increasingly sophisticated, agents have demonstrated remarkable proficiency in complex, long-horizon coding tasks and even autonomous experiment execution.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.