Benchmark Radar
AI BENCHMARK PROFILE

AARRI-Bench

General AIKnowledge & ReasoningAARR-bench

AARRI-Bench evaluates LLM agents on entry-level research intern tasks, measuring success rate in containerized environments with fixed tasks and scoring.

Released
2026-06-05
Readiness
Runnable
Primary field
General AI

Why it matters

Provides a reproducible benchmark for agentic research behavior, highlighting gaps in nuanced reasoning and offering a public comparison platform.

Motivation

As foundation models advance and agent scaffolding becomes increasingly sophisticated, agents have demonstrated remarkable proficiency in complex, long-horizon coding tasks and even autonomous experiment execution.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.