Benchmark Radar
AI BENCHMARK PROFILE

AgentHPOBench

General AIKnowledge & ReasoningAgentHPOBench team

AgentHPOBench evaluates LLM agents as sequential hyperparameter optimizers across 30 executable ML tasks. Agents observe accumulated configurations, metrics, and logs, then propose the next configuration. Scoring compares agents and conventional HPO baselines under a unified protocol.

Released
2026-07-31
Readiness
Paper only
Primary field
General AI

Why it matters

Existing benchmarks do not directly assess whether agents can interpret experimental evidence and use it to guide subsequent hyperparameter decisions. AgentHPOBench provides a repeatable protocol for measuring iterative decision-making, filling a gap in evaluating autonomous scientific agents.

Motivation

As LLMs evolve from code completion systems into autonomous scientific agents, evaluating their ability to conduct experiments has become increasingly important.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.