AgentHPOBench
AgentHPOBench evaluates LLM agents as sequential hyperparameter optimizers across 30 executable ML tasks. Agents observe accumulated configurations, metrics, and logs, then propose the next configuration. Scoring compares agents and conventional HPO baselines under a unified protocol.
- Released
- 2026-07-31
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Existing benchmarks do not directly assess whether agents can interpret experimental evidence and use it to guide subsequent hyperparameter decisions. AgentHPOBench provides a repeatable protocol for measuring iterative decision-making, filling a gap in evaluating autonomous scientific agents.
Motivation
As LLMs evolve from code completion systems into autonomous scientific agents, evaluating their ability to conduct experiments has become increasingly important.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.