Benchmark Radar
AI BENCHMARK PROFILE

Opti-Agent-Bench

General AIKnowledge & Reasoning

Opti-Agent-Bench evaluates LLM agents on the end-to-end optimization R&D pipeline, from business-language description to mathematical modeling, algorithm selection, code implementation, and report generation, with modular evaluation.

Released
2026-07-12
Readiness
Paper only
Primary field
General AI

Why it matters

Optimization benchmarks typically test pre-structured formulations. Opti-Agent-Bench assesses the full pipeline, exposing failure modes like constraint omission and model-code inconsistency invisible under single-metric evaluation.

Motivation

LLM-based agents are increasingly deployed to solve optimization problems, yet existing benchmarks evaluate them on pre-structured mathematical formulations that bypass the most critical challenge: translating complex business requirements into correct models and solve efficiently.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.