Opti-Agent-Bench
Opti-Agent-Bench evaluates LLM agents on the end-to-end optimization R&D pipeline, from business-language description to mathematical modeling, algorithm selection, code implementation, and report generation, with modular evaluation.
- Released
- 2026-07-12
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Optimization benchmarks typically test pre-structured formulations. Opti-Agent-Bench assesses the full pipeline, exposing failure modes like constraint omission and model-code inconsistency invisible under single-metric evaluation.
Motivation
LLM-based agents are increasingly deployed to solve optimization problems, yet existing benchmarks evaluate them on pre-structured mathematical formulations that bypass the most critical challenge: translating complex business requirements into correct models and solve efficiently.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.