Benchmark Radar
AI BENCHMARK PROFILE

ORAgentBench

General AIAgents

ORAgentBench evaluates autonomous agents on end-to-end operations research tasks. Each task involves a natural-language brief, multi-file data, configuration artifacts, and a required submission schema. Agents write and run solution code, validated for schema, hard constraints, and objective quality.

Released
2026-06-18
Readiness
Paper only
Primary field
General AI

Why it matters

Existing OR evaluations often decouple modeling from solving and rarely test full workflows from artifacts to validated decisions. ORAgentBench provides a realistic, execution-grounded protocol for measuring practical decision-making in operations research, helping to identify strategic weaknesses in current models.

Motivation

Large language models are increasingly deployed as autonomous agents for multi-step tasks in executable environments, yet their ability to perform realistic operations research (OR) work remains unclear.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.