ORAgentBench
ORAgentBench evaluates autonomous agents on end-to-end operations research tasks. Each task involves a natural-language brief, multi-file data, configuration artifacts, and a required submission schema. Agents write and run solution code, validated for schema, hard constraints, and objective quality.
- Released
- 2026-06-18
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Existing OR evaluations often decouple modeling from solving and rarely test full workflows from artifacts to validated decisions. ORAgentBench provides a realistic, execution-grounded protocol for measuring practical decision-making in operations research, helping to identify strategic weaknesses in current models.
Motivation
Large language models are increasingly deployed as autonomous agents for multi-step tasks in executable environments, yet their ability to perform realistic operations research (OR) work remains unclear.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.