OR-Space
A full-lifecycle workspace benchmark for industrial optimization agents, evaluating model construction, revision, and grounded explanation across three task modes (Build, Revise, Explain) using executable multi-file workspaces with task-specific evaluators.
- Released
- 2026-05-27
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
Fills the gap in benchmarking LLM agents for real industrial OR workflows, where persistent multi-artifact workspaces and multi-stage lifecycles are central, offering a more realistic evaluation of practical readiness beyond single-shot formulation tasks.
Motivation
Large language model (LLM) agents are increasingly used to assist with operations research (OR) modeling, yet existing OR-oriented benchmarks often reduce evaluation to one-shot translation from a self-contained problem statement into a mathematical formulation or solver program.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.