PHREEQC-MCQ-200
PHREEQC-MCQ-200 evaluates tool-augmented agents on 200 multiple-choice questions derived from 21 validated PHREEQC scenarios. Agents must construct simulator inputs, execute PHREEQC, inspect structured outputs, and commit to final answers. Scoring is based on exact match of selected answer.
- Released
- 2026-07-01
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
This benchmark addresses the lack of standardized evaluation for tool-augmented agents in scientific simulation, measuring not only accuracy but also item-level retention, output-access sensitivity, and trajectory failures. It provides a diagnostic lens on when tool access improves or degrades performance, informing design of reliable scientific agents.
Motivation
Large language model agents are increasingly connected to scientific software, yet it remains unclear when tool access makes scientific computation more reliable rather than merely more complex.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.