Benchmark Radar
AI BENCHMARK PROFILE

PHREEQC-MCQ-200

General AIAgentsThe authors

PHREEQC-MCQ-200 evaluates tool-augmented agents on 200 multiple-choice questions derived from 21 validated PHREEQC scenarios. Agents must construct simulator inputs, execute PHREEQC, inspect structured outputs, and commit to final answers. Scoring is based on exact match of selected answer.

Released
2026-07-01
Readiness
Paper only
Primary field
General AI

Why it matters

This benchmark addresses the lack of standardized evaluation for tool-augmented agents in scientific simulation, measuring not only accuracy but also item-level retention, output-access sensitivity, and trajectory failures. It provides a diagnostic lens on when tool access improves or degrades performance, informing design of reliable scientific agents.

Motivation

Large language model agents are increasingly connected to scientific software, yet it remains unclear when tool access makes scientific computation more reliable rather than merely more complex.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.