Benchmark Radar
AI BENCHMARK PROFILE

ChemCoTBench-V2

Science & ResearchKnowledge & Reasoning

ChemCoTBench-V2 evaluates chemical reasoning at the process level, with 5,620 samples across molecular understanding, editing, optimization, and reaction prediction. It checks intermediate steps using deterministic chemistry rules and reference traces.

Released
2026-06-02
Readiness
Paper only
Primary field
Science & Research

Why it matters

Existing chemistry benchmarks only score final answers, missing violations in reasoning. ChemCoTBench-V2 provides auditable, verifiable process-level evaluation with three separate signals, enabling fine-grained model comparison.

Motivation

Large language models are increasingly used as chemistry assistants, yet most chemistry benchmarks still score only final answers.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.