Benchmark Radar
AI BENCHMARK PROFILE

PHITSBench

General AIKnowledge & Reasoning

PHITSBench evaluates AI-assisted generation of PHITS radiation-transport input via natural language across 282 tasks in three workflows: Edit, Repair, and Reproduce. Scoring uses a Composite Metric Score combining execution success and agreement with reference transport observables.

Released
2026-07-08
Readiness
Paper only
Primary field
General AI

Why it matters

Provides an execution-grounded benchmark for a niche task, highlighting the need for machine-readable knowledge bases and curated training data in AI-assisted radiation-transport modeling.

Motivation

We introduce PHITSBench, an execution-scored benchmark for the Monte Carlo Particle and Heavy Ion Transport code System (PHITS).

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.