AI BENCHMARK PROFILE
NuclearQAv2
NuclearQAv2 evaluates LLMs on nuclear engineering knowledge with approximately 1,240 QA pairs covering boolean, numeric, and verbal questions. It uses structured prompting for automated question generation and response evaluation.
- Released
- 2026-06-25
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Technical domains like nuclear engineering require reliable LLM evaluation. NuclearQAv2 provides a scalable benchmark to assess factual knowledge, quantitative reasoning, and conceptual understanding in this domain.
Motivation
Large language models (LLMs) have demonstrated strong performance across a wide range of tasks, but ensuring their reliability in highly technical domains remains a significant challenge.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.