AI BENCHMARK PROFILE
InterFLOPBench
InterFLOPBench is a benchmark of 90 C kernels and 1,130 test samples for evaluating LLMs on floating-point error classification across six categories: cancellation, comparison, division by zero, overflow, underflow, and NaN.
- Released
- 2026-06-30
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Provides a targeted evaluation for LLM capabilities in static floating-point error detection, a niche but important area. Enables comparison across models and error types.
Motivation
This paper investigates the capability of Large Language Models (LLMs) to detect and classify floating-point errors statically in software code.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.