AI BENCHMARK PROFILE
MCR-Bench
MCR-Bench is a benchmark for evaluating reproducibility in mission-critical LLM tasks, measuring output consistency across heterogeneous hardware.
- Released
- 2026-06-19
- Readiness
- Paper only
- Primary field
- Finance & Economics
Why it matters
LLM deployments in finance, medicine, and law require reproducible outputs. MCR-Bench addresses the lack of standardized evaluation for numerical instability in 16-bit inference, helping practitioners select methods that balance reproducibility and performance.
Motivation
As Large Language Models (LLMs) deploy into mission-critical domains (e.g., finance, medicine, and law), output reproducibility has become a strict system requirement.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.