Benchmark Radar
AI BENCHMARK PROFILE

HardMTBench

General AIKnowledge & Reasoning

HardMTBench is a difficulty-aware diagnostic benchmark for Chinese-English domain translation, covering 12 domains with 20,000 directional test items and annotated hardness scores based on domain knowledge, translation difficulty, and terminology load.

Released
2026-05-27
Readiness
Runnable
Primary field
General AI

Why it matters

Addresses the saturation of general MT benchmarks on Chinese-English by widening score separation, exposing domain-specific weaknesses in knowledge-intensive areas that quality-only metrics miss.

Motivation

General-purpose machine translation benchmarks such as FLORES-200 have reached a saturation regime on Chinese-English pairs, where modern large language models cluster within a narrow band of high scores.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.