HardMTBench
HardMTBench is a difficulty-aware diagnostic benchmark for Chinese-English domain translation, covering 12 domains with 20,000 directional test items and annotated hardness scores based on domain knowledge, translation difficulty, and terminology load.
- Released
- 2026-05-27
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
Addresses the saturation of general MT benchmarks on Chinese-English by widening score separation, exposing domain-specific weaknesses in knowledge-intensive areas that quality-only metrics miss.
Motivation
General-purpose machine translation benchmarks such as FLORES-200 have reached a saturation regime on Chinese-English pairs, where modern large language models cluster within a narrow band of high scores.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.