MTM-Bench
MTM-Bench is a controlled benchmark for language-conditioned task execution in multilingual settings, enumerating all 27 instruction-content-response language triplets across English, Spanish, and Chinese. It contains 2,430 instances per model across semantic reversal, final-state extraction, and language purity tasks, with decomposed metrics for semantic correctness, language adherence, constraint satisfaction, and joint success.
- Released
- 2026-05-26
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Multilingual LLMs are used when instruction, source content, and response languages differ, yet existing evaluations rarely isolate these roles. MTM-Bench provides a fully crossed design to attribute degradation to specific language roles, revealing that response-slot mismatch drives most performance loss and that mismatch count is not a monotonic predictor of difficulty.
Motivation
Multilingual LLMs are increasingly used when instruction, source content, and required response languages do not coincide.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.