MTEB-BR
MTEB-BR is a text embedding benchmark for Brazilian Portuguese with 22 native tasks across seven categories, including classification, clustering, retrieval, and reranking. It evaluates models on native Portuguese data and provides a public leaderboard.
- Released
- 2026-07-06
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
Portuguese embedding evaluation lacked native benchmarks. This benchmark provides statistically grounded model comparison, showing that multilingual leaderboards only moderately predict Portuguese performance.
Motivation
Text embeddings for Portuguese have no dedicated benchmark: evaluation rests on translated corpora such as English MS MARCO or on thin multilingual coverage, with native tasks scattered and unconsolidated.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.