AI BENCHMARK PROFILE
VSRo-200
A large-scale dataset for visual speech recognition in Romanian, with 200 hours of video and annotations. It studies supervision quality, robustness under domain shift, and multimodal fusion.
- Released
- 2026-07-09
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
The dataset enables research in low-resource visual speech recognition, but the paper primarily focuses on studying supervision and robustness rather than defining a fixed benchmark with a scoring contract. It is a dataset resource for training, not a standalone evaluation benchmark.
Motivation
We introduce VSRo-200, the first large-scale dataset for visual speech recognition (lip reading) in Romanian, comprising 200 hours of real-world podcast videos.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.