VideoFDB
VideoFDB is a benchmark presented for evaluating full-duplex audio-visual conversational agents. It includes 237 dyadic clips with 11 nonverbal conversational dynamics from real-world video calls, along with a taxonomy and rubric-based LM-as-judge evaluation framework.
- Released
- 2026-05-28
- Readiness
- Inspectable
- Primary field
- Cybersecurity
Why it matters
Existing full-duplex benchmarks only evaluate speech, missing the audio-visual nature of natural conversation. VideoFDB aims to fill this gap by evaluating agents that must produce and interpret nonverbal cues.
Motivation
Natural human conversation is full-duplex and audio-visual: people simultaneously speak and listen while continuously interpreting and producing nonverbal cues, such as nods, smiles, and gestures.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.