Joseph Roth
Papers
2
Total Citations
34
H-Index
2
About
Joseph Roth is a leading researcher in audio-visual machine learning, with a primary focus on active speaker detection—a critical component for advancing video analysis, speaker diarization, and human-robot interaction. His most significant contribution is the creation of the **AVA-ActiveSpeaker dataset**, the first large-scale, carefully labeled audio-visual benchmark for this task. Introduced in his highly cited 2020 paper (19 citations) and its 2019 supplementary material (15 citations), this dataset addresses a long-standing gap in the field by providing synchronized face tracks and audio streams with precise active/inactive speaker labels. Roth’s work has enabled robust training and evaluation of multimodal models, directly impacting applications like speech enhancement and video re-targeting for meetings. By systematically curating and annotating this resource, he has established a foundational standard for the community, accelerating progress in real-world, multi-speaker environments. His contributions are essential for any researcher developing algorithms that require reliable, context-aware speaker detection from video.
Research Focus
Key Achievements
Top Papers
- 1Ava Active Speaker: An Audio-Visual Dataset for Active Speaker Detection19 citations · 2020
- 2