Joseph Roth

Google (United States)

Papers

2

Total Citations

34

H-Index

2

About

Joseph Roth is a leading researcher in audio-visual machine learning, with a primary focus on active speaker detection—a critical component for advancing video analysis, speaker diarization, and human-robot interaction. His most significant contribution is the creation of the **AVA-ActiveSpeaker dataset**, the first large-scale, carefully labeled audio-visual benchmark for this task. Introduced in his highly cited 2020 paper (19 citations) and its 2019 supplementary material (15 citations), this dataset addresses a long-standing gap in the field by providing synchronized face tracks and audio streams with precise active/inactive speaker labels. Roth’s work has enabled robust training and evaluation of multimodal models, directly impacting applications like speech enhancement and video re-targeting for meetings. By systematically curating and annotating this resource, he has established a foundational standard for the community, accelerating progress in real-world, multi-speaker environments. His contributions are essential for any researcher developing algorithms that require reliable, context-aware speaker detection from video.

Research Focus

Key Achievements

2
H-Index
2
Papers
34
Total Citations
17
Avg Citations/Paper
🏆 Most Cited Paper
Ava Active Speaker: An Audio-Visual Dataset for Active Speaker Detection
19 citations · 2020
📈 Most Prolific Year: 2020 (1 Papers)
🤝 Key Collaborators: 12
🏛 Institutions: Google (United States)

Top Papers

  1. 1
  2. 2

Key Collaborators

Contact & Links

Available for collaboration
Content generated · 12 days ago