Sourish Chaudhuri
Papers
2
Total Citations
34
H-Index
2
About
Sourish Chaudhuri is a leading researcher in audio-visual machine learning, with a core focus on active speaker detection—a critical task for enabling machines to understand who is speaking in a video. His major contribution is the creation of the **AVA-ActiveSpeaker dataset**, a large-scale, carefully annotated audio-visual benchmark that filled a crucial gap in the field. Prior to this work, the absence of such a dataset constrained progress in applications ranging from speaker diarization and video re-targeting for meetings to speech enhancement and human-robot interaction. Chaudhuri’s dataset, introduced in his 2020 paper (19 citations) and its 2019 supplementary material (15 citations), provides the standardized evaluation framework that has since become foundational for the community. By providing synchronized face tracks and audio streams with precise active-speaker labels, his work has enabled the development of more robust, real-world capable models. Chaudhuri’s contributions are essential for any researcher building systems that must seamlessly integrate visual and auditory cues to understand human communication.
Research Focus
Key Achievements
Top Papers
- 1Ava Active Speaker: An Audio-Visual Dataset for Active Speaker Detection19 citations · 2020
- 2