Sharadh Ramaswamy
Papers
2
Total Citations
34
H-Index
2
About
Sharadh Ramaswamy is a researcher whose work sits at the critical intersection of audio and visual analysis, with a primary focus on active speaker detection. His major contribution to the field is the creation of the **AVA-ActiveSpeaker** dataset, a large-scale, carefully labeled audio-visual resource designed to train and evaluate algorithms that determine who is speaking in a video. This dataset directly addressed a key bottleneck in the field: the absence of a robust benchmark for this task. By providing this resource, Ramaswamy enabled significant advances in applications ranging from speaker diarization and video re-targeting for meetings to speech enhancement and human-robot interaction. His foundational papers on this work have garnered over 30 citations, establishing the dataset as a standard benchmark. Ramaswamy’s contribution is not just a dataset, but a catalyst that has empowered the broader research community to build more intelligent, context-aware systems for analyzing human communication in video.
Research Focus
Key Achievements
Top Papers
- 1Ava Active Speaker: An Audio-Visual Dataset for Active Speaker Detection19 citations · 2020
- 2