Papers
3
Total Citations
22
H-Index
2
About
Yoshiki Masuyama is a researcher specializing in audio-visual perception, robot audition, and self-supervised machine learning, with a particular focus on enabling autonomous systems to better understand their acoustic and visual environments. His most notable contribution, "Self-supervised Neural Audio-Visual Sound Source Localization via Probabilistic Spatial Modeling" (2020), has garnered 16 citations and addresses a fundamental challenge in robotics: detecting and localizing sounding objects without requiring exhaustive manual labeling. By leveraging self-supervised learning techniques paired with probabilistic spatial modeling, Masuyama's approach allows systems to learn from naturally co-occurring audio-visual signals — a significant step toward scalable, real-world robot perception. Beyond audio-visual integration, Masuyama has contributed to robust sound source localization and separation through his work on probabilistic integration of MUSIC and Complex Gaussian Mixture Models (CGMM), published in 2021. This research tackles the practical challenge of accurately estimating directions of arrival for multiple sound sources, even under varying acoustic conditions. Collectively, his work pushes the boundaries of how intelligent systems perceive and interpret multi-modal sensory data, making him a promising voice in the fields of robot audition and audio-visual scene understanding.
Research Focus
Key Achievements
Top Papers
- 1
- 2
- 3