Papers
7
Total Citations
132
H-Index
4
About
Xinyuan Qian is a leading researcher in multi-modal perception for human-robot interaction, specializing in audio-visual speaker tracking and localization. His major contributions lie in developing robust, real-time systems that fuse complementary audio and visual signals to overcome challenging acoustic environments, such as noisy and reverberant spaces. His seminal work, "Multi-Speaker Tracking From an Audio–Visual Sensing Device" (60 citations), introduced a compact platform for portable robotics, while his "Audio-Visual Cross-Attention Network for Robotic Speaker Tracking" (37 citations) advanced multi-modal fusion using deep learning. Qian also pioneered speech-oriented attention mechanisms for sound source localization, as seen in his GCC-PHAT-based approach (20 citations). Beyond tracking, he has explored privacy-preserving SSL with analytic class incremental learning and contributed to responsive listening head generation with his non-autoregressive Transformer model, ListenFormer. His interdisciplinary work extends to CMOS image sensors with polarization pixels for machine vision. With a growing citation impact exceeding 130, Qian’s research is pivotal for next-generation robotics, surveillance, and assistive technologies, bridging signal processing and deep learning to enable more intuitive human-robot collaboration.
Research Focus
Key Achievements
Top Papers
- 1Multi-Speaker Tracking From an Audio–Visual Sensing Device60 citations · 2019
- 2Audio-Visual Cross-Attention Network for Robotic Speaker Tracking37 citations · 2022
- 3
- 4
- 5
- 6
- 7