Speech recognition
Related papers: 20
About
Speech recognition is the technology that enables machines to identify, interpret, and transcribe spoken human language into text or actionable commands. It works by processing audio signals through a pipeline that typically involves acoustic feature extraction (such as mel-frequency cepstral coefficients), language modeling, and pattern matching algorithms — increasingly powered by deep learning architectures like recurrent neural networks and transformer models. In robotics and AI, speech recognition serves as a foundational input modality, enabling voice-controlled robots, conversational agents, and human-robot interaction systems where natural spoken commands replace physical interfaces. It is closely integrated with speaker localization, noise suppression, and turn-taking systems to function effectively in real-world environments. Beyond command recognition, it extends into emotion and prosody detection, enriching a robot's contextual understanding of human intent. Speech recognition matters because it dramatically lowers the barrier for intuitive human-machine communication, making robotic systems more accessible to non-expert users and enabling deployment in assistive, service, and social robotics contexts where natural interaction is essential.
Top Researchers
Top Institutes
Top Cited Papers
Derivative Dynamic Time Warping
Eamonn Keogh, Michael J. Pazzani
Citations: 1124 • 2001
Deep Learning with Convolutional Neural Networks Applied to Electromyography Data: A Resource for the Classification of Movements for Prosthetic Hands
Manfredo Atzori, Matteo Cognolato, Henning Müller
Citations: 662 • 2016
Max-pooling convolutional neural networks for vision-based hand gesture recognition
Jawad Nagi, Frederick Ducatelle, Gianni A. Di, Dan Cireşan, Ueli Meier, Alessandro Giusti, Farrukh Nagi, Jürgen Schmidhuber, Luca Maria Gambardella
Citations: 648 • 2011
Real Time Face Detection and Facial Expression Recognition: Development and Applications to Human Computer Interaction.
Marian Stewart Bartlett, Gwen Littlewort, Ian Fasel, Javier R. Movellan
Citations: 568 • 2003
The production and recognition of emotions in speech: features and algorithms
Oudeyer Pierre-Yves
Citations: 441 • 2003
Clustering-Based Speech Emotion Recognition by Incorporating Learned Features and Deep BiLSTM
Mustaqeem Mustaqeem, Muhammad Sajjad, Soonil Kwon
Citations: 396 • 2020
Error-Related EEG Potentials Generated During Simulated Brain–Computer Interaction
Pierre W. Ferrez, José del R. Millán
Citations: 363 • 2008
Confusions Among Visually Perceived Consonants
Cletus G. Fisher
Citations: 338 • 1968
Perturbation of Vowel Articulations By Consonantal Context: An Acoustical Study
Kenneth N. Stevens, Arthur S. House
Citations: 305 • 1963
Speech gesture generation from the trimodal context of text, audio, and speaker identity
Youngwoo Yoon, Bok Cha, Joo-Haeng Lee, Minsu Jang, Jaeyeon Lee, Jaehong Kim, Geehyuk Lee
Citations: 300 • 2020
On the Acoustics of Emotion in Audio: What Speech, Music, and Sound have in Common
Felix Weninger, Florian Eyben, Björn W. Schuller, Marcello Mortillaro, Klaus R. Scherer
Citations: 277 • 2013
Multimodal Gesture Recognition Using 3-D Convolution and Convolutional LSTM
Guangming Zhu, Liang Zhang, Peiyi Shen, Juan Song
Citations: 277 • 2017
SAE+LSTM: A New Framework for Emotion Recognition From Multi-Channel EEG
Xiaofen Xing, Zhiyang Li, Tianyuan Xu, Lin Shu, Bin Hu, Xiangmin Xu
Citations: 268 • 2019
Brain–computer interfaces for speech communication
Jonathan S. Brumberg, Alfonso Nieto-Castañón, Philip R. Kennedy, Frank H. Guenther
Citations: 262 • 2010
Speech emotion recognition based on feature selection and extreme learning machine decision tree
Zhentao Liu, Min Wu, Weihua Cao, Jun-Wei Mao, Jianping Xu, Guanzheng Tan
Citations: 260 • 2017
Turn-taking in Conversational Systems and Human-Robot Interaction: A Review
Gabriel Skantze
Citations: 255 • 2020
Online Myoelectric Control of a Dexterous Hand Prosthesis by Transradial Amputees
Christian Cipriani, Christian Antfolk, Marco Controzzi, Göran Lundborg, Birgitta Rosén, Maria Chiara Carrozza, Fredrik Sebelius
Citations: 245 • 2011
Multi-Sensor Guided Hand Gesture Recognition for a Teleoperated Robot Using a Recurrent Neural Network
Wen Qi, Salih Ertug Ovur, Zhijun Li, Aldo Marzullo, Rong Song
Citations: 245 • 2021
Sensory Preference in Speech Production Revealed by Simultaneous Alteration of Auditory and Somatosensory Feedback
Daniel R. Lametti, Sazzad M. Nasir, David J. Ostry
Citations: 233 • 2012
Head gesture recognition for hands‐free control of an intelligent wheelchair
Pei Jia, Huosheng Hu, Tao Lü, Kui Yuan
Citations: 225 • 2007