Home /Research /Unsupervised phoneme and word acquisition from continuous speech based on a hierarchical probabilistic generative model
OTHER

Unsupervised phoneme and word acquisition from continuous speech based on a hierarchical probabilistic generative model

Masatoshi Nagano, Tomoaki Nakamura

Year
2023
Citations
3

Abstract

Humans can divide the perceived continuous speech signals, which exhibit double articulation structure, into phonemes and words without explicit boundary points or labels and thus learn a language. In constructive developmental studies, learning the double articulation structure of speech signals is important for realizing robots with human-like language learning abilities. In this study, we propose a novel probabilistic generative model called the Gaussian process-hidden semi-Markov model-based double articulation analyzer (GP-HSMM-DAA), which can learn phonemes and words from continuous speech signals by hierarchically connecting two probabilistic generative models (PGMs), namely, the Gaussian process-hidden semi-Markov model and hidden semi-Markov model. In the proposed model, the parameters of each PGM are mutually and complementarily updated and learned, enabling accurate learning of the phonemes and words. The experimental results reveal that GP-HSMM-DAA can segment continuous speech into phonemes and words with higher accuracy than the conventional method.

Keywords

Hidden Markov modelComputer scienceSpeech recognitionProbabilistic logicGenerative grammarGenerative modelArtificial intelligenceWord (group theory)Articulation (sociology)Natural language processing

Related papers

Browse all OTHER papers