Home /Research /Controlling the voice quality dimension of prosody in synthetic speech using an acoustic glottal model
MANIPULATION

Controlling the voice quality dimension of prosody in synthetic speech using an acoustic glottal model

Andrew W. Murphy

Year
2021
Citations
3
Access
Open access

Abstract

Statistical parametric speech synthesis (SPSS) offers a means of generating synthetic speech without the need for complex and extensive rules. One way in which this approach is sometimes lacking is through the use of simple excitation models that may result in unnatural, robotic sounding synthetic speech. A further limitation is that these systems tend to lack prosodic variation. For example, they do not capture the expressive nature of human spoken interaction. This is something that would be highly desirable for applications utilising speech synthesis,such as educational games or synthetic voices for people with disordered speech. The use of a more complex excitation model could offer the flexibility in the voice source that could provide a basis for more adequate modelling of prosody. However, acoustic models of the voice source entail many potentially important parameters and controlling these could be a challenge. The main aims of this work were to: investigate how an acoustic glottal model could be used to manipulate aspects of linguistic and paralinguistic prosody of synthetic speech using a minimal set of control parameters; implement the knowledge gained from this investigation into an analysis-and-synthesis system; use this system in SPSS; and conduct pilot tests to demonstrate how the system can be used to explore the voice source correlates of prosody,through user-driven manipulation tasks. To achieve the first goal,experiments were carried out to explore how the global wave shape parameter, Rd (Fant, 1995), could be used to control aspects of linguistic and paralinguistic prosody. This parameter can be used to generate glottal pulse shapes that result in voice qualities ranging from breathy to tense. As the tense-lax dimension of voice quality is important in prosodic modulation, Rd appears to be ideal for minimising the number of control parameters needed to transform voice quality.To allow control of this parameter, the glottal source, and the vocal tract filter that shapes it, must be modelled using the principles of the source-filter model of speech production. These speech components must be separated effectively to do this. Inverse filtering was used to obtain estimates of the source signal by removing the effects of the vocal tract transfer function from the speech signal. An acoustic glottal model, the Liljencrants-Fant (LF) model (Fant et al., 1985), was then used to parameterise the glottal source. \nThree experiments were carried out,using manually inverse filtered data,to investigate how Rd could be used as a control parameter for linguistic and paralinguistic prosody,even in the absence of f0modulation. Experiment1 examined how manipulating Rd could be used to control where focal prominence (an aspect of linguistic prosody) occurs in an utterance. Experiment 2 explored how Rd could be used as a control parameter for perceived affect. Experiment 3 built upon the results of Experiment 1 to optimise the implementation of the Rd parameter contour.The results of these experiments confirmed that Rd can serve as a control parameter to generate linguistic prominence as well as paralinguistic modification of affective colouring. The results confirmed, and elaborated on,the findings of earlier research, suggesting that tense-lax modulation of voice quality is important in prosodic expression. They indicated that a more tense phonation on the focally accented item can be used to signal prominence, while laxer phonation of post-focal material provides source deaccentuation that further enhances the perceived prominence.These experiments provided information concerning Rd ranges and settings that fed into the development of the second goal of this work,i.e.an analysis-and-synthesis system, called GlórCáil,for the control of parameters for prosodic variation in synthesis. The system also allows for some speaker characteristic transformation,letting the user manipulate both voice source and vocal tract parameters be

Keywords

ProsodySpeech recognitionDimension (graph theory)Speech synthesisQuality (philosophy)AcousticsComputer scienceMathematicsPhysics

Related papers

Browse all MANIPULATION papers