qTA model →
A computational realization of the TA model. Its parameters can be automatically extracted by a Praat script: qTAtrainer (Prom-on, Xu & Thipakorn, 2009).
Overview
How exactly does human speech transmit multiple layers of communicative meanings through an articulation process? This is the central concern of my research. To address this issue, some fundamental questions need to be answered: What are the kinds of meanings transmitted by speech? What are the encoding mechanisms? What are the decoding mechanisms? Since it is impossible to answer these questions all at once, a realistic strategy is to divide and conquer. That is, to always prioritize the kind of questions for which other things are relatively established.
My research priority has been based on the following understanding of the state of the art in speech science:
My early work was therefore focused on Mandarin tones in continuous speech. The functional meaning of tone is clear: to distinguish morphemes that are otherwise identical in terms of CV structure. The canonical forms of Mandarin tones had also been previously well established. What my work further established is the basic patterns of contextual tonal variation (Xu 1993, 1994, 1997, 1998, 2001a). This has led to the Target Approximation (TA) model of tone production (Xu & Wang, 2001). The TA model was then applied to intonation of both Mandarin and English (Xu, 1999; Xu & Xu, 2005). The success of these applications led to further expansion of the approach in a number of new directions.
New directions
Directions that grew out of the Target Approximation approach.
A computational realization of the TA model. Its parameters can be automatically extracted by a Praat script: qTAtrainer (Prom-on, Xu & Thipakorn, 2009).
A model of speech prosody that allows parallel marking of multiple layers of communicative meanings with articulatory targets that can generate surface f0 contours (Xu, 2005). The intonational version of PENTA is realized through PENTAtrainer, a Praat script that can automatically extract function-loaded pitch targets (Xu & Prom-on, 2014).
Speech is driven by the need to convey information at the fastest rate possible (Xu & Prom-on, 2019). As a result, speech articulation is executed near an overall performance ceiling in terms of articulatory effort (Xu & Sun, 2002; Xu & Wang, 2009; Cheng & Xu, 2013). This view differs from the widely accepted principle of economy of effort, especially in the form of the H&H theory (Lindblom, 1990).
A new conceptualization of the syllable as a synchronization mechanism that initiates the articulation of consonant, vowel and tone simultaneously at the onset of the syllable. It offers a drastically different view not only on the nature of the syllable, but also on issues such as coarticulation, coarticulation resistance, locus equation, time interval of segments and temporal alignment of segmental and tonal events (Kang & Xu, forthcoming; Liu & Xu, 2023; Liu, Xu & Hsieh, 2022; Xu, 2020; Xu & Liu, 2006).
A theory of vocal expression of emotions (Xu, Kelly & Smillie, 2013). Emotional and attitudinal meanings are vocal expressed by simultaneously manipulating a number of bio-informational dimensions -- size projection, dynamicity, audibility and association. At least the first two dimensions have been found to be highly relevant for a number of emotions (Chuenwattanapranithi et al., 2008; Xu & Kelly, 2010; Xu, Kelly & Smillie, 2013).
The use of post-focus compression (PFC) as a prosodic marker of focus is likely to have a single historical origin, possibly the hypothetical proto-Nostratic language (Xu, 2011; Xu, Chen & Wang, 2012).
Consonants, vowels and lexical tones can be recognized directly from raw speech signals, without extracting subcategorical features such as distinctive features, articulatory gestures or tone levels in the case of tone. (Chen, Gao & Xu, 2022; Gauthier, Shi & Xu, 2007a, 2007b).
Evidence
Further reading