← Yi Xu's homepage

Research

Overview

Divide and conquer

How exactly does human speech transmit multiple layers of communicative meanings through an articulation process? This is the central concern of my research. To address this issue, some fundamental questions need to be answered: What are the kinds of meanings transmitted by speech? What are the encoding mechanisms? What are the decoding mechanisms? Since it is impossible to answer these questions all at once, a realistic strategy is to divide and conquer. That is, to always prioritize the kind of questions for which other things are relatively established.

My research priority has been based on the following understanding of the state of the art in speech science:

  • With regard to encoding and decoding mechanisms, the static aspects of speech sounds, whether in terms of acoustic patterns or articulatory correlates, are relatively well established. What remains unclear, even to this day, is the basic dynamic mechanisms of speech production and their processing in perception.
  • With regard to meanings, lexical meanings are the most easily established; all other meanings are up for grabs.

My early work was therefore focused on Mandarin tones in continuous speech. The functional meaning of tone is clear: to distinguish morphemes that are otherwise identical in terms of CV structure. The canonical forms of Mandarin tones had also been previously well established. What my work further established is the basic patterns of contextual tonal variation (Xu 1993, 1994, 1997, 1998, 2001a). This has led to the Target Approximation (TA) model of tone production (Xu & Wang, 2001). The TA model was then applied to intonation of both Mandarin and English (Xu, 1999; Xu & Xu, 2005). The success of these applications led to further expansion of the approach in a number of new directions.

New directions

Models and theories

Directions that grew out of the Target Approximation approach.

03

Maximum rate of information hypothesis

Speech is driven by the need to convey information at the fastest rate possible (Xu & Prom-on, 2019). As a result, speech articulation is executed near an overall performance ceiling in terms of articulatory effort (Xu & Sun, 2002; Xu & Wang, 2009; Cheng & Xu, 2013). This view differs from the widely accepted principle of economy of effort, especially in the form of the H&H theory (Lindblom, 1990).

04

Synchronization model of the syllable →

A new conceptualization of the syllable as a synchronization mechanism that initiates the articulation of consonant, vowel and tone simultaneously at the onset of the syllable. It offers a drastically different view not only on the nature of the syllable, but also on issues such as coarticulation, coarticulation resistance, locus equation, time interval of segments and temporal alignment of segmental and tonal events (Kang & Xu, forthcoming; Liu & Xu, 2023; Liu, Xu & Hsieh, 2022; Xu, 2020; Xu & Liu, 2006).

06

Single Origin of PFC hypothesis

The use of post-focus compression (PFC) as a prosodic marker of focus is likely to have a single historical origin, possibly the hypothetical proto-Nostratic language (Xu, 2011; Xu, Chen & Wang, 2012).

07

Direct perception theory

Consonants, vowels and lexical tones can be recognized directly from raw speech signals, without extracting subcategorical features such as distinctive features, articulatory gestures or tone levels in the case of tone. (Chen, Gao & Xu, 2022; Gauthier, Shi & Xu, 2007a, 2007b).

Evidence

Major empirical findings

  1. Consonant, vowel and tone are fully synchronized at the syllable onset ( Kang & Xu, 2024; Liu, Xu & Hsieh, 2022; Liu & Xu, 2021; Xu & Liu, 2006, 2007)
  2. English is not stress-timed, because its segment duration is inflexible. But Mandarin is both phrase-timed and syllable-timed, because it shows a weak tendency toward both equal syllable duration and equal phrase duration (Wang, Xu & Zhang, 2023; Xu & Wang, 2009).
  3. Maximum speed of articulation is regularly reached in normal speech, contrary to the notion of economy of effort (Xu & Prom-on, 2019; Xu & Sun, 2002).
  4. Post-focus Compression (PFC) is absent in Taiwanese, Taiwan Mandarin and Cantonese (Chen, Wang & Xu, 2009; Wu & Xu, 2010)
  5. Whispered Mandarin has no production-enhanced cues for tone and intonation (Jiao & Xu, 2019).
  6. Emotional prosody as well as human vocal attractiveness are expressed mainly through body-size projection (Chuenwattanapranithi et al., 2008; Xu et al., 2013).
  7. Post-low bouncing -- After a Low tone, f0 bounces back in the subsequent syllables, especially when they are in the neutral tone (Chen & Xu, 2006; Prom-on, Liu & Xu, 2012).
  8. Mandarin neutral tone is not toneless, but likely has a [mid] pitch target accompanied by weak articulatory strength (Chen & Xu, 2006; Xu & Prom-on, 2014).
  9. F0 peak delay is closely related to the interaction of tonal targets articulatory constraints (Xu, 2001a).
  10. Prosodic focus is encoded via tri-zone pitch range manipulation in Mandarin (Xu, 1999) and English (Xu & Xu, 2005).
  11. Tonal targets are synchronously implemented with the entire syllable rather than with only nucleus vowel or syllable rime (Xu, 1998).
  12. Contextual tonal variations are robustly asymmetrical: Carryover effects are strong and assimilatory, whereas anticipatory effects are weak and largely dissimilatory (Xu, 1993, 1994, 1997, 1999).
  13. Lexical tones in Mandarin that are distorted due to articulatory constraints are still perceptually identifiable. No categorical changes therefore is likely to have taken place (Xu, 1993, 1994).
  14. Perceptual compensation for contextual tonal variations is not complete. (Xu, 1993, 1994).
  15. Mandarin tone 3 (the Low tone) sandhi is applied in short-term memory, indicating the phonetic nature of human working memory (Xu, 1991).
  16. In a Mandarin syllable with a final nasal, the duration of the nucleus vowel is inversely related to vowel height: the higher the vowel, the shorter the duration; This duration variation is compensated for by the duration of the nasal murmur: the shorter the vowel, the longer the nasal murmur; Thus in a syllable with a low vowel such as /bang/, there is often hardly any nasal murmur, whereas in a syllable with a high vowel such as /bing/, the nasal murmur can be longer than the vowel (Xu, 1986);
  17. Mandarin final nasals are realized as nasalization on the preceding nucleus vowel with no nasal murmur if the following syllable begins with a vowel or a glide (Xu, 1986);
  18. In a Mandarin disyllabic word or phrase, the initial consonant position in syllable 2 is much less "consonantal" than the initial consonant position in syllable 1:  where the consonant is shorter, more likely to become voiced if voiceless, a stop or affricate is more likely to lose its closure and become a fricative, and a fricative is more likely to lose its frication (Xu, 1986).

Further reading

Research philosophies

Read more →