Skip to main navigation Skip to search Skip to main content

Synthesizing 3D Trump: Predicting and Visualizing the Relationship between Text, Speech, and Articulatory Movements

  • University of Science and Technology of China
  • Tsinghua University

Research output: Contribution to journalArticlepeer-review

2 Scopus citations

Abstract

The movements of articulators, such as lips, tongue and teeth, play an important role in increasing the language expression capability by unmasking the information hid in text or speech. Hence, it is necessary to deeply mine and visualize the relationship between text, speech and articulatory movements for understanding language in multi-modality and multi-level. As a case study, given text and audio of President Donald John Trump, this paper synthesizes a high quality 3D animation of him speaking with accurate synchronicity between speech and articulators. First, visual co-articulation is modeled by predicting the mapping from text/speech to articulatory movements. Then, based on a reconstructed 3D head model, physiological characteristics and statistical learning are combined to visualize each phoneme. Finally, the visualization results of consecutive phonemes are fused by visual co-articulation model to generate synchronized articulatory animations. Experiments show that the system can not only produce photo-realistic results in front but also distinguish the visual differences among phonemes from unconstrained views.

Original languageEnglish
Article number8805138
Pages (from-to)2223-2233
Number of pages11
JournalIEEE/ACM Transactions on Audio Speech and Language Processing
Volume27
Issue number12
DOIs
StatePublished - Dec 2019

Keywords

  • speech animation
  • Visual co-articulation

Fingerprint

Dive into the research topics of 'Synthesizing 3D Trump: Predicting and Visualizing the Relationship between Text, Speech, and Articulatory Movements'. Together they form a unique fingerprint.

Cite this