Skip to main navigation Skip to search Skip to main content

3D singing head for music VR: Learning external and internal articulatory synchronicity from lyric, audio and notes

  • University of Science and Technology of China

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

7 Scopus citations

Abstract

We propose a real-time 3D singing head system to enhance the talking head on model integrity, keyframe generation and song synchronicity. The individual head appearance meshes are first obtained by matching multi-view visible images with face prior for accuracy, and then used to reconstruct entire head model by integrating with generic internal articulatory meshes for efficiency. After embedding physiology, the keyframes of each phoneme-music note correspondence are substantially synthesized from real articulation data. The song synchronicity of articulators is learned using a deep neural network to train visual co-articulation model (VCM) on parallel audio-visual data. Finally, the keyframes of adjacent phoneme-music note correspondences are blended by VCM to produce song synchronized animation. Compared to state-of-the-art baselines, our system can not only clearly distinguish phonemes and notes, but also significantly reduce the dependence on training data .

Original languageEnglish
Title of host publicationMM 2019 - Proceedings of the 27th ACM International Conference on Multimedia
PublisherAssociation for Computing Machinery, Inc
Pages945-952
Number of pages8
ISBN (Electronic)9781450368896
DOIs
StatePublished - Oct 15 2019
Event27th ACM International Conference on Multimedia, MM 2019 - Nice, France
Duration: Oct 21 2019Oct 25 2019

Publication series

NameMM 2019 - Proceedings of the 27th ACM International Conference on Multimedia

Conference

Conference27th ACM International Conference on Multimedia, MM 2019
Country/TerritoryFrance
CityNice
Period10/21/1910/25/19

Keywords

  • Articulatory animation
  • Virtual reality
  • Visual song synthesis
  • Visual speech synthesis

Fingerprint

Dive into the research topics of '3D singing head for music VR: Learning external and internal articulatory synchronicity from lyric, audio and notes'. Together they form a unique fingerprint.

Cite this