TY - GEN
T1 - 3D singing head for music VR
T2 - 27th ACM International Conference on Multimedia, MM 2019
AU - Yu, Jun
AU - Chen, Chang Wen
AU - Wang, Zengfu
N1 - Publisher Copyright:
© 2019 Copyright held by the owner/author(s). Publication rights licensed to ACM.
PY - 2019/10/15
Y1 - 2019/10/15
N2 - We propose a real-time 3D singing head system to enhance the talking head on model integrity, keyframe generation and song synchronicity. The individual head appearance meshes are first obtained by matching multi-view visible images with face prior for accuracy, and then used to reconstruct entire head model by integrating with generic internal articulatory meshes for efficiency. After embedding physiology, the keyframes of each phoneme-music note correspondence are substantially synthesized from real articulation data. The song synchronicity of articulators is learned using a deep neural network to train visual co-articulation model (VCM) on parallel audio-visual data. Finally, the keyframes of adjacent phoneme-music note correspondences are blended by VCM to produce song synchronized animation. Compared to state-of-the-art baselines, our system can not only clearly distinguish phonemes and notes, but also significantly reduce the dependence on training data .
AB - We propose a real-time 3D singing head system to enhance the talking head on model integrity, keyframe generation and song synchronicity. The individual head appearance meshes are first obtained by matching multi-view visible images with face prior for accuracy, and then used to reconstruct entire head model by integrating with generic internal articulatory meshes for efficiency. After embedding physiology, the keyframes of each phoneme-music note correspondence are substantially synthesized from real articulation data. The song synchronicity of articulators is learned using a deep neural network to train visual co-articulation model (VCM) on parallel audio-visual data. Finally, the keyframes of adjacent phoneme-music note correspondences are blended by VCM to produce song synchronized animation. Compared to state-of-the-art baselines, our system can not only clearly distinguish phonemes and notes, but also significantly reduce the dependence on training data .
KW - Articulatory animation
KW - Virtual reality
KW - Visual song synthesis
KW - Visual speech synthesis
UR - https://www.scopus.com/pages/publications/85074827826
U2 - 10.1145/3343031_3350865
DO - 10.1145/3343031_3350865
M3 - Conference contribution
AN - SCOPUS:85074827826
T3 - MM 2019 - Proceedings of the 27th ACM International Conference on Multimedia
SP - 945
EP - 952
BT - MM 2019 - Proceedings of the 27th ACM International Conference on Multimedia
PB - Association for Computing Machinery, Inc
Y2 - 21 October 2019 through 25 October 2019
ER -