TY - GEN
T1 - A Large-Scale 3D Representation Dataset and Benchmark for Continuous Sign Language Understanding
AU - Chaudhary, Lipisha
AU - Hoq, Enjamamul
AU - Dong, Lu
AU - Adler, Henry
AU - Nwogu, Ifeoma
N1 - Publisher Copyright:
© 2026 IEEE.
PY - 2026
Y1 - 2026
N2 - American Sign Language (ASL) is a visually rich 3D language essential to the Deaf and Hard-of-Hearing (D/HH) communities, where 3D mesh representations have increasingly surpassed traditional skeletal poses in capturing expressive detail and articulatory nuance. Yet, the scarcity of large-scale 3D ASL datasets continues to constrain model generalization and the enhancement of cross-linguistic knowledge. To address this gap, we present a large-scale, expressive, and privacy-conscious 3D mesh dataset comprising 250+ hours of data curated from various sign language datasets1. Building upon this resource, we establish comprehensive benchmarks for ASL translation and generation, facilitating standardized evaluation across tasks. In addition, we propose a full-body fine-grained refinement method, FusePose, which jointly integrates high-quality hand and body representations to improve articulation fidelity. Extensive experiments demonstrate that even a small amount of high-quality data can foster better cross-linguistic generation. Furthermore, we show a downstream application in controllable character animation, highlighting the broader impact of our dataset and methodology on 3D ASL research. Our experiments further show that fused hand-body representations provide stronger performance for both Sign Language Translation and Sign Language Generation/Production.
AB - American Sign Language (ASL) is a visually rich 3D language essential to the Deaf and Hard-of-Hearing (D/HH) communities, where 3D mesh representations have increasingly surpassed traditional skeletal poses in capturing expressive detail and articulatory nuance. Yet, the scarcity of large-scale 3D ASL datasets continues to constrain model generalization and the enhancement of cross-linguistic knowledge. To address this gap, we present a large-scale, expressive, and privacy-conscious 3D mesh dataset comprising 250+ hours of data curated from various sign language datasets1. Building upon this resource, we establish comprehensive benchmarks for ASL translation and generation, facilitating standardized evaluation across tasks. In addition, we propose a full-body fine-grained refinement method, FusePose, which jointly integrates high-quality hand and body representations to improve articulation fidelity. Extensive experiments demonstrate that even a small amount of high-quality data can foster better cross-linguistic generation. Furthermore, we show a downstream application in controllable character animation, highlighting the broader impact of our dataset and methodology on 3D ASL research. Our experiments further show that fused hand-body representations provide stronger performance for both Sign Language Translation and Sign Language Generation/Production.
UR - https://www.scopus.com/pages/publications/105043394411
U2 - 10.1109/FG67764.2026.11557028
DO - 10.1109/FG67764.2026.11557028
M3 - Conference contribution
AN - SCOPUS:105043394411
T3 - FG 2026 - 20th IEEE International Conference on Automatic Face and Gesture Recognition
BT - FG 2026 - 20th IEEE International Conference on Automatic Face and Gesture Recognition
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 20th IEEE International Conference on Automatic Face and Gesture Recognition, FG 2026
Y2 - 25 May 2026 through 29 May 2026
ER -