Abstract
Text-to-video (T2V) generation has advanced rapidly in recent years, enabling the synthesis of visually plausible videos from user-input text prompts. As T2V models are increasingly applied in practical scenarios, reliable quality assessment of generated videos becomes essential. While existing subjective T2V quality assessment datasets have provided useful foundations for developing objective assessment methods, motion-related characteristics are still insufficiently represented and analyzed, which is one of the most challenging and critical aspects of video generation. In this paper, we present MoGenVD, a motion-centered dataset for generated video quality assessment. The dataset contains 888 motion-focused text prompts, organized into five motion categories. Based on these prompts, 5,328 videos are generated using six state-of-the-art T2V models and evaluated through controlled in-lab subjective experiments under four fine-grained quality dimensions, including visual quality, motion quality, motion subject authenticity, and semantic consistency. Comprehensive analysis of the subjective annotations reveals distinct performance characteristics and limitations of current T2V models with respect to motion modeling. Furthermore, we benchmark representative video quality assessment (VQA) models on MoGenVD, demonstrating that current VQA approaches still face challenges in accurately evaluating motion-centric generated content. The proposed dataset provides a challenging benchmark for T2V quality assessment and is expected to facilitate further research in this area. The dataset will be made publicly available at https://github.com/gracezhangyx/MoGenVD.
| Original language | English |
|---|---|
| Pages (from-to) | 12097-12110 |
| Number of pages | 14 |
| Journal | IEEE Transactions on Circuits and Systems for Video Technology |
| Volume | 36 |
| Issue number | 8 |
| DOIs | |
| State | Published - Aug 1 2026 |
Keywords
- Video quality assessment
- motion quality
- subjective quality assessment
- text-to-video generation
Fingerprint
Dive into the research topics of 'MoGenVD: A Motion-Centered Quality Assessment Benchmark for Text-to-Video Generation'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver