Skip to main navigation Skip to search Skip to main content

MoGenVD: A Motion-Centered Quality Assessment Benchmark for Text-to-Video Generation

  • Tianjin University of Science & Technology
  • Hong Kong Polytechnic University

Research output: Contribution to journalArticlepeer-review

Abstract

Text-to-video (T2V) generation has advanced rapidly in recent years, enabling the synthesis of visually plausible videos from user-input text prompts. As T2V models are increasingly applied in practical scenarios, reliable quality assessment of generated videos becomes essential. While existing subjective T2V quality assessment datasets have provided useful foundations for developing objective assessment methods, motion-related characteristics are still insufficiently represented and analyzed, which is one of the most challenging and critical aspects of video generation. In this paper, we present MoGenVD, a motion-centered dataset for generated video quality assessment. The dataset contains 888 motion-focused text prompts, organized into five motion categories. Based on these prompts, 5,328 videos are generated using six state-of-the-art T2V models and evaluated through controlled in-lab subjective experiments under four fine-grained quality dimensions, including visual quality, motion quality, motion subject authenticity, and semantic consistency. Comprehensive analysis of the subjective annotations reveals distinct performance characteristics and limitations of current T2V models with respect to motion modeling. Furthermore, we benchmark representative video quality assessment (VQA) models on MoGenVD, demonstrating that current VQA approaches still face challenges in accurately evaluating motion-centric generated content. The proposed dataset provides a challenging benchmark for T2V quality assessment and is expected to facilitate further research in this area. The dataset will be made publicly available at https://github.com/gracezhangyx/MoGenVD.

Original languageEnglish
Pages (from-to)12097-12110
Number of pages14
JournalIEEE Transactions on Circuits and Systems for Video Technology
Volume36
Issue number8
DOIs
StatePublished - Aug 1 2026

Keywords

  • Video quality assessment
  • motion quality
  • subjective quality assessment
  • text-to-video generation

Fingerprint

Dive into the research topics of 'MoGenVD: A Motion-Centered Quality Assessment Benchmark for Text-to-Video Generation'. Together they form a unique fingerprint.

Cite this