Skip to main navigation Skip to search Skip to main content

Towards Open Domain Text-Driven Synthesis of Multi-person Motions

  • Mengyi Shan
  • , Lu Dong
  • , Yutao Han
  • , Yuan Yao
  • , Tao Liu
  • , Ifeoma Nwogu
  • , Guo Jun Qi
  • , Mitch Hill
  • University of Washington
  • SUNY Buffalo
  • Innopeak Technology
  • University of Rochester
  • Westlake University

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

1 Scopus citations

Abstract

This work aims to generate natural and diverse group motions of multiple humans from textual descriptions. While single-person text-to-motion generation is extensively studied, it remains challenging to synthesize motions for more than one or two subjects from in-the-wild prompts, mainly due to the lack of available datasets. In this work, we curate human pose and motion datasets by estimating pose information from large-scale image and video datasets. Our models use a transformer-based diffusion framework that accommodates multiple datasets with any number of subjects or frames. Experiments explore both generation of multi-person static poses and generation of multi-person motion sequences. To our knowledge, our method is the first to generate multi-subject motion sequences with high diversity and fidelity from a large variety of textual prompts.

Original languageEnglish
Title of host publicationComputer Vision – ECCV 2024 - 18th European Conference, Proceedings
EditorsAleš Leonardis, Elisa Ricci, Stefan Roth, Olga Russakovsky, Torsten Sattler, Gül Varol
PublisherSpringer Science and Business Media Deutschland GmbH
Pages67-86
Number of pages20
ISBN (Print)9783031736490
DOIs
StatePublished - 2025
Event18th European Conference on Computer Vision, ECCV 2024 - Milan, Italy
Duration: Sep 29 2024Oct 4 2024

Publication series

NameLecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics)
Volume15123 LNCS
ISSN (Print)0302-9743
ISSN (Electronic)1611-3349

Conference

Conference18th European Conference on Computer Vision, ECCV 2024
Country/TerritoryItaly
CityMilan
Period09/29/2410/4/24

Keywords

  • Human Motion Generation
  • Human Pose Dataset
  • Multi-Person Motion Generation
  • Text-to-Motion Generation

Fingerprint

Dive into the research topics of 'Towards Open Domain Text-Driven Synthesis of Multi-person Motions'. Together they form a unique fingerprint.

Cite this