Skip to main navigation Skip to search Skip to main content

Object-driven text-to-image synthesis via adversarial training

  • Wenbo Li
  • , Pengchuan Zhang
  • , Lei Zhang
  • , Qiuyuan Huang
  • , Xiaodong He
  • , Siwei Lyu
  • , Jianfeng Gao
  • SUNY Albany
  • Microsoft USA
  • JD AI Research

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

279 Scopus citations

Abstract

In this paper, we propose Object-driven Attentive Generative Adversarial Newtorks (Obj-GANs) that allow attention-driven, multi-stage refinement for synthesizing complex images from text descriptions. With a novel object-driven attentive generative network, the Obj-GAN can synthesize salient objects by paying attention to their most relevant words in the text descriptions and their pre-generated class label. In addition, a novel object-wise discriminator based on the Fast R-CNN model is proposed to provide rich object-wise discrimination signals on whether the synthesized object matches the text description and the pre-generated class label. The proposed Obj-GAN significantly outperforms the previous state of the art in various metrics on the large-scale MS-COCO benchmark, increasing the inception score by 27% and decreasing the FID score by 11%. A thorough comparison between the classic grid attention and the new object-driven attention is provided through analyzing their mechanisms and visualizing their attention layers, showing insights of how the proposed model generates complex scenes in high quality.

Original languageEnglish
Title of host publicationProceedings - 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2019
PublisherIEEE Computer Society
Pages12166-12174
Number of pages9
ISBN (Electronic)9781728132938
DOIs
StatePublished - Jun 2019
Event32nd IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2019 - Long Beach, United States
Duration: Jun 16 2019Jun 20 2019

Publication series

NameProceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition
Volume2019-June
ISSN (Print)1063-6919

Conference

Conference32nd IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2019
Country/TerritoryUnited States
CityLong Beach
Period06/16/1906/20/19

Keywords

  • Image and Video Synthesis
  • Vision + Language

Fingerprint

Dive into the research topics of 'Object-driven text-to-image synthesis via adversarial training'. Together they form a unique fingerprint.

Cite this