Skip to main navigation Skip to search Skip to main content

Expanding the Image Embedding Space for Language-Free Text-to-Face Image Generation

  • Østfold University College
  • Indian Institute of Technology Palakkad

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Recent advancements in text-to-image (T2I) generation have revolutionized image synthesis, but conventional text-image paired training poses challenges when confronted with limited dataset size and narrow descriptive breadth. In particular, limited descriptive breadth can significantly impair a model’s ability to generate unmentioned image features. This issue is also evident in text-to-face image generation, where a method will underperform when rendering an image based on the text description of a specific facial feature not present in the training dataset. Language-free training emerges as a promising solution to this problem, leveraging latent spaces like CLIP to facilitate generalization from image embeddings to text embeddings. However, the modality gap remains a hurdle for language-free trained models. To address this, we propose a Gaussian perturbation-based technique that enhances perturbation coverage and robustness across varying modality gap sizes without the need to train a prior model. Our method achieves new state-of-the-art results on MM-CelebA-HQ in the language-free setting, presenting a novel solution to challenges in text-to-face image generation on limited datasets.

Original languageEnglish
Title of host publicationPattern Recognition - 46th DAGM German Conference, DAGM GCPR 2024, Proceedings
EditorsDaniel Cremers, Zorah Lähner, Michael Moeller, Matthias Nießner, Björn Ommer, Rudolph Triebel
PublisherSpringer Science and Business Media Deutschland GmbH
Pages72-86
Number of pages15
ISBN (Print)9781424469116, 9783031851865
DOIs
StatePublished - 2025
Event46th Annual Conference of the German Association for Pattern Recognition, DAGM GCPR 2024, co-hosted with VMV 2024 - Munich, Germany
Duration: Sep 10 2024Sep 13 2024

Publication series

NameLecture Notes in Computer Science
Volume15298 LNCS
ISSN (Print)0302-9743
ISSN (Electronic)1611-3349

Conference

Conference46th Annual Conference of the German Association for Pattern Recognition, DAGM GCPR 2024, co-hosted with VMV 2024
Country/TerritoryGermany
CityMunich
Period09/10/2409/13/24

Keywords

  • diffusion
  • language-free
  • text-to-face
  • text-to-image

Fingerprint

Dive into the research topics of 'Expanding the Image Embedding Space for Language-Free Text-to-Face Image Generation'. Together they form a unique fingerprint.

Cite this