Skip to main navigation Skip to search Skip to main content

Multi-Modal Multi-Task Unified Embedding Model (M3T-UEM): A Task-Adaptive Representation Learning Framework

  • Rohan Sharma
  • , Changyou Chen
  • , Feng Ju Chang
  • , Seongjun Yun
  • , Xiaohu Xie
  • , Rui Meng
  • , Dehong Xu
  • , Alejandro Mottini
  • , Qingjun Cui
  • Amazon.com, Inc.
  • SUNY Buffalo
  • University of California at Los Angeles

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

1 Scopus citations

Abstract

We present Multi-Modal Multi-Task Unified Embedding Model (M3T-UEM), a framework that advances visionlanguage matching and retrieval by leveraging a large language model (LLM) backbone. While concurrent LLMbased approaches have demonstrated impressive capabilities in multimodal and multitask scenarios; our work introduces novel mechanisms for task-adaptive learning and embedding extraction that further enhance the potential of LLM-based retrieval systems. Our key technical contribution lies in the development of a task-aware contrastive learning framework with an automated Bayesian weighing mechanism. This approach provides a principled way to balance multiple tasks during training, departing from conventional contrastive learning strategies. We further enhance the framework through a multiple token summarization strategy and an auxiliary language modeling objective, which together significantly improve retrieval performance. Comprehensive experiments on M-BEIR and ICinW benchmarks demonstrate the effectiveness of M3T-UEM, showing competitive or superior performance compared to both traditional encoder-based methods and recent LLMbased approaches. Furthermore, we demonstrate particular strengths in handling compositional conceptual changes and multilingual scenarios owing to the incorporation of an LLM backbone where the method drastically outperforms CLIP in zero-shot settings, often by orders of magnitude.∗

Original languageEnglish
Title of host publicationProceedings - 2025 IEEE/CVF International Conference on Computer Vision, ICCV 2025
PublisherInstitute of Electrical and Electronics Engineers Inc.
Pages22783-22793
Number of pages11
ISBN (Electronic)9798331587758
DOIs
StatePublished - 2025
Event2025 IEEE/CVF International Conference on Computer Vision, ICCV 2025 - Honolulu, United States
Duration: Oct 19 2025Oct 23 2025

Publication series

NameProceedings of the IEEE International Conference on Computer Vision
ISSN (Print)1550-5499
ISSN (Electronic)2380-7504

Conference

Conference2025 IEEE/CVF International Conference on Computer Vision, ICCV 2025
Country/TerritoryUnited States
CityHonolulu
Period10/19/2510/23/25

Keywords

  • contrastive learning
  • embedding models
  • llm
  • multi-modal retrieval

Fingerprint

Dive into the research topics of 'Multi-Modal Multi-Task Unified Embedding Model (M3T-UEM): A Task-Adaptive Representation Learning Framework'. Together they form a unique fingerprint.

Cite this