Skip to main navigation Skip to search Skip to main content

MCAD: Multimodal Context-Aware Audio Description Generation for Soccer

  • Lipisha Chaudhary
  • , Trisha Mittal
  • , Subhadra Gopalakrishnan
  • , Ifeoma Nwogu
  • , Jaclyn Pytlarz
  • SUNY Buffalo
  • Dolby Laboratories

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Audio Descriptions (AD) are essential for making visual content accessible to individuals with visual impairments. Recent works have shown a promising step towards automating AD, but they have been limited to describing high-quality movie content using human-annotated ground truth AD in the process. In this work, we present an end-to-end pipeline, MCAD, that extends AD generation beyond movies to the domain of sports, with a focus on soccer games, without relying on ground truth AD. To address the absence of domain-specific AD datasets, we fine-tune a Video Large Language Model on publicly available movie AD datasets so that it learns the narrative structure and conventions of AD. During inference, MCAD incorporates multimodal contextual cues such as player identities, soccer events/actions, and commentary from the game. These cues, combined with input prompts to the fine-tuned VideoLLM, allow the system to produce complete AD text for each video segment. We further introduce a new evaluation metric, A R G E-A D, designed to accurately assess the quality of generated AD. ARGE-AD evaluates the generated AD for the presence of five characteristics: (i) usage of people's names, (ii) mention of actions/events, (iii) appropriate length of AD, (iv) absence of pronouns, and (v) overlap from commentary/subtitles. We present an in-depth analysis of our approach on both movie and soccer datasets. We also validate the use of this metric to quantitatively comment on the quality of generated AD using our metric across domains. Additionally, we contribute audio descriptions for 100 soccer game clips annotated by two AD experts.

Original languageEnglish
Title of host publicationProceedings - 2025 International Symposium on Multimedia, ISM 2025
PublisherInstitute of Electrical and Electronics Engineers Inc.
Pages280-287
Number of pages8
ISBN (Electronic)9798331545901
DOIs
StatePublished - 2025
Event27th International Symposium on Multimedia, ISM 2025 - Naples, Italy
Duration: Dec 8 2025Dec 10 2025

Publication series

NameProceedings - 2025 International Symposium on Multimedia, ISM 2025

Conference

Conference27th International Symposium on Multimedia, ISM 2025
Country/TerritoryItaly
CityNaples
Period12/8/2512/10/25

Keywords

  • Automatic Audio Description Generation
  • MLLM
  • Reference-Free Metric
  • Sports Video Analysis

Fingerprint

Dive into the research topics of 'MCAD: Multimodal Context-Aware Audio Description Generation for Soccer'. Together they form a unique fingerprint.

Cite this