Skip to main navigation Skip to search Skip to main content

Speaker-Conditioned Hierarchical Modeling for Automated Speech Scoring

  • Yaman Kumar Singla
  • , Avyakt Gupta
  • , Shaurya Bagga
  • , Changyou Chen
  • , Balaji Krishnamurthy
  • , Rajiv Ratn Shah
  • Indraprastha Institute of Information Technology Delhi
  • SUNY Buffalo
  • Adobe Systems Incorporated

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

11 Scopus citations

Abstract

Automatic Speech Scoring (ASS) is the computer-assisted evaluation of a candidate's speaking proficiency in a language. ASS systems face many challenges like open grammar, variable pronunciations, and unstructured or semi-structured content. Recent deep learning approaches have shown some promise in this domain. However, most of these approaches focus on extracting features from single audio, making them suffer from the lack of speaker-specific context required to model such a complex task. We propose a novel deep learning technique for non-native ASS, called speaker-conditioned hierarchical modelling. In our technique, we take advantage of the fact that oral proficiency tests rate multiple responses for a candidate. We extract context vectors from these responses and feed them as additional speaker-specific context to our network to score a particular response. We compare our technique with strong baselines and find that such modelling improves the model's average performance by 6.92% (maximum = 12.86%, minimum = 4.51%). We further show both quantitative and qualitative insights into the importance of this additional context in solving the problem of ASS.

Original languageEnglish
Title of host publicationCIKM 2021 - Proceedings of the 30th ACM International Conference on Information and Knowledge Management
PublisherAssociation for Computing Machinery
Pages1681-1691
Number of pages11
ISBN (Electronic)9781450384469
DOIs
StatePublished - Oct 30 2021
Event30th ACM International Conference on Information and Knowledge Management, CIKM 2021 - Virtual, Online, Australia
Duration: Nov 1 2021Nov 5 2021

Publication series

NameInternational Conference on Information and Knowledge Management, Proceedings
ISSN (Print)2155-0751

Conference

Conference30th ACM International Conference on Information and Knowledge Management, CIKM 2021
Country/TerritoryAustralia
CityVirtual, Online
Period11/1/2111/5/21

Keywords

  • ai in education
  • automated speech scoring
  • end-to-end neural networks
  • hierarchical modeling
  • interpretability in ai
  • multi-modal deep learning
  • spontaneous speech

Fingerprint

Dive into the research topics of 'Speaker-Conditioned Hierarchical Modeling for Automated Speech Scoring'. Together they form a unique fingerprint.

Cite this