Skip to main navigation Skip to search Skip to main content

Assessing agreement level between forced alignment models with data from endangered language documentation corpora

  • Christian T. DiCanio
  • , Hosung Nam
  • , D. H. Whalen
  • , H. Timothy Bunnell
  • , Jonathan D. Amith
  • , Rey Castillo García
  • Haskins Laboratories
  • City University of New York
  • Endangered Language Fund, Inc.
  • Alfred I. duPont Hospital for Children
  • University of Delaware
  • Gettysburg College
  • Smithsonian Institution
  • Centro de Investigaciones y Estudios Superiores en Antropología Social

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

9 Scopus citations

Abstract

Automatic forced alignment between transcriptions has achieved high levels of agreement for languages with large corpora, but the technique holds great promise for work on all languages. Here, we apply two forced alignment programs to data from an endangered Mixtecan language of Mexico. Both yielded a majority of boundaries within 20 ms of hand-labeled ones. Phonemes with fairly steady-state elements (e.g. nasals, fricatives) were more accurately labeled than others. Forced alignment thus may increase efficiency of labeling texts from smaller languages, at least in cases where the phoneme inventories are similar to those of the languages of the training.

Original languageEnglish
Title of host publication13th Annual Conference of the International Speech Communication Association 2012, INTERSPEECH 2012
Pages130-133
Number of pages4
StatePublished - 2012
Event13th Annual Conference of the International Speech Communication Association 2012, INTERSPEECH 2012 - Portland, OR, United States
Duration: Sep 9 2012Sep 13 2012

Publication series

Name13th Annual Conference of the International Speech Communication Association 2012, INTERSPEECH 2012
Volume1

Conference

Conference13th Annual Conference of the International Speech Communication Association 2012, INTERSPEECH 2012
Country/TerritoryUnited States
CityPortland, OR
Period09/9/1209/13/12

Keywords

  • Linguistics
  • Phonetics
  • Speech recognition

Fingerprint

Dive into the research topics of 'Assessing agreement level between forced alignment models with data from endangered language documentation corpora'. Together they form a unique fingerprint.

Cite this