Skip to main navigation Skip to search Skip to main content

Assessing the ability of neural TTS systems to model consonant-induced f0 perturbation

  • SUNY Buffalo
  • Australian National University

Research output: Contribution to journalArticlepeer-review

Abstract

This study proposes a segmental-level prosodic probing framework to evaluate neural TTS models’ ability to reproduce consonant-induced f0 perturbation, a fine-grained segmental-prosodic effect that reflects local articulatory mechanisms. We compare synthetic and natural speech realizations for thousands of words, stratified by lexical frequency, using Tacotron 2 and FastSpeech 2 trained on the same speech corpus (LJ Speech). These controlled analyses are then complemented by a large-scale evaluation spanning multiple advanced TTS systems. Results show accurate reproduction for high-frequency words but poor generalization to low-frequency items, suggesting that the examined TTS architectures rely more on lexical-level memorization rather than abstract segmental-prosodic encoding. This finding highlights a limitation in such TTS systems’ ability to generalize prosodic detail beyond seen data. The proposed probe offers a linguistically informed diagnostic framework that may inform future TTS evaluation methods, and has implications for interpretability and authenticity assessment in synthetic speech.

Original languageEnglish
Article number101983
JournalComputer Speech and Language
Volume100
DOIs
StatePublished - Oct 2026

Keywords

  • F0 perturbation
  • Phonetic modeling
  • Pitch prediction
  • Text-to-speech synthesis

Fingerprint

Dive into the research topics of 'Assessing the ability of neural TTS systems to model consonant-induced f0 perturbation'. Together they form a unique fingerprint.

Cite this