A Hybrid HMM-Waveglow Based Text-to-Speech Synthesizer Using Histogram Equalization for Low Resource Indian Languages

Mano Ranjith Kumar M., Sudhanshu Srivastava, Anusha Prakash, Hema A. Murthy


Conventional text-to-speech (TTS) synthesis requires extensive linguistic processing for producing quality output. The advent of end-to-end (E2E) systems has caused a relocation in the paradigm with better synthesized voices. However, hidden Markov model (HMM) based systems are still popular due to their fast synthesis time, robustness to less training data, and flexible adaptation of voice characteristics, speaking styles, and emotions.

This paper proposes a technique that combines the classical parametric HMM-based TTS framework (HTS) with the neural-network-based Waveglow vocoder using histogram equalization (HEQ) in a low resource environment. The two paradigms are combined by performing HEQ across mel-spectrograms extracted from HTS generated audio and source spectra of training data. During testing, the synthesized mel-spectrograms are mapped to the source spectrograms using the learned HEQ. Experiments are carried out on Hindi male and female dataset of the Indic TTS database. Systems are evaluated based on degradation mean opinion scores (DMOS). Results indicate that the synthesis quality of the hybrid system is better than that of the conventional HTS system. These results are quite promising as they pave way to good quality TTS systems with less data compared to E2E systems.


 DOI: 10.21437/Interspeech.2020-3180

Cite as: M., M.R.K., Srivastava, S., Prakash, A., Murthy, H.A. (2020) A Hybrid HMM-Waveglow Based Text-to-Speech Synthesizer Using Histogram Equalization for Low Resource Indian Languages. Proc. Interspeech 2020, 2037-2041, DOI: 10.21437/Interspeech.2020-3180.


@inproceedings{M.2020,
  author={Mano Ranjith Kumar M. and Sudhanshu Srivastava and Anusha Prakash and Hema A. Murthy},
  title={{A Hybrid HMM-Waveglow Based Text-to-Speech Synthesizer Using Histogram Equalization for Low Resource Indian Languages}},
  year=2020,
  booktitle={Proc. Interspeech 2020},
  pages={2037--2041},
  doi={10.21437/Interspeech.2020-3180},
  url={http://dx.doi.org/10.21437/Interspeech.2020-3180}
}