Ninth International Conference on Spoken Language Processing

Pittsburgh, PA, USA
September 17-21, 2006

Efficient Gaussian Mixture Model Evaluation in Voice Conversion

Jilei Tian, Jani Nurminen, Victor Popa

Nokia Research Center, Finland

Voice conversion refers to the adaptation of the characteristics of a source speakerís voice to those of a target speaker. Gaussian mixture models (GMM) have been found to be efficient in the voice conversion task. The GMM parameters are estimated from a training set with the goal to minimize the mean squared error (MSE) between the transformed and target vectors. Obviously, the quality of the GMM model plays an important role in achieving better voice conversion quality. This paper presents a very efficient approach for the evaluation of GMM models directly from the model parameters without using any test data, facilitating the improvement of the transformation performance especially in the case of embedded implementations. Though the proposed approach can be used in any application that utilizes GMM based transformation, we take voice conversion as an example application throughout the paper. The proposed approach is experimented with in this context and evaluated against an MSE based evaluation method. The results show that the proposed method is in line with all subjective observations and MSE results.

Full Paper

Bibliographic reference.  Tian, Jilei / Nurminen, Jani / Popa, Victor (2006): "Efficient Gaussian mixture model evaluation in voice conversion", In INTERSPEECH-2006, paper 1533-Thu1BuP.9.