Phonological Feature Based Mispronunciation Detection and Diagnosis Using Multi-Task DNNs and Active Learning

Vipul Arora, Aditi Lahiri, Henning Reetz


This paper presents a phonological feature based computer aided pronunciation training system for the learners of a new language (L2). Phonological features allow analysing the learners’ mispronunciations systematically and rendering the feedback more effectively. The proposed acoustic model consists of a multi-task deep neural network, which uses a shared representation for estimating the phonological features and HMM state probabilities. Moreover, an active learning based scheme is proposed to efficiently deal with the cost of annotation, which is done by expert teachers, by selecting the most informative samples for annotation. Experimental evaluations are carried out for German and Italian native-speakers speaking English. For mispronunciation detection, the proposed feature-based system outperforms conventional GOP measure and classifier based methods, while providing more detailed diagnosis. Evaluations also demonstrate the advantage of active learning based sampling over random sampling.


 DOI: 10.21437/Interspeech.2017-1350

Cite as: Arora, V., Lahiri, A., Reetz, H. (2017) Phonological Feature Based Mispronunciation Detection and Diagnosis Using Multi-Task DNNs and Active Learning. Proc. Interspeech 2017, 1432-1436, DOI: 10.21437/Interspeech.2017-1350.


@inproceedings{Arora2017,
  author={Vipul Arora and Aditi Lahiri and Henning Reetz},
  title={Phonological Feature Based Mispronunciation Detection and Diagnosis Using Multi-Task DNNs and Active Learning},
  year=2017,
  booktitle={Proc. Interspeech 2017},
  pages={1432--1436},
  doi={10.21437/Interspeech.2017-1350},
  url={http://dx.doi.org/10.21437/Interspeech.2017-1350}
}