Transfer Learning of Articulatory Information Through Phone Information

Abdolreza Sabzi Shahrebabaki, Negar Olfati, Sabato Marco Siniscalchi, Giampiero Salvi, Torbjørn Svendsen


Articulatory information has been argued to be useful for several speech tasks. However, in most practical scenarios this information is not readily available. We propose a novel transfer learning framework to obtain reliable articulatory information in such cases. We demonstrate its reliability both in terms of estimating parameters of speech production and its ability to enhance the accuracy of an end-to-end phone recognizer. Articulatory information is estimated from speaker independent phonemic features, using a small speech corpus, with electro-magnetic articulography (EMA) measurements. Next, we employ a teacher-student model to learn estimation of articulatory features from acoustic features for the targeted phone recognition task. Phone recognition experiments, demonstrate that the proposed transfer learning approach outperforms the baseline transfer learning system acquired directly from an acoustic-to-articulatory (AAI) model. The articulatory features estimated by the proposed method, in conjunction with acoustic features, improved the phone error rate (PER) by 6.7% and 6% on the TIMIT core test and development sets, respectively, compared to standalone static acoustic features. Interestingly, this improvement is slightly higher than what is obtained by static+dynamic acoustic features, but with a significantly less. Adding articulatory features on top of static+dynamic acoustic features yields a small but positive PER improvement.


 DOI: 10.21437/Interspeech.2020-1139

Cite as: Shahrebabaki, A.S., Olfati, N., Siniscalchi, S.M., Salvi, G., Svendsen, T. (2020) Transfer Learning of Articulatory Information Through Phone Information. Proc. Interspeech 2020, 2877-2881, DOI: 10.21437/Interspeech.2020-1139.


@inproceedings{Shahrebabaki2020,
  author={Abdolreza Sabzi Shahrebabaki and Negar Olfati and Sabato Marco Siniscalchi and Giampiero Salvi and Torbjørn Svendsen},
  title={{Transfer Learning of Articulatory Information Through Phone Information}},
  year=2020,
  booktitle={Proc. Interspeech 2020},
  pages={2877--2881},
  doi={10.21437/Interspeech.2020-1139},
  url={http://dx.doi.org/10.21437/Interspeech.2020-1139}
}