ISCA Archive Interspeech 2013
ISCA Archive Interspeech 2013

Deep segmental neural networks for speech recognition

Ossama Abdel-Hamid, Li Deng, Dong Yu, Hui Jiang

Hybrid systems which integrate the deep neural network (DNN) and hidden Markov model (HMM) have recently achieved remarkable performance in many large vocabulary speech recognition tasks. These systems, however, remain to rely on the HMM and assume the acoustic scores for the (windowed) frames are independent given the state, suffering from the same difficulty as in the previous GMM-HMM systems. In this paper, we propose the deep segmental neural network (DSNN), a segmental model that uses DNNs to estimate the acoustic scores of phonemic or sub-phonemic segments with variable lengths. This allows the DSNN to represent each segment as a single unit, in which frames are made dependent on each other. We describe the architecture of the DSNN, as well as its learning and decoding algorithms. Our evaluation experiments demonstrate that the DSNN can outperform the DNN/HMM hybrid systems and two existing segmental models including the segmental conditional random field and the shallow segmental neural network.

doi: 10.21437/Interspeech.2013-455

Cite as: Abdel-Hamid, O., Deng, L., Yu, D., Jiang, H. (2013) Deep segmental neural networks for speech recognition. Proc. Interspeech 2013, 1849-1853, doi: 10.21437/Interspeech.2013-455

  author={Ossama Abdel-Hamid and Li Deng and Dong Yu and Hui Jiang},
  title={{Deep segmental neural networks for speech recognition}},
  booktitle={Proc. Interspeech 2013},