ISCA Archive Interspeech 2013
ISCA Archive Interspeech 2013

Strategies for high accuracy keyword detection in noisy channels

Arindam Mandal, Julien van Hout, Yik-Cheung Tam, Vikramjit Mitra, Yun Lei, Jing Zheng, Dimitra Vergyri, Luciana Ferrer, Martin Graciarena, Andreas Kathol, Horacio Franco

We present design strategies for a keyword spotting (KWS) system that operates in highly degraded channel conditions with very low signal-to-noise ratio levels. We employ a system combination approach by combining the outputs of multiple large vocabulary automatic speech recognition (LVCSR) systems, each of which employs a different system design approach targeting three different levels of information: front-end signal processing features (standard cepstra-based, noise-robust modulation and multi layer perceptron features), statistical acoustic models (gaussian mixtures models (GMM) and subspace GMMs) and keyword search strategies (word-based and phone-based). We also use keyword-aware capabilities in the system at two levels: in the LVCSR language models by assigning higher weights to n-grams with keywords in them and in LVCSR search by using a relaxed pruning threshold for keywords. The LVCSR system outputs are represented as lattice-based unigram indices whose scores are fused by a logistic-regression based classifier to produce the final system combination output. We present the performance of our system in the phase II evaluations of DARPA's Robust Automatic Transcription of Speech (RATS) program for both Levantine Arabic and Farsi conversational speech corpora.

doi: 10.21437/Interspeech.2013-4

Cite as: Mandal, A., Hout, J.v., Tam, Y.-C., Mitra, V., Lei, Y., Zheng, J., Vergyri, D., Ferrer, L., Graciarena, M., Kathol, A., Franco, H. (2013) Strategies for high accuracy keyword detection in noisy channels. Proc. Interspeech 2013, 15-19, doi: 10.21437/Interspeech.2013-4

  author={Arindam Mandal and Julien van Hout and Yik-Cheung Tam and Vikramjit Mitra and Yun Lei and Jing Zheng and Dimitra Vergyri and Luciana Ferrer and Martin Graciarena and Andreas Kathol and Horacio Franco},
  title={{Strategies for high accuracy keyword detection in noisy channels}},
  booktitle={Proc. Interspeech 2013},