13th Annual Conference of the International Speech Communication Association

Portland, OR, USA
September 9-13, 2012

Using Broad Phonetic Classes to Guide Search in Automatic Speech Recognition

Stefan Ziegler, Bogdan Ludusan, Guillaume Gravier

CNRS – IRISA, Campus de Beaulieu, Rennes, France

This work presents a novel framework to guide the Viterbi decoding process of a hidden Markov model based speech recognition system by means of broad phonetic classes. In a first step, decision trees are employed, along with frame and segment based attributes, in order to detect broad phonetic classes in the speech signal. Then, the detected phonetic classes are used to reinforce paths in the search process, either at every frame or at phonetically significant landmarks. Results obtained on French broadcast news data show a relative improvement in word error rate of about 2% with respect to the baseline.

Index Terms: Viterbi decoding, broad phonetic classes, landmarks

Full Paper

Bibliographic reference.  Ziegler, Stefan / Ludusan, Bogdan / Gravier, Guillaume (2012): "Using broad phonetic classes to guide search in automatic speech recognition", In INTERSPEECH-2012, 1023-1026.