13th Annual Conference of the International Speech Communication Association

Portland, OR, USA
September 9-13, 2012

Discriminatively Trained Phoneme Confusion Model for Keyword Spotting

Panagiota Karanasou (1), Lukas Burget (2), Dimitra Vergyri (2), Murat Akbacak (2), Arindam Mandal (2)

(1) LIMSI/CNRS, Université Paris-Sud, Orsay, France
(2) Speech Technology and Research Laboratory, SRI International, Menlo Park, CA, USA

Keyword Spotting (KWS) aims at detecting speech segments that contain a given query within large amounts of audio data. Typically, a speech recognizer is involved in a first indexing step. One of the challenges of KWS is how to handle recognition errors and out-of-vocabulary (OOV) terms. This work proposes the use of discriminative training to construct a phoneme confusion model, which expands the phonemic index of a KWS system by adding phonemic variation to handle the above-mentioned problems. The objective function that is optimized is the Figure of Merit (FOM), which is directly related to the KWS performance. The experiments conducted on English data sets show some improvement on the FOM and are promising for the use of such technique.

Index Terms: keyword spotting, confusion model, discriminative training, Figure of Merit

Full Paper

Bibliographic reference.  Karanasou, Panagiota / Burget, Lukas / Vergyri, Dimitra / Akbacak, Murat / Mandal, Arindam (2012): "Discriminatively trained phoneme confusion model for keyword spotting", In INTERSPEECH-2012, 2434-2437.