13th Annual Conference of the International Speech Communication Association

Portland, OR, USA
September 9-13, 2012

Bag-of-Audio-Words Approach for Multimedia Event Classification

Stephanie Pancoast (1,2), Murat Akbacak (1)

(1) Speech Technology and Research Lab, SRI International, Menlo Park, CA, USA
(2) Department of Electrical Engineering, Stanford University, Stanford, CA, USA

With the popularity of online multimedia videos, there has been much interest in recent years in acoustic event detection and classification for the improvement of online video search. The audio component of a video has the potential to contribute significantly to multimedia event classification. Recent research in audio document classification has drawn parallels to text and image document retrieval by employing what is referred to as the bag-of-audio words (BoAW) method. Compared to supervised approaches where audio concept detectors are trained using annotated data and extracted labels are used as lowlevel features for multimedia event classification. The BoAW approach extracts audio concepts in an unsupervised fashion. Hence this method has the advantage that it can be employed easily for a new set of audio concepts in multimedia videos without going through a laborious annotation effort. In this paper, we explore variations of the BoAW method and present results on NIST 2011 multimedia event detection (MED) dataset.

Index Terms: Bag-of-audio-words, multimedia event detection

Full Paper

Bibliographic reference.  Pancoast, Stephanie / Akbacak, Murat (2012): "Bag-of-audio-words approach for multimedia event classification", In INTERSPEECH-2012, 2105-2108.