First Workshop on Speech, Language and Audio in Multimedia (SLAM 2013)

Marseille, France
August 22-23, 2013

LMELECTURES: A Multimedia Corpus of Academic Spoken English

Korbinian Riedhammer, Martin Gropp, Tobias Bocklet, Florian Hönig, Elmar Nöth, Stefan Steidl

Pattern Recognition Lab, University of Erlangen-Nuremberg, Germany

This paper describes the acquisition, transcription and annotation of a multi-media corpus of academic spoken English, the LMELectures. It consists of two lecture series that were read in the summer term 2009 at the computer science department of the University of Erlangen- Nuremberg, covering topics in pattern analysis, machine learning and interventional medical image processing. In total, about 40 hours of high-definition audio and video of a single speaker was acquired in a constant recording environment. In addition to the recordings, the presentation slides are available in machine readable (PDF) format. The manual annotations include a suggested segmentation into speech turns and a complete manual transcription that was done using BLITZSCRIBE2, a new tool for the rapid transcription. For one lecture series, the lecturer assigned key words to each recordings; one recording of that series was further annotated with a list of ranked key phrases by five human annotators each. The corpus is available for non-commercial purpose upon request.

Index Terms: corpus description, academic spoken English, e-learning

Full Paper

Bibliographic reference.  Riedhammer, Korbinian / Gropp, Martin / Bocklet, Tobias / Hönig, Florian / Nöth, Elmar / Steidl, Stefan (2013): "LMELECTURES: a multimedia corpus of academic spoken English", In SLAM-2013, 102-107.