5th International Conference on Spoken Language Processing

Sydney, Australia
November 30 - December 4, 1998

Collection and Detailed Transcription of a Speech Database for Development of Language Learning Technologies

Harry Bratt, Leonardo Neumeyer, Elizabeth Shriberg, Horacio Franco

SRI International, USA

We describe the methodologies for collecting and annotating a Latin-American Spanish speech database. The database includes recordings by native and nonnative speakers. The nonnative recordings are annotated with ratings of pronunciation quality and detailed phonetic transcriptions. We use the annotated database to investigate rater reliability, the effect of each phone on overall perceived nonnativeness, and the frequency of specific pronunciation errors.

