4th International Conference on Spoken Language Processing

Philadelphia, PA, USA
October 3-6, 1996

Inclusion of Temporal Information into Features for Speech Recognition

Ben Milner

Speech Technology Unit, BT Laboratories, Martlesham Heath, Suffolk, England, UK

Conventional methods for incorporating temporal information into speech features apply regression to a series of successive cepstral vectors to generate differential cepstra, or apply a cosine transform to generate cepstral-time matrices. This paper aims to generalise these techniques such that a series of stacked cepstral vectors is multiplied by a temporal transform matrix to produce the final speech feature. This can made to incorporate both static and dynamic speech information. Using this method, the coding of temporal information is not restricted to regression or cosine coefficients - any suitable transform may used. Results are presented for a variety of transforms, such as Legendre, Karhunen-Loeve, Cosine, Rectangle, where it is shown that the transform based techniques offer higher performance than conventional differential cepstrum.

Full Paper

Bibliographic reference.  Milner, Ben (1996): "Inclusion of temporal information into features for speech recognition", In ICSLP-1996, 256-259.