Fourth European Conference on Speech Communication and Technology

Madrid, Spain
September 18-21, 1995

Optimising Selection of Units from Speech Databases for Concatenative Synthesis

Alan W. Black, Nick Campbell

ATR Interpreting Telecommunications Research Laboratories, Soraku-gun, Kyoto, Japan

Concatenating units of natural speech is one method of speech synthesis1. Most such systems use an inventory of fixed length units, typically diphones or triphones with one instance of each type. An alternative is to use more varied, non-uniform units extracted from large speech databases containing multiple instances of each. The greater variability in such natural speech segments allows closer modeling of naturalness and differences in speaking styles, and eliminates the need for specially-recorded, single-use databases. However, with the greater variability comes the problem of how to select between the many instances of units in the database. This paper addresses that issue and presents a general method for unit selection.

