HMM-based speech synthesis with various degrees of articulation: A perceptual study

详细信息查看全文

作者：Benjamin Picart ; ^{benjamin.picart@umons.ac.be" class="auth_mail}Author Vitae ; Thomas DrugmanAuthor Vitae ; Thierry DutoitAuthor Vitae
关键词：Speech synthesis ; Expressive speech ; Speaking style adaptation ; Perceptual effects ; Voice quality
刊名：Neurocomputing
出版年：20 May, 2014
年：2014
卷：132
期：Complete
页码：142-147
全文大小：811 K

文摘

HMM-based speech synthesis is very convenient for creating a synthesizer whose speaker characteristics and speaking styles can be easily modified. This can be obtained by adapting a source speaker's model to a target speaker's model, using intra-speaker voice adaptation techniques. In this paper, we focus on high-quality HMM-based speech synthesis integrating various degrees of articulation, and more specifically on the internal mechanisms leading to the perception of the degrees of articulation by listeners. Therefore the process of adapting a neutral speech synthesizer to generate hypo and hyperarticulated speech is broken down into four factors: cepstrum, prosody, phonetic transcription adaptation as well as the complete adaptation. The impact of these factors on the perceived degree of articulation is studied. Moreover, this study is complemented with an Absolute Category Rating (ACR) evaluation, allowing the subjective assessment of hypo/hyperarticulated speech through various dimensions: comprehension, non-monotony, fluidity and pronunciation. This paper quantifies the importance of prosody and cepstrum adaptation as well as the use of a Natural Language Processor able to generate realistic hypo and hyperarticulated phonetic transcriptions.

地址：北京市海淀区学院路29号邮编：100083

电话：办公室：(+86 10)66554848；文献借阅、咨询服务、科技查新：66554700