Stochastic feature compensation methods for speaker verification in noisy environments

详细信息查看全文

作者：Sourjya Sarkar ; ^{sourjyasarkar@gmail.com" class="auth_mail} ; K. Sreenivasa Rao ^{ksrao@iitkgp.ac.in" class="auth_mail}
关键词：Speaker verification ; Noisy environment ; Minimum mean squared error ; Maximum likelihood estimate ; Expectation Maximization algorithm ; Gaussian Mixture Models
刊名：Applied Soft Computing Journal
出版年：June, 2014
年：2014
卷：19
期：Complete
页码：198-214
全文大小：2980 K

文摘

This paper explores the significance of stereo-based stochastic feature compensation (SFC) methods for robust speaker verification (SV) in mismatched training and test environments. Gaussian Mixture Model (GMM)-based SFC methods developed in past has been solely restricted for speech recognition tasks. Application of these algorithms in a SV framework for background noise compensation is proposed in this paper. A priori knowledge about the test environment and availability of stereo training data is assumed. During the training phase, Mel frequency cepstral coefficient (MFCC) features extracted from a speaker's noisy and clean speech utterance (stereo data) are used to build front end GMMs. During the evaluation phase, noisy test utterances are transformed on the basis of a minimum mean squared error (MMSE) or maximum likelihood (MLE) estimate, using the target speaker GMMs. Experiments conducted on the NIST-2003-SRE database with clean speech utterances artificially degraded with different types of additive noises reveal that the proposed SV systems strictly outperform baseline SV systems in mismatched conditions across all noisy background environments.

地址：北京市海淀区学院路29号邮编：100083

电话：办公室：(+86 10)66554848；文献借阅、咨询服务、科技查新：66554700