Speaker clustering for speech recognition using the parameters characterizing vocal tract dimensions

Masaki Naito; Li Deng; Yoshinori Sagisaka

Speaker clustering for speech recognition using the parameters characterizing vocal tract dimensions

Masaki Naito ,
Li Deng ,
Yoshinori Sagisaka

Proc. ICASSP | March 1998

Download BibTex

We propose speaker clustering methods based on the vocal tractsize related articulatory parameters associated with individual speakers. Two parameters characterizing gross vocal-tract dimensions are first derived from formants of speaker-specific Japanese vowels, and are then used to cluster a total of 148 male Japanese speakers. The resultant speaker clusters are found to be significantly different from the speaker clusters obtained by conventional acoustic criteria. Japanese phoneme recognition experiments are carried out using speaker-clustered tied-state HMMs (HMNets) trained for each cluster. Compared with the baseline gender dependent model, 5.7% of recognition error reduction has been achieved based on the clustering method using vocal-tract parameters.