Title :
Speaker Diarization Exploiting the Eigengap Criterion and Cluster Ensembles
Author :
Bassiou, Nikoletta ; Moschou, Vassiliki ; Kotropoulos, Constantine
Author_Institution :
Dept. of Inf., Aristotle Univ. of Thessaloniki, Thessaloniki, Greece
Abstract :
A novel system for speaker diarization is proposed that combines the eigengap criterion and cluster ensembles. No explicit assumptions on the number of speakers are made. Two variants of the system are developed. The first variant does not cluster the speech segments that are detected as outliers, while the second one does. The aforementioned system variants are assessed with respect to various metrics, such as the overall classification error, the average cluster purity, and the average speaker purity. Experiments are conducted on two-person dialogue scenes in movies as well as on news broadcasts from MDE RT-03 Training Data Speech Corpus released by the U.S. National Institute of Standards and Technology. In the latter case, the diarization error rate is also reported. It is demonstrated that the clustering performance does not degrade when outliers are present. Moreover, thanks to the eigengap criterion, the evaluation metrics are improved.
Keywords :
pattern clustering; speaker recognition; speech processing; MDE RT-03 Training Data Speech Corpus; U.S. National Institute of Standards and Technology; average cluster purity; average speaker purity; classification error; cluster ensembles; eigengap criterion; speaker diarization; speech segments; Audio recording; Degradation; Layout; Loudspeakers; MPEG 7 Standard; Model driven engineering; Motion pictures; NIST; Speech; TV broadcasting; Broadcasts; cluster ensembles; eigengap criterion; movie scene analysis; speaker clustering; speaker diarization; two-person dialogues;
Journal_Title :
Audio, Speech, and Language Processing, IEEE Transactions on
DOI :
10.1109/TASL.2010.2042121