Scoring Metrics of Assessing Voiceprint Distinctiveness Based on Speech Content and Rate

Nov.-Dec. 2024, pp. 5582-5599, vol. 21
DOI Bookmark: 10.1109/TDSC.2024.3380603

Authors

Ruiwen He, College of Electrical Engineering, Zhejiang University, Hangzhou, China
Yushi Cheng, Ubiquitous System Security Lab (USSLAB), ZJU-UIUC Institute, Zhejiang University, Hangzhou, China
Junning Ze, College of Electrical Engineering, Zhejiang University, Hangzhou, China
Xinfeng Li, College of Electrical Engineering, Zhejiang University, Hangzhou, China
Xiaoyu Ji, College of Electrical Engineering, Zhejiang University, Hangzhou, China
Wenyuan Xu, College of Electrical Engineering, Zhejiang University, Hangzhou, China

Keywords

Spectrogram, Speech Recognition, Phonetics, Personal Voice Assistants, Internet, Analytical Models, Authentication, Speaker Verification, AI Security, Statistical Analysis, Speech Rate, Speech Content, False Rate, Test Samples, Commercial Products, Virtual Assistant, English Language, Training Dataset, Number Of Types, Test Dataset, Distribution Of Rates, Model Verification, Model In This Paper, Enrolment Rates, Common Observation, Real Words, Audio Data, Speaker Recognition, Threat Model, Speech Analysis, False Acceptance Rate, Sensitivity Scenarios, Form Of Boxplots, Interjections, Speech Detection, DNN Model, Effective Rate, Speech Training, Equal Error Rate, Part Of Speech

Abstract

A voiceprint is the distinctive pattern of human voices widely used for authentication in voice assistants. This article investigates the impact of speech contents and speech rates on the distinctiveness of voiceprint, and has obtained answers to three questions by studying 2457 speakers and 21,500,000 test samples: 1) What are the influential factors that users can control to affect the distinctiveness of voiceprints? 2) How to quantify the distinctiveness for given speeches, e.g., the speech of wake-up words when activating voice assistants? 3) How to help users select wake-up words and adjust the speech rate to improve distinctiveness levels? To answer those questions, we break down speeches into phones, and experimentally obtain the correlation between false recognition rates and the richness, order, length, and elements of the phones. Then, we define the PROLE Score that can reflect the voice distinctiveness, and evaluate 30 wake-up words of 19 commercial voice assistant products to provide recommendations on selecting secure voiceprint words. We also measure the correlation between false recognition rates and speech rates, and define the TER Score that reveals the distance of distinctiveness from the secure voiceprint, and it guides users to adjust their speech rate to a secure value.

References

1 PROLE: phonetic richness, order, length, elements.