| 研究生: |
錡軒誼 Chi, Hsuan-Yi |
|---|---|
| 論文名稱: |
基於深度學習框架之耳語語者辨識研究 Whispered Speech Speaker Recognition Based on Deep Learning Frameworks |
| 指導教授: |
廖文宏
Liao, Wen-Hung |
| 口試委員: |
彭彥璁
Peng, Yan-Tsung 陳駿丞 Chen, Jun-Cheng |
| 學位類別: |
碩士
Master |
| 系所名稱: |
資訊學院 - 資訊科學系 Department of Computer Science |
| 論文出版年: | 2026 |
| 畢業學年度: | 115 |
| 語文別: | 中文 |
| 論文頁數: | 72 |
| 中文關鍵詞: | 耳語語音 、語音辨識 、自監督式語音表徵 、領域偏移 、特徵轉換 |
| 外文關鍵詞: | Whispered Speech, Speaker Recognition, Self-Supervised Speech Representation, Domain Shift, Feature Transformation |
| 相關次數: | 點閱:52 下載:3 |
| 分享至: |
| 查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報 |
本研究旨在探討並改善正常語音與耳語語音之間的跨語音型態語者辨識問題。耳語語音因缺乏聲帶振動,導致基頻與諧波結構消失,使其嵌入特徵與正常語音之間存在顯著的分布落差,進而影響語者辨識模型於跨型態條件下之穩定性與準確性。
為改善此問題,本研究結合自監督式學習(Self-Supervised Learning, SSL)語音模型所擷取之高階語音嵌入,並提出一個基於 Transformer 架構之特徵轉換模型,用以學習耳語與正常語音嵌入之間的非線性映射關係。透過均方誤差損失(Mean Squared Error, MSE)與三元組損失(Triplet Loss)之共同優化,模型可對齊跨語音型態之嵌入分布,同時保留語者之區辨特性。
實驗採用自建之 Whisper-100 語料庫,並於 Normal → Whisper 與 Whisper → Normal 兩種跨型態情境下進行評估。本研究同時從語者識別(Identification)與語者驗證(Verification)兩個面向進行分析,分別以分類準確率與等錯誤率(Equal Error Rate, EER)作為評估指標。
實驗結果顯示,所提出之特徵轉換方法能縮小正常語音與耳語語音之間的嵌入分布差異,提升跨型態語者辨識表現,並改善未見語者條件下的語者驗證能力。此外,本研究亦分析性別因素對嵌入分布與特徵轉換效果之影響,結果顯示不同性別資料在跨型態轉換與辨識表現上存在差異,說明資料分布特性對耳語語者辨識模型之學習具有重要影響。整體而言,本研究驗證了以自監督式語音嵌入結合特徵轉換策略處理跨語音型態語者辨識問題之可行性,並為後續耳語語音辨識與跨型態語音處理研究提供參考。
This study aims to investigate and improve the problem of cross-phonation speaker recognition between normal speech and whispered speech. Since whispered speech lacks vocal fold vibration, the fundamental frequency and harmonic structure are largely absent, resulting in a significant distribution gap between whispered and normal speech embeddings. This mismatch further affects the stability and accuracy of speaker recognition models under cross-phonation conditions. To address this issue, this study leverages high-level speech embeddings extracted from self-supervised learning (SSL) speech models and proposes a Transformer-based feature mapping model to learn the nonlinear mapping relationship between whispered and normal speech embeddings. By jointly optimizing Mean Squared Error (MSE) loss and Triplet Loss, the model can align the embedding distributions across phonation types while preserving speaker-discriminative characteristics. Experiments were conducted on the self-constructed Whisper-100 corpus and evaluated under two cross-phonation scenarios: Normal→Whisper and Whisper→Normal. This study analyzes the performance from both speaker identification and speaker verification perspectives, using classification accuracy and Equal Error Rate (EER) as the evaluation metrics, respectively. Experimental results show that the proposed feature transformation method can reduce the embedding distribution gap between normal and whispered speech, improve cross-phonation speaker recognition performance, and enhance speaker verification capability under unseen-speaker conditions. In addition, this study analyzes the influence of gender factors on embedding distributions and feature transformation effectiveness. The results show that gender-related differences exist in cross-phonation transformation and recognition performance, indicating that data distribution characteristics play an important role in learning whispered speech speaker recognition models. Overall, this study verifies the feasibility of using self-supervised speech embeddings combined with a feature transformation strategy to address cross-phonation speaker recognition, and provides a reference for future research on whispered speech recognition and crossphonation speech processing.
摘要 I
Abstract II
目錄 IV
表次 VII
圖次 VIII
第一章 緒論 1
1.1 研究背景與動機 1
1.2 研究目的 3
1.3 論文架構 4
第二章 技術背景與相關研究 6
2.1 語者辨識技術 6
2.1.1 語者辨識模型架構 7
2.2 自監督式學習(Self-SupervisedLearning,SSL)語音模型 8
2.2.1 Wav2Vec2.0 9
2.2.2 HuBERT 10
2.2.3 BEATs 11
2.2.4 OpenBEATs 12
2.3 耳語與正常語音的聲學差異 13
2.4 耳語語音研究現況與應用 14
2.4.1 耳語語音自動辨識 14
2.4.2 耳語語音轉換與合成 15
2.5 跨語音型態語者辨識研究 16
2.6 小結 18
第三章 研究方法 19
3.1 跨語音型態資料集介紹以及前處理 19
3.1.1 Whisper-100 19
3.1.2 資料集前處理 20
3.2 特徵擷取方法 20
3.3 特徵轉換模型設計 22
3.3.1 均方誤差損失(MeanSquaredError,MSELoss) 24
3.3.2 三元組損失(TripletLoss) 25
3.4 語者辨識模型 26
3.5 評估指標 28
3.5.1 等錯誤率(EqualErrorRate,EER) 28
3.6 性別分析實驗 29
3.7 延伸實驗:基於ASR與TTS之耳語正常語音重建 30
第四章 實驗結果 31
4.1 參數設定 31
4.2 語者嵌入特徵辨識結果 31
4.2.1 Whisper-100資料集結果 32
4.3 跨型態語音特徵轉換結果 35
4.3.1 Normal→Whisper特徵轉換 36
4.3.2 Whisper→Normal特徵轉換 38
4.3.3 損失函數消融與權重比例分析 43
4.3.4 特徵轉換基準方法比較 48
4.3.5 小結 52
4.4 性別分析實驗結果 52
4.4.1 男女語者聲學層級差異 53
4.4.2 不同性別模型之特徵轉換結果比較 54
4.4.3 小結 58
第五章 結論與未來工作 60
參考文獻 62
附錄A 基於ASR與TTS之耳語正常語音重建實驗 65
A.1 實驗目的 65
A.2 系統流程 65
A.3 評估指標 66
A.3.1 字符錯誤率(CharacterErrorRate,CER) 66
A.3.2 語者相似度與等錯誤率 67
A.4 實驗結果 67
A.4.1 內容正確性分析 67
A.4.2 語者特徵保留分析 68
A.5 小結 71
[1] Preeti Wadhwani. Conversational system market size: By component, by technology, by deployment mode, by application, growth forecast, 2025–2034. Global Market Insights Inc. (Report ID: GMI13898), May 2025.
[2] T. Itoh, K. Takeda, and F. Itakura. Acoustic analysis and recognition of whispered speech. In IEEE Workshop on Automatic Speech Recognition and Understanding, 2001. ASRU ’01., pages 429–432, 2001.
[3] Zhaofeng Lin, Tanvina Patel, and Odette Scharenborg. Improving whispered speech recognition performance using pseudo-whispered based data augmentation. In 2023 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU), pages 1–8. IEEE, 2023.
[4] Xing Fan and John HL Hansen. Speaker identification within whispered speech audio streams. IEEE transactions on audio, speech, and language processing, 19(5):1408–1421, 2010.
[5] Lu Yi and Man Wai Mak. Cross-domain adaptation in distance space for speaker verification. In 2023 Asia Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC), pages 2238–2243, 2023.
[6] Abdelrahman Mohamed, Hung-yi Lee, Lasse Borgholt, Jakob D. Havtorn, Joakim Edin, Christian Igel, Katrin Kirchhoff, Shang-Wen Li, Karen Livescu, Lars Maaløe, Tara N. Sainath, and Shinji Watanabe. Self-supervised speech representation learning: A review. IEEE Journal of Selected Topics in Signal Processing, 16(6):1179–1210, 2022.
[7] Alexei Baevski, Henry Zhou, Abdelrahman Mohamed, and Michael Auli. wav2vec 2.0: a framework for self-supervised learning of speech representations. In Proceedings of the 34th International Conference on Neural Information Processing Systems, NIPS ’20, Red Hook, NY, USA, 2020. Curran Associates Inc.
[8] Wei-Ning Hsu, Benjamin Bolte, Yao-Hung Hubert Tsai, Kushal Lakhotia, Ruslan Salakhutdinov, and Abdelrahman Mohamed. Hubert: Self-supervised speech representation learning by masked prediction of hidden units. IEEE/ACM Trans. Audio, Speech and Lang. Proc., 29:3451–3460, October 2021.
[9] Sanyuan Chen, Yu Wu, Chengyi Wang, Shujie Liu, Daniel Tompkins, Zhuo Chen, Wanxiang Che, Xiangzhan Yu, and Furu Wei. Beats: audio pre-training with acoustic tokenizers. In Proceedings of the 40th International Conference on Machine Learning, ICML’23. JMLR.org, 2023.
[10] Shikhar Bharadwaj, Samuele Cornell, Kwanghee Choi, Satoru Fukayama, Hye jin Shim, Soham Deshmukh, and Shinji Watanabe. Openbeats: A fully open-source general-purpose audio encoder. 2025 IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA), pages 1–5, 2025.
[11] Najim Dehak, Patrick J. Kenny, Réda Dehak, Pierre Dumouchel, and Pierre Ouellet. Front-endfactor analysisforspeakerverification. IEEETransactionsonAudio, Speech, and Language Processing, 19(4):788–798, 2011.
[12] Li Wan, Quan Wang, Alan Papir, and Ignacio Lopez Moreno. Generalized end to-end loss for speaker verification. In 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 4879–4883, 2018.
[13] David Snyder, Daniel Garcia-Romero, Gregory Sell, Daniel Povey, and Sanjeev Khudanpur. X-vectors: Robust dnn embeddings for speaker recognition. In 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 5329–5333, 2018.
[14] Ville Vestman, Dhananjaya Gowda, Md Sahidullah, Paavo Alku, and Tomi Kinnunen. Speaker recognition from whispered speech: A tutorial survey and an application of time-varying linear prediction. Speech Communication, 99:62–79, 2018.
[15] Aref Farhadipour, Homa Asadi, and Volker Dellwo. Leveraging self-supervised models for automatic whispered speech recognition. In 2024 14th International Conference on Computer and Knowledge Engineering (ICCKE), pages 188–193, 2024.
[16] Matthias Janke, Michael Wand, Till Heistermann, Tanja Schultz, and K Prahallad. Fundamental frequency generation for whisper-to-audible speech conversion. In 2014 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 2579–2583. IEEE, 2014.
[17] Jun Rekimoto. Wesper: Zero-shot and realtime whisper to normal voice conversion for whisper-based speech interactions. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems, CHI ’23, New York, NY, USA, 2023. Association for Computing Machinery.
[18] Chi Zhang and John HL Hansen. Analysis and classification of speech mode: whispered through shouted. In Interspeech, volume 7, pages 2289–2292, 2007.
[19] Xing Fan and John HL Hansen. Acoustic analysis and feature transformation from neutral to whisper for speaker identification within whispered speech audio streams. Speech communication, 55(1):119–134, 2013.
[20] Jakaria Islam Emon, Md Abu Salek, and Kazi Tamanna Alam. Whisper speaker identification: Leveraging pre-trained multilingual transformers for robust speaker embeddings. arXiv preprint arXiv:2503.10446, 2025.
[21] Dominik Wagner, Ilja Baumann, and Tobias Bocklet. Generative adversarial networks for whispered to voiced speech conversion: a comparative study. International Journal of Speech Technology, 27(4):1093–1110, 2024.
[22] Songting Liu. Zero-shot voice conversion with diffusion transformers. arXiv preprint arXiv:2411.09943, 2024.
[23] Santi Prieto, Alfonso Ortega Giménez, Iván López-Espejo, and EDUARDO LLEIDA SOLANO. Shouted speech compensation for speaker verification robust to vocal effort conditions. In Interspeech, 2020.
[24] Florian Schroff, Dmitry Kalenichenko, and James Philbin. Facenet: A unified embedding for face recognition and clustering. 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Jun 2015.