| 研究生: |
周俊和 Chou, Chun-He |
|---|---|
| 論文名稱: |
以音訊特徵為基礎之可解釋音樂曲風分類方法 An Interpretable Music Genre Classification Method Based on Audio Features |
| 指導教授: |
余清祥
Yue, Ching-Syang |
| 口試委員: |
張志浩
Chang, Chih-Hao 黃郁芬 Huang, Yu-Fen 徐南蓉 Hsu, Nan-Jung |
| 學位類別: |
碩士
Master |
| 系所名稱: |
商學院 - 統計學系 Department of Statistics |
| 論文出版年: | 2026 |
| 畢業學年度: | 114 |
| 語文別: | 中文 |
| 論文頁數: | 80 |
| 中文關鍵詞: | 音樂曲風分類 、音訊特徵 、特徵選取 、統計機器學習 、跨資料集測試 |
| 外文關鍵詞: | Music Genre Classification, Audio Features, Feature Selection, Statistical Machine Learning, Cross-Dataset Testing |
| 相關次數: | 點閱:3 下載:0 |
| 分享至: |
| 查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報 |
隨著數位音樂產業與串流平台快速發展,歌曲推薦、音樂檢索與播放清單生成等應用日益普及,穩定一致的曲風標註需求更為殷切。然而,人工分類不僅耗時且成本較高,也容易受到標註者音樂經驗、文化背景與曲風認知差異影響;此外,單一標籤亦難以完整描述具有多重曲風特徵的歌曲。近年研究多採用深度學習方法(Choi et al., 2017),雖具有良好分類效能,卻也需要較高運算資源,且分類依據不易對應至具體聲學特性,限制模型的可解釋性。因此本文聚焦可解釋音訊特徵的統計方法,探討其曲風二元分類能力及跨資料集穩定性。
本文採用 GTZAN 與 FMA 兩個公開音樂資料集,比較「爵士與古典」、「搖滾與鄉村」兩組二元分類。音訊特徵涵蓋音色、頻譜、和聲、節奏與能量等面向,除平均值與變異數外,也納入穩健離群值比例(Robust Outlier Ratio, ROR)與波峰因數(Crest Factor, CF),以捕捉瞬時峰值、動態範圍與分布尾端資訊。研究透過探索性資料分析、2效應量與模型特徵重要度篩選關鍵特徵,再以統計與機器學習模型進行分類,並透過跨資料集測試評估泛化能力。
分析顯示,以上述特徵結合統計與機器學習模型在特定曲風的分類結果,可達到與卷積循環神經網路(Convolutional Recurrent Neural Network, CRNN)相近的準確率,所需解釋變數可用於實質詮釋、且數量較少。以 GTZAN「搖滾與鄉村」為例,使用全部特徵時準確率約85%,低於 CRNN 的87%;經篩選音色與頻譜等關鍵特徵後,準確率提升至93%。換言之,適當的特徵選擇可有效提升分類效能,不同曲風的聲學差異亦無法由固定特徵組合完整描述。此外,ROR 與 CF 可補充平均值與變異數難以呈現的瞬時峰值及動態範圍資訊,有助於提升分類穩定性。整體而言,本研究證實可解釋音訊特徵結合統計機器學習,兼具分類效能、模型簡約性與可解釋性,並具有跨資料集應用的潛力。
With the rapid growth of digital music and streaming platforms, applications such as music recommendation, retrieval, and playlist generation increasingly rely on stable and consistent genre annotations. However, manual genre classification is time-consuming and costly, and may be influenced by annotators’ musical experience, cultural background, and interpretation of genres. Moreover, a single label may not fully represent songs containing multiple genre characteristics. Although deep learning has been widely adopted for music genre classification (Choi et al., 2017), it generally requires substantial computational resources, and its decisions are often difficult to relate to specific acoustic properties, limiting interpretability. This study investigates the effectiveness of interpretable audio features combined with statistical methods for binary genre classification and evaluates their stability across datasets.
Experiments use the GTZAN and FMA datasets, focusing on two binary classification tasks: “Jazz vs. Classical” and “Rock vs. Country.” Audio features cover timbral, spectral, harmonic, rhythmic, and energy-related characteristics. In addition to mean and variance, the Robust Outlier Ratio (ROR) and Crest Factor (CF) are included to capture transient peaks, dynamic range, and distributional extremes. Key features are identified through exploratory data analysis, η² effect size analysis (Lakens, 2013), and model-specific feature importance. Statistical and machine learning models are then evaluated using cross-dataset testing to assess generalizability.
The results show that feature-based statistical and machine learning models can achieve accuracy comparable to a Convolutional Recurrent Neural Network (CRNN) for specific genre classification tasks while using substantially fewer and more interpretable variables. For the GTZAN “Rock vs. Country” task, accuracy is approximately 85% using all features, compared with 87% for the CRNN; after selecting key timbral and spectral features, accuracy increases to 93%. These findings indicate that appropriate feature selection can improve classification performance and that acoustic differences between genre pairs cannot be fully characterized by a fixed set of features. ROR and CF further provide information on transient peaks and dynamic characteristics beyond mean and variance, contributing to greater classification stability. Overall, interpretable audio features combined with statistical machine learning can achieve effective classification while maintaining model parsimony and interpretability, with potential for application across music datasets.
第一章 緒論 1
第一節 研究動機 1
第二節 研究目的 3
第二章 文獻探討 6
第一節 文獻回顧 6
第二節 音樂曲風資料集與資料依賴性 9
第三節 音訊特徵與曲風辨識方法 11
第三章 研究方法 14
第一節 曲風類別選擇與二分類任務設計 16
第二節 資料前處理 18
第三節 特徵提取 20
第四節 特徵統計量 26
第五節 特徵選取 31
第六節 分類模型與訓練流程 36
第四章 研究結果與分析 42
第一節 不同重疊設定對分類效果之影響 43
第二節 資料依賴性檢查與二分類探索性資料分析 45
第三節 η² 效果量分析 47
第四節 模型特徵選取策略與分類模型結果 58
第五節 離群值相關統計量之分類效果 63
第六節 跨資料集泛化能力分析 66
第七節 小結 69
第五章 結論與建議 72
第一節 研究發現 72
第二節 研究限制與建議 74
參考文獻 76
一、 中文文獻
[1] 林振瑋(2009)。「二級式音樂類型分類之研究」,國立臺北科技大學電機工程系研究所碩士論文。
[2] 洪暐桓(2012)。「基於感知訊號處理之強健型貝氏音樂資訊檢索與分析」,國立交通大學工學院聲音與音樂創意科技碩士學位學程碩士論文。
[3] 黃哲彥(2026)。「可視化音訊特徵用於深度學習之音樂曲風分類研究」,中原大學智慧運算與大數據碩士學位學程碩士論文。
二、 英文文獻
[1] Ba, T. C., Le, T. D. T. & Van, L. T. (2025). Music genre classification using deep neural networks and data augmentation. Entertainment Computing, 53, 100929. https://doi.org/10.1016/j.entcom.2025.100929
[2] Bello, J. P., Daudet, L., Abdallah, S., Duxbury, C., Davies, M. & Sandler, M. B. (2005). A tutorial on onset detection in music signals. IEEE Transactions on Speech and Audio Processing, 13(5), 1035-1047. https://doi.org/10.1109/TSA.2005.851998
[3] Bogdanov, D., Porter, A., Herrera, P. & Serra, X. (2016). Cross-collection evaluation for music classification tasks. In Proceedings of the 17th International Society for Music Information Retrieval Conference (ISMIR 2016).
[4] Breiman, L. (2001). Random forests. Machine Learning, 45(1), 5-32.-https://doi.org/10.1023/A:1010933404324
[5] Chen, T. & Guestrin, C. (2016). XGBoost: A scalable tree boosting system. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (pp. 785-794). https://doi.org/10.1145/2939672.2939785
[6] Choi, K., Fazekas, G., Sandler, M. & Cho, K. (2017). Convolutional recurrent neural networks for music classification. In 2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) (pp. 2392-2396). https://doi.org/10.1109/ICASSP.2017.7952585
[7] Cohen, J. (1988). Statistical power analysis for the behavioral sciences (2nd ed.). Lawrence Erlbaum Associates.
[8] Davis, S. B. & Mermelstein, P. (1980). Comparison of parametric representations for monosyllabic word recognition in continuously spoken sentences. IEEE Transactions on Acoustics, Speech, and Signal Processing, 28(4), 357-366. https://doi.org/10.1109/TASSP.1980.1163420
[9] Defferrard, M., Benzi, K., Vandergheynst, P. & Bresson, X. (2017). FMA: A dataset for music analysis. In Proceedings of the 18th International Society for Music Information Retrieval Conference (ISMIR 2017) (pp. 316-323).
[10] Fabbri, F. (1982). A theory of musical genres: Two applications. In D. Horn & P. Tagg (Eds.), Popular music perspectives (pp. 52-81). International Association for the Study of Popular Music.
[11] FitzGerald, D. (2010). Harmonic/percussive separation using median filtering. In Proceedings of the 13th International Conference on Digital Audio Effects (DAFx-10).
[12] Flexer, A. (2007). A closer look on artist filters for musical genre classification. In Proceedings of the 8th International Conference on Music Information Retrieval (ISMIR 2007) (pp. 341-344).
[13] Friedman, J. H. (2001). Greedy function approximation: A gradient boosting machine. The Annals of Statistics, 29(5), 1189-1232. https://doi.org/10.1214/aos/1013203451
[14] Grosche, P., Müller, M. & Kurth, F. (2010). Cyclic tempogram: A mid-level tempo representation for music signals. In 2010 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) (pp. 5522-5525). https://doi.org/10.1109/ICASSP.2010.5495219
[15] Grosche, P. & Müller, M. (2011). Tempogram toolbox: MATLAB implementations for tempo and pulse analysis of music recordings. In Late-Breaking and Demo Session of the International Society for Music Information Retrieval Conference (ISMIR 2011).
[16] Harte, C., Sandler, M. & Gasser, M. (2006). Detecting harmonic change in musical audio. In Proceedings of the 1st ACM Workshop on Audio and Music Computing Multimedia (pp. 21-26). https://doi.org/10.1145/1178723.1178727
[17] Iglewicz, B. & Hoaglin, D. C. (1993). How to detect and handle outliers. ASQC Quality Press.
[18] Jiang, D.-N., Lu, L., Zhang, H.-J., Tao, J.-H. & Cai, L.-H. (2002). Music type classification by spectral contrast feature. In Proceedings of the 2002 IEEE International Conference on Multimedia and Expo (ICME) (pp. 113-116). https://doi.org/10.1109/ICME.2002.1035731
[19] Kedem, B. (1986). Spectral analysis and discrimination by zero-crossings. Proceedings of the IEEE, 74(11), 1477-1493. https://doi.org/10.1109/PROC.1986.13624
[20] Kirchberger, M. & Russo, F. A. (2016). Dynamic range across music genres and the perception of dynamic compression in hearing-impaired listeners. Trends in Hearing, 20, 1-16. https://doi.org/10.1177/2331216516630549
[21] Lakens, D. (2013). Calculating and reporting effect sizes to facilitate cumulative science: A practical primer for t-tests and ANOVAs. Frontiers in Psychology, 4, 863. https://doi.org/10.3389/fpsyg.2013.00863
[22] Lefaivre, A. & Zhang, J. Z. (2018). Music genre classification: Genre-specific characterization and pairwise evaluation. In Proceedings of the Audio Mostly 2018 on Sound in Immersion and Emotion. https://doi.org/10.1145/3243274.3243310
[23] Logan, B. (2000). Mel frequency cepstral coefficients for music modeling. In Proceedings of the 1st International Symposium on Music Information Retrieval (ISMIR 2000).
[24] Lukashevich, H. M., Abeßer, J., Dittmar, C. & Großmann, H. (2009). From multi-labeling to multi-domain-labeling: A novel two-dimensional approach to music genre classification. In Proceedings of the 10th International Society for Music Information Retrieval Conference (ISMIR 2009).
[25] Müller, M., Kurth, F. & Clausen, M. (2005). Chroma-based statistical audio features for audio matching. In Proceedings of the IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA) (pp. 275-278).
[26] Müller, M. & Ewert, S. (2011). Chroma toolbox: MATLAB implementations for extracting variants of chroma-based audio features. In Proceedings of the 12th International Society for Music Information Retrieval Conference (ISMIR 2011) (pp. 215-220).
[27] Oppenheim, A. V. & Schafer, R. W. (2004). From frequency to quefrency: A history of the cepstrum. IEEE Signal Processing Magazine, 21(5), 95-106. https://doi.org/10.1109/MSP.2004.1328092
[28] Oramas, S., Barbieri, F., Nieto, O. & Serra, X. (2017). Multi-label music genre classification from audio, text, and images using deep features. In Proceedings of the 18th International Society for Music Information Retrieval Conference (ISMIR 2017). https://arxiv.org/abs/1707.04916
[29] Peeters, G. (2004). A large set of audio features for sound description (similarity and classification) in the CUIDADO project. IRCAM.
[30] Rabiner, L. R. & Schafer, R. W. (1978). Digital processing of speech signals. Prentice-Hall.
[31] Rabiner, L. R. & Schafer, R. W. (2011). Theory and applications of digital speech processing. Pearson.
[32] Sturm, B. L. (2013). The GTZAN dataset: Its contents, its faults, their effects on evaluation, and its future use. arXiv. https://arxiv.org/abs/1306.1461
[33] Sturm, B. L. (2014). The state of the art ten years after a state of the art: Future research in music information retrieval. Journal of New Music Research, 43(2), 147-172. https://doi.org/10.1080/09298215.2014.894533
[34] Tzanetakis, G. & Cook, P. (2002). Musical genre classification of audio signals. IEEE Transactions on Speech and Audio Processing, 10(5), 293-302. https://doi.org/10.1109/TSA.2002.800560
[35] Wiener, N. (1949). Extrapolation, interpolation, and smoothing of stationary time series: With engineering applications. MIT Press.
[36] Xu, W. (2024). Music genre classification using deep learning: A comparative analysis of CNNs and RNNs. Applied Mathematics and Nonlinear Sciences, 9, 1-16. https://doi.org/10.2478/amns-2024-3309
[37] Yeh, C.-K., Su, L. & Yang, Y.-H. (2013). Dual-layer bag-of-frames model for music genre classification. In 2013 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). https://doi.org/10.1109/ICASSP.2013.6637646
[38] Yu, Y., Luo, S., Liu, S., Qiao, H., Liu, Y. & Feng, L. (2020). Deep attention based music genre classification. Neurocomputing, 372, 84-91. https://doi.org/10.1016/j.neucom.2019.09.054
[39] Zhang, W. (2022). Music genre classification based on deep learning. Computational Intelligence and Neuroscience, 2022, 2376888. https://doi.org/10.1155/2022/2376888
全文公開日期 2031/08/21