| 研究生: |
趙于傑 Jhao, Yu-Jie |
|---|---|
| 論文名稱: |
基於混合式特徵工程與決策樹之詐騙簡訊偵測模型 A Detection Model for Smishing Based on Hybrid Feature Engineering and Decision Tree Algorithm |
| 指導教授: |
許志堅
廖峻鋒 |
| 口試委員: | 王聖銘 |
| 學位類別: |
碩士
Master |
| 系所名稱: |
傳播學院 - 數位內容碩士學位學程 Digital Content and Technologies |
| 論文出版年: | 2026 |
| 畢業學年度: | 114 |
| 語文別: | 中文 |
| 論文頁數: | 65 |
| 中文關鍵詞: | 簡訊詐騙偵測 、決策樹 、預訓練語言模型 、混合式特徵工程 、可解釋性人工智慧 、關聯規則 |
| 外文關鍵詞: | Smishing Detection, Decision Tree, Pre-trained Language Model, Hybrid Feature Engineering, Explainable AI, Association Rules |
| 相關次數: | 點閱:25 下載:0 |
| 分享至: |
| 查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報 |
隨著行動通訊之普及,簡訊詐騙已成為當前對社會大眾威脅最為顯著之詐騙型態之一。傳統過濾機制多仰賴靜態黑名單與關鍵字比對,惟其在面對持續演化之規避技術時往往力有未逮;另一方面,近年廣泛採用之深度學習模型雖具備優異之偵測正確性,然其「黑箱」特性使模型決策過程難以解釋,於高度要求透明度與可稽核性之資安防禦場域中,應用上面臨相當侷限。
有鑑於此,本研究提出一套「混合式特徵工程」之簡訊詐騙偵測模型,整合 RoBERTa 預訓練語言模型與 Improved-ID3 決策樹演算法,期能同時兼顧高偵測正確性與高決策透明度。本研究於資料前處理階段,運用 RoBERTa 進行語意特徵萃取,擴充關鍵字詞庫;並定義涵蓋語意意圖、規避與混淆手法、文本結構與統計等三大類特徵。訓練階段採用 Improved-ID3 演算法,引入「純度」與「支持度」雙指標進行動態剪枝,有效抑制傳統決策樹易生之過度擬合問題;並進一步將決策樹路徑轉化為具備量化配分之 IF-THEN 關聯規則,供執行階段進行彈性加權判定。
本研究以 Salman, Ikram, and Kaafar (2024)所發佈之大型英文簡訊資料集 super_sms_dataset(約 67,000 筆樣本)進行實驗驗證,模型於測試集上達成 98.08% 之準確率(Accuracy)、97.94% 之精確率(Precision)、97.14% 之召回率(Recall),以及 97.54% 之 F1-Score,分類正確性優於傳統機器學習方法及多數詞嵌入結合深度學習之比較模型。進一步印證本研究於詐騙偵測任務上之有效性。
With the widespread adoption of mobile communications, SMS scams have become one of the most significant forms of fraud threatening the general public. Traditional filtering mechanisms primarily rely on static blacklists and keyword matching; however, they often struggle to cope with continuously evolving evasion techniques. In contrast, although deep learning models widely adopted in recent years have demonstrated excellent detection accuracy, their black-box nature makes the decision-making process difficult to interpret. This limitation poses challenges to their application in cybersecurity defense scenarios where transparency and auditability are highly valued.
To address these issues, this study proposes an SMS scam detection model based on a hybrid feature engineering approach, integrating the RoBERTa pretrained language model with the Improved-ID3 decision tree algorithm to achieve both high detection accuracy and decision transparency. During the data preprocessing stage, RoBERTa is employed for semantic feature extraction to expand the keyword lexicon. The proposed approach further defines three major categories of features, including semantic intent, evasion and obfuscation techniques, and textual structure and statistical characteristics. During the training stage, the Improved-ID3 algorithm incorporates purity and support as dual criteria for dynamic pruning, thereby effectively mitigating the overfitting problem commonly encountered in conventional decision trees. Furthermore, the decision tree paths are transformed into quantifiable IF-THEN association rules, which enable flexible weighted decision-making during the execution phase.
To evaluate the proposed approach, experiments were conducted on the large-scale English SMS dataset super_sms_dataset, released by Salman, Ikram, and Kaafar (2024), containing approximately 67,000 messages. On the test set, the proposed model achieved an Accuracy of 98.08%, Precision of 97.94%, Recall of 97.14%, and F1-Score of 97.54%. The classification performance outperformed conventional machine learning methods and most comparative models combining word embeddings with deep learning. These results demonstrate the effectiveness of the proposed approach for SMS scam detection while maintaining a high degree of decision transparency.
致謝 2
Abstract 4
目次 5
圖次 5
表次 6
第一章 緒論 7
第一節 研究背景與動機 8
第二節 研究目的與特色 11
第三節 研究範圍與限制 12
第四節 論文架構 13
第二章 文獻回顧 14
第一節 資料探勘 (Data Mining) 之定義與內涵 14
第二節 簡訊詐騙現況之分析 15
第三節 詐騙簡訊現有過濾機制 21
第四節 決策樹演算法 23
第三章 研究方法 30
第一節 研究架構 30
第二節 第一階段:輸入資料前處理 31
第三節 第二階段:機器學習之訓練階段 35
第四節 第三階段:機器學習之執行階段 39
第四章 實驗設計與結果分析 40
第一節 實驗資料集 40
第二節 實驗分類正確性評估指標 41
第三節 實驗參數設定與最佳化 43
第四節 實驗設計與流程對照 50
第五節 詳細實驗結果與分析 52
第六節 與其他過濾方法之比較 54
第五章 結論與建議 57
第一節 研究結論與貢獻 57
第二節 學術與社會貢獻 58
第三節 研究限制與未來建議 58
參考文獻 61
中文文獻 61
英文文獻 61
中文文獻
內政部警政署 165 打詐儀錶板. (2025). 詐騙手法 [Fraud methods]. Retrieved December 29, 2025, from https://165dashboard.tw/fraud-method
Trend Micro. (2016, January 4). 什麼是魚叉式網路釣魚(Spear Phishing)?[What is spear phishing?]. 資安趨勢部落格. https://blog.trendmicro.com.tw
Trend Micro. (2025, January 15). 破解簡訊詐騙四大招,保障你的錢包安全![Four tips to beat SMS scams]. 資安趨勢部落格. https://blog.trendmicro.com.tw/?p=85049
Whoscall. (2025). Whoscall 2024 年度報告 [Whoscall 2024 annual report].
英文文獻
Adebowale, M. A., Lwin, K. T., & Hossain, M. A. (2023). Intelligent phishing detection scheme using deep learning algorithms. Journal of Enterprise Information Management, 36(3), 747-766.
Aggarwal, C. C. (2015). Data mining: The textbook. Springer.
Al-Kaabi, H., Darroudi, A. D., & Jasim, A. K. (2024). Survey of SMS spam detection techniques: A taxonomy. AlKadhim Journal for Computer Science, 2(4), 23-34.
Ali, M., Khalid, S., & Aslam, M. H. (2017). Pattern based comprehensive Urdu stemmer and short text classification. IEEE Access, 6, 7374-7389.
Almeida, T. A., & Hidalgo, J. M. G. (2011). SMS spam collection [Data set]. https://doi.org/10.24432/C5CC84
Atkins, B., & Huang, W. (2013). A study of social engineering in online frauds. Open Journal of Social Sciences, 1(3), 23-32.
Berry, M. J. A., & Linoff, G. S. (2009). Data mining techniques (3rd ed.). John Wiley & Sons.
Cabena, P., Hadjinian, P., Stadler, R., Verhees, J., & Zanasi, A. (1998). Discovering data mining: From concept to implementation. Prentice Hall.
Çekik, R. (2024). A new filter feature selection method for text classification. IEEE Access, 12, 139316–139335. https://doi.org/10.1109/ACCESS.2024.3468001
Coleman, M., & Liau, T. L. (1975). A computer readability formula designed for machine scoring. Journal of Applied Psychology, 60(2), 283-284.
Cross, C., & Holt, T. J. (2025). Examining phishing attempts on data breach victims. Social Science Computer Review. Advance online publication. https://doi.org/10.1177/08944393251399841
Devlin, J., Chang, M.-W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Vol. 1, pp. 4171-4186). Association for Computational Linguistics.
Doshi-Velez, F., & Kim, B. (2017). Towards a rigorous science of interpretable machine learning. arXiv. https://arxiv.org/abs/1702.08608
Fayyad, U., Piatetsky-Shapiro, G., & Smyth, P. (1996). From data mining to knowledge discovery in databases. AI Magazine, 17(3), 37-54.
Flesch, R. (1948). A new readability yardstick. Journal of Applied Psychology, 32(3), 221-233.
Global Anti-Scam Alliance, & Feedzai. (2025). The global state of scams 2025 report. https://gasa.org/knowledge-base/reports/global-state-of-scams-2025
Global Anti-Scam Alliance, & Whoscall. (2025). State of scams in Taiwan report - 2025. https://gasa.org/knowledge-base/reports/state-of-scams-in-taiwan-report-2025
Goenka, R., Chawla, M., & Tiwari, N. (2025). Enhanced phishing detection approach using a layered model: Domain squatting and URL obfuscation identification and lexical feature-based classification. IEEE Access, 13, 187285-187306.
Grupe, F. H., & Owrang, M. M. (1995). Data base mining discovering new knowledge and competitive advantage. Information Systems Management, 12(4), 26-31.
Han, J., Kamber, M., & Pei, J. (2012). Data mining: Concepts and techniques (3rd ed.). Morgan Kaufmann.
Hao, M., Wang, W., & Zhou, F. (2021). Joint representations of texts and labels with compositional loss for short text classification. Journal of Web Engineering, 20(3), 669-688.
Henke, M., Carvalho, A., Covões, T. F., & Plastino, A. (2021). Spam detection based on feature evolution to deal with concept drift. Journal of Universal Computer Science, 27(4), 364-386.
Hochreiter, S., & Schmidhuber, J. (1997). Long short-term memory. Neural Computation, 9(8), 1735-1780.
Holdsworth, J. (n.d.). What is data mining? Retrieved December 29, 2025, from https://www.ibm.com/think/topics/data-mining
Huq, A., & Pervin, M. T. (2020). Adversarial attacks and defense on texts: A survey. arXiv. https://arxiv.org/abs/2005.14108
International Organization for Standardization, & International Electrotechnical Commission. (2022). Information technology — Artificial intelligence — Concepts and terminology (ISO/IEC 22989:2022). https://www.iso.org/standard/74296.html
Kincaid, J. P., Fishburne, R. W., Jr., Rogers, R. L., & Chissom, B. S. (1975). Derivation of new readability formulas (Automated Readability Index, Fog Count, and Flesch Reading Ease formula) for Navy enlisted personnel (Research Branch Report 8-75). Naval Technical Training Command.
Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., ... Stoyanov, V. (2019). RoBERTa: A robustly optimized BERT pretraining approach. arXiv. https://arxiv.org/abs/1907.11692
Mambina, I. S., Waweru, M. N., Rimiru, R., & Kagiri, C. (2024). Uncovering SMS spam in Swahili text using deep learning approaches. IEEE Access, 12, 25164-25175.
Marchal, S., François, J., State, R., & Engel, T. (2014). PhishStorm: Detecting phishing with streaming analytics. IEEE Transactions on Network and Service Management, 11(4), 458-471.
McLaughlin, G. H. (1969). SMOG grading—A new readability formula. Journal of Reading, 12(8), 639-646.
Mehmood, M. K., Arshad, H., Alawida, M., & Mehmood, A. (2024). Enhancing smishing detection: A deep learning approach for improved accuracy and reduced false positives. IEEE Access, 12, 137176–137193. https://doi.org/10.1109/ACCESS.2024.3463871
Ohmann, C., Moustakis, V., Yang, Q., & Lang, K. (1996). Evaluation of automatic knowledge acquisition techniques in the diagnosis of acute abdominal pain. Artificial Intelligence in Medicine, 8(1), 23-36.
Pyle, D. (1999). Data preparation for data mining. Morgan Kaufmann.
Salloum, S., Khan, R., & Shaalan, K. (2022). A systematic literature review on phishing email detection using natural language processing techniques. IEEE Access, 10, 65703-65727.
Salman, M., Ikram, M., Basta, N., & Kaafar, M. A. (2024). SpaLLM-Guard: Pairing SMS spam detection using open-source and commercial LLMs. In Proceedings of the 30th Annual International Conference on Mobile Computing and Networking (MobiCom) (pp. 1-15). ACM.
Salman, M., Ikram, M., & Kaafar, M. A. (2024). Investigating evasive techniques in SMS spam filtering: A comparative analysis of machine learning models. IEEE Access, 12, 24306-24324.
Senter, R. J., & Smith, E. A. (1967). Automated readability index. AMRL-TR-66-220. Aerospace Medical Research Laboratories.
Seo, J. W., Kim, J., & Lee, J. (2024). On-device smishing classifier resistant to text evasion attack. IEEE Access, 12, 4762-4779.
Sheu, J. J., Su, Y. H., & Chu, K. T. (2009). Segmenting online game customers: The perspective of experiential marketing. Expert Systems with Applications, 36(4), 8487-8495.
Sonowal, G. (2020). Phishing email detection based on binary search feature selection. SN Computer Science, 1, 191.
Tang, S., Li, Y., Ni, X., & Xu, L. (2022). Clues in tweets: Twitter-guided discovery and analysis of SMS spam. In Proceedings of the 2022 ACM SIGSAC Conference on Computer and Communications Security (pp. 2122-2136). ACM.
The Business Research Company. (2025). Information technology global market report 2025. Retrieved from MarketResearch.com
Wang, W., Wang, R., Wang, L., Wang, Z., & Ye, A. (2019). Towards a robust deep neural network in texts: A survey. arXiv. https://arxiv.org/abs/1902.07285
West, J., & Bhattacharya, M. (2016). Intelligent financial fraud detection: A comprehensive review. Computers & Security, 57, 47-66.
Witten, I. H., Frank, E., & Hall, M. A. (2011). Data mining: Practical machine learning tools and techniques (3rd ed.). Morgan Kaufmann.
Gupta, M., Bakliwal, A., Agarwal, S., & Mehndiratta, P. (2018). A comparative study of spam SMS detection using machine learning classifiers. In Proceedings of the 2018 Eleventh International Conference on Contemporary Computing (IC3) (pp. 1–7). IEEE.
Baaqeel, H., & Zagrouba, R. (2020). Hybrid SMS spam filtering system using machine learning techniques. In Proceedings of the 2020 21st International Arab Conference on Information Technology (ACIT) (pp. 1–8). IEEE.
全文公開日期 2031/08/19