| 研究生: |
陸人瑋 Lu, Ren Wei |
|---|---|
| 論文名稱: |
針對傳統中文在地化大型語言模型的雙軌紅隊測試與評估 Dual-Track Red Teaming and Evaluation for Traditional Chinese Localized LLMs |
| 指導教授: |
郁方
Yu, Fang |
| 口試委員: |
郁方
Yu, Fang 洪智鐸 Hong, Chih-Duo 陳琬萍 Chen, Wan-Ping |
| 學位類別: |
碩士
Master |
| 系所名稱: |
資訊學院 - 資訊安全碩士學位學程 Master Program in Information Security |
| 論文出版年: | 2026 |
| 畢業學年度: | 114 |
| 語文別: | 英文 |
| 論文頁數: | 59 |
| 中文關鍵詞: | 繁體中文 、大型語言模型 、安全對齊 、紅隊測試 、越獄評估 、提示語模糊測試 、可解釋性 |
| 外文關鍵詞: | Traditional Chinese, Large language models, Safety alignment, Red teaming, Jailbreak evaluation, Prompt Fuzzing, Interpretability |
| 相關次數: | 點閱:24 下載:0 |
| 分享至: |
| 查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報 |
本研究提出一套雙軌紅隊測試框架,旨在評估繁體中文在地化大型語言模型在安全對齊(Safety Alignment)上的失效模式。儘管多數大型語言模型已歷經安全對齊訓練,然而在繁體中文語境、在地化用語與特定區域使用情境下,模型仍可能出現拒答機制不穩定或對有害請求提供實質協助之盲點。本研究設計Track A與Track B兩種互補的測試設定:Track A(靜態多輪測試)採用固定的多輪對話,並以良性情境包裝惡意意圖,用以檢驗模型在持續互動中是否會逐步放寬安全邊界;Track B(動態單輪測試)則將相同情境轉換為單輪種子提示詞,並結合蒙地卡羅樹搜尋(MCTS)導引的Prompt Fuzz變體,系統性探索因提示語改寫(Prompt Reformulation)所誘發的安全風險。本研究最終構建之可報告資料集包含10個目標模型、共計36,637筆模型回覆,均歷經回覆品質過濾與第二階段LLM-as-a-judge評分機制。實驗結果顯示,不同模型家族的攻擊成功率(ASR)存在顯著差異:GPT-OSS系列整體的ASR維持在9%以下,展現較高韌性;而Gemma-3、Qwen3、DeepSeek-Qwen與TAIDE等模型在多項設定下則呈現較高的安全風險。整體而言,Track A的ASR為60.15%,Track B則為50.12%,證實了固定多輪情境與自動化提示語改寫能揭露不同維度的安全弱點。此外,本研究進一步提出僅依賴回覆之DIVI-SHAP(Response-only DIVI-SHAP)分析方法,對模型輸出進行非監督式分群與詞彙歸因,藉此識別出高風險回覆的群聚現象及其代表性語意特徵,旨在為後續的回覆端防護與在地化安全評估提供具體之診斷依據。
This study introduces a dual-track red-teaming framework designed to evaluate safety alignment failures in Traditional Chinese(TC)localized Large Language Models(LLMs). While mainstream LLMs undergo extensive safety alignment, their localized TC variants frequently exhibit vulnerabilities when exposed to language-specific phrasings and regional usage con texts. To systematically assess these vulnerabilities, we establish two complementary evaluation tracks: Track A utilizes fixed, multi-turn, benign-masked scenarios to examine whether harmful intent can bypass safeguards across sustained interactions. Conversely, Track B trans forms these identical scenarios into single-turn seeds, employing a Monte Carlo Tree Search (MCTS)-guided Prompt Fuzz variant to evaluate model sensitivity to prompt reformulations without conversational carryover.
The final evaluation corpus comprises 36,637 reportable responses derived from 10 target models, all post-processed through a rigid response-quality filter and a two-stage LLM-as-a judge scoring paradigm. Experimental results reveal substantial divergence across model families: while the GPT-OSS series maintains an overall Attack Success Rate (ASR) below 9%, variants such as Gemma-3, Qwen3, DeepSeek-Qwen, and TAIDE exhibit pronounced vulnerabilities. Aggregate data indicates a 60.15% ASR for Track A compared to 50.12% for Track B, demonstrating that fixed multi-turn narratives and automated prompt mutations expose distinct failure modes. Furthermore, a response-only DIVI-SHAP diagnostic pipeline is implemented to isolate refusal behaviors from high-ASR clusters, providing empirical taxonomies for localized safety validation and response-side filtering.
誌謝 i
摘要 ii
Abstract iii
Contents iv
1 Introduction 1
1.1 Contributions 2
1.2 Thesis Organization 3
2 Related Work 4
2.1 Evolution of LLM Evaluation: From Static to Dynamic 4
2.2 Adversarial Attacks and Red Teaming 5
2.3 Cross-Lingual Safety and Interpretability 6
3 Methodology 8
3.1 Task Definition and Threat Model 8
3.2 Adversarial System Configurations 10
3.3 Dual-Track Evaluation Strategy 12
3.4 Post-hoc Diagnostic Pipeline: DIVI-SHAP 16
4 Evaluation Setup 19
4.1 Target Models and Adversarial Dataset Curation 19
4.2 Environment Settings & Evaluation Rubric 19
4.3 Ethical Considerations and Responsible Disclosure 21
5 Experimental Results 22
5.1 Overall Model Robustness 22
5.2 Supplemental English Baseline Comparison 23
5.3 Supplementary Online Model Probe Validation 24
5.4 Multi-turn, Scenario, and Mutation Effects 26
5.5 Response-only DIVI-SHAP Diagnosis 29
6 Conclusion 32
Bibliography 33
A Supplementary Materials 36
A.1 Supplemental English Baseline Notes 36
A.2 Supplementary Online Model Probe Notes 36
A.3 Model-wise Cluster Distribution Heatmaps 36
B Prompt Materials and Judge Rubrics 41
B.1 Full System Prompt Configurations 41
B.2 Track A Scenario Prompts 45
B.3 Track B Seed Prompts 50
B.4 Second-Judge Prompt and Rubric 54
[1] TAIDE Project Team, Taide: Trustworthy ai dialogue engine for taiwan, https://taide.tw/, Accessed: 2026-01-15, 2024 (cit. pp. 1, 2, 5, 12).
[2] A.Wei,N.Haghtalab, and J. Steinhardt, “Jailbroken: How does llm safety training fail?”Advances in Neural Information Processing Systems, vol. 36, pp. 80079–80110, 2023(cit. p. 1).
[3] L.Zhengetal.,“Judgingllm-as-a-judgewithmt-bench and chatbot arena,”arXivpreprint arXiv:2306.05685, 2023 (cit. p. 1).
[4] Z.-X.Yong,C.Menghini,andS.H.Lu,“Low-resourcelanguagesjailbreak gpt-4,” arXiv preprint arXiv:2310.02446, 2023 (cit. pp. 1, 6).
[5] M.Mazeikaetal.,“Harmbench:Astandardizedevaluation framework for automated red teaming and robust refusal,” arXiv preprint arXiv:2402.04249, 2024 (cit. pp. 1, 4).
[6] M. Research, C.-J. Hsu, C.-S. Liu, et al., “The breeze 2 herd of models: Traditional chinese llms based on llama with vision-aware and function-calling capabilities,” arXiv preprint arXiv:2501.13921, 2025 (cit. p. 2).
[7] GemmaTeamGoogleDeepMind,“Gemma3technicalreport,”arXivpreprintarXiv:2503.19786, 2025 (cit. p. 2).
[8] M.AI, Llama 3 model card, Accessed: 2024-05-20, 2024 (cit. p. 2).
[9] Q.Team, Qwen3 technical report, 2025. arXiv: 2505.09388 [cs.CL] (cit. p. 2).
[10] OpenAI, Gpt-oss-120b & gpt-oss-20b model card, 2025. arXiv: 2508.10925 [cs.CL] (cit. p. 2).
[11] DeepSeek-AI, Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning, 2025. arXiv: 2501.12948 [cs.CL] (cit. p. 2).
[12] Y. Liu et al., “G-eval: Nlg evaluation using gpt-4 with better human alignment,” arXiv preprint arXiv:2303.16634, 2023 (cit. p. 4).
[13] A.Souly et al., “A strongreject for empty jailbreaks,” arXiv preprint arXiv:2402.10260, 2024 (cit. p. 4).
[14] H. Xiao, X. Li, R. Wang, et al., “M2DE: A multi-stressor multi-dimensional dynamic evaluation framework for the trustworthiness of llms,” Pattern Recognition, vol. 178, p. 113428, 2026 (cit. pp. 4, 12).
[15] S. Liu, S. Cui, H. Bu, Y. Shang, and X. Zhang, Jailbench: A comprehensive chinese security assessment benchmark for large language models, 2025. arXiv: 2502.18935 [cs.CL] (cit. p. 4).
[16] P.-C. Hsu, M.-H. Chen, T. L. Chao, C. T. Han, and D.-s. Shiu, Taiwan safety benchmark and breeze guard: Toward trustworthy ai for taiwanese mandarin, 2026. arXiv: 2603. 07286 [cs.CL] (cit. p. 4).
[17] A. Zou, Z. Wang, N. Carlini, et al., “Universal and transferable adversarial attacks on aligned language models,” arXiv preprint arXiv:2307.15043, 2023 (cit. p. 5).
[18] Y. Shao, J. Yu, H. Miao, et al., “Promptfuzz: Harnessing fuzzing techniques for robust testing of prompt injection in llms,” IEEE Transactions on Information Forensics and Security, 2026, Accepted for publication (cit. pp. 5, 12).
[19] X. Liu, N. Xu, M. Chen, and C. Xiao, “Autodan: Generating stealthy jailbreak prompts on aligned large language models,” arXiv preprint arXiv:2310.04451, 2023 (cit. p. 5).
[20] P. Chao, A. Robey, E. Dobriban, et al., “Jailbreaking black box large language models in twenty queries,” arXiv preprint arXiv:2310.03684, 2023 (cit. p. 5).
[21] A.Mehrotra, M. Zampetakis, P. Cassano, et al., “Tree of attacks: Jailbreaking black-box llms automatically,” arXiv preprint arXiv:2312.02119, 2024 (cit. p. 5).
[22] X. Li et al., “Deepinception: Hypnotize large language model to be jailbreaker,” arXiv preprint arXiv:2311.03191, 2023 (cit. p. 5).
[23] Z. Weng, X. Jin, J. Jia, and X. Zhang, Foot-in-the-door: A multi-turn jailbreak for llms, 2025. arXiv: 2502.19820 [cs.CL] (cit. p. 5).
[24] A.KumarappanandA.Mujoo,Automatingdeception:Scalablemulti-turnllmjailbreaks, 2026. arXiv: 2511.19517 [cs.LG] (cit. p. 5).
[25] X. Yang, J. Lee, A.-K. Dick, et al., Multi-turn jailbreaks are simpler than they seem, 2025. arXiv: 2508.07646 [cs.LG] (cit. p. 5).
[26] P. Röttger, P. Schott, Y. Ruan, D. Hovy, and J. Lovelace, “The english-centric shadow: A systematic audit of protocols and data distributions in llm safety alignment,” arXiv preprint arXiv:2502.10432, 2025 (cit. p. 6).
[27] A. Zou, L. Phan, S. Chen, et al., “Representation engineering: A top-down approach to ai transparency,” arXiv preprint arXiv:2310.01405, 2023 (cit. p. 6).
[28] W.P.Chen,Adata-informedvariationalclusteringframeworkfornoisyhigh-dimensional data, 2026. arXiv: 2604.06864 [stat.ML] (cit. pp. 6, 16).
[29] F. Perez and I. Ribeiro, “Ignore previous prompt: Attack techniques for language models,” arXiv preprint arXiv:2211.09527, 2022 (cit. p. 10).
[30] X. Shen, Z. Chen, M. Backes, Y. Shen, and Y. Zhang, “” do anything now”: Characterizing and evaluating in-the-wild jailbreak prompts on large language models,” in Proceedings of the 2024 on ACM SIGSAC Conference on Computer and Communications Security, 2024, pp. 1671–1685 (cit. p. 10).
[31] S. M. Lundberg and S.-I. Lee, “A unified approach to interpreting model predictions,” Advances in neural information processing systems, vol. 30, 2017 (cit. p. 17).
[32] S.M.Lundberg,G.Erion,H.Chen,etal.,“Fromlocalexplanationstoglobalunderstanding with explainable ai for trees,” Nature Machine Intelligence, vol. 2, no. 1, pp. 56–67, 2020 (cit. p. 17).