跳到主要內容

簡易檢索 / 詳目顯示

研究生: 張睿軒
Chang, Jui-Hsuan
論文名稱: 大型語言模型水印偏好詞的特徵分析
Characterization of Watermark-favored Words: A Diagnostic Framework for Analyzing Token-Level Behavior in Decoding-Time LLM Watermarking
指導教授: 郁方
Yu, Fang
口試委員: 洪智鐸
Hong, Chih Duo
江介宏
Jiang, Roland
學位類別: 碩士
Master
系所名稱: 商學院 - 資訊管理學系
Department of Management Information System
論文出版年: 2026
畢業學年度: 115
語文別: 英文
論文頁數: 41
中文關鍵詞: 大型語言模型浮水印解碼階段浮水印偏好詞領域轉移語意分歧WordNet
外文關鍵詞: LLM watermarking, decoding-time watermarking, favored tokens, domain shift, semantic divergence, WordNet
相關次數: 點閱:26下載:0
分享至:
查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報
  • 大型語言模型(Large Language Models, LLMs)的解碼階段浮水印(decoding-time watermarking)技術已成為內容來源歸屬(content attribution)的一種實用機制,然而,其在不同領域(domain)下的行為特性,尤其是詞元(token)層級的表現,仍缺乏充分探討。本研究聚焦於偏好詞元(favored tokens),即其選取機率會因浮水印演算法而提高的詞元,分析其分布如何隨不同領域及浮水印方法而變化。為了進行詞彙與語意層面的比較,我們將偏好詞元映射至 WordNet 同義詞集(synset)表示,並利用 Relaxed Word Mover's Distance(RWMD)量測跨領域語意差異。本研究針對五個不同領域及五種具代表性的浮水印方法進行跨領域分析,重點探討偏好詞元集中程度、跨領域語意差異,以及考量偏好詞元子集合的偵測效能。實驗結果顯示,浮水印訊號的分布具有高度不均勻性,少數高頻偏好詞元即貢獻了大部分可觀測到的浮水印訊號。其中,此集中現象在 Unigram 方法中最為明顯,而在 EXP 方法中最不顯著,顯示偏好詞元的集中程度主要受到浮水印演算法設計的影響。此外,跨領域語意差異亦高度依賴於浮水印方法。Unigram 因採用全域固定的 greenlist 設計,在不同領域間呈現最低的語意差異;相較之下,KGW 與 SynthID 等具情境敏感性的浮水印方法則展現較高的領域專屬性。然而,子集合感知(subset-aware)偵測實驗顯示,偏好詞元的高度集中並不必然代表當偵測僅依賴高頻偏好詞元時能獲得穩健的偵測效能。此外,在 Avoid/Favor 診斷設定下,偽陽性率顯著提升,顯示高頻偏好詞元雖可能影響偵測器分數,但不足以單獨作為浮水印存在的可靠證據。綜合而言,本研究發現,不同解碼階段浮水印方法對領域轉移(domain shift)的反應受到其演算法設計影響,而詞元層級的集中程度、語意穩定性及偵測穩健性分別反映了浮水印行為的不同面向。


    Decoding-time watermarking for large language models (LLMs) has emerged as a practical mechanism for content attribution, yet its behavior under domain variation remains poorly understood, particularly at the token level. We investigate favored tokens—token-level units whose selection probabilities are increased by watermarking algorithms—and analyze how their distributions vary across domains and watermarking methods. To enable lexical and semantic comparison, we map favored tokens to WordNet synset representations and measure cross-domain divergence using Relaxed Word Mover's Distance (RWMD). We conduct a cross-domain study across five domains and five representative watermarking methods, focusing on favored-token concentration, cross-domain semantic divergence, and subset-aware detection performance. Results show that watermark evidence is highly unevenly distributed: a small subset of high-frequency favored tokens accounts for a substantial portion of the observed watermark signal. This concentration is strongest for Unigram and weakest for EXP, indicating that favored-token concentration is primarily shaped by the watermarking algorithm. Cross-domain divergence is also strongly method-dependent. Unigram exhibits the lowest semantic divergence across domains, consistent with its globally fixed greenlist design, whereas context-sensitive methods such as KGW and SynthID show greater domain-specific specialization. However, subset-aware detection experiments show that high favored-token concentration does not necessarily imply robust detection when scoring is restricted to frequent favored tokens. In addition, an Avoid/Favor diagnostic setting substantially increases false positives, indicating that frequent favored tokens can influence detector scores but are not reliable evidence of watermark presence on their own. These findings show that decoding-time watermarking methods respond differently to domain shift depending on their algorithmic design, and that token-level concentration, semantic stability, and detection robustness capture distinct aspects of watermark behavior.

    致謝 i
    摘要 iii
    Abstract v
    Contents vii
    List of Figures ix
    List of Tables xi
    Introduction 1
    Related Work 5
    2.1 Responsible AI, provenance, and LLM watermarking 5
    2.2 Robustness, attacks, and detection 6
    2.3 Interpretability, diagnostics, and domain shift 7
    2.4 Research gap 7
    Methodology 9
    3.1 Favored-Token Measurability Across Injection Paradigms 9
    3.2 Favored Token 9
    3.3 Token Frequency Regimes 10
    Quantifying Semantic Divergence 11
    4.1 Rationale for Cross-Domain Metric Transfer 11
    4.2 Favored-word Synset Profiles 11
    4.3 Ground Cost Between Favored-token Entries 13
    4.4 Profile-level Divergence via Relaxed Word Mover’s Distance 14
    Experimental Setup 15
    5.1 Model and Watermarking Configuration 15
    5.2 Cross-domain Dataset Construction 16
    5.3 Evaluation and Statistical Analysis 16
    Results and Discussion 19
    6.1 RQ1: Distributional Unevenness of Favored Tokens 19
    6.2 RQ2: Cross-domain Semantic Divergence of Favored-token Profiles 20
    6.3 RQ3: Detection Signal Concentration in Frequently Favored Tokens 23
    6.4 Security and Deployment Implications 24
    6.5 Limitations 26
    Conclusion 27
    Bibliography 29
    A Appendix 35

    [BDP07] J. Blitzer, M. Dredze, and F. C. Pereira, “Biographies, bollywood, boom-boxes and blenders: Domain adaptation for sentiment classification,” in Annual Meeting of the Association for Computational Linguistics, 2007 (cit. p. 7).
    [BLL⁺10] S. Ben-David, T. Lu, T. Luu, and D. Pál, “Impossibility theorems for domain adaptation,” in International Conference on Artificial Intelligence and Statistics, 2010 (cit. p. 7).
    [Can25] F. Cano, “Towards responsible ai: Advances in safety, fairness, and accountability of autonomous systems,” arXiv preprint arXiv:2506.10192, vol. abs/2506.10192, 2025 (cit. p. 5).
    [CG24] M. Christ and S. Gunn, “Pseudorandom error-correcting codes,” IACR Cryptol. ePrint Arch., vol. 2024, p. 235, 2024 (cit. pp. 2, 6).
    [CGL⁺25] Y. Cheng, H. Guo, Y. Li, and L. Sigal, “Revealing weaknesses in text watermarking through self-information rewrite attacks,” arXiv preprint arXiv:2505.05190, 2025 (cit. p. 6).
    [CGZ23] M. Christ, S. Gunn, and O. Zamir, “Undetectable watermarks for language models,” IACR Cryptol. ePrint Arch., vol. 2023, p. 763, 2023 (cit. p. 5).
    [CHS25] H. Chang, H. Hassani, and R. Shokri, “Watermark smoothing attacks against language models,” in Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, 2025 (cit. p. 6).
    [CKH⁺24] Y. Chang, K. Krishna, A. Houmansadr, J. Wieting, and M. Iyyer, “Postmark: A robust blackbox watermark for large language models,” ArXiv, vol. abs/2406.14517, 2024 (cit. p. 5).
    [CKL⁺19] K. Clark, U. Khandelwal, O. Levy, and C. D. Manning, “What does bert look at? an analysis of bert’s attention,” in BlackboxNLP@ACL, 2019 (cit. p. 7).
    [CWG⁺24] R. Chen, Y. Wu, J. Guo, and H. Huang, “De-mark: Watermark removal in large language models,” arXiv preprint arXiv:2410.13808, 2024 (cit. p. 6).
    [FCT⁺23] P. Fernandez, A. Chaffin, K. Tit, V. Chappelier, and T. Furon, “Three bricks to consolidate watermarks for large language models,” 2023 IEEE International Workshop on Information Forensics and Security (WIFS), pp. 1–6, 2023 (cit. p. 6).
    [FZY⁺24] J. Fu, X. Zhao, R. Yang, et al., “Gumbelsoft: Diversified language model watermarking via the gumbelmax-trick,” in Annual Meeting of the Association for Computational Linguistics, 2024 (cit. p. 5).
    [GCW⁺22] M. Geva, A. Caciularu, K. Wang, and Y. Goldberg, “Transformer feed-forward layers build predictions by promoting concepts in the vocabulary space,” ArXiv, vol. abs/2203.14680, 2022 (cit. p. 7).
    [GM24] N. Golowich and A. Moitra, “Edit distance robust watermarks for language models,” ArXiv, vol. abs/2406.02633, 2024 (cit. pp. 2, 6).
    [GMS⁺20] S. Gururangan, A. Marasović, S. Swayamdipta, et al., “Don’t stop pretraining: Adapt language models to domains and tasks,” ArXiv, vol. abs/2004.10964, 2020 (cit. p. 7).
    [GTB24] S. Goellner, M. Tropmann-Frick, and B. Brumen, “Responsible artificial intelligence: A structured literature review,” arXiv preprint arXiv:2403.06910, vol. abs/2403.06910, 2024 (cit. p. 5).
    [Gu24] J. Gu, “Responsible generative ai: What to generate and what not,” ArXiv, vol. abs/2404.05783, 2024 (cit. p. 5).
    [HAR19] K. Huang, J. Altosaar, and R. Ranganath, “Clinicalbert: Modeling clinical notes and predicting hospital readmission,” ArXiv, vol. abs/1904.05342, 2019 (cit. p. 1).
    [HBK⁺21] D. Hendrycks, C. Burns, S. Kadavath, et al., “Measuring mathematical problem solving with the math dataset,” ArXiv, vol. abs/2103.03874, 2021 (cit. p. 7).
    [HCY26] C.-D. Hong, Y.-P. Chen, and F. Yu, “Signature filtering: A lightweight enhancement for statistical watermark detection in large language models,” arXiv preprint arXiv:2606.18430, 2026 (cit. p. 1).
    [HPW25] B. Huang, X. Pu, and X. Wan, “B⁴: A black-box scrubbing attack on LLM watermarks,” in Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies, 2025, pp. 9113–9126 (cit. p. 6).
    [HSL⁺24] M. Huo, S. A. Somayajula, Y. Liang, et al., “Token-specific watermarking with enhanced detectability and semantic coherence for large language models,” ArXiv, vol. abs/2402.18059, 2024 (cit. p. 5).
    [HZZ⁺25] H. Huang, Y. Zhang, H. Zheng, et al., “Rlcracker: Exposing the vulnerability of llm watermarks with adaptive rl attacks,” arXiv preprint arXiv:2509.20924, 2025 (cit. p. 6).
    [JSV24] N. Jovanović, R. Staab, and M. Vechev, “Watermark stealing in large language models,” in International Conference on Machine Learning, 2024, pp. 22 570–22 593 (cit. p. 6).
    [KGW⁺23] J. Kirchenbauer, J. Geiping, Y. Wen, et al., “A watermark for large language models,” in International Conference on Machine Learning, 2023 (cit. pp. 1, 5).
    [LGW⁺23] J. Lai, W. Gan, J. Wu, Z. Qi, and P. S. Yu, “Large language models in law: A survey,” ArXiv, vol. abs/2312.03718, 2023 (cit. p. 1).
    [LHA⁺23] T. Lee, S. Hong, J. Ahn, et al., “Who wrote this code? watermarking for code generation,” in Annual Meeting of the Association for Computational Linguistics, 2023 (cit. pp. 1, 5).
    [LPL⁺23] A. Liu, L. Pan, Y. Lu, et al., “A survey of text watermarking in the era of large language models,” ArXiv, vol. abs/2312.07913, 2023 (cit. p. 6).
    [LWH⁺25] J. Liang, Z. Wang, S. Hong, S. Ji, and T. Wang, “Watermark under fire: A robustness evaluation of LLM watermarking,” in Findings of the Association for Computational Linguistics: EMNLP 2025, 2025, pp. 21 050–21 074 (cit. p. 6).
    [MPW⁺25] C. Meister, T. Pimentel, G. Wiher, and R. Cotterell, Locally typical sampling, 2025. arXiv: 2202.00666 [cs.CL] (cit. p. 7).
    [NDG⁺25] N. Nashid, D. Ding, K. Gallaba, A. E. Hassan, and A. Mesbah, “Characterizing multi-hunk patches: Divergence, proximity, and llm repair challenges,” ArXiv, vol. abs/2506.04418, 2025 (cit. p. 11).
    [NF24] G. Nicholas and P. Friedl, “Regulating large language models: A roundtable report,” ArXiv, vol. abs/2403.15397, 2024 (cit. p. 1).
    [nos20] nostalgebraist, Interpreting gpt: The logit lens, Blog post, LessWrong, Jul. 2020 (cit. p. 7).
    [PHZ⁺24] Q. Pang, S. Hu, W. Zheng, and V. Smith, “Attacking LLM watermarks by exploiting their strengths,” in ICLR 2024 Workshop on Secure and Trustworthy Large Language Models, 2024 (cit. p. 6).
    [PRB23] J. Parker, V. Richard, and K. Becker, “Guidelines for the integration of large language models in developing and refining interview protocols,” The Qualitative Report, 2023 (cit. p. 1).
    [RCM⁺25] A. Reuel, P. Connolly, K. Meimandi, et al., “Responsible ai in the global context: Maturity model and survey,” in Proceedings of the ACM Conference on Fairness, Accountability, and Transparency, 2025 (cit. p. 5).
    [Rit24] R. Rithika, “Recent advances in large language models: An upshot,” International Journal of Research Publication and Reviews, 2024 (cit. p. 1).
    [TDP19] I. Tenney, D. Das, and E. Pavlick, “Bert rediscovers the classical nlp pipeline,” in Annual Meeting of the Association for Computational Linguistics, 2019 (cit. p. 7).
    [ZAL⁺24] X. Zhao, P. V. Ananth, L. Li, and Y.-X. Wang, “Provable robust watermarking for AI-generated text,” in The Twelfth International Conference on Learning Representations, 2024 (cit. pp. 1, 5).
    [ZJZ⁺23] J. Zhang, X. Ji, Z. Zhao, X. S. Hei, and K.-K. R. Choo, “Ethical considerations and policy implications for large language models: Guiding responsible development and deployment,” ArXiv, vol. abs/2308.02678, 2023 (cit. p. 1).
    [ZLW⁺24] X. Zhao, C. Liao, Y.-X. Wang, and L. Li, “Efficiently identifying watermarked segments in mixed-source texts,” ArXiv, vol. abs/2410.03600, 2024 (cit. p. 6).
    [ZZZ⁺24] Z. Zhang, X. Zhang, Y. Zhang, et al., “Large language model watermark stealing with mixed integer programming,” arXiv preprint arXiv:2405.19677, 2024 (cit. p. 6).

    無法下載圖示 全文公開日期 2031/07/21
    QR CODE
    :::