跳到主要內容

簡易檢索 / 詳目顯示

研究生: 郭家榛
Guo, Jia-Jhen
論文名稱: 基於單層探測方法的繁體中文語言模型更加預訓練研究
Single Layer Probing Approach to Further Pre-training for Traditional Chinese Language Model
指導教授: 蔡瑞煌
Tsaih, Rua-Huan
口試委員: 蔡瑞煌
Tsaih, Rua-Huan
張家銘
Chang, Jia-Ming
王永鐘
Wang, Yung-Chung
李志宏
Lee, Jie Haun
學位類別: 碩士
Master
系所名稱: 商學院 - 資訊管理學系
Department of Management Information System
論文出版年: 2026
畢業學年度: 115
語文別: 英文
論文頁數: 77
中文關鍵詞: 大型語言模型領域適應金融語言模型繁體中文模型參數效率訓練單層探測方法人工智慧持續預訓練
外文關鍵詞: Large Language Model, Domain Adaptation, Financial Language Modeling, Traditional Chinese Language Model, Single-Layer Probing, Parameter-Efficient Training, TransformerLayers, Further Pre-training
相關次數: 點閱:8下載:0
分享至:
查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報
  • 一般領域語料訓練的大型語言模型,通常缺乏特定應用領域所需的領域性。以金融問答情境為例,模型回答必須能反映專業術語、法規語言、機構文件寫作慣例與金融專業知識。然而,將大型語言模型適應至金融領域仍面臨挑戰:全參數 pre-training 需要大量運算資源與大規模資料,而 supervised fine-tuning 則仰賴昂貴且耗時的專家標註資料。為解決此問題,本研究提出 Single-Layer Probing (SLP),作為一種具參數效率的 further pre-training 方法,用於繁體中文金融語言建模。

    本研究以Gemma-3-TAIDE-12B 為基礎模型,並自台灣公開金融資源建置約11.16 億tokens 的繁體中文金融語料庫,資料來源包含金融法規、企業財務報告、金融新聞與專業參考資料。不同於更新完整模型,SLP 凍結大部分模型參數,僅更新單一Transformer decoder layer 中的 self-attention module。此設計將layerdepth視為可控制的介入變數,同時大幅降低可訓練參數量。

    實驗比較shallow、 middle 與deep 三種layer probing settings,並進一步分析不同train-token-to-parameter ratios 的影響。結果顯示,middle-layer probing 具有最穩定的領域適應表現,而在測試範圍內,增加 training-token scale 可降低validation loss。問卷式人工評估進一步顯示,在選擇金融領域適應模型時,不應僅依賴 validation loss,也應結合以使用者為中心的回答品質評估。

    整體而言,本研究證明layer placement 與training-token scale 皆是具參數效率之金融領域適應的重要因素。所提出的 SLP 架構提供了一種可控制且節省運算資源的方法,用以分析大型語言模型的 layer-wise specialization,並支援繁體中文金融語言模型的領域適應。


    Large language models trained on general-domain corpora often lack the domain specificity required for specialized applications. In financial question-answering scenarios, for example, model responses must reflect specialized terminology, regulatory language, institutional writing conventions, and financial domain knowledge. However, adapting large language models to the financial domain remains challenging: full-parameter pre-training requires substantial computational resources and large-scale data, while supervised fine-tuning relies on costly and time-consuming expert-labeled datasets. To address these challenges, this study proposes Single-Layer Probing (SLP), a parameter-efficient further pre-training framework for Traditional Chinese financial language modeling.

    Using Gemma-3-TAIDE-12B as the base model, this study constructs a Traditional Chinese financial corpus of approximately 1.116 billion tokens from Taiwanese public financial resources, including regulatory documents, corporate financial reports, financial news, and professional reference materials. Instead of updating the entire model, SLP freezes most parameters and updates only the self-attention module of one selected Transformer decoder layer. This design allows layer depth to be treated as a controllable intervention variable while substantially reducing the number of trainable parameters.

    The experiments compare shallow-, middle-, and deep-layer probing settings under a fixed trainable-parameter budget, and further examine different train-token-to-parameter ratios. The results show that the middle-layer probing achieves the most stable adaptation performance, and that increasing the training-token scale improves validation loss within the tested range. The questionnaire-based human evaluation further highlights that model selection for financial-domain adaptation should not rely solely on validation loss, but should also incorporate user-centered assessment of response quality.

    Overall, this study demonstrates that layer placement and training-token scale are important factors in parameter-efficient financial-domain adaptation. The proposed SLP framework provides a controlled and resource-efficient approach for analyzing layer-wise specialization and adapting Traditional Chinese financial language models.

    Chapter 1 Introduction 1
    1.1 Research Background 1
    1.2 Research Objectives and Questions 2
    1.3 Significance and Contributions 3
    1.4 Research Overview 4

    Chapter 2 Related Work 5
    2.1 Self-Supervised Learning and Transfer Learning in Large Language Models 5
    2.1.1 The Foundation Model Paradigm and Self-Supervision 5
    2.1.2 Transfer Learning in Large Language Models 5
    2.2 Layer-Wise Representation Learning in Transformer Models 7
    2.2.1 Linguistic Hierarchy: From BERTology to Decoder-Only Models 7
    2.2.2 Knowledge Localization and the “Knowledge Neuron” Hypothesis 7
    2.2.3 Layer Depth as a Controllable Variable in Adaptation 7
    2.3 Parameter Restriction and Scaling Considerations 8
    2.3.1 Alignment with Compute-Optimal Scaling Laws 8
    2.3.2 Selective Parameter Updating as Efficient Adaptation 8
    2.4 Localized and Sovereign Language Models in Financial NLP 9
    2.4.1 The Rise of Sovereign AI and Cultural Intelligence 9
    2.4.2 Challenges of Traditional Chinese in Global Foundation Models 9
    2.4.3 Domain-Specific Localization as a Strategic Asset 9
    2.5 Foundational Infrastructure and Baseline Objective: Gemma-3-TAIDE-12B 10
    2.5.1 Gemma-3: Architecture and Model Specifications 10
    2.5.2 TAIDE: Pre-trained Parameters and Localized Knowledge Sources 10
    2.5.3 Baseline Objective: Standard Causal Language Modeling Pre-training 11

    Chapter 3 Methodology 13
    3.1 Problem Formulation and Our Motivation 13
    3.2 Base Model and Operational Behavior 13
    3.2.1 Architectural Backbone 14
    3.2.2 Forward Computation 14
    3.2.3 Pre-trained Parameters and Domain Alignment 14
    3.2.4 Learning Distinction in Further Pre-training 14
    3.3 Further Pre-training Algorithm 16
    3.3.1 Parameter Set Definition 16
    3.3.2 Loss Function and Optimization Objective 17
    3.3.3 SLP Training Procedure 17

    Chapter 4 Experimental Setup 19
    4.1 Financial Corpus Construction and Preprocessing 19
    4.1.1 Data Acquisition and Corpus Construction 19
    4.1.2 Data Preprocessing Pipeline 20
    4.2 Hardware Infrastructure 22
    4.2.1 Hardware and Software Environment 22
    4.2.2 Training Configuration 23
    4.3 Design Hypotheses 23
    4.4 Experimental Design: Layer-Wise Further Pre-training 24
    4.4.1 Design Objective 24
    4.4.2 Category Groups 25
    4.5 Parameter Restriction and Train-Token Ratio 26
    4.6 Evaluation Metrics 29
    4.6.1 Quantitative Evaluation 29
    4.6.2 Questionnaire-Based Human Evaluation 30

    Chapter 5 Results and Discussion 31
    5.1 H1: Layer Placement Effects in Single-Layer Probing 31
    5.1.1 Shallow-Layer Hyperparameter Search 31
    5.1.2 Middle-Layer Hyperparameter Search 33
    5.1.3 Deep-Layer Hyperparameter Search 34
    5.1.4 Comparison of the Best Layer-Wise Configurations 35
    5.1.5 Discussion of Hypothesis 1 36
    5.2 H2: Effects of the Train-Token-to-Parameter Ratio 37
    5.2.1 Hyperparameter Search for the 1:18.7 Ratio 37
    5.2.2 The 1:20.0 Ratio Configuration 38
    5.2.3 Hyperparameter Search for the 1:21.3 Ratio 39
    5.2.4 Comparison Across Train-Token Ratios 40
    5.2.5 Discussion of Hypothesis 2 41
    5.3 Execution Time of the Final Training Runs 42
    5.3.1 Insight 42
    5.4 Questionnaire-Based Human Evaluation 43
    5.4.1 Evaluation Procedure and Participant Profile 43
    5.4.2 Overall Model Comparison 44
    5.4.3 Domain Background Comparison 46
    5.4.4 Age-Based Comparison 48
    5.4.5 Educational Background Comparison 50
    5.4.6 Qualitative Case Comparison 52
    5.4.7 Discussion of Human Evaluation Results 53

    Chapter 6 Conclusion 57
    6.1 Conclusion 57
    6.2 Limitations and Future Work 58

    REFERENCE 63

    Appendix A Questionnaire Interface and Evaluation Records 67
    A.1 Questionnaire Interface 68
    A.1.1 System Introduction and Consent Page 68
    A.1.2 Participant Background Questionnaire 69
    A.1.3 Financial Question Input Interface 69
    A.1.4 Anonymized Model-Response Comparison Interface 70
    A.1.5 Overall Best and Worst Response Selection Interface 70
    A.1.6 Response-Quality Evaluation Interface 71
    A.2 Detailed Participant Demographic Profile 72
    A.3 Raw Questionnaire Rating Scores 73
    A.3.1 Response-Level Selection and Rating Records 73

    Appendix B Questionnaire Case Question 75
    B.1 Question 75
    B.2 H1-best Response 75
    B.3 TAIDE-baseline Response 76
    B.4 Participant Feedback on H1-best 76
    B.5 Participant Feedback on TAIDE-baseline 77
    B.6 English Translation of Question and Participant Feedback 77

    Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., ... & McGrew, B. (2023). GPT-4 technical report. arXiv preprint arXiv:2303.08774.

    Almazrouei, E., Alobeidli, H., Alshamsi, A., Cappelli, A., Cojocaru, R., Alhammadi, M., ... & Penedo, G. (2023). The falcon series of language models: Towards open frontier models. Hugging Face repository.

    Bommasani, R., Hudson, D. A., Adeli, E., Altman, R., Arora, S., von Arx, S., ... & Liang, P. (2021). On the opportunities and risks of foundation models. arXiv preprint arXiv:2108.07258.

    Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., ... & Amodei, D. (2020). Language models are few-shot learners. Advances in Neural Information Processing Systems, 33, 1877–1901.

    Devlin, J., Chang, M. W., Lee, K., & Toutanova, K. (2019, June). BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers) (pp. 4171–4186).

    Geva, M., Schuster, R., Berant, J., & Levy, O. (2021, November). Transformer feed-forward layers are key-value memories. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing (pp. 5484–5495).

    Gemma Team, Mesnard, T., Hardin, C., Dadashi, R., Bhupatiraju, S., Pathak, S., ... & Kenealy, K. (2024). Gemma: Open models based on Gemini research and technology. arXiv preprint arXiv:2403.08295.

    Gururangan, S., Marasović, A., Swayamdipta, S., Lo, K., Beltagy, I., Downey, D., & Smith, N. A. (2020, July). Don't stop pretraining: Adapt language models to domains and tasks. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (pp. 8342–8360).

    Hoffmann, J., Borgeaud, S., Mensch, A., Buchatskaya, E., Cai, T., Rutherford, E., ... & Sifre, L. (2022). An empirical analysis of compute-optimal large language model training. Advances in Neural Information Processing Systems, 35, 30016–30030.

    Jiang, A. Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D. S., de Las Casas, D., Bressand, F., Lengyel, G., Lample, G., Saulnier, L., Renard Lavaud, L., Lachaux, M.-A., Stock, P., Le Scao, T., Lavril, T., Wang, T., Lacroix, T., & El Sayed, W. (2023). Mistral 7B. arXiv preprint arXiv:2310.06825, 3.

    Jin, M., Yu, Q., Huang, J., Zeng, Q., Wang, Z., Hua, W., ... & Zhang, Y. (2025, January). Exploring concept depth: How large language models acquire knowledge and concept at different layers? In Proceedings of the 31st International Conference on Computational Linguistics (pp. 558–573).

    Liu, P., Yuan, W., Fu, J., Jiang, Z., Hayashi, H., & Neubig, G. (2023). Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing. ACM Computing Surveys, 55(9), 1–35.

    Lyu, H., Luo, J., Kang, J., & Koenecke, A. (2025, June). Characterizing bias: Benchmarking large language models in simplified versus Traditional Chinese. In Proceedings of the 2025 ACM Conference on Fairness, Accountability, and Transparency (pp. 2815–2846).

    Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., ... & Lowe, R. (2022). Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems, 35, 27730–27744.

    Pan, S. J., & Yang, Q. (2009). A survey on transfer learning. IEEE Transactions on Knowledge and Data Engineering, 22(10), 1345–1359.

    Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., & Sutskever, I. (2019). Language models are unsupervised multitask learners. OpenAI blog, 1(8), 9.

    Rogers, A., Kovaleva, O., & Rumshisky, A. (2020). A primer in BERTology: What we know about how BERT works. Transactions of the Association for Computational Linguistics, 8, 842–866.

    Tenney, I., Das, D., & Pavlick, E. (2019, July). BERT rediscovers the classical NLP pipeline. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics (pp. 4593–4601).

    TAIDE Project Team. (2024). TAIDE: A large language model for Traditional Chinese. Taiwan AI Labs. https://huggingface.co/taide/Gemma-3-TAIDE-12b-Chat

    Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., ... & Scialom, T. (2023). Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288.

    Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., ... & Polosukhin, I. (2017). Attention is all you need. Advances in Neural Information Processing Systems, 30.

    Wei, J., Tay, Y., Bommasani, R., Raffel, C., Zoph, B., Borgeaud, S., ... & Fedus, W. (2022). Emergent abilities of large language models. arXiv preprint arXiv:2206.07682.

    Xie, Q., Han, W., Chen, Z., Xiang, R., Zhang, X., He, Y., ... & Huang, J. (2024). FinBen: A holistic financial benchmark for large language models. Advances in Neural Information Processing Systems, 37, 95716–95743.

    Zhao, W. X., Zhou, K., Li, J., Tang, T., Wang, X., Hou, Y., ... & Wen, J. R. (2023). A survey of large language models. arXiv preprint arXiv:2303.18223, 1(2), 1–124.

    Zhou, C., Liu, P., Xu, P., Iyer, S., Sun, J., Mao, Y., ... & Levy, O. (2023). LIMA: Less is more for alignment. Advances in Neural Information Processing Systems, 36, 55006–55021.

    QR CODE
    :::