| 研究生: |
駱泳誌 Lo, Yung-Chih |
|---|---|
| 論文名稱: |
增強式殘差適配器預訓練方法 The Augmentative Residual Adapter Approach to Pre-training |
| 指導教授: |
蔡瑞煌
Tsaih, Rua Huan |
| 口試委員: |
王永鐘
張家銘 |
| 學位類別: |
碩士
Master |
| 系所名稱: |
商學院 - 資訊管理學系 Department of Management Information System |
| 論文出版年: | 2026 |
| 畢業學年度: | 115 |
| 語文別: | 英文 |
| 論文頁數: | 83 |
| 中文關鍵詞: | 增強式預訓練 、增強式殘差適配器 、凍結骨幹模型 、因果語言建模 、金融領域適配 、問題與語料對齊 |
| 外文關鍵詞: | Augmentative Pre-Training, augmentative residual adapter, frozen backbone adaptation, causal language modeling, financial domain adaptation, Question–Corpus Alignment |
| 相關次數: | 點閱:10 下載:0 |
| 分享至: |
| 查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報 |
以未標註語料進行專業領域適配時,適配器位置、固定可訓練容量下的資料配置,以及語料與預期應用的關係,都是關鍵設計問題。本研究提出增強式預訓練(Augmentative Pre-Training, APT),以自監督因果語言建模訓練增強式殘差適配器(augmentative residual adapter, AuRa),並凍結 Llama-3.1-TAIDE-LX-8B-Chat 骨幹模型。本研究使用未標註臺灣金融語料進行兩項控制實驗與盲測。
H1 比較淺層、中層與深層插入位置;一個訓練週期後,淺層區塊 1–2 取得最低驗證 CLM 損失 1.7890。H2 固定淺層位置與 33.55M 個可訓練 AuRa 參數,比較r18、r20 與 r22 三種資料量;r22 取得最低損失 1.7385,而 r18 優於 r20,顯示資料量與最佳化設定共同影響結果。20 位參與者的 100 筆盲測紀錄中,原始TAIDE、H1-best 與 H2-best 分別獲得 63、21 與 16 筆最佳回答;在較接近主要公司揭露語料的問題中,H1-best 與 H2-best 被選為最佳回答的合計比例較高。
整體而言,AuRa 插入位置、固定容量下的資料配置與最佳化設定,共同影響APT 對目標語料的建模結果;下游表現則與語料對預期問題的涵蓋程度相關。據此,本研究將 APT 的設計與評估連結為兩個層次:在訓練層次比較插入位置與資料配置,在應用層次依問題分布規劃語料組成與下游評估。
Adapting a large language model to a specialized domain with unlabeled text raises three design questions: where adapters should be placed, how training data should be allocated relative to a fixed trainable capacity, and how corpus adaptation relates to the intended application. This thesis proposes Augmentative Pre-Training (APT), a parameter-efficient pre-training method that trains augmentative residual adapter (AuRa) modules through self-supervised causal language modeling while keeping the Llama-3.1-TAIDE-LX-8B-Chat backbone frozen. Using an unlabeled Taiwanese financial corpus, the study conducts two controlled experiments and a blind human preference test.
H1 compares shallow, middle, and deep placements. After one epoch, shallow blocks 1–2 achieve the lowest validation CLM loss of 1.7890. H2 fixes the shallow placement and 33.55 million trainable AuRa parameters while comparing r18, r20, and r22. The r22 condition reaches the lowest loss of 1.7385, while r18 outperforms r20, indicating that data scale and the selected optimization settings jointly shape the observed result. In 100 records from a blind preference test completed by 20 participants, the original TAIDE baseline, H1-best, and H2-best receive 63, 21, and 16 best-answer selections, respectively. Descriptively, the two APT checkpoints account for 50.0% of best selections among questions closer to the dominant disclosure corpus, compared with 32.9% among questions less represented in that corpus.
Overall, AuRa placement, data allocation under a fixed capacity, and optimization settings jointly shape how APT models the target corpus. The downstream analysis further associates model preference with how well the corpus covers the intended question distribution. Accordingly, this thesis connects APT design and evaluation at two levels: comparing placement and data allocation at the training level, and aligning corpus composition and downstream evaluation with the expected questions at the application level.
Contents iii
List of Figures v
List of Tables vii
Chapter 1 Introduction 1
Chapter 2 Related Work 5
2.1 Overview 5
2.2 Domain Adaptation for LLMs 6
2.3 Layer Representations in LLMs 7
2.4 Scaling Laws in Transfer Learning 7
2.5 Llama-3.1-TAIDE-LX-8B-Chat 8
Chapter 3 Augmentative Pre-Training 13
Chapter 4 Experimental Setup and Evaluation Design 19
4.1 Corpus and Training Setup 19
4.2 Experimental Design 23
4.3 Evaluation and Test Design 26
Chapter 5 Results and Analysis 31
5.1 Validation CLM Loss Results 31
5.2 Human Preference Test Results 41
5.3 Question–Corpus Alignment Analysis 44
Chapter 6 Conclusion 47
Bibliography 53
Appendix A Human Evaluation Metadata and Participant Records 57
A.1 Human Preference Test Protocol 57
A.2 Participant Profile Records 58
A.3 Decoded Record-Level Outcome Index 59
Appendix B Question–Corpus Alignment Coding 63
B.1 Coding Scope and Operational Criteria 63
B.2 Complete Question Coding Index 64
Appendix C Representative Cases from the Human Preference Test 73
Aguda, T. D., Siddagangappa, S., Kochkina, E., Kaur, S., Wang, D., & Smiley, C. (2024). Large language models as financial data annotators: A study on effectiveness and efficiency. In N. Calzolari, M.-Y. Kan, V. Hoste, A. Lenci, S. Sakti, & N. Xue (Eds.), Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024) (pp. 10124–10145). ELRA and ICCL. https://aclanthology.org/2024.lrec-main.885/
Chen, T., Tan, Z., Gong, T., Wu, Y., Chu, Q., Liu, B., Ye, J., & Yu, N. (2024). Llama Slayer 8B: Shallow layers hold the key to knowledge injection. In Y. Al-Onaizan, M. Bansal, & Y.-N. Chen (Eds.), Findings of the Association for Computational Linguistics: EMNLP 2024 (pp. 5991–6002). Association for Computational Linguistics. https://doi.org/10.18653/v1/2024.findings-emnlp.347
Colombo, P., Pires, T. P., Boudiaf, M., Culver, D., Melo, R., Corro, C., Martins, A. F. T., Esposito, F., Raposo, V. L., Morgado, S., & Desa, M. (2024). SaulLM-7B: A pioneering large language model for law. https://arxiv.org/abs/2403.03883
Dubey, A., Grattafiori, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Vaughan, A., Yang, A., Fan, A., Goyal, A., Hartshorn, A., Yang, A., Mitra, A., Sravankumar, A., Korenev, A., Hinsvark, A., … Ma, Z. (2024, November 23). The Llama 3 herd of models. arXiv:2407.21783 [cs]. https://doi.org/10.48550/arXiv.2407.21783
French, R. M. (1999). Catastrophic forgetting in connectionist networks. Trends in Cognitive Sciences, 3(4), 128–135. https://doi.org/10.1016/S1364-6613(99)01294-2
Gupta, K., Thérien, B., Ibrahim, A., Richter, M. L., Anthony, Q., Belilovsky, E., Rish, I., & Lesort, T. (2023, September 6). Continual pre-training of large language models: How to (re)warm your model? arXiv:2308.04014 [cs]. https://doi.org/10.48550/arXiv.2308.04014
Hernandez, D., Kaplan, J., Henighan, T. J., & McCandlish, S. (2021). Scaling laws for transfer. ArXiv, abs/2102.01293. https://api.semanticscholar.org/CorpusID:231749962
Hoffmann, J., Borgeaud, S., Mensch, A., Buchatskaya, E., Cai, T., Rutherford, E., de Las Casas, D., Hendricks, L. A., Welbl, J., Clark, A., Hennigan, T., Noland, E., Millican, K., van den Driessche, G., Damoc, B., Guy, A., Osindero, S., Simonyan, K., Elsen, E., … Sifre, L. (2022). Training compute-optimal large language models. Proceedings of the 36th International Conference on Neural Information Processing Systems.
Kaplan, J., McCandlish, S., Henighan, T. J., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., & Amodei, D. (2020). Scaling laws for neural language models. ArXiv, abs/2001.08361. https://api.semanticscholar.org/CorpusID:210861095
Li, C.-A., & Lee, H.-Y. (2024). Examining forgetting in continual pre-training of aligned large language models. https://arxiv.org/abs/2401.03129
Ling, C., Zhao, X., Lu, J., Deng, C., Zheng, C., Wang, J., Chowdhury, T., Li, Y., Cui, H., Zhang, X., Zhao, T., Panalkar, A., Mehta, D., Pasquali, S., Cheng, W., Wang, H., Liu, Y., Chen, Z., Chen, H., … Zhao, L. (2025). Domain specialization as the key to make large language models disruptive: A comprehensive survey. ACM Computing Surveys, 58(3), 79. https://doi.org/10.1145/3764579
Meng, K., Bau, D., Andonian, A., & Belinkov, Y. (2022). Locating and editing factual associations in GPT. Proceedings of the 36th International Conference on Neural Information Processing Systems.
Skean, O., Arefin, M. R., Zhao, D., Patel, N., Naghiyev, J., LeCun, Y., & Shwartz-Ziv, R. (2025). Layer by layer: Uncovering hidden representations in language models. ArXiv, abs/2502.02013. https://api.semanticscholar.org/CorpusID:276107264
Song, Z., Yan, B., Liu, Y., Fang, M., Li, M., Yan, R., & Chen, X. (2025). Injecting domain-specific knowledge into large language models: A comprehensive survey. In C. Christodoulopoulos, T. Chakraborty, C. Rose, & V. Peng (Eds.), Findings of the Association for Computational Linguistics: EMNLP 2025 (pp. 25297–25311). Association for Computational Linguistics. https://doi.org/10.18653/v1/2025.findings-emnlp.1379
TAIDE Team. (2024a). Llama-3.1-TAIDE-LX-8B-Chat. Hugging Face. https://huggingface.co/taide/Llama-3.1-TAIDE-LX-8B-Chat
TAIDE Team. (2024b). Trustworthy AI Dialogue Engine (TAIDE). https://taide.tw/
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., & Polosukhin, I. (2017). Attention is all you need. In I. Guyon, U. V. Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, & R. Garnett (Eds.), Advances in Neural Information Processing Systems (Vol. 30). Curran Associates, Inc. https://proceedings.neurips.cc/paper_files/paper/2017/file/3f5ee243547dee91fbd053c1c4a845aa-Paper.pdf
Wu, S., Irsoy, O., Lu, S., Dabravolski, V., Dredze, M., Gehrmann, S., Kambadur, P., Rosenberg, D., & Mann, G. (2023). BloombergGPT: A large language model for finance. ArXiv, abs/2303.17564. https://api.semanticscholar.org/CorpusID:257833842
Xie, Q., Chen, Q., Chen, A., Peng, C., Hu, Y., Lin, F., Peng, X., Huang, J., Zhang, J., Keloth, V., Zhou, X., He, H., Ohno-Machado, L., Wu, Y., Xu, H., & Bian, J. (2024). Me-LLaMA: Foundation large language models for medical applications [Version 1.0.0]. PhysioNet. https://doi.org/10.13026/wwfd-2t39
Xie, Q., Han, W., Zhang, X., Lai, Y., Peng, M., Lopez-Lira, A., & Huang, J. (2023). PIXIU: A large language model, instruction data and evaluation benchmark for finance. Proceedings of the 37th International Conference on Neural Information Processing Systems.
Zhang, D., Feng, T., Xue, L., Wang, Y., Dong, Y., & Tang, J. (2025). Parameter-efficient fine-tuning for foundation models. arXiv preprint arXiv:2501.13787.
Zhao, W. X., Zhou, K., Li, J., Tang, T., Wang, X., Hou, Y., Min, Y., Zhang, B., Zhang, J., Dong, Z., Du, Y., Yang, C., Chen, Y., Chen, Z., Jiang, J., Ren, R., Li, Y., Tang, X., Liu, Z., … Wen, J.-R. (2023). A survey of large language models. https://arxiv.org/abs/2303.18223