| 研究生: |
謝育穎 Hsieh, Yu-Ying |
|---|---|
| 論文名稱: |
辨識 AI 生成之金融新聞:以彭博新聞為基礎的統計分類研究 Identifying AI-Generated Financial News: A Statistical Classification Study Based on Bloomberg News |
| 指導教授: |
楊曉文
Yang, Sheau-Wen 余清祥 Yue, Ching-Syang |
| 口試委員: |
李百靈
Li, Pai-Ling 林新沛 Lam, San-Pui |
| 學位類別: |
碩士
Master |
| 系所名稱: |
商學院 - 統計學系 Department of Statistics |
| 論文出版年: | 2026 |
| 畢業學年度: | 114 |
| 語文別: | 中文 |
| 論文頁數: | 98 |
| 中文關鍵詞: | 文字分析 、大型語言模型 、探索性資料分析 、AI 生成文本 、彭博新聞 |
| 外文關鍵詞: | Text analysis, large language models, exploratory data analysis, AI-generated text, Bloomberg News |
| 相關次數: | 點閱:13 下載:0 |
| 分享至: |
| 查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報 |
文字是人類傳遞思想、紀錄事件與延續文化的重要媒介。相較於語言具有即逝性,文字能突破時空限制得以跨世代累積知識。不同領域的文本因使用情境與受眾需求,在形式、風格與資訊密度上呈現不同特徵,財經新聞作為投資者、分析師及政策制定者掌握市場動態的重要資訊來源,其內容品質與可信度直接影響金融決策。近年大型語言模型快速普及,部分財經媒體已開始使用 AI 生成新聞與理財內容,像是專為金融領域設計的 BloombergGPT 於 2023 年發表,然而曾發生AI 文章因計算錯誤與事實不符,凸顯建立文本辨識機制的重要性。本研究以 2023至 2025 年彭博(Bloomberg)新聞為研究對象,蒐集人工與 AI 生成新聞,透過探索性資料分析與統計方法比較兩類文本的寫作風格。
本文依據兩類文本在句長、詞彙多樣性及共現詞等基本文字特徵,加上依存距離、句法排列效率及音韻近似等語言特徵,代入羅吉斯迴歸與機器學習模型,再與深度學習模型比較模型優劣。另外,由於真人、電腦文字報導的篇幅差異很大,本文也探討如何公平選取文章位置及長度,有效篩選辨識報導風格的重要變數。結果顯示,人工新聞的句法結構較為穩定,AI 新聞則呈現長、短距離依存交雜,句法組織較鬆散,且較少使用插入語、從屬子句及修飾結構。以探索性資料分析篩選出的重要特徵後,羅吉斯迴歸、隨機森林與 XGBoost 均能在低維度特徵下,達到接近深度學習模型的辨識效能,挑選出的文字特徵也具有可解釋性,顯示本文方法兼具統計可信度、運算效率及解釋透明度,可作為金融資訊受眾辨識新聞寫作風格與區分作者的參考。
Text is an important medium for communicating ideas, recording events, and pre- serving cultural heritage. Among different types of texts, financial news serves as a vital source for investors, analysts, and policymakers, and its quality and credibility directly influence financial decision-making. With the rapid development of large language models (LLMs), some finan cial media have begun using artificial intelligence (AI) to generate news and financial content. BloombergGPT, a domain-specific large language model for finance released in 2023, exemplifies this trend. However, incidents involving AI-generated articles containing calculation errors and factual inconsistencies have highlighted the importance of developing reliable methods for distinguishing AI-generated text from human-written content.
This study investigates Bloomberg financial news published between 2023 and 2025 and by collecting both human-written and AI-generated articles. Exploratory data analysis and statistical methods are employed to compare the writing styles of the two types of texts. The analysis incorporates fundamental textual features, including sentence length, lexical diversity, and word co-occurrence, together with linguistic features such as dependency distance, syntactic arrangement efficiency, and phonological approximation. These features are incorporated into Logistic Regression and several machine learning models and are further compared with deep learning models. In addition, because human-written and AI-generated news articles differ substantially in length, this study further investigates how to fairly select article positions and text lengths to effectively identify discriminative writing-style features.
The results indicate that human-written news exhibits more stable syntactic structures, whereas AI-generated news is characterized by a mixture of short- and long-distance dependency relations, resulting in looser syntactic organization and less frequent use of parenthetical expressions, subordinate clauses, and modifying structures. Using the key features selected through exploratory data analysis, Logistic Regression, Random Forest, and XGBoost achieve classification performance comparable to that of deep learning models while relying on a low-dimensional feature set. The selected features are also highly interpretable, demonstrating that the proposed approach combines statistical reliability, computational efficiency, and interpretability for distinguishing between human-written and AI-generated financial news.
第一章 緒論 1
第一節 研究背景與動機 1
第二節 研究目的 3
第二章 文獻探討與資料介紹 5
第一節 文獻回顧 5
第二節 資料介紹 7
第三章 研究方法 13
第一節 斷詞與文本前處理 13
第二節 探索性資料分析 15
第三節 實驗設計與資料切分 15
第四節 特徵工程 18
第五節 分類模型 32
第六節 大語言模型 35
第四章 新聞文本分析 37
第一節 探索性資料分析 37
第二節 特徵選擇與分類結果-全文 47
第三節 切分實驗結果 56
第四節 特徵重要性與解釋 63
第五章 驗證性分析 70
第一節 驗證設計與資料設定 70
第二節 年度內穩定性分析 72
第三節 跨年度泛化分析 74
第四節 年度內與跨年度結果之綜合比較 76
第六章 結論與建議 79
第一節 結論 79
第二節 研究限制與建議 80
參考文獻 83
附錄一、全文特徵之分布 86
附錄二、Segment 特徵之分布 94
一、中文文獻
余清祥(1998)。統計在紅樓夢的應用。國立政治大學學報,76,303– 327。
余清祥、葉昱廷(2020)。以文字探勘技術分析臺灣四大報文字風格。數位典藏與數位人文,(6),69– 96。
唐士哲(2024)。生成式人工智慧、新聞室自動化與變遷中的新聞樣貌。文化:政策・管理・新創,3(1),9– 27。
二、英文文獻
Araci, D. (2019). FinBERT: Financial sentiment analysis with pre-trained language mod-els. arXiv.
Associated Press. (2015, June 15). Automated earnings stories multiply. The AssociatedPress.
Chalkidis, I., Fergadiotis, M., Malakasiotis, P., Aletras, N., & Androutsopoulos, I. (2020).LEGAL-BERT: The Muppets straight out of law school. In Findings of the Association for Computational Linguistics: EMNLP 2020 (pp. 2898– 2904). Association for Computational Linguistics.
Devlin, J., Chang, M.-W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics:Human Language Technologies, Volume 1 (Long and Short Papers) (pp. 4171– 4186).
Association for Computational Linguistics.
Dickens, C. (1856). The wreck of the Golden Mary. Household Words, Extra Christmas Number.
Ferrer-i-Cancho, R., Gómez-Rodríguez, C., Esteban, J. L., & Alemany-Puig, L. (2022).Optimality of syntactic dependency distances. Physical Review E, 105(1), 014308.
Firth, J. R. (1957). A synopsis of linguistic theory, 1930– 1955. In Studies in linguistic analysis (pp. 1– 32). Blackwell.
Georgiou, G. P. (2025). Differentiating between human-written and AI-generated texts using automatically extracted linguistic features. Information, 16(11), Article 979.
Gibson, E. (1998). Linguistic complexity: Locality of syntactic dependencies. Cognition,68(1), 1– 76.
Gu, Y., Tinn, R., Cheng, H., Lucas, M., Usuyama, N., Liu, X., Naumann, T., Gao, J.,& Poon, H. (2021). Domain-specific language model pretraining for biomedical natural language processing. ACM Transactions on Computing for Healthcare, 3(1), Article 2.
Lee, J., Yoon, W., Kim, S., Kim, D., Kim, S., So, C. H., & Kang, J. (2020). BioBERT: A pre-trained biomedical language representation model for biomedical text mining. Bioinformatics, 36(4), 1234– 1240.
Liu, H. (2008). Dependency distance as a metric of language comprehension difficulty. Journal of Cognitive Science, 9(2), 159– 191.
Loughran, T., & McDonald, B. (2011). When is a liability not a liability? Textual analysis, dictionaries, and 10-Ks. The Journal of Finance, 66(1), 35– 65.
Lundberg, S. M., & Lee, S.-I. (2017). A unified approach to interpreting model predictions. In Advances in Neural Information Processing Systems, 30 (pp. 4768– 4777).
Mendenhall, T. C. (1887). The characteristic curves of composition. Science, 9(214S), 237– 246.
Mosteller, F., & Wallace, D. L. (1964). Inference and disputed authorship: The Federalist. Addison-Wesley.
Muñoz-Ortiz, A., Gómez-Rodríguez, C., & Vilares, D. (2024). Contrasting linguistic patterns in human- and LLM-generated news text. Artificial Intelligence Review, 57, Article 265.
OpenAI. (2023, January 31). New AI classifier for indicating AI-written text.
Peterson, K. (2008, September 9). UAL shares walloped by new posting of old news. Reuters.
Ribeiro, M. T., Singh, S., & Guestrin, C. (2016). “Why should I trust you?”: Explaining the predictions of any classifier. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (pp. 1135– 1144). Association for Computing Machinery.
Selyukh, A. (2013, April 23). Hackers send fake market-moving AP tweet on White House explosions. Reuters.
Stamatatos, E. (2009). A survey of modern authorship attribution methods. Journal of the American Society for Information Science and Technology, 60(3), 538– 556.
Tetlock, P. C. (2007). Giving content to investor sentiment: The role of media in the stock market. The Journal of Finance, 62(3), 1139– 1168.
Wu, S., Irsoy, O., Lu, S., Dabravolski, V., Dredze, M., Gehrmann, S., Kambadur, P., Rosenberg, D., & Mann, G. (2023). BloombergGPT: A large language model for finance. arXiv.
Xiao, C., Hu, X., Liu, Z., Tu, C., & Sun, M. (2021). Lawformer: A pre-trained language model for Chinese legal long documents. AI Open, 2, 79– 84.
全文公開日期 2031/08/24