| 研究生: |
陳泓霖 Chen, Hung-Lin |
|---|---|
| 論文名稱: |
引導式多概念挖掘模型之應用、模擬與驗證 Guided Diverse Concept Miner: Application, Simulation and Validation |
| 指導教授: |
莊皓鈞
Chuang, Howard Hao-Chun 周彥君 Chou, Yen-Chun |
| 口試委員: |
陳柏安
Chen, Po-An |
| 學位類別: |
碩士
Master |
| 系所名稱: |
商學院 - 資訊管理學系 Department of Management Information System |
| 論文出版年: | 2026 |
| 畢業學年度: | 114 |
| 語文別: | 中文 |
| 論文頁數: | 54 |
| 中文關鍵詞: | 財金文本分析 、引導式主題模型 、電話會議文本探勘 、模擬實驗 、估計偏誤 |
| 外文關鍵詞: | Financial Text Analysis, Guided Topic Model, Earnings Call Text Mining, Simulation Study, Estimation Bias |
| 相關次數: | 點閱:145 下載:0 |
| 分享至: |
| 查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報 |
主題模型長期是解析大規模非結構化文本的核心工具,但傳統模型難以同時兼顧可解釋性與預測能力。Lee et al.(2025)提出的引導式多概念挖掘模型(Guided Diverse Concept Miner, GDCM),透過將概念萃取與外部管理結果變數聯合最佳化,成功在電商評論資料上同時達成高度可解釋性與預測力。然而,該模型是否能推廣至篇幅更長、雜訊更高、議題更為交織的財金文本,仍屬未知。本研究以美國零售業上市公司之電話會議逐字稿為對象,檢驗 GDCM 在此類複雜語料下的適用邊界,並提出第一個研究問題(RQ1):GDCM 應用於高雜訊真實企業文本時,其概念萃取與預測效度受到何種程度影響?
為回答 RQ1,本研究蒐集 2011 年至 2020 年間 117 家北美零售業上市公司、共 3,673 筆「公司—季度」層級之電話會議逐字稿,並以存貨營收比作為引導標籤,設計三組對照實驗:以真實財務指標引導之有效監督組、以隨機雜訊為標籤之安慰劑組,以及完全關閉監督機制之無監督組。結果顯示,有效監督組可萃取出語意連貫且與存貨營收比之經濟意涵相符的概念(如亞太觀光客流、關稅與供應鏈壓力),測試集 R² 達 0.170;然而,隨機雜訊組雖然預測力明顯低落,卻仍對部分概念賦予非零的解釋權重,顯示模型在面對無效標籤時可能產生具方向性的虛假關聯。由於真實文本缺乏已知的語意基準(Ground Truth),本研究無法判定此一現象究竟反映模型捕捉到微弱的真實訊號,抑或僅是最佳化過程中的過度擬合,此一方法論疑慮促使本研究進一步提出第二個研究問題(RQ2):在具備已知真實結構的模擬環境下,GDCM 之概念萃取與係數估計表現為何?
為回答 RQ2 並釐清上述疑慮,本研究進一步設計具備已知潛在主題結構的模擬實驗,透過操弄監督訊號強度(ρ)與主題—概念映射複雜度,系統性檢驗模型的結構還原能力。結果顯示,測試集 R² 與概念分布之真實還原程度存在高度相關(r > 0.97),可作為實務上缺乏基準時判斷概念萃取可信度之代理指標;但係數估計之精確度高度依賴映射複雜度:在簡單映射下接近無偏,在複雜映射下即使提高監督權重仍存在無法消除之系統性偏誤,僅能反映概念間之相對重要性排序。此外,在監督係數為零的對照情境下,模型並未產生具方向性的系統性虛假權重,部分緩解了 RQ1 中隨機雜訊組所引發之方法論疑慮,惟此一結論仍受限於模擬設定之低維度與簡化假設,推廣至實務時應保持審慎。
本研究之貢獻在於:一、首次將 GDCM 應用於高雜訊、長篇幅之財金文本,檢驗其在跨語料情境下的泛化能力;二、透過具備已知基準之模擬實驗,釐清引導式主題模型在監督訊號薄弱時可能產生的估計偏誤與其邊界條件;三、提出以測試集預測表現作為篩選引導標籤與解讀概念可信度之操作準則,供後續研究者於缺乏基準的實證情境中參考。
Topic models remain a core tool for analyzing large-scale unstructured text, yet conventional models struggle to balance interpretability and predictive power. The Guided Diverse Concept Miner (GDCM), proposed by Lee et al. (2025), jointly optimizes concept extraction with an external managerial outcome variable and has demonstrated strong interpretability and predictive performance on e-commerce review data. Whether this performance generalizes to longer, noisier, and topically more entangled financial text, however, remains untested. This thesis examines GDCM's applicability boundary using earnings conference call transcripts from publicly listed U.S. retail firms, guided by an initial research question (RQ1): To what extent does GDCM's concept extraction and predictive validity hold up when applied to noisy real-world corporate transcripts?
To address RQ1, this study collects 3,673 firm-quarter earnings call transcripts from 117 North American retail firms between 2011 and 2020, using the inventory-to-sales ratio as the guiding label under three experimental conditions: a validly supervised condition guided by a genuine financial indicator, a placebo condition guided by random noise, and an unsupervised condition with supervision disabled entirely. Results show that the validly supervised condition extracts semantically coherent concepts consistent with the economic intuition of inventory-to-sales (e.g., Asia-Pacific tourist flows, tariff and supply-chain pressure), achieving a test set R² of 0.170. The random-noise condition, despite markedly lower predictive power, still assigns non-zero explanatory weights to some concepts—suggesting the model may produce directionally spurious associations under uninformative labels. Because real-world text lacks a known semantic ground truth, this study cannot determine whether this phenomenon reflects a weak but genuine signal or mere overfitting during optimization; this methodological ambiguity motivates a second research question (RQ2): In a controlled simulation environment with known ground truth, how does GDCM perform in concept recovery and coefficient estimation?
To address RQ2 and resolve this ambiguity, a simulation study with known latent topic structure is designed, systematically varying supervision strength (ρ) and topic-to-concept mapping complexity. Results show that test R² is strongly correlated with the true recovery of the concept distribution (r > 0.97), supporting its use as an observable proxy for concept reliability when no ground truth is available in practice. However, the precision of coefficient estimation depends heavily on mapping complexity: estimates are nearly unbiased under simple one-to-one mappings, but exhibit a systematic bias under complex mappings that persists even at high supervision weights, such that coefficients remain informative only for relative concept ranking rather than precise magnitude. Furthermore, under a null-coefficient benchmark condition, the model does not produce directionally systematic spurious weights, partially alleviating the methodological concern raised by the random-noise condition in RQ1—though this conclusion is bounded by the simulation's low-dimensional, simplified setting and should be generalized to practice with caution.
This thesis contributes by: (1) providing the first application of GDCM to noisy, long-form financial text, testing its generalizability across corpus types; (2) using a simulation with known ground truth to clarify the estimation bias and boundary conditions of guided topic models under weak supervision; and (3) proposing test-set predictive performance as a practical criterion for screening guiding labels and assessing concept reliability in the absence of ground truth.
摘要 i
Abstract iii
目次 v
表次 vii
圖次 viii
第一章 緒論 1
第二章 文獻探討 4
第一節 主題模型之定位與演進 4
第二節 Guided Diverse Concept Miner (GDCM) 7
一、文字、文件及概念的空間映射 8
二、鼓勵概念稀疏性 10
三、鼓勵概念多樣性 11
四、使用管理結果變數引導模型 11
第三章 實證案例分析 14
第一節 研究資料 14
第二節 實證設計 16
第三節 結果與討論 19
第四章 模擬實驗分析 24
第一節 數據生成過程與實驗設定 24
第二節 監督訊號強度與結構還原分析 34
第三節 無效標籤與模型穩健性測試 45
第五章 結論 49
第一節 研究結論 49
第二節 管理意涵與研究限制 50
參考文獻 52
Blei, D. M., Ng, A. Y., & Jordan, M. I. (2003). Latent Dirichlet allocation. Journal of Machine Learning Research, 3, 993–1022.
Devlin, J., Chang, M.-W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of deep bidirectional transformers for language understanding. Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), 4171–4186. https://doi.org/10.18653/v1/N19-1423
Dieng, A. B., Ruiz, F. J. R., & Blei, D. M. (2020). Topic modeling in embedding spaces. Transactions of the Association for Computational Linguistics, 8, 439–453. https://doi.org/10.1162/tacl_a_00325
Grootendorst, M. (2022). BERTopic: Neural topic modeling with a class-based TF-IDF procedure. arXiv. https://doi.org/10.48550/arXiv.2203.05794
Jagarlamudi, J., Daumé III, H., & Udupa, R. (2012). Incorporating lexical priors into topic models. Proceedings of the 13th Conference of the European Chapter of the Association for Computational Linguistics, 204–213.
Kesavan, S., Gaur, V., & Raman, A. (2010). Do inventory and gross margin data improve sales forecasts for U.S. public retailers? Management Science, 56(9), 1519–1533. https://doi.org/10.1287/mnsc.1100.1209
Kesavan, S., & Mani, V. (2013). The relationship between abnormal inventory growth and future earnings for U.S. public retailers. Manufacturing & Service Operations Management, 15(1), 6–23. https://doi.org/10.1287/msom.1120.0389
Kingma, D. P., & Welling, M. (2022). Auto-encoding variational Bayes. arXiv. https://doi.org/10.48550/arXiv.1312.6114
Lee, D. "Dk", Cheng, Z. "Zq", Mao, C., & Manzoor, E. (2025). Guided Diverse Concept Miner (GDCM): Uncovering relevant constructs for managerial insights from text. Information Systems Research, 36(1), 370–393. https://doi.org/10.1287/isre.2020.0494
Matsumoto, D., Pronk, M., & Roelofsen, E. (2011). What makes conference calls useful? The information content of managers' presentations and analysts' discussion sessions. The Accounting Review, 86(4), 1383–1414. https://doi.org/10.2308/accr-10034
McAuliffe, J. D., & Blei, D. M. (2007). Supervised topic models. Advances in Neural Information Processing Systems 20 (NIPS 2007), 20, 121–128.
Mikolov, T., Chen, K., Corrado, G., & Dean, J. (2013). Efficient estimation of word representations in vector space. arXiv. https://doi.org/10.48550/arXiv.1301.3781
Röder, M., Both, A., & Hinneburg, A. (2015). Exploring the space of topic coherence measures. Proceedings of the Eighth ACM International Conference on Web Search and Data Mining, 399–408. https://doi.org/10.1145/2684822.2685324
Sridhar, D., Daumé, H., & Blei, D. (2022). Heterogeneous supervised topic models. Transactions of the Association for Computational Linguistics, 10, 732–745. https://doi.org/10.1162/tacl_a_00487
Srivastava, A., & Sutton, C. (2017). Autoencoding variational inference for topic models. arXiv. https://doi.org/10.48550/arXiv.1703.01488
Valle, D., Mintz, J., & Brack, I. V. (2024). Estimation and interpretation problems and solutions when using proportion covariates in linear regression models. Ecology, 105(4), e4256. https://doi.org/10.1002/ecy.4256
Wu, D. (2022). Text-based measure of supply chain risk exposure (SSRN Scholarly Paper No. 4158073). https://doi.org/10.2139/ssrn.4158073
Yang, X., Zhao, H., Xu, W., Qi, Y., Lu, J., Phung, D., & Du, L. (2025). Neural topic modeling with large language models in the loop. arXiv. https://doi.org/10.48550/arXiv.2411.08534
Yang, Y., Zhang, K., & Fan, Y. (2023). sDTM: A supervised Bayesian deep topic model for text analytics. Information Systems Research, 34(1), 137–156. https://doi.org/10.1287/isre.2022.1124
全文公開日期 2031/07/25