| 研究生: |
許証育 Hsu, Cheng-Yu |
|---|---|
| 論文名稱: |
使用決策樹填補遺失值 : 不確定性導向疊代填補與雙層聯合分割最佳化之新方法 A Decision Tree-Based Approach to Missing Value Imputation: Uncertainty-Guided Iterative Imputation and Two-Layer Joint Split Optimization |
| 指導教授: |
張育瑋
Chang, Yu-Wei |
| 口試委員: |
陳怡如
Chen, Yi-Ju 簡立欣 Chien, Li-Hsin |
| 學位類別: |
碩士
Master |
| 系所名稱: |
商學院 - 統計學系 Department of Statistics |
| 論文出版年: | 2026 |
| 畢業學年度: | 114 |
| 語文別: | 中文 |
| 論文頁數: | 74 |
| 中文關鍵詞: | CART 、決策樹 、疊代插補 、遺失值插補 、遺失機制 、遺失納入屬性(MIA) |
| 外文關鍵詞: | CART, Decision Tree, Iterative Imputation, Missing Value Imputation, Missing Data Mechanism, Missingness Incorporated in Attributes |
| 相關次數: | 點閱:63 下載:0 |
| 分享至: |
| 查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報 |
資料中的遺失值是統計分析時常見的現象,若未妥善處理,可能使後續推論產生偏誤。以插補方式將遺失值補齊,是常見的做法之一。本研究延續以決策樹插補遺失值的方向,提出兩種新的插補方法,並比較文獻上數個既有做法 ; 除沿用經典的 CART 外,也會嘗試使用卡方檢定來找分割變數。相較於 Rahman與Islam (2013) 的 DMI (Decision tree-based Missing value Imputation) 方法僅採用毫無遺失的觀測值,本研究採取待插補變數有觀測到的資料皆可當作訓練資料,能更充分地運用觀測值。又因建樹時所選的分割變數本身也可能含有遺失值,本研究除採用文獻上以平均數或眾數讓其通過的做法外,亦運用 MIA (Missingness Incorporated in Attributes ; Twala et al., 2008) 將遺失本身直接納入分割判斷。此外,本研究提出兩種新方法。在本研究中,第一個方法是在 MIA 的基礎上,結合疊代插補與葉節點不確定性指標,將各變數的遺失值疊代插補直到收斂,以應用變數之間的關係,並加入葉節點不確定性指標排除不可信的預測,避免偏差較大的插補值在疊代中被反覆引用而放大誤差 ; 另一方法則在建樹前兩層改採聯合搜尋,取代傳統的逐層貪婪做法,期望於建樹初期即掌握變數間的交互作用。本研究透過四種相關性與三種遺失機制所構成的情境進行模擬研究,以比較各方法的表現,並將兩個所提方法應用於心臟病真實資料,檢視其實際表現。
Missing values are common in statistical analysis and, if not properly handled, may bias subsequent inferences; imputation is one common remedy. This study follows the line of research on decision-tree-based imputation, proposing two new imputation methods and comparing several existing methods. In addition to the classical CART method, chi-square tests are also used to identify splitting variables. Unlike the DMI method of ~Rahman and Islam (2013), which uses only fully observed cases, this study treats any observation with an observed target variable as training data, making fuller use of the sample. Moreover, since a chosen splitting variable may itself be missing, this study adopts both the conventional mean or mode substitution and MIA, which incorporates missingness directly into the splitting decision. Two new methods are further proposed. In the current study, the first extends MIA with iterative imputation, which updates each variable until convergence to exploit inter-variable relationships, and adds a leaf-node uncertainty measure that withholds unreliable predictions, preventing biased values from being reused and amplifying error across iterations. The second replaces the layer-by-layer greedy strategy with a joint search over the first two layers, so that interactions among variables can be captured early in developing the CART. Simulation studies covering four levels of correlation and three missing-data mechanisms have been conducted to compare these methods across scenarios. The practical performance of the two proposed methods is further evaluated using a real-world heart disease dataset.
第一章緒論 1
第二章背景知識 4
2.1 分類樹與迴歸樹 4
2.1.1 分類樹(Classfication Tree) 4
2.1.2 迴歸樹(Regression Tree) 7
2.2 三種遺失值的機制 9
第三章研究方法 11
3.1 MIA 結合疊代與節點內不確定性插補方法(方法八) 11
3.1.1 方法八演算法流程 12
3.1.2 Missingness Incorporated in Attributes, MIA 14
3.1.3 疊代插補法(Iterative Imputation Techniques) 15
3.2 兩層聯合搜尋決策樹插補法(方法九) 17
3.2.1 候選分割點的產生 17
3.2.2 遺失值通過節點的處理策略 18
3.2.3 方法九的完整演算法 19
3.3 方法零至方法七介紹 20
第四章模擬研究 22
4.1 模擬設定 22
4.2 模擬結果之評估準則 25
4.3 M0 至M8 模擬結果 26
4.4 M9 模擬結果 44
第五章資料分析 46
5.1 遺失值之設定 47
5.2 插補結果 50
5.3 邏輯斯迴歸分析結果 53
第六章結論與建議 55
參考文獻 57
附錄A 數值變數之最低RMSE 次數表(M0–M8) 60
附錄B 類別變數之最高準確度次數表(M0–M8) 63
附錄C 類別變數之最高準確度次數表(M0–M9) 66
附錄D 方法9 之準確度箱型圖(M0–M9) 69
Beaulac, C. and Rosenthal, J. S. (2020). BEST: A decision tree algorithm that handles missing values. Computational Statistics, 35(3):1001–1026.
Breiman, L., Friedman, J. H., Olshen, R. A., and Stone, C. J. (1984). Classification and Regression Trees. Wadsworth International Group, Belmont, CA.
Burgette, L. F. and Reiter, J. P. (2010). Multiple imputation for missing data via sequential regression trees. American Journal of Epidemiology, 172(9):1070–1076.
Chen, C.-Y. and Chang, Y.-W. (2024). Missing data imputation using classification and regression trees. PeerJ Computer Science, 10:e2119.
Chen, Y.-W. (2024). 使用決策樹進行遺失值填補之研究. 碩士論文, 國立政治大學統計學系.
Chipman, H. A., George, E. I., and McCulloch, R. E. (2010). BART: Bayesian additive regression trees. The Annals of Applied Statistics, 4(1):266–298.
Fazakis, N., Kostopoulos, G., Kotsiantis, S., and Mporas, I. (2020). Iterative robust semisupervised missing data imputation. IEEE Access, 8:90555–90569.
Hapfelmeier, A., Hothorn, T., and Ulm, K. (2012). Recursive partitioning on incomplete data using surrogate decisions and multiple imputation. Computational Statistics & Data Analysis, 56(6):1552–1565.
Hastie, T., Tibshirani, R., and Friedman, J. (2001). The Elements of Statistical Learning: Data Mining, Inference, and Prediction. Springer, New York.
James, G., Witten, D., Hastie, T., and Tibshirani, R. (2013). An Introduction to Statistical Learning: With Applications in R. Springer, New York.
Kim, H. and Loh, W.-Y. (2001). Classification trees with unbiased multiway splits. Journal of the American Statistical Association, 96(454):589–604.
Little, R. J. A. and Rubin, D. B. (2002). Statistical Analysis with Missing Data. John Wiley & Sons, Hoboken, NJ, 2 edition.
Loh, W.-Y. and Shih, Y.-S. (1997). Split selection methods for classification trees. Statistica Sinica, 7(4):815–840.
Nikfalazar, S., Yeh, C.-H., Bedingfield, S., and Khorshidi, H. A. (2020). Missing data imputation using decision trees and fuzzy clustering with iterative learning. Knowledge and Information Systems, 62(6):2419–2437.
Quinlan, J. R. (1986). Induction of decision trees. Machine Learning, 1(1):81–106.
Quinlan, J. R. (1993). C4.5: Programs for Machine Learning. Morgan Kaufmann Publishers, San Mateo, CA.
Rahman, M. G. and Islam, M. Z. (2013). Missing value imputation using decision trees and decision forests by splitting and merging records: Two novel techniques. Knowledge-Based Systems, 53:51–65.
Rodgers, D. M., Jacobucci, R., and Grimm, K. J. (2021). A multiple imputation approach for handling missing data in classification and regression trees. Journal of Behavioral Data Science, 1(1):127–153.
Rubin, D. B. (1976). Inference and missing data. Biometrika, 63(3):581–592.
Stekhoven, D. J. and Bühlmann, P. (2012). MissForest—non-parametric missing value imputation for mixed-type data. Bioinformatics, 28(1):112–118.
Strobl, C., Boulesteix, A.-L., and Augustin, T. (2007). Unbiased split selection for classification trees based on the Gini index. Computational Statistics & Data Analysis, 52(1):483–501.
Therneau, T. M. and Atkinson, E. J. (2009). An introduction to recursive partitioning using the RPART routines. Technical report, Mayo Foundation.
Twala, B. E. T. H., Jones, M. C., and Hand, D. J. (2008). Good methods for coping with missing data in decision trees. Pattern Recognition Letters, 29(7):950–956.
van Buuren, S. and Groothuis-Oudshoorn, K. (2011). mice: Multivariate imputation by chained equations in R. Journal of Statistical Software, 45(3):1–67.
Xu, D., Daniels, M. J., and Winterstein, A. G. (2016). Sequential BART for imputation of missing covariates. Biostatistics, 17(3):589–602.
全文公開日期 2028/07/28