| 研究生: |
戴衣伶 Dai, Yi-Ling |
|---|---|
| 論文名稱: |
局部加權樹模型分量迴歸於空間異質性之分析 Locally Weighted Tree-Based Quantile Regression for Spatial Heterogeneity Analysis |
| 指導教授: |
陳怡如
Chen, Yi-Ju |
| 口試委員: |
吳漢銘
Wu, Han-Ming 李百靈 Li, Pai-Ling |
| 學位類別: |
碩士
Master |
| 系所名稱: |
商學院 - 統計學系 Department of Statistics |
| 論文出版年: | 2026 |
| 畢業學年度: | 114 |
| 語文別: | 中文 |
| 論文頁數: | 113 |
| 中文關鍵詞: | 分量迴歸 、隨機森林 、極限梯度提升 、空間異質性 、SHAP 、地理加權分量迴歸 |
| 外文關鍵詞: | Quantile Regression, Random Forest, XGBoost, Spatial heterogeneity, SHAP, Geographically Weighted Quantile Regression |
| 相關次數: | 點閱:114 下載:0 |
| 分享至: |
| 查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報 |
空間資料廣泛應用於社會科學、公共衛生、區域經濟及環境研究等不同領域,其分析除須考量空間異質性外,反應變數於不同條件分位數下之分布異質性亦為重要研究議題。為此,地理加權分量迴歸(geographically weighted quantile regression, GWQR)透過局部建模與條件分位數估計,可同時描述變數關係隨空間位置及條件分布之變化。然而,GWQR 仍建立於線性模型架構之上,較難處理複雜非線性關係、高階變數交互作用及高維度資料。另一方面,分量迴歸森林(quantile regression forest, QRF)與分量極限梯度提升(quantile XGBoost, QXGBoost)具備非線性分量學習能力,但既有研究將兩者應用於空間資料分析時,多採全域建模方式,或以距離加權及空間內插方式處理空間資訊,而非於各地理位置建立局部模型,因此難以反映變數關係於不同地理位置及不同條件分位數下之局部變化,亦較少著墨於局部模型估計與變數效果分析。有鑑於此,本研究提出一套局部加權樹模型分量迴歸(locally weighted tree-based quantile regression)分析架構,整合地理加權局部建模、樹模型集成學習及分量估計之概念,發展出地理加權分量迴歸森林(geographically weighted quantile regression forest, GW-QRF)及地理加權分量極限梯度提升(geographically weighted quantile XGBoost, GW-QXGB)兩種方法。兩者皆透過空間權重建構局部分量樹模型,使模型得以同時捕捉空間異質性、非線性關係及條件分布異質性。此外,本研究結合 SHapley Additive exPlanations(SHAP),探討不同地理位置及不同條件分位數下各解釋變數之局部貢獻,並並透過拔靴法評估其穩定性。本研究進一步透過模擬實驗,評估所提方法於不同空間異質情境下之估計與預測表現,並以四組公開空間資料進行實證分析,與傳統 QR、GWQR、QRF 及 QXGBoost 等既有方法進行比較。結果顯示,GW-QRF 與 GW-QXGB 於多數情境下具有較佳之預測表現,並可降低模型殘差之空間自相關;SHAP 分析亦能呈現不同空間位置及不同條件分位數下解釋變數之局部貢獻差異。本研究所建立之局部加權樹模型分量迴歸,能同時處理空間異質性、非線性關係及條件分布異質性,並結合以 SHAP 為基礎之局部模型可解釋性分析,不僅擴展地理加權分量迴歸之建模彈性,亦為空間機器學習提供一套新的局部分量分析方法。
Spatial data are widely encountered in social sciences, public health, regional economics, environmental studies, and many other disciplines. In addition to spatial heterogeneity, distributional heterogeneity of the response variable across different conditional quantiles is also an important consideration in spatial data analysis. Geographically weighted quantile regression (GWQR) was developed to jointly address both issues by combining localized modeling with conditional quantile estimation, allowing variable relationships to vary across both geographic location and the conditional distribution. However, GWQR is fundamentally based on a linear modeling framework, which limits its flexibility in capturing complex nonlinear relationships, high-order interactions, and high-dimensional data structures. Meanwhile, quantile regression forest (QRF) and quantile XGBoost (QXGBoost) possess powerful nonlinear quantile learning capabilities. Existing spatial extensions of these methods either adopt global learning strategies or incorporate geographical information through sampling weighting or spatial interpolation, rather than calibrating localized models at each geographic location. Consequently, they are less effective in characterizing spatially varying quantile relationships and provide limited insight into localized model estimation and variable effects. To address these gaps, this study proposes a locally weighted tree-based quantile regression framework that integrates geographically weighted local modeling, tree-based ensemble learning, and quantile estimation. Within this framework, two localized methods are developed: geographically weighted quantile regression forest (GW-QRF) and geographically weighted quantile XGBoost (GW-QXGB). Both methods construct localized quantile tree models using spatial kernel weights and enable the simultaneous modeling of spatial heterogeneity, nonlinear relationships, and conditional distributional heterogeneity. In addition, SHapley Additive exPlanations (SHAP) are devoted to quantify localized variable contributions across geographic locations and quantile levels, while bootstrap procedures are employed to assess the stability of the corresponding SHAP values. The proposed methods are evaluated through simulation experiments under various spatial heterogeneity scenarios and further illustrated using four publicly available spatial datasets, with comparisons against conventional QR, GWQR, QRF, and QXGBoost. The results demonstrate that GW-QRF and GW-QXGB generally achieve superior predictive performance and yield weaker spatial autocorrelation in model residuals than competing methods. Furthermore, the SHAP analysis reveals substantial spatial variation in variable contributions across different conditional quantiles. The proposed locally weighted tree-based quantile regression framework not only extends the modeling capability of GWQR but also offers a new analytical approach for spatial machine learning.
摘要 i
Abstract iii
目錄 v
圖目錄 viii
表目錄 xi
第一章 緒論 1
1.1研究動機 1
1.2研究目的 4
1.3研究架構 5
第二章 文獻探討 7
2.1機器學習概述 7
2.1.1監督式機器學習與迴歸問題 8
2.1.2機器學習模型特性與評估 10
2.2樹模型與集成學習 12
2.2.1隨機森林 14
2.2.2極限梯度提升 16
2.3空間資料分析 18
2.4地理加權迴歸 19
2.4.1估計 19
2.4.2帶寬及核函數 20
2.5地理加權機器學習模型 22
2.5.1地理加權隨機森林 23
2.5.2地理加權極限梯度提升 25
2.6機器學習模型的可解釋性–SHAP 26
2.7小結 27
第三章 研究方法 29
3.1分量迴歸 29
3.2地理加權分量迴歸 31
3.3分量隨機森林 33
3.4分量極限梯度提升 34
3.5地理加權分量隨機森林 36
3.6地理加權分量極限梯度提升 37
3.7 GW-QRF、GW-QXGB超參數調校與帶寬選擇 38
3.8集成預測方法 40
3.9機器學習分量預測模型之SHAP解釋與拔靴法檢定 41
第四章 模擬實驗 43
4.1模擬資料 43
4.2實驗設定 44
4.2.1參數設定 44
4.2.2評估指標 46
4.3模擬結果 49
4.3.1各模式比較 49
4.3.2 GW-QXGB變數重要性分析 56
4.4小結 57
第五章 實證資料分析 58
5.1資料來源與介紹 58
5.2調參與分析設定 61
5.3實證結果比較 61
5.4特徵重要性與模型解釋 80
5.4.1全樣本SHAP與空間分布解釋 80
5.4.2 SHAP穩定度 87
第六章 結論 91
6.1總結與討論 91
6.2未來研究方向 93
參考文獻 95
附錄A 100
Adadi, A. and Berrada, M. (2018). Peeking inside the black-box: A survey on explainable artificial intelligence (XAI). IEEE Access, 6:52138–52160.
Akaike, H. (1974). A new look at the statistical model identification. IEEE Transactions on Automatic Control, 19(6):716–723.
Akiba, T., Sano, S., Yanase, T., Ohta, T., and Koyama, M. (2019). Optuna: A next generation hyperparameter optimization framework. In Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, page2623–2631. Association for Computing Machinery.
Amegbor, P. M. and Rosenberg, M. W. (2019). What geography can tell us? effect of higher education on intimate partner violence against women in uganda. Applied Geography, 106:71–81.
Anselin, L. (1988). Spatial Econometrics: Methods and Models, volume 4. Kluwer Academic Publishers, Dordrecht, The Netherlands.
Breiman, L. (1996). Bagging predictors. Machine Learning, 24:123–140.
Breiman, L. (2001). Random forests. Machine Learning, 45:5–32.
Breiman, L., Friedman, J. H., Olshen, R. A., and Stone, C. J. (1984). Classification and Regression Trees. Wadsworth.
Brunsdon, C., Fotheringham, S., and Charlton, M. (2000). Geographically weighted regression as a statistical model. Technical report, Spatial Analysis Research Group, Department of Geography, University of Newcastle-upon-Tyne, Newcastle-upon-Tyne,UK.
Burnham, K. P. and Anderson, D. R. (2004). Multimodel inference: understanding aic and bic in model selection. Sociological methods research, 33(2):261–304.
Chen, C., Wang, J., Li, D., Sun, X., Zhang, J., Yang, C., and Zhang, B. (2024). Unraveling nonlinear effects of environment features on green view index using multiple data sources and explainable machine learning. Scientific Reports, 14:30189.
Chen, T. and Guestrin, C. (2016). XGBoost: A scalable tree boosting system. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, pages 785–794. ACM.
Chen, V. Y.-J., Deng, W.-S., Yang, T.-C., and Matthews, S. A. (2012). Geographically weighted quantile regression (gwqr): An application to u.s. mortality data. Geographical Analysis, 44(2):134–150.
Chen, W., Shen, Y., and Wang, Y. (2018). Does industrial land price lead to industrial diffusion in china? an empirical study from a spatial perspective. Sustainable Cities and Society, 40:307–316.
Cui, P., Abdel-Aty, M., Wang, C., Yang, X., and Song, D. (2025). Examining the impact of spatial inequality in socio-demographic and commute patterns on traffic crash rates:Insights from interpretable machine learning and spatial statistical models. Transport Policy, 167:222–245.
Córdoba, M., Carranza, J. P., Piumetto, M., Monzani, F., and Balzarini, M. (2021). A spatially based quantile regression forest model for mapping rural land values. Journal of Environmental Management, 289:112509.
Davison, A. C. and Hinkley, D. V. (1997). Bootstrap Methods and their Application. Cambridge Series in Statistical and Probabilistic Mathematics. Cambridge University Press.
Freund, Y. and Schapire, R. E. (1997). A decision-theoretic generalization of on-line learning and an application to boosting. Journal of Computer and System Sciences, 55(1):119–139.
Friedman, J. H. (2001). Greedy function approximation: A gradient boosting machine. The Annals of Statistics, 29(5):1189–1232.
Gao, F., He, S., and Kwan, M.-P. (2025). Mixed geographically weighted xgboost (mgwxgb) model: A new spatially explicit machine learning model. Annals of the American Association of Geographers.
Georganos, S., Grippa, T., Gadiaga, A. N., Linard, C., Lennert, M., Vanhuysse, S., Mboga,N., Wolff, E., and Kalogirou, S. (2021). Geographical random forests: a spatial extension of the random forest algorithm to address spatial heterogeneity in remote sensing and population modelling. Geocarto International, 36(2):121–136.
Georganos, S.andKalogirou, S.(2022). A forest of forests:A spatially weighted and computationally efficient formulation of geographical random forests. ISPRS International Journal of Geo-Information, 11(9).
Grekousis, G. (2025). Geographical-xgboost: a new ensemble model for spatially local regression based on gradient-boosted trees. Journal of Geographical Systems, 27(1):169-195.
Hastie, T., Tibshirani, R., and Friedman, J. (2009). The Elements of Statistical Learning:Data Mining, Inference, and Prediction. Springer, 2 edition.
Hwang, C. and Shim, J. (2017). Geographically weighted least squares-support vector machine. Journal of the Korean Data and Information Science Society, 28(1):227–235.
James, G., Witten, D., Hastie, T., and Tibshirani, R. (2021). An Introduction to Statistical Learning: With Applications in R. Springer, 2 edition.
Koenker, R. and Hallock, K. F. (2001). Quantile regression. Journal of Economic Perspectives, 15(4):143–156.111
Koenker, R. and Machado, J. A. F. (1999). Goodness of fit and related inference processes for quantile regression. Journal of the American Statistical Association, 94(448):1296-1310.
Li, Z. (2022). Extracting spatial effects from machine learning model using local interpretation method: An example of shap and xgboost. Computers,Environment and UrbanSystems, 96:101845.
Lundberg, S. M. and Lee, S.-I. (2017). A unified approach to interpreting model predictions. In Proceedings of the 31st International Conference on Neural Information Processing Systems (NIPS 2017), pages 4765–4774, Long Beach, CA, USA.
Maxwell, K., Rajabi, M., and Esterle, J. (2021). Spatial interpolation of coal properties using geographic quantile regression forest. International Journal of Coal Geology,248:103869.
Meinshausen, N. (2006). Quantile regression forests. Journal of Machine Learning Research, 7:983–999.
Moran, P. A. P. (1950). Notes on continuous stochastic phenomena. Biometrika,37(1/2):17–23.
North, B., Curtis, D., and Sham, P. (2002). A note on the calculation of empirical p values from monte carlo procedures. American journal of human genetics, 71:439–41.
Qin, Z., Peng, Q., Jin, C., Xu, J., Xing, S., Zhu, P., and Yang, G. (2025). Geographically weighted random forest fusing multi-source environmental covariates for spatial prediction of soil heavy metals. Environmental pollution (Barking, Essex : 1987),385:127135.
Shapley, L. S. (1953). A value for n-person games. In Kuhn, H. W. and Tucker, A. W.,editors, Contributions to the Theory of Games II, pages 307–317. Princeton University Press, Princeton.112
Sluijterman, L., Kreuwel, F., Cator, E., and Heskes, T. (2024). Composite quantile regression with xgboost using the novel arctan pinball loss. International Journal of Machine Learning and Cybernetics, 16:7575– 7589.
Sun, K., Zhou, R., Kim, J., and Hu, Y. (2024). Pygrf: An improved python geographical random forest model and case studies in public health and natural disasters. Transactions in GIS, 28(7):2476–2491.
Yin, X., Fallah-Shorshani, M., McConnell, R., Fruin, S., Chiang, Y.-Y., and Franklin,M. (2023). Quantile extreme gradient boosting for uncertainty quantification. ArXiv,abs/2304.11732.
Zhou, R. Z., Hu, Y., Tirabassi, J. N., Ma, Y., and Xu, Z. (2022). Deriving neighborhood level diet and physical activity measurements from anonymized mobile phone location data for enhancing obesity estimation. International Journal of Health Geographics,21:22.
Zhou, Y., Sarabi, S., West, T. A., Xie, S., and Han, Q. (2025). Exploring spatial dependency and heterogeneity in forest land dynamics in the randstad metropolitan region:A combined spatial nonlinear modeling approach. Urban Forestry & Urban Greening,112:128976.
盧冠仁、吳元維(2023). 基於多元線性迴歸模型及可解釋機器學習模型之精準執法成效分析. 交通學報,23(2):59–82.
賴品霖(2016). 運用分量迴歸模型分析不動產價格因子-以夜市與捷運站為例.Master’s thesis, 國立臺灣科技大學財務金融研究所. 碩士論文.
全文公開日期 2028/08/10