跳到主要內容

簡易檢索 / 詳目顯示

研究生: 陳育聰
Chen, Yu-Tsung
論文名稱: 基於 KL 散度估計之兩個分布是否相同的檢定
A test for detecting whether two distributions are the same based on KL divergence estimation
指導教授: 黃子銘
Huang, Tzee-Ming Huang
口試委員: 翁久幸
Weng, Chiu-Hsing Weng
鄭宇翔
Cheng, Yu-Hsiang
學位類別: 碩士
Master
系所名稱: 商學院 - 統計學系
Department of Statistics
論文出版年: 2026
畢業學年度: 115
語文別: 中文
論文頁數: 32
中文關鍵詞: 兩樣本檢定KL 散度密度比估計B-spline
外文關鍵詞: two-sample test, KL divergence, density-ratio estimation, B-spline
相關次數: 點閱:73下載:0
分享至:
查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報
  • 本文旨在建立一個用於檢定兩組一維資料是否可視為來自相同分布的兩階段程序。第一階段先比較兩組資料在左右兩端不重疊區域所占的比例,並以單一樣本比例檢定判斷取值範圍是否已有顯著差異;若第一階段未達顯著,第二階段再於共同範圍內進行密度比估計,以兩個方向的 Kullback–Leibler(KL)散度加總形成對稱 KL 統計量,並透過 bootstrap 建立經驗虛無分布與 p 值。
    在估計方法上,本文以三階 B-spline 基底表示密度比函數,使未知的分布差異能以有限維參數形式進行估計。為降低共同區間端點附近資料出現機率過小時對密度比估計造成的不穩定影響,本文亦討論加入虛擬樣本後的替代分析結果。模擬研究包含四組情境:均勻分布與常態分布、兩組相同常態分布、相同截斷區間內不同形狀之常態分布,以及外觀相近但理論上仍不同的 Beta 分布與截斷常態分布。
    模擬結果顯示,當兩組分布差異明顯時,本文方法能穩定拒絕兩組分布相同的假設;在虛無情境下,未加入虛擬樣本時拒絕比例為 19%,加入後降為 4%,接近設定的顯著水準。對於外觀相近且差異較小的情境,加入虛擬樣本後拒絕比例由 77% 降為 9%,表示此處理具有平滑與保守化效果,但也可能降低對細微差異的敏感度。整體而言,本文流程可作為比較兩組一維資料分布差異的工具,但實際解讀時仍須同時考量樣本數、基底設定、重抽樣次數與虛擬樣本處理方式。


    In this thesis, a two-stage procedure is proposed for testing whether two one-dimensional samples can be regarded as being generated from the same distribution. In the first stage, the proportions of observations falling in the non-overlapping regions at the two endpoints are examined using one-sample proportion tests. If no significant endpoint difference is detected, the second stage estimates the density ratio over the common range of the two samples. Two directional Kullback–Leibler (KL) divergence estimates are then combined to form a symmetric KL statistic, and bootstrap resampling is used to construct the empirical null distribution and the corresponding p-values.
    The proposed density-ratio estimator is based on approximating the unknown ratio function with cubic B-spline basis functions, allowing distributional differences to be estimated through a finite-dimensional model. To reduce instability caused by very small occurrence probabilities near the endpoints of the common interval, this thesis also investigates an alternative procedure in which additional virtual samples are included for density estimation. Four simulation settings are considered: a uniform distribution versus a normal distribution, two identical normal distributions, two truncated normal distributions with different shapes over the same interval, and a visually similar but theoretically different pair consisting of a beta distribution and a truncated normal distribution.
    The simulation results show that the proposed procedure consistently rejects the null hypothesis when the distributional difference is pronounced. Under the null setting, the rejection rate decreases from 19% to 4% after adding uniform virtual samples, which is close to the nominal significance level. For the visually similar setting with a small theoretical difference, the rejection rate decreases from 77% to 9% after adding virtual samples, indicating that the virtual-sample procedure has a smoothing and conservative effect but may also reduce sensitivity to subtle differences. Overall, the proposed workflow provides a practical tool for comparing two one-dimensional distributions, while its conclusions should be interpreted together with the sample size, basis specification, number of bootstrap replications, and the use of virtual samples.

    誌謝 i
    摘要 iii
    Abstract iv

    第一章 緒論 1
    1.1 研究背景與動機 1

    第二章 文獻探討 3
    2.1 以散度衡量分配差異 3
    2.2 Kullback–Leibler 散度 4
    2.3 密度估計與密度比估計 5
    2.4 密度比模型與估計 5
    2.5 Spline 與 B-spline 基底函數 7

    第三章 研究方法 8
    3.1 研究架構與問題設定 8
    3.2 第一階段兩端單一樣本比例檢定 9
    3.2.1 最小值區間 9
    3.2.2 最大值區間 10
    3.2.3 統計假設與檢定統計量 10
    3.3 第二階段密度比估計 12
    3.3.1 第二階段分析樣本設定 12
    3.3.2 問題設定 13
    3.3.3 B-spline 基底建構 13
    3.3.4 最佳化目標函數 14
    3.4 對稱 KL 散度 15
    3.5 Bootstrap 推論程序 16
    3.5.1 以第一組資料為母體之 bootstrap 16
    3.5.2 以第二組資料為母體之 bootstrap 17
    3.5.3 判定方式 17

    第四章 資料分析 18
    4.1 分析說明 18
    4.2 模擬情境 18
    4.3 一般分析結果 19
    4.4 加入虛擬樣本後之結果 20
    4.5 KS 檢定結果 21
    4.6 直接積分結果 22
    4.7 本章小結 23

    第五章 結論與建議 24
    參考文獻 28
    附錄 A 附錄 30

    Ali, S. M., & Silvey, S. D. (1966). A general class of coefficients of divergence of one distribution from another. *Journal of the Royal Statistical Society: Series B (Methodological), 28*(1), 131–142. https://doi.org/10.1111/j.2517-6161.1966.tb00626.x

    Basseville, M., & Nikiforov, I. V. (1993). *Detection of abrupt changes: Theory and application*. Prentice Hall.

    Chernoff, H. (1952). A measure of asymptotic efficiency for tests of a hypothesis based on the sum of observations. *The Annals of Mathematical Statistics, 23*(4), 493–507. https://doi.org/10.1214/aoms/1177729330

    Cover, T. M., & Thomas, J. A. (2006). *Elements of information theory* (2nd ed.). Wiley. https://doi.org/10.1002/047174882X

    Cox, M. G. (1972). The numerical evaluation of B-splines. *Journal of the Institute of Mathematics and Its Applications, 10*(2), 134–149. https://doi.org/10.1093/imamat/10.2.134

    de Boor, C. (1972). On calculating with B-splines. *Journal of Approximation Theory, 6*(1), 50–62. https://doi.org/10.1016/0021-9045(72)90080-9

    de Boor, C. (2001). *A practical guide to splines* (Revised ed.). Springer.

    Eilers, P. H. C., & Marx, B. D. (1996). Flexible smoothing with B-splines and penalties. *Statistical Science, 11*(2), 89–121. https://doi.org/10.1214/ss/1038425655

    Kanamori, T., Suzuki, T., & Sugiyama, M. (2012). f-divergence estimation and two-sample homogeneity test under semiparametric density-ratio models. *IEEE Transactions on Information Theory, 58*(2), 708–720. https://doi.org/10.1109/TIT.2011.2163380

    Kullback, S., & Leibler, R. A. (1951). On information and sufficiency. *The Annals of Mathematical Statistics, 22*(1), 79–86. https://doi.org/10.1214/aoms/1177729694

    Liu, S., Yamada, M., Collier, N., & Sugiyama, M. (2013). Change-point detection in time-series data by relative density-ratio estimation. *Neural Networks, 43*, 72–83. https://doi.org/10.1016/j.neunet.2013.01.012

    Møllersen, K., Dhar, S. S., & Godtliebsen, F. (2016). On data-independent properties for density-based dissimilarity measures in hybrid clustering. *Applied Mathematics, 7*(15), 1674–1706. https://doi.org/10.4236/am.2016.715143

    Nguyen, X., Wainwright, M. J., & Jordan, M. I. (2010). Estimating divergence functionals and the likelihood ratio by convex risk minimization. *IEEE Transactions on Information Theory, 56*(11), 5847–5861. https://doi.org/10.1109/TIT.2010.2068870

    Sugiyama, M., Nakajima, S., Kashima, H., von Bünau, P., & Kawanabe, M. (2008). Direct importance estimation with model selection and its application to covariate shift adaptation. *Neural Networks, 21*(10), 1396–1405. https://doi.org/10.1016/j.neunet.2008.09.004

    Sugiyama, M., Suzuki, T., Itoh, Y., Kanamori, T., & Kimura, M. (2011). Least-squares two-sample test. *Neural Networks, 24*(7), 735–751. https://doi.org/10.1016/j.neunet.2011.04.003

    Sugiyama, M., Suzuki, T., & Kanamori, T. (2012). *Density ratio estimation in machine learning*. Cambridge University Press. https://doi.org/10.1017/CBO9781139035613

    無法下載圖示 全文公開日期 2031/07/16
    QR CODE
    :::