| 研究生: |
管漢程 Kuan, Han-Cheng |
|---|---|
| 論文名稱: |
Hi-Clover : 基於三元組神經網路的 Hi-C 數據分類方法依據重複樣本和不同實驗條件進行分類 Hi-Clover: A Triplet Neural Network for Classifying Hi-C Data Across Replicates and Conditions |
| 指導教授: | 張家銘 |
| 口試委員: |
古倫維
陳世淯 |
| 學位類別: |
碩士
Master |
| 系所名稱: |
資訊學院 - 資訊科學系 Department of Computer Science |
| 論文出版年: | 2026 |
| 畢業學年度: | 115 |
| 語文別: | 中文 |
| 論文頁數: | 43 |
| 中文關鍵詞: | Hi-C 、孿生神經網路 、三元組神經網路 、技術噪聲 、生物變異 、再現性分 析 、深度學習 、度量學習 |
| 外文關鍵詞: | Hi-C, Siamese neural network, Triplet neural network, Technical noise, Biological variation, Reproducibility analysis, Deep learning, Metric learning |
| 相關次數: | 點閱:20 下載:0 |
| 分享至: |
| 查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報 |
Hi-C 技術是研究基因組三維結構的重要方法,能夠產生高解析度的染色質交互作用圖譜,從而揭示細胞中基因調控與結構之間潛在的關聯。然而,Hi-C 數據常受到技術噪聲 (technical noise) 以及不同實驗條件所造成之生物變異的影響。技術噪聲可能降低資料的準確性與再現性 (reproducibility),而生物條件間的差異則會增加跨樣本比較與後續下游分析 (downstream analysis) 的複雜度與不確定性。因此,如何精確且穩健地評估 Hi-C 數據的再現性,成為當前基因組學研究中一項亟需解決的挑戰。針對此問題,先前已有研究嘗試應用深度學習方法,例如:基於孿生神經網路 (Siamese Neural Network, SNN) 的架構,透過比較成對的 Hi-C 接觸矩陣 (contact map) 來學習其潛在相似性結構,進而區分重複樣本與不同實驗條件的資料。
延續此研究方向,本研究提出一套三元組神經網路架構,稱為 Hi-Clover,旨在提升 Hi-C 數據在區分重複樣本與不同實驗條件樣本時的分類能力與再現性評估準確度。該架構參考三元組神經網路 (Triplet Neural Network) 在蛋白質摺疊辨識任務中的成功應用, Hi-Clover 模型以 anchor、positive 與 negative 三元組樣本進行訓練,透過學習其相對距離以優化嵌入空間 (embedding space),進而提升樣本間的可分性。
在模型評估方面,本研究採用孿生形式的資料進行測試,並以平均效能 (mean performance) 與分離指數 (separation index) 作為主要衡量標準,評估模型在辨識重複樣本與區分不同實驗條件樣本時的穩定性與判別能力。實驗結果顯示,Hi-Clover 在 Liver 資料集上的測試集平均效能達 0.9893、分離指數為 0.9786,在 NPC 資料集上的測試集平均效能為 0.9913、分離指數為 0.9874,兩者均高於基準模型 Twins。然而,在反映自然細胞分化過程的 T Cell 資料集上,Hi-Clover 的測試集平均效能為 0.8243、分離指數為 0.6623,仍低於 Twins 的 0.8467 與 0.7010,顯示模型在條件間差異較為細微的資料集上仍面臨挑戰。整體而言,Hi-Clover 在條件差異較明確的人為擾動資料集上展現良好的分類能力,但在自然分化資料集上的表現仍有進一步提升空間。
Hi-C is an important technique for studying the three-dimensional organization of the genome by generating chromatin contact maps. However, Hi-C data are often affected by technical noise, experimental variation, and differences between biological conditions, which may reduce reproducibility and increase uncertainty in downstream analyses. Therefore, developing robust methods for evaluating Hi-C data similarity and reproducibility remains an important task in genomics research.
In this study, we propose Hi-Clover, a triplet neural network-based framework for classifying Hi-C data across replicates and experimental conditions. Hi-Clover is trained using triplets consisting of an anchor sample, a positive sample from the same condition, and a negative sample from a different condition. By learning relative distances among samples, the model aims to map replicate samples closer together while separating samples from different conditions.
For evaluation, Hi-Clover was tested using a Siamese-style framework, in which pairs of Hi-C contact map patches were compared via Euclidean distance in the embedding space. Mean performance and separation index were used as the main evaluation metrics. Hi-Clover achieved a test mean performance of 0.9893 and a separation index of 0.9786 on the Liver dataset, and 0.9913 and 0.9874 on the NPC dataset, yielding higher metric values than Twins in both cases. However, on the T Cell dataset, Hi-Clover achieved a test mean performance of 0.8243 and a separation index of 0.6623, which remained lower than those of Twins.
These results indicate that Hi-Clover is effective for Hi-C datasets with clear experimental perturbations, such as Liver and NPC, but still faces challenges in datasets with more subtle biological differences, such as T Cell. Overall, this study shows that triplet-based metric learning is feasible for comparing Hi-C samples.
第一章 緒論 1
1.1 研究背景與動機 1
1.2 高通量染色體捕獲技術 Hi-C 2
1.3 Hi-C 資料特性對再現性分析的影響 2
1.4 重複樣本與實驗條件 (Replicates and Conditions) 4
1.5 孿生神經網路 (Siamese Neural Network) 4
1.6 三元組神經網路 (Triplet Neural Network) 6
第二章 方法 9
2.1 概覽 9
2.2 資料來源與前處理 10
2.2.1 資料集 10
2.2.2 Hi-C 資料處理與子圖擷取 11
2.2.3 資料切分與樣本數量 12
2.2.4 三元組訓練資料生成 13
2.2.5 孿生測試資料生成 14
2.2.6 資料增強與輸入處理 15
2.3 模型架構 16
2.4 模型訓練與參數設定 18
2.5 評估指標 20
2.5.1 平均效能 (Mean Performance) 20
2.5.2 分離指數 (Separation Index) 21
2.5.3 補充評估指標:AUROC 與 AUPRC 22
2.6 實驗環境設定 22
2.6.1 實驗平台與硬體環境 23
2.6.2 軟體環境與套件版本 23
第三章 結果 25
3.1 訓練過程監控分析 25
3.1.1 訓練輪數與訓練測試時間 26
3.2 模型分類效能與基準比較 27
3.3 嵌入距離分布分析 30
3.4 嵌入向量之 UMAP 降維視覺化分析 31
3.5 模型設定與訓練策略之選擇 32
3.5.1 Margin 之選擇 33
3.5.2 水平翻轉與 semi-hard negative mining 之效果比較 34
3.5.3 semi-hard negative mining 與 joint loss 之效果比較 35
3.5.4 最終模型設定 36
第四章 討論 37
4.1 Hi-Clover 於人為擾動資料集上的表現 37
4.2 T Cell 資料集的效能落差分析 37
4.3 Hi-Clover 與 Twins 的架構差異比較 39
第五章 結論 40
參考文獻 41
1. E. Lieberman-Aiden et al., “Comprehensive mapping of long-range interactions reveals folding principles of the human genome,” Science, vol. 326, no. 5950, pp. 289–293, 2009.
2. S. S. P. Rao et al., “A 3D map of the human genome at kilobase resolution reveals principles of chromatin looping,” Cell, vol. 159, no. 7, pp. 1665–1680, 2014.
3. M. Forcato et al., “Comparison of computational methods for Hi-C data analysis,” Nature Methods, vol. 14, no. 7, pp. 679–685, 2017.
4. T. Yang et al., “HiCRep: assessing the reproducibility of Hi-C data using a stratum adjusted correlation coefficient,” Genome Research, vol. 27, no. 11, pp. 1939–1949, 2017.
5. E. Al-jibury, J. W. D. King, Y. Guo, et al., “A deep learning method for replicate based analysis of chromosome conformation contacts using Siamese neural networks,” Nature Communications, vol. 14, Art. no. 5007, 2023.
6. Y. Liu, K. Han, Y.-H. Zhu, Y. Zhang, L.-C. Shen, J. Song, and D.-J. Yu, “Improving protein fold recognition using triplet network and ensemble deep learning,” Briefings in Bioinformatics, vol. 22, no. 6, Art. no. bbab248, 2021.
7. A. D. Schmitt, M. Hu, and B. Ren, “Genome-wide mapping and analysis of chromosome architecture,” Nature Reviews Molecular Cell Biology, vol. 17, pp. 743 755, 2016.
8. M. Imakaev et al., “Iterative correction of Hi-C data reveals hallmarks of chromosome organization,” Nature Methods, vol. 9, no. 10, pp. 999–1003, 2012.
9. P. A. Knight and D. Ruiz, “A fast algorithm for matrix balancing,” IMA Journal of Numerical Analysis, vol. 33, no. 3, pp. 1029–1047, 2013.
10. G. G. Yardımcı et al., “Measuring the reproducibility and quality of Hi-C data,” Genome Biology, vol. 20, no. 1, p. 57, 2019.
11. J. Bromley, I. Guyon, Y. LeCun, E. Säckinger, and R. Shah, “Signature verification using a ‘Siamese’ time delay neural network,” in Advances in Neural Information Processing Systems, vol. 6, 1993.
12. S. Chopra, R. Hadsell, and Y. LeCun, “Learning a similarity metric discriminatively, with application to face verification,” 2005 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 539–546, 2005.
13. R. Hadsell, S. Chopra, and Y. LeCun, “Dimensionality Reduction by Learning an Invariant Mapping,” 2006 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR), pp. 1735–1742, 2006.
14. E. Hoffer and N. Ailon, “Deep metric learning using triplet network,” in Similarity Based Pattern Recognition, Cham, Switzerland: Springer, 2015, pp. 84–92.
15. F. Schroff, D. Kalenichenko, and J. Philbin, “FaceNet: A unified embedding for face recognition and clustering,” 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 815–823, 2015.
16. Z. Ma, Y. Y. Lu, Y. Wang, R. Lin, Z. Yang, F. Zhang, and Y. Wang, “Metric learning for comparing genomic data with triplet network,” Briefings in Bioinformatics, vol. 23, no. 6, bbac421, 2022.
17. Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner, “Gradient-based learning applied to document recognition,” Proceedings of the IEEE, vol. 86, no. 11, pp. 2278–2324, 1998.
18. S. Ioffe and C. Szegedy, “Batch normalization: Accelerating deep network training by reducing internal covariate shift,” Proceedings of the 32nd International Conference on Machine Learning (ICML), pp. 448–456, 2015.
19. J. Hu, L. Shen, and G. Sun, “Squeeze-and-Excitation Networks,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 7132–7141, 2018.
20. J. Duchi, E. Hazan, and Y. Singer, “Adaptive subgradient methods for online learning and stochastic optimization,” Journal of Machine Learning Research, vol. 12, pp. 2121–2159, 2011.
21. I. Loshchilov and F. Hutter, “Decoupled Weight Decay Regularization,” International Conference on Learning Representations (ICLR), 2019.
22. I. Loshchilov and F. Hutter, “SGDR: Stochastic Gradient Descent with Warm Restarts,” International Conference on Learning Representations (ICLR), 2017.
23. L. McInnes, J. Healy, and J. Melville, “UMAP: Uniform Manifold Approximation and Projection for Dimension Reduction,” arXiv preprint arXiv:1802.03426, 2018.
全文公開日期 2031/08/03