| 研究生: |
劉亭妤 Liu, Ting-Yu |
|---|---|
| 論文名稱: |
修復而非重訓:一種約束導引的神經網路修復框架 Repair Instead of Retraining: A Constraint-Guided Framework for Neural Network Repair |
| 指導教授: |
郁方
Yu, Fang |
| 口試委員: |
江介宏
Jiang, Roland 洪智鐸 Hong, Chih Duo |
| 學位類別: |
碩士
Master |
| 系所名稱: |
商學院 - 資訊管理學系 Department of Management Information System |
| 論文出版年: | 2026 |
| 畢業學年度: | 114 |
| 語文別: | 英文 |
| 論文頁數: | 49 |
| 中文關鍵詞: | 神經網路修復 、可解釋人工智慧 、DeepSHAP 、Concolic Execution 、Max-SMT 、Z3 |
| 外文關鍵詞: | neural network repair, explainable AI, DeepSHAP, concolic execution, Max-SMT, Z3 |
| 相關次數: | 點閱:25 下載:0 |
| 分享至: |
| 查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報 |
深度神經網路(DNN)仍易受對抗性擾動與後門攻擊影響,然而事後修復必須在無需完整重訓的前提下,兼顧可解釋性、行為保真度與計算成本。本文提出一個約束導引(但非形式化可證)的修復框架:以 DeepSHAP 定位稀疏的故障相關權重,透過具體執行從良性與對抗性執行中萃取符號路徑約束,並以稀疏權重遮罩上的 Top-K Max-SMT 搜尋權重更新;候選修復一律在原始部署模型上評估,因此所報告的準確率與攻擊成功率反映實際運行行為,而非簡化分析模型。我們不侷限於單一啟發式方法,而是刻畫可配置的修復設計空間——包括修改哪些權重、如何蒐集約束,以及如何排序候選——並在六個基準上進行 4,700 次以上實驗。結果顯示,修復成效關鍵取決於權重方向、約束蒐集所用代理網路,以及簡化強度的聯合校準:以後門基準而言,外向權重修復搭配激活閘控代理網路,可在Fashion-MNIST 與 MNIST-BD 上近乎完全移除後門(攻擊成功率由超過 98% 降至低於 8%);以對抗基準而言,僅調整偏置項搭配激活閘控代理網路,則在MNIST-6(100% 降至 5.81%)與 ResNet18(85.51% 降至 40.42%)上大幅減少對抗性誤行為。CIFAR-10 與 GTSRB 在單層修復下仍具挑戰性,揭示可解性與保真度之間的權衡。上述結果指出,偏置方向修復是對抗性故障修正中一個重要且先前較少探索的維度,唯有透過系統性的設計空間探索方能發現。
Deep neural networks remain vulnerable to adversarial perturbations and backdoor attacks, yet repairing a deployed model without full retraining must balance interpretability, behavioral fidelity, and computational cost. We present a constraint-guided—but not formally certified—framework that uses DeepSHAP to localize a sparse set of fault-relevant weights, collects symbolic path constraints from benign and adversarial executions through concolic testing, and searches for weight updates with Top-K Max-SMT over a sparse repair mask. Candidate repairs are always evaluated on the original deployed model, so reported accuracy and attack success rate reflect actual runtime behavior rather than simplified analysis models. Rather than committing to a single fixed recipe, we characterize a configurable repair design space—including which weights to modify, how constraints are collected, and how candidates are ranked—and study it through more than 4,700 runs on six benchmarks. Repair success depends critically on jointly choosing the weight direction, the surrogate model used for constraint collection, and the simplification strength. For backdoor bench-marks, outgoing-weight repair with an activation-gated surrogate achieves near-complete backdoor removal on Fashion-MNIST and MNIST-BD (attack success rate from above 98% to below 8%). For adversarial benchmarks, bias-only repair with the same activation-gated surrogate substantially reduces misbehavior on MNIST-6 (100% to 5.81%) and ResNet18 (85.51% to 40.42%). CIFAR-10 and GTSRB remain challenging under single-layer repair, highlighting a tractability–fidelity trade-off. These results show that bias-direction repair is a critical and previously underexplored option for adversarial correction, discoverable only through systematic design-space exploration.
致謝 i
摘要 iii
Abstract v
Contents vii
List of Figures x
List of Tables xi
1 Introduction 1
2 Related Work 5
2.1 Neural Network Repair 5
2.2 XAI-Guided Fault Localization and Adversarial Detection 7
2.3 Constraint-Based Neural Network Repair 7
2.4 Provable and Verification-Guided Repair 8
2.5 Program Repair and Symbolic Constraint Extraction 9
2.6 Position of Our Work 10
3 Methodology 13
3.1 Fault Localization 14
3.2 Path-Constraint Specification 17
3.3 Rigid Simplification 20
3.4 Ternary Network Formulations 21
3.5 Evaluation 24
4 Experiments 27
4.1 Experimental Setup 27
4.2 Overall Repair Performance 31
4.3 Effect of Repair Mask Direction and Selection 32
4.4 Effect of Ternary and Gated Surrogates 34
4.5 Effect of Abstraction Strength 35
4.6 Comparison with CARE 36
5 Discussion 39
5.1 Constraint Count as the Primary Cost Driver 39
5.2 Design-Space Implications 39
5.3 Threats to Validity 40
5.4 Path-Level vs. Point-Wise Repair 41
5.5 Limitations and Future Directions 41
6 Conclusion 43
Bibliography 45
[AKV+15] A. Angelova, A. Krizhevsky, V. Vanhoucke, A.S. Ogale,and D. Ferguson, “Real-time pedestrian detection with deep network cascades,” in Proceedings of the British Machine Vision Conference (BMVC),2015(cit.p. 1).
[AM18] N. Akhtar and A. Mian, “Threat of adversarial attacks on deep learning in computer vision: A survey,”, IEEE access : practical innovations, open solutions, vol.6, pp.14410–14430,2018(cit.p. 1).
[AOS+16] D. Amodei, C. Olah, J. Steinhardt, et al.,“ConcreteproblemsinAIsafety,” 2016. arXiv: 1606.06565(cit.p. 1).
[BCB15] D. Bahdanau, K. Cho,and Y. Bengio,“Neuralmachinetranslationbyjointly learningtoalignandtranslate,”, in Proceedings of the International Conference on Learning Representations (ICLR),2015(cit.p. 1).
[BF23] N. Bjørner and K. Fazekas, “On incremental pre-processing for SMT,” in Automated Deduction – CADE 29, Cham: Springer, 2023, pp. 41–60 (cit. pp.20,21).
[CMY+25] Z. Chi, J. Ma, P. Yang, et al. “PatchSynthesisforPropertyRepairofDeep Neural Networks.” arXiv:2404.01642 [cs]. [Online]. Available: http://arxiv.org/abs/2404.01642pre-published(cit.p. 8).
[CTW+21] Y.-F. Chen, W.-L. Tsai, W.-C. Wu, D.-D. Yen,and F. Yu,“PyCT:APython concolictester,”, in Programming Languages and Systems, H. Oh, Ed.,Cham: SpringerInternationalPublishing,2021, pp.38–46(cit.pp. 10,17).
[CZS+24] Z. Chen, J. Zhou, Y. Sun, et al.,“InterpretabilityBasedNeuralNetworkRepair,”, in Proceedings of the 33rd ACM SIGSOFT International Symposium on Software Testing and Analysis, ser. ISSTA 2024, New York, NY, USA: AssociationforComputingMachinery,2024, pp.908–919(cit.pp. 2,6,10).
[FL22] F. Fu and W. Li. “REASSURE: Sound and Complete Neural Network Repair with Minimality and Locality Guarantees.” arXiv:2110.07682 [cs]. [Online].Available: http://arxiv.org/abs/2110.07682 pre-published (cit.p. 8).
[GLD+19] T. Gu, K. Liu, B. Dolan-Gavitt, and S. Garg, “Badnets: Evaluating backdooring attacks on deep neural networks,”, Ieee Access, vol. 7, pp. 47230– 47244,2019(cit.pp. 1,28).
[GMH13] A. Graves, A.-r. Mohamed, and G. Hinton, “Speech recognition with deep recurrentneuralnetworks,”, in 2013 IEEE International Conference on Acoustics, Speech and Signal Processing, IEEE,2013, pp.6645–6649(cit.p. 1).
[GPC+16] V. Gulshan, L. Peng, M. Coram,et al., “Development and validation of a deeplearningalgorithmfordetectionofdiabeticretinopathyinretinalfundusphotographs,” JAMA : the journal of the American Medical Association, vol.316, no.22, pp.2402–2410,2016(cit.p. 1).
[GSS14] I. J. Goodfellow, J. Shlens, and C. Szegedy, “Explaining and harnessing adversarialexamples,”2014. arXiv: 1412.6572(cit.p. 1).
[LDD+24] F. Liu, X. Du, H. Ding, and J. Qian, “Towards robust neural networks: Exploring counterfactual causality-based repair,”, Expert Systems With Applications, vol.257, p.125082,2024(cit.pp. 6,10).
[LL17] S.M. Lundbergand S.-I. Lee,“Aunifiedapproachtointerpretingmodelpredictions,” in Advances in Neural Information Processing Systems, vol. 30, 2017, pp.4765–4774(cit.p. 7).
[LLW+22] F. Li, B. Liu, X. Wang, B. Zhang, and J. Yan. “Ternary Weight Networks.” arXiv: 1605.04711 [cs].[Online].Available: http://arxiv.org/abs/ 1605.04711pre-published(cit.p. 22).
[LNF+12] C. LeGoues, T. Nguyen, S. Forrest,and W. Weimer,“GenProg:Ageneric methodforautomaticsoftwarerepair,” IEEE Transactions on Software Engineering, vol.38, no.1, pp.54–72,2012(cit.p. 9).
[LXS+24] J. Liu, Y. Xing, X. Shi, et al.,“Abstractionandrefinement:Towardsscalable and exact verification of neural networks,”, ACM Transactions on Software Engineering and Methodology, vol. 33, no. 5, pp. 1–35, 2024 (cit. pp.24, 42).
[LY23] Y.-C. Lin and F. Yu, “DeepSHAP summary for adversarial example detection,” in 2023 IEEE/ACM International Workshop on Deep Learning for Testing and Testing for Deep Learning (DeepTest), IEEE, 2023, pp. 17–24 (cit.pp. 2,7,15).
[MKM+30] N. Mellempudi, A. Kundu, D. Mudigere,et al. “Ternary Neural Networks withFine-GrainedQuantization.”, arXiv: 1705.01462 [cs].[Online].Available: http://arxiv.org/abs/1705.01462pre-published(cit.p. 22).
[MWX+25a] J. Ma, J. Wang, Q. Xuan,and Z. Wang,“ProRepair:ProvableRepairofDeep NeuralNetworkDefectsbyPreimageSynthesisand PropertyRefinement,” in Proceedings of the 2025 ACM SIGSAC Conference on Computer and Communications Security, ser. CCS’25,NewYork, NY,USA:Association forComputingMachinery,2025, pp.4169–4183(cit.p. 8).
[MWX+25b] J. Ma, J. Wang, Q. Xuan,and Z. Wang,“ProvableFairnessRepairforDeep Neural Networks,” in 2025 40th IEEE/ACM International Conference on Automated Software Engineering (ASE),2025, pp.508–520(cit.p. 1).
[MYR15] S. Mechtaev, J. Yi, and A. Roychoudhury, “DirectFix: Looking for simple programrepairs,”, in Proceedings of the 37th IEEE/ACM International Conference on Software Engineering, IEEE,2015, pp.448–458(cit.p. 9).
[MYR16] S. Mechtaev, J. Yi,and A. Roychoudhury,“Angelix:Scalablemultilineprogrampatchsynthesisviasymbolicanalysis,”, in Proceedings of the 38th International Conference on Software Engineering, ACM,2016, pp.691–701 (cit.p. 9).
[MYW+24] J. Ma, P. Yang, J. Wang, et al.,“VeRe:VerificationGuidedSynthesisforRepairingDeepNeuralNetworks,”, in Proceedings of the IEEE/ACM 46th International Conference on Software Engineering, ser. ICSE’24,NewYork, NY,USA:AssociationforComputingMachinery,2024, pp.1–13(cit.pp. 2, 8,10,42).
[MZA+21] K. Majd, S. Zhou, H.B. Amor, G. Fainekos,and S. Sankaranarayanan.“Local Repair of Neural Networks Using Optimization.” arXiv:2109.14041 [cs]. [Online]. Available: http : / / arxiv. org / abs / 2109. 14041prepublished(cit.pp. 8,42).
[NQR+13] H.D.T. Nguyen, D. Qi, A. Roychoudhury,and S. Chandra,“SemFix:Programrepairviasemanticanalysis,”, in Proceedings of the 35th International Conference on Software Engineering, IEEE,2013, pp.772–781(cit.p. 9).
[NSM15] P. Nightingale, P. Spracklen,and I. Miguel,“AutomaticallyimprovingSAT encoding of constraint problems through common subexpression eliminationinSavileRow,”, in Principles and Practice of Constraint Programming, Cham:Springer,2015, pp.330–340(cit.p. 20).
[SGM19] E. Strubell, A. Ganesh, and A. McCallum, “Energy and policy considerationsfordeeplearninginNLP,”, in Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics (ACL), 2019, pp. 3645– 3650(cit.p. 1).
[SHK+19] Y. Sun, X. Huang, D. Kroening, et al.,“DeepConcolic:Testingand DebuggingDeepNeuralNetworks,”, in 2019 IEEE/ACM 41st International Conference on Software Engineering: Companion Proceedings (ICSE-Companion), 2019, pp.111–114(cit.p. 24).
[SLW+25] X. Sun, W. Liu, S. Wang,et al., “AutoRIC: Automated Neural Network Repairing Based on Constrained Optimization,”, ACM Trans. Softw. Eng. Methodol., vol.34, no.2,50:1–50:29,2025(cit.pp. 1,8).
[SSP+22] B. Sun, J. Sun, H. L. Pham, and J. Shi. “CARE: Causality-based Neural Network Repair.” arXiv:2204. 09274 [cs]. [Online]. Available: http : //arxiv.org/abs/2204.09274pre-published(cit.pp. 1,6,10,28,36).
[ST21] M. Sotoudeh and A. V. Thakur, “PRDNN: Provable repair of deep neural networks,” in Proceedings of the 42nd ACM SIGPLAN International Conference on Programming Language Design and Implementation, ser. PLDI 2021, New York, NY, USA: Association for Computing Machinery, 2021, pp.588–603(cit.pp. 8,10,41).
[SWR+18] Y. Sun, M. Wu, W. Ruan, et al.,“Concolictestingfordeepneuralnetworks,” in Proceedings of the 33rd ACM/IEEE International Conference on Automated Software Engineering,2018, pp.109–119(cit.p. 18).
[SZS+13] C. Szegedy, W. Zaremba, I. Sutskever, et al.,“Intriguingpropertiesofneural networks,”2013. arXiv: 1312.6199(cit.p. 1).
[TNM+23] Z. Tao, S. Nawas, J. Mitchell, and A. V. Thakur, “APRNN: ArchitecturePreservingProvableRepairofDeepNeuralNetworks,” Reproduction Package for the PLDI 2023 Article ”, Architecture-Preserving Provable Repair of Deep Neural Networks”, vol.7,124:443–124:467,PLDI2023(cit.pp. 1,8, 10).
[UGS+21] M. Usman, D. Gopinath, Y. Sun, Y. Noller, and C. Pasareanu. “NNrepair: Constraint-basedRepairofNeuralNetworkClassifiers.”, arXiv: 2103.12535 [cs]. [Online]. Available: http : / / arxiv. org / abs / 2103. 12535prepublished(cit.pp. 1,7,41).
[VJ25] F. Vares and B. Johnson, “Causality-driven neural network repair: Challenges and opportunities,” in Proceedings of the 33rd ACM International Conference on the Foundations of Software Engineering, Trondheim, Norway:ACM,2025, pp.1406–1409(cit.p. 6).
[YBZ+26] A. Yang, P. Bergsträßer, G. Zetzsche, D. Chiang, and A. W. Lin, “Length generalization bounds for transformers,”, arXiv preprint arXiv:2603.02238, 2026(cit.pp. 1,5).
[YD25] T. Yuviler and D. Drachsler-Cohen, “Enhancing neural network robustness via synthesis of repair programs,” in Static Analysis: 32nd International Symposium (SAS 2025), Berlin, Heidelberg: Springer, 2025, pp. 221–248 (cit.p. 9).
全文公開日期 2031/07/20