| 研究生: |
夏永紳 Hsia, Yung-Shen |
|---|---|
| 論文名稱: |
可認證的法律修正:用於代理式人工智慧最佳化合規的神經符號流程 Certifiable Legal Corrections: A Neuro-Symbolic Pipeline for Optimized Compliance in Agentic AI |
| 指導教授: | 郁方 |
| 口試委員: |
江介宏
洪智鐸 |
| 學位類別: |
碩士
Master |
| 系所名稱: |
商學院 - 資訊管理學系 Department of Management Information System |
| 論文出版年: | 2026 |
| 畢業學年度: | 114 |
| 語文別: | 英文 |
| 論文頁數: | 56 |
| 中文關鍵詞: | 國立政治大學 、神經符號合規 、決策支持 、可信賴人工智慧 、優化 、SMT 約束求解 |
| 外文關鍵詞: | NCCU, Neuro-symbolic compliance, Decision support, Trustworthy AI, Optimization, SMT constraint solving |
| 相關次數: | 點閱:58 下載:0 |
| 分享至: |
| 查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報 |
法規合規性修正(Regulatory compliance correction)是一項自動化決策支持問題。在此問題中,分析人員必須判定法律規則是否可執行、記錄的個案是否違反已編碼的法定限制,以及何種補救行動能以最小的干涉度恢復合規。現有基於大型語言模型(LLM)的方法雖能解讀法律文本,但其給出的建議往往缺乏透明度且在邏輯上無法驗證。本研究提出了一個可驗證的自動化決策支持系統,將大語言模型的法律語義詮釋與可滿足性模數理論(SMT)的驗證及優化進行有機整合。該系統將法規與個案敘述轉換為求解器可檢查的表徵、診斷不可行的合規狀態,並在不可變限制下計算出具備成本意識的最小補救計畫。當求解器找到可行的修復方案時,它會返回一個滿足解(satisfying assignment);當限制條件發生衝突時,則會返回一個不可滿足核心(unsatisfiable core)。這些求解器生成的產物共同為可驗證的決策提供了機器可檢查的證據。
本研究針對台灣金融監督管理委員會(金管會)發布的 479 件裁罰案件進行實驗,結果顯示所提系統達到了 99.79% 的形式驗證率。在研究問題二與三(RQ2–RQ3)的基線測試中,獨立運作的 LLM 僅達到 72.0% 的可行性準確率,且在生成合規補救計畫的成功率僅為 15.14%。此外,在相同的階段性檢查點下,將Codex基線視為黑盒子生成器進行評估,其端到端違規檢測準確率達到 88.31%,且在 479個案件中有 423 個案件被正確分類為 UNSAT。這些實驗結果表明,將神經網路的語義詮釋與符號約束求解相結合,能為法規決策支持提供一個具備可擴展性、可解釋性與可審計性的堅實基礎。
Regulatory compliance correction is an automated decision-support problem in which analysts must determine whether legal rules are executable, whether documented cases violate encoded statutory constraints, and which remediation actions can restore compliance with minimal intervention. Existing LLM-based approaches can interpret legal text, but their recommendations are often opaque and logically unverifiable. This study proposes a verifiable, automated decision-support system that integrates large language model (LLM) legal interpretation with satisfiability modulo theories (SMT) verification and optimization. The system converts statutes and case narratives into solver-checkable representations, diagnoses infeasible compliance states, and computes cost-aware minimal remediation plans under immutable constraints. When the solver finds a feasible repair, it returns a satisfying assignment; when constraints conflict, it returns an unsatisfiable core. Together, these solver artifacts provide the machine-checkable evidence required for verifiable decision making. Experiments on 479 enforcement cases issued by Taiwan’s Financial Supervisory Commission show that the proposed system attains a 99.79% formal verification rate. For the RQ2–RQ3 baselines, standalone LLMs achieve only 72.0% feasibility accuracy and 15.14% success in generating compliant remediation plans. In addition, a Codex baseline, treated as a black-box generator and evaluated under the same staged checkpoints, reaches 88.31% end-to-end violation-detection accuracy, with 423 out of 479 cases correctly classified as UNSAT. These results demonstrate that coupling neural semantic interpretation with symbolic constraint solving yields a scalable, explainable, and auditable foundation for regulatory decision support.
致謝 i
摘要 ii
Abstract iii
Contents iv
List of Figures vi
List of Tables vii
1. Introduction 1
2. Related Work 5
2.1 LLMs for Legal and Financial Compliance 5
2.2 SMT Solvers and Symbolic Approaches 5
2.3 Hybrid Neuro-Symbolic Systems and Verification 6
2.4 LLM-Augmented Formal Methods and the Validation–Repair Paradigm 7
3. Methodology 9
3.1 Neuro-Symbolic System Overview 9
3.2 Core Reasoning Tasks for Automated Compliance 11
3.3 Law Interpretation and Legality Consistency Checking 13
3.4 Case Fact Understanding and Illegality Consistency Checking 19
3.5 RAG-Based Legal Context Enrichment 25
4. System Prototype and Deployment 32
4.1 Prototype Workflow Overview 33
4.2 Chatbot Interaction and Session-State Management 34
4.3 Interactive Constraint Refinement and Re-solving 35
4.4 Tool-Augmented Formal Reasoning 37
4.5 Continuous Pipeline and System Orchestration 37
5. Experiments 39
5.1 RQ1: Formal Representation Construction 40
5.2 Codex Baseline 41
5.3 RQ2: Feasibility Judgment for Compliance Decision Support 43
5.4 RQ3: Minimal Remediation Planning 47
6. Conclusions 52
References 53
Arner, Douglas W, Janos Barberis, and Ross P Buckley. 2015. “The Evolution of Fintech: A New Post-Crisis Paradigm.” Geo. J. Int’l L. 47: 1271.
Athul, S, Anshul Saxena, Jayant Mahajan, Lekha Panikulangara, Shruti Kulkarni, and Sanjay Bang. 2024. “LegalMind System and the LLM-Based Legal Judgment Query System.” 2024 International Conference on Trends in Quantum Computing and Emerging Business Technologies, 1–5.
Barrett, Clark, Christopher L Conway, Morgan Deters, et al. 2011. “Cvc4.” Computer Aided Verification: 23rd International Conference, CAV 2011, Snowbird, UT, USA, July 14-20, 2011. Proceedings 23, 171–77.
Barrett, Clark, Thomas A Henzinger, and Sanjit A Seshia. 2026. “Certificates in Ai: Learn but Verify.” Communications of the ACM 69 (1): 66–75.
Berger, Armin, Lars Hillebrand, David Leonhard, et al. 2023. “Towards Automated Regulatory Compliance Verification in Financial Auditing with Large Language Models.” 2023 IEEE International Conference on Big Data (BigData), 4626–35. https://doi.org/10.1109/BigData59044.2023.10386518.
Bolton, Richard J, and David J Hand. 2002. “Statistical Fraud Detection: A Review.” Statistical Science 17 (3): 235–55.
Bradley, Curtis A, and Trevor W Morrison. 2013. “Presidential Power, Historical Practice, and Legal Constraint.” Columbia Law Review, 1097–161.
Chen, Minyu, Guoqiang Li, Ling-I Wu, et al. 2024. “Can Language Models Pretend Solvers? Logic Code Simulation with Llms.” International Symposium on Dependable Software Engineering: Theories, Tools, and Applications, 102–21.
De Moura, Leonardo, and Nikolaj Bjørner. 2008. “Z3: An Efficient SMT Solver.” International Conference on Tools and Algorithms for the Construction and Analysis of Systems, 337–40.
Feng, Nick, Lina Marsso, Mehrdad Sabetzadeh, and Marsha Chechik. 2023. “Early Verification of Legal Compliance via Bounded Satisfiability Checking.” International Conference on Computer Aided Verification, 374–96.
Handler, Abram, Kai R Larsen, and Richard Hackathorn. 2024. “Large Language Models Present New Questions for Decision Support.” International Journal of Information Management 79: 102811.
Hao, Yilun, Yang Zhang, and Chuchu Fan. 2024. “Planning Anything with Rigor: General-Purpose Zero-Shot Planning with Llm-Based Formalized Programming.” arXiv Preprint arXiv:2410.12112.
Hitzler, Pascal, Aaron Eberhart, Monireh Ebrahimi, Md Kamruzzaman Sarker, and Lu Zhou. 2022. “Neuro-Symbolic Approaches in Artificial Intelligence.” National Science Review 9 (6): nwac035.
Ibrahimzada, Ali Reza, Brandon Paulsen, Reyhaneh Jabbarvand, Joey Dodds, and Daniel Kroening. 2025. “MatchFixAgent: Language-Agnostic Autonomous Repository-Level Code Translation Validation and Repair.” arXiv Preprint arXiv:2509.16187.
Jin, Xulei, Lihua Huang, Tan Cheng, Shuaiyong Xiao, Chenghong Zhang, and Yajing Wang. 2025. “Uncertainty-Aware Augmented Generation (UAG): A Novel Deep Learning Method for Enriching in-Conversation User Intent Toward Improved LLM Generation.” Decision Support Systems, 114558.
Judson, Samuel, Matthew Elacqua, Filip Cano, et al. 2024. “Soid: A Tool for Legal Accountability for Automated Decision Making.” International Conference on Computer Aided Verification, 233–46.
Liao, Huanxuan, Shizhu He, Yao Xu, Yuanzhe Zhang, Kang Liu, and Jun Zhao. 2025. “Neural-Symbolic Collaborative Distillation: Advancing Small Language Models for Complex Reasoning Tasks.” Proceedings of the AAAI Conference on Artificial Intelligence 39: 24567–75.
Libal, Tomer, and Tereza Novotná. 2021. “Towards Transparent Legal Formalization.” Explainable and Transparent AI and Multi-Agent Systems: Third International Workshop, EXTRAAMAS 2021, Virtual Event, May 3–7, 2021, Revised Selected Papers 3, 296–313.
Liu, Yizhou, Pengfei Gao, Xinchen Wang, et al. 2024. “Marscode Agent: Ai-Native Automated Bug Fixing.” arXiv Preprint arXiv:2409.00899.
Luo, Zhengxiong, Huan Zhao, Dylan Wolff, Cristian Cadar, and Abhik Roychoudhury. 2026. “Agentic Concolic Execution.” Proceedings of the IEEE Symposium on Security and Privacy (s&p), 1–19.
Maddila, Chandra, Adam Tait, Claire Chang, et al. 2025. “Agentic Program Repair from Test Failures at Scale: A Neuro-Symbolic Approach with Static Analysis and Test Execution Feedback.” arXiv Preprint arXiv:2507.18755.
Mao, Jiayuan, Joshua B Tenenbaum, and Jiajun Wu. 2025. “Neuro-Symbolic Concepts.” arXiv Preprint arXiv:2505.06191.
Mishra, Venkatesh, Bimsara Pathiraja, Mihir Parmar, et al. 2025. “Investigating the Shortcomings of LLMs in Step-by-Step Legal Reasoning.” arXiv Preprint arXiv:2502.05675.
Mohsen, Sara Ebrahim, Allam Hamdan, and Haneen Mohammad Shoaib. 2024. “Digital Transformation and Integration of Artificial Intelligence in Financial Institutions.” Journal of Financial Reporting and Accounting.
Ning, Maizhen, Zihao Zhou, Qiufeng Wang, Xiaowei Huang, and Kaizhu Huang. 2025. “GNS: Solving Plane Geometry Problems by Neural-Symbolic Reasoning with Multi-Modal LLMs.” Proceedings of the AAAI Conference on Artificial Intelligence 39: 24957–65.
Ponkin, Igor, and Alena Redkina. 2019. “Digital Formalization of Law.” International Journal of Open Information Technologies 7 (1): 39–48.
Raza, Mohammad, and Natasa Milic-Frayling. 2025. “Instantiation-Based Formalization of Logical Reasoning Tasks Using Language Models and Logical Solvers.” arXiv Preprint arXiv:2501.16961.
Shi, Jihao, Xiao Ding, Hengwei Zhao, Ting Liu, and Bing Qin. 2025. “Bridging Neural and Symbolic Reasoning: A Dual-System Framework for Interpretable Question Answering.” ICASSP 2025 - 2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 1–5. https://doi.org/10.1109/ICASSP49660.2025.10888935.
Shu, Dong, Haoran Zhao, Xukun Liu, David Demeter, Mengnan Du, and Yongfeng Zhang. 2024. “LawLLM: Law Large Language Model for the US Legal System.” Proceedings of the 33rd ACM International Conference on Information and Knowledge Management, 4882–89.
Strunk, Jobin, Anika Nissen, and Stefan Smolnik. 2025. “All Risks Ain’t the Same–a Risk Facets Perspective on AI-Based Decision Support Systems.” Decision Support Systems, 114557.
Surden, Harry. 2024. “Computable Law and Artificial Intelligence.” U of Colorado Law Legal Studies Research Paper Forthcoming, Cambridge Handbook of Private Law and Artificial Intelligence (Forthcoming 2024).
Tsigkanos, Christos, Alessio Arleo, Johannes Sorger, and Schahram Dustdar. 2019. “How Do Firms Transact? Guesstimation and Validation of Financial Transaction Networks with Satisfiability.” 2019 IEEE 20th International Conference on Information Reuse and Integration for Data Science (IRI), 15–22.
Wan, Yuwei, Zheyuan Chen, Ying Liu, Chong Chen, and Michael Packianather. 2025. “Prompting Large Language Models Based on Semantic Schema for Text-to-Cypher Transformation Towards Domain q&a.” Decision Support Systems, 114553.
Xia, Chunqiu Steven, Yinlin Deng, Soren Dunn, and Lingming Zhang. 2024. “Agentless: Demystifying Llm-Based Software Engineering Agents.” arXiv Preprint arXiv:2407.01489.
Xia, Chunqiu Steven, Zhe Wang, Yan Yang, Yuxiang Wei, and Lingming Zhang. 2025. “Live-SWE-Agent: Can Software Engineering Agents Self-Evolve on the Fly?” arXiv Preprint arXiv:2511.13646.
Ye, Xi, Qiaochu Chen, Isil Dillig, and Greg Durrett. 2023. “Satlm: Satisfiability-Aided Language Models Using Declarative Prompting.” Advances in Neural Information Processing Systems 36: 45548–80.
Zhang, Yedi, Yufan Cai, Xinyue Zuo, et al. 2024. “The Fusion of Large Language Models and Formal Methods for Trustworthy AI Agents: A Roadmap.” arXiv Preprint arXiv:2412.06512.
全文公開日期 2031/07/16