| 研究生: |
苟頡書 Kou, Chieh-Shu |
|---|---|
| 論文名稱: |
原子事實驅動之通俗摘要生成: 具幻覺與遺漏控制之結構化框架 ATLAS (Atomic fact To LAy Summary): A Structured Generation Framework with Hallucination and Omission Control |
| 指導教授: |
林怡伶
Lin, Yi-Ling |
| 口試委員: |
蕭舜文
Hsiao, Shun-Wen 黃瀚萱 Huang, Hen-Hsen |
| 學位類別: |
碩士
Master |
| 系所名稱: |
商學院 - 資訊管理學系 Department of Management Information System |
| 論文出版年: | 2026 |
| 畢業學年度: | 114 |
| 語文別: | 英文 |
| 論文頁數: | 85 |
| 中文關鍵詞: | 大型語言模型 、通俗摘要生成 、幻覺 、資訊遺漏 、原子事實 、醫學文本簡化 |
| 外文關鍵詞: | Large Language Models, Lay Summarization, Hallucination, Omission, Atomic Facts, Biomedical Text Simplification |
| 相關次數: | 點閱:5 下載:0 |
| 分享至: |
| 查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報 |
大型語言模型(Large Language Models, LLMs)能將生物醫學研究文章轉寫為一般讀者可理解的通俗摘要。然而,提升可讀性的過程可能降低內容對來源的忠實度,造成幻覺與遺漏:前者是生成來源未支持的內容,後者是漏掉應保留的重要資訊。既有研究多聚焦於幻覺,且常採用事後檢核,只能評估模型已生成的內容,無法直接發現從未被寫入摘要的資訊。
為同時控制這兩類問題,本研究提出 ATLAS(Atomic Fact To LAy Summary),一個由三個模組組成的結構化生成框架。Module 1 從完整文章抽取候選原子事實,在生成前明確列出重要內容;Module 2 以文章摘要(abstract)篩選候選事實,建立具有來源依據的核心事實集合;Module 3 將每項保留事實轉為「一正三誤加上以上皆非」(1T3F+NOTA)選擇題,檢查生成摘要是否正確表達該事實,並只改寫不正確或遺漏的部分。
本研究使用 BioLaySumm 2025 的 284 篇 PLOS 與 eLife 文章及三個大型語言模型進行評估。相較於直接生成,ATLAS 將 AlignScore 由 0.375–0.503 提升0.833–0.924,SummaC 由 0.485–0.558 提升至 0.807–0.930,同時維持對專家摘要事實的高度覆蓋。消融實驗顯示,Module 1 建立資訊覆蓋,Module 2 在不降低覆蓋率的情況下帶來主要的事實性提升,Module 3 則進行最後的局部修正。使用近期推理模型、其他公開骨幹模型與大型語言模型評分者的補充實驗亦支持相同結論。整體結果顯示,透過在生成前選取核心原子事實,並在生成過程中進行事實驗證,ATLAS 能有效控制幻覺與遺漏,提升生物醫學通俗摘要對來源文章的忠實度。
Large Language Models (LLMs) can simplify biomedical research articles into lay summaries for non-expert readers. However, improving readability can reduce faithfulness to the source through hallucination, the generation of unsupported content, and omission, the loss of important information. Existing work focuses mainly on hallucination and often relies on post-hoc checking, which evaluates only content that has already been generated and therefore cannot directly identify information that was never included.
This study proposes ATLAS (Atomic Fact To LAy Summary), a structured generation framework that controls both failure modes through three modules. Module 1 extracts candidate atomic facts from the full article to make important content explicit before generation. Module 2 filters these facts against the article’s abstract and retains a grounded set of core facts. Module 3 converts each retained fact into a One-True-Three-False plus None-of-the-Above (1T3F+NOTA) question, verifies whether the generated summary expresses it correctly, and rewrites only inaccurate or missing content.
ATLAS is evaluated on 284 BioLaySumm 2025 articles from PLOS and eLife using three LLMs. Compared with direct generation, it raises AlignScore from 0.375–0.503 to 0.833–0.924 and SummaC from 0.485–0.558 to 0.807–0.930, while maintaining high coverage of the facts contained in expert summaries. The ablation study shows that Module 1 establishes coverage, Module 2 produces the main factuality gain without reducing coverage, and Module 3 provides a final targeted refinement. Supplementary experiments with a recent reasoning model, additional public backbone models, and an LLM judge support the same overall conclusion. Overall, the results show that selecting core atomic facts before generation and verifying them throughout the generation process enable ATLAS to control both hallucination and omission, thereby improving the faithfulness of biomedical lay summaries to their source articles.
致謝 i
摘要 ii
Abstract iii
Table of Contents iv
List of Figures vi
List of Tables vii
1 Introduction 1
2 Related Work 3
2.1 Large Language Models in Healthcare and Lay Summarization 3
2.2 Trust Risks in Medical Text Generation: Hallucination and Omission 5
2.3 Limitations of Current Factuality Evaluation and Generation Control 8
2.4 Atomic Facts as the Foundation for Structured Generation 10
2.5 Theoretical Foundations of the ATLAS Framework Modules 11
3 Methodology 14
3.1 ATLAS Framework Overview 14
3.2 Atomic Fact: Definition 15
3.3 Module 1 — Candidate Atomic Fact Extraction 19
3.4 Module 2 — Filtering Atomic Facts Against the Abstract 20
3.5 Module 3 — Verification and Targeted Rewriting 22
3.6 Experimental Setup 26
4 Experiments and Results 31
4.1 Experiment Overview 31
4.2 Pre-Experiment 1 — Reliability of the 1T3F+NOTA Question Format 32
4.3 Pre-Experiment 2 — Error Detection: Multiple-Choice vs. Direct Judgment 35
4.4 Pre-Experiment 3 — Fact Coverage of Module 1 Extraction 37
4.5 Main Experiment 39
4.5.1 Direct Generation vs. ATLAS 39
4.5.2 Recent Reasoning-Model Baseline 43
4.5.3 Ablation Study 45
4.5.4 Supplementary Backbone Validation 48
4.5.5 LLM-as-a-Judge Evaluation 50
4.5.6 Leaderboard Comparison 53
5 Discussion and Conclusion 55
5.1 Discussion 55
5.1.1 Structuring facts before generation 56
5.1.2 Making Omission Visible 56
5.1.3 Readability, Factuality, and Safety 57
5.2 Conclusion 58
5.3 Limitation 60
5.4 Future Work 61
6 References 62
7 Appendix 70
Alden, D. L., Friend, J., & Chun, M. B. J. (2013). Shared decision making and patient decision aids: Knowledge, attitudes, and practices among Hawai'i physicians. Hawai'i Journal of Medicine & Public Health, 72(11), 396–400.
Asgari, E., Montaña-Brown, N., Dubois, M., Khalil, S., Balloch, J., Yeung, J. A., & Pimenta, D. (2025). A framework to assess clinical safety and hallucination rates of LLMs for medical text summarisation. npj Digital Medicine, 8(1), Article 274. https://doi.org/10.1038/s41746-025-01670-7
Aydin, S., Karabacak, M., Vlachos, V., & Margetis, K. (2024). Large language models in patient education: A scoping review of applications in medicine. Frontiers in Medicine, 11, Article 1477898. https://doi.org/10.3389/fmed.2024.1477898
Banerjee, S., & Lavie, A. (2005). METEOR: An automatic metric for MT evaluation with improved correlation with human judgments. In Proceedings of the ACL Workshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation and/or Summarization (pp. 65–72). Association for Computational Linguistics. https://aclanthology.org/W05-0909/
Brådland, H., Goodwin, M., Andersen, P. A., Nossum, A. S., & Gupta, A. (2025). A new HOPE: Domain-agnostic automatic evaluation of text chunking. In Proceedings of the 48th International ACM SIGIR Conference on Research and Development in Information Retrieval (pp. 170–179). Association for Computing Machinery. https://doi.org/10.1145/3726302.3729882
Buhnila, I., Sinha, A., Agarwal, R., Prasad, D. K., & Constant, M. (2026). Evaluating LLM-as-a-judge for medical term simplification. In D. Demner-Fushman, S. Ananiadou, K. Roberts, & J. Tsujii (Eds.), BioNLP 2026 (pp. 687–694). Association for Computational Linguistics. https://doi.org/10.18653/v1/2026.bionlp-1.55
Cohan, A., Dernoncourt, F., Kim, D. S., Bui, T., Kim, S., Chang, W., & Goharian, N. (2018). A discourse-aware attention model for abstractive summarization of long documents. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 2 (Short Papers) (pp. 615–621). Association for Computational Linguistics. https://doi.org/10.18653/v1/N18-2097
Coleman, M., & Liau, T. L. (1975). A computer readability formula designed for machine scoring. Journal of Applied Psychology, 60(2), 283–284. https://doi.org/10.1037/h0076540
Dale, E., & Chall, J. S. (1948). A formula for predicting readability. Educational Research Bulletin, 27(1), 11–20, 28.
Dou, C., Zhang, Y., Chen, Y., Jin, Z., Jiao, W., Zhao, H., Zhao, Y., Tao, Z., & Huang, Y. (2024). Detection, diagnosis, and explanation: A benchmark for Chinese medical hallucination evaluation. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (pp. 4784–4794). ELRA and ICCL. https://aclanthology.org/2024.lrec-main.428/
Feng, S., Shi, W., Wang, Y., Ding, W., Balachandran, V., & Tsvetkov, Y. (2024). Don't hallucinate, abstain: Identifying LLM knowledge gaps via multi-LLM collaboration. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics, Volume 1: Long Papers (pp. 14664–14690). Association for Computational Linguistics. https://doi.org/10.18653/v1/2024.acl-long.786
Filippova, K. (2020). Controlled hallucinations: Learning to generate faithfully from noisy data. In Findings of the Association for Computational Linguistics: EMNLP 2020 (pp. 864–870). Association for Computational Linguistics. https://doi.org/10.18653/v1/2020.findings-emnlp.76
Hosseini, M. J., Gao, Y., Baumgärtner, T., Fabrikant, A., & Amplayo, R. K. (2024). Scalable and domain-general abstractive proposition segmentation. In Findings of the Association for Computational Linguistics: EMNLP 2024 (pp. 8856–8872). Association for Computational Linguistics. https://doi.org/10.18653/v1/2024.findings-emnlp.517
Huang, L., Yu, W., Ma, W., Zhong, W., Feng, Z., Wang, H., Chen, Q., Peng, W., Feng, X., Qin, B., & Liu, T. (2025). A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions. ACM Transactions on Information Systems, 43(2), Article 42. https://doi.org/10.1145/3703155
Kianian, R., Sun, D., Crowell, E. L., & Tsui, E. (2024). The use of large language models to generate education materials about uveitis. Ophthalmology Retina, 8(2), 195–201. https://doi.org/10.1016/j.oret.2023.09.008
Kincaid, J. P., Fishburne, R. P., Jr., Rogers, R. L., & Chissom, B. S. (1975). Derivation of new readability formulas (Automated Readability Index, Fog Count and Flesch Reading Ease Formula) for Navy enlisted personnel (Research Branch Report 8-75). Naval Air Station Memphis, Chief of Naval Technical Training.
Kim, Y., Jeong, H., Chen, S., Li, S. S., Park, C., Lu, M., Alhamoud, K., Mun, J., Grau, C., Jung, M., Gameiro, R., Fan, L., Park, E., Lin, T., Yoon, J., Yoon, W., Sap, M., Tsvetkov, Y., Liang, P., … Breazeal, C. (2025). Medical hallucinations in foundation models and their impact on healthcare. arXiv. https://arxiv.org/abs/2503.05777
Krenn, C., Loder, C., Berger, N., Jeitler, K., Semlitsch, T., Siebenhofer, A., & Wilfling, D. (2026). Automated approaches of text simplification of patient education materials: Scoping review. Journal of Medical Internet Research, 28, Article e88365. https://doi.org/10.2196/88365
Laban, P., Schnabel, T., Bennett, P. N., & Hearst, M. A. (2022). SummaC: Re-visiting NLI-based models for inconsistency detection in summarization. Transactions of the Association for Computational Linguistics, 10, 163–177. https://doi.org/10.1162/tacl_a_00453
Lin, C.-Y. (2004). ROUGE: A package for automatic evaluation of summaries. In Text Summarization Branches Out (pp. 74–81). Association for Computational Linguistics. https://aclanthology.org/W04-1013/
Liu, Y., Iter, D., Xu, Y., Wang, S., Xu, R., & Zhu, C. (2023). G-Eval: NLG evaluation using GPT-4 with better human alignment. In H. Bouamor, J. Pino, & K. Bali (Eds.), Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (pp. 2511–2522). Association for Computational Linguistics. https://doi.org/10.18653/v1/2023.emnlp-main.153
Longpre, S., Perisetla, K., Chen, A., Ramesh, N., DuBois, C., & Singh, S. (2021). Entity-based knowledge conflicts in question answering. In M.-F. Moens, X. Huang, L. Specia, & S. W.-T. Yih (Eds.), Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing (pp. 7052–7063). Association for Computational Linguistics. https://doi.org/10.18653/v1/2021.emnlp-main.565
Maddela, M., Dou, Y., Heineman, D., & Xu, W. (2023). LENS: A learnable evaluation metric for text simplification. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics, Volume 1: Long Papers (pp. 16383–16408). Association for Computational Linguistics. https://doi.org/10.18653/v1/2023.acl-long.905
McMahon, H. V., & McMahon, B. D. (2024). Automating untruths: ChatGPT, self-managed medication abortion, and the threat of misinformation in a post-Roe world. Frontiers in Digital Health, 6, Article 1287186. https://doi.org/10.3389/fdgth.2024.1287186
Mehenni, G., Lamarche, F., Rios-Ibacache, O., Kildea, J., & Zouaq, A. (2025). MedHal: An evaluation dataset for medical hallucination detection. arXiv. https://arxiv.org/abs/2504.08596
Min, S., Krishna, K., Lyu, X., Lewis, M., Yih, W.-T., Koh, P. W., Iyyer, M., Zettlemoyer, L., & Hajishirzi, H. (2023). FActScore: Fine-grained atomic evaluation of factual precision in long form text generation. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (pp. 12076–12100). Association for Computational Linguistics. https://doi.org/10.18653/v1/2023.emnlp-main.741
Neeman, E., Aharoni, R., Honovich, O., Choshen, L., Szpektor, I., & Abend, O. (2023). DisentQA: Disentangling parametric and contextual knowledge with counterfactual question answering. In A. Rogers, J. Boyd-Graber, & N. Okazaki (Eds.), Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics, Volume 1: Long Papers (pp. 10056–10070). Association for Computational Linguistics. https://doi.org/10.18653/v1/2023.acl-long.559
Ni, C., Qadir, S., Steitz, B., Vaidya, M. S., Song, Q., Xia, L., Mulvaney, S., Liu, S., Ryu, H., Hecht, L., Bucher, A., Symons, C., Novak, L., Rose, S. L., Kantarcioglu, M., Malin, B., & Yin, Z. (2026). Disentangling prompt element level risk factors for hallucinations and omissions in mental health LLM responses. arXiv. https://arxiv.org/abs/2604.00014
Oukelmoun, A., Semmar, N., de Chalendar, G., Cormi, C., Oukelmoun, M., Vibert, E., & Allard, M.-A. (2025). Detecting omissions in LLM-generated medical summaries. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing: Industry Track (pp. 325–337). Association for Computational Linguistics. https://doi.org/10.18653/v1/2025.emnlp-industry.22
Pagnoni, A., Balachandran, V., & Tsvetkov, Y. (2021). Understanding factuality in abstractive summarization with FRANK: A benchmark for factuality metrics. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (pp. 4812–4829). Association for Computational Linguistics. https://doi.org/10.18653/v1/2021.naacl-main.383
Pal, A., Umapathi, L. K., & Sankarasubbu, M. (2023). Med-HALT: Medical domain hallucination test for large language models. In Proceedings of the 27th Conference on Computational Natural Language Learning (CoNLL) (pp. 314–334). Association for Computational Linguistics. https://doi.org/10.18653/v1/2023.conll-1.21
Pandit, S., Xu, J., Hong, J., Wang, Z., Chen, T., Xu, K., & Ding, Y. (2025). MedHallu: A comprehensive benchmark for detecting medical hallucinations in large language models. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing (pp. 2858–2873). Association for Computational Linguistics. https://doi.org/10.18653/v1/2025.emnlp-main.143
Papineni, K., Roukos, S., Ward, T., & Zhu, W.-J. (2002). BLEU: A method for automatic evaluation of machine translation. In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics (pp. 311–318). Association for Computational Linguistics. https://doi.org/10.3115/1073083.1073135
Patel, N., Grewal, H., Buddhavarapu, V., & Dhillon, G. (2025). OpenEvidence: Enhancing medical student clinical rotations with AI but with limitations. Cureus, 17(1), Article e76867. https://doi.org/10.7759/cureus.76867
Pradeep, R., Thakur, N., Upadhyay, S., Campos, D., Craswell, N., Soboroff, I., Dang, H. T., & Lin, J. (2025). The great nugget recall: Automating fact extraction and RAG evaluation with large language models. In Proceedings of the 48th International ACM SIGIR Conference on Research and Development in Information Retrieval (pp. 180–190). Association for Computing Machinery. https://doi.org/10.1145/3726302.3730090
Rooney, M. K., Santiago, G., Perni, S., Horowitz, D. P., McCall, A. R., Einstein, A. J., Jagsi, R., & Golden, D. W. (2021). Readability of patient education materials from high-impact medical journals: A 20-year analysis. Journal of Patient Experience, 8, Article 2374373521998847. https://doi.org/10.1177/2374373521998847
Salvi, R. C., Panigrahi, S., Jain, D., Yadav, S., & Akhtar, M. S. (2025). Towards understanding LLM-generated biomedical lay summaries. In S. Ananiadou, D. Demner-Fushman, D. Gupta, & P. Thompson (Eds.), Proceedings of the Second Workshop on Patient-Oriented Language Processing (CL4Health) (pp. 260–268). Association for Computational Linguistics. https://doi.org/10.18653/v1/2025.cl4health-1.22
Sunjaya, A. P., Bao, L., Martin, A., DiTanna, G. L., & Jenkins, C. R. (2022). Systematic review of effectiveness and quality assessment of patient education materials and decision aids for breathlessness. BMC Pulmonary Medicine, 22(1), Article 237. https://doi.org/10.1186/s12890-022-02032-9
Vishwanath, P. R., Tiwari, S., Naik, T. G., Gupta, S., Thai, D. N., Zhao, W., Kwon, S., Ardulov, V., Tarabishy, K., McCallum, A., Salloum, W., & Ai, M. (2024). Faithfulness hallucination detection in healthcare AI. In Proceedings of the KDD 2024 Workshop on Artificial Intelligence and Data Science for Healthcare (KDD-AIDSH 2024). https://openreview.net/forum?id=6eMIzKFOpJ
Vladika, J., Domres, A., Nguyen, M., Moser, R., Nano, J., Busch, F., Adams, L. C., Bressem, K. K., Bernhardt, D., Combs, S. E., Borm, K. J., Matthes, F., & Peeken, J. C. (2025). Improving reliability and explainability of medical question answering through atomic fact checking in retrieval-augmented LLMs. arXiv. https://arxiv.org/abs/2505.24830
Voorhees, E. M. (2004). Overview of the TREC 2003 question answering track. In E. M. Voorhees & L. P. Buckland (Eds.), The Twelfth Text REtrieval Conference (TREC 2003) (NIST Special Publication 500-255, pp. 54–68). National Institute of Standards and Technology. https://trec.nist.gov/pubs/trec12/papers/QA.OVERVIEW.pdf
Xiao, C., Zhao, K., Wang, X., Wu, S., Yan, S., Goldsack, T., Ananiadou, S., Al Moubayed, N., Zhan, L., Cheung, W. K., & Lin, C. (2025). Overview of the BioLaySumm 2025 shared task on lay summarization of biomedical research articles and radiology reports. In Proceedings of the 24th Workshop on Biomedical Language Processing (pp. 365–377). Association for Computational Linguistics. https://doi.org/10.18653/v1/2025.bionlp-1.31
Xie, Q., Luo, Z., Wang, B., & Ananiadou, S. (2023). A survey for biomedical text summarization: From pre-trained to large language models. arXiv. https://arxiv.org/abs/2304.08763
Zaretsky, J., Kim, J. M., Baskharoun, S., Zhao, Y., Austrian, J., Aphinyanaphongs, Y., Gupta, R., Blecker, S. B., & Feldman, J. (2024). Generative artificial intelligence to transform inpatient discharge summaries to patient-friendly language and format. JAMA Network Open, 7(3), Article e240357. https://doi.org/10.1001/jamanetworkopen.2024.0357
Zha, Y., Yang, Y., Li, R., & Hu, Z. (2023). AlignScore: Evaluating factual consistency with a unified alignment function. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics, Volume 1: Long Papers (pp. 11328–11348). Association for Computational Linguistics. https://doi.org/10.18653/v1/2023.acl-long.634
Zhang, T., Kishore, V., Wu, F., Weinberger, K. Q., & Artzi, Y. (2020). BERTScore: Evaluating text generation with BERT. In International Conference on Learning Representations. https://openreview.net/forum?id=SkeHuCVFDr
Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E. P., Zhang, H., Gonzalez, J. E., & Stoica, I. (2023). Judging LLM-as-a-judge with MT-Bench and Chatbot Arena. Advances in Neural Information Processing Systems, 36, 46595–46623. https://proceedings.neurips.cc/paper_files/paper/2023/hash/91f18a1287b398d378ef22505bf41832-Abstract-Datasets_and_Benchmarks.html
全文公開日期 2031/08/10