| 研究生: |
徐慧妤 Hsu, Hui-Yu |
|---|---|
| 論文名稱: |
情境喜劇多模態幽默解釋生成:基於外部 Chain-of-Thought 增強的方法 Multimodal Humor Explanation Generation in Sitcoms: An External Chain-of-Thought Augmentation Approach |
| 指導教授: |
向倩儀
Hsiang, Chine-Yi |
| 口試委員: |
吳浩庠
Wu, Hao-Hsiang 楊錦生 Yang, Chin-Sheng |
| 學位類別: |
碩士
Master |
| 系所名稱: |
商學院 - 資訊管理學系 Department of Management Information System |
| 論文出版年: | 2026 |
| 畢業學年度: | 114 |
| 語文別: | 英文 |
| 論文頁數: | 70 |
| 中文關鍵詞: | 多模態幽默解釋生成 、外部 Chain-of-Thought 增強 、設計科學研究 、情境喜劇 、多模態大型語言模型 |
| 外文關鍵詞: | Multimodal Humor Explanation Generation, External Chain-of-Thought Augmentation, Design Science Research, Sitcom, Multimodal Large Language Models |
| 相關次數: | 點閱:19 下載:0 |
| 分享至: |
| 查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報 |
本研究探討在多模態情境喜劇場景中,模型如何理解幽默產生之成因。幽默的形成涉及語境預期與後續語意發展之不一致,並於組織溝通與社交互動情境中具有重要應用價值。然而,現有 Multimodal Large Language Models(MLLMs) 在此類語用理解任務上仍與人類表現存在顯著差距,且多數方法之推理過程不可觀察,亦缺乏與幽默認知理論之系統性對應。針對此問題,本研究提出之核心研究問題為:如何在容量受限之 MLLMs 中,設計並操作化一套以理論為基礎之推理結構,以提升其多模態幽默解釋生成能力。本研究依循 Design Science Research 取徑,以 Suls (1972) Incongruity-Resolution 模型為理論基礎,將幽默形成之歷程拆解為 Expectation Formation、Incongruity Detection、Resolution 與 Explanation Generation 四個順序性推理階段,並採用 External CoT Augmentation 架構,由 Qwen2.5-7B 生成前三階段推理描述,再以 prefix injection 方式提供予 VideoLLaMA3-2B 以進行幽默解釋生成。實驗以 SMILE 資料集之 sitcom 子集進行評估,結果顯示,人工評估中本研究方法被選為最符合人類撰寫風格之比例達 52.5%,明顯高於基準模型的 5.0%,甚至高於人工標註答案的 42.5%,顯示模型輸出具備高度語用一致性。此外,自動評估指標與消融實驗結果亦進一步支持本方法之有效性。本研究之貢獻包括一套理論對齊之 Design Artifact、三項可遷移之 Design Principles,以及明確界定之 Boundary Conditions,為資訊系統領域之語用理解相關研究提供理論基礎與實作參考。
This study investigates how models understand humor generation in multimodal situational comedy scenes. Humor comprehension involves incongruity between contextual expectations and subsequent semantic development and is important in organizational communication and social interaction. However, existing Multimodal Large Language Models (MLLMs) still exhibit a significant gap compared to human performance in pragmatic understanding, and their reasoning processes are often unobservable and weakly aligned with theories of humor cognition. To address this, we propose a theory-driven reasoning framework for capacity-constrained MLLMs to improve multimodal humor explanation generation. Following a Design Science Research (DSR) approach and based on Suls (1972) Incongruity-Resolution theory, we decompose humor understanding into four sequential stages: Expectation Formation, Incongruity Detection, Resolution, and Explanation Generation. We introduce an External Chain-of-Thought (CoT) Augmentation framework, where Qwen2.5-7B generates the reasoning descriptions for the first three stages, which are injected via prefix injection into VideoLLaMA3-2B for final explanation generation. Experiments on the SMILE Sitcom dataset show that the proposed method is selected as the most human-like output in 52.5% of cases, outperforming the baseline (5.0%) and even the Ground Truth (42.5%), indicating strong pragmatic consistency. Automatic metrics and ablation studies further confirm the effectiveness of the proposed approach. This work contributes a theory-aligned design artifact, transferable design principles, and clearly defined boundary conditions, offering practical guidance for pragmatic understanding in Information Systems research.
1 Introduction 1
2 Background and Related Work 5
2.1 Humor Theories and Cognitive Mechanisms 5
2.2 Computational Approaches to Humor Tasks 6
2.3 Multimodal Large Language Models 8
2.4 Chain-of-Thought Reasoning and Rationale Distillation 10
2.5 Design Science Research 11
2.6 Research Gap 12
3 Design Artifact 15
3.1 Design Objectives 15
3.2 Design Rationale 16
3.3 Artifact Architecture 17
3.4 Component Specification 19
3.5 Implementation Mechanisms 22
4 Evaluation Framework 29
4.1 Validity Framework 29
4.2 Criterion Validity Methods 29
4.3 Causal Validity Methods 31
4.4 Context Validity Methods 33
5 Experimental Results 34
5.1 Experimental Setup 34
5.2 Results of Criterion Validity 37
5.3 Results of Causal Validity 40
5.4 Results of Context Validity 44
5.5 Case Study 45
5.6 Self-CoT Exploration as Boundary Condition 47
5.7 Cross-Domain Additional Analysis 48
6 Discussion and Conclusion 49
6.1 Design Principles 49
6.2 Boundary Conditions 52
6.3 Theoretical Implications 54
6.4 Practical Implications 55
6.5 Limitations 56
6.6 Ethical Considerations 58
6.7 Future Work 59
6.8 Conclusion 60
References 62
Appendix A: Prompt Design for the Four-Stage CoT Reasoning Framework 67
Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., Avila, R., Babuschkin, I., Balaji, S., Balcom, V., Baltescu, P., Bao, H., Bavarian, M., Belgum, J., Bello, I., … Zoph, B. (2023). GPT-4 technical report. arXiv. https://doi.org/10.48550/arXiv.2303.08774
Anil, R., Borgeaud, S., Wu, Y., Alayrac, J.-B., Yu, J., Soricut, R., Schalkwyk, J., Dai, A. M., Hauth, A., Millican, K., Silver, D., Johnson, M., Antonoglou, I., Schrittwieser, J., Glaese, A., Chen, J., Pitler, E., Lillicrap, T., Lazaridou, A., … Vinyals, O. (2023). Gemini: A family of highly capable multimodal models. arXiv. https://doi.org/10.48550/arXiv.2312.11805
Attardo, S., Pickering, L., & Baker, A. (2011). Prosodic and multimodal markers of humor in conversation. Pragmatics & Cognition, 19(2), 224–247. https://doi.org/10.1075/pc.19.2.03att
Bai, S., Chen, K., Liu, X., Wang, J., Ge, W., Song, S., Dang, K., Wang, P., Wang, S., Tang, J., Zhong, H., Zhu, Y., Yang, M., Li, Z., Wan, J., Wang, P., Ding, W., Fu, Z., Xu, Y., … Lin, J. (2025). Qwen2.5-VL technical report. arXiv. https://doi.org/10.48550/arXiv.2502.13923
Banerjee, S., & Lavie, A. (2005). METEOR: An automatic metric for MT evaluation with improved correlation with human judgments. In J. Goldstein, A. Lavie, C.-Y. Lin, & C. Voss (Eds.), In Proceedings of the ACL workshop on intrinsic and extrinsic evaluation measures for machine translation and/or summarization (pp. 65–72). Association for Computational Linguistics.
Berger, A. A. (2017). An anatomy of humor. Routledge. (Original work published 1993)
Bitterly, T. B., Brooks, A. W., & Schweitzer, M. E. (2017). Risky business: When humor increases and decreases status. Journal of Personality and Social Psychology, 112(3), 431–455. https://doi.org/10.1037/pspi0000079
Burgoon, J. K. (1993). Interpersonal expectations, expectancy violations, and emotional communication. Journal of Language and Social Psychology, 12(1–2), 30–48. https://doi.org/10.1177/0261927X93121003
Caffagni, D., Cocchi, F., Barsellotti, L., Moratelli, N., Sarto, S., Baraldi, L., Baraldi, L., Cornia, M., & Cucchiara, R. (2024). The revolution of multimodal large language models: A survey. In Findings of the Association for Computational Linguistics: ACL 2024 (pp. 13590–13618). Association for Computational Linguistics. https://doi.org/10.18653/v1/2024.findings-acl.807
Chandrasekaran, A., Vijayakumar, A. K., Antol, S., Bansal, M., Batra, D., Zitnick, C. L., & Parikh, D. (2016). We are humor beings: Understanding and predicting visual humor. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 4603–4612). IEEE. https://doi.org/10.1109/CVPR.2016.498
Chen, Y., Yuan, Y., Liu, P., Liu, D., Guan, Q., Guo, M., Peng, H., Liu, B., Li, Z., & Xiao, Y. (2024). Talk funny! A large-scale humor response dataset with chain-of-humor interpretation. In Proceedings of the AAAI Conference on Artificial Intelligence, 38(16), 17826–17834. https://doi.org/10.1609/aaai.v38i16.29736
Choube, A., & Soleymani, M. (2020). Punchline detection using context-aware hierarchical multimodal fusion. In Proceedings of the 2020 International Conference on Multimodal Interaction (ICMI '20) (pp. 675–679). Association for Computing Machinery. https://doi.org/10.1145/3382507.3418891
Hasan, M. K., Rahman, W., Bagher Zadeh, A., Zhong, J., Tanveer, M. I., Morency, L.-P., & Hoque, M. E. (2019). UR-FUNNY: A multimodal language dataset for understanding humor. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP) (pp. 2046–2056). Association for Computational Linguistics. https://doi.org/10.18653/v1/D19-1211
Hessel, J., Marasović, A., Hwang, J. D., Lee, L., Da, J., Zellers, R., Mankoff, R., & Choi, Y. (2023). Do androids laugh at electric sheep? Humor "understanding" benchmarks from the New Yorker Caption Contest. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) (pp. 688–714). Association for Computational Linguistics. https://doi.org/10.18653/v1/2023.acl-long.41
Hossain, N., Krumm, J., Gamon, M., & Kautz, H. (2020). SemEval-2020 Task 7: Assessing humor in edited news headlines. In Proceedings of the Fourteenth Workshop on Semantic Evaluation (pp. 746–758). International Committee for Computational Linguistics. https://doi.org/10.18653/v1/2020.semeval-1.98
Hsieh, C.-Y., Li, C.-L., Yeh, C.-K., Nakhost, H., Fujii, Y., Ratner, A., Krishna, R., Lee, C.-Y., & Pfister, T. (2023). Distilling step-by-step! Outperforming larger language models with less training data and smaller model sizes. In Findings of the Association for Computational Linguistics: ACL 2023 (pp. 8003–8017). Association for Computational Linguistics. https://doi.org/10.18653/v1/2023.findings-acl.507
Hwang, E., West, P., & Shwartz, V. (2025). BottleHumor: Self-informed humor explanation using the information bottleneck principle. In Findings of the Association for Computational Linguistics: ACL 2025. Association for Computational Linguistics. https://doi.org/10.18653/v1/2025.findings-acl.1163
Hyun, L., Kim, S.-B., Han, S., Yu, Y., & Oh, T.-H. (2024). SMILE: Multimodal dataset for understanding laughter in video with language models. In Findings of the Association for Computational Linguistics: NAACL 2024 (pp. 1149–1167). Association for Computational Linguistics. https://doi.org/10.18653/v1/2024.findings-naacl.73
Jiang, T., Li, H., & Hou, Y. (2019). Cultural differences in humor perception, usage, and implications. Frontiers in Psychology, 10, Article 123. https://doi.org/10.3389/fpsyg.2019.00123
Larsen, K. R., Lukyanenko, R., Mueller, R. M., Storey, V. C., Parsons, J., VanderMeer, D., & Hovorka, D. S. (2025). Validity in design science. MIS Quarterly, 49(4), 1267–1294. https://doi.org/10.25300/MISQ/2024/18064
Lin, C.-Y. (2004). ROUGE: A package for automatic evaluation of summaries. In Text summarization branches out (pp. 74–81). Association for Computational Linguistics.
Magister, L. C., Mallinson, J., Adamek, J., Malmi, E., & Severyn, A. (2023). Teaching small language models to reason. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers) (pp. 1773–1781). Association for Computational Linguistics. https://doi.org/10.18653/v1/2023.acl-short.151
Ouyang, K., Liu, Y., Li, S., Liu, Y., Zhou, H., Meng, F., Zhou, J., & Sun, X. (2025). PunchBench: Benchmarking MLLMs in multimodal punchline comprehension. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) (pp. 986–1008). Association for Computational Linguistics. https://doi.org/10.18653/v1/2025.acl-long.49
Papineni, K., Roukos, S., Ward, T., & Zhu, W.-J. (2002). BLEU: A method for automatic evaluation of machine translation. In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics (pp. 311–318). Association for Computational Linguistics. https://doi.org/10.3115/1073083.1073135
Peffers, K., Tuunanen, T., Rothenberger, M. A., & Chatterjee, S. (2007). A design science research methodology for information systems research. Journal of Management Information Systems, 24(3), 45–77. https://doi.org/10.2753/MIS0742-1222240302
Potash, P., Romanov, A., & Rumshisky, A. (2017). SemEval-2017 Task 6: #HashtagWars: Learning a sense of humor. In Proceedings of the 11th International Workshop on Semantic Evaluation (SemEval-2017) (pp. 49–57). Association for Computational Linguistics. https://doi.org/10.18653/v1/S17-2004
Provine, R. R. (2000). Laughter: A scientific investigation. Viking.
Sennrich, R., Haddow, B., & Birch, A. (2016). Improving neural machine translation models with monolingual data. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) (pp. 86–96). Association for Computational Linguistics. https://doi.org/10.18653/v1/P16-1009
Suls, J. M. (1972). A two-stage model for the appreciation of jokes and cartoons: An information-processing analysis. In J. H. Goldstein & P. E. McGhee (Eds.), The psychology of humor: Theoretical perspectives and empirical issues (pp. 81–100). Academic Press.
Tuunanen, T., Winter, R., & vom Brocke, J. (2024). Dealing with complexity in design science research: A methodology using design echelons. MIS Quarterly, 48(2), 427–458. https://doi.org/10.25300/MISQ/2023/16700
Wang, X., Wei, J., Schuurmans, D., Le, Q., Chi, E., Narang, S., Chowdhery, A., & Zhou, D. (2022). Self-consistency improves chain of thought reasoning in language models. arXiv. https://doi.org/10.48550/arXiv.2203.11171
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Xia, F., Chi, E., Le, Q. V., & Zhou, D. (2022). Chain-of-thought prompting elicits reasoning in large language models. Advances in Neural Information Processing Systems, 35, 24824–24837. https://doi.org/10.48550/arXiv.2201.11903
Wei, J., & Zou, K. (2019). EDA: Easy data augmentation techniques for boosting performance on text classification tasks. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP) (pp. 6382–6388). Association for Computational Linguistics. https://doi.org/10.18653/v1/D19-1670
Yang, A., Yang, B., Zhang, B., Hui, B., Zheng, B., Yu, B., Li, C., Liu, D., Huang, F., Wei, H., Lin, H., Yang, J., Tu, J., Zhang, J., Yang, J., Yang, J., Zhou, J., Lin, J., Dang, K., … Qiu, Z. (2024). Qwen2.5 technical report. arXiv. https://doi.org/10.48550/arXiv.2412.15115
Yao, B., Zhang, Y., Li, Q., & Qin, J. (2025). Is sarcasm detection a step-by-step reasoning process in large language models? In Proceedings of the AAAI Conference on Artificial Intelligence, 39(24), 25651–25659. https://doi.org/10.1609/aaai.v39i24.34756
Zelikman, E., Wu, Y., Mu, J., & Goodman, N. (2022). STaR: Bootstrapping reasoning with reasoning. In Advances in Neural Information Processing Systems, 35, 15476–15488. https://doi.org/10.48550/arXiv.2203.14465
Zhang, B., Jian, M., Liu, Y., Huang, Y., Li, H., Liu, Y., Chen, G., Leng, S., Jiang, Y., Zhang, H., Li, X., Jin, P., Zhang, W., Wang, F., Bing, L., & Zhao, D. (2025). VideoLLaMA 3: Frontier multimodal foundation models for image and video understanding. arXiv. https://doi.org/10.48550/arXiv.2501.13106
Zhang, J., Luo, S., Zhang, R., & Su, Q. (2025). HUMORCHAIN: Theory-guided multi-stage reasoning for interpretable multimodal humor generation. arXiv. https://doi.org/10.48550/arXiv.2511.21732
Zhang, T., Kishore, V., Wu, F., Weinberger, K. Q., & Artzi, Y. (2020). BERTScore: Evaluating text generation with BERT. In Proceedings of the 8th International Conference on Learning Representations (ICLR 2020). https://doi.org/10.48550/arXiv.1904.09675
Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E. P., Zhang, H., Gonzalez, J. E., & Stoica, I. (2023). Judging LLM-as-a-judge with MT-bench and chatbot arena. In Advances in Neural Information Processing Systems, 36, 46595–46623. https://doi.org/10.48550/arXiv.2306.05685