跳到主要內容

簡易檢索 / 詳目顯示

研究生: 邱湘芝
Chiu, Hsiang-Chih
論文名稱: 人類與 AI 的後設認知於人機協作中的影響
The reciprocal influence of human and AI metacognition in human–AI collaboration
指導教授: 陳宜秀
Chen, Yih-siu
蔡炎龍
Tsai, Yen-Lung
口試委員: 黃從仁
Huang, Tsung-Ren
學位類別: 碩士
Master
系所名稱: 傳播學院 - 數位內容碩士學位學程
Digital Content and Technologies
論文出版年: 2026
畢業學年度: 114
語文別: 中文
論文頁數: 356
中文關鍵詞: 後設認知人工智慧後設認知人機協作以人為中心的人工智慧
外文關鍵詞: Metacognition, AI Metacognition, Human-AI Collaboration, Human- Centered AI
相關次數: 點閱:10下載:3
分享至:
查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報
  • 本研究以後設認知為核心,探討人類後設認知能力與 AI 後設認知線索在人機協作歷程中的作用,並且關注三項問題:AI 是否展現後設認知線索,是否會影響人類的信任判斷、互動行為與任務表現;人類自身後設認知能力是否影響其與 AI 協作時的監控與調整行為;以及不同人機後設認知組合是否形成不同的協作型態。

    本研究採 2 × 2 組間準實驗設計,以人類後設認知高低與 AI 後設認知有無作為主要自變項。第一階段先以圓點判斷任務估計受試者之 M-ratio,作為人類後設認知效率指標,並據此區分高、低後設認知組。第二階段則邀請受試者進行中文 Connections 關鍵詞分組任務,並與有或無後設認知線索的 AI 協作者互動。

    最終共納入 65 位受試者,分為(高/低人類後設認知)× (有/無 AI 後設認知)
    四組。研究蒐集任務績效、對話互動行為、主觀問卷與半結構式訪談資料,並以 ANOVA、ANCOVA、相關分析與質化主題整理進行分析。
    研究結果顯示,操弄之有/無 AI 後設認知能被受試者明確感知。AI 後設認知
    條件下,受試者較能察覺 AI 會表達信心水準、不確定性、共同判斷需求與可能犯錯,並且讓受試者找出較多正確詞組,錯誤提交次數也較少。然而,AI 後設認知能力並未直接提升整體信任或主觀滿意度,顯示 AI 後設認知的主要價值並非單純提高信任,而是使 AI 建議的可靠性變得更可被評估。在人類後設認知方面,高、低人類後設認知者雖然在任務正確率上未呈現差異,但在互動策略上展現不同型態。高後設認知者較常出現主動批判監控、探索協商與拒絕 AI 先前提出的分組方向等行為,將 AI 視為可被檢查與調整的協作資源,而非單純答案來源。相對地,低後設認知者的互動方式較容易受到 AI 條件影響;當 AI 缺乏後設認知線索時,較容易出現依賴求助、要求具體字詞或將任務交由 AI 主導的傾向。

    本研究同時進一步觀察人機後設認知組合後,發現「高後設認知人類 + 有後
    設認知 AI」並不必然為最佳的績效組合,但是在主觀經驗與互動型態上,則可觀察到後設認知相匹配與不相匹配組合的差異。後設認知相匹配的組合在部分合作感受上較佳,而不相匹配組合則可能出現較高信心與較差感受並存的錯位現象。此結果顯示,人機協作的品質不僅取決於人或 AI 個別能力高低,也取決於雙方如何理解、回應與調整彼此的判斷狀態。

    整體而言,本研究指出,後設認知在人機協作中的作用並非單純提升信任或直接改善表現。AI 後設認知的作用機制是「使建議可被判斷」而非「提升信任」;人類後設認知則影響使用者是否能有效運用 AI 所提供的線索。人機後設認知不存在簡單加乘關係,而是呈現『相匹配/不相匹配』的協作型態。本研究結果補充 Human-AI Teaming、AI 透明度與信任校準等相關研究,並為未來設計具後設認知線索的 AI 協作系統提供參考。


    This study centers on metacognition and examines the roles of human metacognitive ability and AI metacognitive cues in human–AI collaboration. It addresses three main questions: whether the presence of metacognitive cues in AI affects human trust judgments, interaction behaviors, and task performance; whether individuals' metacognitive ability influences how they monitor and adjust their behavior when collaborating with AI; and whether different combinations of human and AI metacognition give rise to distinct patterns of collaboration.

    A 2 × 2 between-subjects quasi-experimental design was adopted, with human metacognition level and the presence or absence of AI metacognitive cues as the primary independent variables. In the first stage, participants completed a dot discrimination task, from which their M-ratio was estimated as an indicator of metacognitive efficiency. Based on this measure, participants were classified into high- and low-metacognition groups. In the second stage, participants completed a Chinese-language Connections task involving keyword grouping while interacting with an AI collaborator that either provided or did not provide metacognitive cues. A total of 65 participants were included and assigned to four conditions defined by human metacognition level (high vs. low) and AI metacognitive cues (present vs. absent). Data were collected on task performance, conversational interaction behaviors, subjective questionnaire responses, and semi-structured interviews. The data were analyzed using ANOVA, ANCOVA, correlation analysis, and qualitative thematic analysis.

    The results showed that participants were able to clearly perceive the manipulation of AI metacognitive cues. When such cues were present, participants were more likely to recognize that the AI communicated its confidence level, uncertainty, need for joint judgment, and possibility of making errors. Participants in this condition also identified more correct word groups and made fewer incorrect submissions. However, AI metacognitive cues did not directly increase overall trust or subjective satisfaction. This suggests that the primary value of AI metacognition lies not in simply increasing trust, but in making the reliability of AI recommendations more assessable. With respect to human metacognition, participants with high and low metacognitive ability did not differ significantly in task accuracy, but they exhibited different interaction strategies. Participants with higher metacognitive ability more frequently engaged in active critical monitoring, exploratory negotiation, and rejection of previously proposed AI grouping suggestions. They treated the AI as a collaborative resource that could be examined and adjusted, rather than merely as a source of answers. In contrast, participants with lower metacognitive ability were more strongly influenced by the AI condition. When the AI did not provide metacognitive cues, they were more likely to rely on the AI for assistance, request specific words or answers, or allow the AI to take the lead in the task.

    Further examination of different human–AI metacognitive combinations showed that the combination of a highly metacognitive human and an AI that provided metacognitive cues did not necessarily produce the best task performance. Nevertheless, differences between metacognitively matched and mismatched combinations were observed in subjective experience and interaction patterns. Matched combinations were associated with more positive collaboration experiences on some measures, whereas mismatched combinations sometimes exhibited a misalignment in which higher confidence coexisted with less favorable subjective experiences. These findings indicate that the quality of human–AI collaboration depends not only on the individual abilities of the human or the AI, but also on how both parties interpret, respond to, and adjust to each other's judgment states.

    Overall, this study demonstrates that the role of metacognition in human–AI collaboration cannot be reduced to simply increasing trust or directly improving performance. The mechanism through which AI metacognition operates is to make its recommendations more assessable rather than merely more trustworthy, whereas human metacognition influences whether users can effectively interpret and use the cues provided by AI. Human and AI metacognition do not have a simple additive relationship; instead, they form distinct patterns of metacognitive match and mismatch. These findings extend existing research on Human–AI Teaming, AI transparency, and trust calibration, and provide implications for the future design of AI collaboration systems that incorporate metacognitive cues.

    致謝 i
    摘要 iv
    Abstract vi
    目錄 ix
    圖目錄 xvi
    表目錄 xviii
    第一章 緒論 1
    1.1 研究背景與動機 1
    1.2 研究目標 2
    第二章 文獻回顧 4
    2.1 協作 4
    2.1.1 協作的本質與定義 4
    2.1.2 人與人的協作 5
    2.1.3 人機之間的協作 7
    2.1.3.1 人機協作與 Human–AI Teaming、TIMS 7
    2.1.3.2 影響人機協作的因素 8
    2.2 後設認知 10
    2.2.1 後設認知的兩層次:顯性與隱性,以及四種知識情境 10
    2.2.2 AI 的後設認知 12
    2.3 後設認知會如何影響人與 AI 的協作? 14
    2.3.1 AI 與人後設認知的交互影響 14
    2.3.2 後設認知如何影響人機協作:高度自動化與高度人參與系統 16
    2.4 研究問題與研究假設 18
    第三章 研究方法 20
    3.1 實驗概述 20
    3.1.1 實驗設計 20
    3.1.2 獨立變項 21
    3.2 實驗系統 21
    3.2.1 實驗系統架構 22
    3.2.2 AI 後設認知操弄與後設認知控制架構設計 24
    3.2.2.1 任務複雜度與初步測試發現之困難 24
    3.2.2.2 共用前處理:遊戲狀態摘要階段 24
    3.2.2.3 有無後設認知條件開發說明 25
    3.2.3 前端介面與資料流程 27
    3.3 前導研究 29
    3.3.1 前導研究一:協作任務題目難度檢驗 29
    3.3.2 前導研究二:AI 後設認知操弄檢驗 31
    3.3.3 前導研究結果與正式實驗調整 35
    3.4 正式實驗 35
    3.4.1 第一階段:圓點判斷任務與人類後設認知分組 36
    3.4.2 第二階段:人機協作任務及互動方式 37
    3.4.2.1 任務流程 37
    3.4.2.2 任務指導語 40
    3.4.2.3 實驗流程 42
    3.5 資料蒐集 43
    3.5.1 任務行為資料:互動意圖指標建構 43
    3.5.2 依變項分述 46
    3.6 資料分析 51
    3.6.1 量化分析 51
    3.6.2 質化分析:人類訪談與 AI 對話過程分析 52
    第四章 研究結果與分析 53
    4.1 樣本描述統計 53
    4.1.1 整體樣本數量與後設認知能力分布 54
    4.1.2 性別與年齡分布 55
    4.1.3 教育程度 55
    4.1.4 AI 使用行為分布 56
    4.2 問卷信度分析 57
    4.3 AI 操弄檢查 59
    4.3.1 AI 後設認知操弄 60
    4.3.2 AI 能力與一般印象控制題 61
    4.3.3 AI 客觀行為控制檢驗 62
    4.3.4 操弄檢測總結 64
    4.4 實驗設計的檢驗 64
    4.4.1 人類與 AI 後設認知對任務績效之影響 65
    4.4.1.1 整體任務績效分析 65
    4.4.1.2 不同題目難度下的任務績效分析 68
    4.4.1.3 不同題目的任務績效分析 74
    4.4.1.4 小結 74
    4.4.2 人類與 AI 後設認知對互動行為之影響 75
    4.4.2.1 整體互動投入、強度分析 75
    4.4.2.2 七項互動意圖分析 77
    4.4.2.3 互動模式指標建構 80
    4.4.2.4 人類與 AI 後設認知對互動模式之影響 82
    4.4.3 題目難度對互動行為之層次補充分析 85
    4.4.3.1 互動投入與貢獻指標在難易問題中的差別 86
    4.4.3.2 互動策略與意圖模式在難易問題中的差別 93
    4.4.3.3 不同題目之補充分析 98
    4.4.3.4 互動模式與意圖發現總結 99
    4.4.4 人類與 AI 後設認知對主觀評估之影響 100
    4.4.4.1 整體主觀感受 100
    4.4.4.2 主觀信心分析 104
    4.4.4.3 題目主觀難度分析 106
    4.4.4.4 主觀感受總結 107
    4.5 探索分析:任務績效、互動行為與主觀評估之關聯 108
    4.5.1 任務績效與互動投入行為之關聯 108
    4.5.2 任務績效與互動模式之關聯 109
    4.5.3 主觀信心與作答正確性之關聯 110
    4.5.3.1 受試者層次分析 110
    4.5.3.2 提交層次分析 110
    4.5.3.3 不同組別的信心辨別力在不同難度下是否有差異 112
    4.5.4 主觀難度與任務績效/互動行為之關聯 113
    4.5.5 整合迴歸分析 114
    4.6 探索分析:AI 後設認知、信任與採納行為之關聯 115
    4.6.1 知覺 AI 後設認知與信任、依賴之關聯 115
    4.6.2 先驗信任(Q39)與採納行為、任務績效之關聯 117
    4.6.3 先驗 AI 信任(Q39)與主觀感受之關聯 119
    4.7 小結與對 RQ 的回應 120
    第五章 討論 123
    5.1 不同難度中的人機組合差異 123
    5.1.1 簡單題情境中的人機組合差異 123
    5.1.2 困難題情境中的人機組合差異 125
    5.2 RQ1 回應:AI 後設認知如何影響人機協作 126
    5.2.1 AI 後設認知對任務績效與互動方式的影響 126
    5.2.2 AI 後設認知對信任校準與適當依賴的影響 127
    5.2.3 AI 建議的可評估性與作用邊界 128
    5.3 RQ2 回應:人類後設認知如何影響人機協作 129
    5.3.1 人類後設認知對互動策略的影響 130
    5.3.2 人類後設認知對 AI 判斷與信任調整的影響 131
    5.3.3 人類後設認知在 AI 可信度判斷與知識分工中的作用 132
    5.4 RQ3 回應:不同人機組合如何影響人機協作 134
    5.4.1 績效層次:不是「高人 × 高 AI」就一定最好 135
    5.4.2 主觀感受層次:關鍵不只是高低,而是人機是否相匹配 135
    5.4.3 信心層次:不匹配組合反而出現校準風險 137
    第六章 研究結論、限制與未來方向 139
    6.1 研究發現 139
    6.2 研究貢獻 140
    6.2.1 對 Human-AI Teaming 的理論貢獻 140
    6.2.2 對 Explainable AI 與 AI 後設認知設計的實務啟示 142
    6.2.3 AI 使用者素養與教育訓練的啟示 143
    6.3 研究限制 144
    6.3.1 人類後設認知量測與 AI 協作情境之間的落差 144
    6.3.2 先驗 AI 信任可能影響人類後設認知的發揮 145
    6.3.3 互動行為指標的詮釋限制 146
    6.3.4 樣本數與任務情境的限制 147
    6.3.5 質化分析作為補充資料的限制 148
    6.3.6 任務難度與 AI 解題能力限制 148
    6.4 未來發展 149
    6.4.1 發展 AI 協作情境下的人類後設認知量測 149
    6.4.2 擴展不同任務難度與應用情境 150
    6.4.3 優化 AI 後設認知的表達與互動設計 150
    6.4.4 探討長期人機協作中的信任校準 151
    6.4.5 將 AI 後設認知納入使用者教育與素養訓練 151
    參考文獻 153
    Appendix A 第三章研究材料補充 163
    A.1 AI 協作提示詞 163
    A.1.1 遊戲狀態整理提示詞(stateSummary) 163
    A.1.2 無後設認知條件提示詞(nonMeta) 166
    A.1.3 後設認知條件第一階段提示詞(meta hypothesis) 171
    A.1.4 後設認知條件第二階段提示詞(meta critique) 173
    A.1.5 後設認知條件第三階段提示詞(meta finalize) 175
    A.2 中文 Connections 正式實驗題目 180
    Appendix B 第四章補充統計結果 182
    B.1 任務績效分析 182
    B.1.1 整體任務績效描述統計原始表 182
    B.1.2 整體任務績效變異數與共變數分析原始表 183
    B.1.3 不同難度下任務績效描述統計原始表 187
    B.1.4 不同難度下任務績效混合設計變異數與共變數分析原始表 189
    B.1.5 不同難度下平均作答時間三因子交互作用之簡單主效果與成對比較原始表 198
    B.1.6 不同題目任務績效描述統計原始表 199
    B.1.7 不同題目任務績效混合設計變異數與共變數分析原始表 202
    B.2 整體互動行為分析 210
    B.2.1 互動投入與貢獻指標描述統計原始表 210
    B.2.2 互動投入與貢獻指標變異數與共變數分析原始表 211
    B.2.3 互動模式與意圖類型採用數描述統計原始表 216
    B.2.4 互動模式與意圖類型採用數變異數與共變數分析原始表 218
    B.3 七項互動意圖分析 223
    B.3.1 七項互動意圖描述統計原始表 223
    B.3.2 七項互動意圖變異數與共變數分析原始表 225
    B.3.3 要求字詞意圖三因子交互作用之簡單主效果與成對比較原始表 232
    B.3.4 七項互動意圖 Spearman 相關分析原始表 234
    B.4 不同題目難度與不同題目之互動行為分析 235
    B.4.1 不同難度下互動投入與貢獻指標描述統計原始表 235
    B.4.2 不同難度下互動投入與貢獻指標混合設計變異數與共變數分析原始表 238
    B.4.3 人類貢獻比率三因子交互作用之簡單主效果與成對比較原始表 250
    B.4.4 不同難度下互動模式與意圖類型採用數描述統計原始表 253
    B.4.5 不同難度下互動模式與意圖類型採用數混合設計變異數與共變數分析原始表 255
    B.4.6 意圖類型採用數三因子交互作用之簡單主效果與成對比較原始表 265
    B.4.7 不同題目下互動投入與貢獻指標描述統計原始表 266
    B.4.8 不同題目下互動投入與貢獻指標混合設計變異數與共變數分析原始表 270
    B.4.9 不同題目下互動模式與意圖類型採用數描述統計原始表 284
    B.4.10 不同題目下互動模式與意圖類型採用數混合設計變異數與共變數分析原始表 288
    B.5 主觀評估分析 303
    B.5.1 主觀感受七構面描述統計原始表 303
    B.5.2 主觀感受七構面變異數與共變數分析原始表 305
    B.5.3 事後問卷各題描述統計原始表 312
    B.5.4 事後問卷各題變異數與共變數分析原始表 319
    B.5.5 整體主觀信心描述統計原始表 342
    B.5.6 整體主觀信心變異數與共變數分析原始表 343
    B.5.7 不同題目難度下主觀信心與主觀難度描述統計原始表 344
    B.5.8 不同題目難度下主觀信心與主觀難度混合設計變異數與共變數分析原始表 345
    B.5.9 不同題目下主觀信心與主觀難度描述統計原始表 349
    B.5.10 不同題目下主觀信心與主觀難度混合設計變異數與共變數分析原始表 350
    B.5.11 知覺 AI 後設認知與七項主觀感受構面之 Spearman 相關分析 354
    B.5.12 先驗 AI 信任(Q39)與七項主觀感受構面之 Pearson 相關分析 356

    [1] Saleh Afroogh, Ali Akbari, Emmie Malone, Mohammadali Kargar, and Hananeh Alambeigi. 2024. Trust in AI: progress, challenges, and future directions. Humanities and Social Sciences Communications 11, 1 (2024), 1568. doi: 10.1057/ s41599-024-04044-8
    [2] Ann E. Austin and Roger G. Baldwin. 1991. Faculty Collaboration: Enhancing the Quality of Scholarship and Teaching. ASHE-ERIC Higher Edu cation Report No. 7, 1991. Technical Report. ASHE-ERIC Higher Education Report No. 7. ERIC Number: ED346805. https://eric.ed.gov/?id=ED346805
    [3] Lisanne Bainbridge. 1983. Ironies of automation. Automatica 19, 6 (1983), 775–779. doi: 10.1016/0005-1098(83)90046-8
    [4] Gabriele Bammer, Michael Smithson, and The Goolabri Group. 2008. The Nature of Uncertainty. In Uncertainty and Risk: Multidisciplinary Perspectives, Gabriele Bammer and Michael Smithson (Eds.). Earthscan, London, 289–303. doi: 10.4324/ 9781849773607
    [5] David Bani-Harouni, Chantal Pellegrini, Paul Stangel, Ege Özsoy, Kamilia Zaripova, Nassir Navab, and Matthias Keicher. 2026. Rewarding Doubt: A Reinforcement Learning Approach to Calibrated Confidence Expression of Large Language Models. arXiv:2503.02623 [cs.CL]. doi:10.48550/arXiv.2503.02623
    [6] Sophie Berretta, Alina Tausch, Greta Ontrup, Björn Gilles, Corinna Peifer, and Annette Kluge. 2023. Defining human-AI teaming the human-centered way: a scoping review and network analysis. Frontiers in Artificial Intelligence 6 (2023), 1250725.doi: 10.3389/frai.2023.1250725
    [7] Fanni Biró and Csaba Csíkos. 2025. The Broadwell-Burch Four-Stage Model of Competence Development and the Mathematics Teaching Profession. Mathematics Teaching Research Journal 17, 3 (2025), 5–25. https://eric.ed.gov/?id= EJ1481729
    [8] Jessica Y Bo, Sophia Wan, and Ashton Anderson. 2025. To Rely or Not to Rely? Evaluating Interventions for Appropriate Reliance on Large Language Models. In Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems (CHI ’25). Association for Computing Machinery, New York, NY, USA, 1–23. doi: 10.1145/3706598.3714097
    [9] Michael E. Bratman. 1992. Shared Cooperative Activity. The Philosophical Review 101, 2 (1992), 327–341. doi: 10.2307/2185537
    [10] Martin M. Broadwell. 1969. Teaching for Learning (XVI.). The Gospel Guardian 20, 41(Feb.1969), 1–3a.https://wordsfitlyspoken.org/gospel_guardian/ v20/v20n41p1-3a.html
    [11] Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu,Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, SamMcCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners. In Proceedings of the 34th International Conference on Neural Information Processing Systems (NIPS ’20). Curran Associates Inc., Red Hook, NY, USA, 1877–1901. doi: 10.5555/3495724.3495883
    [12] Janis A. Cannon-Bowers, Eduardo Salas, and Sharolyn Converse. 1993. Shared Mental Models in Expert Team Decision Making. In Individual and Group Decision Making: Current Issues, Jr. Castellan, N. John (Ed.). Lawrence Erlbaum Associates,Hillsdale, NJ, USA, 221–245. doi: 10.4324/9780203772744-16
    [13] Qi Cao, Yufan Wang, Peijia Qin, Shuhao Zhang, and Pengtao Xie. 2026. LLMs Know When They Know, but Do Not Act on It: A Metacognitive Harness for Test-time Scaling. arXiv preprint arXiv:2605.14186. doi: 10.48550/arXiv.2605.14186
    [14] Nancy J. Cooke. 2015. Team Cognition as Interaction. Current Directions in Psychological Science 24, 6 (2015), 415–419. doi:10.1177/0963721415602474
    [15] Zoltan Dienes and Josef Perner. 1999. A theory of implicit and explicit knowledge. Behavioral and Brain Sciences 22, 5 (1999), 735–808. doi: 10.1017/ S0140525X99002186
    [16] Finale Doshi-Velez and Been Kim. 2017. Towards A Rigorous Science of Interpretable Machine Learning. arXiv:1702.08608 [stat.ML]. doi:10.48550/arXiv. 1702.08608
    [17] Embodied Computation Group. 2024. metadPy: Metacognitive Efficiency Modelling in Python.Software.https://github.com/subjectivitylab/metadPy
    [18] Mica R. Endsley. 1995. Toward a Theory of Situation Awareness in Dynamic Systems. Human Factors 37, 1 (1995), 32–64.doi:10.1518/001872095779049543
    [19] M. R. Endsley and D. B. Kaber. 1999. Level of automation effects on performance, situation awareness and workload in a dynamic control task. Ergonomics 42, 3 (March 1999), 462–492. doi: 10.1080/001401399185595
    [20] Md Meftahul Ferdaus, Mahdi Abdelguerfi, Elias Loup, Kendall N. Niles, Ken Pathak, and Steven Sloan. 2026. Towards Trustworthy AI: A Review of Ethical and Robust Large Language Models. Comput. Surveys 58, 7 (2026), 176:1–176:43. doi: 10.1145/3777382
    [21] John H. Flavell. 1979. Metacognition and cognitive monitoring: A new area of cognitive–developmental inquiry. American Psychologist 34, 10 (1979), 906–911. doi: 10.1037/0003-066X.34.10.906
    [22] Chris D. Frith. 2012. The role of metacognition in human social interactions. Philosophical Transactions of the Royal Society B: Biological Sciences 367, 1599 (2012), 2213–2223. doi: 10.1098/rstb.2012.0123
    [23] Chris D. Frith and Uta Frith. 2022. The mystery of the brain–culture interface. Trends in Cognitive Sciences 26, 12 (2022), 1023–1025. doi: 10.1016/j.tics. 2022.08.013
    [24] Frederic Gmeiner, Kaitao Luo, Ye Wang, Kenneth Holstein, and Nikolas Martelaro. 2025. Exploring the Potential of Metacognitive Support Agents for Human-AI Co-Creation. In Proceedings of the 2025 ACM Designing Interactive Systems Conference (DIS ’25). Association for Computing Machinery, New York, NY, USA, 1244–1269. doi: 10.1145/3715336.3735785
    [25] Barbara Gray. 1989. Collaborating: Finding Common Ground for Multiparty Problems. Jossey-Bass, San Francisco, CA, USA. 329 pages.
    [26] Eetu Haataja, Muhterem Dindar, Jonna Malmberg, and Sanna Järvelä. 2022. Individuals in a group: Metacognitive and regulatory predictors of learning achievement in collaborative learning. Learning and Individual Differences 96 (2022), 102146. doi: 10.1016/j.lindif.2022.102146
    [27] Arthur Turovh Himmelman. 1996. On the Theory and Practice of Transformational Collaboration: From Social Service to Social Justice. In Creating Collaborative Advantage, Chris Huxham (Ed.). SAGE Publications, London, UK, 20–43. doi: 10.4135/9781446221600.n2
    [28] Guy Hoffman. 2019. Evaluating Fluency in Human–Robot Collaboration. IEEE Transactions on Human-Machine Systems 49, 3 (2019), 209–218. doi: 10.1109/ THMS.2019.2904558
    [29] Konstantin Hopf, Nora Nahr, Thorsten Staake, and Franz Lehner. 2025. The group mind of hybrid teams with humans and intelligent agents in knowledge-intense work. Journal of Information Technology 40, 1 (2025), 9–34. doi: 10.1177/02683962241296883
    [30] Lei Huang, Weijiang Yu, Weitao Ma, Weihong Zhong, Zhangyin Feng, Haotian Wang, Qianglong Chen, Weihua Peng, Xiaocheng Feng, Bing Qin, and Ting Liu. 2025. A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions. ACM Transactions on Information Systems 43, 2 (2025), 42:1–42:55. doi: 10.1145/3703155
    [31] Mohammad Hossein Jarrahi. 2018. Artificial intelligence and the future of work: Human-AI symbiosis in organizational decision making. Business Horizons 61, 4 (2018), 577–586. doi: 10.1016/j.bushor.2018.03.007
    [32] S Mo Jones-Jang and Yong Jin Park. 2023. How do people react to AI failure? Automation bias, algorithmic aversion, and perceived controllability. Journal of Computer-Mediated Communication 28, 1 (2023), zmac029. doi: 10.1093/jcmc/ zmac029
    [33] Saurav Kadavath, Tom Conerly, Amanda Askell, Tom Henighan, Dawn Drain, Ethan Perez, Nicholas Schiefer, Zac Hatfield-Dodds, Nova DasSarma, Eli Tran-Johnson, Scott Johnston, Sheer El-Showk, Andy Jones, Nelson Elhage, Tristan Hume, Anna Chen, Yuntao Bai, Sam Bowman, Stanislav Fort, Deep Ganguli, Danny Hernandez, Josh Jacobson, Jackson Kernion, Shauna Kravec, Liane Lovitt, Kamal Ndousse, Catherine Olsson, Sam Ringer, Dario Amodei, Tom Brown, Jack Clark, Nicholas Joseph, Ben Mann, Sam McCandlish, Chris Olah, and Jared Kaplan. 2022. Language Models (Mostly) Know What They Know. arXiv:2207.05221 [cs.CL] doi:10.48550/arXiv.2207.05221
    [34] Ann Kerwin. 1993. None Too Solid: Medical Ignorance. Knowledge 15, 2 (1993), 166–185. doi: 10.1177/107554709301500204
    [35] Gary Klein, Sterling Wiggins, and Cynthia O. Dominguez. 2010. Team Sensemaking. Theoretical Issues in Ergonomics Science 11, 4 (2010), 304–320. doi: 10.1080/14639221003729177
    [36] Andrew Stuart Lane and Chris Roberts. 2022. Contextualised reflective competence:a new learning model promoting reflective practice for clinical training. BMC Medical Education 22 (2022), 71. doi: 10.1186/s12909-022-03112-4
    [37] Doyeon Lee, Joseph Pruitt, Tianyu Zhou, Jing Du, and Brian Odegaard. 2025. Metacognitive sensitivity: The key to calibrating trust and optimal decision making with AI. PNAS Nexus 4, 5 (2025), pgaf133. doi: 10.1093/pnasnexus/pgaf133
    [38] John D. Lee and Katrina A. See. 2004. Trust in Automation: Designing for Appropriate Reliance. Human Factors 46, 1 (2004), 50–80. doi: 10.1518/hfes.46.1. 50_30392
    [39] Kyle Lewis. 2003. Measuring Transactive Memory Systems in the Field: Scale Development and Validation. Journal of Applied Psychology 88 (2003), 587–604. doi: 10.1037/0021-9010.88.4.587
    [40] Rongxing Liu, Kumar Shridhar, Manish Prajapat, Patrick Xia, and Mrinmaya Sachan. 2024. SMART: Self-learning Meta-strategy Agent for Reasoning Tasks.arXiv:2410.16128 [cs.AI]. doi: 10.48550/arXiv.2410.16128
    [41] Shuai Ma, Xinru Wang, Ying Lei, Chuhan Shi, Ming Yin, and Xiaojuan Ma. 2024. “Are You Really Sure?"Understanding the Effects of Human Self-Confidence Calibration in AI-Assisted Decision Making. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems (CHI ’24). Association for Computing Machinery, New York, NY, USA, 1–20. doi: 10.1145/3613904.3642671
    [42] Laura R. Marusich, Jonathan Z. Bakdash, Yan Zhou, and Murat Kantarcioglu. 2024. Using AI Uncertainty Quantification to Improve Human Decision-Making.arXiv:2309.10852 [cs.AI]. doi: 10.48550/arXiv.2309.10852
    [43] John E. Mathieu, Tonia S. Heffner, Gerald F. Goodwin, Eduardo Salas, and Janis A. Cannon-Bowers. 2000. The influence of shared mental models on team process and performance. Journal of Applied Psychology 85, 2 (2000), 273–283. doi: 10.1037/ 0021-9010.85.2.273
    [44] Nat McAleese, Rai Michael Pokorny, Juan Felipe Ceron Uribe, Evgenia Nitishinskaya, Maja Trebacz, and Jan Leike. 2024. LLM Critics Help Catch LLM Bugs.arXiv:2407.00215 [cs.SE]. doi: 10.48550/arXiv.2407.00215
    [45] Susan Mohammed, Lori Ferzandi, and Katherine Hamilton. 2010. Metaphor No More: A 15-Year Review of the Team Mental Model Construct. Journal of Management 36, 4 (2010), 876–910. doi:10.1177/0149206309356804
    [46] Thomas O. Nelson. 1990. Metamemory: A Theoretical Framework and New Findings. In Psychology of Learning and Motivation. Vol. 26. Academic Press, San Diego, CA, 125–173. doi: 10.1016/S0079-7421(08)60053-5
    [47] Sergei Nirenburg, Marjorie McShane, and Thomas M. Ferguson. 2025. Mutual Trust in Human–AI Teams Relies on Metacognition. In Metacognitive Artificial Intelligence, Paulo Shakarian and Hua Wei (Eds.). Cambridge University Press, Cambridge, 56–80. doi: 10.1017/9781009522472.007
    [48] Raja Parasuraman and Victor Riley. 1997. Humans and Automation: Use, Misuse, Disuse, Abuse. Human Factors 39, 2 (1997), 230–253. doi: 10.1518/ 001872097778543886
    [49] R. Parasuraman, T.B. Sheridan, and C.D. Wickens. 2000. A model for types and levels of human interaction with automation. IEEE Transactions on Systems, Man, and Cybernetics - Part A: Systems and Humans 30, 3 (2000), 286–297. doi: 10. 1109/3468.844354
    [50] Vesa Peltokorpi and Mervi Hasu. 2016. Transactive memory systems in research team innovation: A moderated mediation analysis. Journal of Engineering and Technology Management 39(2016), 1–12. doi:10.1016/j.jengtecman.2015.11.001
    [51] Dobromir Rahnev. 2025. A comprehensive assessment of current methods for measuring metacognition. Nature Communications 16, 1 (2025), 701. doi: 10.1038/ s41467-025-56117-0
    [52] Iyad Rahwan, Manuel Cebrian, Nick Obradovich, Josh Bongard, Jean-François Bonnefon, Cynthia Breazeal, Jacob W. Crandall, Nicholas A. Christakis, Iain D. Couzin, Matthew O. Jackson, Nicholas R. Jennings, Ece Kamar, Isabel M. Kloumann, Hugo Larochelle, David Lazer, Richard McElreath, Alan Mislove, David C. Parkes, Alex`Sandy'Pentland, MargaretE.Roberts, AzimShariff, JoshuaB.Tenenbaum,and Michael Wellman. 2019. Machine behaviour. Nature 568, 7753 (2019), 477– 486. doi: 10.1038/s41586-019-1138-y
    [53] Max Schemmer, Niklas Kuehl, Carina Benz, Andrea Bartos, and Gerhard Satzger. 2023. Appropriate Reliance on AI Advice: Conceptualization and the Effect of Explanations. In Proceedings of the 28th International Conference on Intelligent User Interfaces (IUI ’23). Association for Computing Machinery, New York, NY, USA, 410–422. doi: 10.1145/3581641.3584066
    [54] Alan H. Schoenfeld. 1987. What’s All the Fuss About Metacognition? In Cognitive Science and Mathematics Education. Lawrence Erlbaum Associates, Hillsdale, NJ, 189–215. Num Pages: 27. https://www.taylorfrancis.com/chapters/edit/ 10.4324/9780203062685-8/fuss-metacognition-alan-schoenfeld
    [55] Isabella Seeber, Eva Bittner, Robert O. Briggs, Triparna de Vreede, Gert-Jan de Vreede, Aaron Elkins, Ronald Maier, Alexander B. Merz, Sarah Oeste-Reiß, Nils Randrup, Gerhard Schwabe, and Matthias Söllner. 2020. Machines as teammates: A research agenda on AI in team collaboration. Information & Management 57, 2 (2020), 103174. doi: 10.1016/j.im.2019.103174
    [56] Thomas B. Sheridan and William L. Verplank. 1978. Human and Computer Control of Undersea Teleoperators. Technical Report ADA057655. MIT Man-Machine Systems Laboratory, Cambridge, MA, USA. 188 pages. https://ntrl.ntis.gov/NTRL/dashboard/searchResults/titleDetail/ADA057655.xhtml
    [57] Noah Shinn, Federico Cassano, Edward Berman, Ashwin Gopinath, Karthik Narasimhan, and Shunyu Yao. 2023. Reflexion: Language Agents with Verbal Reinforcement Learning. arXiv:2303.11366 [cs.AI]. doi: 10.48550/arXiv.2303. 11366
    [58] Ben Shneiderman. 2020. Human-Centered Artificial Intelligence: Reliable, Safe & Trustworthy. International Journal of Human–Computer Interaction 36, 6 (2020),495–504. doi: 10.1080/10447318.2020.1741118
    [59] Tejas Srinivasan and Jesse Thomason. 2026. Adjust for Trust: Mitigating TrustInduced Inappropriate Reliance on AI Assistance. arXiv:2502.13321 [cs.HC]. doi: 10.48550/arXiv.2502.13321
    [60] Mark Steyvers, Heliodoro Tejeda, Aakriti Kumar, Catarina Belem, Sheer Karny, Xinyue Hu, Lukas W. Mayer, and Padhraic Smyth. 2025. What large language models know and what people think they know. Nature Machine Intelligence 7, 2 (2025), 221–231. doi: 10.1038/s42256-024-00976-7
    [61] Andreas Stolcke, Noah Coccaro, Rebecca Bates, Paul Taylor, Carol Van EssDykema, Klaus Ries, Elizabeth Shriberg, Daniel Jurafsky, Rachel Martin, and Marie Meteer. 2000. Dialogue act modeling for automatic tagging and recognition of conversational speech. Computational Linguistics 26, 3 (2000), 339–373. doi: 10.1162/089120100561737
    [62] Lev Tankelevitch, Viktor Kewenig, Auste Simkute, Ava Elizabeth Scott, Advait Sarkar, Abigail Sellen, and Sean Rintel. 2024. The Metacognitive Demands and Opportunities of Generative AI. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems (CHI ’24). Association for Computing Machinery, New York, NY, USA, 1–24. doi: 10.1145/3613904.3642902
    [63] United Nations Educational, Scientific and Cultural Organization. 2021. Recommendation on the Ethics of Artificial Intelligence. Adopted 23 November 2021. https://www.unesco.org/en/legal-affairs/recommendation-ethics-artificial-intelligence?hub=1063
    [64] Mor Vered, Tali Livni, Piers Douglas Lionel Howe, Tim Miller, and Liz Sonenberg. 2023. The effects of explanations on automation bias. Artificial Intelligence 322 (2023), 103952. doi: 10.1016/j.artint.2023.103952
    [65] Peter B. Walker, Jonathan J. Haase, Melissa L. Mehalick, Christopher T. Steele, Dale W. Russell, and Ian N. Davidson. 2025. Harnessing Metacognition for Safe and Responsible AI. Technologies 13, 3 (2025), 107. doi: 10.3390/ technologies13030107
    [66] Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou. 2023. Self-Consistency Improves Chain of Thought Reasoning in Language Models. arXiv:2203.11171 [cs.CL]. doi: 10.48550/arXiv.2203.11171
    [67] Daniel M. Wegner. 1995. A computer network model of human transactive memory. Social Cognition 13, 3 (1995), 319–339. doi: 10.1521/soco.1995.13.3.319
    [68] Jiexi Xu and Qianhui Lu. 2026. Agentic Metacognition: Designing a Self-aware Low-Code Agent for Failure Prediction and Human Handoff. In Focus on Artificial Intelligence in Intelligent Systems Design, Radek Silhavy and Petr Silhavy (Eds.).Springer Nature Switzerland, Cham, 410–421. doi: 10.1007/978-3-032-22236-7_33
    [69] Hamed Zamani, Johanne R. Trippas, Jeff Dalton, and Filip Radlinski. 2023. Conversational Information Seeking. Foundations and Trends in Information Retrieval 17, 3–4 (2023), 244–456. doi: 10.1561/1500000081
    [70] Xuanming Zhang, Yuxuan Chen, Samuel Yeh, and Sharon Li. 2025. MetaMind: Modeling Human Social Thoughts with Metacognitive Multi-Agent Systems. arXiv:2505.18943 [cs.CL]. doi: 10.48550/arXiv.2505.18943
    [71] Zhi-Xue Zhang, Paul S. Hempel, Yu-Lan Han, and Dean Tjosvold. 2007. Transactive Memory System Links Work Team Characteristics and Performance. Journal of Applied Psychology 92, 6 (2007), 1722–1730. doi: 10.1037/0021-9010.92.6. 1722
    [72] Michelle Zhao, Reid Simmons, and Henny Admoni. 2025. The Role of Adaptation in Collective Human–AI Teaming. Topics in Cognitive Science 17, 2 (2025), 291–323. doi: 10.1111/tops.12633

    QR CODE
    :::