跳到主要內容

簡易檢索 / 詳目顯示

研究生: 張敦皓
Chang, Tun-Hao
論文名稱: 真實退化情境下 AI 生成影像偵測之計算感知局部證據分析
Compute-Aware Local Evidence Analysis for AI Image Detection under Real Degradations
指導教授: 廖文宏
Liao, Wen-Hung
口試委員: 紀明德
Chi, Ming-Te
劉遠楨
Liu, Yuan-Chen
學位類別: 碩士
Master
系所名稱: 資訊學院 - 資訊科學系碩士在職專班
Excutive Master Program of Computer Science
論文出版年: 2026
畢業學年度: 115
語文別: 中文
論文頁數: 84
中文關鍵詞: AI 生成影像偵測局部證據真實世界穩健性再數位化計算感知推論
外文關鍵詞: AI-generated image detection, local evidence, real-world robustness, re-digitization, compute-aware inference
相關次數: 點閱:34下載:0
分享至:
查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報
  • 近年生成式人工智慧(Generative AI)影像技術快速發展,文字生成影像模型(text-to-image models)已能產生高度逼真的視覺內容,並廣泛影響數位內容生態,使數位影像真偽鑑識與內容信任面臨更複雜的挑戰。既有整張影像偵測模型(full-image detectors)雖可取得良好的評估效能,但較難說明其判斷所依賴的局部證據(local evidence),也不易量化推論過程所需的證據預算(evidence budget),對其在壓縮、轉傳與再數位化等真實流通條件下的失效行為與邊界解釋也相對有限。為回應此問題,本文以 RRDataset / RRBench 的真實退化條件為基礎,提出「計算感知局部證據選擇」(Compute-Aware Evidence Selection, CAES)分析框架。此框架串接候選區塊建構、特徵排序與隨時可停式(anytime)聚合機制,以追蹤局部線索的可用性並量化局部證據需求。
    實驗結果顯示,局部證據的可用性具有明顯的情境相依性。在未經退化的原始情境中,加入提前停止機制的影像區塊分支(Patch Early Stopping, Patch ES)呈現較強的鑑識能力,AUC 達 0.9532,並明顯高於匹配式整張影像基準模型;相對地,再數位化流程會削弱高頻與局部殘差特徵,使影像區塊分支效能降至 AUC 0.6424,凸顯局部鑑識線索在此情境下的失效邊界。
    進一步的分數層級互補性分析(CAES-G)顯示,在固定測試切分中結合局部敏感度與整體影像脈絡可形成互補的診斷訊號,整體 AUC 達 0.8875。本文將此結果作為事後分數層級後融合診斷,用以觀察局部與整體證據在偵測任務中的互補關係。本論文因此將局部證據的動態累積、證據預算權衡與真實世界失效邊界納入同一分析視角,為後續 AI 生成影像偵測方法的證據分析提供受控實證依據。


    Recent advances in generative artificial intelligence have enabled text-to-image models to produce highly realistic visual content, making them increasingly common in digital media and creating more complex challenges for digital image forensics and content trust. Although full-image detectors can achieve strong evaluation performance, they provide limited visibility into the local evidence on which decisions depend, the evidence budget required during inference, and the failure patterns and boundaries that emerge under real-world circulation conditions such as compression, forwarding, and re-digitization. To examine these issues, this thesis proposes Compute-Aware Evidence Selection (CAES), an analysis framework grounded in the real-degradation conditions of RRDataset / RRBench. This framework integrates candidate-region construction, feature-based ranking, and anytime aggregation to track the availability of local cues and quantify local evidence requirements.
    The results indicate that the usefulness of local evidence is strongly scenario-dependent. In the original, non-degraded scenario, the early-stopping patch branch (Patch ES) shows strong forensic capability, reaching an AUC of 0.9532 and outperforming the matched full-image baseline. In contrast, re-digitization weakens high-frequency and local residual cues, reducing patch-branch performance to an AUC of 0.6424 and highlighting a failure boundary for local forensic evidence under this scenario.
    A further score-level complementarity analysis (CAES-G) shows that combining local sensitivity with global context produces a complementary diagnostic signal on the fixed test split, achieving an overall AUC of 0.8875. This result is used as a post-hoc score-level late-fusion diagnostic that supports the complementary value of local and global evidence in the detection task. Overall, this study brings the dynamic accumulation of local evidence, evidence-budget trade-offs, and real-world failure boundaries into one analytical perspective, providing controlled empirical grounding for future evidence analysis in AI-generated image detection.

    摘要 i
    Abstract ii
    目次 iii
    表次 vi
    圖次 viii
    式次 ix
    第一章 緒論 1
    1.1 研究背景 1
    1.2 研究動機 2
    1.3 研究目的 3
    1.4 研究貢獻 4
    1.5 論文架構 5
    1.6 研究範圍與限制 6
    第二章 技術背景與相關研究 7
    2.1 AI 生成影像偵測與被動式影像內容鑑識 7
    2.2 真實世界退化與 RRDataset / RRBench 8
    2.3 低階、頻域、殘差與重建誤差偵測路線 9
    2.4 局部證據、影像區塊偵測與局部/整體互補性 10
    2.5 Foundation / CLIP / VLM / 混合式偵測器作為外部脈絡 12
    2.6 浮水印、來源憑證與影像內容鑑識的互補關係 12
    2.7 計算感知推論、校準、提前停止與證據預算之概念脈絡 13
    2.8 CAES 之研究定位與本章小結 13
    第三章 研究方法 14
    3.1 CAES 整體架構 14
    3.2 問題定義 15
    3.3 候選影像區塊集合建構 16
    3.4 低成本局部特徵擷取 17
    3.5 選擇器排序機制 19
    3.6 影像區塊編碼器 20
    3.7 隨時可停式聚合器 20
    3.8 溫度校準與推論階段提前停止 21
    3.9 匹配式整張影像基準模型 22
    3.10 CAES-G 50:50 分數層級後融合診斷 23
    3.11 小結 23
    第四章 實驗設計 24
    4.1 資料集與情境定義 24
    4.2 清單檔與資料切分 25
    4.3 訓練流程 26
    4.4 評估指標 26
    4.5 基準方法與實驗組 28
    4.6 推論與評估流程 29
    4.7 公平比較與解讀範圍 30
    4.8 可重現性摘要 31
    4.9 K=26 尺度組合與 K>26 證據範圍敏感度分析 31
    第五章 實驗結果與分析 33
    5.1 主要效能與情境相依性 33
    5.2 證據預算與提前停止行為 36
    5.3 CAES-G 50:50 分數層級後融合診斷 41
    5.4 Gated 事後分數診斷 42
    5.5 候選池建構與選擇器排序分析 43
    5.6 96/128 較大局部上下文後續設定 47
    5.7 K=26 以後的證據範圍敏感度 49
    5.8 特徵退化與視覺化案例分析 51
    5.9 外部基準脈絡定位 54
    5.10 研究問題小結 55
    第六章 結論與未來工作 57
    6.1 回答 RQ1 57
    6.2 回答 RQ2 57
    6.3 回答 RQ3 57
    6.4 研究貢獻 57
    6.5 研究限制 58
    6.6 未來工作 59
    參考文獻 60
    附錄 A 重現性摘要與實驗紀錄 64
    A.1 正式實驗紀錄索引 64
    A.2 硬體與軟體環境摘要 64
    A.3 Dependency 與主要套件版本 65
    A.4 Training config 與 preflight metadata 摘要 65
    A.5 Session 2/3/4/5 補強 artifact 摘要 66
    A.6 K=26 以後補充分析紀錄 73
    A.7 凍結分數尺度組合診斷 75
    A.8 未保存 metadata 與可重現性限制 76
    附錄 B 方法偽程式碼 77
    附錄 C 整張影像 saliency 與影像區塊關注區域分析 78

    [1] L. Verdoliva, “Media forensics and deepfakes: An overview,” IEEE J. Sel. Topics Signal Process., vol. 14, no. 5, pp. 910–932, 2020. https://doi.org/10.1109/JSTSP.2020.3002101
    [2] S.-Y. Wang, O. Wang, R. Zhang, A. Owens, and A. A. Efros, “CNN-generated images are surprisingly easy to spot... for now,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), pp. 8692–8701, 2020. https://doi.org/10.1109/CVPR42600.2020.00872
    [3] C. Li et al., “Bridging the gap between ideal and real-world evaluation: Benchmarking AI-generated image detection in challenging scenarios,” in Proc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV), pp. 20379–20389, 2025. https://doi.org/10.1109/ICCV51701.2025.01895
    [4] P. Fernandez, G. Couairon, H. Jégou, M. Douze, and T. Furon, “The Stable Signature: Rooting watermarks in latent diffusion models,” in Proc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV), pp. 22409–22420, 2023. https://doi.org/10.1109/ICCV51070.2023.02053
    [5] K. Balan, S. Agarwal, S. Jenni, A. Parsons, A. Gilbert, and J. Collomosse, “EKILA: Synthetic media provenance and attribution for generative art,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. Workshops (CVPRW), pp. 913–922, 2023. https://doi.org/10.1109/CVPRW59228.2023.00098
    [6] M. Guo et al., “AI-generated image detection: Passive or watermark?” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. Workshops (CVPRW), pp. 400–410, 2026. https://openaccess.thecvf.com/content/CVPR2026W/SAFE/html/Guo_AI-generated_Image_Detection_Passive_or_Watermark_CVPRW_2026_paper.html
    [7] J. Frank, T. Eisenhofer, L. Schönherr, A. Fischer, D. Kolossa, and T. Holz, “Leveraging frequency analysis for deep fake image recognition,” in Proc. 37th Int. Conf. Mach. Learn. (ICML), Proc. Mach. Learn. Res., vol. 119, pp. 3247–3258, 2020. https://proceedings.mlr.press/v119/frank20a.html
    [8] H. Farid, “Image forgery detection,” IEEE Signal Process. Mag., vol. 26, no. 2, pp. 16–25, 2009. https://doi.org/10.1109/MSP.2008.931079
    [9] A. Piva, “An overview on image forensics,” ISRN Signal Process., vol. 2013, Art. no. 496701, 22 pp., 2013. https://doi.org/10.1155/2013/496701
    [10] L. Yuan, X. Li, Y. Zhang, J. Zhang, H. Li, and X. Gao, “MLEP: Multi-granularity local entropy patterns for generalized AI-generated image detection,” in Adv. Neural Inf. Process. Syst., vol. 38, pp. 68981–69000, 2025. https://papers.nips.cc/paper_files/paper/2025/hash/63db94e06ce85dde10871c99dd39ad98-Abstract-Conference.html
    [11] S. Liang, J. Liu, R. Chen, and Q. Guan, “FerretNet: Efficient synthetic image detection via local pixel dependencies,” in Adv. Neural Inf. Process. Syst., vol. 38, pp. 42109–42133, 2025. https://papers.nips.cc/paper_files/paper/2025/hash/3bfcc45018b4156b92197db47e9742bd-Abstract-Conference.html
    [12] H. Liu, Z. Tan, C. Tan, Y. Wei, J. Wang, and Y. Zhao, “Forgery-aware adaptive transformer for generalizable synthetic image detection,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), pp. 10770–10780, 2024. https://doi.org/10.1109/CVPR52733.2024.01024
    [13] Z. Yang et al., “All patches matter, more patches better: Enhance AI-generated image detection via panoptic patch learning,” in Proc. Int. Conf. Learn. Represent. (ICLR), 2026. https://iclr.cc/virtual/2026/poster/10007395
    [14] H.-S. Chen, S. Hu, S. You, and C.-C. J. Kuo, “DefakeHop++: An enhanced lightweight deepfake detector,” APSIPA Trans. Signal Inf. Process., vol. 11, no. 2, pp. 1–21, 2022. https://doi.org/10.1561/116.00000126
    [15] A. Howard et al., “Searching for MobileNetV3,” in Proc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV), pp. 1314–1324, 2019. https://doi.org/10.1109/ICCV.2019.00140
    [16] A. Gushchin et al., “NTIRE 2026 challenge on robust AI-generated image detection in the wild,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. Workshops (CVPRW), pp. 1895–1913, 2026. https://openaccess.thecvf.com/content/CVPR2026W/NTIRE/html/Gushchin_NTIRE_2026_Challenge_on_Robust_AI-Generated_Image_Detection_in_the_CVPRW_2026_paper.html
    [17] C. Tan et al., “Rethinking the up-sampling operations in CNN-based generative network for generalizable deepfake detection,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), pp. 28130–28139, 2024. https://doi.org/10.1109/CVPR52733.2024.02657
    [18] Z. Wang et al., “DIRE for diffusion-generated image detection,” in Proc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV), pp. 22388–22398, 2023. https://doi.org/10.1109/ICCV51070.2023.02051
    [19] J. Ricker, D. Lukovnikov, and A. Fischer, “AEROBLADE: Training-free detection of latent diffusion images using autoencoder reconstruction error,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), pp. 9130–9140, 2024. https://doi.org/10.1109/CVPR52733.2024.00872
    [20] Y. Ju, S. Jia, L. Ke, H. Xue, K. Nagano, and S. Lyu, “Fusing global and local features for generalized AI-synthesized image detection,” in Proc. IEEE Int. Conf. Image Process. (ICIP), pp. 3465–3469, 2022. https://doi.org/10.1109/ICIP46576.2022.9897820
    [21] Y. Ju, S. Jia, J. Cai, H. Guan, and S. Lyu, “GLFF: Global and local feature fusion for AI-synthesized image detection,” IEEE Trans. Multimedia, vol. 26, pp. 4073–4085, 2024. https://doi.org/10.1109/TMM.2023.3313503
    [22] U. Ojha, Y. Li, and Y. J. Lee, “Towards universal fake image detectors that generalize across generative models,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), pp. 24480–24489, 2023. https://doi.org/10.1109/CVPR52729.2023.02345
    [23] S. Yan et al., “A sanity check for AI-generated image detection,” in Proc. Int. Conf. Learn. Represent. (ICLR), pp. 70702–70720, 2025. https://proceedings.iclr.cc/paper_files/paper/2025/hash/b0303773962ea1b5394c3a83cc7dd066-Abstract-Conference.html
    [24] S. Teerapittayanon, B. McDanel, and H. T. Kung, “BranchyNet: Fast inference via early exiting from deep neural networks,” in Proc. 23rd Int. Conf. Pattern Recognit. (ICPR), pp. 2464–2469, 2016. https://doi.org/10.1109/ICPR.2016.7900006
    [25] C. Guo, G. Pleiss, Y. Sun, and K. Q. Weinberger, “On calibration of modern neural networks,” in Proc. 34th Int. Conf. Mach. Learn. (ICML), Proc. Mach. Learn. Res., vol. 70, pp. 1321–1330, 2017. https://proceedings.mlr.press/v70/guo17a.html
    [26] R. Schwartz, J. Dodge, N. A. Smith, and O. Etzioni, “Green AI,” Commun. ACM, vol. 63, no. 12, pp. 54–63, 2020. https://doi.org/10.1145/3381831
    [27] M. Zhu et al., “GenImage: A million-scale benchmark for detecting AI-generated image,” in Adv. Neural Inf. Process. Syst., vol. 36, Datasets and Benchmarks Track, pp. 77771–77782, 2023. https://doi.org/10.52202/075280-3398
    [28] Z. Li et al., “Is artificial intelligence generated image detection a solved problem?” in Adv. Neural Inf. Process. Syst., vol. 38, Datasets and Benchmarks Track, 2025. https://papers.nips.cc/paper_files/paper/2025/hash/fb693c67f61e5321746ffce8b6fdd2d0-Abstract-Datasets_and_Benchmarks_Track.html
    [29] S. Gye, J. Ko, H. Shon, M. Kwon, and J. Kim, “Reducing the content bias for AI-generated image detection,” in Proc. IEEE/CVF Winter Conf. Appl. Comput. Vis. (WACV), pp. 399–408, 2025. https://doi.org/10.1109/WACV61041.2025.00049
    [30] B. Efron and R. J. Tibshirani, An Introduction to the Bootstrap. New York, NY, USA: Chapman & Hall, 1993.
    [31] M. Mitchell et al., “Model cards for model reporting,” in Proc. Conf. Fairness, Accountability, Transparency (FAT*), pp. 220–229, 2019. https://doi.org/10.1145/3287560.3287596
    [32] R. R. Selvaraju et al., “Grad-CAM: Visual explanations from deep networks via gradient-based localization,” in Proc. IEEE Int. Conf. Comput. Vis. (ICCV), pp. 618–626, 2017. https://doi.org/10.1109/ICCV.2017.74
    [33] Z. Ye and F. Farnia, “Gaussian smoothing in saliency maps: The stability-fidelity trade-off in neural network interpretability,” in Proc. 28th Int. Conf. Artif. Intell. Stat. (AISTATS), Proc. Mach. Learn. Res., vol. 258, pp. 2125–2133, 2025. https://proceedings.mlr.press/v258/ye25a.html
    [34] V. Petsiuk, A. Das, and K. Saenko, “RISE: Randomized input sampling for explanation of black-box models,” in Proc. Brit. Mach. Vis. Conf. (BMVC), paper no. 151, pp. 1–13, 2018. https://bmva-archive.org.uk/bmvc/2018/contents/papers/1064.pdf
    [35] M. D. Zeiler and R. Fergus, “Visualizing and understanding convolutional networks,” in Computer Vision—ECCV 2014, Part I, Lecture Notes in Computer Science, vol. 8689, pp. 818–833, 2014. https://doi.org/10.1007/978-3-319-10590-1_53

    QR CODE
    :::