| 研究生: |
黃柏憲 Huang, Bo-Xian |
|---|---|
| 論文名稱: |
基於對比學習的水下顯著目標檢測 Underwater Salient Object Detection based on Contrastive Learning |
| 指導教授: |
彭彥璁
Yan-Tsung Peng |
| 口試委員: |
紀明德
Ming-Te Chi 黃士嘉 Shih-Chia Huang |
| 學位類別: |
碩士
Master |
| 系所名稱: |
資訊學院 - 資訊科學系碩士在職專班 Excutive Master Program of Computer Science |
| 論文出版年: | 2026 |
| 畢業學年度: | 114 |
| 語文別: | 中文 |
| 論文頁數: | 67 |
| 中文關鍵詞: | 水下顯著目標檢測 、對比表徵學習 、InfoNCE 、特徵層級正則化 、視覺基礎模型 |
| 外文關鍵詞: | Underwater Salient Object Detection, Contrastive Representation Learning, InfoNCE, Feature-level Regularisation, Visual Foundation Model |
| 相關次數: | 點閱:15 下載:0 |
| 分享至: |
| 查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報 |
水下顯著目標檢測(Underwater Salient Object Detection, USOD)由於水體對光線之選擇性吸收與後向散射,前景物體之色彩與對比受到衰減,與背景水色在低層次特徵上之可分性顯著下降。在純粹像素級分割監督下,編碼器之特徵表徵是否能在前景區域內穩定地依賴前景像素並無顯式約束。本論文提出AquaCL(Aquatic Contrastive Learning)——一套以平均色構造正負樣本(Easy-Fill)之跨架構統一對比學習機制:以背景平均色填補背景可得正樣本、填補前景可得簡單負樣本,並於編碼器最粗特徵上施加InfoNCE 約束,作為像素級分割損失之外的特徵層級正則化。所提機制以加法式設計同時植入五個不同架構之既有顯著目標檢測模型——Dual-SAM、HFUSOD、WaterFlow、TC-USOD 與RMFormer;於USOD10K 主測與USOD-300、UFO-120 兩個跨資料集評估上,所提機制相較於各自Origin 在絕大多數對照中維持或改善MAE、最大F-measure、S-measure 與最大E-measure 等四項指標;所提機制不取代任何基底之原始骨幹與分割損失,僅以輕量之LoRA / adapter 與投影頭作加法式擴充。本論文亦討論:簡單合成的負樣本相對於LaMa、MAT 等生成式修補方案具備實作簡潔與每epoch 構造確定之雙重優勢;對比約束須限制於最粗特徵層方可避免破壞解碼器之空間重建路徑。
Underwater Salient Object Detection (USOD) suffers from reduced low-level discriminability between foreground objects and the surrounding water-body colour, caused by wavelength-dependent absorption and backscatter that attenuate foreground colour and compress foreground-background contrast. Under purely pixellevel segmentation supervision, there is no explicit constraint that representations within the foreground region must depend on foreground pixels. This thesis proposes AquaCL (Aquatic Contrastive Learning), a unified, cross-architecture contrastivelearning framework centred on an Easy-Fill construction that symmetrically fills the foreground or the background with the mean background colour to produce positive and easy-negative samples, and imposes an InfoNCE constraint at the coarsest encoder feature as a feature-level regulariser in addition to the pixel-level segmentation loss. The proposed framework is plugged into five salient-object-detection baselines with different architectures —Dual-SAM, HFUSOD, WaterFlow, TCUSOD, and RMFormer —in an additive fashion. Systematic experiments on USOD10K and on cross-dataset evaluations (USOD-300, UFO-120) show that the proposed framework maintains or improves the MAE, maximum-F, S-measure and maximum-E scores relative to each baseline in the vast majority of comparisons, leaving each baseline’s backbone and original segmentation loss intact and adding only lightweight LoRA / adapter modules and a projection head. Ablations further discuss why simple colour-only negatives remain preferable to LaMa / MAT inpainted alternatives, and why the contrastive constraint must be confined to the coarsest feature layer to avoid corrupting the decoder’s spatial reconstruction path.
誌謝 i
摘要 ii
Abstract iii
目錄 v
圖目錄 viii
表目錄 xii
第一章 緒論 1
第一節 研究背景 1
第二節 研究動機與目的 2
第二章 相關工作 5
第一節 章節概覽 5
第二節 顯著目標檢測 5
一、 偽裝目標檢測與其遷移 6
第三節 水下顯著目標檢測 6
一、 水下成像之物理挑戰 6
二、 USOD10K與相關資料集 7
三、 本論文採用之五個基底模型 8
四、 像素級分割監督之特徵層級盲區 11
第四節 對比表徵學習 12
一、 由實例判別到InfoNCE 12
二、 監督式對比學習與密集預測 12
三、 負樣本構造:從複雜到簡單 13
第五節 視覺基礎模型與物理先驗適配 14
一、 SegmentAnythingModel之下游適配 14
二、 修正流與條件式生成式分割 15
三、 水下視覺中之物理先驗整合 15
第六節 本論文之研究定位 16
第三章 研究方法 17
第一節 章節 概覽 17
第二節 基底模型I:Dual-SAM 17
第三節 基底模型II:WaterFlow 18
第四節 基底模型III:HFUSOD 19
第五節 基底模型IV:TC-USOD 20
第六節 基底模型V:RMFormer 21
第七節 統一的AquaCL對比學習框架 22
一、 正負樣本構造(Easy-Fill) 22
二、 最粗尺度接入點與投影頭 24
三、 InfoNCE損失與總體訓練目標 25
第八節 在五個基底模型上的具體接入 26
一、 Anchor梯度回傳策略 27
二、 Dual-SAM接入 28
三、 WaterFlow接入 28
四、 HFUSOD接入 29
五、 TC-USOD接入(微調模式) 30
六、 RMFormer接入(微調模式) 30
七、 五模型一致性比較 30
第九節 訓練設定與超參數 31
第十節 本章小結 32
第四章 實驗 34
第一節 實驗設定 34
一、 資料集 34
二、 評估指標 35
三、 實作細節 35
第二節 主結果 36
一、 USOD10K主測試集 36
第三節 跨域泛化 37
一、 五基底之USOD-300與UFO-120評估 37
二、 USOD-300跨資料集評估(TC-USOD) 39
三、 UFO-120延伸跨域評估(TC-USOD )41
第四節 定性分析 41
一、 Dual-SAM定性結果 42
二、 HFUSOD定性結果 44
三、 WaterFlow定性結果 46
四、 TC-USOD定性結果 48
五、 RMFormer定性結果 50
第五節 消融研究 52
一、 設定 52
二、 兩種編碼器之Origin對照 52
三、 元件與填補構造消融 53
四、 生成式填補之對照(LaMa/MAT) 54
五、 溫度τ之敏感度 55
六、 接入尺度之選擇 56
第五章 結論與未來工作 57
第一節 結論 57
第二節 未來工作:水色估計之改進 58
參考文獻 60
[1] L. Hong, X. Wang, G. Zhang, and M. Zhao, “USOD10K: A new benchmark dataset for underwater salient object detection,” IEEE Transactions on Image Processing, vol. 34, pp.1602–1615, 2025.
[2] P. Zhang, T. Yan, Y. Liu, and H. Lu, “Fantastic animals and where to find them: Segment any marine animal with dual SAM,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024, pp. 2578–2587.
[3] W. Huang and D. Zhu, “HFUSOD: Hierarchical fusion of Swin Transformer with CNN network for underwater salient object detection,” IET Image Processing, vol. 20, no. 1, p.e70331, 2026.
[4] R. Li, S. Lian, H. Li, Y. Li, W. Wu, and S. Kwong, “WaterFlow: Explicit physics-prior rectified flow for underwater saliency mask generation,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2026.
[5] X. Deng, P. Zhang, W. Liu, and H. Lu, “Recurrent multi-scale transformer for highresolution salient object detection,” in ACM International Conference on Multimedia (ACM MM), 2023.
[6] Z. Zhong, L. Zheng, G. Kang, S. Li, and Y. Yang, “Random erasing data augmentation,” in AAAI Conference on Artificial Intelligence, 2020, pp. 13001–13008.
[7] A. Kirillov, E. Mintun, N. Ravi, H. Mao, C. Rolland, L. Gustafson, T. Xiao, S. Whitehead, A. C. Berg, W.-Y. Lo, P. Dollár, and R. Girshick, “Segment anything,” in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2023.
[8] E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen, “LoRA: Low-rank adaptation of large language models,” in International Conference on Learning Representations (ICLR), 2022.
[9] X. Liu, C. Gong, and Q. Liu, “Flow straight and fast: Learning to generate and transfer data with rectified flow,” in International Conference on Learning Representations (ICLR),2023.
[10] D. Akkaynak and T. Treibitz, “Sea-thru: A method for removing water from underwater images,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2019.
[11] H. Dong, J. Wu, C. Xing, H. Xi, H. Cui, and J. Zhu, “Treasure in the background: Improve saliency object detection by self-supervised contrast learning,” Expert Systems with Applications, vol. 267, p. 126244, 2025.
[12] R. Suvorov, E. Logacheva, A. Mashikhin, A. Remizova, A. Ashukha, A. Silvestrov, N. Kong, H. Goka, K. Park, and V. Lempitsky, “Resolution-robust large mask inpainting with Fourier convolutions,” in Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), 2022, pp. 2149–2159.
[13] W. Li, Z. Lin, K. Zhou, L. Qi, Y. Wang, and J. Jia, “MAT: Mask-aware transformer for large hole image inpainting,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022, pp. 10758–10768.
[14] A. van den Oord, Y. Li, and O. Vinyals, “Representation learning with contrastive predictive coding,” arXiv preprint arXiv:1807.03748, 2018.
[15] J. Harel, C. Koch, and P. Perona, “Graph-based visual saliency,” in Advances in Neural Information Processing Systems (NIPS), vol. 19, 2006.
[16] F. Perazzi, P. Krähenbühl, Y. Pritch, and A. Hornung, “Saliency filters: Contrast based filtering for salient region detection,” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2012.
[17] J. Long, E. Shelhamer, and T. Darrell, “Fully convolutional networks for semantic segmentation,” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR),2015.
[18] S. Xie and Z. Tu, “Holistically-nested edge detection,” in IEEE International Conference on Computer Vision (ICCV), 2015.
[19] L. Wang, H. Lu, Y. Wang, M. Feng, D. Wang, B. Yin, and X. Ruan, “Learning to detect salient objects with image-level supervision,” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017.
[20] C. Yang, L. Zhang, H. Lu, X. Ruan, and M.-H. Yang, “Saliency detection via graph-based manifold ranking,” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2013.
[21] G. Li and Y. Yu, “Visual saliency based on multiscale deep features,” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2015.
[22] J. Shi, Q. Yan, L. Xu, and J. Jia, “Hierarchical image saliency detection on extended CSSD,” IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI), vol. 38, no. 4, pp. 717–729, 2016.
[23] Y. Li, X. Hou, C. Koch, J. M. Rehg, and A. L. Yuille, “The secrets of salient object segmentation,” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2014.
[24] Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo, “Swin transformer: Hierarchical vision transformer using shifted windows,” in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2021.
[25] W. Wang, E. Xie, X. Li, D.-P. Fan, K. Song, D. Liang, T. Lu, P. Luo, and L. Shao, “Pyramid vision transformer: A versatile backbone for dense prediction without convolutions,” in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2021.
[26] ——, “PVTv2: Improved baselines with pyramid vision transformer,” Computational Visual Media, 2022.
[27] C. Ryali, Y.-T. Hu, D. Bolya, C. Wei, H. Fan, P.-Y. Huang, V. Aggarwal, A. Chowdhury, O. Poursaeed, J. Hoffman, J. Malik, Y. Li, and C. Feichtenhofer, “Hiera: A hierarchical vision transformer without the bells-and-whistles,” in International Conference on Machine Learning (ICML), 2023.
[28] X. Qin, Z. Zhang, C. Huang, C. Gao, M. Dehghan, and M. Jagersand, “BASNet:
Boundary-aware salient object detection,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2019.
[29] J. Wei, S. Wang, and Q. Huang, “F3Net: Fusion, feedback and focus for salient object detection,” Proceedings of the AAAI Conference on Artificial Intelligence, vol. 34, no. 07, pp. 12321–12328, Apr. 2020.
[30] M. Zhuge, D.-P. Fan, N. Liu, D. Zhang, D. Xu, and L. Shao, “Salient object detection via integrity learning,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 45, no. 3, pp. 3738–3752, 2023.
[31] N. Liu, N. Zhang, K. Wan, J. Han, and L. Shao, “Visual saliency transformer,” in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2021.
[32] Y. K. Yun and W. Lin, “SelfReformer: Self-refined network with transformer for salient object detection,” arXiv preprint arXiv:2205.11283, 2022.
[33] F. Milletari, N. Navab, and S.-A. Ahmadi, “V-Net: Fully convolutional neural networks for volumetric medical image segmentation,” in International Conference on 3D Vision (3DV), 2016.
[34] M. A. Rahman and Y. Wang, “Optimizing intersection-over-union in deep neural networks for image segmentation,” in International Symposium on Visual Computing (ISVC), 2016.
[35] Z. Yang, S. Soltanian-Zadeh, and S. Farsiu, “BiconNet: An edge-preserved connectivitybased approach for salient object detection,” Pattern Recognition, vol. 121, p. 108231, 2022.
[36] D.-P. Fan, G.-P. Ji, G. Sun, M.-M. Cheng, J. Shen, and L. Shao, “Camouflaged object detection,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2020.
[37] H. Mei, G.-P. Ji, Z. Wei, X. Yang, X. Wei, and D.-P. Fan, “Camouflaged object segmentation with distraction mining,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2021.
[38] Z. Chen, K. Sun, X. Lin, and R. Ji, “CamoDiffusion: Camouflaged object detection via conditional diffusion models,” in AAAI Conference on Artificial Intelligence, vol. 38, no. 2, 2024, pp. 1272–1280.
[39] J. Ho, A. Jain, and P. Abbeel, “Denoising diffusion probabilistic models,” in Advances in Neural Information Processing Systems (NeurIPS), 2020.
[40] K. He, J. Sun, and X. Tang, “Single image haze removal using dark channel prior,” IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI), 2010.
[41] R. Ranftl, A. Bochkovskiy, and V. Koltun, “Vision transformers for dense prediction,” in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2021.
[42] M. J. Islam, P. Luo, and J. Sattar, “Simultaneous enhancement and super-resolution of underwater imagery for improved visual perception,” in Robotics: Science and Systems (RSS), 2020.
[43] M. J. Islam, R. Wang, and J. Sattar, “SVAM: Saliency-guided visual attention modeling by autonomous underwater robots,” in Robotics: Science and Systems (RSS), 2022, arXiv:2011.06252.
[44] M. Xu, J. Su, and Y. Liu, “AquaSAM: Underwater image foreground segmentation,” arXiv preprint arXiv:2308.04218, 2023.
[45] S. Lian, Z. Zhang, H. Li, W. Li, L. T. Yang, S. Kwong, and R. Cong, “Diving into underwater: Segment anything model guided underwater salient instance segmentation and a large-scale dataset,” in International Conference on Machine Learning (ICML), 2024.
[46] T. Chen, L. Zhu, C. Deng, R. Cao, Y. Wang, S. Zhang, Z. Li, L. Sun, Y. Zang, and P. Mao, “SAM-Adapter: Adapting segment anything in underperformed scenes,” in Proceedings of the IEEE/CVF International Conference on Computer Vision Workshops (ICCVW), 2023.
[47] R. Cong, W. Yang, W. Zhang, C. Li, C.-L. Guo, Q. Huang, and S. Kwong, “PUGAN: Physical model-guided underwater image enhancement using GAN with dual-discriminators,” IEEE Transactions on Image Processing, vol. 32, pp. 4472–4485, 2023.
[48] W. Chen, Y. Lei, S. Luo, Z. Zhou, M. Li, and C.-M. Pun, “UWFormer: Underwater image enhancement via a semi-supervised multi-scale transformer,” arXiv preprint arXiv:2310.20210, 2024.
[49] Z. Wu, Y. Xiong, S. X. Yu, and D. Lin, “Unsupervised feature learning via non-parametric instance discrimination,” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018.
[50] T. Chen, S. Kornblith, M. Norouzi, and G. Hinton, “A simple framework for contrastive learning of visual representations,” in International Conference on Machine Learning (ICML), 2020.
[51] K. He, H. Fan, Y. Wu, S. Xie, and R. Girshick, “Momentum contrast for unsupervised visual representation learning,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2020.
[52] X. Chen, H. Fan, R. Girshick, and K. He, “Improved baselines with momentum contrastive learning,” arXiv preprint arXiv:2003.04297, 2020.
[53] J.-B. Grill, F. Strub, F. Altché, C. Tallec, P. H. Richemond, E. Buchatskaya, C. Doersch, B. A. Pires, Z. D. Guo, M. G. Azar et al., “Bootstrap your own latent: A new approach to self-supervised learning,” in Advances in Neural Information Processing Systems (NeurIPS), 2020.
[54] X. Chen and K. He, “Exploring simple siamese representation learning,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2021.
[55] P. Khosla, P. Teterwak, C. Wang, A. Sarna, Y. Tian, P. Isola, A. Maschinot, C. Liu, and D. Krishnan, “Supervised contrastive learning,” in Advances in Neural Information Processing Systems (NeurIPS), 2020.
[56] W. Wang, T. Zhou, F. Yu, J. Dai, E. Konukoglu, and L. Van Gool, “Exploring cross-image pixel contrast for semantic segmentation,” in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2021.
[57] S. Liu, S. Zhi, E. Johns, and A. J. Davison, “Bootstrapping semantic segmentation with regional contrast (ReCo),” in International Conference on Learning Representations (ICLR), 2022.
[58] Q. Yan, X. Du, C. Li, and X. Tian, “CLIB: Contrastive learning of ignoring background for underwater fish image classification,” Frontiers in Neurorobotics, vol. 18, p. 1423848, 2024.
[59] J. Robinson, C.-Y. Chuang, S. Sra, and S. Jegelka, “Contrastive learning with hard negative samples,” in International Conference on Learning Representations (ICLR), 2021.
[60] Y. Kalantidis, M. B. Sariyildiz, N. Pion, P. Weinzaepfel, and D. Larlus, “Hard negative mixing for contrastive learning (MoCHi),” in Advances in Neural Information Processing Systems (NeurIPS), 2020.
[61] T. Huynh, S. Kornblith, M. R. Walter, M. Maire, and M. Khademi, “Boosting contrastive self-supervised learning with false negative cancellation,” in Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), 2022.
[62] R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer, “High-resolution image synthesis with latent diffusion models,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022, pp. 10684–10695.
[63] J. Cheng, J. Ye, Z. Deng, J. Chen, T. Li, H. Wang, Y. Su, Z. Huang, J. Chen, L. Jiang, H. Sun, J. He, S. Zhang, M. Zhu, and Y. Qiao, “SAM-Med2D,” arXiv preprint arXiv:2308.16184, 2023.
[64] D. Baranchuk, I. Rubachev, A. Voynov, V. Khrulkov, and A. Babenko, “Label-efficient semantic segmentation with diffusion models,” in International Conference on Learning Representations (ICLR), 2022.
[65] T. Amit, T. Shaharbany, E. Nachmani, and L. Wolf, “SegDiff: Image segmentation with diffusion probabilistic models,” arXiv preprint arXiv:2112.00390, 2021.
[66] J. Wu, H. Fang, Y. Zhang, Y. Yang, and Y. Xu, “MedSegDiff: Medical image segmentation with diffusion probabilistic model,” arXiv preprint arXiv:2211.00611, 2022.
[67] C. Li, S. Anwar, J. Hou, R. Cong, C. Guo, and W. Ren, “Underwater image enhancement via medium transmission-guided multi-color space embedding,” IEEE Transactions on Image Processing, vol. 30, pp. 4985–5000, 2021.
[68] C. Li, C. Guo, W. Ren, R. Cong, J. Hou, S. Kwong, and D. Tao, “An underwater image enhancement benchmark dataset and beyond,” IEEE Transactions on Image Processing (TIP), 2020.
[69] L. Yang, B. Kang, Z. Huang, Z. Zhao, X. Xu, J. Feng, and H. Zhao, “Depth anything V2,” in Advances in Neural Information Processing Systems (NeurIPS), 2024.