| 研究生: |
陳凱輝 Chen, Kai-Hui |
|---|---|
| 論文名稱: |
基於語意引導的俯視影像多任務學習動態損失調度方法 Semantic-Guided Dynamic Loss Scheduling for Multi-Task Learning in Overhead Imagery |
| 指導教授: |
廖文宏
Liao, Wen-Hung |
| 口試委員: |
彭彥璁
Peng, Yan-Tsung 陳駿丞 Chen, Jun-Cheng |
| 學位類別: |
碩士
Master |
| 系所名稱: |
資訊學院 - 資訊科學系 Department of Computer Science |
| 論文出版年: | 2026 |
| 畢業學年度: | 114 |
| 語文別: | 中文 |
| 論文頁數: | 111 |
| 中文關鍵詞: | 俯視影像 、多任務學習 、DINOv2 、動態損失調度 、語意分割 、深度估計 、物件偵測 、AirSim 、合成資料集 |
| 外文關鍵詞: | Overhead Imagery, Multi-Task Learning, DINOv2, Dynamic Loss Scheduling, Semantic Segmentation, Depth Estimation, Object Detection, AirSim, Synthetic Dataset |
| 相關次數: | 點閱:39 下載:1 |
| 分享至: |
| 查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報 |
隨著無人機技術普及,俯視影像(Overhead Imagery)的即時感知能力越發重要。然而,俯視視角缺乏傳統透視幾何線索,且多任務學習(Multi-Task Learning, MTL)常面臨任務間梯度競爭與權重失衡的瓶頸。為此,本研究提出一套以 DINOv2 為骨幹網路的「語意優先動態損失調度(Semantic-Guided Dynamic Loss Scheduling, SGC-DOLS)」框架。此處「語意引導」並非傳統特徵層級之空間先驗注入,而是將語意任務視為「損失層級之排程最佳化支點(Loss-level Scheduling Pivot)」。相較於現有動態加權策略多屬被動式回應訓練,本研究引入「時序課程式損失調度」概念,建構雙層權重調整機制:在宏觀輪次(macro-epoch)層級將語意分割設定為優先參照,隨訓練推進有序釋放特徵空間;並在輪次末期(epoch-end)依據絕對損失量級落差,週期性動態補償進度落後之幾何任務。我們假設語意特徵可作為較為穩定的共享資源池,透過此主動式課程設計,期能有效緩解語意分割、單目深度估計與旋轉物件偵測在聯合訓練時的資源賽局。
在資料工程方面,針對高品質航拍標註資料稀缺且難以對齊的困境,我們基於 Microsoft AirSim 模擬環境建置自動化生成管線,利用正規表示式映射與強制渲染同步機制,產出像素級精確對齊的多模態合成資料集。經 36 組消融與對照實驗,結果發現傳統損失量級校準會引發「梯度壓縮陷阱(Gradient Compression Trap, GCT)」,在正確排除該機制後,最終配置(Ours)於所有對照組中取得最佳的深度估計絕對相對誤差(AbsRel),數值達 0.1975,較固定等權重基準(0.2293)降低 0.0318。為因應不同部署需求,本系統引入雙指標評估框架:在強調幾何任務的測繪導向之幾何精度分數(Geometric Precision Score, GeoScore)上,最終配置以 1.0447 取得最佳表現(優於同條件之不確定性加權版本的 1.0310),且為本文所比較之設定中,唯一達到優於 0.20 深度誤差門檻之設定;而在通用均衡分數(Balance Score, BS)上,不確定性加權版本則以 1.0318 略勝一籌(Ours 為 1.0201),清晰描繪出相異策略在 Pareto 效率前沿上的取捨關係。
此外,與單任務基線(Single-Task Baseline)的對照結果顯示出多任務架構下的正向遷移效應:在聯合訓練下,深度估計 AbsRel 進步 0.2141,物件偵測 mAP@50 提升 0.0322,而語意分割僅付出 0.0696 mIoU 的效能代價。此量化證據在本研究之合成資料設定下,支持了在 DINOv2 強韌表徵下,將語意特徵視為穩定資源以輔助幾何任務學習的方法論價值。
As Unmanned Aerial Vehicle (UAV) technology proliferates, real-time perception in overhead imagery becomes increasingly critical. However, overhead perspectives lack traditional linear perspective cues, and Multi-Task Learning (MTL) often suffers from gradient competition and task weight imbalance. To address these bottlenecks, this study proposes a Semantic-Guided Dynamic Loss Scheduling (SGC-DOLS) framework using a DINOv2 backbone network. In this work, Semantic-Guided refers to treating semantic segmentation as the optimization pivot for loss scheduling rather than traditional feature-level guidance. Unlike existing dynamic weighting strategies that reactively adapt to training dynamics, this framework introduces a temporal curriculum for loss scheduling based on a dual-level weight adjustment mechanism. At the macro-epoch level, semantic segmentation is designated as a prior reference, progressively reallocating optimization capacity for geometric tasks as training proceeds. At the epoch-end level, an absolute magnitude signal is employed to periodically compensate for lagging geometric tasks. Under the hypothesis that semantic features function as a relatively stable shared resource pool, this proactive curriculum design effectively mitigates the optimization conflicts among semantic segmentation, monocular depth estimation, and oriented object detection during joint training.
To overcome the scarcity and alignment difficulties of high-quality aerial annotations, an automated data generation pipeline is established using the Microsoft AirSim simulator, producing a pixel-perfect aligned multi-task synthetic dataset via regular expression mapping and forced rendering synchronization. Through 36 sets of ablation and comparative experiments, this study identifies a ”gradient compression trap” (GCT) induced by traditional loss scale auto-calibration. After properly discarding this mechanism, the final configuration (Ours) yields the best performance with a depth estimation absolute relative error (AbsRel) of 0.1975, achieving a reduction of 0.0318 compared to the fixed equal-weight baseline (0.2293). To accommodate distinct deployment requirements, a dual-metric evaluation framework is introduced. Particularly, Geometric Precision Score (GeoScore) is proposed as a deployment-oriented metric emphasizing spatial geometric tasks. Under this metric, the final configuration achieves the best performance with a score of 1.0447 (outperforming the non-calibrated uncertainty weighting baseline of 1.0310) and remains the only configuration among the evaluated settings to achieve a depth AbsRel below the 0.20 threshold. Conversely, under the Balance Score (BS) for overall multi-task equilibrium, the uncertainty weighting baseline slightly leads with 1.0318 compared to Ours (1.0201), clearly illustrating the trade-offs of varying strategies on the Pareto front.
Furthermore, comparisons against single-task baselines indicate a measurable positive transfer within the multi-task architecture: joint training achieves an improvement of 0.2141 in depth estimation AbsRel and an increase of 0.0322 in object detection mAP@50, with a minor performance compromise of 0.0696 in semantic segmentation mIoU. Under the synthetic data setting of this study, these quantitative findings support the methodological value of leveraging semantic features as a stable resource to assist geometric learning under the robust representation of DINOv2.
謝誌 I
摘要 II
Abstract IV
目錄 VI
表次 XI
圖次 XII
第一章 緒論 1
1.1 研究背景與動機 1
1.2 研究目的與核心假設 3
1.3 主要貢獻 4
1.4 論文架構 6
第二章 相關研究與技術背景 7
2.1 視覺基礎模型 7
2.1.1 Vision Transformer 之演進 7
2.1.2 自監督學習與 DINOv2 8
2.2 多任務學習與動態損失調度 9
2.2.1 多任務學習的參數共享機制 9
2.2.2 損失加權策略 10
2.2.3 梯度調控策略 12
2.2.4 多任務評估方法論 13
2.2.5 現有 MTL 方法之缺口與本研究定位 14
2.3 俯視影像的密集預測 15
2.3.1 語意分割 16
2.3.2 單目深度估計 16
2.3.3 旋轉物件偵測 16
2.4 合成資料與模擬環境 17
2.5 本章小結 18
第三章 研究方法與架構設計 19
3.1 合成資料生成管線 19
3.1.1 模擬環境設定與渲染最佳化 19
3.1.2 語意映射與標註生成 20
3.1.3 智慧採集策略 23
3.2 網路架構設計 25
3.2.1 特徵擷取骨幹網路 26
3.2.2 高解析度空間旁路編碼器 27
3.2.3 特徵融合頸部網路 28
3.2.4 多任務解碼頭 28
3.3 多任務損失函數 30
3.3.1 總損失定義 31
3.3.2 語意分割損失 32
3.3.3 深度估計損失 32
3.3.4 物件偵測損失 33
3.4 語意優先雙層動態損失調度機制(SGC-DOLS) 35
3.4.1 損失量級校準之潛阱與主動捨棄 36
3.4.2 策略一:語意優先課程式退火 37
3.4.3 策略二:難度差距補位與訊號機制改良 37
3.4.4 階層式架構視角:SGC-DOLS 的主動與被動機制 40
3.4.5 最終參數配置(Ours) 41
3.5 本章小結 42
第四章 實驗結果與分析 43
4.1 實驗設定與評估指標 43
4.1.1 資料集與類別設定 43
4.1.2 訓練環境與超參數細節 44
4.1.3 多任務評估指標 44
4.2 核心消融實驗矩陣設計 46
4.3 核心因子設計結果 47
4.4 超參數最佳化與訊號機制消融 49
4.4.1 超參數敏感度分析 50
4.4.2 訊號機制改良:從 Ratio 到 Magnitude 51
4.4.3 Magnitude 訊號之幾何實質突破 53
4.4.4 保護機制之過度約束驗證 53
4.4.5 最終最佳配置(Ours)之確立 54
4.4.6 跨策略同條件對照與 Pareto 前沿分析 55
4.4.7 實體損失收斂軌跡與最佳化行為驗證 61
4.5 校準—調度耦合效應與時序消融觀測 62
4.6 單任務基線比較與正向遷移效應 63
4.7 定性視覺化與空間分布分析 67
4.8 本章小結 70
第五章 討論 72
5.1 核心發現與設計規則 72
5.1.1 發現一:語意分支作為穩定資源池 72
5.1.2 發現二:動態調度的三項必要條件 73
5.1.3 設計空間邊界的方法論意義 74
5.2 同條件對照下的調度策略定位 80
5.2.1 時序課程式調度之文獻定位 81
5.3 校準耦合陷阱之歸因分析 81
5.3.1 梯度量級干擾與梯度壓縮陷阱 82
5.3.2 共享骨幹網路之表徵容量限制 83
5.3.3 動態調度與靜態正規化之結構性互斥 84
5.4 多任務正向遷移效應 85
5.5 機制組合的非加性交互效應 87
5.6 本章小結 88
第六章 結論與未來工作 90
6.1 研究總結 90
6.2 核心研究發現 91
6.3 主要貢獻 92
6.4 研究限制 93
6.5 未來工作方向 96
參考文獻 99
附錄 A 多任務學習完整實驗結果總表 105
附錄 B 任務類別效能細部解析 107
B.1 各語意類別分割效能 107
B.2 各偵測類別效能 108
附錄 C 網路架構演進與設計決策 110
C.1 架構演進歷程 110
C.2 各分支架構量化比較 111
[1] Aleksei Bochkovskii, Amael Delaunoy, Hugo Germain, Marcel Santos, Yichao Zhou, Stephan R. Richter, and Vladlen Koltun. Depth pro: Sharp monocular metric depth in less than a second. In International Conference on Learning Representations (ICLR), 2025.
[2] Christian Bohn, Ido Freeman, Hasan Tercan, and Tobias Meisen. Task weighting through gradient projection for multitask learning. In arXiv preprint arXiv:2409.01793, 2024.
[3] Christopher F. Brown, Michal R. Kazmierski, Valerie J. Pasquarella, William J. Rucklidge, Masha Samsikova, Chenhui Zhang, Evan Shelhamer, Estefania Lahera, Olivia Wiles, Simon Ilyushchenko, Noel Gorelick, Lihui Lydia Zhang, Sophia Alj, Emily Schechter, Sean Askay, Oliver Guinan, Rebecca Moore, Alexis Boukouvalas, and Pushmeet Kohli. AlphaEarth Foundations: An embedding field model for accurate and efficient global mapping from sparse label data. arXiv preprint arXiv:2507.22291, 2025.
[4] Mathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou, Julien Mairal, Piotr Bojanowski, and Armand Joulin. Emerging properties in self-supervised vision transformers. In Proceedings of the IEEE/CVF international conference on computer vision, pages 9650–9660, 2021.
[5] Rich Caruana. Multitask learning. Machine Learning, 28(1):41–75, 1997.
[6] Zhan Chen, Yidan Zhang, Xiyu Qi, Yongqiang Mao, Xin Zhou, Lulu Niu, Hui Wu, Lei Wang, and Yunping Ge. HeightFormer: A multilevel interaction and image-adaptive classification-regression network for monocular height estimation with aerial images. Remote Sensing, 16(2):295, 2024.
[7] Zhao Chen, Vijay Badrinarayanan, Chen-Yu Lee, and Andrew Rabinovich. GradNorm: Gradient normalization for adaptive loss balancing in deep multitask networks. In International conference on machine learning (ICML), pages 794–803. PMLR, 2018.
[8] Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale. In International Conference on Learning Representations (ICLR), 2021.
[9] David Eigen, Christian Puhrsch, and Rob Fergus. Depth map prediction from a single image using a multi-scale deep network. Advances in Neural Information Processing Systems, 27, 2014.
[10] Michelle Guo, Albert Haque, De-An Huang, Serena Yeung, and Li Fei-Fei. Dynamic task prioritization for multitask learning. In Proceedings of the European Conference on Computer Vision (ECCV), pages 282–299, 2018.
[11] Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollár, and Ross Girshick. Masked autoencoders are scalable vision learners. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 16000–16009, 2022.
[12] Alex Kendall, Yarin Gal, and Roberto Cipolla. Multi-task learning using uncertainty to weigh losses for scene geometry and semantics. In Proceedings of the IEEE conference on computer vision and pattern recognition (CVPR), pages 7482–7491, 2018.
[13] Sahil Khose, Anisha Pal, Aayushi Agarwal, Deepanshi, Judy Hoffman, and Prithvijit Chattopadhyay. SKYSCENES: A synthetic dataset for aerial scene understanding. In Proceedings of the European Conference on Computer Vision (ECCV), 2024. doi: 10.1007/978-3-031-72986-7_2.
[14] Baijiong Lin, Feiyang Ye, Yu Zhang, and Ivor W Tsang. Reasonable effectiveness of random weighting: A litmus test for multi-task learning. Transactions on Machine Learning Research, 2022.
[15] Bo Liu, Xingchao Liu, Xiaojie Jin, Peter Stone, and Qiang Liu. Conflict-averse gradient descent for multi-task learning. In Advances in Neural Information Processing Systems, volume 34, pages 18878–18890, 2021.
[16] Bo Liu, Yihao Feng, Peter Stone, and Qiang Liu. FAMO: Fast adaptive multitask optimization. In Advances in Neural Information Processing Systems (NeurIPS), volume 36, 2024.
[17] Dong Liu and Yanxuan Yu. MT2ST: Adaptive multi-task to single-task learning. arXiv preprint arXiv:2406.18038, 2025. Presented at MAGMaR Workshop, ACL 2025.
[18] Shikun Liu, Edward Johns, and Andrew J Davison. End-to-end multi-task learning with attention. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 1871–1880, 2019.
[19] Shikun Liu, Stephen James, Andrew J. Davison, and Edward Johns. Auto-lambda: Disentangling dynamic task relationships. Transactions on Machine Learning Research, 2022.
[20] Ye Lyu, George Vosselman, Gui-Song Xia, Alper Yilmaz, and Michael Ying Yang. Uavid: A semantic segmentation dataset for uav imagery. ISPRS Journal of Photogrammetry and Remote Sensing, 165:108–119, 2020.
[21] Jiaqi Ma, Zhe Zhao, Xinyang Yi, Jilin Chen, Lichan Hong, and Ed H Chi. Modeling task relationships in multi-task learning with multi-gate mixture-of-experts. In Proceedings of the 24th ACM SIGKDD international conference on knowledge discovery & data mining, pages 1930–1939, 2018.
[22] Logambal Madhuanand, Francesco Nex, and Michael Ying Yang. Self-supervised monocular depth estimation from oblique uav videos. ISPRS Journal of Photogrammetry and Remote Sensing, 176:1–14, 2021. doi: 10.1016/j.isprsjprs.2021.03.024.
[23] Ishan Misra, Abhinav Shrivastava, Abhinav Gupta, and Martial Hebert. Cross-stitch networks for multi-task learning. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 3994–4003, 2016.
[24] Aviv Navon, Aviv Shamsian, Idan Achituve, Haggai Maron, Kenji Kawaguchi, Gal Chechik, and Ethan Fetaya. Multi-task learning as a bargaining game. In International Conference on Machine Learning, pages 16428–16446. PMLR, 2022.
[25] Francesco Nex, Costas Armenakis, Michael Cramer, et al. UAV in the advent of the twenties: Where we stand and what is next. ISPRS Journal of Photogrammetry and Remote Sensing, 184:215–242, 2022.
[26] Maxime Oquab, Timothée Darcet, Théo Moutakanni, Huy Vo, Marc Szafraniec, Vasil Khalidov, et al. DINOv2: Learning robust visual features without supervision. Transactions on Machine Learning Research (TMLR), 2024.
[27] Mohammad Pezeshki, Oumar Kaba, Yoshua Bengio, Aaron Courville, Doina Precup, and Guillaume Lajoie. Gradient starvation: A learning proclivity in neural networks. In Advances in Neural Information Processing Systems (NeurIPS), volume 34, pages 1256–1272, 2021.
[28] René Ranftl, Alexey Bochkovskiy, and Vladlen Koltun. Vision transformers for dense prediction. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 12179–12188, 2021.
[29] Ramprasaath R. Selvaraju, Michael Cogswell, Abhishek Das, Ramakrishna Vedantam, Devi Parikh, and Dhruv Batra. Grad-CAM: Visual explanations from deep networks via gradient-based localization. In Proceedings of the IEEE International Conference on Computer Vision (ICCV), pages 618–626, 2017.
[30] Ozan Sener and Vladlen Koltun. Multi-task learning as multi-objective optimization. Advances in Neural Information Processing Systems, 31, 2018.
[31] Shital Shah, Debadeepta Dey, Chris Lovett, and Ashish Kapoor. AirSim: High-fidelity visual and physical simulation for autonomous vehicles. In Field and Service Robotics, pages 621–635. Springer, 2018.
[32] Oriane Siméoni, Huy V. Vo, Maximilian Seitzer, Federico Baldassarre, Maxime Oquab, Cijo Jose, Vasil Khalidov, Marc Szafraniec, Seungeun Yi, Michaël Ramamonjisoa, Francisco Massa, Daniel Haziza, Luca Wehrstedt, Jianyuan Wang, Timothée Darcet, Théo Moutakanni, Leonel Sentana, Claire Roberts, Andrea Vedaldi, Jamie Tolan, John Brandt, Camille Couprie, Julien Mairal, Hervé Jégou, Patrick Labatut, and Piotr Bojanowski. DINOv3. arXiv preprint arXiv:2508.10104, 2025. doi: 10.48550/arXiv.2508.10104.
[33] Trevor Standley, Amir R. Zamir, Dawn Chen, Leonidas Guibas, Jitendra Malik, and Silvio Savarese. Which tasks should be learned together in multi-task learning? In Proceedings of the International Conference on Machine Learning (ICML), pages 9120–9132. PMLR, 2020.
[34] Simon Vandenhende, Stamatios Georgoulis, Wouter Van Gansbeke, Marc Proesmans, Dengxin Dai, and Luc Van Gool. Multi-task learning for dense prediction tasks: A survey. IEEE Transactions on Pattern Analysis and Machine Intelligence, 44(7):3614–3633, 2022.
[35] Gui-Song Xia, Xiang Bai, Jian Ding, Zhen Zhu, Serge Belongie, Jiebo Luo, Mihai Datcu, Marcello Pelillo, and Liangpei Zhang. Dota: A large-scale dataset for object detection in aerial images. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 3974–3983, 2018.
[36] Zhitong Xiong, Wei Huang, Jingtao Hu, and Xiao Xiang Zhu. THE benchmark: Transferable representation learning for monocular height estimation. IEEE Transactions on Geoscience and Remote Sensing, 61, 2023. doi: 10.1109/TGRS.2023.3311764.
[37] Xue Yang, Junchi Yan, Qi Ming, Wentao Wang, Xiaopeng Zhang, and Qi Tian. Rethinking rotated object detection with gaussian wasserstein distance loss. In International Conference on Machine Learning, pages 11830–11841. PMLR, 2021.
[38] Xue Yang, Xiaojiang Yang, Jirui Yang, Qi Ming, Wentao Wang, Qi Tian, and Junchi Yan. Learning high-precision bounding box for rotated object detection via kullback-leibler divergence. In Advances in Neural Information Processing Systems (NeurIPS), volume 34, pages 18381–18394, 2021.
[39] Jun Yu, Yutong Dai, Xiaokang Liu, Jin Huang, Yishan Shen, Ke Zhang, Rong Zhou, Eashan Adhikarla, Wenxuan Ye, Yixin Liu, Zhaoming Kong, Kai Zhang, Yilong Yin, Vinod Namboodiri, Brian D. Davison, Jason H. Moore, and Yong Chen. Unleashing the power of multi-task learning: A comprehensive survey spanning traditional, deep, and pretrained foundation model eras. arXiv preprint arXiv:2404.18961, 2024.
[40] Tianhe Yu, Saurabh Kumar, Abhishek Gupta, Sergey Levine, Karol Hausman, and Chelsea Finn. Gradient surgery for multi-task learning. In Advances in Neural Information Processing Systems (NeurIPS), volume 33, pages 5824–5836, 2020.
[41] Amir R. Zamir, Alexander Sax, William Shen, Leonidas Guibas, Jitendra Malik, and Silvio Savarese. Taskonomy: Disentangling task transfer learning. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018.
[42] Jinghao Zhou, Chen Wei, Huiyu Wang, Wei Shen, Cihang Xie, Alan Yuille, and Tao Kong. iBOT: Image bert pre-training with online tokenizer. In International Conference on Learning Representations (ICLR), 2022.
[43] Xingyi Zhou, Dequan Wang, and Philipp Krähenbühl. Objects as points. In arXiv preprint arXiv:1904.07850, 2019.