跳到主要內容

簡易檢索 / 詳目顯示

研究生: 楊聲遠
Yang, Sheng-Yuan
論文名稱: CAWS-DA:一種具部署感知之容量調適多寬度排程策略於尾端風險網路流量預測
CAWS-DA : A Deployment-Aware Capacity-Adaptive Multi-Width Scheduling Strategy for Tail-Risk-Aware Internet Traffic Forecasting
指導教授: 張宏慶
Hung-Chen Jang
口試委員: 吳曉光
胡誌麟
馮輝文
學位類別: 碩士
Master
系所名稱: 資訊學院 - 資訊科學系
Department of Computer Science
論文出版年: 2026
畢業學年度: 114
語文別: 中文
論文頁數: 82
中文關鍵詞: 深度學習時空網路流量預測模型壓縮寬度排程部署導向訓練尾端風險
外文關鍵詞: Deep Learning, Spatio-Temporal Internet Traffic Forecasting, Model Compression, Width Scheduling, Deployment-aware Training, Tail Risk
相關次數: 點閱:10下載:0
分享至:
查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報
  • 隨著行動通訊網路與智慧城市相關應用的快速發展,城市級時空流量預測已成為網路資源配置與系統規劃中的重要課題。近年來,時空圖神經網路(Spatio-Temporal Graph Neural Networks, STGNNs)在此類預測任務上展現出優異的建模能力,然而其高模型複雜度與龐大的計算需求,往往限制了模型在實際部署環境中的可行性。如何在維持預測效能的同時,降低模型部署成本,成為一項具挑戰性的研究問題。
      本論文提出一種容量感知之寬度排程與部署導向訓練策略(Capacity-Aware Width Scheduling with Deployment-Aware Training, CAWS-DA),以解決時空網路流量預測模型在部署階段所面臨的效能與效率取捨問題。CAWS-DA 以單一可調寬度的 STGCN 超網路為核心,透過訓練過程中的寬度共享,使不同容量之子網路能夠共用模型參數,從而避免多模型訓練所帶來的額外成本。
      在訓練策略上,本方法引入 warm-up 與線性寬度遞減排程機制,於訓練初期固定使用最大寬度模型以學習穩定的時空特徵表示,並於 warm-up 階段結束後逐步將訓練寬度縮減至目標部署寬度。此外,為提升不同寬度子網路之學習穩定性,本研究採用三寬度 sandwich 監督機制,同時對最大寬度、排程寬度與部署寬度子網路進行監督學習。進一步地,引入一致性學習(consistency learning),藉由限制部署寬度子網路與最大寬度子網路輸出的一致性,促進低容量模型對高容量模型表徵的有效繼承。
      本研究以 Milan 城市行動通訊流量資料集進行實驗評估,結果顯示 CAWS-DA 在顯著降低模型參數數量與計算複雜度的同時,仍能在 RMSE 及高分位誤差指標(P95、P99)上維持與完整模型相近的預測效能。相較於傳統知識蒸餾或後處理模型壓縮方法,CAWS-DA 無須額外預訓練教師模型,且能在單一訓練流程中同時支援不同部署需求,展現其在實際時空網路流量預測系統中的實用性與彈性。


    With the rapid development of mobile communication networks and smart city applications, urban-scale spatio-temporal traffic forecasting has become an important task for network resource allocation and system planning. Although spatio-temporal graph neural networks provide strong forecasting capability, their model complexity and computational requirements often limit deployment in resource-constrained environments.
    This study proposes CAWS-DA (Capacity-Adaptive Width Scheduling with Deployment-Aware Training), a deployment-oriented multi-width training strategy built on a width-adjustable STGCN supernetwork. CAWS-DA combines warm-up training, linear width scheduling, sandwich supervision, and consistency learning to jointly optimize subnetworks with different capacities while aligning training with the target deployment width.
    Experiments are conducted on the Milan mobile communication traffic dataset using RMSE, MAE, P95, and P99 to evaluate both overall prediction error and tail risk. The results show that CAWS-DA with β=0.2 achieves the lowest RMSE and P99 among the main comparison methods. On the top 10% high-traffic subset, CAWS-DA also achieves the best RMSE, MAE, and P95. Deployment-cost analysis further shows that CAWS-DA can maintain competitive forecasting performance without increasing the structural cost of models at the same deployment width, while providing flexible operating points under different resource constraints.
    Overall, CAWS-DA provides an effective trade-off among forecasting accuracy, tail-risk control, and deployment cost, making it suitable for resource-constrained spatio-temporal network traffic forecasting.

    第一章 緒論 1
    1.1 研究背景 1
    1.2 研究動機 3
    1.3 研究目標 3
    1.4 研究缺口 4
    1.5 論文架構 5
    第二章 相關研究 6
    2.1 時空預測問題 6
    2.2 基於圖結構的時空模型(Graph-based Spatio-Temporal Models) 7
    2.2.1 STGCN (Spatio-Temporal Graph Convolutional Network) 9
    2.3 長尾分布和尾部感知預測 10
    2.4 模型壓縮方法 12
    2.4.1 知識蒸餾 12
    2.4.2 模型剪枝 14
    2.4.3 Slimmable Neural Network 15
    2.4.4 Once For ALL (OFA) 16
    2.5 小結 17
    第三章 研究方法 18
    3.1 問題定義 18
    3.1.1 網路流量預測任務 18
    3.1.2 部署資源限制 19
    3.1.3 時間序列切分與資料正規化 20
    3.2 主幹模型架構: STGCN 21
    3.2.1 空間卷積(Graph Convolution) 21
    3.2.2 時間卷積(Temporal Convolution) 22
    3.2.3 ST-Block 23
    3.3 CAWS-DA 24
    3.3.1 多寬度共享架構 24
    3.3.2 寬度排程 (CAWS) 24
    3.3.3 三寬度Sandwich 監督與一致性學習 25
    3.4 整體訓練流程 26
    3.5 顯著圖 (Saliency map) 28
    3.6 評估指標 29
    第四章 實驗與結果分析 31
    4.1 實驗環境 31
    4.2 資料集介紹與前處理 32
    4.3 參數配置與比較方法 34
    4.4 模型間預測效能比較 35
    4.4.1 各模型之實驗結果比較 35
    4.4.2 各方法之訓練複雜度分析 38
    4.4.3 CDF 和 CCDF 分析 39
    4.4.4 Trentino跨資料集之各模型比較分析 41
    4.5 後設梯度參數梯度行為分析 43
    4.6 消融實驗 45
    4.6.1 移除中間層權重(without w_mid) 46
    4.6.2 移除寬度排程(without scheduling) 48
    4.6.3 移除一致性約束(without consistency) 49
    4.6.4 移除暖機機制(without warmup) 50
    4.7 敏感度分析 52
    4.7.1 部署寬度敏感度分析 52
    4.7.2 暖機參數敏感度分析 54
    4.7.3 一致性Beta敏感度分析 55
    4.8 基於顯著圖的差異分析 56
    4.9 高流量情境下之壓縮方法性能比較 60
    4.10 部署成本與效能之權衡分析 62
    第五章 結論與未來研究 72
    5.1 結論 72
    5.2 未來研究 74
    參考文獻 76

    [1] W. Jiang, “Cellular traffic prediction with machine learning: A survey,” Expert Systems with Applications, vol. 201, p. 117163, Sep. 2022, doi: 10.1016/j.eswa.2022.117163.
    [2] M. J. Neve and G. B. Rowe, “Mobile radio propagation in irregular cellular topographies using ray methods,” Inst. Elect. Eng. Proc. Microw., vol. 142, no. 6, pp. 447–451, Dec. 1995.
    [3] O. Aouedi, V. A. Le, K. Piamrat, and Y. Ji, “Deep Learning on Network Traffic Prediction: Recent Advances, Analysis, and Future Directions,” ACM Comput. Surv., vol. 57, no. 6, pp. 1–37, Jun. 2025, doi: 10.1145/3703447.
    [4] K. Cho et al., “Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation,” 2014, arXiv. doi: 10.48550/ARXIV.1406.1078.
    [5] J. C. B. Gamboa, “Deep Learning for Time-Series Analysis,” 2017, arXiv. doi: 10.48550/ARXIV.1701.01887.
    [6] D. Blalock, J. J. G. Ortiz, J. Frankle, and J. Guttag, “What is the State of Neural Network Pruning?,” 2020, arXiv. doi: 10.48550/ARXIV.2003.03033.
    [7] J. Gou, B. Yu, S. J. Maybank, and D. Tao, “Knowledge distillation: A survey,” International journal of computer vision, vol. 129, no. 6, pp. 1789–1819, 2021.
    [8] T. Elsken, J. H. Metzen, and F. Hutter, “Neural architecture search: A survey,” Journal of Machine Learning Research, vol. 20, no. 55, pp. 1–21, 2019.
    [9] S. Wang, J. Cao, and S. Y. Philip, “Deep learning for spatio-temporal data mining: A survey,” IEEE transactions on knowledge and data engineering, vol. 34, no. 8, pp. 3681–3700, 2020.
    [10] X. Shi, Z. Chen, H. Wang, D.-Y. Yeung, W.-K. Wong, and W. Woo, “Convolutional LSTM network: A machine learning approach for precipitation nowcasting,” Advances in neural information processing systems, vol. 28, 2015.
    [11] A. Vaswani et al., “Attention is all you need,” Advances in neural information processing systems, vol. 30, 2017.
    [12] T. N. Kipf, “Semi-supervised classification with graph convolutional networks,” arXiv preprint arXiv:1609.02907, 2016.
    [13] J. Gilmer, S. S. Schoenholz, P. F. Riley, O. Vinyals, and G. E. Dahl, “Neural message passing for quantum chemistry,” in International conference on machine learning, Pmlr, 2017, pp. 1263–1272.
    [14] B. Yu, H. Yin, and Z. Zhu, “Spatio-Temporal Graph Convolutional Networks: A Deep Learning Framework for Traffic Forecasting,” in Proceedings of the Twenty-Seventh International Joint Conference on Artificial Intelligence, Stockholm, Sweden: International Joint Conferences on Artificial Intelligence Organization, Jul. 2018, pp. 3634–3640. doi: 10.24963/ijcai.2018/505.
    [15] L. Zhao et al., “T-GCN: A temporal graph convolutional network for traffic prediction,” IEEE transactions on intelligent transportation systems, vol. 21, no. 9, pp. 3848–3858, 2019.
    [16] Li, Y.; Yu, R.; Shahabi, C.; Liu, Y. Diffusion Convolutional Recurrent Neural Network: Data-Driven Traffic Forecasting. In International Conference on Learning Representations (ICLR); 2018.
    [17] Z. Wu, S. Pan, G. Long, J. Jiang, and C. Zhang, “Graph WaveNet for Deep Spatial-Temporal Graph Modeling,” in Proceedings of the Twenty-Eighth International Joint Conference on Artificial Intelligence, IJCAI-19, International Joint Conferences on Artificial Intelligence Organization, Jul. 2019, pp. 1907–1913. doi: 10.24963/ijcai.2019/264.
    [18] S. Guo, Y. Lin, N. Feng, C. Song, and H. Wan, “Attention based spatial-temporal graph convolutional networks for traffic flow forecasting,” in Proceedings of the AAAI conference on artificial intelligence, 2019, pp. 922–929.
    [19] L. Bai, L. Yao, C. Li, X. Wang, and C. Wang, “Adaptive graph convolutional recurrent network for traffic forecasting,” Advances in neural information processing systems, vol. 33, pp. 17804–17815, 2020.
    [20] C. Zheng, X. Fan, C. Wang, and J. Qi, “Gman: A graph multi-attention network for traffic prediction,” in Proceedings of the AAAI conference on artificial intelligence, 2020, pp. 1234–1241.
    [21] I. Sutskever, O. Vinyals, and Q. V. Le, “Sequence to sequence learning with neural networks,” Advances in neural information processing systems, vol. 27, 2014.
    [22] H. Zhou et al., “Informer: Beyond efficient transformer for long sequence time-series forecasting,” in Proceedings of the AAAI conference on artificial intelligence, 2021, pp. 11106–11115.
    [23] G. Hinton, O. Vinyals, and J. Dean, “Distilling the knowledge in a neural network,” arXiv preprint arXiv:1503.02531, 2015.
    [24] Y. Chebotar and A. Waters, “Distilling knowledge from ensembles of neural networks for speech recognition.,” in Interspeech, 2016, pp. 3439–3443.
    [25] A. Romero, N. Ballas, S. E. Kahou, A. Chassang, C. Gatta, and Y. Bengio, “FitNets: Hints for Thin Deep Nets,” 2014, arXiv. doi: 10.48550/ARXIV.1412.6550.
    [26] M. Takamoto, Y. Morishita, and H. Imaoka, “An Efficient Method of Training Small Models for Regression Problems with Knowledge Distillation,” 2020, arXiv. doi: 10.48550/ARXIV.2002.12597.
    [27] Zeng, A.; Chen, M.; Zhang, L.; Xu, Q. Are Transformers Effective for Time Series Forecasting? In Proceedings of the AAAI Conference on Artificial Intelligence; 2023.
    [28] Zheng, Y.; Capra, L.; Wolfson, O.; Yang, H. Urban Computing: Concepts, Methodologies, and Applications. ACM Trans. Intell. Syst. Technol. 2014, 5, 38:1-38:55.
    [29] Wang, X.; Zhou, Z.; Xiao, F.; Xing, K.; Yang, Z.; Liu, Y.; Peng, C. Spatio-Temporal Analysis and Prediction of Cellular Traffic in Metropolis. IEEE Transactions on Mobile Computing 2019, 18 (9), 2190–2202. https://doi.org/10.1109/TMC.2018.2870135.
    [30] Koenker, R.; Bassett Jr, G. Regression Quantiles. Econometrica: journal of the Econometric Society 1978, 33–50.
    [31] Scheepens, D.; Schicker, I.; Hlavackova-Schindler, K.; Plant, C. Adapting a Deep Convolutional RNN Model with Imbalanced Regression Loss for Improved Spatio-Temporal Forecasting of Extreme Wind Speed Events in the Short to Medium Range. Geoscientific Model Development 2023, 16, 251–270. https://doi.org/10.5194/gmd-16-251-2023.
    [32] Yang, Y.; Zha, K.; Chen, Y.-C.; Wang, H.; Katabi, D. Delving into Deep Imbalanced Regression. In International Conference on Machine Learning; 2021.
    [33] Lecun, Y.; Denker, J.; Solla, S. Optimal Brain Damage. In Advances in Neural Information Processing Systems; 1989; Vol. 2, pp 598–605.
    [34] Y. Wang et al., “Pruning from Scratch,” AAAI, vol. 34, no. 07, pp. 12273–12280, Apr. 2020, doi: 10.1609/aaai.v34i07.6910.
    [35] Cai, H.; Gan, C.; Wang, T.; Zhang, Z.; Han, S. Once for All: Train One Network and Specialize It for Efficient Deployment. In International Conference on Learning Representations; 2020.
    [36] Yu, J.; Yang, L.; Xu, N.; Yang, J.; Huang, T. Slimmable Neural Networks. In International Conference on Learning Representations (ICLR); 2019.
    [37] Yu, J.; Huang, T. S. Universally Slimmable Networks and Improved Training Techniques. In Proceedings of the IEEE/CVF international conference on computer vision; 2019; pp 242–251.
    [38] K. Simonyan, A. Vedaldi, and A. Zisserman, “Deep Inside Convolutional Networks: Visualising Image Classification Models and Saliency Maps,” 2013, arXiv. doi: 10.48550/ARXIV.1312.6034.
    [39] G. Barlacchi et al., “A multi-source dataset of urban life in the city of Milan and the Province of Trentino,” Sci Data, vol. 2, no. 1, p. 150055, Oct. 2015, doi: 10.1038/sdata.2015.55.
    [40] A. A. Hussien, H. Nashaat, and R. F. Abdel-Kader, “Machine learning techniques for spatiotemporal traffic prediction in 5G cellular networks,” Discov Appl Sci, vol. 7, no. 10, p. 1047, Sep. 2025, doi: 10.1007/s42452-025-06746-3.
    [41] J. Zhang, Y. Zheng, and D. Qi, “Deep Spatio-Temporal Residual Networks for Citywide Crowd Flows Prediction,” AAAI, vol. 31, no. 1, Feb. 2017, doi: 10.1609/aaai.v31i1.10735.
    [42] Belt, Eline A., Thomas Koch, and Elenna R. Dugundji. "Hourly forecasting of traffic flow rates using spatial temporal graph neural networks." Procedia Computer Science 220 (2023): 102-109.
    [43] Howard, A. G.; Zhu, M.; Chen, B.; Kalenichenko, D.; Wang, W.; Weyand, T.; Andreetto, M.; Adam, H. MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications. ArXiv 2017, abs/1704.04861.
    [44] Zhang, Li Lyna, et al. "Nn-meter: Towards accurate latency prediction of deep-learning model inference on diverse edge devices." Proceedings of the 19th Annual International Conference on Mobile Systems, Applications, and Services. 2021.
    [45] Jain, P.; Mo, X.; Jain, A.; Subbaraj, H.; Durrani, R.; Tumanov, A.; Gonzalez, J. E.; Stoica, I. Dynamic Space-Time Scheduling for GPU Inference. ArXiv 2018, abs/1901.00041.
    [46] MA, Ningning, et al. Shufflenet v2: Practical guidelines for efficient cnn architecture design. In: Proceedings of the European conference on computer vision (ECCV). 2018. p. 116-131.
    [47] Karbachevsky, A.; Baskin, C.; Zheltonozhskii, E.; Yermolin, Y.; Gabbay, F.; Bronstein, A.M.; Mendelson, A. Early-Stage Neural Network Hardware Performance Analysis. Sustainability 2021, 13, 717. https://doi.org/10.3390/su13020717
    [48] Nie, Y.; H. Nguyen, N.; Sinthong, P.; Kalagnanam, J. A Time Series Is Worth 64 Words: Long-Term Forecasting with Transformers. In International Conference on Learning Representations; 2023.
    [49] Wang, S.; Wu, H.; Shi, X.; Hu, T.; Luo, H.; Ma, L.; Zhang, J. Y.; ZHOU, J. TimeMixer: Decomposable Multiscale Mixing for Time Series Forecasting. In International Conference on Learning Representations (ICLR); 2024.
    [50] Y. Yang et al., “A Survey on Diffusion Models for Time Series and Spatio-Temporal Data,” 2024, arXiv. doi: 10.48550/ARXIV.2404.18886.
    [51] Frankle, J., & Carbin, M. (2018). The lottery ticket hypothesis: Finding sparse, trainable neural networks. arXiv preprint arXiv:1803.03635.
    [52] Jin, S., Zhang, C., Jiang, X., Feng, Y., Guan, H., Li, G., Song, S. L., & Tao, D. (2021). COMET: a novel Memory-Efficient Deep Learning Training Framework by using Error-Bounded Lossy Compression. arXiv (Cornell University). https://doi.org/10.48550/arxiv.2111.09562
    [53] B. Jacob et al., “Quantization and Training of Neural Networks for Efficient Integer-Arithmetic-Only Inference,” arXiv:1712.05877 [cs, stat], Dec. 2017, Available: https://arxiv.org/abs/1712.05877
    [54] Esser, Steven K., et al. “Learned Step Size Quantization.” arXiv, 2019. DOI.org (Datacite), https://doi.org/10.48550/ARXIV.1902.08153.
    [55] Z. Chen, V. Badrinarayanan, C.-Y. Lee, and A. Rabinovich, “GradNorm: Gradient Normalization for Adaptive Loss Balancing in Deep Multitask Networks,” 2017, doi: 10.48550/ARXIV.1711.02257.

    QR CODE
    :::