| 研究生: |
胡元亨 Hu, Yuan-Hen |
|---|---|
| 論文名稱: |
分散式任務系統的協調與故障恢復管理機制: 以 Celery 為例 On Efficient Coordination and Fault Recovery of Distributed Task Queue Systems: A Case Study of Celery |
| 指導教授: |
廖峻鋒
Liao, Chun-Feng |
| 口試委員: |
馬尚彬
MA,SHANG-PIN 江玥慧 Chiang, Yueh-Hui |
| 學位類別: |
碩士
Master |
| 系所名稱: |
資訊學院 - 資訊科學系碩士在職專班 Excutive Master Program of Computer Science |
| 論文出版年: | 2026 |
| 畢業學年度: | 114 |
| 語文別: | 中文 |
| 論文頁數: | 46 |
| 中文關鍵詞: | Celery 、Redis 、容錯 、故障恢復 、事件驅動 、分散式鎖 |
| 外文關鍵詞: | Celery, Redis, Fault Tolerance, Failure Recovery, Event-driven, Distributed Lock |
| 相關次數: | 點閱:42 下載:0 |
| 分享至: |
| 查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報 |
作為 Python 生態系常用的分散式任務佇列框架,Celery 常搭配 RabbitMQ 或 Redis 作為 Broker。這種搭配在實務上也造成一些問題。首先,當以 Redis 作為 broker 時,其本身資料儲存機制缺乏在消費者連線中斷後主動退回未 ack 訊息的機制。其次,Redis 僅負責訊息的儲存與傳送,無法掌握任務的實際執行狀態與資源持有情形。在此情況下,Celery 只能依賴一個全域的任務逾時參數設定並搭配週期性掃描以恢復長時間未 ack 的訊息,以致系統難以同時縮短短任務的恢復時間並降低長任務的誤判風險。因此,本研究提出一個由 Registry、Monitor 與 Guardian 構成的協調與故障恢復架構。Registry 作為任務協調與故障恢復的共享狀態(shared state),獨立於個別 Worker 之外,使恢復決策得以跨越單一 Worker 的生命週期;Monitor 將 Celery 提供的監控事件用於分散式鎖管理與主動任務恢復,使系統由被動等待週期性掃描轉為事件驅動的秒級恢復;Guardian 可依任務的預期執行時間設定個別的逾時門檻,使預期執行時間不同的任務不必共用單一全域逾時設定。實驗顯示,Monitor 將容器重啟後的任務恢復時間大幅降低,在 Worker child process 崩潰場景亦可將原 Worker 持有的分散式鎖迅速釋放;Guardian 將短時間任務恢復所需時間大幅縮短,長時間任務之恢復誤判率則大幅降低。在正常運作的固定負載下,啟用三個設計元件後,整體 CPU 使用率由 4.69% 增至 5.67%,記憶體用量由 528.13 MiB 增至 664.18 MiB。根據實驗結果,本研究所提機制可有效融合於既有 Celery 系統上,彌補 Redis broker 被動恢復與任務狀態掌握不足所造成的限制。
Celery, a distributed task queue framework widely used in the Python ecosystem, is commonly paired with RabbitMQ or Redis as its broker. This configuration also creates several practical problems. First, when Redis is used as the broker, its storage mechanism lacks a way to proactively return unacknowledged messages after a consumer connection is interrupted. Second, Redis is responsible only for message storage and delivery and cannot track actual task execution states or resource ownership. Under these conditions, Celery must rely on a single global timeout setting together with periodic polling to recover messages that remain unacknowledged for an extended period. As a result, the system has difficulty both shortening the recovery time of short-running tasks and reducing the risk of falsely recovering long-running tasks. This study therefore proposes a coordination and failure recovery architecture composed of a Registry, Monitor, and Guardian. The Registry serves as the shared state for task coordination and failure recovery independently of individual Workers, allowing recovery decisions to persist beyond the lifecycle of a single Worker. Monitor uses Celery events for distributed lock management and proactive task recovery, shifting the system from passively waiting for periodic polling to event-driven recovery within seconds. Guardian can set an individual timeout threshold for each task according to its expected execution time, so tasks with different expected execution times do not need to share a single global timeout setting. Experimental results show that Monitor substantially reduces task recovery time after a Worker container restart and promptly releases distributed locks left behind by the original Worker when a Worker child process crashes. Guardian substantially shortens the recovery time of short-running tasks while greatly reducing the false-positive recovery rate for long-running tasks. Under a fixed workload during normal operation, enabling the three components increased total container CPU usage from 4.69% to 5.67% and memory usage from 528.13 MiB to 664.18 MiB. Based on these results, the proposed mechanisms can be integrated into existing Celery systems to compensate for the limitations caused by the Redis broker's passive recovery behavior and its lack of task-state awareness.
誌謝 I
摘要 II
Abstract III
目次 V
表次 VII
圖次 VIII
第一章 緒論 1
1.1 研究背景 1
1.2 研究動機 2
1.3 研究目標 4
1.4 論文架構 5
第二章 技術背景與相關研究 6
2.1 分散式任務系統與 Celery 故障處理背景 6
2.2 相關機制限制 9
2.3 小結 11
第三章 系統設計與故障恢復機制 12
3.1 研究範圍與設計目標 12
3.2 整體系統架構 14
3.3 Registry:任務協調註冊表 16
3.4 Monitor:事件驅動式恢復控制器 18
3.5 Guardian:任務守護控制器 20
3.6 故障場景與恢復策略 21
3.7 小結 22
第四章 系統實作與執行環境 23
第五章 系統評估 25
5.1 VT Polling 實證 25
5.2 Monitor 實驗評估 27
5.3 Guardian 實驗評估 34
5.4 Registry 外部重派次數控制實驗評估 37
5.5 資源開銷評估 38
5.6 綜合比較與討論 40
第六章 結論與未來工作 43
參考文獻 44
[1] S. Ghemawat, H. Gobioff, and S.-T. Leung, "The Google file system," in Proc. 19th ACM Symp. Operating Systems Principles (SOSP), Bolton Landing, NY, USA, Oct. 2003, pp. 29–43.
[2] G. DeCandia et al., "Dynamo: Amazon's highly available key-value store," ACM SIGOPS Operating Systems Review, vol. 41, no. 6, pp. 205–220, Dec. 2007, doi: 10.1145/1323293.1294281.
[3] T. D. Chandra and S. Toueg, "Unreliable failure detectors for reliable distributed systems," J. ACM, vol. 43, no. 2, pp. 225–267, Mar. 1996.
[4] N. Hayashibara, X. Défago, R. Yared, and T. Katayama, "The φ accrual failure detector," in Proc. 23rd IEEE Int. Symp. Reliable Distributed Systems (SRDS), Florianópolis, Brazil, Oct. 2004, pp. 66–78.
[5] D. Gelernter, "Generative communication in Linda," ACM Trans. Program. Lang. Syst., vol. 7, no. 1, pp. 80–112, Jan. 1985.
[6] G. Hohpe and B. Woolf, Enterprise Integration Patterns: Designing, Building, and Deploying Messaging Solutions. Boston, MA, USA: Addison-Wesley, 2003.
[7] P. Helland, "Idempotence is not a medical condition," Commun. ACM, vol. 55, no. 5, pp. 56–65, May 2012, doi: 10.1145/2160718.2160734.
[8] M. Kleppmann, Designing Data-Intensive Applications: The Big Ideas Behind Reliable, Scalable, and Maintainable Systems. Sebastopol, CA, USA: O'Reilly Media, 2017.
[9] Celery Project, "Tasks," Celery 5.3.2 Documentation. [Online]. Available: https://docs.celeryq.dev/en/v5.3.3/userguide/tasks.html. [Accessed: Jun. 30, 2026].
[10] Celery Project, "Using Redis," Celery 5.3.3 Documentation. [Online]. Available: https://docs.celeryq.dev/en/v5.3.3/getting-started/backends-and-brokers/redis.html. [Accessed: Jun. 30, 2026].
[11] Celery Project, "Monitoring and Management Guide," Celery 5.3.3 Documentation. [Online]. Available: https://docs.celeryq.dev/en/v5.3.3/userguide/monitoring.html. [Accessed: Jun. 30, 2026].
[12] RabbitMQ Team, "Reliability Guide," RabbitMQ Documentation. [Online]. Available: https://www.rabbitmq.com/docs/reliability. [Accessed: Jul. 20, 2026].
[13] P. Dobbelaere and K. S. Esmaili, "Kafka versus RabbitMQ: A comparative study of two industry reference publish/subscribe implementations," in Proc. 11th ACM Int. Conf. Distributed and Event-based Systems (DEBS), Barcelona, Spain, Jun. 2017, pp. 227–238, doi: 10.1145/3093742.3093908.
[14] J. Kreps, N. Narkhede, and J. Rao, "Kafka: A distributed messaging system for log processing," in Proc. 6th Int. Workshop Networking Meets Databases (NetDB), Athens, Greece, Jun. 2011, pp. 1–7. [Online]. Available: https://cwiki.apache.org/confluence/download/attachments/27822226/Kafka-netdb-06-2011.pdf
[15] G. Hohpe, "Conversation Patterns: Interactions between Loosely Coupled Services," in Proc. 12th European Conference on Pattern Languages of Programs (EuroPLoP), Kloster Irsee, Germany, 2007.
[16] M. Kawazoe Aguilera, W. Chen, and S. Toueg, "Heartbeat: A timeout-free failure detector for quiescent reliable communication," in Distributed Algorithms. Berlin, Heidelberg: Springer, 1997, pp. 126–140, doi: 10.1007/BFb0030680.
[17] C. G. Gray and D. R. Cheriton, "Leases: An efficient fault-tolerant mechanism for distributed file cache consistency," in Proc. 12th ACM Symp. Operating Systems Principles (SOSP), Litchfield Park, AZ, USA, Dec. 1989, pp. 202–210.
[18] L. Lamport, R. Shostak, and M. Pease, "The Byzantine generals problem," ACM Trans. Program. Lang. Syst., vol. 4, no. 3, pp. 382–401, Jul. 1982, doi: 10.1145/357172.357176.
[19] G. Coulouris, J. Dollimore, T. Kindberg, and G. Blair, Distributed Systems: Concepts and Design, 5th ed. Boston, MA, USA: Addison-Wesley, 2011.
[20] S. Gilbert and N. Lynch, "Brewer's conjecture and the feasibility of consistent, available, partition-tolerant web services," ACM SIGACT News, vol. 33, no. 2, pp. 51–59, Jun. 2002, doi: 10.1145/564585.564601.
[21] Celery Project, "Backends and Brokers," Celery 5.3.3 Documentation. [Online]. Available: https://docs.celeryq.dev/en/v5.3.3/getting-started/backends-and-brokers/index.html. [Accessed: Jun. 1, 2026].
[22] Redis, "Distributed locks with Redis," Redis Documentation. [Online]. Available: https://redis.io/docs/latest/develop/clients/patterns/distributed-locks/. [Accessed: Jun. 30, 2026].
[23] P. Hunt, M. Konar, F. P. Junqueira, and B. Reed, "ZooKeeper: Wait-free coordination for Internet-scale systems," in Proc. USENIX Annu. Tech. Conf. (USENIX ATC), Boston, MA, USA, Jun. 2010, pp. 145–158.
[24] M. Burrows, "The Chubby lock service for loosely-coupled distributed systems," in Proc. 7th USENIX Symp. Operating Systems Design and Implementation (OSDI), Seattle, WA, USA, Nov. 2006, pp. 335–350.
[25] M.-C. Hsueh, T. K. Tsai, and R. K. Iyer, "Fault injection techniques and tools," Computer, vol. 30, no. 4, pp. 75–82, Apr. 1997.
[26] Kombu Contributors, "kombu/transport/redis.py," kombu 5.6.2 source code. [Online]. Available: https://github.com/celery/kombu/blob/v5.6.2/kombu/transport/redis.py. [Accessed: Jun. 30, 2026].
全文公開日期 2027/08/06