Deakin University
Browse

Checkpointing schemes for grid workflow systems

Version 2 2024-06-17, 21:45
Version 1 2017-05-11, 14:58
journal contribution
posted on 2024-06-17, 21:45 authored by Z Li, Y Xiang
One of the major challenges in wide use of Grid workflow systems is fault tolerance and avoidance. Checkpointing schemes provide a way of fault detection and recovery. In our research, we focus on the performance optimization of checkpointing schemes and dynamic voltage scaling (DVS) for Grid workflow systems. We propose offline checkpointing schemes with DVS and online adaptive checkpointing schemes that dynamically adjust the checkpointing intervals by using store checkpoints and compare checkpoints. When combined with DVS, offline adaptive checkpointing schemes not only are fault tolerant but also lead to reduce average execution time of tasks. These schemes can efficiently utilize comparison and storage operations and significantly improve the performance. Further, these schemes can calculate the optimal numbers of checkpoints by which the mean execution time can be minimized. We also expand the online adaptive checkpointing schemes from single-task execution scenarios to multi-task execution scenarios. Simulation results show that these online schemes outstandingly increase the likelihood of timely task completion when faults occur.

History

Journal

Concurrency computation practice and experience

Volume

20

Pagination

1773-1790

Location

Chichester, United Kingdom

ISSN

1532-0626

eISSN

1532-0634

Language

eng

Publication classification

C1.1 Refereed article in a scholarly journal

Copyright notice

2008 John Wiley & Sons

Issue

15

Publisher

John Wiley & Sons