FastCheck: fast checkpointing and recovery for DNN training via parallel transmission and compression
Yun TENG , Dawei SUN , Shipeng HU , Zhiyue LI , Guangyan ZHANG , Haidong TIAN , Rui CHANG
Eng Inform Technol Electron Eng ›› 2026, Vol. 27 ›› Issue (2) : 250034
Training large-scale deep neural networks (DNNs) is prone to software and hardware failures, with critical failures often requiring full-machine reboots that substantially prolong training. Existing checkpoint-recovery solutions either cannot tolerate such critical failures or suffer from slow checkpointing and recovery due to constrained input/output bandwidth. In this paper, we propose FastCheck, a checkpoint-recovery framework that accelerates checkpointing and recovery through parallel transmission and tailored compression. First, FastCheck partitions checkpoints into shards and leverages multiple nodes for parallel checkpointing and recovery. Second, it further reduces checkpoint size and overhead with delta compression for weights and index compression for momentum. Third, FastCheck employs lightweight and consistent health status maintenance that accurately tracks node health, preventing checkpoint transmission to failed nodes. We implement FastCheck in PyTorch and evaluate it on multiple DNN models against two baselines. Experimental results show that FastCheck reduces the checkpointing time by up to 78.42% and the recovery time by up to 77.41%, while consistently improving efficiency across different training stages.
Deep neural network models / Critical failures / Parallel transmission / Data compression / Checkpointing and recovery
The Authors. Published by Zhejiang University Press Co., Ltd.
/
| 〈 |
|
〉 |