Repairing Multiple Failures with Coordinated and Adaptive Regenerating Codes
arXiv:1102.0204
Abstract
Erasure correcting codes are widely used to ensure data persistence in distributed storage systems. This paper addresses the simultaneous repair of multiple failures in such codes. We go beyond existing work (i.e., regenerating codes by Dimakis et al.) by describing (i) coordinated regenerating codes (also known as cooperative regenerating codes) which support the simultaneous repair of multiple devices, and (ii) adaptive regenerating codes which allow adapting the parameters at each repair. Similarly to regenerating codes by Dimakis et al., these codes achieve the optimal tradeoff between storage and the repair bandwidth. Based on these extended regenerating codes, we study the impact of lazy repairs applied to regenerating codes and conclude that lazy repairs cannot reduce the costs in term of network bandwidth but allow reducing the disk-related costs (disk bandwidth and disk I/O).
Update to previous version adding (i) study of lazy repairs, (ii) adaptive codes at the MBR point, and (iii) discussion of related work. Extended from a regular paper at NetCod 2011 available at http://dx.doi.org/10.1109/ISNETCOD.2011.5978920 . First version: "Beyond Regenerating Codes", September 2010 on http://hal.inria.fr/inria-00516647/
Cited by in corpus (29)
- Cooperative Local Repair in Distributed Storage
- Distributed Data Storage Systems with Opportunistic Repair
- In-Network Redundancy Generation for Opportunistic Speedup of Backup
- Byzantine Fault Tolerance of Regenerating Codes
- Security Concerns in Minimum Storage Cooperative Regenerating Codes
- Cooperative Regenerating Codes
- Rack-Aware Cooperative Regenerating Codes
- Bandwidth Adaptive & Error Resilient MBR Exact Repair Regenerating Codes
- An Empirical Study of the Repair Performance of Novel Coding Schemes for Networked Distributed Storage Systems
- Centralized Multi-Node Repair Regenerating Codes
- Erasure Coding for Distributed Storage: An Overview
- A Note on Secure Minimum Storage Regenerating Codes
- Secure Cooperative Regenerating Codes for Distributed Storage Systems
- CORE: Augmenting Regenerating-Coding-Based Recovery for Single and Concurrent Failures in Distributed Storage Systems
- New MDS codes with small sub-packetization and near-optimal repair bandwidth
- An Overview of Codes Tailor-made for Better Repairability in Networked Distributed Storage Systems
- RapidRAID: Pipelined Erasure Codes for Fast Data Archival in Distributed Storage Systems
- Repair for Distributed Storage Systems in Packet Erasure Networks
- On the Achievability Region of Regenerating Codes for Multiple Erasures
- Cooperative Repair of Multiple Node Failures in Distributed Storage Systems
- Storage-Repair Bandwidth Trade-off for Wireless Caching with Partial Failure and Broadcast Repair
- Functional Broadcast Repair of Multiple Partial Failures in Wireless Distributed Storage Systems
- Repairing Multiple Failures in the Suh-Ramchandran Regenerating Codes
- New constructions of cooperative MSR codes: Reducing node size to
- Capacity of Wireless Distributed Storage Systems with Broadcast Repair
- Product Matrix MSR Codes with Bandwidth Adaptive Exact Repair
- Concurrent Regenerating Codes and Scalable Application in Network Storage
- Exact Cooperative Regenerating Codes with Minimum-Repair-Bandwidth for Distributed Storage
- Erasure Codes for Distributed Storage: Tight Bounds and Matching Constructions