Low latency via redundancy
arXiv:1306.3707
Abstract
Low latency is critical for interactive networked applications. But while we know how to scale systems to increase capacity, reducing latency --- especially the tail of the latency distribution --- can be much more difficult. In this paper, we argue that the use of redundancy is an effective way to convert extra capacity into reduced latency. By initiating redundant operations across diverse resources and using the first result which completes, redundancy improves a system's latency even under exceptional conditions. We study the tradeoff with added system utilization, characterizing the situations in which replicating all tasks reduces mean latency. We then demonstrate empirically that replicating all operations can result in significant mean and tail latency reduction in real-world systems including DNS queries, database servers, and packet forwarding within networks.
Cited by in corpus (16)
- RepFlow: Minimizing Flow Completion Times with Replicated Flows in Data Centers
- On Delay-Optimal Scheduling in Queueing Systems with Replications
- Provably Delay Efficient Data Retrieving in Storage Clouds
- On the stability of redundancy models
- Improving the performance of heterogeneous data centers through redundancy
- RepNet: Cutting Tail Latency in Data Center Networks with Flow Replication
- Efficient Straggler Replication in Large-scale Parallel Computing
- Contrasting Effects of Replication in Parallel Systems: From Overload to Underload and Back
- Optimizing Redundancy Levels in Master-Worker Compute Clusters for Straggler Mitigation
- A cost-benefit analysis of low latency via added utilization
- Approximate Networking for Global Access to the Internet for All (GAIA)
- Towards a Speed of Light Internet
- Efficient Task Replication for Fast Response Times in Parallel Computation
- Will 5G See its Blind Side? Evolving 5G for Universal Internet Access
- Online Job Scheduling with Redundancy and Opportunistic Checkpointing: A Speedup-Function-Based Analysis
- Tars: Timeliness-aware Adaptive Replica Selection for Key-Value Stores