9 citations · 10 across the 4 of their papers we have counts for
5 papers
ACT now: Aggregate Comparison of Traces for Incident Localization
Kamala Ramasubramanian, Ashutosh Raina, Jonathan Mace +1
Incidents in production systems are common and downtime is expensive. Applying an appropriate mitigating action quickly, such as changing a specific firewall rule, reverting a chan…
Mapping Datasets to Object Storage System
Xiaowei, Chu, Jeff LeFevre +5
Access libraries such as ROOT and HDF5 allow users to interact with datasets using high level abstractions, like coordinate systems and associated slicing operations. Unfortunately…
Elle: Inferring Isolation Anomalies from Experimental Observations
Kyle Kingsbury, Peter Alvaro
Users who care about their data store it in databases, which (at least in principle) guarantee some form of transactional isolation. However, experience shows [Kleppmann 2019, King…
Co-evolving Tracing and Fault Injection with Box of Pain
Daniel Bittman, Ethan L. Miller, Peter Alvaro
Distributed systems are hard to reason about largely because of uncertainty about what may go wrong in a particular execution, and about whether the system will mitigate those faul…
Keeping CALM: When Distributed Consistency is Easy
Joseph M. Hellerstein, Peter Alvaro
A key concern in modern distributed systems is to avoid the cost of coordination while maintaining consistent semantics. Until recently, there was no answer to the question of when…