3 citations · 6 across the 4 of their papers we have counts for
5 papers · 1 filter
ACT now: Aggregate Comparison of Traces for Incident Localization
Kamala Ramasubramanian, Ashutosh Raina, Jonathan Mace +1
Incidents in production systems are common and downtime is expensive. Applying an appropriate mitigating action quickly, such as changing a specific firewall rule, reverting a chan…
The Benefit of Hindsight: Tracing Edge-Cases in Distributed Systems
Lei Zhang, Vaastav Anand, Zhiqiang Xie +2
Today's distributed tracing frameworks are ill-equipped to troubleshoot rare edge-case requests. The crux of the problem is a trade-off between specificity and overhead. On the one…
Aggregate-Driven Trace Visualizations for Performance Debugging
Vaastav Anand, Matheus Stolet, Thomas Davidson +3
Performance issues in cloud systems are hard to debug. Distributed tracing is a widely adopted approach that gives engineers visibility into cloud systems. Existing trace analysis…
Serving DNNs like Clockwork: Performance Predictability from the Bottom Up
Arpan Gujarati, Reza Karimi, Safya Alzayat +4
Machine learning inference is becoming a core building block for interactive web applications. As a result, the underlying model serving systems on which these applications depend…
No DNN Left Behind: Improving Inference in the Cloud with Multi-Tenancy
Amit Samanta, Suhas Shrinivasan, Antoine Kaufmann +1
With the rise of machine learning, inference on deep neural networks (DNNs) has become a core building block on the critical path for many cloud applications. Applications today re…