Publications (6)
SLIDE : In Defense of Smart Algorithms over Hardware Acceleration for Large-Scale Deep Learning Systems
Beidi Chen, Tharun Medini, James Farwell +3
Deep Learning (DL) algorithms are the central focus of modern machine learning systems. As data volumes keep growing, it has become customary to train large neural networks with hu…
PROMPT: Learning Dynamic Resource Allocation Policies for Network Applications
Drew Penney, Bin Li, Jaroslaw Sydir +5
A growing number of service providers are exploring methods to improve server utilization and reduce power consumption by co-scheduling high-priority latency-critical workloads wit…
RAPID: Enabling Fast Online Policy Learning in Dynamic Public Cloud Environments
Drew Penney, Bin Li, Lizhong Chen +6
Resource sharing between multiple workloads has become a prominent practice among cloud service providers, motivated by demand for improved resource utilization and reduced cost of…
ORCA: A Network and Architecture Co-design for Offloading us-scale Datacenter Applications
Yifan Yuan, Jinghan Huang, Yan Sun +7
Responding to the "datacenter tax" and "killer microseconds" problems for datacenter applications, diverse solutions including Smart NIC-based ones have been proposed. Nonetheless,…
Accelerating SLIDE Deep Learning on Modern CPUs: Vectorization, Quantizations, Memory Optimizations, and More
Shabnam Daghaghi, Nicholas Meisburger, Mengnan Zhao +4
Deep learning implementations on CPUs (Central Processing Units) are gaining more traction. Enhanced AI capabilities on commodity x86 architectures are commercially appealing due t…
IOCA: High-Speed I/O-Aware LLC Management for Network-Centric Multi-Tenant Platform
Yifan Yuan, Mohammad Alian, Yipeng Wang +4
In modern server CPUs, last-level cache (LLC) is a critical hardware resource that exerts significant influence on the performance of the workloads, and how to manage LLC is a key…