papers

Publications (6)

cs.DC2020

SLIDE : In Defense of Smart Algorithms over Hardware Acceleration for Large-Scale Deep Learning Systems

Beidi Chen, Tharun Medini, James Farwell +3

Deep Learning (DL) algorithms are the central focus of modern machine learning systems. As data volumes keep growing, it has become customary to train large neural networks with hu…

cs.LG2023

PROMPT: Learning Dynamic Resource Allocation Policies for Network Applications

Drew Penney, Bin Li, Jaroslaw Sydir +5

A growing number of service providers are exploring methods to improve server utilization and reduce power consumption by co-scheduling high-priority latency-critical workloads wit…

cs.LG2023

RAPID: Enabling Fast Online Policy Learning in Dynamic Public Cloud Environments

Drew Penney, Bin Li, Lizhong Chen +6

Resource sharing between multiple workloads has become a prominent practice among cloud service providers, motivated by demand for improved resource utilization and reduced cost of…

cs.AR2022

ORCA: A Network and Architecture Co-design for Offloading us-scale Datacenter Applications

Yifan Yuan, Jinghan Huang, Yan Sun +7

Responding to the "datacenter tax" and "killer microseconds" problems for datacenter applications, diverse solutions including Smart NIC-based ones have been proposed. Nonetheless,…

cs.LG2021

Accelerating SLIDE Deep Learning on Modern CPUs: Vectorization, Quantizations, Memory Optimizations, and More

Shabnam Daghaghi, Nicholas Meisburger, Mengnan Zhao +4

Deep learning implementations on CPUs (Central Processing Units) are gaining more traction. Enhanced AI capabilities on commodity x86 architectures are commercially appealing due t…

cs.AR2021

IOCA: High-Speed I/O-Aware LLC Management for Network-Centric Multi-Tenant Platform

Yifan Yuan, Mohammad Alian, Yipeng Wang +4

In modern server CPUs, last-level cache (LLC) is a critical hardware resource that exerts significant influence on the performance of the workloads, and how to manage LLC is a key…