4 citations · 13 across the 6 of their papers we have counts for
12 papers
Eiger: An Efficient Library for GPU-based Data Analytics
Bowen Wu, Marko Kabić, Sven Hepkema +3
GPUs have become an increasingly attractive platform for accelerating analytical workloads due to their massive parallelism and high memory bandwidth. Recent studies show that in s…
SOL: Safe On-Node Learning in Cloud Platforms
Yawen Wang, Daniel Crankshaw, Neeraja J. Yadwadkar +3
Cloud platforms run many software agents on each server node. These agents manage all aspects of node operation, and in some cases frequently collect data and make decisions. Unfor…
RecShard: Statistical Feature-Based Memory Optimization for Industry-Scale Neural Recommendation
Geet Sethi, Bilge Acun, Niket Agarwal +3
We propose RecShard, a fine-grained embedding table (EMB) partitioning and placement technique for deep learning recommendation models (DLRMs). RecShard is designed based on two ke…
Faa$T: A Transparent Auto-Scaling Cache for Serverless Applications
Francisco Romero, Gohar Irfan Chaudhry, Íñigo Goiri +6
Function-as-a-Service (FaaS) has become an increasingly popular way for users to deploy their applications without the burden of managing the underlying infrastructure. However, ex…
Llama: A Heterogeneous & Serverless Framework for Auto-Tuning Video Analytics Pipelines
Francisco Romero, Mark Zhao, Neeraja J. Yadwadkar +1
The proliferation of camera-enabled devices and large video repositories has led to a diverse set of video analytics applications. These applications rely on video pipelines, repre…
RackSched: A Microsecond-Scale Scheduler for Rack-Scale Computers (Technical Report)
Hang Zhu, Kostis Kaffes, Zixu Chen +4
Low-latency online services have strict Service Level Objectives (SLOs) that require datacenter systems to support high throughput at microsecond-scale tail latency. Dataplane oper…