176 citations · 699 across the 41 of their papers we have counts for
13 papers · 1 filter
Dynamic Space-Time Scheduling for GPU Inference
Paras Jain, Xiangxi Mo, Ajay Jain +5
Serving deep neural networks in latency critical interactive settings often requires GPU acceleration. However, the small batch sizes typical in online inference results in poor GP…
Using Multitask Learning to Improve 12-Lead Electrocardiogram Classification
J. Weston Hughes, Taylor Sittler, Anthony D. Joseph +3
We develop a multi-task convolutional neural network (CNN) to classify multiple diagnoses from 12-lead electrocardiograms (ECGs) using a dataset comprised of over 40,000 ECGs, with…
InferLine: ML Prediction Pipeline Provisioning and Management for Tight Latency Objectives
Daniel Crankshaw, Gur-Eyal Sela, Corey Zumar +4
Serving ML prediction pipelines spanning multiple models and hardware accelerators is a key challenge in production machine learning. Optimally configuring these pipelines to meet…
On the Computational Inefficiency of Large Batch Sizes for Stochastic Gradient Descent
Noah Golmant, Nikita Vemuri, Zhewei Yao +5
Increasing the mini-batch size for stochastic gradient descent offers significant opportunities to reduce wall-clock training time, but there are a variety of theoretical and syste…
ReXCam: Resource-Efficient, Cross-Camera Video Analytics at Scale
Samvit Jain, Xun Zhang, Yuhao Zhou +4
Enterprises are increasingly deploying large camera networks for video analytics. Many target applications entail a common problem template: searching for and tracking an object or…
Inter-BMV: Interpolation with Block Motion Vectors for Fast Semantic Segmentation on Video
Samvit Jain, Joseph E. Gonzalez
Models optimized for accuracy on single images are often prohibitively slow to run on each frame in a video. Recent work exploits the use of optical flow to warp image features for…