20 citations · 27 across the 4 of their papers we have counts for
4 papers · 1 filter
Is the GPU Half-Empty or Half-Full? Practical Scheduling Techniques for LLMs
Ferdi Kossmann, Bruce Fontaine, Daya Khudia +2
Serving systems for Large Language Models (LLMs) improve throughput by processing several requests concurrently. However, multiplexing hardware resources between concurrent request…
Low-Precision Hardware Architectures Meet Recommendation Model Inference at Scale
Zhaoxia, Deng, Jongsoo Park +17
Tremendous success of machine learning (ML) and the unabated growth in ML model complexity motivated many ML-specific designs in both CPU and accelerator architectures to speed up…
FBGEMM: Enabling High-Performance Low-Precision Deep Learning Inference
Daya Khudia, Jianyu Huang, Protonu Basu +4
Deep learning models typically use single-precision (FP32) floating point data types for representing activations and weights, but a slew of recent research work has shown that com…
Deep Learning Inference in Facebook Data Centers: Characterization, Performance Optimizations and Hardware Implications
Jongsoo Park, Maxim Naumov, Protonu Basu +25
The application of deep learning techniques resulted in remarkable improvement of machine learning models. In this paper provides detailed characterizations of deep learning models…