9 citations · 9 across the 4 of their papers we have counts for
4 papers
Beyond the Buzz: A Pragmatic Take on Inference Disaggregation
Tiyasa Mitra, Ritika Borkar, Nidhi Bhatia +10
As inference scales to multi-node deployments, disaggregation - splitting inference into distinct phases - offers a promising path to improving the throughput-interactivity Pareto…
Key, Value, Compress: A Systematic Exploration of KV Cache Compression Techniques
Neusha Javidnia, Bita Darvish Rouhani, Farinaz Koushanfar
Large language models (LLMs) have demonstrated exceptional capabilities in generating text, images, and video content. However, as context length grows, the computational cost of a…
Microscaling Data Formats for Deep Learning
Bita Darvish Rouhani, Ritchie Zhao, Ankit More +30
Narrow bit-width data formats are key to reducing the computational and storage costs of modern deep learning applications. This paper evaluates Microscaling (MX) data formats that…
With Shared Microexponents, A Little Shifting Goes a Long Way
Bita Rouhani, Ritchie Zhao, Venmugil Elango +19
This paper introduces Block Data Representations (BDR), a framework for exploring and evaluating a wide spectrum of narrow-precision formats for deep learning. It enables compariso…