activity
20172026
most citedScaleCom: Scalable Sparsified Gradient Compression for Communication-Efficient Distributed Training

13 citations · 25 across the 7 of their papers we have counts for

collaborators

11 papers

cs.LG2026

Is Finer Better? The Limits of Microscaling Formats in Large Language Models

Andrea Fasoli, Monodeep Kar, Chi-Chun Liu +4

Microscaling data formats leverage per-block tensor quantization to enable aggressive model compression with limited loss in accuracy. Unlocking their potential for efficient train…

cs.AR2025

Making Strong Error-Correcting Codes Work Effectively for HBM in AI Inference

Rui Xie, Yunhua Fang, Asad Ul Haq +5

LLM inference is increasingly memory bound, and HBM cost per GB dominates system cost. Current HBM stacks include short on-die ECC that tightens binning, raises price, and fixes re…

cs.AR2025

Breaking the HBM Bit Cost Barrier: Domain-Specific ECC for AI Inference Infrastructure

Rui Xie, Asad Ul Haq, Yunhua Fang +5

High-Bandwidth Memory (HBM) delivers exceptional bandwidth and energy efficiency for AI workloads, but its high cost per bit, driven in part by stringent on-die reliability require…

cs.LG2024

SmartQuant: CXL-based AI Model Store in Support of Runtime Configurable Weight Quantization

Rui Xie, Asad Ul Haq, Linsen Ma +5

Recent studies have revealed that, during the inference on generative AI models such as transformer, the importance of different weights exhibits substantial context-dependent vari…

cs.AR20227 cited

Approximate Computing and the Efficient Machine Learning Expedition

Jörg Henkel, Hai Li, Anand Raghunathan +4

Approximate computing (AxC) has been long accepted as a design alternative for efficient system implementation at the cost of relaxed accuracy requirements. Despite the AxC researc…

cs.CL2021

4-bit Quantization of LSTM-based Speech Recognition Models

Andrea Fasoli, Chia-Yu Chen, Mauricio Serrano +9

We investigate the impact of aggressive low-precision representations of weights and activations in two families of large LSTM-based architectures for Automatic Speech Recognition…