activity
20242026
collaborators

9 papers

cs.LG2026

FASQ: Flexible Accelerated Subspace Quantization for Calibration-Free LLM Compression

Ye Qiao, Yian Wang, Zhiheng Chen +2

Compressing large language models (LLMs) for deployment on commodity GPUs remains challenging: conventional scalar quantization is limited to fixed bit-widths (e.g., 8/4/3-bit), of…

cs.AR2026

Characterizing State Space Model and Hybrid Language Model Performance with Long Context

Saptarshi Mitra, Rachid Karami, Haocheng Xu +2

Emerging applications such as AR are driving demands for machine intelligence capable of processing continuous and/or long-context inputs on local devices. However, currently domin…

cs.PF2025

RAPID-LLM: Resilience-Aware Performance analysis of Infrastructure for Distributed LLM Training and Inference

George Karfakis, Faraz Tahmasebi, Binglu Chen +5

RAPID-LLM is a unified performance modeling framework for distributed large language model (LLM) training and inference on GPU clusters, without relying on deployment-specific trac…

cs.AR2025

D-com: Accelerating Iterative Processing to Enable Low-rank Decomposition of Activations

Faraz Tahmasebi, Michael Pelluer, Hyoukjun Kwon

The computation and memory costs of large language models kept increasing over last decade, which reached over the scale of 1T parameters. To address the challenges from the large…

cs.LG2025

Exploring the Dynamic Scheduling Space of Real-Time Generative AI Applications on Emerging Heterogeneous Systems

Rachid Karami, Rajeev Patwari, Hyoukjun Kwon +1

The integration of generative AI models, particularly large language models (LLMs), into real-time multi-model AI applications such as video conferencing and gaming is giving rise…

cs.AR2025

FlexiBit: Fully Flexible Precision Bit-parallel Accelerator Architecture for Arbitrary Mixed Precision AI

Faraz Tahmasebi, Yian Wang, Benji Y. H. Huang +1

Recent research has shown that large language models (LLMs) can utilize low-precision floating point (FP) quantization to deliver high efficiency while maintaining original model a…