Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
FASQ: Flexible Accelerated Subspace Quantization for Calibration-Free LLM Compression
Ye Qiao, Yian Wang, Zhiheng Chen +2
Compressing large language models (LLMs) for deployment on commodity GPUs remains challenging: conventional scalar quantization is limited to fixed bit-widths (e.g., 8/4/3-bit), of…
cs.LG2025
Exploring the Dynamic Scheduling Space of Real-Time Generative AI Applications on Emerging Heterogeneous Systems
Rachid Karami, Rajeev Patwari, Hyoukjun Kwon +1
The integration of generative AI models, particularly large language models (LLMs), into real-time multi-model AI applications such as video conferencing and gaming is giving rise…