3 papers
cs.LG2026
InnerQ: Hardware-Aware Tuning-Free Quantization of KV Cache for Large Language Models
Sayed Mohammadreza Tayaranian Hosseini, Amir Ardakani, Warren J. Gross
When transformer-based language models are deployed for text generation, most of the inference time is spent in the decoding stage, where output tokens are generated sequentially.…
cs.AR2026
Memory-Efficient FPGA Implementation of Stochastic Simulated Annealing
Duckgyu Shin, Naoya Onizawa, Warren J. Gross +1
Simulated annealing (SA) is a well-known algorithm for solving combinatorial optimization problems. However, the computation time of SA increases rapidly, as the size of the proble…
cs.CV2024
BD-KD: Balancing the Divergences for Online Knowledge Distillation
Ibtihel Amara, Nazanin Sepahvand, Brett H. Meyer +2
We address the challenge of producing trustworthy and accurate compact models for edge devices. While Knowledge Distillation (KD) has improved model compression in terms of achievi…