2 papers
cs.CL2025
SASQ: Static Activation Scaling for Quantization-Aware Training in Large Language Models
Shizhuo Mao, Song Chen, Yi Kang
Large language models (LLMs) excel at natural language tasks but face deployment challenges due to their growing size outpacing GPU memory advancements. Model quantization mitigate…
cs.DC2024
SparseMap: Loop Mapping for Sparse CNNs on Streaming Coarse-grained Reconfigurable Array
Xiaobing Ni, Mengke Ge, Jiaheng Ruan +2
Streaming coarse-grained reconfgurable array (CGRA) is a promising architecture for data/computing-intensive applications because of its fexibility, high throughput and efcient mem…