2 papers
cs.AR2025
Ecco: Improving Memory Bandwidth and Capacity for LLMs via Entropy-aware Cache Compression
Feng Cheng, Cong Guo, Chiyue Wei +7
Large language models (LLMs) have demonstrated transformative capabilities across diverse artificial intelligence applications, yet their deployment is hindered by substantial memo…
cs.AR2019
Thread Batching for High-performance Energy-efficient GPU Memory Design
Bing Li, Mengjie Mao, Xiaoxiao Liu +6
Massive multi-threading in GPU imposes tremendous pressure on memory subsystems. Due to rapid growth in thread-level parallelism of GPU and slowly improved peak memory bandwidth, t…