2 papers
cs.DC2025
SpeedMalloc: Improving Multi-threaded Applications via a Lightweight Core for Memory Allocation
Ruihao Li, Qinzhe Wu, Krishna Kavi +4
Memory allocation, though constituting only a small portion of the executed code, can have a "butterfly effect" on overall program performance, leading to significant and far-reach…
cs.AR2019
AMOEBA: A Coarse Grained Reconfigurable Architecture for Dynamic GPU Scaling
Xianwei Cheng, Hui Zhao, Mahmut Kandemir +2
Different GPU applications exhibit varying scalability patterns with network-on-chip (NoC), coalescing, memory and control divergence, and L1 cache behavior. A GPU consists of seve…