3 papers
cs.AR2026
Pickle: Precise, Flexible Cross-Core Last-level Cache Data Prefetching for Irregular Memory Accesses
Hoa Nguyen, Pongstorn Maidee, Jason Lowe-Power +1
Graph analytics and sparse scientific workloads are dominated by parallel chains of data-dependent, long-latency memory accesses whose patterns are difficult for hardware to infer…
cs.AR2026
CXL-ClusterSim: Modeling CXL-based Disaggregated Memory Cluster for Pooling and Sharing using gem5 and SST
Kaustav Goswami, Maryam Babaie, Hoa Nguyen +2
Large-scale AI training and inference require hundreds of gigabytes to terabytes of DRAM with high peak to average utilization ratios, resulting in overprovisioning. In cloud compu…
cs.AR2025
Choreographer: A Full-System Framework for Fine-Grained Tasks in Cache Hierarchies
Hoa Nguyen, Pongstorn Maidee, Jason Lowe-Power +1
In this paper, we introduce Choreographer, a simulation framework that enables a holistic system-level evaluation of fine-grained accelerators designed for latency-sensitive tasks.…