activity
20242026
collaborators

5 papers

cs.AR2026

Pickle: Precise, Flexible Cross-Core Last-level Cache Data Prefetching for Irregular Memory Accesses

Hoa Nguyen, Pongstorn Maidee, Jason Lowe-Power +1

Graph analytics and sparse scientific workloads are dominated by parallel chains of data-dependent, long-latency memory accesses whose patterns are difficult for hardware to infer…

cs.AR2026

Portable Targeted Sampling Framework Using LLVM

Zhantong Qiu, Mahyar Samani, Jason Lowe-Power

Evaluating architectural ideas on realistic workloads is increasingly challenging due to the prohibitive cost of detailed simulation and the lack of portable sampling tools. Existi…

cs.PF2025

Characterizing GPU Energy Usage in Exascale-Ready Portable Science Applications

William F. Godoy, Oscar Hernandez, Paul R. C. Kent +8

We characterize the GPU energy usage of two widely adopted exascale-ready applications representing two classes of particle and mesh solvers: (i) QMCPACK, a quantum Monte Carlo pac…

cs.AR2025

Choreographer: A Full-System Framework for Fine-Grained Tasks in Cache Hierarchies

Hoa Nguyen, Pongstorn Maidee, Jason Lowe-Power +1

In this paper, we introduce Choreographer, a simulation framework that enables a holistic system-level evaluation of fine-grained accelerators designed for latency-sensitive tasks.…

cs.CR2024

FP-Rowhammer: DRAM-Based Device Fingerprinting

Hari Venugopalan, Kaustav Goswami, Zainul Abi Din +3

Device fingerprinting leverages attributes that capture heterogeneity in hardware and software configurations to extract unique and stable fingerprints. Fingerprinting countermeasu…