8 citations · 16 across the 3 of their papers we have counts for
6 papers
MANA: Microarchitecting an Instruction Prefetcher
Ali Ansari, Fatemeh Golshan, Pejman Lotfi-Kamran +1
L1 instruction (L1-I) cache misses are a source of performance bottleneck. Sequential prefetchers are simple solutions to mitigate this problem; however, prior work has shown that…
A Survey on Recent Hardware Data Prefetching Approaches with An Emphasis on Servers
Mohammad Bakhshalipour, Mehran Shakerinava, Fatemeh Golshan +3
Data prefetching, i.e., the act of predicting application's future memory accesses and fetching those that are not in the on-chip caches, is a well-known and widely-used approach t…
ORIGAMI: A Heterogeneous Split Architecture for In-Memory Acceleration of Learning
Hajar Falahati, Pejman Lotfi-Kamran, Mohammad Sadrosadati +1
Memory bandwidth bottleneck is a major challenges in processing machine learning (ML) algorithms. In-memory acceleration has potential to address this problem; however, it needs to…
Die-Stacked DRAM: Memory, Cache, or MemCache?
Mohammad Bakhshalipour, HamidReza Zare, Pejman Lotfi-Kamran +1
Die-stacked DRAM is a promising solution for satisfying the ever-increasing memory bandwidth requirements of multi-core processors. Manufacturing technology has enabled stacking se…
Making Belady-Inspired Replacement Policies More Effective Using Expected Hit Count
Seyed Armin Vakil Ghahani, Sara Mahdizadeh Shahri, Mohammad Bakhshalipour +2
Memory-intensive workloads operate on massive amounts of data that cannot be captured by last-level caches (LLCs) of modern processors. Consequently, processors encounter frequent…
Scale-Out Processors & Energy Efficiency
Pouya Esmaili-Dokht, Mohammad Bakhshalipour, Behnam Khodabandeloo +2
Scale-out workloads like media streaming or Web search serve millions of users and operate on a massive amount of data, and hence, require enormous computational power. As the numb…