3 papers
cs.AR2026
Rethinking LLM Inference Bottlenecks: Insights from Latent Attention and Mixture-of-Experts
Sungmin Yun, Seonyong Park, Hwayong Nam +10
Computational workloads composing traditional transformer models are starkly bifurcated. Multi-Head Attention (MHA) and Grouped-Query Attention are memory-bound due to low arithmet…
cs.CR2025
SoK: Systematizing a Decade of Architectural RowHammer Defenses Through the Lens of Streaming Algorithms
Michael Jaemin Kim, Seungmin Baek, Jumin Kim +3
A decade after its academic introduction, RowHammer (RH) remains a moving target that continues to challenge both the industry and academia. With its potential to serve as a critic…
cs.CR2024
DRAMScope: Uncovering DRAM Microarchitecture and Characteristics by Issuing Memory Commands
Hwayong Nam, Seungmin Baek, Minbok Wi +5
The demand for precise information on DRAM microarchitectures and error characteristics has surged, driven by the need to explore processing in memory, enhance reliability, and mit…