2 papers
cs.AR2024
Duplex: A Device for Large Language Models with Mixture of Experts, Grouped Query Attention, and Continuous Batching
Sungmin Yun, Kwanhee Kyung, Juhwan Cho +6
Large language models (LLMs) have emerged due to their capability to generate high-quality content across diverse contexts. To reduce their explosively increasing demands for compu…
cs.CR2024
DRAMScope: Uncovering DRAM Microarchitecture and Characteristics by Issuing Memory Commands
Hwayong Nam, Seungmin Baek, Minbok Wi +5
The demand for precise information on DRAM microarchitectures and error characteristics has surged, driven by the need to explore processing in memory, enhance reliability, and mit…