6 citations · 8 across the 2 of their papers we have counts for
3 papers
cs.AR2025★ 2 cited
SSD Offloading for LLM Mixture-of-Experts Weights Considered Harmful in Energy Efficiency
Kwanhee Kyung, Sungmin Yun, Jung Ho Ahn
Large Language Models (LLMs) applying Mixture-of-Experts (MoE) scale to trillions of parameters but require vast memory, motivating a line of research to offload expert weights fro…
cs.AR2025
Rethinking LLM Inference Bottlenecks: Insights from Latent Attention and Mixture-of-Experts
Sungmin Yun, Seonyong Park, Hwayong Nam +10
Computational workloads composing traditional transformer models are starkly bifurcated. Multi-Head Attention (MHA) and Grouped-Query Attention are memory-bound due to low arithmet…
cs.AR2025★ 6 cited
Cosmos: A CXL-Based Full In-Memory System for Approximate Nearest Neighbor Search
Seoyoung Ko, Hyunjeong Shim, Wanju Doh +8
Retrieval-Augmented Generation (RAG) is crucial for improving the quality of large language models by injecting proper contexts extracted from external sources. RAG requires high-t…