2 papers
cs.AR2025
Hardware-based Heterogeneous Memory Management for Large Language Model Inference
Soojin Hwang, Jungwoo Kim, Sanghyeon Lee +2
A large language model (LLM) is one of the most important emerging machine learning applications nowadays. However, due to its huge model size and runtime increase of the memory fo…
cs.DC2025
Efficient LLM Inference with Activation Checkpointing and Hybrid Caching
Sanghyeon Lee, Hongbeen Kim, Soojin Hwang +3
Recent large language models (LLMs) with enormous model sizes use many GPUs to meet memory capacity requirements incurring substantial costs for token generation. To provide cost-e…