2 papers
cs.PF2024
AutoChunk: Automated Activation Chunk for Memory-Efficient Long Sequence Inference
Xuanlei Zhao, Shenggan Cheng, Guangyang Lu +5
Large deep learning models have achieved impressive performance across a range of applications. However, their large memory requirements, including parameter memory and activation…
cs.PF2024
HeteGen: Heterogeneous Parallel Inference for Large Language Models on Resource-Constrained Devices
Xuanlei Zhao, Bin Jia, Haotian Zhou +3
In recent times, the emergence of Large Language Models (LLMs) has resulted in increasingly larger model size, posing challenges for inference on low-resource devices. Prior approa…