2 papers
cs.CV2025
edgeVLM: Cloud-edge Collaborative Real-time VLM based on Context Transfer
Chen Qian, Xinran Yu, Zewen Huang +6
Vision-Language Models (VLMs) are increasingly deployed in real-time applications such as autonomous driving and human-computer interaction, which demand fast and reliable response…
cs.LG2025
SpecOffload: Unlocking Latent GPU Capacity for LLM Inference on Resource-Constrained Devices
Xiangwen Zhuge, Xu Shen, Zeyu Wang +6
Efficient LLM inference on resource-constrained devices presents significant challenges in compute and memory utilization. Due to limited GPU memory, existing systems offload model…