Showing cs.LGShow all
2 papers · 1 filter
cs.LG2025
Efficient Edge LLMs Deployment via HessianAware Quantization and CPU GPU Collaborative
Tuo Zhang, Ning Li, Xin Yuan +4
With the breakthrough progress of large language models (LLMs) in natural language processing and multimodal tasks, efficiently deploying them on resource-constrained edge devices…
cs.LG2024
A QoE-Aware Split Inference Accelerating Algorithm for NOMA-based Edge Intelligence
Xin Yuan, Ning Li, Quan Chen +3
Even the AI has been widely used and significantly changed our life, deploying the large AI models on resource limited edge devices directly is not appropriate. Thus, the model spl…