10 papers
Task-Oriented Communication for Human Action Understanding via Edge-Cloud Co-Inference
Jingyi Liu, Cheng Yuan, Lijun He +2
The expanding application of smart sensing has created a growing demand for the accurate understanding of human action at the network edge. Traditional approaches require massive v…
VLMQ: Token Saliency-Driven Post-Training Quantization for Vision-language Models
Yufei Xue, Yushi Huang, Jiawei Shao +4
Post-training quantization (PTQ) has emerged as an effective technique for compressing large models and accelerating inference without retraining. While PTQ has been extensively st…
Feed-Forward 3D Gaussian Splatting Compression with Long-Context Modeling
Zhening Liu, Rui Song, Yushi Huang +5
3D Gaussian Splatting (3DGS) has emerged as a revolutionary 3D representation. However, its substantial data size poses a major barrier to widespread adoption. While feed-forward 3…
CAS-Spec: Cascade Adaptive Self-Speculative Decoding for On-the-Fly Lossless Inference Acceleration of LLMs
Zhiyuan Ning, Jiawei Shao, Ruge Xu +4
Speculative decoding has become a widely adopted as an effective technique for lossless inference acceleration when deploying large language models (LLMs). While on-the-fly self-sp…
Task-Oriented Feature Compression for Multimodal Understanding via Device-Edge Co-Inference
Cheng Yuan, Zhening Liu, Jiashu Lv +4
With the rapid development of large multimodal models (LMMs), multimodal understanding applications are emerging. As most LMM inference requests originate from edge devices with li…
WirelessAgent: Large Language Model Agents for Intelligent Wireless Networks
Jingwen Tong, Wei Guo, Jiawei Shao +4
The rapid evolution of wireless networks presents unprecedented challenges in managing complex and dynamic systems. Existing methods are increasingly facing fundamental limitations…