3 papers
cs.DC2026
From Cloud to Crowd: Democratizing LLM Service with Decentralized Edge Collaboration for RAG
Jiaxing Li, Hengzhi Wang, Feng Wang +5
The rapid advancement of large language models (LLMs) has increased demand for scalable and cost-effective deployment, especially for mobile and edge devices. Cloud-hosted LLMs are…
cs.CV2026
QuickGrasp: Responsive Video-Language Querying Service via Accelerated Tokenization and Edge-Augmented Inference
Miao Zhang, Ruixiao Zhang, Jianxin Shi +3
Video-language models (VLMs) are reshaping video querying services, bringing unified solutions to complex perception and reasoning tasks. However, deploying large VLMs in real-worl…
cs.NI2026
ViTMAlis: Towards Latency-Critical Mobile Video Analytics with Vision Transformers
Miao Zhang, Guanzhen Wu, Hao Fang +4
Edge-assisted mobile video analytics (MVA) applications are increasingly shifting from using vision models based on convolutional neural networks (CNNs) to those built on vision tr…