collaborators

5 papers

cs.CV2026

SwiftVLM: Efficient Vision-Language Model Inference via Cross-Layer Token Bypass

Chen Qian, Xinran Yu, Danyang Li +4

Visual token pruning is a promising approach for reducing the computational cost of vision-language models (VLMs), and existing methods often rely on early pruning decisions to imp…

cs.CV2025

edgeVLM: Cloud-edge Collaborative Real-time VLM based on Context Transfer

Chen Qian, Xinran Yu, Zewen Huang +6

Vision-Language Models (VLMs) are increasingly deployed in real-time applications such as autonomous driving and human-computer interaction, which demand fast and reliable response…

cs.CV2025

OpenMoCap: Rethinking Optical Motion Capture under Real-world Occlusion

Chen Qian, Danyang Li, Xinran Yu +2

Optical motion capture is a foundational technology driving advancements in cutting-edge fields such as virtual reality and film production. However, system performance suffers sev…

cs.RO2025

OpenMap: Instruction Grounding via Open-Vocabulary Visual-Language Mapping

Danyang Li, Zenghui Yang, Guangpeng Qi +4

Grounding natural language instructions to visual observations is fundamental for embodied agents operating in open-world environments. Recent advances in visual-language mapping h…

cs.LG2025

SpecOffload: Unlocking Latent GPU Capacity for LLM Inference on Resource-Constrained Devices

Xiangwen Zhuge, Xu Shen, Zeyu Wang +6

Efficient LLM inference on resource-constrained devices presents significant challenges in compute and memory utilization. Due to limited GPU memory, existing systems offload model…