4 papers
Fast and Accurate: An Adaptive VLA Inference Framework through Environment-aware Model Selection
Yuewei Sun, Lang Qin, Zechuan Tian +11
Embodied intelligence demands both long-horizon reasoning and real-time closed-loop responsiveness. Recent dual-system Vision-Language-Action (VLA) architectures combine fast react…
MAR: Efficient Large Language Models via Module-aware Architecture Refinement
Junhong Cai, Guiqin Wang, Kejie Zhao +6
Large Language Models (LLMs) excel across diverse domains but suffer from high energy costs due to quadratic attention and dense Feed-Forward Network (FFN) operations. To address t…
EdgeSync: Accelerating Edge-Model Updates for Data Drift through Adaptive Continuous Learning
Runchu Donga, Peng Zhao, Guiqin Wang +2
Real-time video analytics systems typically deploy lightweight models on edge devices to reduce latency. However, the distribution of data features may change over time due to vari…
Generative Model-Based Feature Attention Module for Video Action Analysis
Guiqin Wang, Peng Zhao, Cong Zhao +3
Video action analysis is a foundational technology within the realm of intelligent video comprehension, particularly concerning its application in Internet of Things(IoT). However,…