6 papers
Fast and Accurate: An Adaptive VLA Inference Framework through Environment-aware Model Selection
Yuewei Sun, Lang Qin, Zechuan Tian +11
Embodied intelligence demands both long-horizon reasoning and real-time closed-loop responsiveness. Recent dual-system Vision-Language-Action (VLA) architectures combine fast react…
Adaptive Depth Sparse Framework: Similarity-Driven Resource Allocation for Pre-Trained LLMs
Yidu Wu, Xiang Wang, Kejie Zhao +3
Large language models (LLMs) achieve strong generation and reasoning performance, but the Transformer architecture incurs high inference cost. Existing acceleration methods often r…
Threshold Modulation for Online Test-Time Adaptation of Spiking Neural Networks
Kejie Zhao, Wenjia Hua, Aiersi Tuerhong +3
Recently, spiking neural networks (SNNs), deployed on neuromorphic chips, provide highly efficient solutions on edge devices in different scenarios. However, their ability to adapt…
MAR: Efficient Large Language Models via Module-aware Architecture Refinement
Junhong Cai, Guiqin Wang, Kejie Zhao +6
Large Language Models (LLMs) excel across diverse domains but suffer from high energy costs due to quadratic attention and dense Feed-Forward Network (FFN) operations. To address t…
Hebbian Learning with Global Direction
Wenjia Hua, Kejie Zhao, Luziwei Leng +3
Backpropagation algorithm has driven the remarkable success of deep neural networks, but its lack of biological plausibility and high computational costs have motivated the ongoing…
SPikE-SSM: A Sparse, Precise, and Efficient Spiking State Space Model for Long Sequences Learning
Yan Zhong, Ruoyu Zhao, Chao Wang +4
Spiking neural networks (SNNs) provide an energy-efficient solution by utilizing the spike-based and sparse nature of biological systems. Since the advent of Transformers, SNNs hav…