4 papers
VLN-Cache: Enabling Token Caching for VLN Models with Visual/Semantic Dynamics Awareness
Zihao Zheng, Zhihao Mao, Xingyue Zhou +9
Vision-and-Language Navigation (VLN) increasingly relies on large vision-language models, but their inference cost conflicts with real-time deployment. Token caching is a promising…
GOMA: Geometrically Optimal Mapping via Analytical Modeling for Spatial Accelerators
Wulve Yang, Hailong Zou, Rui Zhou +5
General matrix multiplication (GEMM) on spatial accelerators is highly sensitive to mapping choices in both execution efficiency and energy consumption. However, the mapping space…
DyQ-VLA: Temporal-Dynamic-Aware Quantization for Embodied Vision-Language-Action Models
Zihao Zheng, Hangyu Cao, Sicheng Tian +9
Vision-Language-Action (VLA) models are dominant in embodied intelligence but are constrained by inference overheads. While model quantization alleviates these bottlenecks for edge…
RAPID: Redundancy-Aware and Compatibility-Optimal Edge-Cloud Partitioned Inference for Diverse VLA Models
Zihao Zheng, Sicheng Tian, Hangyu Cao +7
Vision Language Action (VLA) models are mainstream in embodied intelligence but face high inference costs. Edge-Cloud Collaborative (ECC) inference offers an effective fix by easin…