4 papers · 1 filter
LEEVLA: Seeing What Matters in Latent Environment Evolution for Vision-Language-Action
Qi Lyu, Baicheng Liu, Xudong Wang +3
Vision-language-action (VLA) models aim to map multimodal inputs to robot actions. However, most existing approaches struggle to cover complex dynamic scenarios due to treating all…
SAB-LVLM: Significance-Aware Binarization for Large Vision-Language Models
Qi Lyu, Jiahua Dong, Baichen Liu +7
Large Vision-Language Models (LVLMs) have achieved remarkable progress in multimodal understanding, yet their enormous parameter scale and cross-modal computation incur substantial…
CRISP: Contrastive Residual Injection and Semantic Prompting for Continual Video Instance Segmentation
Baichen Liu, Qi Lyu, Xudong Wang +3
Continual video instance segmentation demands both the plasticity to absorb new object categories and the stability to retain previously learned ones, all while preserving temporal…
Review helps learn better: Temporal Supervised Knowledge Distillation
Dongwei Wang, Zhi Han, Yanmei Wang +3
Reviewing plays an important role when learning knowledge. The knowledge acquisition at a certain time point may be strongly inspired with the help of previous experience. Thus the…