3 papers
cs.CV2026
Listening makes Vision Clear for VLMs
Yiyang Chen, Yixin Tan, Binrui Shen
Recent work typically assesses vision--language consistency using attention distributions of answer-side tokens. However, we observe that highest attention regions are not always c…
cs.RO2026
World-Language-Action Model for Unified World Modeling, Language Reasoning, and Action Synthesis
Yi Yang, Zhihong Liu, Siqi Kou +9
We propose world-language-action (WLA) models as a new class of embodied foundation models. WLA takes textual instructions, images, and robot states as inputs to jointly predict te…
cs.LG2025
LibContinual: A Comprehensive Library towards Realistic Continual Learning
Wenbin Li, Shangge Liu, Borui Kang +7
A fundamental challenge in Continual Learning (CL) is catastrophic forgetting, where adapting to new tasks degrades the performance on previous ones. While the field has evolved wi…