10 papers
Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation
Hongyu Qu, Jianzhe Gao, Xiaobin Hu +6
Mainstream Vision-Language-Action (VLA) models predict actions primarily from the current observation under a Markovian assumption, thus struggling with long-horizon, temporally de…
Learning 3D Representations for Spatial Intelligence from Unposed Multi-View Images
Bo Zhou, Qiuxia Lai, Zeren Sun +3
Robust 3D representation learning forms the perceptual foundation of spatial intelligence, enabling downstream tasks in scene understanding and embodied AI. However, learning such…
Bilingual Text-to-Motion Generation: A New Benchmark and Baselines
Wanjiang Weng, Xiaofeng Tan, Xiangbo Shu +3
Text-to-motion generation holds significant potential for cross-linguistic applications, yet it is hindered by the lack of bilingual datasets and the poor cross-lingual semantic un…
Beyond Quadratic: Linear-Time Change Detection with RWKV
Zhenyu Yang, Gensheng Pei, Tao Chen +4
Existing paradigms for remote sensing change detection are caught in a trade-off: CNNs excel at efficiency but lack global context, while Transformers capture long-range dependenci…
PCA-Seg: Revisiting Cost Aggregation for Open-Vocabulary Semantic and Part Segmentation
Jianjian Yin, Tao Chen, Yi Chen +4
Recent advances in vision-language models (VLMs) have garnered substantial attention in open-vocabulary semantic and part segmentation (OSPS). However, existing methods extract ima…
Spatio-temporal Decoupled Knowledge Compensator for Few-Shot Action Recognition
Hongyu Qu, Xiangbo Shu, Rui Yan +3
Few-Shot Action Recognition (FSAR) is a challenging task that requires recognizing novel action categories with a few labeled videos. Recent works typically apply semantically coar…