5 papers
PEAfowl: Perception-Enhanced Multi-View Vision-Language-Action for Bimanual Manipulation
Qingyu Fan, Zhaoxiang Li, Yi Lu +7
Bimanual manipulation in cluttered scenes requires policies that remain stable under occlusions, viewpoint and scene variations. Existing vision-language-action models often fail t…
QuoTA: Query-oriented Token Assignment via CoT Query Decouple for Long Video Comprehension
Yongdong Luo, Wang Chen, Xiawu Zheng +8
Recent advances in long video understanding typically mitigate visual redundancy through visual token pruning based on attention distribution. However, while existing methods emplo…
LaVida Drive: Vision-Text Interaction VLM for Autonomous Driving with Token Selection, Recovery and Enhancement
Siwen Jiao, Yangyi Fang, Baoyun Peng +2
Recent advancements in Visual Language Models (VLMs) have made them crucial for visual question answering (VQA) in autonomous driving, enabling natural human-vehicle interactions.…
Point-PRC: A Prompt Learning Based Regulation Framework for Generalizable Point Cloud Analysis
Hongyu Sun, Qiuhong Ke, Yongcai Wang +4
This paper investigates the 3D domain generalization (3DDG) ability of large 3D models based on prevalent prompt learning. Recent works demonstrate the performances of 3D point clo…
Multi-Task Model Merging via Adaptive Weight Disentanglement
Feng Xiong, Runxi Cheng, Wang Chen +4
Model merging has recently gained attention as an economical and scalable approach to incorporate task-specific weights from various tasks into a unified multi-task model. For exam…