6 papers
Seeing Touch from Motion: A Unified Modality-Aware Visuo-Tactile Policy with Tactile Motion Correlation
Shengqi Xu, Guojin Zhong, Yang Liu +7
Visuo-Tactile policies leveraging optical tactile sensors have shown great promise in contact-rich manipulation. These sensors achieve high spatial resolution and multi-dimensional…
CCRC: A Change-Aware Captioning and Reasoning Chain for Image Change Captioning and Segmentation
Jinhong Hu, Xiaoping Wang, Shuyin Huang +3
Understanding and localizing subtle changes between paired images is critical for tasks such as surveillance and image editing. However, traditional Image Change Captioning (ICC) m…
ThinkingVLA: Interleaved Vision and Language Reasoning for Robotic Manipulation
Tianyi Lu, Hui Zhang, Zijie Diao +8
Most Vision-Language-Action (VLA) models map observations directly to actions without explicit reasoning, limiting their capacity for reasoning-intensive long-horizon tasks. To add…
ActiveMimic: Egocentric Video Pretraining with Active Perception
Xingyao Lin, Guojin Zhong, Tianyi Lu +4
Egocentric human video offers a scalable alternative to robot data for pretraining, yet models pretrained on such video consistently underperform those pretrained on robot data. We…
SGDiff: Scene Graph Guided Diffusion Model for Image Collaborative SegCaptioning
Xu Zhang, Jin Yuan, Hanwang Zhang +4
Controllable image semantic understanding tasks, such as captioning or segmentation, necessitate users to input a prompt (e.g., text or bounding boxes) to predict a unique outcome,…
AVAM: Universal Training-free Adaptive Visual Anchoring Embedded into Multimodal Large Language Model for Multi-image Question Answering
Kang Zeng, Guojin Zhong, Jintao Cheng +2
The advancement of Multimodal Large Language Models (MLLMs) has driven significant progress in Visual Question Answering (VQA), evolving from Single to Multi Image VQA (MVQA). Howe…