collaborators

6 papers

cs.RO2026

Seeing Touch from Motion: A Unified Modality-Aware Visuo-Tactile Policy with Tactile Motion Correlation

Shengqi Xu, Guojin Zhong, Yang Liu +7

Visuo-Tactile policies leveraging optical tactile sensors have shown great promise in contact-rich manipulation. These sensors achieve high spatial resolution and multi-dimensional…

cs.CV2026

CCRC: A Change-Aware Captioning and Reasoning Chain for Image Change Captioning and Segmentation

Jinhong Hu, Xiaoping Wang, Shuyin Huang +3

Understanding and localizing subtle changes between paired images is critical for tasks such as surveillance and image editing. However, traditional Image Change Captioning (ICC) m…

cs.RO2026

ThinkingVLA: Interleaved Vision and Language Reasoning for Robotic Manipulation

Tianyi Lu, Hui Zhang, Zijie Diao +8

Most Vision-Language-Action (VLA) models map observations directly to actions without explicit reasoning, limiting their capacity for reasoning-intensive long-horizon tasks. To add…

cs.RO2026

ActiveMimic: Egocentric Video Pretraining with Active Perception

Xingyao Lin, Guojin Zhong, Tianyi Lu +4

Egocentric human video offers a scalable alternative to robot data for pretraining, yet models pretrained on such video consistently underperform those pretrained on robot data. We…

cs.CV2025

SGDiff: Scene Graph Guided Diffusion Model for Image Collaborative SegCaptioning

Xu Zhang, Jin Yuan, Hanwang Zhang +4

Controllable image semantic understanding tasks, such as captioning or segmentation, necessitate users to input a prompt (e.g., text or bounding boxes) to predict a unique outcome,…

cs.CV2025

AVAM: Universal Training-free Adaptive Visual Anchoring Embedded into Multimodal Large Language Model for Multi-image Question Answering

Kang Zeng, Guojin Zhong, Jintao Cheng +2

The advancement of Multimodal Large Language Models (MLLMs) has driven significant progress in Visual Question Answering (VQA), evolving from Single to Multi Image VQA (MVQA). Howe…