6 papers
SSI-Policy: Learning Structured Scene Interfaces for Vision-Language Robotic Manipulation
Kaijun Wang, Zikai Ouyang, Xuping Wu +6
Real-world robotic manipulation demands spatial grounding, task-aware reasoning, and precise control. Learning such capabilities becomes particularly challenging in the low-data re…
From Video to Control: A Survey of Learning Manipulation Interfaces from Temporal Visual Data
Linfang Zheng, Zikai Ouyang, Chen Wang +2
Video is a scalable observation of physical dynamics: it captures how objects move, how contact unfolds, and how scenes evolve under interaction -- all without requiring robot acti…
State-Anchored Complete-View Distillation for Robust Conversational Multimodal Emotion Recognition
Zhaoyan Pan, Xiangdong Li, Wenke Wu +5
Conversational multimodal emotion recognition (MER) requires reliable prediction when language, acoustic, or visual observations are missing or unreliable. Many missing-modality me…
Masks Can Talk: Extracting Structured Text Information from Single-Modal Images for Remote Sensing Change Detection
Kai Zheng, Hang-Cheng Dong, Jiatong Pan +3
Remote sensing change detection is pivotal for urban monitoring, disaster assessment, and environmental resource management. Yet, unimodal deep learning methods frequently confuse…
Beyond Isolated Utterances: Cue-Guided Interaction for Context-Dependent Conversational Multimodal Understanding
Zhaoyan Pan, Hengyang Zhou, Xiangdong Li +5
Conversational multimodal understanding aims to infer the meaning or label of the current utterance from its preceding dialogue context together with textual, acoustic, and visual…
Generative Artificial Intelligence in Robotic Manipulation: A Survey
Kun Zhang, Peng Yun, Jun Cen +11
This survey provides a comprehensive review on recent advancements of generative learning models in robotic manipulation, addressing key challenges in the field. Robotic manipulati…