Showing 2025Show all
2 papers · 1 filter
cs.RO2025
Object-Centric Mobile Manipulation through SAM2-Guided Perception and Imitation Learning
Wang Zhicheng, Satoshi Yagi, Satoshi Yamamori +1
Imitation learning for mobile manipulation is a key challenge in the field of robotic manipulation. However, current mobile manipulation frameworks typically decouple navigation an…
cs.CV2025
IPFormer-VideoLLM: Enhancing Multi-modal Video Understanding for Multi-shot Scenes
Yujia Liang, Jile Jiao, Xuetao Feng +3
Video Large Language Models (VideoLLMs) have demonstrated remarkable understanding capabilities, but are found struggling to tackle multi-shot scenarios,e.g., video clips with vary…