5 papers
Image Generators are Generalist Vision Learners
Valentin Gabeur, Shangbang Long, Songyou Peng +22
Recent works show that image and video generators exhibit zero-shot visual understanding behaviors, in a way reminiscent of how LLMs develop emergent capabilities of language under…
Compression Tells Intelligence: Visual Coding, Visual Token Technology, and the Unification
Xin Jin, Jinming Liu, Yuntao Wei +6
"Compression Tells Intelligence", is supported by research in artificial intelligence, particularly concerning (multimodal) large language models (LLMs/MLLMs), where compression ef…
Object-Centric Mobile Manipulation through SAM2-Guided Perception and Imitation Learning
Wang Zhicheng, Satoshi Yagi, Satoshi Yamamori +1
Imitation learning for mobile manipulation is a key challenge in the field of robotic manipulation. However, current mobile manipulation frameworks typically decouple navigation an…
IPFormer-VideoLLM: Enhancing Multi-modal Video Understanding for Multi-shot Scenes
Yujia Liang, Jile Jiao, Xuetao Feng +3
Video Large Language Models (VideoLLMs) have demonstrated remarkable understanding capabilities, but are found struggling to tackle multi-shot scenarios,e.g., video clips with vary…
FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering
Zheng Cheng, Rendong Wang, Zhicheng Wang
Recently, multi-modal large language models have made significant progress. However, visual information lacking of guidance from the user's intention may lead to redundant computat…