5 papers
Compression Tells Intelligence: Visual Coding, Visual Token Technology, and the Unification
Xin Jin, Jinming Liu, Yuntao Wei +6
"Compression Tells Intelligence", is supported by research in artificial intelligence, particularly concerning (multimodal) large language models (LLMs/MLLMs), where compression ef…
Object-Centric Mobile Manipulation through SAM2-Guided Perception and Imitation Learning
Wang Zhicheng, Satoshi Yagi, Satoshi Yamamori +1
Imitation learning for mobile manipulation is a key challenge in the field of robotic manipulation. However, current mobile manipulation frameworks typically decouple navigation an…
IPFormer-VideoLLM: Enhancing Multi-modal Video Understanding for Multi-shot Scenes
Yujia Liang, Jile Jiao, Xuetao Feng +3
Video Large Language Models (VideoLLMs) have demonstrated remarkable understanding capabilities, but are found struggling to tackle multi-shot scenarios,e.g., video clips with vary…
FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering
Zheng Cheng, Rendong Wang, Zhicheng Wang
Recently, multi-modal large language models have made significant progress. However, visual information lacking of guidance from the user's intention may lead to redundant computat…
Takin: A Cohort of Superior Quality Zero-shot Speech Generation Models
Sijing Chen, Yuan Feng, Laipeng He +22
With the advent of the big data and large language model era, zero-shot personalized rapid customization has emerged as a significant trend. In this report, we introduce Takin Audi…