activity
20242026
collaborators

5 papers

cs.CV2026

Compression Tells Intelligence: Visual Coding, Visual Token Technology, and the Unification

Xin Jin, Jinming Liu, Yuntao Wei +6

"Compression Tells Intelligence", is supported by research in artificial intelligence, particularly concerning (multimodal) large language models (LLMs/MLLMs), where compression ef…

cs.RO2025

Object-Centric Mobile Manipulation through SAM2-Guided Perception and Imitation Learning

Wang Zhicheng, Satoshi Yagi, Satoshi Yamamori +1

Imitation learning for mobile manipulation is a key challenge in the field of robotic manipulation. However, current mobile manipulation frameworks typically decouple navigation an…

cs.CV2025

IPFormer-VideoLLM: Enhancing Multi-modal Video Understanding for Multi-shot Scenes

Yujia Liang, Jile Jiao, Xuetao Feng +3

Video Large Language Models (VideoLLMs) have demonstrated remarkable understanding capabilities, but are found struggling to tackle multi-shot scenarios,e.g., video clips with vary…

cs.CV2024

FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering

Zheng Cheng, Rendong Wang, Zhicheng Wang

Recently, multi-modal large language models have made significant progress. However, visual information lacking of guidance from the user's intention may lead to redundant computat…

cs.SD2024

Takin: A Cohort of Superior Quality Zero-shot Speech Generation Models

Sijing Chen, Yuan Feng, Laipeng He +22

With the advent of the big data and large language model era, zero-shot personalized rapid customization has emerged as a significant trend. In this report, we introduce Takin Audi…