7 citations · 7 across the 3 of their papers we have counts for
4 papers · 1 filter
OmniTokenizer: A Joint Image-Video Tokenizer for Visual Generation
Junke Wang, Yi Jiang, Zehuan Yuan +3
Tokenizer, serving as a translator to map the intricate visual data into a compact latent space, lies at the core of visual generative models. Based on the finding that existing to…
OmniVid: A Generative Framework for Universal Video Understanding
Junke Wang, Dongdong Chen, Chong Luo +4
The core of video understanding tasks, such as recognition, captioning, and tracking, is to automatically detect objects or actions in a video and analyze their temporal evolution.…
MouSi: Poly-Visual-Expert Vision-Language Models
Xiaoran Fan, Tao Ji, Changhao Jiang +21
Current large vision-language models (VLMs) often encounter challenges such as insufficient capabilities of a single visual component and excessively long visual tokens. These issu…
ChatVideo: A Tracklet-centric Multimodal and Versatile Video Understanding System
Junke Wang, Dongdong Chen, Chong Luo +4
Existing deep video models are limited by specific tasks, fixed input-output spaces, and poor generalization capabilities, making it difficult to deploy them in real-world scenario…