7 citations · 7 across the 3 of their papers we have counts for
3 papers
cs.CV2024
OmniVid: A Generative Framework for Universal Video Understanding
Junke Wang, Dongdong Chen, Chong Luo +4
The core of video understanding tasks, such as recognition, captioning, and tracking, is to automatically detect objects or actions in a video and analyze their temporal evolution.…
cs.CV2024
MouSi: Poly-Visual-Expert Vision-Language Models
Xiaoran Fan, Tao Ji, Changhao Jiang +21
Current large vision-language models (VLMs) often encounter challenges such as insufficient capabilities of a single visual component and excessively long visual tokens. These issu…
cs.CV2023★ 7 cited
ChatVideo: A Tracklet-centric Multimodal and Versatile Video Understanding System
Junke Wang, Dongdong Chen, Chong Luo +4
Existing deep video models are limited by specific tasks, fixed input-output spaces, and poor generalization capabilities, making it difficult to deploy them in real-world scenario…