69 citations · 83 across the 8 of their papers we have counts for
6 papers · 1 filter
ChatVideo: A Tracklet-centric Multimodal and Versatile Video Understanding System
Junke Wang, Dongdong Chen, Chong Luo +4
Existing deep video models are limited by specific tasks, fixed input-output spaces, and poor generalization capabilities, making it difficult to deploy them in real-world scenario…
LaMD: Latent Motion Diffusion for Image-Conditional Video Generation
Yaosi Hu, Zhenzhong Chen, Chong Luo
The video generation field has witnessed rapid improvements with the introduction of recent diffusion models. While these models have successfully enhanced appearance quality, they…
Streaming Video Model
Yucheng Zhao, Chong Luo, Chuanxin Tang +3
Video understanding tasks have traditionally been modeled by two separate architectures, specially tailored for two distinct tasks. Sequence-based video tasks, such as action recog…
Universal Few-shot Learning of Dense Prediction Tasks with Visual Token Matching
Donggyun Kim, Jinwoo Kim, Seongwoong Cho +2
Dense prediction tasks are a fundamental class of problems in computer vision. As supervised methods suffer from high pixel-wise labeling cost, a few-shot learning solution that ca…
OmniTracker: Unifying Object Tracking by Tracking-with-Detection
Junke Wang, Zuxuan Wu, Dongdong Chen +4
Visual Object Tracking (VOT) aims to estimate the positions of target objects in a video sequence, which is an important vision task with various real-world applications. Depending…
OmniVL:One Foundation Model for Image-Language and Video-Language Tasks
Junke Wang, Dongdong Chen, Zuxuan Wu +7
This paper presents OmniVL, a new foundation model to support both image-language and video-language tasks using one universal architecture. It adopts a unified transformer-based v…