9 citations · 9 across the 6 of their papers we have counts for
4 papers · 1 filter
Repeating Words for Video-Language Retrieval with Coarse-to-Fine Objectives
Haoyu Zhao, Jiaxi Gu, Shicong Wang +4
The explosive growth of video streaming presents challenges in achieving high accuracy and low training costs for video-language retrieval. However, existing methods rely on large-…
AutoTVG: A New Vision-language Pre-training Paradigm for Temporal Video Grounding
Xing Zhang, Jiaxi Gu, Haoyu Zhao +6
Temporal Video Grounding (TVG) aims to localize a moment from an untrimmed video given the language description. Since the annotation of TVG is labor-intensive, TVG under limited s…
Reuse and Diffuse: Iterative Denoising for Text-to-Video Generation
Jiaxi Gu, Shicong Wang, Haoyu Zhao +7
Inspired by the remarkable success of Latent Diffusion Models (LDMs) for image synthesis, we study LDM for text-to-video generation, which is a formidable challenge due to the comp…
Towards Universal Vision-language Omni-supervised Segmentation
Bowen Dong, Jiaxi Gu, Jianhua Han +2
Existing open-world universal segmentation approaches usually leverage CLIP and pre-computed proposal masks to treat open-world segmentation tasks as proposal classification. Howev…