Publications (6)
TokenFlow: Rethinking Fine-grained Cross-modal Alignment in Vision-Language Retrieval
Xiaohan Zou, Changqiao Wu, Lele Cheng +1
Most existing methods in vision-language retrieval match two modalities by either comparing their global feature vectors which misses sufficient information and lacks interpretabil…
Decouple Content and Motion for Conditional Image-to-Video Generation
Cuifeng Shen, Yulu Gan, Chen Chen +4
The goal of conditional image-to-video (cI2V) generation is to create a believable new video by beginning with the condition, i.e., one image and text.The previous cI2V generation…
Paragraph-to-Image Generation with Information-Enriched Diffusion Model
Weijia Wu, Zhuang Li, Yefei He +6
Text-to-image (T2I) models have recently experienced rapid development, achieving astonishing performance in terms of fidelity and textual alignment capabilities. However, given a…
LaT: Latent Translation with Cycle-Consistency for Video-Text Retrieval
Jinbin Bai, Chunhui Liu, Feiyue Ni +4
Video-text retrieval is a class of cross-modal representation learning problems, where the goal is to select the video which corresponds to the text query between a given text quer…
Weakly Supervised Learning with Side Information for Noisy Labeled Images
Lele Cheng, Xiangzeng Zhou, Liming Zhao +5
In many real-world datasets, like WebVision, the performance of DNN based classifier is often limited by the noisy labeled data. To tackle this problem, some image related side inf…
Learning from Large-scale Noisy Web Data with Ubiquitous Reweighting for Image Classification
Jia Li, Yafei Song, Jianfeng Zhu +5
Many advances of deep learning techniques originate from the efforts of addressing the image classification task on large-scale datasets. However, the construction of such clean da…