papers

Publications (6)

cs.CV2022

TokenFlow: Rethinking Fine-grained Cross-modal Alignment in Vision-Language Retrieval

Xiaohan Zou, Changqiao Wu, Lele Cheng +1

Most existing methods in vision-language retrieval match two modalities by either comparing their global feature vectors which misses sufficient information and lacks interpretabil…

cs.CV2023

Decouple Content and Motion for Conditional Image-to-Video Generation

Cuifeng Shen, Yulu Gan, Chen Chen +4

The goal of conditional image-to-video (cI2V) generation is to create a believable new video by beginning with the condition, i.e., one image and text.The previous cI2V generation…

cs.CV2025

Paragraph-to-Image Generation with Information-Enriched Diffusion Model

Weijia Wu, Zhuang Li, Yefei He +6

Text-to-image (T2I) models have recently experienced rapid development, achieving astonishing performance in terms of fidelity and textual alignment capabilities. However, given a…

cs.CV2023

LaT: Latent Translation with Cycle-Consistency for Video-Text Retrieval

Jinbin Bai, Chunhui Liu, Feiyue Ni +4

Video-text retrieval is a class of cross-modal representation learning problems, where the goal is to select the video which corresponds to the text query between a given text quer…

cs.CV2020

Weakly Supervised Learning with Side Information for Noisy Labeled Images

Lele Cheng, Xiangzeng Zhou, Liming Zhao +5

In many real-world datasets, like WebVision, the performance of DNN based classifier is often limited by the noisy labeled data. To tackle this problem, some image related side inf…

cs.CV2019

Learning from Large-scale Noisy Web Data with Ubiquitous Reweighting for Image Classification

Jia Li, Yafei Song, Jianfeng Zhu +5

Many advances of deep learning techniques originate from the efforts of addressing the image classification task on large-scale datasets. However, the construction of such clean da…