activity
20232025
most citedCompositional Prompt Tuning with Motion Cues for Open-vocabulary Video Relation Detection

9 citations · 9 across the 5 of their papers we have counts for

collaborators

5 papers

cs.CV2025

Generalized Visual Relation Detection with Diffusion Models

Kaifeng Gao, Siqi Chen, Hanwang Zhang +3

Visual relation detection (VRD) aims to identify relationships (or interactions) between object pairs in an image. Although recent VRD models have achieved impressive performance,…

cs.CV2025

Towards Better Alignment: Training Diffusion Models with Reinforcement Learning Against Sparse Rewards

Zijing Hu, Fengda Zhang, Long Chen +6

Diffusion models have achieved remarkable success in text-to-image generation. However, their practical applications are hindered by the misalignment between generated images and c…

cs.CV2024

ViD-GPT: Introducing GPT-style Autoregressive Generation in Video Diffusion Models

Kaifeng Gao, Jiaxin Shi, Hanwang Zhang +2

With the advance of diffusion models, today's video generation has achieved impressive quality. But generating temporal consistent long videos is still challenging. A majority of v…

cs.CV2023

Triple Correlations-Guided Label Supplementation for Unbiased Video Scene Graph Generation

Wenqing Wang, Kaifeng Gao, Yawei Luo +5

Video-based scene graph generation (VidSGG) is an approach that aims to represent video content in a dynamic graph by identifying visual entities and their relationships. Due to th…

cs.CV20239 cited

Compositional Prompt Tuning with Motion Cues for Open-vocabulary Video Relation Detection

Kaifeng Gao, Long Chen, Hanwang Zhang +2

Prompt tuning with large-scale pretrained vision-language models empowers open-vocabulary predictions trained on limited base categories, e.g., object classification and detection.…