activity
20142023
most citedControlVideo: Training-free Controllable Text-to-Video Generation

34 citations · 64 across the 14 of their papers we have counts for

collaborators

14 papers

cs.LG20232 cited

Continual Vision-Language Representation Learning with Off-Diagonal Information

Zixuan Ni, Longhui Wei, Siliang Tang +2

Large-scale multi-modal contrastive learning frameworks like CLIP typically require a large amount of image-text samples for training. However, these samples are always collected c…

cs.CV202334 cited

ControlVideo: Training-free Controllable Text-to-Video Generation

Yabo Zhang, Yuxiang Wei, Dongsheng Jiang +3

Text-driven diffusion models have unlocked unprecedented abilities in image generation, whereas their video counterpart still lags behind due to the excessive training cost of temp…

cs.CV20231 cited

ShiftDDPMs: Exploring Conditional Diffusion Models by Shifting Diffusion Trajectories

Zijian Zhang, Zhou Zhao, Jun Yu +1

Diffusion models have recently exhibited remarkable abilities to synthesize striking image samples since the introduction of denoising diffusion probabilistic models (DDPMs). Their…

cs.CV20235 cited

Multi-modal Prompting for Low-Shot Temporal Action Localization

Chen Ju, Zeqian Li, Peisen Zhao +5

In this paper, we consider the problem of temporal action localization under low-shot (zero-shot & few-shot) scenario, with the goal of detecting and classifying the action instanc…

cs.CV20233 cited

Lformer: Text-to-Image Generation with L-shape Block Parallel Decoding

Jiacheng Li, Longhui Wei, ZongYuan Zhan +4

Generative transformers have shown their superiority in synthesizing high-fidelity and high-resolution images, such as good diversity and training stability. However, they suffer f…

cs.CV2022

Dilated Context Integrated Network with Cross-Modal Consensus for Temporal Emotion Localization in Videos

Juncheng Li, Junlin Xie, Linchao Zhu +8

Understanding human emotions is a crucial ability for intelligent robots to provide better human-robot interactions. The existing works are limited to trimmed video-level emotion c…