34 citations · 64 across the 14 of their papers we have counts for
14 papers
Continual Vision-Language Representation Learning with Off-Diagonal Information
Zixuan Ni, Longhui Wei, Siliang Tang +2
Large-scale multi-modal contrastive learning frameworks like CLIP typically require a large amount of image-text samples for training. However, these samples are always collected c…
ControlVideo: Training-free Controllable Text-to-Video Generation
Yabo Zhang, Yuxiang Wei, Dongsheng Jiang +3
Text-driven diffusion models have unlocked unprecedented abilities in image generation, whereas their video counterpart still lags behind due to the excessive training cost of temp…
ShiftDDPMs: Exploring Conditional Diffusion Models by Shifting Diffusion Trajectories
Zijian Zhang, Zhou Zhao, Jun Yu +1
Diffusion models have recently exhibited remarkable abilities to synthesize striking image samples since the introduction of denoising diffusion probabilistic models (DDPMs). Their…
Multi-modal Prompting for Low-Shot Temporal Action Localization
Chen Ju, Zeqian Li, Peisen Zhao +5
In this paper, we consider the problem of temporal action localization under low-shot (zero-shot & few-shot) scenario, with the goal of detecting and classifying the action instanc…
Lformer: Text-to-Image Generation with L-shape Block Parallel Decoding
Jiacheng Li, Longhui Wei, ZongYuan Zhan +4
Generative transformers have shown their superiority in synthesizing high-fidelity and high-resolution images, such as good diversity and training stability. However, they suffer f…
Dilated Context Integrated Network with Cross-Modal Consensus for Temporal Emotion Localization in Videos
Juncheng Li, Junlin Xie, Linchao Zhu +8
Understanding human emotions is a crucial ability for intelligent robots to provide better human-robot interactions. The existing works are limited to trimmed video-level emotion c…