25 citations · 35 across the 12 of their papers we have counts for
Showing cs.CVShow all
2 papers · 1 filter
cs.CV2024
MAVIN: Multi-Action Video Generation with Diffusion Models via Transition Video Infilling
Bowen Zhang, Xiaofei Xie, Haotian Lu +3
Diffusion-based video generation has achieved significant progress, yet generating multiple actions that occur sequentially remains a formidable task. Directly generating a video w…
cs.CV2024
MIP: CLIP-based Image Reconstruction from PEFT Gradients
Peiheng Zhou, Ming Hu, Xiaofei Xie +3
Contrastive Language-Image Pre-training (CLIP) model, as an effective pre-trained multimodal neural network, has been widely used in distributed machine learning tasks, especially…