5 papers
Image Editing with Diffusion Models: A Survey
Jia Wang, Jie Hu, Xiaoqi Ma +3
With deeper exploration of diffusion model, developments in the field of image generation have triggered a boom in image creation. As the quality of base-model generated images con…
Denoising with a Joint-Embedding Predictive Architecture
Dengsheng Chen, Jie Hu, Xiaoming Wei +1
Joint-embedding predictive architectures (JEPAs) have shown substantial promise in self-supervised representation learning, yet their application in generative modeling remains und…
Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input
Jiajun Liu, Yibing Wang, Hanghang Ma +6
Rapid advancements have been made in extending Large Language Models (LLMs) to Large Multi-modal Models (LMMs). However, extending input modality of LLMs to video data remains a ch…
Fine-gained Zero-shot Video Sampling
Dengsheng Chen, Jie Hu, Xiaoming Wei +1
Incorporating a temporal dimension into pretrained image diffusion models for video generation is a prevalent approach. However, this method is computationally demanding and necess…
Deformable 3D Shape Diffusion Model
Dengsheng Chen, Jie Hu, Xiaoming Wei +1
The Gaussian diffusion model, initially designed for image generation, has recently been adapted for 3D point cloud generation. However, these adaptations have not fully considered…