14 papers
Track the Noise, Move the World:3D-Grounded Motion-Consistent Noise for Controllable Video Generation
Long Vu, Tan Ngo, Animesh Karnewar +5
Modern image-and-text-to-video diffusion models can synthesize highly realistic videos by iteratively denoising an initial Gaussian noise tensor conditioned on reference image and…
SwiftTailor: Efficient 3D Garment Generation with Geometry Image Representation
Phuc Pham, Uy Dieu Tran, Binh-Son Hua +1
Realistic and efficient 3D garment generation remains a longstanding challenge in computer vision and digital fashion. Existing methods typically rely on large vision- language mod…
Catch Me If You Can Describe Me: Open-Vocabulary Camouflaged Instance Segmentation with Diffusion
Tuan-Anh Vu, Duc Thanh Nguyen, Qing Guo +4
Text-to-image diffusion techniques have shown exceptional capabilities in producing high-quality, dense visual predictions from open-vocabulary text. This indicates a strong correl…
FROMAT: Multiview Material Appearance Transfer via Few-Shot Self-Attention Adaptation
Hubert Kompanowski, Varun Jampani, Aaryaman Vasishta +1
Multiview diffusion models have rapidly emerged as a powerful tool for content creation with spatial consistency across viewpoints, offering rich visual realism without requiring e…
Power of Boundary and Reflection: Semantic Transparent Object Segmentation using Pyramid Vision Transformer with Transparent Cues
Tuan-Anh Vu, Hai Nguyen-Truong, Ziqiang Zheng +4
Glass is a prevalent material among solid objects in everyday life, yet segmentation methods struggle to distinguish it from opaque materials due to its transparency and reflection…
Text-to-3D Generation using Jensen-Shannon Score Distillation
Khoi Do, Binh-Son Hua
Score distillation sampling is an effective technique to generate 3D models from text prompts, utilizing pre-trained large-scale text-to-image diffusion models as guidance. However…