7 papers
SAM-Flow: Source-Anchored Masked Flow for Training-Free Image Editing
Haowang Cui, Rui Chen, Tao Luo +3
Training-free image editing has recently attracted increasing attention due to its ability to modify real images using powerful pre-trained diffusion and flow-matching models witho…
UniView: Enhancing Novel View Synthesis From A Single Image By Unifying Reference Features
Haowang Cui, Rui Chen, Jiaze Wang +2
The task of synthesizing novel views from a single image is highly ill-posed due to multiple explanations for unobserved areas. Most current methods tend to generate unseen regions…
MOSAIC: Generating Consistent, Privacy-Preserving Scenes from Multiple Depth Views in Multi-Room Environments
Zhixuan Liu, Haokun Zhu, Rui Chen +4
We introduce a diffusion-based approach for generating privacy-preserving digital twins of multi-room indoor environments from depth images only. Central to our approach is a novel…
VoxHammer: Training-Free Precise and Coherent 3D Editing in Native 3D Space
Lin Li, Zehuan Huang, Haoran Feng +4
3D local editing of specified regions is crucial for game industry and robot interaction. Recent methods typically edit rendered multi-view images and then reconstruct 3D models, b…
Disentanglement in T-space for Faster and Distributed Training of Diffusion Models with Fewer Latent-states
Samarth Gupta, Raghudeep Gadde, Rui Chen +1
We challenge a fundamental assumption of diffusion models, namely, that a large number of latent-states or time-steps is required for training so that the reverse generative proces…
LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs
Jiaze Wang, Rui Chen, Haowang Cui
Recent spatial control methods for text-to-image (T2I) diffusion models have shown compelling results. However, these methods still fail to precisely follow the control conditions…