4 papers
Resolving Representation Ambiguity in Feedforward Novel View Synthesis Transformer via Semantic-Spatial Decoupling
Yihang Wu, Yihang Sun, Shaofeng Zhang +4
Transformer-based models have advanced feedforward novel view synthesis (NVS). Current architectures such as GS-LRM and LVSM mix semantic information (e.g., RGB) and spatial inform…
VGGT-CD: Training-Free Robust Registration for 3D Change Detection
Wei Zhang, Songhua Li, Yihang Wu +2
3D change detection from multi-view images is essential for urban monitoring, disaster assessment, and autonomous driving. However, existing methods predominantly operate in the 2D…
MTADiffusion: Mask Text Alignment Diffusion Model for Object Inpainting
Jun Huang, Ting Liu, Yihang Wu +3
Advancements in generative models have enabled image inpainting models to generate content within specific regions of an image based on provided prompts and masks. However, existin…
Towards Better Text-to-Image Generation Alignment via Attention Modulation
Yihang Wu, Xiao Cao, Kaixin Li +4
In text-to-image generation tasks, the advancements of diffusion models have facilitated the fidelity of generated results. However, these models encounter challenges when processi…