3 papers
cs.CV2026
Direct 3D-Aware Object Insertion via Decomposed Visual Proxies
Jingbo Gong, Yikai Wang, Yushi Lan +6
Object insertion aims to seamlessly composite a reference object into a specified region of a background image. Recent diffusion-based methods achieve high visual quality but formu…
cs.CV2026
Thinking with Camera: A Unified Multimodal Model for Camera-Centric Understanding and Generation
Kang Liao, Size Wu, Zhonghua Wu +6
Camera-centric understanding and generation are two cornerstones of spatial intelligence, yet they are typically studied in isolation. We present Puffin, a unified camera-centric m…
cs.CV2026
Next Visual Granularity Generation
Yikai Wang, Zhouxia Wang, Zhonghua Wu +3
We propose a novel approach to image generation by decomposing an image into a structured sequence, where each element in the sequence shares the same spatial resolution but differ…