7 papers
HumanNOVA: Photorealistic, Universal and Rapid 3D Human Avatar Modeling from a Single Image
Hezhen Hu, Wangbo Zhao, Lanqing Guo +6
In this paper, we present HumanNOVA, a photorealistic, universal, and rapid model for generating 3D human avatars from a single RGB image. Achieving both photorealism and generaliz…
Scale Where It Matters: Training-Free Localized Scaling for Diffusion Models
Qin Ren, Yufei Wang, Lanqing Guo +3
Diffusion models have become the dominant paradigm in text-to-image generation, and test-time scaling (TTS) improves sample quality by allocating additional computation at inferenc…
Controlling Your Image via Simplified Vector Graphics
Lanqing Guo, Xi Liu, Yufei Wang +2
Recent advances in image generation have achieved remarkable visual quality, while a fundamental challenge remains: Can image generation be controlled at the element level, enablin…
D2-Mamba: Dual-Scale Fusion and Dual-Path Scanning with SSMs for Shadow Removal
Linhao Li, Boya Jin, Zizhe Li +4
Shadow removal aims to restore images that are partially degraded by shadows, where the degradation is spatially localized and non-uniform. Unlike general restoration tasks that as…
Mastering Regional 3DGS: Locating, Initializing, and Editing with Diverse 2D Priors
Lanqing Guo, Yufei Wang, Hezhen Hu +4
Many 3D scene editing tasks focus on modifying local regions rather than the entire scene, except for some global applications like style transfer, and in the context of 3D Gaussia…
Demystifying the Visual Quality Paradox in Multimodal Large Language Models
Shuo Xing, Lanqing Guo, Hongyuan Hua +5
Recent Multimodal Large Language Models (MLLMs) excel on benchmark vision-language tasks, yet little is known about how input visual quality shapes their responses. Does higher per…