6 papers
D2-Mamba: Dual-Scale Fusion and Dual-Path Scanning with SSMs for Shadow Removal
Linhao Li, Boya Jin, Zizhe Li +4
Shadow removal aims to restore images that are partially degraded by shadows, where the degradation is spatially localized and non-uniform. Unlike general restoration tasks that as…
Mastering Regional 3DGS: Locating, Initializing, and Editing with Diverse 2D Priors
Lanqing Guo, Yufei Wang, Hezhen Hu +4
Many 3D scene editing tasks focus on modifying local regions rather than the entire scene, except for some global applications like style transfer, and in the context of 3D Gaussia…
Demystifying the Visual Quality Paradox in Multimodal Large Language Models
Shuo Xing, Lanqing Guo, Hongyuan Hua +5
Recent Multimodal Large Language Models (MLLMs) excel on benchmark vision-language tasks, yet little is known about how input visual quality shapes their responses. Does higher per…
Person Recognition at Altitude and Range: Fusion of Face, Body Shape and Gait
Feng Liu, Nicholas Chimitt, Lanqing Guo +16
We address the problem of whole-body person recognition in unconstrained environments. This problem arises in surveillance scenarios such as those in the IARPA Biometric Recognitio…
Oscillation Inversion: Understand the structure of Large Flow Model through the Lens of Inversion Method
Yan Zheng, Zhenxiao Liang, Xiaoyan Cong +4
We explore the oscillatory behavior observed in inversion methods applied to large-scale text-to-image diffusion models, with a focus on the "Flux" model. By employing a fixed-poin…
HiPrompt: Tuning-free Higher-Resolution Generation with Hierarchical MLLM Prompts
Xinyu Liu, Yingqing He, Lanqing Guo +10
The potential for higher-resolution image generation using pretrained diffusion models is immense, yet these models often struggle with issues of object repetition and structural a…