11 papers
From Synthesis to Removal: Physics-Grounded Reflection Simulation and Diffusion-Based Video Dereflection
Zepeng Wang, Jiagao Hu, Fuhao Li +3
Videos captured through glass often contain reflections that degrade visual quality and interfere with downstream vision tasks. Although single-image reflection removal has been ex…
TIGER: Taming Identity, Geometry, and Generative Priors for High-Quality Face Video Restoration
Yang Zhou, Wenxue Li, Peng Zhang +3
Face Video Restoration (FVR) aims to recover high-fidelity facial videos from degraded input while preserving identity and semantic consistency across frames. Existing methods ofte…
RS-Gen: A Multi-Stage Agentic Framework for Reasoning and Search-Augmented Image Generation
Feifei Bian, Zhimin Zheng, Wei Deng +2
Recent years have witnessed remarkable progress in image generation and editing, particularly regarding instruction following and visual fidelity. However, when handling ambiguous…
STAR: SpatioTemporal Adaptive Reward Allocation for Text-to-Image RL Post-Training
Jinjie Shen, Wei Deng, Xian Hu +2
Existing RL post-training methods for text-to-image generation usually convert the final-image reward into a single scalar advantage and apply it with the same strength to the enti…
PixelWizard: Towards Efficient High-Fidelity Video Generation at Ultra-Large Spatial Resolution
Wenxue Li, Jingjing Ren, Peng Zhang +4
High-resolution video generation faces a coupled bottleneck of optimization instability and prohibitive computational costs. The massive expansion of the token sequence not only bi…
PROVE: A Perceptual RemOVal cohErence Benchmark for Visual Media
Fuhao Li, Shaofeng You, Jiagao Hu +6
Evaluating object removal in images and videos remains challenging because the task is inherently one-to-many, yet existing metrics frequently disagree with human perception. Full-…