6 papers
YOSE: You Only Select Essential Tokens for Efficient DiT-based Video Object Removal
Chenyang Wu, Lina Lei, Fan Li +6
Recent advances in Diffusion Transformer (DiT)-based video generation technologies have shown impressive results for video object removal. However, these methods still suffer from…
HP-Edit: A Human-Preference Post-Training Framework for Image Editing
Fan Li, Chonghuinan Wang, Lina Lei +9
Common image editing tasks typically adopt powerful generative diffusion models as the leading paradigm for real-world content editing. Meanwhile, although reinforcement learning (…
ColorFLUX: A Structure-Color Decoupling Framework for Old Photo Colorization
Bingchen Li, Zhixin Wang, Fan Li +5
Old photos preserve invaluable historical memories, making their restoration and colorization highly desirable. While existing restoration models can address some degradation issue…
PlanViz: Evaluating Planning-Oriented Image Generation and Editing for Computer-Use Tasks
Junxian Li, Kai Liu, Leyang Chen +7
Unified multimodal models (UMMs) have shown impressive capabilities in generating natural images and supporting multimodal reasoning. However, their potential in supporting compute…
PocketSR: The Super-Resolution Expert in Your Pocket Mobiles
Haoze Sun, Linfeng Jiang, Fan Li +9
Real-world image super-resolution (RealSR) aims to enhance the visual quality of in-the-wild images, such as those captured by mobile phones. While existing methods leveraging larg…
UI2V-Bench: An Understanding-based Image-to-video Generation Benchmark
Ailing Zhang, Lina Lei, Dehong Kong +7
Generative diffusion models are developing rapidly and attracting increasing attention due to their wide range of applications. Image-to-Video (I2V) generation has become a major f…