6 papers
TextWand: A Unified Framework for Scene Text Editing
Shuyu Wang, Zhile Guan, Hongxiu Chen +5
We propose TextWand, a general-purpose framework that unifies scene text removal, generation, and replacement into a single model. By decomposing complex editing tasks into the ato…
RealOSR: Latent Guidance Boosts Diffusion-based Real-world Omnidirectional Image Super-Resolutions
Xuhan Sheng, Runyi Li, Bin Chen +3
Omnidirectional image super-resolution (ODISR) aims to upscale low-resolution (LR) omnidirectional images (ODIs) to high-resolution (HR), catering to the growing demand for detaile…
TalkPhoto: A Versatile Training-Free Conversational Assistant for Intelligent Image Editing
Yujie Hu, Zecheng Tang, Xu Jiang +2
Thanks to the powerful language comprehension capabilities of Large Language Models (LLMs), existing instruction-based image editing methods have introduced Multimodal Large Langua…
Label-guided Facial Retouching Reversion
Guanhua Zhao, Yu Gu, Xuhan Sheng +2
With the popularity of social media platforms and retouching tools, more people are beautifying their facial photos, posing challenges for fields requiring photo authenticity. To a…
MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection
Shuyu Wang, Weiqi Li, Qian Wang +2
Recent advances in AI-generated content (AIGC) have significantly accelerated image editing techniques, driving increasing demand for diverse and fine-grained edits. Despite these…
OmniDrag: Enabling Motion Control for Omnidirectional Image-to-Video Generation
Weiqi Li, Shijie Zhao, Chong Mou +6
As virtual reality gains popularity, the demand for controllable creation of immersive and dynamic omnidirectional videos (ODVs) is increasing. While previous text-to-ODV generatio…