activity
20242026
collaborators

6 papers

cs.CV2026

TextWand: A Unified Framework for Scene Text Editing

Shuyu Wang, Zhile Guan, Hongxiu Chen +5

We propose TextWand, a general-purpose framework that unifies scene text removal, generation, and replacement into a single model. By decomposing complex editing tasks into the ato…

eess.IV2026

RealOSR: Latent Guidance Boosts Diffusion-based Real-world Omnidirectional Image Super-Resolutions

Xuhan Sheng, Runyi Li, Bin Chen +3

Omnidirectional image super-resolution (ODISR) aims to upscale low-resolution (LR) omnidirectional images (ODIs) to high-resolution (HR), catering to the growing demand for detaile…

cs.CV2026

TalkPhoto: A Versatile Training-Free Conversational Assistant for Intelligent Image Editing

Yujie Hu, Zecheng Tang, Xu Jiang +2

Thanks to the powerful language comprehension capabilities of Large Language Models (LLMs), existing instruction-based image editing methods have introduced Multimodal Large Langua…

cs.CV2025

Label-guided Facial Retouching Reversion

Guanhua Zhao, Yu Gu, Xuhan Sheng +2

With the popularity of social media platforms and retouching tools, more people are beautifying their facial photos, posing challenges for fields requiring photo authenticity. To a…

cs.CV2025

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection

Shuyu Wang, Weiqi Li, Qian Wang +2

Recent advances in AI-generated content (AIGC) have significantly accelerated image editing techniques, driving increasing demand for diverse and fine-grained edits. Despite these…

cs.CV2024

OmniDrag: Enabling Motion Control for Omnidirectional Image-to-Video Generation

Weiqi Li, Shijie Zhao, Chong Mou +6

As virtual reality gains popularity, the demand for controllable creation of immersive and dynamic omnidirectional videos (ODVs) is increasing. While previous text-to-ODV generatio…