7 papers
TextWand: A Unified Framework for Scene Text Editing
Shuyu Wang, Zhile Guan, Hongxiu Chen +5
We propose TextWand, a general-purpose framework that unifies scene text removal, generation, and replacement into a single model. By decomposing complex editing tasks into the ato…
RealOSR: Latent Guidance Boosts Diffusion-based Real-world Omnidirectional Image Super-Resolutions
Xuhan Sheng, Runyi Li, Bin Chen +3
Omnidirectional image super-resolution (ODISR) aims to upscale low-resolution (LR) omnidirectional images (ODIs) to high-resolution (HR), catering to the growing demand for detaile…
TalkPhoto: A Versatile Training-Free Conversational Assistant for Intelligent Image Editing
Yujie Hu, Zecheng Tang, Xu Jiang +2
Thanks to the powerful language comprehension capabilities of Large Language Models (LLMs), existing instruction-based image editing methods have introduced Multimodal Large Langua…
TalkFashion: Intelligent Virtual Try-On Assistant Based on Multimodal Large Language Model
Yujie Hu, Xuanyu Zhang, Weiqi Li +1
Virtual try-on has made significant progress in recent years. This paper addresses how to achieve multifunctional virtual try-on guided solely by text instructions, including full…
D3C2-Net: Dual-Domain Deep Convolutional Coding Network for Compressive Sensing
Weiqi Li, Bin Chen, Shuai Liu +4
By mapping iterative optimization algorithms into neural networks (NNs), deep unfolding networks (DUNs) exhibit well-defined and interpretable structures and achieve remarkable suc…
Invertible Diffusion Models for Compressed Sensing
Bin Chen, Zhenyu Zhang, Weiqi Li +5
While deep neural networks (NN) significantly advance image compressed sensing (CS) by improving reconstruction quality, the necessity of training current CS NNs from scratch const…