Publications (35)
Unbalanced Feature Transport for Exemplar-based Image Translation
Fangneng Zhan, Yingchen Yu, Kaiwen Cui +7
Despite the great success of GANs in images translation with different conditioned inputs such as semantic segmentation and edge maps, generating high-fidelity realistic images wit…
TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding
Zuhao Yang, Yingchen Yu, Yunqing Zhao +2
Video Temporal Grounding (VTG) aims to precisely identify video event segments in response to textual queries. The outputs of VTG tasks manifest as sequences of events, each define…
Diverse Image Inpainting with Bidirectional and Autoregressive Transformers
Yingchen Yu, Fangneng Zhan, Rongliang Wu +6
Image inpainting is an underdetermined inverse problem, which naturally allows diverse contents to fill up the missing or corrupted regions realistically. Prevalent approaches usin…
Monocular Normal Estimation via Shading Sequence Estimation
Zongrui Li, Xinhua Ma, Minghui Hu +6
Monocular normal estimation aims to estimate the normal map from a single RGB image of an object under arbitrary lights. Existing methods rely on deep models to directly predict no…
GMLight: Lighting Estimation via Geometric Distribution Approximation
Fangneng Zhan, Yingchen Yu, Changgong Zhang +6
Inferring the scene illumination from a single image is an essential yet challenging task in computer vision and computer graphics. Existing works estimate lighting by regressing r…
Let ViT Speak: Generative Language-Image Pre-training
Yan Fang, Mengcheng Lan, Zilong Huang +7
In this paper, we present \textbf{Gen}erative \textbf{L}anguage-\textbf{I}mage \textbf{P}re-training (GenLIP), a minimalist generative pretraining framework for Vision Transformers…
WaveNeRF: Wavelet-based Generalizable Neural Radiance Fields
Muyu Xu, Fangneng Zhan, Jiahui Zhang +5
Neural Radiance Field (NeRF) has shown impressive performance in novel view synthesis via implicit scene representation. However, it usually suffers from poor scalability as requir…
Latent Multi-Relation Reasoning for GAN-Prior based Image Super-Resolution
Jiahui Zhang, Fangneng Zhan, Yingchen Yu +3
Recently, single image super-resolution (SR) under large scaling factors has witnessed impressive progress by introducing pre-trained generative adversarial networks (GANs) as prio…
Accelerating DETR Convergence via Semantic-Aligned Matching
Gongjie Zhang, Zhipeng Luo, Yingchen Yu +2
The recently developed DEtection TRansformer (DETR) establishes a new object detection paradigm by eliminating a series of hand-crafted components. However, DETR suffers from extre…
WaveFill: A Wavelet-based Generation Network for Image Inpainting
Yingchen Yu, Fangneng Zhan, Shijian Lu +4
Image inpainting aims to complete the missing or corrupted regions of images with realistic contents. The prevalent approaches adopt a hybrid objective of reconstruction and percep…
Auto-regressive Image Synthesis with Integrated Quantization
Fangneng Zhan, Yingchen Yu, Rongliang Wu +4
Deep generative models have achieved conspicuous progress in realistic image synthesis with multifarious conditional inputs, while generating diverse yet high-fidelity images remai…
Blind Image Super-Resolution via Contrastive Representation Learning
Jiahui Zhang, Shijian Lu, Fangneng Zhan +1
Image super-resolution (SR) research has witnessed impressive progress thanks to the advance of convolutional neural networks (CNNs) in recent years. However, most existing SR meth…
SteerVTE: Seamless Video Text Editing with Style and Glyph Control
Kai Zeng, Moran Li, Zhengwei Wang +6
Visual text editing aims to precisely modify text in images and videos while preserving stylistic consistency and visual realism. Despite significant advances in the image domain,…
Marginal Contrastive Correspondence for Guided Image Generation
Fangneng Zhan, Yingchen Yu, Rongliang Wu +3
Exemplar-based image translation establishes dense correspondences between a conditional input and an exemplar (from two different domains) for leveraging detailed exemplar styles…
Debiasing Text-to-Image Diffusion Models
Ruifei He, Chuhui Xue, Haoru Tan +4
Learning-based Text-to-Image (TTI) models like Stable Diffusion have revolutionized the way visual content is generated in various domains. However, recent research has shown that…
VMRF: View Matching Neural Radiance Fields
Jiahui Zhang, Fangneng Zhan, Rongliang Wu +5
Neural Radiance Fields (NeRF) have demonstrated very impressive performance in novel view synthesis via implicitly modelling 3D representations from multi-view 2D images. However,…
EMLight: Lighting Estimation via Spherical Distribution Approximation
Fangneng Zhan, Changgong Zhang, Yingchen Yu +4
Illumination estimation from a single image is critical in 3D rendering and it has been investigated extensively in the computer vision and computer graphic research community. On…
Multimodal Image Synthesis and Editing: The Generative AI Era
Fangneng Zhan, Yingchen Yu, Rongliang Wu +6
As information exists in various modalities in real world, effective interaction and fusion among multimodal information plays a key role for the creation and perception of multimo…
POCE: Pose-Controllable Expression Editing
Rongliang Wu, Yingchen Yu, Fangneng Zhan +3
Facial expression editing has attracted increasing attention with the advance of deep neural networks in recent years. However, most existing methods suffer from compromised editin…
Weakly Supervised 3D Open-vocabulary Segmentation
Kunhao Liu, Fangneng Zhan, Jiahui Zhang +6
Open-vocabulary segmentation of 3D scenes is a fundamental function of human perception and thus a crucial objective in computer vision research. However, this task is heavily impe…
Learning to Animate Images from A Few Videos to Portray Delicate Human Actions
Haoxin Li, Yingchen Yu, Qilong Wu +3
Despite recent progress, video generative models still struggle to animate static images into videos that portray delicate human actions, particularly when handling uncommon or nov…
KD-DLGAN: Data Limited Image Generation via Knowledge Distillation
Kaiwen Cui, Yingchen Yu, Fangneng Zhan +3
Generative Adversarial Networks (GANs) rely heavily on large-scale training data for training high-quality image generation models. With limited training data, the GAN discriminato…
Towards Counterfactual Image Manipulation via CLIP
Yingchen Yu, Fangneng Zhan, Rongliang Wu +6
Leveraging StyleGAN's expressivity and its disentangled latent codes, existing methods can achieve realistic editing of different visual attributes such as age and gender of facial…
StyleRF: Zero-shot 3D Style Transfer of Neural Radiance Fields
Kunhao Liu, Fangneng Zhan, Yiwen Chen +5
3D style transfer aims to render stylized novel views of a 3D scene with multi-view consistency. However, most existing work suffers from a three-way dilemma over accurate geometry…
CodeDance: A Dynamic Tool-integrated MLLM for Executable Visual Reasoning
Qi Song, Honglin Li, Yingchen Yu +6
Recent releases such as o3 highlight human-like "thinking with images" reasoning that combines tool use with stepwise verification, yet most open-source approaches still rely on te…
Modulated Contrast for Versatile Image Synthesis
Fangneng Zhan, Jiahui Zhang, Yingchen Yu +2
Perceiving the similarity between images has been a long-standing and fundamental problem underlying various visual generation tasks. Predominant approaches measure the inter-image…
Versatile Transition Generation with Image-to-Video Diffusion
Zuhao Yang, Jiahui Zhang, Yingchen Yu +2
Leveraging text, images, structure maps, or motion trajectories as conditional guidance, diffusion models have achieved great success in automated and high-quality video generation…
Pose-Free Neural Radiance Fields via Implicit Pose Regularization
Jiahui Zhang, Fangneng Zhan, Yingchen Yu +5
Pose-free neural radiance fields (NeRF) aim to train NeRF with unposed multi-view images and it has achieved very impressive success in recent years. Most existing works share the…
TextSculptor: Training and Benchmarking Scene Text Editing
Yiheng Lin, Siyu Jiao, Xiaohan Lan +12
Recent advances in Multimodal Large Language Models (MLLMs) and diffusion-based generative models have substantially improved prompt-driven image editing. However, scene text editi…
GUI-Rise: Structured Reasoning and History Summarization for GUI Navigation
Tao Liu, Chongyu Wang, Rongjie Li +3
While Multimodal Large Language Models (MLLMs) have advanced GUI navigation agents, current approaches face limitations in cross-domain generalization and effective history utiliza…
Bi-level Feature Alignment for Versatile Image Translation and Manipulation
Fangneng Zhan, Yingchen Yu, Rongliang Wu +5
Generative adversarial networks (GANs) have achieved great success in image translation and manipulation. However, high-fidelity image generation with faithful style control remain…
Text4Seg++: Advancing Image Segmentation via Generative Language Modeling
Mengcheng Lan, Chaofeng Chen, Jiaxing Xu +6
Multimodal Large Language Models (MLLMs) have shown exceptional capabilities in vision-language tasks. However, effectively integrating image segmentation into these models remains…
Preference-Based Self-Distillation: Beyond KL Matching via Reward Regularization
Xin Yu, Liuchen Liao, Yiwen Zhang +3
On-policy distillation is an efficient alternative to reinforcement learning, offering dense token-level training signals. However, its reliance on a stronger external teacher has…
Audio-Driven Talking Face Generation with Diverse yet Realistic Facial Animations
Rongliang Wu, Yingchen Yu, Fangneng Zhan +3
Audio-driven talking face generation, which aims to synthesize talking faces with realistic facial animations (including accurate lip movements, vivid facial expression details and…
ThinkGen: Generalized Thinking for Visual Generation
Siyu Jiao, Yiheng Lin, Yujie Zhong +9
Recent progress in Multimodal Large Language Models (MLLMs) demonstrates that Chain-of-Thought (CoT) reasoning enables systematic solutions to complex understanding tasks. However,…