papers

Publications (35)

cs.CV2021

Unbalanced Feature Transport for Exemplar-based Image Translation

Fangneng Zhan, Yingchen Yu, Kaiwen Cui +7

Despite the great success of GANs in images translation with different conditioned inputs such as semantic segmentation and edge maps, generating high-fidelity realistic images wit…

cs.CV2025

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding

Zuhao Yang, Yingchen Yu, Yunqing Zhao +2

Video Temporal Grounding (VTG) aims to precisely identify video event segments in response to textual queries. The outputs of VTG tasks manifest as sequences of events, each define…

cs.CV2021

Diverse Image Inpainting with Bidirectional and Autoregressive Transformers

Yingchen Yu, Fangneng Zhan, Rongliang Wu +6

Image inpainting is an underdetermined inverse problem, which naturally allows diverse contents to fill up the missing or corrupted regions realistically. Prevalent approaches usin…

cs.CV2026

Monocular Normal Estimation via Shading Sequence Estimation

Zongrui Li, Xinhua Ma, Minghui Hu +6

Monocular normal estimation aims to estimate the normal map from a single RGB image of an object under arbitrary lights. Existing methods rely on deep models to directly predict no…

cs.CV2022

GMLight: Lighting Estimation via Geometric Distribution Approximation

Fangneng Zhan, Yingchen Yu, Changgong Zhang +6

Inferring the scene illumination from a single image is an essential yet challenging task in computer vision and computer graphics. Existing works estimate lighting by regressing r…

cs.CV2026

Let ViT Speak: Generative Language-Image Pre-training

Yan Fang, Mengcheng Lan, Zilong Huang +7

In this paper, we present \textbf{Gen}erative \textbf{L}anguage-\textbf{I}mage \textbf{P}re-training (GenLIP), a minimalist generative pretraining framework for Vision Transformers…

cs.CV2023

WaveNeRF: Wavelet-based Generalizable Neural Radiance Fields

Muyu Xu, Fangneng Zhan, Jiahui Zhang +5

Neural Radiance Field (NeRF) has shown impressive performance in novel view synthesis via implicit scene representation. However, it usually suffers from poor scalability as requir…

cs.CV2022

Latent Multi-Relation Reasoning for GAN-Prior based Image Super-Resolution

Jiahui Zhang, Fangneng Zhan, Yingchen Yu +3

Recently, single image super-resolution (SR) under large scaling factors has witnessed impressive progress by introducing pre-trained generative adversarial networks (GANs) as prio…

cs.CV2022

Accelerating DETR Convergence via Semantic-Aligned Matching

Gongjie Zhang, Zhipeng Luo, Yingchen Yu +2

The recently developed DEtection TRansformer (DETR) establishes a new object detection paradigm by eliminating a series of hand-crafted components. However, DETR suffers from extre…

cs.CV2021

WaveFill: A Wavelet-based Generation Network for Image Inpainting

Yingchen Yu, Fangneng Zhan, Shijian Lu +4

Image inpainting aims to complete the missing or corrupted regions of images with realistic contents. The prevalent approaches adopt a hybrid objective of reconstruction and percep…

cs.CV2022

Auto-regressive Image Synthesis with Integrated Quantization

Fangneng Zhan, Yingchen Yu, Rongliang Wu +4

Deep generative models have achieved conspicuous progress in realistic image synthesis with multifarious conditional inputs, while generating diverse yet high-fidelity images remai…

cs.CV2021

Blind Image Super-Resolution via Contrastive Representation Learning

Jiahui Zhang, Shijian Lu, Fangneng Zhan +1

Image super-resolution (SR) research has witnessed impressive progress thanks to the advance of convolutional neural networks (CNNs) in recent years. However, most existing SR meth…

cs.CV2026

SteerVTE: Seamless Video Text Editing with Style and Glyph Control

Kai Zeng, Moran Li, Zhengwei Wang +6

Visual text editing aims to precisely modify text in images and videos while preserving stylistic consistency and visual realism. Despite significant advances in the image domain,…

cs.CV2022

Marginal Contrastive Correspondence for Guided Image Generation

Fangneng Zhan, Yingchen Yu, Rongliang Wu +3

Exemplar-based image translation establishes dense correspondences between a conditional input and an exemplar (from two different domains) for leveraging detailed exemplar styles…

cs.CV2024

Debiasing Text-to-Image Diffusion Models

Ruifei He, Chuhui Xue, Haoru Tan +4

Learning-based Text-to-Image (TTI) models like Stable Diffusion have revolutionized the way visual content is generated in various domains. However, recent research has shown that…

cs.CV2023

VMRF: View Matching Neural Radiance Fields

Jiahui Zhang, Fangneng Zhan, Rongliang Wu +5

Neural Radiance Fields (NeRF) have demonstrated very impressive performance in novel view synthesis via implicitly modelling 3D representations from multi-view 2D images. However,…

cs.CV2020

EMLight: Lighting Estimation via Spherical Distribution Approximation

Fangneng Zhan, Changgong Zhang, Yingchen Yu +4

Illumination estimation from a single image is critical in 3D rendering and it has been investigated extensively in the computer vision and computer graphic research community. On…

cs.CV2023

Multimodal Image Synthesis and Editing: The Generative AI Era

Fangneng Zhan, Yingchen Yu, Rongliang Wu +6

As information exists in various modalities in real world, effective interaction and fusion among multimodal information plays a key role for the creation and perception of multimo…

cs.CV2023

POCE: Pose-Controllable Expression Editing

Rongliang Wu, Yingchen Yu, Fangneng Zhan +3

Facial expression editing has attracted increasing attention with the advance of deep neural networks in recent years. However, most existing methods suffer from compromised editin…

cs.CV2024

Weakly Supervised 3D Open-vocabulary Segmentation

Kunhao Liu, Fangneng Zhan, Jiahui Zhang +6

Open-vocabulary segmentation of 3D scenes is a fundamental function of human perception and thus a crucial objective in computer vision research. However, this task is heavily impe…

cs.CV2025

Learning to Animate Images from A Few Videos to Portray Delicate Human Actions

Haoxin Li, Yingchen Yu, Qilong Wu +3

Despite recent progress, video generative models still struggle to animate static images into videos that portray delicate human actions, particularly when handling uncommon or nov…

cs.CV2023

KD-DLGAN: Data Limited Image Generation via Knowledge Distillation

Kaiwen Cui, Yingchen Yu, Fangneng Zhan +3

Generative Adversarial Networks (GANs) rely heavily on large-scale training data for training high-quality image generation models. With limited training data, the GAN discriminato…

cs.CV2022

Towards Counterfactual Image Manipulation via CLIP

Yingchen Yu, Fangneng Zhan, Rongliang Wu +6

Leveraging StyleGAN's expressivity and its disentangled latent codes, existing methods can achieve realistic editing of different visual attributes such as age and gender of facial…

cs.CV2023

StyleRF: Zero-shot 3D Style Transfer of Neural Radiance Fields

Kunhao Liu, Fangneng Zhan, Yiwen Chen +5

3D style transfer aims to render stylized novel views of a 3D scene with multi-view consistency. However, most existing work suffers from a three-way dilemma over accurate geometry…

cs.CV2026

CodeDance: A Dynamic Tool-integrated MLLM for Executable Visual Reasoning

Qi Song, Honglin Li, Yingchen Yu +6

Recent releases such as o3 highlight human-like "thinking with images" reasoning that combines tool use with stepwise verification, yet most open-source approaches still rely on te…

cs.CV2022

Modulated Contrast for Versatile Image Synthesis

Fangneng Zhan, Jiahui Zhang, Yingchen Yu +2

Perceiving the similarity between images has been a long-standing and fundamental problem underlying various visual generation tasks. Predominant approaches measure the inter-image…

cs.CV2025

Versatile Transition Generation with Image-to-Video Diffusion

Zuhao Yang, Jiahui Zhang, Yingchen Yu +2

Leveraging text, images, structure maps, or motion trajectories as conditional guidance, diffusion models have achieved great success in automated and high-quality video generation…

cs.CV2023

Pose-Free Neural Radiance Fields via Implicit Pose Regularization

Jiahui Zhang, Fangneng Zhan, Yingchen Yu +5

Pose-free neural radiance fields (NeRF) aim to train NeRF with unposed multi-view images and it has achieved very impressive success in recent years. Most existing works share the…

cs.CV2026

TextSculptor: Training and Benchmarking Scene Text Editing

Yiheng Lin, Siyu Jiao, Xiaohan Lan +12

Recent advances in Multimodal Large Language Models (MLLMs) and diffusion-based generative models have substantially improved prompt-driven image editing. However, scene text editi…

cs.AI2025

GUI-Rise: Structured Reasoning and History Summarization for GUI Navigation

Tao Liu, Chongyu Wang, Rongjie Li +3

While Multimodal Large Language Models (MLLMs) have advanced GUI navigation agents, current approaches face limitations in cross-domain generalization and effective history utiliza…

cs.CV2022

Bi-level Feature Alignment for Versatile Image Translation and Manipulation

Fangneng Zhan, Yingchen Yu, Rongliang Wu +5

Generative adversarial networks (GANs) have achieved great success in image translation and manipulation. However, high-fidelity image generation with faithful style control remain…

cs.CV2025

Text4Seg++: Advancing Image Segmentation via Generative Language Modeling

Mengcheng Lan, Chaofeng Chen, Jiaxing Xu +6

Multimodal Large Language Models (MLLMs) have shown exceptional capabilities in vision-language tasks. However, effectively integrating image segmentation into these models remains…

cs.LG2026

Preference-Based Self-Distillation: Beyond KL Matching via Reward Regularization

Xin Yu, Liuchen Liao, Yiwen Zhang +3

On-policy distillation is an efficient alternative to reinforcement learning, offering dense token-level training signals. However, its reliance on a stronger external teacher has…

cs.CV2023

Audio-Driven Talking Face Generation with Diverse yet Realistic Facial Animations

Rongliang Wu, Yingchen Yu, Fangneng Zhan +3

Audio-driven talking face generation, which aims to synthesize talking faces with realistic facial animations (including accurate lip movements, vivid facial expression details and…

cs.CV2025

ThinkGen: Generalized Thinking for Visual Generation

Siyu Jiao, Yiheng Lin, Yujie Zhong +9

Recent progress in Multimodal Large Language Models (MLLMs) demonstrates that Chain-of-Thought (CoT) reasoning enables systematic solutions to complex understanding tasks. However,…