papers

Publications (185)

cs.CV2025

Self-Evaluation Unlocks Any-Step Text-to-Image Generation

Xin Yu, Xiaojuan Qi, Zhengqi Li +6

We introduce the Self-Evaluating Model (Self-E), a novel, from-scratch training approach for text-to-image generation that supports any-step inference. Self-E learns from data simi…

cs.CV2020

Meticulous Object Segmentation

Chenglin Yang, Yilin Wang, Jianming Zhang +3

Compared with common image segmentation tasks targeted at low-resolution images, higher resolution detailed image segmentation receives much less attention. In this paper, we propo…

cs.CV2013

GPU Asynchronous Stochastic Gradient Descent to Speed Up Neural Network Training

Thomas Paine, Hailin Jin, Jianchao Yang +2

The ability to train large-scale neural networks has resulted in state-of-the-art performance in many areas of computer vision. These results have largely come from computational b…

cs.CV2024

UniHuman: A Unified Model for Editing Human Images in the Wild

Nannan Li, Qing Liu, Krishna Kumar Singh +4

Human image editing includes tasks like changing a person's pose, their clothing, or editing the image according to a text prompt. However, prior work often tackles these tasks sep…

cs.CV2022

Shape-guided Object Inpainting

Yu Zeng, Zhe Lin, Vishal M. Patel

Previous works on image inpainting mainly focus on inpainting background or partially missing objects, while the problem of inpainting an entire missing object remains unexplored.…

cs.CV2021

Multimodal Contrastive Training for Visual Representation Learning

Xin Yuan, Zhe Lin, Jason Kuen +5

We develop an approach to learning visual representations that embraces multimodal data, driven by a combination of intra- and inter-modal similarity preservation objectives. Unlik…

cs.CV2016

Video Scene Parsing with Predictive Feature Learning

Xiaojie Jin, Xin Li, Huaxin Xiao +9

In this work, we address the challenging video scene parsing problem by developing effective representation learning methods given limited parsing annotations. In particular, we co…

cs.CV2023

SmartMask: Context Aware High-Fidelity Mask Generation for Fine-grained Object Insertion and Layout Control

Jaskirat Singh, Jianming Zhang, Qing Liu +3

The field of generative image inpainting and object insertion has made significant progress with the recent advent of latent diffusion models. Utilizing a precise object mask can g…

cs.CV2022

CoGS: Controllable Generation and Search from Sketch and Style

Cusuh Ham, Gemma Canet Tarres, Tu Bui +3

We present CoGS, a novel method for the style-conditioned, sketch-driven synthesis of images. CoGS enables exploration of diverse appearance possibilities for a given sketched obje…

math.LO2014

Residuated Basic Logic II. Interpolation, Decidability and Embedding

Minghui Ma, Zhe Lin

We prove that the sequent calculus for residuated basic logic has strong finite model property, and that intuitionistic logic can be embedded into…

cs.CV2022

Semantic Layout Manipulation with High-Resolution Sparse Attention

Haitian Zheng, Zhe Lin, Jingwan Lu +4

We tackle the problem of semantic image layout manipulation, which aims to manipulate an input image by editing its semantic label map. A core problem of this task is how to transf…

cs.CV2026

LaViDa: A Large Diffusion Language Model for Multimodal Understanding

Shufan Li, Konstantinos Kallidromitis, Hritik Bansal +7

LaViDa introduces a diffusion-based vision-language model that combines a vision encoder with discrete diffusion to enable fast parallel decoding and controllable multimodal genera…

#vision-language models#diffusion models#multimodal understanding#parallel decoding
cs.CV2021

Multi-Scale Aligned Distillation for Low-Resolution Detection

Lu Qi, Jason Kuen, Jiuxiang Gu +5

In instance-level detection tasks (e.g., object detection), reducing input resolution is an easy option to improve runtime efficiency. However, this option traditionally hurts the…

cs.CV2020

Context-Aware Group Captioning via Self-Attention and Contrastive Features

Zhuowan Li, Quan Tran, Long Mai +2

While image captioning has progressed rapidly, existing works focus mainly on describing single images. In this paper, we introduce a new task, context-aware group captioning, whic…

cs.CV2023

Video-P2P: Video Editing with Cross-attention Control

Shaoteng Liu, Yuechen Zhang, Wenbo Li +2

This paper presents Video-P2P, a novel framework for real-world video editing with cross-attention control. While attention control has proven effective for image editing with pre-…

cs.CV2026

Lavida-O: Elastic Large Masked Diffusion Models for Unified Multimodal Understanding and Generation

Shufan Li, Jiuxiang Gu, Kangning Liu +4

Lavida-O is a unified masked diffusion model that combines a lightweight generation branch with a larger understanding branch to perform image understanding, object grounding, imag…

#multimodal diffusion#masked diffusion#object grounding#high‑resolution image synthesis
cs.CV2025

OmniVCus: Feedforward Subject-driven Video Customization with Multimodal Control Conditions

Yuanhao Cai, He Zhang, Xi Chen +11

Existing feedforward subject-driven video customization methods mainly study single-subject scenarios due to the difficulty of constructing multi-subject training data pairs. Anoth…

cs.CV2024

XQ-GAN: An Open-source Image Tokenization Framework for Autoregressive Generation

Xiang Li, Kai Qiu, Hao Chen +5

Image tokenizers play a critical role in shaping the performance of subsequent generative models. Since the introduction of VQ-GAN, discrete image tokenization has undergone remark…

cs.CV2026

UniSER: A Foundation Model for Unified Soft Effects Removal

Jingdong Zhang, Lingzhi Zhang, Qing Liu +12

Digital images are often degraded by soft effects such as lens flare, haze, shadows, and reflections, which reduce aesthetics even though the underlying pixels remain partially vis…

cs.CV2023

XFormer: Fast and Accurate Monocular 3D Body Capture

Lihui Qian, Xintong Han, Faqiang Wang +6

We present XFormer, a novel human mesh and motion capture method that achieves real-time performance on consumer CPUs given only monocular images as input. The proposed network arc…

cs.CV2021

Open-Edit: Open-Domain Image Manipulation with Open-Vocabulary Instructions

Xihui Liu, Zhe Lin, Jianming Zhang +4

We propose a novel algorithm, named Open-Edit, which is the first attempt on open-domain image manipulation with open-vocabulary instructions. It is a challenging task considering…

cs.CV2022

CA-SSL: Class-Agnostic Semi-Supervised Learning for Detection and Segmentation

Lu Qi, Jason Kuen, Zhe Lin +7

To improve instance-level detection/segmentation performance, existing self-supervised and semi-supervised methods extract either task-unrelated or task-specific training signals f…

cs.CV2017

High-Resolution Image Inpainting using Multi-Scale Neural Patch Synthesis

Chao Yang, Xin Lu, Zhe Lin +3

Recent advances in deep learning have shown exciting promise in filling large holes in natural images with semantically plausible and context aware details, impacting fundamental i…

cs.CV2026

LightMover: Generative Light Movement with Color and Intensity Controls

Gengze Zhou, Tianyu Wang, Soo Ye Kim +7

We present LightMover, a framework for controllable light manipulation in single images that leverages video diffusion priors to produce physically plausible illumination changes w…

cs.CV2026

How Long Can Unified Multimodal Models Generate Images Reliably? Taming Long-Horizon Interleaved Image Generation via Context Curation

Haoyu Chen, Qing Liu, Yuqian Zhou +7

Unified multimodal models hold the promise of generating extensive, interleaved narratives, weaving text and imagery into coherent long-form stories. However, current systems suffe…

cs.CV2015

Automatic Content-Aware Color and Tone Stylization

Joon-Young Lee, Kalyan Sunkavalli, Zhe Lin +2

We introduce a new technique that automatically generates diverse, visually compelling stylizations for a photograph in an unsupervised manner. We achieve this by learning style ra…

cs.CV2024

Mixture of Efficient Diffusion Experts Through Automatic Interval and Sub-Network Selection

Alireza Ganjdanesh, Yan Kang, Yuchen Liu +3

Diffusion probabilistic models can generate high-quality samples. Yet, their sampling process requires numerous denoising steps, making it slow and computationally intensive. We pr…

cs.CV2020

Deep Image Compositing

He Zhang, Jianming Zhang, Federico Perazzi +2

Image compositing is a task of combining regions from different images to compose a new image. A common use case is background replacement of portrait images. To obtain high qualit…

cs.CV2026

Frame Guidance: Training-Free Guidance for Frame-Level Control in Video Diffusion Models

Sangwon Jang, Taekyung Ki, Jaehyeong Jo +4

Advancements in diffusion models have significantly improved video quality, directing attention to fine-grained controllability. However, many existing methods depend on fine-tunin…

cs.CV2026

Rolling Sink: Bridging Limited-Horizon Training and Open-Ended Testing in Autoregressive Video Diffusion

Haodong Li, Shaoteng Liu, Zhe Lin +1

Recently, autoregressive (AR) video diffusion models have achieved remarkable performance. However, due to their limited training durations, a train-test gap emerges when testing a…

cs.CV2018

Progressive Attention Networks for Visual Attribute Prediction

Paul Hongsuck Seo, Zhe Lin, Scott Cohen +2

We propose a novel attention model that can accurately attends to target objects of various scales and shapes in images. The model is trained to gradually suppress irrelevant regio…

cs.CV2017

Deep Image Harmonization

Yi-Hsuan Tsai, Xiaohui Shen, Zhe Lin +3

Compositing is one of the most common operations in photo editing. To generate realistic composites, the appearances of foreground and background need to be adjusted to make them c…

cs.CV2023

SimpSON: Simplifying Photo Cleanup with Single-Click Distracting Object Segmentation Network

Chuong Huynh, Yuqian Zhou, Zhe Lin +4

In photo editing, it is common practice to remove visual distractions to improve the overall image quality and highlight the primary subject. However, manually selecting and removi…

cs.LG2022

PowerGear: Early-Stage Power Estimation in FPGA HLS via Heterogeneous Edge-Centric GNNs

Zhe Lin, Zike Yuan, Jieru Zhao +3

Power estimation is the basis of many hardware optimization strategies. However, it is still challenging to offer accurate power estimation at an early stage such as high-level syn…

cs.CV2021

Content-Aware GAN Compression

Yuchen Liu, Zhixin Shu, Yijun Li +3

Generative adversarial networks (GANs), e.g., StyleGAN2, play a vital role in various image generation and synthesis tasks, yet their notoriously high computational cost hinders th…

cs.CV2026

Rethinking Global Text Conditioning in Diffusion Transformers

Nikita Starodubcev, Daniil Pakhomov, Zongze Wu +6

Diffusion transformers typically incorporate textual information via attention layers and a modulation mechanism using a pooled text embedding. Nevertheless, recent approaches disc…

cs.CV2022

HyperNST: Hyper-Networks for Neural Style Transfer

Dan Ruta, Andrew Gilbert, Saeid Motiian +3

We present HyperNST; a neural style transfer (NST) technique for the artistic stylization of images, based on Hyper-networks and the StyleGAN2 architecture. Our contribution is a n…

cs.CV2022

SmartBrush: Text and Shape Guided Object Inpainting with Diffusion Model

Shaoan Xie, Zhifei Zhang, Zhe Lin +2

Generic image inpainting aims to complete a corrupted image by borrowing surrounding information, which barely generates novel content. By contrast, multi-modal inpainting provides…

cs.IR2024

Deep Bag-of-Words Model: An Efficient and Interpretable Relevance Architecture for Chinese E-Commerce

Zhe Lin, Jiwei Tan, Dan Ou +3

Text relevance or text matching of query and product is an essential technique for the e-commerce search system to ensure that the displayed products can match the intent of the qu…

cs.CV2024

Generative Video Propagation

Shaoteng Liu, Tianyu Wang, Jui-Hsien Wang +8

Large-scale video generation models have the inherent ability to realistically model natural scenes. In this paper, we demonstrate that through a careful design of a generative vid…

cs.CL2020

Scene Graph Modification Based on Natural Language Commands

Xuanli He, Quan Hung Tran, Gholamreza Haffari +5

Structured representations like graphs and parse trees play a crucial role in many Natural Language Processing systems. In recent years, the advancements in multi-turn user interfa…

cs.CV2020

Real-time Semantic Segmentation with Fast Attention

Ping Hu, Federico Perazzi, Fabian Caba Heilbron +4

In deep CNN based models for semantic segmentation, high accuracy relies on rich spatial context (large receptive fields) and fine spatial details (high resolution), both of which…

cs.CV2016

Photo Aesthetics Ranking Network with Attributes and Content Adaptation

Shu Kong, Xiaohui Shen, Zhe Lin +2

Real-world applications could benefit from the ability to automatically generate a fine-grained ranking of photo aesthetics. However, previous methods for image aesthetics analysis…

cs.CV2021

Learning to Predict Visual Attributes in the Wild

Khoi Pham, Kushal Kafle, Zhe Lin +4

Visual attributes constitute a large portion of information contained in a scene. Objects can be described using a wide variety of attributes which portray their visual appearance…

cs.LO2014

The Computational Compexity of Decision Problem in Additive Extensions of Nonassociative Lambek Calculus

Zhe Lin, Minghui Ma

We analyze the complexity of decision problems for Boolean Nonassociative Lambek Calculus admitting empty antecedent of sequents (), and the consequence relation o…

cs.CV2026

SNCE: Geometry-Aware Supervision for Scalable Discrete Image Generation

Shufan Li, Jiuxiang Gu, Kangning Liu +3

Recent advancements in discrete image generation showed that scaling the VQ codebook size significantly improves reconstruction fidelity. However, training generative models with a…

cs.CV2022

Inpainting at Modern Camera Resolution by Guided PatchMatch with Auto-Curation

Lingzhi Zhang, Connelly Barnes, Kevin Wampler +4

Recently, deep models have established SOTA performance for low-resolution image inpainting, but they lack fidelity at resolutions associated with modern cameras such as 4K or more…

cs.AI2018

Active Object Perceiver: Recognition-guided Policy Learning for Object Searching on Mobile Robots

Xin Ye, Zhe Lin, Haoxiang Li +2

We study the problem of learning a navigation policy for a robot to actively search for an object of interest in an indoor environment solely from its visual inputs. While scene-dr…

cs.CV2021

CR-Fill: Generative Image Inpainting with Auxiliary Contexutal Reconstruction

Yu Zeng, Zhe Lin, Huchuan Lu +1

Recent deep generative inpainting methods use attention layers to allow the generator to explicitly borrow feature patches from the known region to complete a missing region. Due t…

cs.LO2019

The Finite Model Property of Quasi-transitive Modal Logic

Zhe Lin, Minghui Ma

The finite model property of quasi-transitive modal logic is established. This modal logic is conservatively…

cs.CV2022

GALA: Toward Geometry-and-Lighting-Aware Object Search for Compositing

Sijie Zhu, Zhe Lin, Scott Cohen +3

Compositing-aware object search aims to find the most compatible objects for compositing given a background image and a query bounding box. Previous works focus on learning compati…

cs.CV2025

HBridge: H-Shape Bridging of Heterogeneous Experts for Unified Multimodal Understanding and Generation

Xiang Wang, Zhifei Zhang, He Zhang +11

Recent unified models integrate understanding experts (e.g., LLMs) with generative experts (e.g., diffusion models), achieving strong multimodal performance. However, recent advanc…

cs.CV2024

ControlVAR: Exploring Controllable Visual Autoregressive Modeling

Xiang Li, Kai Qiu, Hao Chen +4

Conditional visual generation has witnessed remarkable progress with the advent of diffusion models (DMs), especially in tasks like control-to-image generation. However, challenges…

cs.CV2024

Attention-Driven Training-Free Efficiency Enhancement of Diffusion Models

Hongjie Wang, Difan Liu, Yan Kang +4

Diffusion Models (DMs) have exhibited superior performance in generating high-quality and diverse images. However, this exceptional performance comes at the cost of expensive archi…

cs.CV2024

MetaShadow: Object-Centered Shadow Detection, Removal, and Synthesis

Tianyu Wang, Jianming Zhang, Haitian Zheng +7

Shadows are often under-considered or even ignored in image editing applications, limiting the realism of the edited results. In this paper, we introduce MetaShadow, a three-in-one…

cs.CV2021

Going Deeper Into Face Detection: A Survey

Shervin Minaee, Ping Luo, Zhe Lin +1

Face detection is a crucial first step in many facial recognition and face analysis systems. Early approaches for face detection were mainly based on classifiers built on top of ha…

cs.CL2019

Expressing Visual Relationships via Language

Hao Tan, Franck Dernoncourt, Zhe Lin +2

Describing images with text is a fundamental problem in vision-language research. Current studies in this domain mostly focus on single image captioning. However, in various real a…

cs.CV2022

SceneComposer: Any-Level Semantic Image Synthesis

Yu Zeng, Zhe Lin, Jianming Zhang +4

We propose a new framework for conditional image synthesis from semantic layouts of any precision levels, ranging from pure text to a 2D semantic canvas with precise shapes. More s…

physics.flu-dyn2021

Effects of Pore-scale on the Macroscopic Properties of Natural Convection in Porous Media

Stefan Gasow, Zhe Lin, Hao Chun Zhang +3

Natural convection in porous media is a fundamental process for the long-term storage of CO2 in deep saline aquifers. Typically, details of mass transfer in porous media are inferr…

cs.CV2024

Refine-by-Align: Reference-Guided Artifacts Refinement through Semantic Alignment

Yizhi Song, Liu He, Zhifei Zhang +8

Personalized image generation has emerged from the recent advancements in generative models. However, these generated personalized images often suffer from localized artifacts such…

cs.DS2018

Efficient and High-Quality Seeded Graph Matching: Employing High Order Structural Information

Haida Zhang, Zengfeng Huang, Xuemin Lin +3

Driven by many real applications, we study the problem of seeded graph matching. Given two graphs and , and a small set of pre-matched node…

cs.CV2024

IMPRINT: Generative Object Compositing by Learning Identity-Preserving Representation

Yizhi Song, Zhifei Zhang, Zhe Lin +7

Generative object compositing emerges as a promising new avenue for compositional image editing. However, the requirement of object identity preservation poses a significant challe…

cs.CV2025

Layer- and Timestep-Adaptive Differentiable Token Compression Ratios for Efficient Diffusion Transformers

Haoran You, Connelly Barnes, Yuqian Zhou +10

Diffusion Transformers (DiTs) have achieved state-of-the-art (SOTA) image generation quality but suffer from high latency and memory inefficiency, making them difficult to deploy o…

cs.CV2026

A Unified and Controllable Framework for Layered Image Generation with Visual Effects

Jinrui Yang, Qing Liu, Yijun Li +5

Recent image generation models produce impressive composites, but often fail to preserve the identity of user-provided content when editing specific elements: the surrounding scene…

cs.CV2019

Multitask Text-to-Visual Embedding with Titles and Clickthrough Data

Pranav Aggarwal, Zhe Lin, Baldo Faieta +1

Text-visual (or called semantic-visual) embedding is a central problem in vision-language research. It typically involves mapping of an image and a text description to a common fea…

cs.CV2025

Multitwine: Multi-Object Compositing with Text and Layout Control

Gemma Canet Tarrés, Zhe Lin, Zhifei Zhang +4

We introduce the first generative model capable of simultaneous multi-object compositing, guided by both text and layout. Our model allows for the addition of multiple objects with…

physics.comp-ph2024

Pseudo grid-based physics-informed convolutional-recurrent network solving the integrable nonlinear lattice equations

Zhe Lin, Yong Chen

Traditional discrete learning methods involve discretizing continuous equations using difference schemes, necessitating considerations of stability and convergence. Integrable nonl…

cs.CV2023

InstantBooth: Personalized Text-to-Image Generation without Test-Time Finetuning

Jing Shi, Wei Xiong, Zhe Lin +1

Recent advances in personalized image generation allow a pre-trained text-to-image model to learn a new concept from a set of images. However, existing personalization approaches u…

cs.CV2016

Top-down Neural Attention by Excitation Backprop

Jianming Zhang, Zhe Lin, Jonathan Brandt +2

We aim to model the top-down attention of a Convolutional Neural Network (CNN) classifier for generating task-specific attention maps. Inspired by a top-down human visual attention…

math.LO2024

On the Cut Elimination of Weak Intuitionistic Tense Logic

Yiheng Wang, Yu Peng, Zhe Lin

In this paper, we use a new method to prove cut-elimination of weak intuitionistic tense logic. This method focuses on splitting the contraction rule and cut rules. Further general…

cs.CL2021

Pushing Paraphrase Away from Original Sentence: A Multi-Round Paraphrase Generation Approach

Zhe Lin, Xiaojun Wan

In recent years, neural paraphrase generation based on Seq2Seq has achieved superior performance, however, the generated paraphrase still has the problem of lack of diversity. In t…

cs.CV2023

AIMS: All-Inclusive Multi-Level Segmentation

Lu Qi, Jason Kuen, Weidong Guo +5

Despite the progress of image segmentation for accurate visual entity segmentation, completing the diverse requirements of image editing applications for different-level region-of-…

cs.CL2021

Making Better Use of Bilingual Information for Cross-Lingual AMR Parsing

Yitao Cai, Zhe Lin, Xiaojun Wan

Abstract Meaning Representation (AMR) is a rooted, labeled, acyclic graph representing the semantics of natural language. As previous works show, although AMR is designed for Engli…

cs.CV2021

SketchEdit: Mask-Free Local Image Manipulation with Partial Sketches

Yu Zeng, Zhe Lin, Vishal M. Patel

Sketch-based image manipulation is an interactive image editing task to modify an image based on input sketches from users. Existing methods typically formulate this task as a cond…

cs.CV2023

Human MotionFormer: Transferring Human Motions with Vision Transformers

Hongyu Liu, Xintong Han, Chengbin Jin +8

Human motion transfer aims to transfer motions from a target dynamic person to a source static one for motion synthesis. An accurate matching between the source person and the targ…

cs.CV2022

CM-GAN: Image Inpainting with Cascaded Modulation GAN and Object-Aware Training

Haitian Zheng, Zhe Lin, Jingwan Lu +7

Recent image inpainting methods have made great progress but often struggle to generate plausible image structures when dealing with large holes in complex images. This is partiall…

cs.CV2018

Learning to Detect Multiple Photographic Defects

Ning Yu, Xiaohui Shen, Zhe Lin +2

In this paper, we introduce the problem of simultaneously detecting multiple photographic defects. We aim at detecting the existence, severity, and potential locations of common ph…

cs.CV2020

Temporally Distributed Networks for Fast Video Semantic Segmentation

Ping Hu, Fabian Caba Heilbron, Oliver Wang +3

We present TDNet, a temporally distributed network designed for fast and accurate video semantic segmentation. We observe that features extracted from a certain high-level layer of…

cs.CV2025

TurboFill: Adapting Few-step Text-to-image Model for Fast Image Inpainting

Liangbin Xie, Daniil Pakhomov, Zhonghao Wang +8

This paper introduces TurboFill, a fast image inpainting model that enhances a few-step text-to-image diffusion model with an inpainting adapter for high-quality and efficient inpa…

cs.LG2018

Stochastic Downsampling for Cost-Adjustable Inference and Improved Regularization in Convolutional Networks

Jason Kuen, Xiangfei Kong, Zhe Lin +4

It is desirable to train convolutional networks (CNNs) to run more efficiently during inference. In many cases however, the computational budget that the system has for inference c…

physics.optics2026

Burst Mode Ultrafast Laser Welding of Sapphire and Fe-36Ni Alloy with Non-optical Contact Condition

Yu Wang, Nan Li, Yuxuan Li +8

Ultrafast laser welding provides a promising approach for high precision integration of transparent and metallic materials. However, its practical application remains constrained b…

cs.CL2022

Visual Information Guided Zero-Shot Paraphrase Generation

Zhe Lin, Xiaojun Wan

Zero-shot paraphrase generation has drawn much attention as the large-scale high-quality paraphrase corpus is limited. Back-translation, also known as the pivot-based method, is ty…

cs.CV2024

Object-level Scene Deocclusion

Zhengzhe Liu, Qing Liu, Chirui Chang +6

Deoccluding the hidden portions of objects in a scene is a formidable task, particularly when addressing real-world scenes. In this paper, we present a new self-supervised PArallel…

cs.LG2020

Shape Adaptor: A Learnable Resizing Module

Shikun Liu, Zhe Lin, Yilin Wang +3

We present a novel resizing module for neural networks: shape adaptor, a drop-in enhancement built on top of traditional resizing layers, such as pooling, bilinear sampling, and st…

cs.CV2019

Towards High-Resolution Salient Object Detection

Yi Zeng, Pingping Zhang, Jianming Zhang +2

Deep neural network based methods have made a significant breakthrough in salient object detection. However, they are typically limited to input images with low resolutions ($400\t…

cs.CV2024

Generative Image Layer Decomposition with Visual Effects

Jinrui Yang, Qing Liu, Yijun Li +7

Recent advancements in large generative models, particularly diffusion-based methods, have significantly enhanced the capabilities of image editing. However, achieving precise cont…

cs.CV2021

Lite Vision Transformer with Enhanced Self-Attention

Chenglin Yang, Yilin Wang, Jianming Zhang +4

Despite the impressive representation capacity of vision transformer models, current light-weight vision transformer models still suffer from inconsistent and incorrect dense predi…

cs.CV2021

ALADIN: All Layer Adaptive Instance Normalization for Fine-grained Style Similarity

Dan Ruta, Saeid Motiian, Baldo Faieta +5

We present ALADIN (All Layer AdaIN); a novel architecture for searching images based on the similarity of their artistic style. Representation learning is critical to visual search…

cs.CV2026

Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models

Shufan Li, Jiuxiang Gu, Kangning Liu +4

Masked Discrete Diffusion Models (MDMs) have achieved strong performance across a wide range of multimodal tasks, including image understanding, generation, and editing. However, t…

cs.CV2024

UniReal: Universal Image Generation and Editing via Learning Real-world Dynamics

Xi Chen, Zhifei Zhang, He Zhang +10

We introduce UniReal, a unified framework designed to address various image generation and editing tasks. Existing solutions often vary by tasks, yet share fundamental principles:…

cs.CV2023

Harnessing the Spatial-Temporal Attention of Diffusion Models for High-Fidelity Text-to-Image Synthesis

Qiucheng Wu, Yujian Liu, Handong Zhao +4

Diffusion-based models have achieved state-of-the-art performance on text-to-image synthesis tasks. However, one critical limitation of these models is the low fidelity of generate…

cs.CV2021

SSH: A Self-Supervised Framework for Image Harmonization

Yifan Jiang, He Zhang, Jianming Zhang +7

Image harmonization aims to improve the quality of image compositing by matching the "appearance" (\eg, color tone, brightness and contrast) between foreground and background image…

cs.CV2022

3D-FM GAN: Towards 3D-Controllable Face Manipulation

Yuchen Liu, Zhixin Shu, Yijun Li +3

3D-controllable portrait synthesis has significantly advanced, thanks to breakthroughs in generative adversarial networks (GANs). However, it is still challenging to manipulate exi…

cs.CV2018

Concept Mask: Large-Scale Segmentation from Semantic Concepts

Yufei Wang, Zhe Lin, Xiaohui Shen +2

Existing works on semantic segmentation typically consider a small number of labels, ranging from tens to a few hundreds. With a large number of labels, training and evaluation of…

cs.SD2024

Efficient Autoregressive Audio Modeling via Next-Scale Prediction

Kai Qiu, Xiang Li, Hao Chen +5

Audio generation has achieved remarkable progress with the advance of sophisticated generative models, such as diffusion models (DMs) and autoregressive (AR) models. However, due t…

cs.CV2024

SegGen: Supercharging Segmentation Models with Text2Mask and Mask2Img Synthesis

Hanrong Ye, Jason Kuen, Qing Liu +3

We propose SegGen, a highly-effective training data generation method for image segmentation, which pushes the performance limits of state-of-the-art segmentation models to a signi…

cs.CV2023

Automatic High Resolution Wire Segmentation and Removal

Mang Tik Chiu, Xuaner Zhang, Zijun Wei +7

Wires and powerlines are common visual distractions that often undermine the aesthetics of photographs. The manual process of precisely segmenting and removing them is extremely te…

cs.CV2018

Reference-Conditioned Super-Resolution by Neural Texture Transfer

Zhifei Zhang, Zhaowen Wang, Zhe Lin +1

With the recent advancement in deep learning, we have witnessed a great progress in single image super-resolution. However, due to the significant information loss of the image dow…

cs.CV2017

Recurrent Multimodal Interaction for Referring Image Segmentation

Chenxi Liu, Zhe Lin, Xiaohui Shen +3

In this paper we are interested in the problem of image segmentation given natural language descriptions, i.e. referring expressions. Existing works tackle this problem by first mo…

cs.CV2018

Contextual-based Image Inpainting: Infer, Match, and Translate

Yuhang Song, Chao Yang, Zhe Lin +4

We study the task of image inpainting, which is to fill in the missing region of an incomplete image with plausible contents. To this end, we propose a learning-based approach to g…