papers

Publications (98)

cs.CV2022

Region Proposal Rectification Towards Robust Instance Segmentation of Biological Images

Qilong Zhangli, Jingru Yi, Di Liu +10

Top-down instance segmentation framework has shown its superiority in object detection compared to the bottom-up framework. While it is efficient in addressing over-segmentation, t…

cs.CV2017

Automatic Liver Segmentation Using an Adversarial Image-to-Image Network

Dong Yang, Daguang Xu, S. Kevin Zhou +5

Automatic liver segmentation in 3D medical images is essential in many clinical applications, such as pathological diagnosis of hepatic diseases, surgical planning, and postoperati…

cs.CV2021

UTNet: A Hybrid Transformer Architecture for Medical Image Segmentation

Yunhe Gao, Mu Zhou, Dimitris Metaxas

Transformer architecture has emerged to be successful in a number of natural language processing tasks. However, its applications to medical vision remain largely unexplored. In th…

cs.LG2025

BLoB: Bayesian Low-Rank Adaptation by Backpropagation for Large Language Models

Yibin Wang, Haizhou Shi, Ligong Han +2

Large Language Models (LLMs) often suffer from overconfidence during inference, particularly when adapted to downstream domain-specific tasks with limited data. Previous work addre…

cs.CV2022

A Dynamic Data Driven Approach for Explainable Scene Understanding

Zachary A Daniels, Dimitris Metaxas

Scene-understanding is an important topic in the area of Computer Vision, and illustrates computational challenges with applications to a wide range of domains including remote sen…

cs.CV2021

Dual Projection Generative Adversarial Networks for Conditional Image Generation

Ligong Han, Martin Renqiang Min, Anastasis Stathopoulos +4

Conditional Generative Adversarial Networks (cGANs) extend the standard unconditional GAN framework to learning joint data-label distributions from samples, and have been establish…

cs.CV2021

CrossNorm and SelfNorm for Generalization under Distribution Shifts

Zhiqiang Tang, Yunhe Gao, Yi Zhu +3

Traditional normalization techniques (e.g., Batch Normalization and Instance Normalization) generally and simplistically assume that training and test data follow the same distribu…

cs.CV2019

Brain Segmentation from k-space with End-to-end Recurrent Attention Network

Qiaoying Huang, Xiao Chen, Dimitris Metaxas +1

The task of medical image segmentation commonly involves an image reconstruction step to convert acquired raw data to images before any analysis. However, noises, artifacts and los…

cs.CV2017

Reconstruction-Based Disentanglement for Pose-invariant Face Recognition

Xi Peng, Xiang Yu, Kihyuk Sohn +2

Deep neural networks (DNNs) trained on large-scale datasets have recently achieved impressive improvements in face recognition. But a persistent challenge remains to develop method…

cs.CV2026

Cross-Space Distillation: Teaching One-Step Students with Modern Diffusion Teachers

Anh Nguyen, Ngan Nguyen, Duc Vu +11

Modern one-step diffusion models achieve impressive quality through distribution-based timestep distillation. Yet, they rely on a critical assumption: Teacher and Student must inha…

cs.CV2024

New Capability to Look Up an ASL Sign from a Video Example

Carol Neidle, Augustine Opoku, Carey Ballard +4

Looking up an unknown sign in an ASL dictionary can be difficult. Most ASL dictionaries are organized based on English glosses, despite the fact that (1) there is no convention for…

eess.IV2022

Modality Bank: Learn multi-modality images across data centers without sharing medical data

Qi Chang, Hui Qu, Zhennan Yan +3

Multi-modality images have been widely used and provide comprehensive information for medical image analysis. However, acquiring all modalities among all institutes is costly and o…

eess.IV2020

Enhanced MRI Reconstruction Network using Neural Architecture Search

Qiaoying Huang, Dong Yang, Yikun Xian +4

The accurate reconstruction of under-sampled magnetic resonance imaging (MRI) data using modern deep learning technology, requires significant effort to design the necessary comple…

cs.CV2019

MRI Reconstruction via Cascaded Channel-wise Attention Network

Qiaoying Huang, Dong Yang, Pengxiang Wu +3

We consider an MRI reconstruction problem with input of k-space data at a very low undersampled rate. This can practically benefit patient due to reduced time of MRI scan, but it i…

cs.CV2022

Contrastive and Selective Hidden Embeddings for Medical Image Segmentation

Zhuowei Li, Zihao Liu, Zhiqiang Hu +5

Medical image segmentation has been widely recognized as a pivot procedure for clinical diagnosis, analysis, and treatment planning. However, the laborious and expensive annotation…

cs.CV2018

Learning to Forecast and Refine Residual Motion for Image-to-Video Generation

Long Zhao, Xi Peng, Yu Tian +2

We consider the problem of image-to-video translation, where an input image is translated into an output video containing motions of a single object. Recent methods for such proble…

cs.CV2022

Automatic Tooth Segmentation from 3D Dental Model using Deep Learning: A Quantitative Analysis of what can be learnt from a Single 3D Dental Model

Ananya Jana, Hrebesh Molly Subhash, Dimitris Metaxas

3D tooth segmentation is an important task for digital orthodontics. Several Deep Learning methods have been proposed for automatic tooth segmentation from 3D dental models or intr…

eess.IV2020

Measure Anatomical Thickness from Cardiac MRI with Deep Neural Networks

Qiaoying Huang, Eric Z. Chen, Hanchao Yu +4

Accurate estimation of shape thickness from medical images is crucial in clinical applications. For example, the thickness of myocardium is one of the key to cardiac disease diagno…

cs.CV2025

RadAlign: Advancing Radiology Report Generation with Vision-Language Concept Alignment

Difei Gu, Yunhe Gao, Yang Zhou +2

Automated chest radiographs interpretation requires both accurate disease classification and detailed radiology report generation, presenting a significant challenge in the clinica…

cs.CV2020

PC-U Net: Learning to Jointly Reconstruct and Segment the Cardiac Walls in 3D from CT Data

Meng Ye, Qiaoying Huang, Dong Yang +4

The 3D volumetric shape of the heart's left ventricle (LV) myocardium (MYO) wall provides important information for diagnosis of cardiac disease and invasive procedure navigation.…

cs.CV2025

DICE: Discrete Inversion Enabling Controllable Editing for Multinomial Diffusion and Masked Generative Models

Xiaoxiao He, Quan Dao, Ligong Han +14

Discrete diffusion models have achieved success in tasks like image generation and masked language modeling but face limitations in controlled content editing. We introduce DICE (D…

cs.CL2022

ASL Video Corpora & Sign Bank: Resources Available through the American Sign Language Linguistic Research Project (ASLLRP)

Carol Neidle, Augustine Opoku, Dimitris Metaxas

The American Sign Language Linguistic Research Project (ASLLRP) provides Internet access to high-quality ASL video data, generally including front and side views and a close-up of…

cs.CV2023

Constructive Assimilation: Boosting Contrastive Learning Performance through View Generation Strategies

Ligong Han, Seungwook Han, Shivchander Sudalairaj +8

Transformations based on domain expertise (expert transformations), such as random-resized-crop and color-jitter, have proven critical to the success of contrastive learning techni…

eess.IV2021

Semi-Supervised Segmentation of Radiation-Induced Pulmonary Fibrosis from Lung CT Scans with Multi-Scale Guided Dense Attention

Guotai Wang, Shuwei Zhai, Giovanni Lasio +7

Computed Tomography (CT) plays an important role in monitoring radiation-induced Pulmonary Fibrosis (PF), where accurate segmentation of the PF lesions is highly desired for diagno…

cs.CV2026

DeDPO: Debiased Direct Preference Optimization for Diffusion Models

Khiem Pham, Quang Nguyen, Tung Nguyen +4

Direct Preference Optimization (DPO) has emerged as a predominant alignment method for diffusion models, facilitating off-policy training without explicit reward modeling. However,…

cs.CV2019

Point Cloud Processing via Recurrent Set Encoding

Pengxiang Wu, Chao Chen, Jingru Yi +1

We present a new permutation-invariant network for 3D point cloud processing. Our network is composed of a recurrent set encoder and a convolutional feature aggregator. Given an un…

cs.CV2024

Score-Guided Diffusion for 3D Human Recovery

Anastasis Stathopoulos, Ligong Han, Dimitris Metaxas

We present Score-Guided Human Mesh Recovery (ScoreHMR), an approach for solving inverse problems for 3D human pose and shape reconstruction. These inverse problems involve fitting…

cs.CV2023

Learning Articulated Shape with Keypoint Pseudo-labels from Web Images

Anastasis Stathopoulos, Georgios Pavlakos, Ligong Han +1

This paper shows that it is possible to learn models for monocular 3D reconstruction of articulated objects (e.g., horses, cows, sheep), using as few as 50-150 images labeled with…

cs.CV2026

Data Augmentation for High-Fidelity Generation of CAR-T/NK Immunological Synapse Images

Xiang Zhang, Boxuan Zhang, Alireza Naghizadeh +4

Chimeric antigen receptor (CAR)-T and NK cell immunotherapies have transformed cancer treatment, and recent studies suggest that the quality of the CAR-T/NK cell immunological syna…

cs.CV2023

Improving Compositional Text-to-image Generation with Large Vision-Language Models

Song Wen, Guian Fang, Renrui Zhang +3

Recent advancements in text-to-image models, particularly diffusion models, have shown significant promise. However, compositional text-to-image models frequently encounter difficu…

cs.LG2022

A Manifold View of Adversarial Risk

Wenjia Zhang, Yikai Zhang, Xiaoling Hu +3

The adversarial risk of a machine learning model has been widely studied. Most previous works assume that the data lies in the whole ambient space. We propose to take a new angle a…

cs.CV2022

Show Me What and Tell Me How: Video Synthesis via Multimodal Conditioning

Ligong Han, Jian Ren, Hsin-Ying Lee +5

Most methods for conditional video synthesis use a single modality as the condition. This comes with major limitations. For example, it is problematic for a model conditioned on an…

eess.IV2020

Deep Learning based NAS Score and Fibrosis Stage Prediction from CT and Pathology Data

Ananya Jana, Hui Qu, Puru Rattan +3

Non-Alcoholic Fatty Liver Disease (NAFLD) is becoming increasingly prevalent in the world population. Without diagnosis at the right time, NAFLD can lead to non-alcoholic steatohep…

cs.CV2018

Jointly Optimize Data Augmentation and Network Training: Adversarial Data Augmentation in Human Pose Estimation

Xi Peng, Zhiqiang Tang, Fei Yang +2

Random data augmentation is a critical technique to avoid overfitting in training deep neural network models. However, data augmentation and network training are usually treated as…

cs.CV2026

MADCrowner: Margin Aware Dental Crown Design with Template Deformation and Refinement

Linda Wei, Chang Liu, Wenran Zhang +9

Dental crown restoration is one of the most common treatment modalities for tooth defect, where personalized dental crown design is critical. While computer-aided design (CAD) syst…

cs.CV2018

StackGAN++: Realistic Image Synthesis with Stacked Generative Adversarial Networks

Han Zhang, Tao Xu, Hongsheng Li +4

Although Generative Adversarial Networks (GANs) have shown remarkable success in various tasks, they still face challenges in generating high quality images. In this paper, we prop…

cs.CV2026

Self-Corrected Flow Distillation for Consistent One-Step and Few-Step Text-to-Image Generation

Quan Dao, Hao Phung, Trung Dao +2

Flow matching has emerged as a promising framework for training generative models, demonstrating impressive empirical performance while offering relative ease of training compared…

cs.LG2026

An Optimal Transport-driven Approach for Cultivating Latent Space in Online Incremental Learning

Quyen Tran, Hai Nguyen, Hoang Phan +6

In online incremental learning, data continuously arrives with substantial distributional shifts, creating a significant challenge because previous samples have limited replay valu…

cs.CV2020

Oriented Object Detection in Aerial Images with Box Boundary-Aware Vectors

Jingru Yi, Pengxiang Wu, Bo Liu +3

Oriented object detection in aerial images is a challenging task as the objects in aerial images are displayed in arbitrary directions and are usually densely packed. Current orien…

cs.CV2020

MotionNet: Joint Perception and Motion Prediction for Autonomous Driving Based on Bird's Eye View Maps

Pengxiang Wu, Siheng Chen, Dimitris Metaxas

The ability to reliably perceive the environmental states, particularly the existence of objects and their motion behavior, is crucial for autonomous driving. In this work, we prop…

cs.LG2021

Global and Local Interpretation of black-box Machine Learning models to determine prognostic factors from early COVID-19 data

Ananya Jana, Carlos D. Minacapelli, Vinod Rustgi +1

The COVID-19 corona virus has claimed 4.1 million lives, as of July 24, 2021. A variety of machine learning models have been applied to related data to predict important factors su…

cs.CV2024

VerSe: Integrating Multiple Queries as Prompts for Versatile Cardiac MRI Segmentation

Bangwei Guo, Meng Ye, Yunhe Gao +3

Despite the advances in learning-based image segmentation approach, the accurate segmentation of cardiac structures from magnetic resonance imaging (MRI) remains a critical challen…

cs.CV2025

Rate-My-LoRA: Efficient and Adaptive Federated Model Tuning for Cardiac MRI Segmentation

Xiaoxiao He, Haizhou Shi, Ligong Han +7

Cardiovascular disease (CVD) and cardiac dyssynchrony are major public health problems in the United States. Precise cardiac image segmentation is crucial for extracting quantitati…

cs.AI2017

Interactive Reinforcement Learning for Object Grounding via Self-Talking

Yan Zhu, Shaoting Zhang, Dimitris Metaxas

Humans are able to identify a referred visual object in a complex scene via a few rounds of natural language communications. Success communication requires both parties to engage a…

cs.LG2025

MedForge: Building Medical Foundation Models Like Open Source Software Development

Zheling Tan, Kexin Ding, Jin Gao +4

Foundational models (FMs) have made significant strides in the healthcare domain. Yet the data silo challenge and privacy concern remain in healthcare systems, hindering safe medic…

cs.CV2026

MPDiT: Multi-Patch Global-to-Local Transformer Architecture For Efficient Flow Matching and Diffusion Model

Quan Dao, Dimitris Metaxas

Transformer architectures, particularly Diffusion Transformers (DiTs), have become widely used in diffusion and flow-matching models due to their strong performance compared to con…

cs.CV2025

AutoEdit: Automatic Hyperparameter Tuning for Image Editing

Chau Pham, Quan Dao, Mahesh Bhosale +3

Recent advances in diffusion models have revolutionized text-guided image editing, yet existing editing methods face critical challenges in hyperparameter identification. To get th…

cs.IT2013

Adaptive low rank and sparse decomposition of video using compressive sensing

Fei Yang, Hong Jiang, Zuowei Shen +2

We address the problem of reconstructing and analyzing surveillance videos using compressive sensing. We develop a new method that performs video reconstruction by low rank and spa…

cs.CV2021

Learning View-Disentangled Human Pose Representation by Contrastive Cross-View Mutual Information Maximization

Long Zhao, Yuxiao Wang, Jiaping Zhao +7

We introduce a novel representation learning method to disentangle pose-dependent as well as view-dependent factors from 2D human poses. The method trains a network using cross-vie…

eess.IV2022

TransFusion: Multi-view Divergent Fusion for Medical Image Segmentation with Transformers

Di Liu, Yunhe Gao, Qilong Zhangli +8

Combining information from multi-view images is crucial to improve the performance and robustness of automated methods for disease diagnosis. However, due to the non-alignment char…

cs.CV2023

SVDiff: Compact Parameter Space for Diffusion Fine-Tuning

Ligong Han, Yinxiao Li, Han Zhang +3

Diffusion models have achieved remarkable success in text-to-image generation, enabling the creation of high-quality images from text prompts or other modalities. However, existing…

cs.CV2026

K-Prism: A Knowledge-Guided and Prompt Integrated Universal Medical Image Segmentation Model

Bangwei Guo, Yunhe Gao, Meng Ye +4

Medical image segmentation is fundamental to clinical decision-making, yet existing models remain fragmented. They are usually trained on single knowledge sources and specific to i…

cs.CV2021

AE-StyleGAN: Improved Training of Style-Based Auto-Encoders

Ligong Han, Sri Harsha Musunuri, Martin Renqiang Min +3

StyleGANs have shown impressive results on data generation and manipulation in recent years, thanks to its disentangled style latent space. A lot of efforts have been made in inver…

cs.LG2020

Maximum-Entropy Adversarial Data Augmentation for Improved Generalization and Robustness

Long Zhao, Ting Liu, Xi Peng +1

Adversarial data augmentation has shown promise for training robust deep neural networks against unforeseen data shifts or corruptions. However, it is difficult to define heuristic…

cs.LG2019

Unsupervised Domain Adaptation via Calibrating Uncertainties

Ligong Han, Yang Zou, Ruijiang Gao +2

Unsupervised domain adaptation (UDA) aims at inferring class labels for unlabeled target domain given a related labeled source dataset. Intuitively, a model trained on source domai…

cs.CV2025

SINE: SINgle Image Editing with Text-to-Image Diffusion Models

Zhixing Zhang, Ligong Han, Arnab Ghosh +2

Recent works on diffusion models have demonstrated a strong capability for conditioning image generation, e.g., text-guided image synthesis. Such success inspires many efforts tryi…

cs.CV2024

Continuous Spatio-Temporal Memory Networks for 4D Cardiac Cine MRI Segmentation

Meng Ye, Bingyu Xin, Leon Axel +1

Current cardiac cine magnetic resonance image (cMR) studies focus on the end diastole (ED) and end systole (ES) phases, while ignoring the abundant temporal information in the whol…

eess.IV2020

Synthetic Learning: Learn From Distributed Asynchronized Discriminator GAN Without Sharing Medical Image Data

Qi Chang, Hui Qu, Yikai Zhang +4

In this paper, we propose a data privacy-preserving and communication efficient distributed GAN learning framework named Distributed Asynchronized Discriminator GAN (AsynDGAN). Our…

eess.IV2023

On the Challenges and Perspectives of Foundation Models for Medical Image Analysis

Shaoting Zhang, Dimitris Metaxas

This article discusses the opportunities, applications and future directions of large-scale pre-trained models, i.e., foundation models, for analyzing medical images. Medical found…

cs.CV2017

Automatic Vertebra Labeling in Large-Scale 3D CT using Deep Image-to-Image Network with Message Passing and Sparsity Regularization

Dong Yang, Tao Xiong, Daguang Xu +10

Automatic localization and labeling of vertebra in 3D medical images plays an important role in many clinical tasks, including pathological diagnosis, surgical planning and postope…

cs.CV2020

Unbiased Auxiliary Classifier GANs with MINE

Ligong Han, Anastasis Stathopoulos, Tao Xue +1

Auxiliary Classifier GANs (AC-GANs) are widely used conditional generative models and are capable of generating high-quality images. Previous work has pointed out that AC-GAN learn…

cs.CV2022

Global Matching with Overlapping Attention for Optical Flow Estimation

Shiyu Zhao, Long Zhao, Zhixing Zhang +2

Optical flow estimation is a fundamental task in computer vision. Recent direct-regression methods using deep neural networks achieve remarkable performance improvement. However, t…

cs.CV2024

Neural Deformable Models for 3D Bi-Ventricular Heart Shape Reconstruction and Modeling from 2D Sparse Cardiac Magnetic Resonance Imaging

Meng Ye, Dong Yang, Mikael Kanski +2

We propose a novel neural deformable model (NDM) targeting at the reconstruction and modeling of 3D bi-ventricular shape of the heart from 2D sparse cardiac magnetic resonance (CMR…

eess.IV2021

Liver Fibrosis and NAS scoring from CT images using self-supervised learning and texture encoding

Ananya Jana, Hui Qu, Carlos D. Minacapelli +3

Non-alcoholic fatty liver disease (NAFLD) is one of the most common causes of chronic liver diseases (CLD) which can progress to liver cancer. The severity and treatment of NAFLD i…

cs.CV2020

Error-Bounded Correction of Noisy Labels

Songzhu Zheng, Pengxiang Wu, Aman Goswami +3

To collect large scale annotated data, it is inevitable to introduce label noise, i.e., incorrect class labels. To be robust against label noise, many successful methods rely on th…

cs.CV2024

Aligning Human Knowledge with Visual Concepts Towards Explainable Medical Image Classification

Yunhe Gao, Difei Gu, Mu Zhou +1

Although explainability is essential in the clinical diagnosis, most deep learning models still function as black boxes without elucidating their decision-making process. In this s…

cs.CV2020

OnlineAugment: Online Data Augmentation with Less Domain Knowledge

Zhiqiang Tang, Yunhe Gao, Leonid Karlinsky +3

Data augmentation is one of the most important tools in training modern deep neural networks. Recently, great advances have been made in searching for optimal augmentation policies…

cs.CV2023

Improving Tuning-Free Real Image Editing with Proximal Guidance

Ligong Han, Song Wen, Qi Chen +13

DDIM inversion has revealed the remarkable potential of real image editing within diffusion-based methods. However, the accuracy of DDIM reconstruction degrades as larger classifie…

cs.CV2021

DeepTag: An Unsupervised Deep Learning Method for Motion Tracking on Cardiac Tagging Magnetic Resonance Images

Meng Ye, Mikael Kanski, Dong Yang +5

Cardiac tagging magnetic resonance imaging (t-MRI) is the gold standard for regional myocardium deformation and cardiac strain estimation. However, this technique has not been wide…

cs.CV2018

Toward Marker-free 3D Pose Estimation in Lifting: A Deep Multi-view Solution

Rahil Mehrizi, Xi Peng, Zhiqiang Tang +3

Lifting is a common manual material handling task performed in the workplaces. It is considered as one of the main risk factors for Work-related Musculoskeletal Disorders. To impro…

stat.ML2019

Self-Attention Generative Adversarial Networks

Han Zhang, Ian Goodfellow, Dimitris Metaxas +1

In this paper, we propose the Self-Attention Generative Adversarial Network (SAGAN) which allows attention-driven, long-range dependency modeling for image generation tasks. Tradit…

stat.ME2009

Learning with Structured Sparsity

Junzhou Huang, Tong Zhang, Dimitris Metaxas

This paper investigates a new learning formulation called structured sparsity, which is a natural extension of the standard sparsity concept in statistical learning and compressive…

cs.CV2025

SnapGen-V: Generating a Five-Second Video within Five Seconds on a Mobile Device

Yushu Wu, Zhixing Zhang, Yanyu Li +11

We have witnessed the unprecedented success of diffusion-based video generation over the past year. Recently proposed models from the community have wielded the power to generate c…

cs.CV2022

DeepRecon: Joint 2D Cardiac Segmentation and 3D Volume Reconstruction via A Structure-Specific Generative Method

Qi Chang, Zhennan Yan, Mu Zhou +8

Joint 2D cardiac segmentation and 3D volume reconstruction are fundamental to building statistical cardiac anatomy models and understanding functional mechanisms from motion patter…

cs.CV2017

StackGAN: Text to Photo-realistic Image Synthesis with Stacked Generative Adversarial Networks

Han Zhang, Tao Xu, Hongsheng Li +4

Synthesizing high-quality images from text descriptions is a challenging problem in computer vision and has many practical applications. Samples generated by existing text-to-image…

cs.CV2022

A Topological Filter for Learning with Label Noise

Pengxiang Wu, Songzhu Zheng, Mayank Goswami +2

Noisy labels can impair the performance of deep neural networks. To tackle this problem, in this paper, we propose a new method for filtering label noise. Unlike most existing meth…

cs.LG2026

Spectral-Aware Analytic Class-Incremental Learning for Long-Tailed Distributions

Quyen Tran, Hai Nguyen, Quan Dao +4

Analytic Continual Learning (ACL) offers a computationally efficient alternative to gradient-based approaches. Recent ACL methods are based on Recursive Least Squares (RLS) and hav…

cs.LG2026

TokUR: Token-Level Uncertainty Estimation for Large Language Model Reasoning

Tunyu Zhang, Haizhou Shi, Yibin Wang +9

While Large Language Models (LLMs) have demonstrated impressive capabilities, their output quality remains inconsistent across various application scenarios, making it difficult to…

cs.CV2025

Discrete Noise Inversion for Next-scale Autoregressive Text-based Image Editing

Quan Dao, Xiaoxiao He, Ligong Han +6

Visual autoregressive models (VAR) have recently emerged as a promising class of generative models, achieving performance comparable to diffusion models in text-to-image generation…

cs.LG2020

Robust Conditional GAN from Uncertainty-Aware Pairwise Comparisons

Ligong Han, Ruijiang Gao, Mun Kim +3

Conditional generative adversarial networks have shown exceptional generation performance over the past few years. However, they require large numbers of annotations. To address th…

cs.CV2024

Resolving Inconsistent Semantics in Multi-Dataset Image Segmentation

Qilong Zhangli, Di Liu, Abhishek Aich +2

Leveraging multiple training datasets to scale up image segmentation models is beneficial for increasing robustness and semantic understanding. Individual datasets have well-define…

cs.LG2018

Improving GANs Using Optimal Transport

Tim Salimans, Han Zhang, Alec Radford +1

We present Optimal Transport GAN (OT-GAN), a variant of generative adversarial nets minimizing a new metric measuring the distance between the generator distribution and the data d…

cs.CV2022

Exploiting Unlabeled Data with Vision and Language Models for Object Detection

Shiyu Zhao, Zhixing Zhang, Samuel Schulter +5

Building robust and generic object detection frameworks requires scaling to larger label spaces and bigger training datasets. However, it is prohibitively costly to acquire annotat…

eess.IV2024

Learning Volumetric Neural Deformable Models to Recover 3D Regional Heart Wall Motion from Multi-Planar Tagged MRI

Meng Ye, Bingyu Xin, Bangwei Guo +2

Multi-planar tagged MRI is the gold standard for regional heart wall motion evaluation. However, accurate recovery of the 3D true heart wall motion from a set of 2D apparent motion…

cs.CV2024

AVID: Any-Length Video Inpainting with Diffusion Model

Zhixing Zhang, Bichen Wu, Xiaoyan Wang +6

Recent advances in diffusion models have successfully enabled text-guided image inpainting. While it seems straightforward to extend such editing capability into the video domain,…

cs.CV2022

Diffusion Guided Domain Adaptation of Image Generators

Kunpeng Song, Ligong Han, Bingchen Liu +2

Can a text-to-image diffusion model be used as a training objective for adapting a GAN generator to another domain? In this paper, we show that the classifier-free guidance can be…

cs.CV2020

Learn distributed GAN with Temporary Discriminators

Hui Qu, Yikai Zhang, Qi Chang +3

In this work, we propose a method for training distributed GAN with sequential temporary discriminators. Our proposed method tackles the challenge of training GAN in the federated…

cs.CV2026

Overcoming the Curvature Bottleneck in MeanFlow

Xinxi Zhang, Shiwei Tan, Quang Nguyen +7

MeanFlow offers a promising framework for one-step generative modeling by directly learning a mean-velocity field, bypassing expensive numerical integration. However, we find that…

cs.CV2021

Enabling Data Diversity: Efficient Automatic Augmentation via Regularized Adversarial Training

Yunhe Gao, Zhiqiang Tang, Mu Zhou +1

Data augmentation has proved extremely useful by increasing training data variance to alleviate overfitting and improve deep neural networks' generalization performance. In medical…

cs.LG2021

Training Federated GANs with Theoretical Guarantees: A Universal Aggregation Approach

Yikai Zhang, Hui Qu, Qi Chang +3

Recently, Generative Adversarial Networks (GANs) have demonstrated their potential in federated learning, i.e., learning a centralized model from data privately hosted by multiple…

cs.CV2018

Quantized Densely Connected U-Nets for Efficient Landmark Localization

Zhiqiang Tang, Xi Peng, Shijie Geng +3

In this paper, we propose quantized densely connected U-Nets for efficient visual landmark localization. The idea is that features of the same semantic meanings are globally reused…

cs.CV2024

SF-V: Single Forward Video Generation Model

Zhixing Zhang, Yanyu Li, Yushu Wu +9

Diffusion-based video generation models have demonstrated remarkable success in obtaining high-fidelity videos through the iterative denoising process. However, these models requir…

cs.CV2026

LUCID-SAE: Learning Unified Vision-Language Sparse Codes for Interpretable Concept Discovery

Difei Gu, Yunhe Gao, Gerasimos Chatzoudis +6

Sparse autoencoders (SAEs) offer a natural path toward comparable explanations across different representation spaces. However, current SAEs are trained per modality, producing dic…

cs.CV2023

OmniLabel: A Challenging Benchmark for Language-Based Object Detection

Samuel Schulter, Vijay Kumar B G, Yumin Suh +4

Language-based object detection is a promising direction towards building a natural interface to describe objects in images that goes far beyond plain category names. While recent…

cs.CV2025

DiMSUM: Diffusion Mamba -- A Scalable and Unified Spatial-Frequency Method for Image Generation

Hao Phung, Quan Dao, Trung Dao +3

We introduce a novel state-space architecture for diffusion models, effectively harnessing spatial and frequency information to enhance the inductive bias towards local features in…

cs.CV2025

Anatomy-VLM: A Fine-grained Vision-Language Model for Medical Interpretation

Difei Gu, Yunhe Gao, Mu Zhou +1

Accurate disease interpretation from radiology remains challenging due to imaging heterogeneity. Achieving expert-level diagnostic decisions requires integration of subtle image fe…

cs.CV2025

Improved Training Technique for Latent Consistency Models

Quan Dao, Khanh Doan, Di Liu +2

Consistency models are a new family of generative models capable of producing high-quality samples in either a single step or multiple steps. Recently, consistency models have demo…

cs.CV2019

Effective 3D Humerus and Scapula Extraction using Low-contrast and High-shape-variability MR Data

Xiaoxiao He, Chaowei Tan, Yuting Qiao +3

For the initial shoulder preoperative diagnosis, it is essential to obtain a three-dimensional (3D) bone mask from medical images, e.g., magnetic resonance (MR). However, obtaining…