Publications (98)
Region Proposal Rectification Towards Robust Instance Segmentation of Biological Images
Qilong Zhangli, Jingru Yi, Di Liu +10
Top-down instance segmentation framework has shown its superiority in object detection compared to the bottom-up framework. While it is efficient in addressing over-segmentation, t…
Automatic Liver Segmentation Using an Adversarial Image-to-Image Network
Dong Yang, Daguang Xu, S. Kevin Zhou +5
Automatic liver segmentation in 3D medical images is essential in many clinical applications, such as pathological diagnosis of hepatic diseases, surgical planning, and postoperati…
UTNet: A Hybrid Transformer Architecture for Medical Image Segmentation
Yunhe Gao, Mu Zhou, Dimitris Metaxas
Transformer architecture has emerged to be successful in a number of natural language processing tasks. However, its applications to medical vision remain largely unexplored. In th…
BLoB: Bayesian Low-Rank Adaptation by Backpropagation for Large Language Models
Yibin Wang, Haizhou Shi, Ligong Han +2
Large Language Models (LLMs) often suffer from overconfidence during inference, particularly when adapted to downstream domain-specific tasks with limited data. Previous work addre…
A Dynamic Data Driven Approach for Explainable Scene Understanding
Zachary A Daniels, Dimitris Metaxas
Scene-understanding is an important topic in the area of Computer Vision, and illustrates computational challenges with applications to a wide range of domains including remote sen…
Dual Projection Generative Adversarial Networks for Conditional Image Generation
Ligong Han, Martin Renqiang Min, Anastasis Stathopoulos +4
Conditional Generative Adversarial Networks (cGANs) extend the standard unconditional GAN framework to learning joint data-label distributions from samples, and have been establish…
CrossNorm and SelfNorm for Generalization under Distribution Shifts
Zhiqiang Tang, Yunhe Gao, Yi Zhu +3
Traditional normalization techniques (e.g., Batch Normalization and Instance Normalization) generally and simplistically assume that training and test data follow the same distribu…
Brain Segmentation from k-space with End-to-end Recurrent Attention Network
Qiaoying Huang, Xiao Chen, Dimitris Metaxas +1
The task of medical image segmentation commonly involves an image reconstruction step to convert acquired raw data to images before any analysis. However, noises, artifacts and los…
Reconstruction-Based Disentanglement for Pose-invariant Face Recognition
Xi Peng, Xiang Yu, Kihyuk Sohn +2
Deep neural networks (DNNs) trained on large-scale datasets have recently achieved impressive improvements in face recognition. But a persistent challenge remains to develop method…
Cross-Space Distillation: Teaching One-Step Students with Modern Diffusion Teachers
Anh Nguyen, Ngan Nguyen, Duc Vu +11
Modern one-step diffusion models achieve impressive quality through distribution-based timestep distillation. Yet, they rely on a critical assumption: Teacher and Student must inha…
New Capability to Look Up an ASL Sign from a Video Example
Carol Neidle, Augustine Opoku, Carey Ballard +4
Looking up an unknown sign in an ASL dictionary can be difficult. Most ASL dictionaries are organized based on English glosses, despite the fact that (1) there is no convention for…
Modality Bank: Learn multi-modality images across data centers without sharing medical data
Qi Chang, Hui Qu, Zhennan Yan +3
Multi-modality images have been widely used and provide comprehensive information for medical image analysis. However, acquiring all modalities among all institutes is costly and o…
Enhanced MRI Reconstruction Network using Neural Architecture Search
Qiaoying Huang, Dong Yang, Yikun Xian +4
The accurate reconstruction of under-sampled magnetic resonance imaging (MRI) data using modern deep learning technology, requires significant effort to design the necessary comple…
MRI Reconstruction via Cascaded Channel-wise Attention Network
Qiaoying Huang, Dong Yang, Pengxiang Wu +3
We consider an MRI reconstruction problem with input of k-space data at a very low undersampled rate. This can practically benefit patient due to reduced time of MRI scan, but it i…
Contrastive and Selective Hidden Embeddings for Medical Image Segmentation
Zhuowei Li, Zihao Liu, Zhiqiang Hu +5
Medical image segmentation has been widely recognized as a pivot procedure for clinical diagnosis, analysis, and treatment planning. However, the laborious and expensive annotation…
Learning to Forecast and Refine Residual Motion for Image-to-Video Generation
Long Zhao, Xi Peng, Yu Tian +2
We consider the problem of image-to-video translation, where an input image is translated into an output video containing motions of a single object. Recent methods for such proble…
Automatic Tooth Segmentation from 3D Dental Model using Deep Learning: A Quantitative Analysis of what can be learnt from a Single 3D Dental Model
Ananya Jana, Hrebesh Molly Subhash, Dimitris Metaxas
3D tooth segmentation is an important task for digital orthodontics. Several Deep Learning methods have been proposed for automatic tooth segmentation from 3D dental models or intr…
Measure Anatomical Thickness from Cardiac MRI with Deep Neural Networks
Qiaoying Huang, Eric Z. Chen, Hanchao Yu +4
Accurate estimation of shape thickness from medical images is crucial in clinical applications. For example, the thickness of myocardium is one of the key to cardiac disease diagno…
RadAlign: Advancing Radiology Report Generation with Vision-Language Concept Alignment
Difei Gu, Yunhe Gao, Yang Zhou +2
Automated chest radiographs interpretation requires both accurate disease classification and detailed radiology report generation, presenting a significant challenge in the clinica…
PC-U Net: Learning to Jointly Reconstruct and Segment the Cardiac Walls in 3D from CT Data
Meng Ye, Qiaoying Huang, Dong Yang +4
The 3D volumetric shape of the heart's left ventricle (LV) myocardium (MYO) wall provides important information for diagnosis of cardiac disease and invasive procedure navigation.…
DICE: Discrete Inversion Enabling Controllable Editing for Multinomial Diffusion and Masked Generative Models
Xiaoxiao He, Quan Dao, Ligong Han +14
Discrete diffusion models have achieved success in tasks like image generation and masked language modeling but face limitations in controlled content editing. We introduce DICE (D…
ASL Video Corpora & Sign Bank: Resources Available through the American Sign Language Linguistic Research Project (ASLLRP)
Carol Neidle, Augustine Opoku, Dimitris Metaxas
The American Sign Language Linguistic Research Project (ASLLRP) provides Internet access to high-quality ASL video data, generally including front and side views and a close-up of…
Constructive Assimilation: Boosting Contrastive Learning Performance through View Generation Strategies
Ligong Han, Seungwook Han, Shivchander Sudalairaj +8
Transformations based on domain expertise (expert transformations), such as random-resized-crop and color-jitter, have proven critical to the success of contrastive learning techni…
Semi-Supervised Segmentation of Radiation-Induced Pulmonary Fibrosis from Lung CT Scans with Multi-Scale Guided Dense Attention
Guotai Wang, Shuwei Zhai, Giovanni Lasio +7
Computed Tomography (CT) plays an important role in monitoring radiation-induced Pulmonary Fibrosis (PF), where accurate segmentation of the PF lesions is highly desired for diagno…
DeDPO: Debiased Direct Preference Optimization for Diffusion Models
Khiem Pham, Quang Nguyen, Tung Nguyen +4
Direct Preference Optimization (DPO) has emerged as a predominant alignment method for diffusion models, facilitating off-policy training without explicit reward modeling. However,…
Point Cloud Processing via Recurrent Set Encoding
Pengxiang Wu, Chao Chen, Jingru Yi +1
We present a new permutation-invariant network for 3D point cloud processing. Our network is composed of a recurrent set encoder and a convolutional feature aggregator. Given an un…
Score-Guided Diffusion for 3D Human Recovery
Anastasis Stathopoulos, Ligong Han, Dimitris Metaxas
We present Score-Guided Human Mesh Recovery (ScoreHMR), an approach for solving inverse problems for 3D human pose and shape reconstruction. These inverse problems involve fitting…
Learning Articulated Shape with Keypoint Pseudo-labels from Web Images
Anastasis Stathopoulos, Georgios Pavlakos, Ligong Han +1
This paper shows that it is possible to learn models for monocular 3D reconstruction of articulated objects (e.g., horses, cows, sheep), using as few as 50-150 images labeled with…
Data Augmentation for High-Fidelity Generation of CAR-T/NK Immunological Synapse Images
Xiang Zhang, Boxuan Zhang, Alireza Naghizadeh +4
Chimeric antigen receptor (CAR)-T and NK cell immunotherapies have transformed cancer treatment, and recent studies suggest that the quality of the CAR-T/NK cell immunological syna…
Improving Compositional Text-to-image Generation with Large Vision-Language Models
Song Wen, Guian Fang, Renrui Zhang +3
Recent advancements in text-to-image models, particularly diffusion models, have shown significant promise. However, compositional text-to-image models frequently encounter difficu…
A Manifold View of Adversarial Risk
Wenjia Zhang, Yikai Zhang, Xiaoling Hu +3
The adversarial risk of a machine learning model has been widely studied. Most previous works assume that the data lies in the whole ambient space. We propose to take a new angle a…
Show Me What and Tell Me How: Video Synthesis via Multimodal Conditioning
Ligong Han, Jian Ren, Hsin-Ying Lee +5
Most methods for conditional video synthesis use a single modality as the condition. This comes with major limitations. For example, it is problematic for a model conditioned on an…
Deep Learning based NAS Score and Fibrosis Stage Prediction from CT and Pathology Data
Ananya Jana, Hui Qu, Puru Rattan +3
Non-Alcoholic Fatty Liver Disease (NAFLD) is becoming increasingly prevalent in the world population. Without diagnosis at the right time, NAFLD can lead to non-alcoholic steatohep…
Jointly Optimize Data Augmentation and Network Training: Adversarial Data Augmentation in Human Pose Estimation
Xi Peng, Zhiqiang Tang, Fei Yang +2
Random data augmentation is a critical technique to avoid overfitting in training deep neural network models. However, data augmentation and network training are usually treated as…
MADCrowner: Margin Aware Dental Crown Design with Template Deformation and Refinement
Linda Wei, Chang Liu, Wenran Zhang +9
Dental crown restoration is one of the most common treatment modalities for tooth defect, where personalized dental crown design is critical. While computer-aided design (CAD) syst…
StackGAN++: Realistic Image Synthesis with Stacked Generative Adversarial Networks
Han Zhang, Tao Xu, Hongsheng Li +4
Although Generative Adversarial Networks (GANs) have shown remarkable success in various tasks, they still face challenges in generating high quality images. In this paper, we prop…
Self-Corrected Flow Distillation for Consistent One-Step and Few-Step Text-to-Image Generation
Quan Dao, Hao Phung, Trung Dao +2
Flow matching has emerged as a promising framework for training generative models, demonstrating impressive empirical performance while offering relative ease of training compared…
An Optimal Transport-driven Approach for Cultivating Latent Space in Online Incremental Learning
Quyen Tran, Hai Nguyen, Hoang Phan +6
In online incremental learning, data continuously arrives with substantial distributional shifts, creating a significant challenge because previous samples have limited replay valu…
Oriented Object Detection in Aerial Images with Box Boundary-Aware Vectors
Jingru Yi, Pengxiang Wu, Bo Liu +3
Oriented object detection in aerial images is a challenging task as the objects in aerial images are displayed in arbitrary directions and are usually densely packed. Current orien…
MotionNet: Joint Perception and Motion Prediction for Autonomous Driving Based on Bird's Eye View Maps
Pengxiang Wu, Siheng Chen, Dimitris Metaxas
The ability to reliably perceive the environmental states, particularly the existence of objects and their motion behavior, is crucial for autonomous driving. In this work, we prop…
Global and Local Interpretation of black-box Machine Learning models to determine prognostic factors from early COVID-19 data
Ananya Jana, Carlos D. Minacapelli, Vinod Rustgi +1
The COVID-19 corona virus has claimed 4.1 million lives, as of July 24, 2021. A variety of machine learning models have been applied to related data to predict important factors su…
VerSe: Integrating Multiple Queries as Prompts for Versatile Cardiac MRI Segmentation
Bangwei Guo, Meng Ye, Yunhe Gao +3
Despite the advances in learning-based image segmentation approach, the accurate segmentation of cardiac structures from magnetic resonance imaging (MRI) remains a critical challen…
Rate-My-LoRA: Efficient and Adaptive Federated Model Tuning for Cardiac MRI Segmentation
Xiaoxiao He, Haizhou Shi, Ligong Han +7
Cardiovascular disease (CVD) and cardiac dyssynchrony are major public health problems in the United States. Precise cardiac image segmentation is crucial for extracting quantitati…
Interactive Reinforcement Learning for Object Grounding via Self-Talking
Yan Zhu, Shaoting Zhang, Dimitris Metaxas
Humans are able to identify a referred visual object in a complex scene via a few rounds of natural language communications. Success communication requires both parties to engage a…
MedForge: Building Medical Foundation Models Like Open Source Software Development
Zheling Tan, Kexin Ding, Jin Gao +4
Foundational models (FMs) have made significant strides in the healthcare domain. Yet the data silo challenge and privacy concern remain in healthcare systems, hindering safe medic…
MPDiT: Multi-Patch Global-to-Local Transformer Architecture For Efficient Flow Matching and Diffusion Model
Quan Dao, Dimitris Metaxas
Transformer architectures, particularly Diffusion Transformers (DiTs), have become widely used in diffusion and flow-matching models due to their strong performance compared to con…
AutoEdit: Automatic Hyperparameter Tuning for Image Editing
Chau Pham, Quan Dao, Mahesh Bhosale +3
Recent advances in diffusion models have revolutionized text-guided image editing, yet existing editing methods face critical challenges in hyperparameter identification. To get th…
Adaptive low rank and sparse decomposition of video using compressive sensing
Fei Yang, Hong Jiang, Zuowei Shen +2
We address the problem of reconstructing and analyzing surveillance videos using compressive sensing. We develop a new method that performs video reconstruction by low rank and spa…
Learning View-Disentangled Human Pose Representation by Contrastive Cross-View Mutual Information Maximization
Long Zhao, Yuxiao Wang, Jiaping Zhao +7
We introduce a novel representation learning method to disentangle pose-dependent as well as view-dependent factors from 2D human poses. The method trains a network using cross-vie…
TransFusion: Multi-view Divergent Fusion for Medical Image Segmentation with Transformers
Di Liu, Yunhe Gao, Qilong Zhangli +8
Combining information from multi-view images is crucial to improve the performance and robustness of automated methods for disease diagnosis. However, due to the non-alignment char…
SVDiff: Compact Parameter Space for Diffusion Fine-Tuning
Ligong Han, Yinxiao Li, Han Zhang +3
Diffusion models have achieved remarkable success in text-to-image generation, enabling the creation of high-quality images from text prompts or other modalities. However, existing…
K-Prism: A Knowledge-Guided and Prompt Integrated Universal Medical Image Segmentation Model
Bangwei Guo, Yunhe Gao, Meng Ye +4
Medical image segmentation is fundamental to clinical decision-making, yet existing models remain fragmented. They are usually trained on single knowledge sources and specific to i…
AE-StyleGAN: Improved Training of Style-Based Auto-Encoders
Ligong Han, Sri Harsha Musunuri, Martin Renqiang Min +3
StyleGANs have shown impressive results on data generation and manipulation in recent years, thanks to its disentangled style latent space. A lot of efforts have been made in inver…
Maximum-Entropy Adversarial Data Augmentation for Improved Generalization and Robustness
Long Zhao, Ting Liu, Xi Peng +1
Adversarial data augmentation has shown promise for training robust deep neural networks against unforeseen data shifts or corruptions. However, it is difficult to define heuristic…
Unsupervised Domain Adaptation via Calibrating Uncertainties
Ligong Han, Yang Zou, Ruijiang Gao +2
Unsupervised domain adaptation (UDA) aims at inferring class labels for unlabeled target domain given a related labeled source dataset. Intuitively, a model trained on source domai…
SINE: SINgle Image Editing with Text-to-Image Diffusion Models
Zhixing Zhang, Ligong Han, Arnab Ghosh +2
Recent works on diffusion models have demonstrated a strong capability for conditioning image generation, e.g., text-guided image synthesis. Such success inspires many efforts tryi…
Continuous Spatio-Temporal Memory Networks for 4D Cardiac Cine MRI Segmentation
Meng Ye, Bingyu Xin, Leon Axel +1
Current cardiac cine magnetic resonance image (cMR) studies focus on the end diastole (ED) and end systole (ES) phases, while ignoring the abundant temporal information in the whol…
Synthetic Learning: Learn From Distributed Asynchronized Discriminator GAN Without Sharing Medical Image Data
Qi Chang, Hui Qu, Yikai Zhang +4
In this paper, we propose a data privacy-preserving and communication efficient distributed GAN learning framework named Distributed Asynchronized Discriminator GAN (AsynDGAN). Our…
On the Challenges and Perspectives of Foundation Models for Medical Image Analysis
Shaoting Zhang, Dimitris Metaxas
This article discusses the opportunities, applications and future directions of large-scale pre-trained models, i.e., foundation models, for analyzing medical images. Medical found…
Automatic Vertebra Labeling in Large-Scale 3D CT using Deep Image-to-Image Network with Message Passing and Sparsity Regularization
Dong Yang, Tao Xiong, Daguang Xu +10
Automatic localization and labeling of vertebra in 3D medical images plays an important role in many clinical tasks, including pathological diagnosis, surgical planning and postope…
Unbiased Auxiliary Classifier GANs with MINE
Ligong Han, Anastasis Stathopoulos, Tao Xue +1
Auxiliary Classifier GANs (AC-GANs) are widely used conditional generative models and are capable of generating high-quality images. Previous work has pointed out that AC-GAN learn…
Global Matching with Overlapping Attention for Optical Flow Estimation
Shiyu Zhao, Long Zhao, Zhixing Zhang +2
Optical flow estimation is a fundamental task in computer vision. Recent direct-regression methods using deep neural networks achieve remarkable performance improvement. However, t…
Neural Deformable Models for 3D Bi-Ventricular Heart Shape Reconstruction and Modeling from 2D Sparse Cardiac Magnetic Resonance Imaging
Meng Ye, Dong Yang, Mikael Kanski +2
We propose a novel neural deformable model (NDM) targeting at the reconstruction and modeling of 3D bi-ventricular shape of the heart from 2D sparse cardiac magnetic resonance (CMR…
Liver Fibrosis and NAS scoring from CT images using self-supervised learning and texture encoding
Ananya Jana, Hui Qu, Carlos D. Minacapelli +3
Non-alcoholic fatty liver disease (NAFLD) is one of the most common causes of chronic liver diseases (CLD) which can progress to liver cancer. The severity and treatment of NAFLD i…
Error-Bounded Correction of Noisy Labels
Songzhu Zheng, Pengxiang Wu, Aman Goswami +3
To collect large scale annotated data, it is inevitable to introduce label noise, i.e., incorrect class labels. To be robust against label noise, many successful methods rely on th…
Aligning Human Knowledge with Visual Concepts Towards Explainable Medical Image Classification
Yunhe Gao, Difei Gu, Mu Zhou +1
Although explainability is essential in the clinical diagnosis, most deep learning models still function as black boxes without elucidating their decision-making process. In this s…
OnlineAugment: Online Data Augmentation with Less Domain Knowledge
Zhiqiang Tang, Yunhe Gao, Leonid Karlinsky +3
Data augmentation is one of the most important tools in training modern deep neural networks. Recently, great advances have been made in searching for optimal augmentation policies…
Improving Tuning-Free Real Image Editing with Proximal Guidance
Ligong Han, Song Wen, Qi Chen +13
DDIM inversion has revealed the remarkable potential of real image editing within diffusion-based methods. However, the accuracy of DDIM reconstruction degrades as larger classifie…
DeepTag: An Unsupervised Deep Learning Method for Motion Tracking on Cardiac Tagging Magnetic Resonance Images
Meng Ye, Mikael Kanski, Dong Yang +5
Cardiac tagging magnetic resonance imaging (t-MRI) is the gold standard for regional myocardium deformation and cardiac strain estimation. However, this technique has not been wide…
Toward Marker-free 3D Pose Estimation in Lifting: A Deep Multi-view Solution
Rahil Mehrizi, Xi Peng, Zhiqiang Tang +3
Lifting is a common manual material handling task performed in the workplaces. It is considered as one of the main risk factors for Work-related Musculoskeletal Disorders. To impro…
Self-Attention Generative Adversarial Networks
Han Zhang, Ian Goodfellow, Dimitris Metaxas +1
In this paper, we propose the Self-Attention Generative Adversarial Network (SAGAN) which allows attention-driven, long-range dependency modeling for image generation tasks. Tradit…
Learning with Structured Sparsity
Junzhou Huang, Tong Zhang, Dimitris Metaxas
This paper investigates a new learning formulation called structured sparsity, which is a natural extension of the standard sparsity concept in statistical learning and compressive…
SnapGen-V: Generating a Five-Second Video within Five Seconds on a Mobile Device
Yushu Wu, Zhixing Zhang, Yanyu Li +11
We have witnessed the unprecedented success of diffusion-based video generation over the past year. Recently proposed models from the community have wielded the power to generate c…
DeepRecon: Joint 2D Cardiac Segmentation and 3D Volume Reconstruction via A Structure-Specific Generative Method
Qi Chang, Zhennan Yan, Mu Zhou +8
Joint 2D cardiac segmentation and 3D volume reconstruction are fundamental to building statistical cardiac anatomy models and understanding functional mechanisms from motion patter…
StackGAN: Text to Photo-realistic Image Synthesis with Stacked Generative Adversarial Networks
Han Zhang, Tao Xu, Hongsheng Li +4
Synthesizing high-quality images from text descriptions is a challenging problem in computer vision and has many practical applications. Samples generated by existing text-to-image…
A Topological Filter for Learning with Label Noise
Pengxiang Wu, Songzhu Zheng, Mayank Goswami +2
Noisy labels can impair the performance of deep neural networks. To tackle this problem, in this paper, we propose a new method for filtering label noise. Unlike most existing meth…
Spectral-Aware Analytic Class-Incremental Learning for Long-Tailed Distributions
Quyen Tran, Hai Nguyen, Quan Dao +4
Analytic Continual Learning (ACL) offers a computationally efficient alternative to gradient-based approaches. Recent ACL methods are based on Recursive Least Squares (RLS) and hav…
TokUR: Token-Level Uncertainty Estimation for Large Language Model Reasoning
Tunyu Zhang, Haizhou Shi, Yibin Wang +9
While Large Language Models (LLMs) have demonstrated impressive capabilities, their output quality remains inconsistent across various application scenarios, making it difficult to…
Discrete Noise Inversion for Next-scale Autoregressive Text-based Image Editing
Quan Dao, Xiaoxiao He, Ligong Han +6
Visual autoregressive models (VAR) have recently emerged as a promising class of generative models, achieving performance comparable to diffusion models in text-to-image generation…
Robust Conditional GAN from Uncertainty-Aware Pairwise Comparisons
Ligong Han, Ruijiang Gao, Mun Kim +3
Conditional generative adversarial networks have shown exceptional generation performance over the past few years. However, they require large numbers of annotations. To address th…
Resolving Inconsistent Semantics in Multi-Dataset Image Segmentation
Qilong Zhangli, Di Liu, Abhishek Aich +2
Leveraging multiple training datasets to scale up image segmentation models is beneficial for increasing robustness and semantic understanding. Individual datasets have well-define…
Improving GANs Using Optimal Transport
Tim Salimans, Han Zhang, Alec Radford +1
We present Optimal Transport GAN (OT-GAN), a variant of generative adversarial nets minimizing a new metric measuring the distance between the generator distribution and the data d…
Exploiting Unlabeled Data with Vision and Language Models for Object Detection
Shiyu Zhao, Zhixing Zhang, Samuel Schulter +5
Building robust and generic object detection frameworks requires scaling to larger label spaces and bigger training datasets. However, it is prohibitively costly to acquire annotat…
Learning Volumetric Neural Deformable Models to Recover 3D Regional Heart Wall Motion from Multi-Planar Tagged MRI
Meng Ye, Bingyu Xin, Bangwei Guo +2
Multi-planar tagged MRI is the gold standard for regional heart wall motion evaluation. However, accurate recovery of the 3D true heart wall motion from a set of 2D apparent motion…
AVID: Any-Length Video Inpainting with Diffusion Model
Zhixing Zhang, Bichen Wu, Xiaoyan Wang +6
Recent advances in diffusion models have successfully enabled text-guided image inpainting. While it seems straightforward to extend such editing capability into the video domain,…
Diffusion Guided Domain Adaptation of Image Generators
Kunpeng Song, Ligong Han, Bingchen Liu +2
Can a text-to-image diffusion model be used as a training objective for adapting a GAN generator to another domain? In this paper, we show that the classifier-free guidance can be…
Learn distributed GAN with Temporary Discriminators
Hui Qu, Yikai Zhang, Qi Chang +3
In this work, we propose a method for training distributed GAN with sequential temporary discriminators. Our proposed method tackles the challenge of training GAN in the federated…
Overcoming the Curvature Bottleneck in MeanFlow
Xinxi Zhang, Shiwei Tan, Quang Nguyen +7
MeanFlow offers a promising framework for one-step generative modeling by directly learning a mean-velocity field, bypassing expensive numerical integration. However, we find that…
Enabling Data Diversity: Efficient Automatic Augmentation via Regularized Adversarial Training
Yunhe Gao, Zhiqiang Tang, Mu Zhou +1
Data augmentation has proved extremely useful by increasing training data variance to alleviate overfitting and improve deep neural networks' generalization performance. In medical…
Training Federated GANs with Theoretical Guarantees: A Universal Aggregation Approach
Yikai Zhang, Hui Qu, Qi Chang +3
Recently, Generative Adversarial Networks (GANs) have demonstrated their potential in federated learning, i.e., learning a centralized model from data privately hosted by multiple…
Quantized Densely Connected U-Nets for Efficient Landmark Localization
Zhiqiang Tang, Xi Peng, Shijie Geng +3
In this paper, we propose quantized densely connected U-Nets for efficient visual landmark localization. The idea is that features of the same semantic meanings are globally reused…
SF-V: Single Forward Video Generation Model
Zhixing Zhang, Yanyu Li, Yushu Wu +9
Diffusion-based video generation models have demonstrated remarkable success in obtaining high-fidelity videos through the iterative denoising process. However, these models requir…
LUCID-SAE: Learning Unified Vision-Language Sparse Codes for Interpretable Concept Discovery
Difei Gu, Yunhe Gao, Gerasimos Chatzoudis +6
Sparse autoencoders (SAEs) offer a natural path toward comparable explanations across different representation spaces. However, current SAEs are trained per modality, producing dic…
OmniLabel: A Challenging Benchmark for Language-Based Object Detection
Samuel Schulter, Vijay Kumar B G, Yumin Suh +4
Language-based object detection is a promising direction towards building a natural interface to describe objects in images that goes far beyond plain category names. While recent…
DiMSUM: Diffusion Mamba -- A Scalable and Unified Spatial-Frequency Method for Image Generation
Hao Phung, Quan Dao, Trung Dao +3
We introduce a novel state-space architecture for diffusion models, effectively harnessing spatial and frequency information to enhance the inductive bias towards local features in…
Anatomy-VLM: A Fine-grained Vision-Language Model for Medical Interpretation
Difei Gu, Yunhe Gao, Mu Zhou +1
Accurate disease interpretation from radiology remains challenging due to imaging heterogeneity. Achieving expert-level diagnostic decisions requires integration of subtle image fe…
Improved Training Technique for Latent Consistency Models
Quan Dao, Khanh Doan, Di Liu +2
Consistency models are a new family of generative models capable of producing high-quality samples in either a single step or multiple steps. Recently, consistency models have demo…
Effective 3D Humerus and Scapula Extraction using Low-contrast and High-shape-variability MR Data
Xiaoxiao He, Chaowei Tan, Yuting Qiao +3
For the initial shoulder preoperative diagnosis, it is essential to obtain a three-dimensional (3D) bone mask from medical images, e.g., magnetic resonance (MR). However, obtaining…