papers

Publications (79)

cs.CV2025

Wan: Open and Advanced Large-Scale Video Generative Models

Team Wan, Ang Wang, Baole Ai +58

This report presents Wan, a comprehensive and open suite of video foundation models designed to push the boundaries of video generation. Built upon the mainstream diffusion transfo…

cs.AI2024

ChatLogic: Integrating Logic Programming with Large Language Models for Multi-Step Reasoning

Zhongsheng Wang, Jiamou Liu, Qiming Bao +2

Large language models (LLMs) such as ChatGPT and GPT-4 have demonstrated impressive capabilities in various generative tasks. However, their performance is often hampered by limita…

cs.LG2021

Learning Diverse-Structured Networks for Adversarial Robustness

Xuefeng Du, Jingfeng Zhang, Bo Han +5

In adversarial training (AT), the main focus has been the objective and optimizer while the model has been less studied, so that the models being used are still those classic ones…

cs.CV2021

DRDF: Determining the Importance of Different Multimodal Information with Dual-Router Dynamic Framework

Haiwen Hong, Xuan Jin, Yin Zhang +4

In multimodal tasks, we find that the importance of text and image modal information is different for different input cases, and for this motivation, we propose a high-performance…

physics.optics2023

Topological Holography and Storage with Optical Knots and Links

Ling-Jun Kong, Jingfeng Zhang, Furong Zhang +1

After more than 70 years of development, holography has become an essential tool of modern optics in many applications. In fact, for various applications of different kinds of holo…

cs.CV2025

Model-agnostic Adversarial Attack and Defense for Vision-Language-Action Models

Haochuan Xu, Yun Sing Koh, Shuhuai Huang +4

Vision-Language-Action (VLA) models have achieved revolutionary progress in robot learning, enabling robots to execute complex physical robot tasks from natural language instructio…

cs.CV2021

RAMS-Trans: Recurrent Attention Multi-scale Transformer forFine-grained Image Recognition

Yunqing Hu, Xuan Jin, Yin Zhang +4

In fine-grained image recognition (FGIR), the localization and amplification of region attention is an important factor, which has been explored a lot by convolutional neural netwo…

cs.CV2024

StyleBooth: Image Style Editing with Multimodal Instruction

Zhen Han, Chaojie Mao, Zeyinzi Jiang +2

Given an original image, image editing aims to generate an image that align with the provided instruction. The challenges are to accept multimodal inputs as instructions and a scar…

cs.CR2023

Assessing Vulnerabilities of Adversarial Learning Algorithm through Poisoning Attacks

Jingfeng Zhang, Bo Song, Bo Han +3

Adversarial training (AT) is a robust learning algorithm that can defend against adversarial attacks in the inference phase and mitigate the side effects of corrupted data in the t…

cs.LG2021

CIFS: Improving Adversarial Robustness of CNNs via Channel-wise Importance-based Feature Selection

Hanshu Yan, Jingfeng Zhang, Gang Niu +3

We investigate the adversarial robustness of CNNs from the perspective of channel-wise activations. By comparing \textit{non-robust} (normally trained) and \textit{robustified} (ad…

cs.LG2021

Understanding the Interaction of Adversarial Training with Noisy Labels

Jianing Zhu, Jingfeng Zhang, Bo Han +5

Noisy labels (NL) and adversarial examples both undermine trained models, but interestingly they have hitherto been studied independently. A recent adversarial training (AT) study…

cs.CV2024

Day-Night Adaptation: An Innovative Source-free Adaptation Framework for Medical Image Segmentation

Ziyang Chen, Yiwen Ye, Yongsheng Pan +3

Distribution shifts widely exist in medical images acquired from different medical centres, hindering the deployment of semantic segmentation models trained on one centre (source d…

cs.LG2020

Hierarchically Fair Federated Learning

Jingfeng Zhang, Cheng Li, Antonio Robles-Kelly +1

When the federated learning is adopted among competitive agents with siloed datasets, agents are self-interested and participate only if they are fairly rewarded. To encourage the…

cs.LG2022

Adversarial Training with Complementary Labels: On the Benefit of Gradually Informative Attacks

Jianan Zhou, Jianing Zhu, Jingfeng Zhang +4

Adversarial training (AT) with imperfect supervision is significant but receives limited attention. To push AT towards more practical scenarios, we explore a brand new yet challeng…

cs.LG2026

Controllable Concept Bottleneck Models

Hongbin Lin, Chenyang Ren, Juangui Xu +7

Concept Bottleneck Models (CBMs) have garnered much attention for their ability to elucidate the prediction process through a human-understandable concept layer. However, most prev…

cs.CV2024

ACE: All-round Creator and Editor Following Instructions via Diffusion Transformer

Zhen Han, Zeyinzi Jiang, Yulin Pan +5

Diffusion models have emerged as a powerful generative technology and have been found to be applicable in various scenarios. Most existing foundational diffusion models are primari…

cs.LG2025

Adversarial Preference Learning for Robust LLM Alignment

Yuanfu Wang, Pengyu Wang, Chenyang Xi +13

Modern language models often rely on Reinforcement Learning from Human Feedback (RLHF) to encourage safe behaviors. However, they remain vulnerable to adversarial attacks due to th…

cs.LG2024

Privacy-Preserving Heterogeneous Federated Learning for Sensitive Healthcare Data

Yukai Xu, Jingfeng Zhang, Yujie Gu

In the realm of healthcare where decentralized facilities are prevalent, machine learning faces two major challenges concerning the protection of data and models. The data-level ch…

cs.CV2024

Text Guided Image Editing with Automatic Concept Locating and Forgetting

Jia Li, Lijie Hu, Zhixian He +3

With the advancement of image-to-image diffusion models guided by text, significant progress has been made in image editing. However, a persistent challenge remains in seamlessly i…

cs.CV2022

Diverse Instance Discovery: Vision-Transformer for Instance-Aware Multi-Label Image Recognition

Yunqing Hu, Xuan Jin, Yin Zhang +5

Previous works on multi-label image recognition (MLIR) usually use CNNs as a starting point for research. In this paper, we take pure Vision Transformer (ViT) as the research base…

cs.CR2025

Goal-oriented Backdoor Attack against Vision-Language-Action Models via Physical Objects

Zirun Zhou, Zhengyang Xiao, Haochuan Xu +3

Recent advances in vision-language-action (VLA) models have greatly improved embodied AI, enabling robots to follow natural language instructions and perform diverse tasks. However…

cs.LG2025

One Stone, Two Birds: Enhancing Adversarial Defense Through the Lens of Distributional Discrepancy

Jiacheng Zhang, Benjamin I. P. Rubinstein, Jingfeng Zhang +1

Statistical adversarial data detection (SADD) detects whether an upcoming batch contains adversarial examples (AEs) by measuring the distributional discrepancies between clean exam…

cs.LG2022

On the Effectiveness of Adversarial Training against Backdoor Attacks

Yinghua Gao, Dongxian Wu, Jingfeng Zhang +4

DNNs' demand for massive data forces practitioners to collect data from the Internet without careful check due to the unacceptable cost, which brings potential risks of backdoor at…

cs.CV2025

ICE-Bench: A Unified and Comprehensive Benchmark for Image Creating and Editing

Yulin Pan, Xiangteng He, Chaojie Mao +4

Image generation has witnessed significant advancements in the past few years. However, evaluating the performance of image generation models remains a formidable challenge. In thi…

cs.CV2025

InstructAttribute: Fine-grained Object Attributes editing with Instruction

Xingxi Yin, Jingfeng Zhang, Yue Deng +3

Text-to-image (T2I) diffusion models are widely used in image editing due to their powerful generative capabilities. However, achieving fine-grained control over specific object at…

cs.CV2022

Accelerating Score-based Generative Models for High-Resolution Image Synthesis

Hengyuan Ma, Li Zhang, Xiatian Zhu +2

Score-based generative models (SGMs) have recently emerged as a promising class of generative models. The key idea is to produce high-quality images by recurrently adding Gaussian…

cs.LG2018

Smooth Inter-layer Propagation of Stabilized Neural Networks for Classification

Jingfeng Zhang, Laura Wynter

Recent work has studied the reasons for the remarkable performance of deep neural networks in image classification. We examine batch normalization on the one hand and the dynamical…

cs.CV2019

Towards Robust ResNet: A Small Step but A Giant Leap

Jingfeng Zhang, Bo Han, Laura Wynter +2

This paper presents a simple yet principled approach to boosting the robustness of the residual network (ResNet) that is motivated by the dynamical system perspective. Namely, a de…

cs.CV2026

Wan-Weaver: Interleaved Multi-modal Generation via Decoupled Training

Jinbo Xing, Zeyinzi Jiang, Yuxiang Tuo +15

Recent unified models have made unprecedented progress in both understanding and generation. However, while most of them accept multi-modal inputs, they typically produce only sing…

cs.LG2025

Editable Concept Bottleneck Models

Lijie Hu, Chenyang Ren, Zhengyu Hu +5

Concept Bottleneck Models (CBMs) have garnered much attention for their ability to elucidate the prediction process through a humanunderstandable concept layer. However, most previ…

cs.AI2026

AutoRAS: Learning Robust Agentic Systems with Primitive Representations

Yang Yue, Xuancheng Zhu, Yuyang Ma +7

The automated design of agentic systems offers a promising pathway for scaling large language models (LLMs) beyond single-agent reasoning. While prior work has advanced task perfor…

cs.LG2019

Where is the Bottleneck of Adversarial Learning with Unlabeled Data?

Jingfeng Zhang, Bo Han, Gang Niu +2

Deep neural networks (DNNs) are incredibly brittle due to adversarial examples. To robustify DNNs, adversarial training was proposed, which requires large-scale but well-labeled da…

cs.LG2020

Robust Federated Recommendation System

Chen Chen, Jingfeng Zhang, Anthony K. H. Tung +2

Federated recommendation systems can provide good performance without collecting users' private data, making them attractive. However, they are susceptible to low-cost poisoning at…

cs.LG2026

Short-length Adversarial Training Helps LLMs Defend Long-length Jailbreak Attacks: Theoretical and Empirical Evidence

Shaopeng Fu, Liang Ding, Jingfeng Zhang +1

Jailbreak attacks against large language models (LLMs) aim to induce harmful behaviors in LLMs through carefully crafted adversarial prompts. To mitigate attacks, one way is to per…

cs.CV2024

An Individual Identity-Driven Framework for Animal Re-Identification

Yihao Wu, Di Zhao, Jingfeng Zhang +1

Reliable re-identification of individuals within large wildlife populations is crucial for biological studies, ecological research, and wildlife conservation. Classic computer visi…

cs.LG2025

Learning without Isolation: Pathway Protection for Continual Learning

Zhikang Chen, Abudukelimu Wuerkaixi, Sen Cui +10

Deep networks are prone to catastrophic forgetting during sequential task learning, i.e., losing the knowledge about old tasks upon learning new tasks. To this end, continual learn…

cs.CV2026

Avatar V: Scaling Video-Reference Avatar Video Generation

Benjamin Liang, Ce Chen, Desmond Lin +20

Generating avatar videos that are not merely visually similar to a target individual but behaviorally recognizable, faithfully reproducing their talking rhythm, gestural tendencies…

cs.LG2023

AutoLoRa: A Parameter-Free Automated Robust Fine-Tuning Framework

Xilie Xu, Jingfeng Zhang, Mohan Kankanhalli

Robust Fine-Tuning (RFT) is a low-cost strategy to obtain adversarial robustness in downstream applications, without requiring a lot of computational resources and collecting signi…

cs.CV2026

A UAV-Based Multi-Modal Vision System for Automated Sideslope Deformation Monitoring and Hazard Detection

Jingfeng Zhang, Yi Li, Xianchong Liang +1

Slope hazards constitute a major safety threat to expressway infrastructure, and their evolution is typically manifested as slow surface deformation. Conventional manual inspection…

cs.CV2024

Improving Accuracy-robustness Trade-off via Pixel Reweighted Adversarial Training

Jiacheng Zhang, Feng Liu, Dawei Zhou +2

Adversarial training (AT) trains models using adversarial examples (AEs), which are natural images modified with specific perturbations to mislead the model. These perturbations ar…

cs.LG2024

Balancing Similarity and Complementarity for Federated Learning

Kunda Yan, Sen Cui, Abudukelimu Wuerkaixi +5

In mobile and IoT systems, Federated Learning (FL) is increasingly important for effectively using data while maintaining user privacy. One key challenge in FL is managing statisti…

cs.CV2024

Towards Multi-dimensional Explanation Alignment for Medical Classification

Lijie Hu, Songning Lai, Wenshuo Chen +5

The lack of interpretability in the field of medical image analysis has significant ethical and legal implications. Existing interpretable methods in this domain encounter several…

cs.LG2026

Benign Overfitting in Adversarial Training for Vision Transformers

Jiaming Zhang, Meng Ding, Shaopeng Fu +2

Despite the remarkable success of Vision Transformers (ViTs) across a wide range of vision tasks, recent studies have revealed that they remain vulnerable to adversarial examples,…

cs.LG2025

Robust Learning of Diffusion Models with Extremely Noisy Conditions

Xin Chen, Gillian Dobbie, Xinyu Wang +3

Conditional diffusion models have the generative controllability by incorporating external conditions. However, their performance significantly degrades with noisy conditions, such…

cs.LG2022

Reliable Adversarial Distillation with Unreliable Teachers

Jianing Zhu, Jiangchao Yao, Bo Han +6

In ordinary distillation, student networks are trained with soft labels (SLs) given by pretrained teacher networks, and students are expected to improve upon teachers since SLs are…

cs.RO2025

A tutorial note on collecting simulated data for vision-language-action models

Heran Wu, Zirun Zhou, Jingfeng Zhang

Traditional robotic systems typically decompose intelligence into independent modules for computer vision, natural language processing, and motion control. Vision-Language-Action (…

cs.CV2026

Wan-Image: Pushing the Boundaries of Generative Visual Intelligence

Chaojie Mao, Chen-Wei Xie, Chongyang Zhong +55

We present Wan-Image, a unified visual generation system explicitly engineered to paradigm-shift image generation models from casual synthesizers into professional-grade productivi…

cs.RO2026

Concept-Based Dictionary Learning for Inference-Time Safety in Vision Language Action Models

Siqi Wen, Shu Yang, Shaopeng Fu +3

Vision Language Action (VLA) models close the perception action loop by translating multimodal instructions into executable behaviors, but this very capability magnifies safety ris…

cs.SD2022

WaveFuzz: A Clean-Label Poisoning Attack to Protect Your Voice

Yunjie Ge, Qian Wang, Jingfeng Zhang +3

People are not always receptive to their voice data being collected and misused. Training the audio intelligence systems needs these data to build useful features, but the cost for…

cs.LG2021

Geometry-aware Instance-reweighted Adversarial Training

Jingfeng Zhang, Jianing Zhu, Gang Niu +3

In adversarial machine learning, there was a common belief that robustness and accuracy hurt each other. The belief was challenged by recent studies where we can maintain the robus…

cs.CV2025

Stable Vision Concept Transformers for Medical Diagnosis

Lijie Hu, Songning Lai, Yuan Hua +3

Transparency is a paramount concern in the medical field, prompting researchers to delve into the realm of explainable AI (XAI). Among these XAI methods, Concept Bottleneck Models…

cs.CR2026

Safety at Scale: A Comprehensive Survey of Large Model and Agent Safety

Xingjun Ma, Yifeng Gao, Yixu Wang +45

The rapid advancement of large models, driven by their exceptional abilities in learning and generalization through large-scale pre-training, has reshaped the landscape of Artifici…

cs.LG2021

Guided Interpolation for Adversarial Training

Chen Chen, Jingfeng Zhang, Xilie Xu +4

To enhance adversarial robustness, adversarial training learns deep neural networks on the adversarial variants generated by their natural data. However, as the training progresses…

cs.LG2022

Adversarial Attack and Defense for Non-Parametric Two-Sample Tests

Xilie Xu, Jingfeng Zhang, Feng Liu +2

Non-parametric two-sample tests (TSTs) that judge whether two sets of samples are drawn from the same distribution, have been widely used in the analysis of critical data. People t…

cs.CR2023

An LLM can Fool Itself: A Prompt-Based Adversarial Attack

Xilie Xu, Keyi Kong, Ning Liu +4

The wide-ranging applications of large language models (LLMs), especially in safety-critical domains, necessitate the proper evaluation of the LLM's adversarial robustness. This pa…

cs.CV2024

Fair Text-to-Image Diffusion via Fair Mapping

Jia Li, Lijie Hu, Jingfeng Zhang +3

In this paper, we address the limitations of existing text-to-image diffusion models in generating demographically fair results when given human-related descriptions. These models…

physics.optics2022

High capacity topological coding based on nested vortex knots and links

Ling-Jun Kong, Weixuan Zhang, Peng +6

Optical knots and links have attracted great attention because of their exotic topological characteristics. Recent investigations have shown that the information encoding based on…

cs.CV2025

Locate, Assign, Refine: Taming Customized Promptable Image Inpainting

Yulin Pan, Chaojie Mao, Zeyinzi Jiang +3

Prior studies have made significant progress in image inpainting guided by either text description or subject image. However, the research on inpainting with flexible guidance or c…

cs.LG2025

Accurate Forgetting for Heterogeneous Federated Continual Learning

Abudukelimu Wuerkaixi, Sen Cui, Jingfeng Zhang +6

Recent years have witnessed a burgeoning interest in federated learning (FL). However, the contexts in which clients engage in sequential learning remain under-explored. Bridging F…

cs.LG2024

BadLabel: A Robust Perspective on Evaluating and Enhancing Label-noise Learning

Jingfeng Zhang, Bo Song, Haohan Wang +4

Label-noise learning (LNL) aims to increase the model's generalization given training data with noisy labels. To facilitate practical LNL algorithms, researchers have proposed diff…

cs.LG2022

NoiLIn: Improving Adversarial Training and Correcting Stereotype of Noisy Labels

Jingfeng Zhang, Xilie Xu, Bo Han +4

Adversarial training (AT) formulated as the minimax optimization problem can effectively enhance the model's robustness against adversarial attacks. The existing AT methods mainly…

cs.CV2025

Make Me Happier: Evoking Emotions Through Image Diffusion Models

Qing Lin, Jingfeng Zhang, Yew-Soon Ong +1

Despite the rapid progress in image generation, emotional image editing remains under-explored. The semantics, context, and structure of an image can evoke emotional responses, mak…

cs.CV2025

VACE: All-in-One Video Creation and Editing

Zeyinzi Jiang, Zhen Han, Chaojie Mao +3

Diffusion Transformer has demonstrated powerful capability and scalability in generating high-quality images and videos. Further pursuing the unification of generation and editing…

cs.LG2024

Privacy-Preserving Low-Rank Adaptation against Membership Inference Attacks for Latent Diffusion Models

Zihao Luo, Xilie Xu, Feng Liu +3

Low-rank adaptation (LoRA) is an efficient strategy for adapting latent diffusion models (LDMs) on a private dataset to generate specific images by minimizing the adaptation loss.…

eess.IV2022

Towards Adversarially Robust Deep Image Denoising

Hanshu Yan, Jingfeng Zhang, Jiashi Feng +2

This work systematically investigates the adversarial robustness of deep image denoisers (DIDs), i.e, how well DIDs can recover the ground truth from noisy observations degraded by…

cs.LG2025

Dissecting Representation Misalignment in Contrastive Learning via Influence Function

Lijie Hu, Chenyang Ren, Huanyi Xie +5

Contrastive learning, commonly applied in large-scale multimodal models, often relies on data from diverse and often unreliable sources, which can include misaligned or mislabeled…

cs.LG2023

Enhancing Adversarial Contrastive Learning via Adversarial Invariant Regularization

Xilie Xu, Jingfeng Zhang, Feng Liu +2

Adversarial contrastive learning (ACL) is a technique that enhances standard contrastive learning (SCL) by incorporating adversarial data to learn a robust representation that can…

physics.optics2023

High-dimensional entanglement-enabled holography for quantum encryption

Ling-Jun Kong, Yifan Sun, Furong Zhang +2

As an important imaging technique, holography has been realized with different physical dimensions of light,including polarization, wavelength, and time. Recently, quantum holograp…

cs.CR2022

FuncFooler: A Practical Black-box Attack Against Learning-based Binary Code Similarity Detection Methods

Lichen Jia, Bowen Tang, Chenggang Wu +6

The binary code similarity detection (BCSD) method measures the similarity of two binary executable codes. Recently, the learning-based BCSD methods have achieved great success, ou…

cs.CV2025

ACE++: Instruction-Based Image Creation and Editing via Context-Aware Content Filling

Chaojie Mao, Jingfeng Zhang, Yulin Pan +4

We report ACE++, an instruction-based diffusion framework that tackles various image generation and editing tasks. Inspired by the input format for the inpainting task proposed by…

cs.CV2023

SCEdit: Efficient and Controllable Image Diffusion Generation via Skip Connection Editing

Zeyinzi Jiang, Chaojie Mao, Yulin Pan +2

Image diffusion models have been utilized in various tasks, such as text-to-image generation and controllable image synthesis. Recent research has introduced tuning methods that ma…

cs.CV2025

ColorEdit: Training-free Image-Guided Color editing with diffusion model

Xingxi Yin, Zhi Li, Jingfeng Zhang +2

Text-to-image (T2I) diffusion models, with their impressive generative capabilities, have been adopted for image editing tasks, demonstrating remarkable efficacy. However, due to a…

cond-mat.mtrl-sci2025

Deteriorated Interlayer Coupling in Twisted Bilayer Cobaltites

Dongke Rong, Xiuqi Chen, Shengru Chen +15

A wealth of remarkable behaviors is observed at the interfaces between magnetic oxides due to the coexistence of Coulomb repulsion and interatomic exchange interactions. While prev…

physics.flu-dyn2014

Lattice Boltzmann Model for The Volume-Averaged Navier-Stokes Equations

Jingfeng Zhang, Limin Wang, Jie Ouyang

A numerical method, based on the discrete lattice Boltzmann equation, is presented for solving the volume-averaged Navier-Stokes equations. With a modified equilibrium distribution…

cs.LG2023

Efficient Adversarial Contrastive Learning via Robustness-Aware Coreset Selection

Xilie Xu, Jingfeng Zhang, Feng Liu +2

Adversarial contrastive learning (ACL) does not require expensive data annotations but outputs a robust representation that withstands adversarial attacks and also generalizes to a…

cs.LG2021

Maximum Mean Discrepancy Test is Aware of Adversarial Attacks

Ruize Gao, Feng Liu, Jingfeng Zhang +4

The maximum mean discrepancy (MMD) test could in principle detect any distributional discrepancy between two datasets. However, it has been shown that the MMD test is unaware of ad…

cs.LG2020

Attacks Which Do Not Kill Training Make Adversarial Learning Stronger

Jingfeng Zhang, Xilie Xu, Bo Han +4

Adversarial training based on the minimax formulation is necessary for obtaining adversarial robustness of trained models. However, it is conservative or even pessimistic so that i…

cond-mat.mtrl-sci2025

Lattice distortions and non-sluggish diffusion in BCC refractory high entropy alloys

Jingfeng Zhang, Xiang Xu, Fritz Körmann +10

Refractory high-entropy alloys (RHEAs) have emerged as promising candidates for extreme high-temperature applications, for example, in next-generation turbines and nuclear reactors…

cs.CV2023

GAT: Guided Adversarial Training with Pareto-optimal Auxiliary Tasks

Salah Ghamizi, Jingfeng Zhang, Maxime Cordy +3

While leveraging additional training data is well established to improve adversarial robustness, it incurs the unavoidable cost of data collection and the heavy computation to trai…