activity
20242026
collaborators

6 papers

cs.AI2026

Fair in Mind, Fair in Action? A Synchronous Benchmark for Understanding and Generation in UMLLMs

Yiran Zhao, Lu Zhou, Xiaogang Xu +3

As artificial intelligence (AI) is increasingly deployed across domains, ensuring fairness has become a core challenge. However, the field faces a "Tower of Babel'' dilemma: fairne…

cs.CL2025

Jailbreaking Commercial Black-Box LLMs with Explicitly Harmful Prompts

Chiyu Zhang, Lu Zhou, Xiaogang Xu +3

Existing black-box jailbreak attacks achieve certain success on non-reasoning models but degrade significantly on recent SOTA reasoning models. To improve attack ability, inspired…

cs.LG2024

DR-Encoder: Encode Low-rank Gradients with Random Prior for Large Language Models Differentially Privately

Huiwen Wu, Deyi Zhang, Xiaohan Li +3

The emergence of the Large Language Model (LLM) has shown their superiority in a wide range of disciplines, including language understanding and translation, relational logic reaso…

cs.AI2024

GADT: Enhancing Transferable Adversarial Attacks through Gradient-guided Adversarial Data Transformation

Yating Ma, Xiaogang Xu, Liming Fang +1

Current Transferable Adversarial Examples (TAE) are primarily generated by adding Adversarial Noise (AN). Recent studies emphasize the importance of optimizing Data Augmentation (D…

cs.CL2024

Iter-AHMCL: Alleviate Hallucination for Large Language Model via Iterative Model-level Contrastive Learning

Huiwen Wu, Xiaohan Li, Xiaogang Xu +3

The development of Large Language Models (LLMs) has significantly advanced various AI applications in commercial and scientific research fields, such as scientific literature summa…

cs.CV2024

Adversarial Attacks of Vision Tasks in the Past 10 Years: A Survey

Chiyu Zhang, Lu Zhou, Xiaogang Xu +2

With the advent of Large Vision-Language Models (LVLMs), new attack vectors, such as cognitive bias, prompt injection, and jailbreaking, have emerged. Understanding these attacks p…