collaborators

7 papers

cs.CR2026

A Multimodal Automatic Redteaming Evaluation based on Atomic Jailbreak Strategy Decoupling and Combination

Shiji Zhao, Yuxuan Zhou, Chen Xiong +3

Multimodal Large Language Models (MLLMs) have achieved impressive progress in image-text comprehension and generation, yet they remain susceptible to jailbreak attacks that can tri…

cs.CR2026

From Role Prompt to Infinite Thinking: Exploiting Persona Conditioning for Inference Cost Attacks in LLMs

Zhiyi Mou, Wangze Ni, Tianfang Xiao +6

LLMs are increasingly deployed in real-world applications, making inference efficiency and service reliability critical concerns due to their substantial computational costs. Howev…

cs.LG2025

Improving Generative Ad Text on Facebook using Reinforcement Learning

Daniel R. Jiang, Alex Nikulkov, Yu-Chia Chen +2

Generative artificial intelligence (AI), in particular large language models (LLMs), is poised to drive transformative economic change. LLMs are pre-trained on vast text data to le…

cs.CR2025

Why does weak-OOD help? A Further Step Towards Understanding Jailbreaking VLMs

Yuxuan Zhou, Yuzhao Peng, Yang Bai +7

Large Vision-Language Models (VLMs) are susceptible to jailbreak attacks: researchers have developed a variety of attack strategies that can successfully bypass the safety mechanis…

cs.CR2025

JPRO: Automated Multimodal Jailbreaking via Multi-Agent Collaboration Framework

Yuxuan Zhou, Yang Bai, Kuofeng Gao +2

The widespread application of large VLMs makes ensuring their secure deployment critical. While recent studies have demonstrated jailbreak attacks on VLMs, existing approaches are…

cs.CL2025

Explore Data Left Behind in Reinforcement Learning for Reasoning Language Models

Chenxi Liu, Junjie Liang, Yuqi Jia +4

Reinforcement Learning with Verifiable Rewards (RLVR) has emerged as an effective approach for improving the reasoning abilities of large language models (LLMs). The Group Relative…