2 papers
cs.SD2025
StyleBreak: Revealing Alignment Vulnerabilities in Large Audio-Language Models via Style-Aware Audio Jailbreak
Hongyi Li, Chengxuan Zhou, Chu Wang +5
Large Audio-language Models (LAMs) have recently enabled powerful speech-based interactions by coupling audio encoders with Large Language Models (LLMs). However, the security of L…
cs.CR2024
JailPO: A Novel Black-box Jailbreak Framework via Preference Optimization against Aligned LLMs
Hongyi Li, Jiawei Ye, Jie Wu +3
Large Language Models (LLMs) aligned with human feedback have recently garnered significant attention. However, it remains vulnerable to jailbreak attacks, where adversaries manipu…