10 papers
ICO: Enhancing Semantic-Shift Jailbreaks via Iterative Context Optimization
Hujian Zhu, Yihao Huang, Felix Juefei-Xu +5
Foundation models have achieved remarkable success across diverse tasks, but they remain vulnerable. To investigate such vulnerabilities, semantic-shift jailbreaks have recently em…
Beyond Pixels: Semantic-aware Typographic Attack for Geo-Privacy Protection
Jiayi Zhu, Yihao Huang, Yue Cao +5
Large Visual Language Models (LVLMs) now pose a serious yet overlooked privacy threat, as they can infer a social media user's geolocation directly from shared images, leading to u…
Pushing the Limits of Safety: A Technical Report on the ATLAS Challenge 2025
Zonghao Ying, Siyang Wu, Run Hao +44
Multimodal Large Language Models (MLLMs) have enabled transformative advancements across diverse applications but remain susceptible to safety threats, especially jailbreak attacks…
Concept Guided Co-salient Object Detection
Jiayi Zhu, Qing Guo, Felix Juefei-Xu +3
Co-salient object detection (Co-SOD) aims to identify common salient objects across a group of related images. While recent methods have made notable progress, they typically rely…
AcTracer: Active Testing of Large Language Model via Multi-Stage Sampling
Yuheng Huang, Jiayang Song, Qiang Hu +2
Performance evaluation plays a crucial role in the development life cycle of large language models (LLMs). It estimates the model's capability, elucidates behavior characteristics,…
A Comprehensive Survey in LLM(-Agent) Full Stack Safety: Data, Training and Deployment
Kun Wang, Guibin Zhang, Zhenhong Zhou +100
The remarkable success of Large Language Models (LLMs) has illuminated a promising pathway toward achieving Artificial General Intelligence for both academic and industrial communi…