10 papers
ImageAuditor: Membership Inference Attack against Image-based Retrieval-Augmented Generation
Jinghuai Zhang, Pengyue Yu, Zhexiao Lin +3
Image-based Retrieval-Augmented Generation (IRAG) conditions a frozen generator on reference images retrieved from an external database, supporting both text-to-image (T2I) and que…
RogueMerge: Robust and Unified Attacks against LLM Model Merging
Jinghuai Zhang, Yetian He, Kunlin Cai +3
Model merging composes specialized capabilities into a single LLM by aggregating task vectors sourced from unverified public platforms, exposing a critical supply-chain attack surf…
What-If World: A Causal Benchmark for General World Models in Embodied Scenarios
Kunlin Cai, Rui Song, Jinghuai Zhang +7
Video generation models are increasingly used as world simulators for tasks like driving and robotic manipulation. What matters in these settings is not whether a single video look…
HIDBench: Benchmarking Large Language Models for Host-Based Intrusion Detection
Danyu Sun, Jinghuai Zhang, Yuan Tian +1
Recent benchmark efforts have advanced the evaluation of large language models (LLMs) in cybersecurity, including tasks such as penetration testing and vulnerability identification…
FineSteer: A Unified Framework for Fine-Grained Inference-Time Steering in Large Language Models
Zixuan Weng, Jinghuai Zhang, Kunlin Cai +3
Large language models (LLMs) often exhibit undesirable behaviors, such as safety violations and hallucinations. Although inference-time steering offers a cost-effective way to adju…
The Landscape of Prompt Injection Threats in LLM Agents: From Taxonomy to Analysis
Peiran Wang, Xinfeng Li, Chong Xiang +5
The evolution of Large Language Models (LLMs) has resulted in a paradigm shift towards autonomous agents, necessitating robust security against Prompt Injection (PI) vulnerabilitie…