4 papers
Activation Steering Meets Preference Optimization: Defense Against Jailbreaks in Vision Language Models
Sihao Wu, Gaojie Jin, Wei Huang +2
Vision Language Models (VLMs) have demonstrated impressive capabilities in integrating visual and textual information for understanding and reasoning, but remain highly vulnerable…
ThermoRL:Structure-Aware Reinforcement Learning for Protein Mutation Design to Enhance Thermostability
Xiangwen Wang, Gaojie Jin, Xiaowei Huang +1
Designing mutations to optimize protein thermostability remains challenging due to the complex relationship between sequence variations, structural dynamics, and thermostability, o…
A Black-Box Evaluation Framework for Semantic Robustness in Bird's Eye View Detection
Fu Wang, Yanghao Zhang, Xiangyu Yin +4
Camera-based Bird's Eye View (BEV) perception models receive increasing attention for their crucial role in autonomous driving, a domain where concerns about the robustness and rel…
BEARD: Benchmarking the Adversarial Robustness for Dataset Distillation
Zheng Zhou, Wenquan Feng, Shuchang Lyu +3
Dataset Distillation (DD) is an emerging technique that compresses large-scale datasets into significantly smaller synthesized datasets while preserving high test performance and e…