8 papers
S2O: Enhancing Adversarial Training with Second-Order Statistics of Weights
Gaojie Jin, Xinping Yi, Wei Huang +2
Adversarial training has emerged as a highly effective way to improve the robustness of deep neural networks (DNNs). It is typically conceptualized as a min-max optimization proble…
Towards A Unified PAC-Bayesian Framework for Norm-based Generalization Bounds
Xinping Yi, Gaojie Jin, Xiaowei Huang +1
Understanding the generalization behavior of deep neural networks remains a fundamental challenge in modern statistical learning theory. Among existing approaches, PAC-Bayesian nor…
Reconcile Certified Robustness and Accuracy for DNN-based Smoothed Majority Vote Classifier
Gaojie Jin, Xinping Yi, Xiaowei Huang
Within the PAC-Bayesian framework, the Gibbs classifier (defined on a posterior ) and the corresponding -weighted majority vote classifier are commonly used to analyze the ge…
Activation Steering Meets Preference Optimization: Defense Against Jailbreaks in Vision Language Models
Sihao Wu, Gaojie Jin, Wei Huang +2
Vision Language Models (VLMs) have demonstrated impressive capabilities in integrating visual and textual information for understanding and reasoning, but remain highly vulnerable…
Shapley Uncertainty in Natural Language Generation
Meilin Zhu, Gaojie Jin, Xiaowei Huang +1
In question-answering tasks, determining when to trust the outputs is crucial to the alignment of large language models (LLMs). Kuhn et al. (2023) introduces semantic entropy as a…
ThermoRL:Structure-Aware Reinforcement Learning for Protein Mutation Design to Enhance Thermostability
Xiangwen Wang, Gaojie Jin, Xiaowei Huang +1
Designing mutations to optimize protein thermostability remains challenging due to the complex relationship between sequence variations, structural dynamics, and thermostability, o…