3 papers
cs.LG2025
Reconcile Certified Robustness and Accuracy for DNN-based Smoothed Majority Vote Classifier
Gaojie Jin, Xinping Yi, Xiaowei Huang
Within the PAC-Bayesian framework, the Gibbs classifier (defined on a posterior ) and the corresponding -weighted majority vote classifier are commonly used to analyze the ge…
cs.CV2025
Activation Steering Meets Preference Optimization: Defense Against Jailbreaks in Vision Language Models
Sihao Wu, Gaojie Jin, Wei Huang +2
Vision Language Models (VLMs) have demonstrated impressive capabilities in integrating visual and textual information for understanding and reasoning, but remain highly vulnerable…
cs.LG2025
POT: Inducing Overthinking in LLMs via Black-Box Iterative Optimization
Xinyu Li, Tianjin Huang, Ronghui Mu +2
Recent advances in Chain-of-Thought (CoT) prompting have substantially enhanced the reasoning capabilities of large language models (LLMs), enabling sophisticated problem-solving t…