3 papers
cs.CV2026
Steal the Patch Size: Adversarially Manipulate Vision-Language Models
Kai Hu, Akash Bharadwaj, Weichen Yu +1
We present a black-box model-stealing attack that recovers private vision-tokenizer configurations of deployed vision-language models (VLMs), including the visual patch size and in…
cs.CL2026
Configurable Reward Model for Balanced Safety Alignment
Zhengping Jiang, Mehran Khodabandeh, Akash Bharadwaj +5
Aligning large language models (LLMs) to heterogeneous and rapidly evolving safety requirements remains a critical challenge. Existing instruction-tuned LLMs and standalone safety…
cs.CL2025
Jailbreak-Zero: A Path to Pareto Optimal Red Teaming for Large Language Models
Kai Hu, Abhinav Aggarwal, Mehran Khodabandeh +6
This paper introduces Jailbreak-Zero, a novel red teaming methodology that shifts the paradigm of Large Language Model (LLM) safety evaluation from a constrained example-based appr…