2 papers
cs.CV2026
Break the Brake, Not the Wheel: Untargeted Jailbreak via Entropy Maximization
Mengqi He, Xinyu Tian, Xin Shen +6
Recent studies show that gradient-based universal image jailbreaks on vision-language models (VLMs) exhibit little or no cross-model transferability, casting doubt on the feasibili…
cs.CV2026
Probing and Bridging Geometry-Interaction Cues for Affordance Reasoning in Vision Foundation Models
Qing Zhang, Xuesong Li, Jing Zhang
What does it mean for a visual system to truly understand affordance? We argue that this understanding hinges on two complementary capacities: geometric perception, which identifie…