direction-based interventions 1inference-time defenses 1jailbreak mitigation 1model robustness 1vision-language models 1
From the 1 of 3 linked papers with an AI index.
3 papers
cs.AI2026
A Cross-Architecture Audit of Direction-Based Inference-Time Defences in Vision-Language Models
Xiangyu Yin, Tora Bodin, Rohan Menon +1
The paper evaluates five direction‑based inference‑time defenses for vision‑language models across multiple architectures, finding that no single method works best for all models a…
cs.AI2025
Hypothesis Network Planned Exploration for Rapid Meta-Reinforcement Learning Adaptation
Maxwell Joseph Jacobson, Rohan Menon, John Zeng +1
Meta-Reinforcement Learning (Meta-RL) learns optimal policies across a series of related tasks. A central challenge in Meta-RL is rapidly identifying which previously learned task…
cs.LG2025
LipShiFT: A Certifiably Robust Shift-based Vision Transformer
Rohan Menon, Nicola Franco, Stephan Günnemann
Deriving tight Lipschitz bounds for transformer-based architectures presents a significant challenge. The large input sizes and high-dimensional attention modules typically prove t…