collaborators

5 papers

cs.CV2026

Visual prompt engineering for video models

Robert Geirhos, Yuxuan Li, Thaddäus Wiedemer +7

In the age of foundation models, a model is only as good as its prompt. For this reason, prompt engineering has become an essential technique for improving language model performan…

cs.CL2026

The ACUTE Protocol: Operationalizing Language Model Activations for Better Calibration, Utility, and Trust

Nishant Subramani, Palash Goyal, Yiwen Song +4

As language models improve and become increasingly deployed to solve a variety of tasks, trustworthiness becomes essential. Calibration is a good proxy for trust: well-calibrated c…

cs.AI2026

Hiding in Plain Text: Detecting Concealed Jailbreaks via Activation Disentanglement

Amirhossein Farzam, Majid Behabahani, Mani Malek +2

Large language models (LLMs) remain vulnerable to jailbreak prompts that are fluent and semantically coherent, and therefore difficult to detect with standard heuristics. A particu…

cs.LG2026

Interpreting and Controlling Model Behavior via Constitutions for Atomic Concept Edits

Neha Kalibhat, Zi Wang, Prasoon Bajpai +4

We introduce a black-box interpretability framework that learns a verifiable constitution: a natural language summary of how changes to a prompt affect a model's specific behavior,…

cs.CV2025

ShieldGemma 2: Robust and Tractable Image Content Moderation

Wenjun Zeng, Dana Kurniawan, Ryan Mullins +14

We introduce ShieldGemma 2, a 4B parameter image content moderation model built on Gemma 3. This model provides robust safety risk predictions across the following key harm categor…