4 papers
MJ1: Multimodal Judgment via Grounded Verification
Bhavesh Kumar, Dylan Feng, Leonard Tang
Multimodal judges struggle to ground decisions in visual evidence. We present MJ1, a multimodal judge trained with reinforcement learning that enforces visual grounding through a s…
Verdict: A Library for Scaling Judge-Time Compute
Nimit Kalra, Leonard Tang
The use of LLMs as automated judges ("LLM-as-a-judge") is now widespread, yet standard judges suffer from a multitude of reliability issues. To address these challenges, we introdu…
Endless Jailbreaks with Bijection Learning
Brian R. Y. Huang, Maximilian Li, Leonard Tang
Despite extensive safety measures, LLMs are vulnerable to adversarial inputs, or jailbreaks, which can elicit unsafe behaviors. In this work, we introduce bijection learning, a pow…
Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teaming
Mrinank Sharma, Meg Tong, Jesse Mu +40
Large language models (LLMs) are vulnerable to universal jailbreaks-prompting strategies that systematically bypass model safeguards and enable users to carry out harmful processes…