7 citations · 7 across the 2 of their papers we have counts for
3 papers
cs.LG2025
Closing the Train-Test Gap in World Models for Gradient-Based Planning
Arjun Parthasarathy, Nimit Kalra, Rohun Agrawal +4
World models paired with model predictive control (MPC) can be trained offline on large-scale datasets of expert trajectories and enable generalization to a wide range of planning…
cs.CL2025
Verdict: A Library for Scaling Judge-Time Compute
Nimit Kalra, Leonard Tang
The use of LLMs as automated judges ("LLM-as-a-judge") is now widespread, yet standard judges suffer from a multitude of reliability issues. To address these challenges, we introdu…
cs.CL2025★ 7 cited
Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teaming
Mrinank Sharma, Meg Tong, Jesse Mu +40
Large language models (LLMs) are vulnerable to universal jailbreaks-prompting strategies that systematically bypass model safeguards and enable users to carry out harmful processes…