4 papers
YRC-Bench: A Benchmark for Learning to Coordinate with Experts
Mohamad H. Danesh, Nguyen X. Khanh, Tu Trinh +1
When deployed in the real world, AI agents will inevitably face challenges that exceed their individual capabilities. A critical component of AI safety is an agent's ability to rec…
Getting By Goal Misgeneralization With a Little Help From a Mentor
Tu Trinh, Mohamad H. Danesh, Nguyen X. Khanh +1
While reinforcement learning (RL) agents often perform well during training, they can struggle with distribution shift in real-world deployments. One particularly severe risk of di…
Probabilities of Chat LLMs Are Miscalibrated but Still Predict Correctness on Multiple-Choice Q&A
Benjamin Plaut, Nguyen X. Khanh, Tu Trinh
We study 15 large language models (LLMs) fine-tuned for chat and find that their maximum softmax probabilities (MSPs) are consistently miscalibrated on multiple-choice Q&A. However…
A StrongREJECT for Empty Jailbreaks
Alexandra Souly, Qingyuan Lu, Dillon Bowen +8
Most jailbreak papers claim the jailbreaks they propose are highly effective, often boasting near-100% attack success rates. However, it is perhaps more common than not for jailbre…