10 citations · 14 across the 5 of their papers we have counts for
5 papers
Trust or Escalate: LLM Judges with Provable Guarantees for Human Agreement
Jaehun Jung, Faeze Brahman, Yejin Choi
We present a principled approach to provide LLM-based evaluation with a rigorous guarantee of human agreement. We first propose that a reliable evaluation method should not uncriti…
How to Train Your Fact Verifier: Knowledge Transfer with Multimodal Open Models
Jaeyoung Lee, Ximing Lu, Jack Hessel +5
Given the growing influx of misinformation across news and social media, there is a critical need for systems that can provide effective real-time verification of news claims. Larg…
WildTeaming at Scale: From In-the-Wild Jailbreaks to (Adversarially) Safer Language Models
Liwei Jiang, Kavel Rao, Seungju Han +8
We introduce WildTeaming, an automatic LLM safety red-teaming framework that mines in-the-wild user-chatbot interactions to discover 5.7K unique clusters of novel jailbreak tactics…
STEER: Unified Style Transfer with Expert Reinforcement
Skyler Hallinan, Faeze Brahman, Ximing Lu +3
While text style transfer has many applications across natural language processing, the core premise of transferring from a single source style is unrealistic in a real-world setti…
The Generative AI Paradox: "What It Can Create, It May Not Understand"
Peter West, Ximing Lu, Nouha Dziri +11
The recent wave of generative AI has sparked unprecedented global attention, with both excitement and concern over potentially superhuman levels of artificial intelligence: models…