4 papers
Measuring Pragmatic Influence in Large Language Model Instructions
Yilin Geng, Omri Abend, Eduard Hovy +1
It is not only what we ask large language models (LLMs) to do that matters, but also how we prompt. Phrases like "This is urgent" or "As your supervisor" can shift model behavior w…
Control Illusion: The Failure of Instruction Hierarchies in Large Language Models
Yilin Geng, Haonan Li, Honglin Mu +5
Large language models (LLMs) are increasingly deployed with hierarchical instruction schemes, where certain instructions (e.g., system-level directives) are expected to take preced…
Community Moderation and the New Epistemology of Fact Checking on Social Media
Isabelle Augenstein, Michiel Bakker, Tanmoy Chakraborty +13
Social media platforms have traditionally relied on internal moderation teams and partnerships with independent fact-checking organizations to identify and flag misleading content.…
Libra-Leaderboard: Towards Responsible AI through a Balanced Leaderboard of Safety and Capability
Haonan Li, Xudong Han, Zenan Zhai +32
To address this gap, we introduce Libra-Leaderboard, a comprehensive framework designed to rank LLMs through a balanced evaluation of performance and safety. Combining a dynamic le…