4 papers · 1 filter
JudgeSense: A Benchmark for Prompt Sensitivity in LLM-as-a-Judge Systems
Rohith Reddy Bellibatlu, Edward Raff, Wenbin Zhang
Large language models are widely adopted as automated evaluation judges, yet the stability of their verdicts under semantically equivalent prompt rephrasings remains largely unexam…
Human-Readable Adversarial Prompts: An Investigation into LLM Vulnerabilities Using Situational Context
Nilanjana Das, Edward Raff, Aman Chadha +1
As the AI systems become deeply embedded in social media platforms, we've uncovered a concerning security vulnerability that goes beyond traditional adversarial attacks. It becomes…
Attribution in Scientific Literature: New Benchmark and Methods
Yash Saxena, Deepa Tilwani, Ali Mohammadi +4
Large language models (LLMs) present a promising yet challenging frontier for automated source citation in scientific communication. Previous approaches to citation generation have…
Human-Interpretable Adversarial Prompt Attack on Large Language Models with Situational Context
Nilanjana Das, Edward Raff, Manas Gaur
Previous research on testing the vulnerabilities in Large Language Models (LLMs) using adversarial attacks has primarily focused on nonsensical prompt injections, which are easily…