2 papers
cs.CL2025
Turning Logic Against Itself : Probing Model Defenses Through Contrastive Questions
Rachneet Sachdeva, Rima Hazra, Iryna Gurevych
Large language models, despite extensive alignment with human values and ethical principles, remain vulnerable to sophisticated jailbreak attacks that exploit their reasoning abili…
cs.CL2025
Localizing and Mitigating Errors in Long-form Question Answering
Rachneet Sachdeva, Yixiao Song, Mohit Iyyer +1
Long-form question answering (LFQA) aims to provide thorough and in-depth answers to complex questions, enhancing comprehension. However, such detailed responses are prone to hallu…