Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
Eliciting Intrinsic Hallucinations in LLMs via Semantically Equivalent Adversarial Attacks
Atri Vivek Sharma, Brian Formento, Alessio Lomuscio
Large language models (LLMs) are often used in conjunction with external knowledge sources to improve their factual accuracy and decrease hallucinations, through methods such as Re…
cs.CL2024
SemRoDe: Macro Adversarial Training to Learn Representations That are Robust to Word-Level Attacks
Brian Formento, Wenjie Feng, Chuan Sheng Foo +2
Language models (LMs) are indispensable tools for natural language processing tasks, but their vulnerability to adversarial attacks remains a concern. While current research has ex…