3 papers
cs.CL2026
Eliciting Intrinsic Hallucinations in LLMs via Semantically Equivalent Adversarial Attacks
Atri Vivek Sharma, Brian Formento, Alessio Lomuscio
Large language models (LLMs) are often used in conjunction with external knowledge sources to improve their factual accuracy and decrease hallucinations, through methods such as Re…
cs.LG2025
Confidence Elicitation: A New Attack Vector for Large Language Models
Brian Formento, Chuan Sheng Foo, See-Kiong Ng
A fundamental issue in deep learning has been adversarial robustness. As these systems have scaled, such issues have persisted. Currently, large language models (LLMs) with billion…
cs.CL2024
SemRoDe: Macro Adversarial Training to Learn Representations That are Robust to Word-Level Attacks
Brian Formento, Wenjie Feng, Chuan Sheng Foo +2
Language models (LMs) are indispensable tools for natural language processing tasks, but their vulnerability to adversarial attacks remains a concern. While current research has ex…