8 papers
Perturbation-based Regional Interpretability through Subtraction Mapping (PRISM): naming-error dissociations in language models and post-stroke aphasia
Xiang Guan, Roger D. Newman-Norlund, Yong Yang +8
Mechanistic interpretability of large language models lacks spatially resolved, falsifiable tools for testing whether internal components are specialized for distinct cognitive ope…
RAG-X: Systematic Diagnosis of Retrieval-Augmented Generation for Medical Question Answering
Aswini Sivakumar, Vijayan Sugumaran, Yao Qiang
Automated question-answering (QA) systems increasingly rely on retrieval-augmented generation (RAG) to ground large language models (LLMs) in authoritative medical knowledge, ensur…
Not All Tokens Are Meant to Be Forgotten
Xiangyu Zhou, Yao Qiang, Saleh Zare Zade +3
Large Language Models (LLMs), pre-trained on massive text corpora, exhibit remarkable human-level language understanding, reasoning, and decision-making abilities. However, they te…
Learning to Poison Large Language Models for Downstream Manipulation
Xiangyu Zhou, Yao Qiang, Saleh Zare Zade +4
The advent of Large Language Models (LLMs) has marked significant achievements in language processing and reasoning capabilities. Despite their advancements, LLMs face vulnerabilit…
Hijacking Large Language Models via Adversarial In-Context Learning
Xiangyu Zhou, Yao Qiang, Saleh Zare Zade +2
In-context learning (ICL) has emerged as a powerful paradigm leveraging LLMs for specific downstream tasks by utilizing labeled examples as demonstrations (demos) in the preconditi…
Automatic Calibration for Membership Inference Attack on Large Language Models
Saleh Zare Zade, Yao Qiang, Xiangyu Zhou +4
Membership Inference Attacks (MIAs) have recently been employed to determine whether a specific text was part of the pre-training data of Large Language Models (LLMs). However, exi…