activity
20242026
collaborators

8 papers

cs.LG2026

Perturbation-based Regional Interpretability through Subtraction Mapping (PRISM): naming-error dissociations in language models and post-stroke aphasia

Xiang Guan, Roger D. Newman-Norlund, Yong Yang +8

Mechanistic interpretability of large language models lacks spatially resolved, falsifiable tools for testing whether internal components are specialized for distinct cognitive ope…

cs.CL2026

RAG-X: Systematic Diagnosis of Retrieval-Augmented Generation for Medical Question Answering

Aswini Sivakumar, Vijayan Sugumaran, Yao Qiang

Automated question-answering (QA) systems increasingly rely on retrieval-augmented generation (RAG) to ground large language models (LLMs) in authoritative medical knowledge, ensur…

cs.LG2025

Not All Tokens Are Meant to Be Forgotten

Xiangyu Zhou, Yao Qiang, Saleh Zare Zade +3

Large Language Models (LLMs), pre-trained on massive text corpora, exhibit remarkable human-level language understanding, reasoning, and decision-making abilities. However, they te…

cs.LG2025

Learning to Poison Large Language Models for Downstream Manipulation

Xiangyu Zhou, Yao Qiang, Saleh Zare Zade +4

The advent of Large Language Models (LLMs) has marked significant achievements in language processing and reasoning capabilities. Despite their advancements, LLMs face vulnerabilit…

cs.LG2025

Hijacking Large Language Models via Adversarial In-Context Learning

Xiangyu Zhou, Yao Qiang, Saleh Zare Zade +2

In-context learning (ICL) has emerged as a powerful paradigm leveraging LLMs for specific downstream tasks by utilizing labeled examples as demonstrations (demos) in the preconditi…

cs.LG2025

Automatic Calibration for Membership Inference Attack on Large Language Models

Saleh Zare Zade, Yao Qiang, Xiangyu Zhou +4

Membership Inference Attacks (MIAs) have recently been employed to determine whether a specific text was part of the pre-training data of Large Language Models (LLMs). However, exi…