Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
How Context Attribution Handles What the Model Already Knows
Quoc-Huy Trinh, Lin Zhu, Sebastian Szyller
Context attribution methods for large language models (LLMs) identify which input context contributes to the model response. Recent works show the initial success in attributing th…
cs.CL2025
Soft Token Attacks Cannot Reliably Audit Unlearning in Large Language Models
Haokun Chen, Sebastian Szyller, Weilin Xu +1
Large language models (LLMs) are trained using massive datasets, which often contain undesirable content such as harmful texts, personal information, and copyrighted material. To a…
cs.CL2024
LLM Self Defense: By Self Examination, LLMs Know They Are Being Tricked
Mansi Phute, Alec Helbling, Matthew Hull +4
Large language models (LLMs) are popular for high-quality text generation but can produce harmful content, even when aligned with human values through reinforcement learning. Adver…