6 papers
Locating and Controlling Implicit Personalization in Large Language Models
Yueru Yan, Siqi Wu, Thai Le
Large language models (LLMs) often shift their outputs in response to implicit demographic cues even when users never state a demographic identity. Previous work has documented thi…
ATLAS: Verifier-Guided Adaptive Latent Activation Steering for Efficient LLM Reasoning
Tuc Nguyen, Thai Le
Recent work on activation and latent steering has demonstrated that modifying internal representations can effectively guide large language models (LLMs) toward improved reasoning…
PreUnlearn: Auditing Collateral Knowledge Damage Before Large Language Model Unlearning
Bo Su, Ankit Shah, Thai Le
Machine unlearning for large language models (LLMs) aims to remove specified knowledge while preserving the rest of the model's capabilities. However, the boundary between knowledg…
Beyond Linear Activation Steering: Invertible Latent Transformations for Controlling LLM Behavior
Tuc Nguyen, Thai Le
Activation steering provides a lightweight inference-time mechanism for controlling large language models (LLMs) by modifying their internal activation vectors toward desired behav…
ShareChat: A Dataset of Chatbot Conversations in the Wild
Yueru Yan, Tuc Nguyen, Bo Su +2
By evaluating Large Language Models (LLMs) through uniform, text-only interfaces, current academic benchmarks obscure how the unique designs and affordances of distinct commercial…
Unraveling Interwoven Roles of Large Language Models in Authorship Privacy: Obfuscation, Mimicking, and Verification
Tuc Nguyen, Yifan Hu, Thai Le
Recent advancements in large language models (LLMs) have been fueled by large scale training corpora drawn from diverse sources such as websites, news articles, and books. These da…