4 papers · 1 filter
ESI: Epistemic Uncertainty Quantification via Semantic-preserving Intervention for Large Language Models
Mingda Li, Xinyu Li, Weinan Zhang +1
Uncertainty Quantification (UQ) is a promising approach to improve model reliability, yet quantifying the uncertainty of Large Language Models (LLMs) is non-trivial. In this work,…
Breaking the Block: Preserving Data Continuity to Train Superior SAEs for Instruct Models
Jiaming Li, Haoran Ye, Yukun Chen +5
Sparse Autoencoders (SAEs) are a cornerstone of mechanistic interpretability. Existing training methods inherit the Block Training paradigm from LLM pre-training, which introduces…
Personalized Language Modeling from Personalized Human Feedback
Xinyu Li, Ruiyang Zhou, Zachary C. Lipton +1
Personalized large language models (LLMs) are designed to tailor responses to individual user preferences. While Reinforcement Learning from Human Feedback (RLHF) is a commonly use…
Large Language Models for Automatic Detection of Sensitive Topics
Ruoyu Wen, Stephanie Elena Crowe, Kunal Gupta +6
Sensitive information detection is crucial in content moderation to maintain safe online communities. Assisting in this traditionally manual process could relieve human moderators…