6 papers
From Insight to Action: A Novel Framework for Interpretability-Guided Data Selection in Large Language Models
Ling Shi, Xinwei Wu, Xiaohu Zhao +7
While mechanistic interpretability tools like Sparse Autoencoders (SAEs) can uncover meaningful features within Large Language Models (LLMs), a critical gap remains in transforming…
TaP: A Taxonomy-Guided Framework for Automated and Scalable Preference Data Generation
Renren Jin, Tianhao Shen, Xinwei Wu +9
Conducting supervised and preference fine-tuning of large language models (LLMs) requires high-quality datasets to improve their ability to follow instructions and align with human…
Finding the Translation Switch: Discovering and Exploiting the Task-Initiation Features in LLMs
Xinwei Wu, Heng Liu, Xiaohu Zhao +6
Large Language Models (LLMs) frequently exhibit strong translation abilities, even without task-specific fine-tuning. However, the internal mechanisms governing this innate capabil…
ConTrans: Weak-to-Strong Alignment Engineering via Concept Transplantation
Weilong Dong, Xinwei Wu, Renren Jin +2
Ensuring large language models (LLM) behave consistently with human goals, values, and intentions is crucial for their safety but yet computationally expensive. To reduce the compu…
Large Language Model Safety: A Holistic Survey
Dan Shi, Tianhao Shen, Yufei Huang +10
The rapid development and deployment of large language models (LLMs) have introduced a new frontier in artificial intelligence, marked by unprecedented capabilities in natural lang…
IRCAN: Mitigating Knowledge Conflicts in LLM Generation via Identifying and Reweighting Context-Aware Neurons
Dan Shi, Renren Jin, Tianhao Shen +3
It is widely acknowledged that large language models (LLMs) encode a vast reservoir of knowledge after being trained on mass data. Recent studies disclose knowledge conflicts in LL…