collaborators

6 papers

cs.AI2026

From Insight to Action: A Novel Framework for Interpretability-Guided Data Selection in Large Language Models

Ling Shi, Xinwei Wu, Xiaohu Zhao +7

While mechanistic interpretability tools like Sparse Autoencoders (SAEs) can uncover meaningful features within Large Language Models (LLMs), a critical gap remains in transforming…

cs.CL2026

TaP: A Taxonomy-Guided Framework for Automated and Scalable Preference Data Generation

Renren Jin, Tianhao Shen, Xinwei Wu +9

Conducting supervised and preference fine-tuning of large language models (LLMs) requires high-quality datasets to improve their ability to follow instructions and align with human…

cs.CL2026

Finding the Translation Switch: Discovering and Exploiting the Task-Initiation Features in LLMs

Xinwei Wu, Heng Liu, Xiaohu Zhao +6

Large Language Models (LLMs) frequently exhibit strong translation abilities, even without task-specific fine-tuning. However, the internal mechanisms governing this innate capabil…

cs.CL2024

ConTrans: Weak-to-Strong Alignment Engineering via Concept Transplantation

Weilong Dong, Xinwei Wu, Renren Jin +2

Ensuring large language models (LLM) behave consistently with human goals, values, and intentions is crucial for their safety but yet computationally expensive. To reduce the compu…

cs.AI2024

Large Language Model Safety: A Holistic Survey

Dan Shi, Tianhao Shen, Yufei Huang +10

The rapid development and deployment of large language models (LLMs) have introduced a new frontier in artificial intelligence, marked by unprecedented capabilities in natural lang…

cs.CL2024

IRCAN: Mitigating Knowledge Conflicts in LLM Generation via Identifying and Reweighting Context-Aware Neurons

Dan Shi, Renren Jin, Tianhao Shen +3

It is widely acknowledged that large language models (LLMs) encode a vast reservoir of knowledge after being trained on mass data. Recent studies disclose knowledge conflicts in LL…