collaborators
Showing cs.CLShow all

5 papers · 1 filter

cs.CL2025

SAGE: An Agentic Explainer Framework for Interpreting SAE Features in Language Models

Jiaojiao Han, Wujiang Xu, Mingyu Jin +1

Large language models (LLMs) have achieved remarkable progress, yet their internal mechanisms remain largely opaque, posing a significant challenge to their safe and reliable deplo…

cs.CL2025

Denoising Concept Vectors with Sparse Autoencoders for Improved Language Model Steering

Haiyan Zhao, Xuansheng Wu, Fan Yang +3

Linear concept vectors effectively steer LLMs, but existing methods suffer from noisy features in diverse datasets that undermine steering robustness. We propose Sparse Autoencoder…

cs.CL2025

Massive Values in Self-Attention Modules are the Key to Contextual Knowledge Understanding

Mingyu Jin, Kai Mei, Wujiang Xu +5

Large language models (LLMs) have achieved remarkable success in contextual knowledge understanding. In this paper, we show that these concentrated massive values consistently emer…

cs.CL2024

Data-centric NLP Backdoor Defense from the Lens of Memorization

Zhenting Wang, Zhizhi Wang, Mingyu Jin +3

Backdoor attack is a severe threat to the trustworthiness of DNN-based language models. In this paper, we first extend the definition of memorization of language models from sample…

cs.CL2024

Exploring Multilingual Probing in Large Language Models: A Cross-Language Analysis

Daoyang Li, Haiyan Zhao, Qingcheng Zeng +1

Probing techniques for large language models (LLMs) have primarily focused on English, overlooking the vast majority of the world's languages. In this paper, we extend these probin…