63 citations · 79 across the 19 of their papers we have counts for
Showing cs.AIShow all
3 papers · 1 filter
cs.AI2026
From Insight to Action: A Novel Framework for Interpretability-Guided Data Selection in Large Language Models
Ling Shi, Xinwei Wu, Xiaohu Zhao +7
While mechanistic interpretability tools like Sparse Autoencoders (SAEs) can uncover meaningful features within Large Language Models (LLMs), a critical gap remains in transforming…
cs.AI2024★ 3 cited
Large Language Model Safety: A Holistic Survey
Dan Shi, Tianhao Shen, Yufei Huang +10
The rapid development and deployment of large language models (LLMs) have introduced a new frontier in artificial intelligence, marked by unprecedented capabilities in natural lang…
cs.AI2023
Is Robustness Transferable across Languages in Multilingual Neural Machine Translation?
Leiyu Pan, Supryadi, Deyi Xiong
Robustness, the ability of models to maintain performance in the face of perturbations, is critical for developing reliable NLP systems. Recent studies have shown promising results…