Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
From Insight to Action: A Novel Framework for Interpretability-Guided Data Selection in Large Language Models
Ling Shi, Xinwei Wu, Xiaohu Zhao +7
While mechanistic interpretability tools like Sparse Autoencoders (SAEs) can uncover meaningful features within Large Language Models (LLMs), a critical gap remains in transforming…
cs.AI2024
Large Language Model Safety: A Holistic Survey
Dan Shi, Tianhao Shen, Yufei Huang +10
The rapid development and deployment of large language models (LLMs) have introduced a new frontier in artificial intelligence, marked by unprecedented capabilities in natural lang…