2 papers
cs.LG2025
AbsTopK: Rethinking Sparse Autoencoders For Bidirectional Features
Xudong Zhu, Mohammad Mahdi Khalili, Zhihui Zhu
Sparse autoencoders (SAEs) have emerged as powerful techniques for interpretability of large language models (LLMs), aiming to decompose hidden states into meaningful semantic feat…
cs.LG2025
From Emergence to Control: Probing and Modulating Self-Reflection in Language Models
Xudong Zhu, Jiachen Jiang, Mohammad Mahdi Khalili +1
Self-reflection -- the ability of a large language model (LLM) to revisit, evaluate, and revise its own reasoning -- has recently emerged as a powerful behavior enabled by reinforc…