most citedInterpreting and Steering LLMs with Mutual Information-based Explanations on Sparse Autoencoders

1 citations · 1 across the 2 of their papers we have counts for

collaborators

5 papers

cs.CL2025

AutoL2S: Auto Long-Short Reasoning for Efficient Large Language Models

Feng Luo, Yu-Neng Chuang, Guanchu Wang +8

Reasoning-capable large language models (LLMs) achieve strong performance on complex tasks but often exhibit overthinking after distillation, generating unnecessarily long chain-of…

cs.CL2025

Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models

Yang Sui, Yu-Neng Chuang, Guanchu Wang +9

Large Language Models (LLMs) have demonstrated remarkable capabilities in complex tasks. Recent advancements in Large Reasoning Models (LRMs), such as OpenAI o1 and DeepSeek-R1, ha…

cs.CL20251 cited

Interpreting and Steering LLMs with Mutual Information-based Explanations on Sparse Autoencoders

Xuansheng Wu, Jiayi Yuan, Wenlin Yao +2

Large language models (LLMs) excel at handling human queries, but they can occasionally generate flawed or unexpected responses. Understanding their internal states is crucial for…

cs.CL2025

The Science of Evaluating Foundation Models

Jiayi Yuan, Jiamu Zhang, Andrew Wen +1

The emergent phenomena of large foundation models have revolutionized natural language processing. However, evaluating these models presents significant challenges due to their siz…

cs.CR2024

Taylor Unswift: Secured Weight Release for Large Language Models via Taylor Expansion

Guanchu Wang, Yu-Neng Chuang, Ruixiang Tang +8

Ensuring the security of released large language models (LLMs) poses a significant dilemma, as existing mechanisms either compromise ownership rights or raise data privacy concerns…