9 papers
AdaJudge: Adaptive Multi-Perspective Judging for Reward Modeling
Yongliang Miao, Yangyang Liang, Mengnan Du
Reward modeling is essential for aligning large language models with human preferences, yet predominant architectures rely on a static pooling strategy to condense sequences into s…
AdaptiveK: Complexity-Driven Sparse Autoencoders for Interpretable Language Model Representations
Yifei Yao, Hanrong Zhang, Mengnan Du
Understanding the internal representations of large language models (LLMs) remains a central challenge for interpretability research. Sparse autoencoders (SAEs) offer a promising s…
A Unified Theory of Sparse Dictionary Learning in Mechanistic Interpretability: Piecewise Biconvexity and Spurious Minima
Yiming Tang, Harshvardhan Saini, Zhaoqian Yao +6
As AI models achieve remarkable capabilities across diverse domains, understanding what representations they learn and how they encode concepts has become increasingly important fo…
DeepSieve: Information Sieving via LLM-as-a-Knowledge-Router
Minghao Guo, Qingcheng Zeng, Xujiang Zhao +5
Large Language Models (LLMs) excel at many reasoning tasks but struggle with knowledge-intensive queries due to their inability to dynamically access up-to-date or domain-specific…
LLM Agents in Law: Taxonomy, Applications, and Challenges
Shuang Liu, Ruijia Zhang, Ruoyun Ma +6
Large language models (LLMs) have precipitated a dramatic improvement in the legal domain, yet the deployment of standalone models faces significant limitations regarding hallucina…
NeuronScope: A Multi-Agent Framework for Explaining Polysemantic Neurons in Language Models
Weiqi Liu, Yongliang Miao, Haiyan Zhao +2
Neuron-level interpretation in large language models (LLMs) is fundamentally challenged by widespread polysemanticity, where individual neurons respond to multiple distinct semanti…