5 papers
An expressivity analysis of hierarchical modelling in deep transformers via bounded-depth grammars
Vinoth Nandakumar, Qiang Qu, Pramod Thebe +2
Deep neural networks are widely believed to derive their expressive power from their ability to form \textbf{hierarchical representations}, capturing progressively more abstract an…
A theoretical model for task routing in mixture-of-expert transformers
Vinoth Nandakumar, Yongli Xiang, Yunzhi Yao +2
Mixture-of-experts (MoE) layers enable the scaling of transformer models while keeping the inference compute fixed. While task-expert specialization has been observed in empirical…
Understanding Diversity Collapse in RLVR via the Lens of Overtraining
Suqin Yuan, Jinkun Chen, Jiyang Zheng +6
Reinforcement learning with verifiable rewards (RLVR) has become a key approach for enhancing the reasoning abilities of large language models. However, RLVR often suffers from \em…
Investigating The Functional Roles of Attention Heads in Vision Language Models: Evidence for Reasoning Modules
Yanbei Jiang, Xueqi Ma, Shu Liu +5
Despite excelling on multimodal benchmarks, vision-language models (VLMs) largely remain a black box. In this paper, we propose a novel interpretability framework to systematically…
Cognitive Mirrors: Exploring the Diverse Functional Roles of Attention Heads in LLM Reasoning
Xueqi Ma, Jun Wang, Yanbei Jiang +3
Large language models (LLMs) have achieved state-of-the-art performance in a variety of tasks, but remain largely opaque in terms of their internal mechanisms. Understanding these…