collaborators

5 papers

cs.CL2026

An expressivity analysis of hierarchical modelling in deep transformers via bounded-depth grammars

Vinoth Nandakumar, Qiang Qu, Pramod Thebe +2

Deep neural networks are widely believed to derive their expressive power from their ability to form \textbf{hierarchical representations}, capturing progressively more abstract an…

cs.LG2026

A theoretical model for task routing in mixture-of-expert transformers

Vinoth Nandakumar, Yongli Xiang, Yunzhi Yao +2

Mixture-of-experts (MoE) layers enable the scaling of transformer models while keeping the inference compute fixed. While task-expert specialization has been observed in empirical…

cs.LG2026

Understanding Diversity Collapse in RLVR via the Lens of Overtraining

Suqin Yuan, Jinkun Chen, Jiyang Zheng +6

Reinforcement learning with verifiable rewards (RLVR) has become a key approach for enhancing the reasoning abilities of large language models. However, RLVR often suffers from \em…

cs.AI2025

Investigating The Functional Roles of Attention Heads in Vision Language Models: Evidence for Reasoning Modules

Yanbei Jiang, Xueqi Ma, Shu Liu +5

Despite excelling on multimodal benchmarks, vision-language models (VLMs) largely remain a black box. In this paper, we propose a novel interpretability framework to systematically…

q-bio.NC2025

Cognitive Mirrors: Exploring the Diverse Functional Roles of Attention Heads in LLM Reasoning

Xueqi Ma, Jun Wang, Yanbei Jiang +3

Large language models (LLMs) have achieved state-of-the-art performance in a variety of tasks, but remain largely opaque in terms of their internal mechanisms. Understanding these…