3 papers
cs.CL2026
An expressivity analysis of hierarchical modelling in deep transformers via bounded-depth grammars
Vinoth Nandakumar, Qiang Qu, Pramod Thebe +2
Deep neural networks are widely believed to derive their expressive power from their ability to form \textbf{hierarchical representations}, capturing progressively more abstract an…
cs.LG2026
A theoretical model for task routing in mixture-of-expert transformers
Vinoth Nandakumar, Yongli Xiang, Yunzhi Yao +2
Mixture-of-experts (MoE) layers enable the scaling of transformer models while keeping the inference compute fixed. While task-expert specialization has been observed in empirical…
cs.CL2025
State space models can express n-gram languages
Vinoth Nandakumar, Qiang Qu, Peng Mi +1
Recent advancements in recurrent neural networks (RNNs) have reinvigorated interest in their application to natural language processing tasks, particularly with the development of…