3 papers
cs.CL2026
An expressivity analysis of hierarchical modelling in deep transformers via bounded-depth grammars
Vinoth Nandakumar, Qiang Qu, Pramod Thebe +2
Deep neural networks are widely believed to derive their expressive power from their ability to form \textbf{hierarchical representations}, capturing progressively more abstract an…
cs.LG2026
A theoretical model for task routing in mixture-of-expert transformers
Vinoth Nandakumar, Yongli Xiang, Yunzhi Yao +2
Mixture-of-experts (MoE) layers enable the scaling of transformer models while keeping the inference compute fixed. While task-expert specialization has been observed in empirical…
cs.CV2023
Why do CNNs excel at feature extraction? A mathematical explanation
Vinoth Nandakumar, Arush Tagade, Tongliang Liu
Over the past decade deep learning has revolutionized the field of computer vision, with convolutional neural network models proving to be very effective for image classification b…