3 citations · 6 across the 5 of their papers we have counts for
6 papers
Training-Free versus Training-Based Intent Classification in LLMs: Accuracy, Robustness, and Failure Modes
Nan Chen, Zhouhao Yang, Soufiane Hayou
Intent classification in Large Language Models (LLMs) involves categorizing user prompts into predefined classes. For instance, given a user prompt, the system must determine wheth…
The Impact of Initialization on LoRA Finetuning Dynamics
Soufiane Hayou, Nikhil Ghosh, Bin Yu
In this paper, we study the role of initialization in Low Rank Adaptation (LoRA) as originally introduced in Hu et al. (2021). Essentially, to start from the pretrained model as in…
How Bad is Training on Synthetic Data? A Statistical Analysis of Language Model Collapse
Mohamed El Amine Seddik, Suei-Wen Chen, Soufiane Hayou +2
The phenomenon of model collapse, introduced in (Shumailov et al., 2023), refers to the deterioration in performance that occurs when new models are trained on synthetic data gener…
Tensor Programs VI: Feature Learning in Infinite-Depth Neural Networks
Greg Yang, Dingli Yu, Chen Zhu +1
By classifying infinite-width neural networks and identifying the *optimal* limit, Tensor Programs IV and V demonstrated a universal way, called P, for *widthwise hyperparameter…
Commutative Width and Depth Scaling in Deep Neural Networks
Soufiane Hayou
This paper is the second in the series Commutative Scaling of Width and Depth (WD) about commutativity of infinite width and depth limits in deep neural networks. Our aim is to und…
On the Connection Between Riemann Hypothesis and a Special Class of Neural Networks
Soufiane Hayou
The Riemann hypothesis (RH) is a long-standing open problem in mathematics. It conjectures that non-trivial zeros of the zeta function all have real part equal to 1/2. The extent o…