2 papers
cs.LG2024
On Generalization Bounds for Neural Networks with Low Rank Layers
Andrea Pinto, Akshay Rangamani, Tomaso Poggio
While previous optimization results have suggested that deep neural networks tend to favour low-rank weight matrices, the implications of this inductive bias on generalization boun…
cs.CL2024
The Fair Language Model Paradox
Andrea Pinto, Tomer Galanti, Randall Balestriero
Large Language Models (LLMs) are widely deployed in real-world applications, yet little is known about their training dynamics at the token level. Evaluation typically relies on ag…