5 papers
An Overview of Low-Rank Structures in the Training and Adaptation of Large Models
Laura Balzano, Tianjiao Ding, Benjamin D. Haeffele +5
The substantial computational demands of modern large-scale deep learning present significant challenges for efficient training and deployment. Recent research has revealed a wides…
Data Distribution as a Lever for Guiding Optimizers Toward Superior Generalization in LLMs
Tushaar Gangavarapu, Jiping Li, Christopher Vattheuer +2
Can modifying the training data distribution guide optimizers toward solutions with improved generalization when training large language models (LLMs)? In this work, we theoretical…
Neon: Negative Extrapolation From Self-Training Improves Image Generation
Sina Alemohammad, Zhangyang Wang, Richard G. Baraniuk
Scaling generative AI models is bottlenecked by the scarcity of high-quality training data. The ease of synthesizing from a generative model suggests using (unverified) synthetic d…
Time-MQA: Time Series Multi-Task Question Answering with Context Enhancement
Yaxuan Kong, Yiyuan Yang, Yoontae Hwang +5
Time series data are foundational in finance, healthcare, and energy domains. However, most existing methods and datasets remain focused on a narrow spectrum of tasks, such as fore…
On How Iterative Magnitude Pruning Discovers Local Receptive Fields in Fully Connected Neural Networks
William T. Redman, Zhangyang Wang, Alessandro Ingrosso +1
Since its use in the Lottery Ticket Hypothesis, iterative magnitude pruning (IMP) has become a popular method for extracting sparse subnetworks that can be trained to high performa…