5 papers
An Effective Gram Matrix Characterizes Generalization in Deep Networks
Rubing Yang, Pratik Chaudhari
We derive a differential equation that governs the evolution of the generalization gap when a deep network is trained by gradient descent. This differential equation is controlled…
From Linearity to Non-Linearity: How Masked Autoencoders Capture Spatial Correlations
Anthony Bisulco, Rahul Ramesh, Randall Balestriero +1
Masked Autoencoders (MAEs) have emerged as a powerful pretraining technique for vision foundation models. Despite their effectiveness, they require extensive hyperparameter tuning…
Prospective Learning in Retrospect
Yuxin Bai, Cecelia Shuai, Ashwin De Silva +3
In most real-world applications of artificial intelligence, the distributions of the data and the goals of the learners tend to change over time. The Probably Approximately Correct…
Many Perception Tasks are Highly Redundant Functions of their Input Data
Rahul Ramesh, Anthony Bisulco, Ronald W. DiTullio +4
We show that many perception tasks, from visual recognition, semantic segmentation, optical flow, depth estimation to vocalization discrimination, are highly redundant functions of…
Prospective Learning: Learning for a Dynamic Future
Ashwin De Silva, Rahul Ramesh, Rubing Yang +3
In real-world applications, the distribution of the data, and our goals, evolve over time. The prevailing theoretical framework for studying machine learning, namely probably appro…