5 papers · 1 filter
Learning through Internalization
Nikolaos Tsilivis, Nirmit Joshi, Marko Medvedev +2
We study internalization processes, by which neural-network-based systems absorb an explicit computational procedure into their own weights, and how they facilitate learning. We in…
Positive Distribution Shift as a Framework for Understanding Tractable Learning
Marko Medvedev, Idan Attias, Elisabetta Cornacchia +3
We study a setting where the goal is to learn a target function f(x) with respect to a target distribution D(x), but training is done on i.i.d. samples from a different training di…
Shift is Good: Mismatched Data Mixing Improves Test Performance
Marko Medvedev, Kaifeng Lyu, Zhiyuan Li +1
We consider training and testing on mixture distributions with different training and test proportions. We show that in many settings, and in some sense generically, distribution s…
Weak-to-Strong Generalization Even in Random Feature Networks, Provably
Marko Medvedev, Kaifeng Lyu, Dingli Yu +3
Weak-to-Strong Generalization (Burns et al., 2024) is the phenomenon whereby a strong student, say GPT-4, learns a task from a weak teacher, say GPT-2, and ends up significantly ou…
Overfitting Behaviour of Gaussian Kernel Ridgeless Regression: Varying Bandwidth or Dimensionality
Marko Medvedev, Gal Vardi, Nathan Srebro
We consider the overfitting behavior of minimum norm interpolating solutions of Gaussian kernel ridge regression (i.e. kernel ridgeless regression), when the bandwidth or input dim…