Showing cs.LGShow all
3 papers · 1 filter
cs.LG2026
Pretraining Curricula Enable Selective Fine-tuning
Sebastian A. Bruijns, Jirko Rubruck, Mia H. Whitefield +3
Transformers follow implicit curricula whereby some tasks are learned before others. However, how explicit pretraining curricula influence learning, generalization, and the selecti…
cs.LG2025
Flexible task abstractions emerge in linear networks with fast and bounded units
Kai Sandbrink, Jan P. Bauer, Alexandra M. Proca +3
Animals survive in dynamic environments changing at arbitrary timescales, but such data distribution shifts are a challenge to neural networks. To adapt to change, neural systems m…
cs.LG2024
Early learning of the optimal constant solution in neural networks and humans
Jirko Rubruck, Jan P. Bauer, Andrew Saxe +1
Deep neural networks learn increasingly complex functions over the course of training. Here, we show both empirically and theoretically that learning of the target function is prec…