63 citations · 170 across the 26 of their papers we have counts for
30 papers · 1 filter
When is Warmstarting Effective for Scaling Language Models?
Neeratyoy Mallik, Maciej Janowski, Johannes Hog +4
Model growth from a given checkpoint aims to accelerate training of a larger model, offering potential resource savings. Despite recent interest, warmstarting has seen limited prac…
Learning to Order: Task Sequencing as In-Context Optimization
Jan Kobiolka, Christian Frey, Arlind Kadra +2
Task sequencing (TS) is one of the core open problems in Deep Learning, arising in a plethora of real-world domains, from robotic assembly lines to autonomous driving. Unfortunatel…
POP: Prior-Fitted First-Order Optimization Policies
Jan Kobiolka, Christian Frey, Gresa Shala +3
Gradient-based optimizers are highly sensitive to design choices in their adaptive learning rate mechanisms. To address this limitation, we introduce POP, a meta-learned Reinforcem…
End-to-End Compression for Tabular Foundation Models
Guri Zabërgja, Rafiq Kamel, Arlind Kadra +2
The long-standing dominance of gradient-boosted decision trees for tabular data has recently been challenged by in-context learning tabular foundation models. In-context learning m…
Warmstarting for Scaling Language Models
Neeratyoy Mallik, Maciej Janowski, Johannes Hog +4
Scaling model sizes to scale performance has worked remarkably well for the current large language models paradigm. The research and empirical findings of various scaling studies l…
Regularized Neural Ensemblers
Sebastian Pineda Arango, Maciej Janowski, Lennart Purucker +3
Ensemble methods are known for enhancing the accuracy and robustness of machine learning models by combining multiple base learners. However, standard approaches like greedy or ran…