3 papers
cs.CL2026
Where to Begin: Efficient Pretraining via Subnetwork Selection and Distillation
Arjun Krishnakumar, Rhea Sanjay Sukthanker, Hannan Javed Mahadik +5
Small Language models (SLMs) offer an efficient and accessible alternative to Large Language Models (LLMs), delivering strong performance while using far fewer resources. We introd…
cs.LG2025
Transferrable Surrogates in Expressive Neural Architecture Search Spaces
Shiwen Qin, Gabriela Kadlecová, Martin Pilát +5
Neural architecture search (NAS) faces a challenge in balancing the exploration of expressive, broad search spaces that enable architectural innovation with the need for efficient…
cs.LG2024
Surprisingly Strong Performance Prediction with Neural Graph Features
Gabriela Kadlecová, Jovita Lukasik, Martin Pilát +4
Performance prediction has been a key part of the neural architecture search (NAS) process, allowing to speed up NAS algorithms by avoiding resource-consuming network training. Alt…