Showing cs.LGShow all
3 papers · 1 filter
cs.LG2026
pscaling small models: Principled warm starts and hyperparameter transfer
Yuxin Ma, Nan Chen, Mateo DÃaz +3
Modern large-scale neural networks are often trained and released in multiple sizes to accommodate diverse inference budgets. To improve efficiency, recent work has explored model…
cs.LG2026
Learning Rate Scaling across LoRA Ranks and Transfer to Full Finetuning
Nan Chen, Soledad Villar, Soufiane Hayou
Low-Rank Adaptation (LoRA) is a standard tool for parameter-efficient finetuning of large models. While it induces a small memory footprint, its training dynamics can be surprising…
cs.LG2025
Exploring Pseudo-Token Approaches in Transformer Neural Processes
Jose Lara-Rangel, Nanze Chen, Fengzhe Zhang
Neural Processes (NPs) have gained attention in meta-learning for their ability to quantify uncertainty, together with their rapid prediction and adaptability. However, traditional…