activity
20242026
collaborators

12 papers

cs.AI2026

Leave it to the Specialist: Repair Sparse LLMs with Sparse Fine-Tuning via Sparsity Evolution

Qiao Xiao, Alan Ansell, Boqian Wu +4

Sparse large language models (LLMs) offer an attractive direction toward efficient deployment, but adapting them to downstream tasks remains challenging. The central difficulty is…

cs.LG2026

When Data Is Scarce: Scaling Sparse Language Models with Repeated Training

Boqian Wu, Qiao Xiao, Patrik Okanovic +6

Scaling laws for dense LLMs under infinite data are well explored, but how sparsity interacts with limited data is not. In this work, we study sparse training in data-constrained r…

cs.LG2026

Memory-Efficient LLM Training with Dynamic Sparsity: From Stability to Practical Scaling

Qiao Xiao, Boqian Wu, Patrik Okanovic +6

Dynamic Sparse Training (DST) offers a promising paradigm for improving the training and inference efficiency of deep neural networks; however, we find that in large language model…

cs.LG2025

Batch Matrix-form Equations and Implementation of Multilayer Perceptrons

Wieger Wesselink, Bram Grooten, Huub van de Wetering +2

Multilayer perceptrons (MLPs) remain fundamental to modern deep learning, yet their algorithmic details are rarely presented in complete, explicit \emph{batch matrix-form}. Rather,…

cs.LG2025

Addressing the Collaboration Dilemma in Low-Data Federated Learning via Transient Sparsity

Qiao Xiao, Boqian Wu, Andrey Poddubnyy +4

Federated learning (FL) enables collaborative model training across decentralized clients while preserving data privacy, leveraging aggregated updates to build robust global models…

cs.LG2025

NeuroTrails: Training with Dynamic Sparse Heads as the Key to Effective Ensembling

Bram Grooten, Farid Hasanov, Chenxiang Zhang +9

Model ensembles have long been a cornerstone for improving generalization and robustness in deep learning. However, their effectiveness often comes at the cost of substantial compu…