activity
20242026
collaborators
Showing cs.LGShow all

6 papers · 1 filter

cs.LG2026

DiffuMamba: High-Throughput Diffusion LMs with Mamba Backbone

Vaibhav Singh, Oleksiy Ostapenko, Pierre-André Noël +2

Diffusion language models (DLMs) have emerged as a promising alternative to autoregressive (AR) generation, yet their reliance on Transformer backbones limits inference efficiency…

cs.LG2026

Dual-Phase Continual Learning: Supervised Adaptation Meets Unsupervised Retention

Vaibhav Singh, Rahaf Aljundi, Eugene Belilovsky

Foundational Vision-Language Models (VLMs) excel across diverse tasks, but adapting them to new domains without forgetting prior knowledge remains a critical challenge. Continual L…

cs.LG2026

Heterogeneous Low-Bandwidth Pre-Training of LLMs

Yazan Obeidi, Amir Sarfi, Joel Lidin +2

Pre-training large language models (LLMs) increasingly requires distributed compute, yet bandwidth constraints make it difficult to scale beyond well-provisioned datacenters-especi…

cs.LG2025

When Data Falls Short: Grokking Below the Critical Threshold

Vaibhav Singh, Eugene Belilovsky, Rahaf Aljundi

In this paper, we investigate the phenomenon of grokking, where models exhibit delayed generalization following overfitting on training data. We focus on data-scarce regimes where…

cs.LG2025

Beyond Cosine Decay: On the effectiveness of Infinite Learning Rate Schedule for Continual Pre-training

Vaibhav Singh, Paul Janson, Paria Mehrbod +4

The ever-growing availability of unlabeled data presents both opportunities and challenges for training artificial intelligence systems. While self-supervised learning (SSL) has em…

cs.LG2025

Incentivizing Permissionless Distributed Learning of LLMs

Joel Lidin, Amir Sarfi, Evangelos Pappas +3

We describe an incentive system for distributed deep learning of foundational models where peers are rewarded for contributions. The incentive system, \textit{Gauntlet}, has been d…