3 papers
cs.LG2026
Heterogeneous Low-Bandwidth Pre-Training of LLMs
Yazan Obeidi, Amir Sarfi, Joel Lidin +2
Pre-training large language models (LLMs) increasingly requires distributed compute, yet bandwidth constraints make it difficult to scale beyond well-provisioned datacenters-especi…
cs.LG2025
Beyond Cosine Decay: On the effectiveness of Infinite Learning Rate Schedule for Continual Pre-training
Vaibhav Singh, Paul Janson, Paria Mehrbod +4
The ever-growing availability of unlabeled data presents both opportunities and challenges for training artificial intelligence systems. While self-supervised learning (SSL) has em…
cs.CV2024
Towards motion from video diffusion models
Paul Janson, Tiberiu Popa, Eugene Belilovsky
Text-conditioned video diffusion models have emerged as a powerful tool in the realm of video generation and editing. But their ability to capture the nuances of human movement rem…