3 papers
cs.LG2026
LLM Pretraining Shapes a Generalizable Manifold: Insights into Cross-Modal Transfer to Time Series
Alexis Roger, Prateek Humane, Zhenghan Tai +4
Can language-pretrained transformers become effective time-series forecasters, and why? In this paper, we show that cross-modal transfer arises because language pretraining precond…
cs.LG2025
Small Vocabularies, Big Gains: Pretraining and Tokenization in Time Series Models
Alexis Roger, Gwen Legate, Kashif Rasul +2
Tokenization and transfer learning are two critical components in building state of the art time series foundation models for forecasting. In this work, we systematically study the…
cs.LG2025
Warming Up for Zeroth-Order Federated Pre-Training with Low Resource Clients
Gwen Legate, Irina Rish, Eugene Belilovsky
Federated learning enables collaborative model training across numerous edge devices without requiring participants to share data; however, memory and communication constraints on…