4 papers
Domain-Aware Scaling Laws Uncover Data Synergy
Kimia Hamidieh, Lester Mackey, David Alvarez-Melis
Machine learning progress is often attributed to scaling model size and dataset volume, yet the composition of data can be just as consequential. Empirical findings repeatedly show…
Express Language Modeling
Albert Gong, Annabelle Michael Carrell, Raaz Dwivedi +1
We introduce a new tool, Express, for converting a non-causal attention approximation into a causal approximation with matching approximation guarantees. When combined with the sta…
KerJEPA: Kernel Discrepancies for Euclidean Self-Supervised Learning
Eric Zimmermann, Harley Wiltzer, Justin Szeto +2
Recent breakthroughs in self-supervised Joint-Embedding Predictive Architectures (JEPAs) have established that regularizing Euclidean representations toward isotropic Gaussian prio…
Adapting Language Models via Token Translation
Zhili Feng, Tanya Marwah, Nicolo Fusi +2
Modern large language models use a fixed tokenizer to effectively compress text drawn from a source domain. However, applying the same tokenizer to a new target domain often leads…