Showing 2026Show all
2 papers · 1 filter
cs.LG2026
Domain-Aware Scaling Laws Uncover Data Synergy
Kimia Hamidieh, Lester Mackey, David Alvarez-Melis
Machine learning progress is often attributed to scaling model size and dataset volume, yet the composition of data can be just as consequential. Empirical findings repeatedly show…
cs.LG2026
Express Language Modeling
Albert Gong, Annabelle Michael Carrell, Raaz Dwivedi +1
We introduce a new tool, Express, for converting a non-causal attention approximation into a causal approximation with matching approximation guarantees. When combined with the sta…