Showing cs.LGShow all
3 papers · 1 filter
cs.LG2025
Influence Functions for Efficient Data Selection in Reasoning
Prateek Humane, Paolo Cudrano, Daniel Z. Kaplan +3
Fine-tuning large language models (LLMs) on chain-of-thought (CoT) data shows that a small amount of high-quality data can outperform massive datasets. Yet, what constitutes "quali…
cs.LG2025
Small Vocabularies, Big Gains: Pretraining and Tokenization in Time Series Models
Alexis Roger, Gwen Legate, Kashif Rasul +2
Tokenization and transfer learning are two critical components in building state of the art time series foundation models for forecasting. In this work, we systematically study the…
cs.LG2025
Spectra 1.1: Scaling Laws and Efficient Inference for Ternary Language Models
Tejas Vaidhya, Ayush Kaushal, Vineet Jain +5
Large language models (LLMs) are increasingly used across research and industry applications, yet their inference efficiency remains a significant challenge. As the computational p…