activity
20212026
most citedAdaptive Sampling: Algorithmic vs. Human Waypoint Selection

2 citations · 2 across the 11 of their papers we have counts for

collaborators
Showing cs.LGShow all

7 papers · 1 filter

cs.LG2026

Thinking at the Right Size: Amortized Distillation Across Post-Trained LLMs

Yan Zhou, Sara Kangaslahti, Jonathan Geuter +4

Practical deployment of large language models (LLMs) requires families of post-trained variants---instruction-tuned, reasoning-tuned, and chat-style models---each at multiple sizes…

cs.LG2026

Understanding Layer Patching in Model Size Interpolation

Sara Kangaslahti, Jonathan Geuter, Nihal V. Nayak +3

Zero-shot model size interpolation aims to create new models of intermediate target sizes by combining existing models without additional training. Recent work on boomerang distill…

cs.LG2026

Inverse Depth Scaling From Most Layers Being Similar

Yizhou Liu, Sara Kangaslahti, Ziming Liu +1

Neural scaling laws relate loss to model size in large language models (LLMs), yet depth and width may contribute to performance differently, requiring more detailed studies. Here,…

cs.LG2025

Analyzing Political Text at Scale with Online Tensor LDA

Sara Kangaslahti, Danny Ebanks, Jean Kossaifi +3

This paper proposes a topic modeling method that scales linearly to billions of documents. We make three core contributions: i) we present a topic modeling method, Tensor Latent Di…

cs.LG2025

Boomerang Distillation Enables Zero-Shot Model Size Interpolation

Sara Kangaslahti, Nihal V. Nayak, Jonathan Geuter +3

Large language models (LLMs) are typically deployed under diverse memory and compute constraints. Existing approaches build model families by training each size independently, whic…

cs.LG2025

Hidden Breakthroughs in Language Model Training

Sara Kangaslahti, Elan Rosenfeld, Naomi Saphra

Loss curves are smooth during most of model training, so visible discontinuities stand out as possible conceptual breakthroughs. Studying these breakthroughs enables a deeper under…