Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Thinking at the Right Size: Amortized Distillation Across Post-Trained LLMs
Yan Zhou, Sara Kangaslahti, Jonathan Geuter +4
Practical deployment of large language models (LLMs) requires families of post-trained variants---instruction-tuned, reasoning-tuned, and chat-style models---each at multiple sizes…
cs.LG2026
Black-Box Assisted Regression: Phase Transitions and Minimax Optimality
Yan Zhou
Foundation models are often used as fixed black-box predictors for downstream tasks with limited labeled data, but their predictions may be biased and unsafe to trust blindly. We s…