Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
FragileFlow: Spectral Control of Correct-but-Fragile Predictions for Foundation Model Robustness
Zhuoyun Li, Boxuan Wang, Jinwei Hu +2
Robust adaptation of LLMs and VLMs is often evaluated by average accuracy or average consistency under perturbations. However, these averages can hide a structured failure mode: a…
cs.CL2025
Metadata Conditioning Accelerates Language Model Pre-training
Tianyu Gao, Alexander Wettig, Luxi He +3
The vast diversity of styles, domains, and quality levels present in language model pre-training corpora is essential in developing general model capabilities, but efficiently lear…
cs.CL2025
Cache Me If You Can: How Many KVs Do You Need for Effective Long-Context LMs?
Adithya Bhaskar, Alexander Wettig, Tianyu Gao +2
Language models handle increasingly long contexts for tasks such as book summarization, but this leads to growing memory costs for the key-value (KV) cache. Many prior works have p…