3 papers
cs.CL2026
Retrievable Gradients: Continual Post-Training Without Cumulative Weight Drift
Weihang Su, Jiacheng Kang, Jingyan Xu +7
Continual post-training enables models to absorb emerging knowledge after deployment, but repeatedly updating shared parameters can accumulate weight drift, potentially causing cat…
cs.LG2026
Continuous Latent Contexts Enable Efficient Online Learning in Transformers
Emile Anand, Abdullah Ateyeh, Xinyuan Cao +1
Large language models (LLMs) exhibit a strong capacity for in-context learning: Given labeled examples, they can generate good predictions without parameter updates. However, many…
cs.LG2025
Provable Long-Range Benefits of Next-Token Prediction
Xinyuan Cao, Santosh S. Vempala
Why do modern language models, trained to do well on next-word prediction, appear to generate coherent documents and capture long-range structure? Here we show that next-token pred…