Showing cs.LGShow all
2 papers · 1 filter
cs.LG2025
End-to-End Test-Time Training for Long Context
Arnuv Tandon, Karan Dalal, Xinhao Li +11
We formulate long-context language modeling as a problem in continual learning rather than architecture design. Under this formulation, we only use a standard architecture -- a Tra…
cs.LG2025
Beyond Scores: Proximal Diffusion Models
Zhenghan Fang, Mateo DÃaz, Sam Buchanan +1
Diffusion models have quickly become some of the most popular and powerful generative models for high-dimensional data. The key insight that enabled their development was the reali…