3 papers
cs.CL2026
Global PIQA: Evaluating Commonsense Reasoning Across 100+ Languages and Cultures
Tyler A. Chang, Catherine Arnett, Abdelrahman Sadallah +377
To date, there exist almost no culturally-specific evaluation benchmarks for large language models (LLMs) that cover a large number of languages and cultures. In this paper, we pre…
cs.LG2025
End-to-End Test-Time Training for Long Context
Arnuv Tandon, Karan Dalal, Xinhao Li +11
We formulate long-context language modeling as a problem in continual learning rather than architecture design. Under this formulation, we only use a standard architecture -- a Tra…
cs.LG2025
Beyond Scores: Proximal Diffusion Models
Zhenghan Fang, Mateo DÃaz, Sam Buchanan +1
Diffusion models have quickly become some of the most popular and powerful generative models for high-dimensional data. The key insight that enabled their development was the reali…