2 papers
cs.CL2026
Cracks in the Foundation: Seemingly Minor Architectural Choices Impact Long Context Extension
Amanda Bertsch, Luca Soldaini, Matthew R. Gormley +4
One might imagine that architectural variations within the dense transformer paradigm have a limited effect on accuracy. However, we demonstrate that this is not the case in the lo…
cs.LG2026
Olmo Hybrid: From Theory to Practice and Back
William Merrill, Yanhong Li, Tyler Romero +19
Recent work has demonstrated the potential of non-transformer language models, especially linear recurrent neural networks (RNNs) and hybrid models that mix recurrence and attentio…