3 papers
cs.CL2026
Cracks in the Foundation: Seemingly Minor Architectural Choices Impact Long Context Extension
Amanda Bertsch, Luca Soldaini, Matthew R. Gormley +4
One might imagine that architectural variations within the dense transformer paradigm have a limited effect on accuracy. However, we demonstrate that this is not the case in the lo…
cs.CL2025
Oolong: Evaluating Long Context Reasoning and Aggregation Capabilities
Amanda Bertsch, Adithya Pratapa, Teruko Mitamura +2
As model context lengths continue to grow, concerns about whether models effectively use the full context length have persisted. While several carefully designed long-context evalu…
cs.CL2025
In-Context Learning with Long-Context Models: An In-Depth Exploration
Amanda Bertsch, Maor Ivgi, Emily Xiao +4
As model context lengths continue to increase, the number of demonstrations that can be provided in-context approaches the size of entire training datasets. We study the behavior o…