9 papers
Olmo 3
Team Olmo, :, Allyson Ettinger +66
We introduce Olmo 3, a family of state-of-the-art, fully-open language models at the 7B and 32B parameter scales. Olmo 3 model construction targets long-context reasoning, function…
Not-Just-Scaling Laws: Towards a Better Understanding of the Downstream Impact of Language Model Design Decisions
Emmy Liu, Amanda Bertsch, Lintang Sutawika +9
Improvements in language model capabilities are often attributed to increasing model size or training data, but in some cases smaller models trained on curated data or with differe…
FicSim: A Dataset for Multi-Faceted Semantic Similarity in Long-Form Fiction
Natasha Johnson, Amanda Bertsch, Maria-Emil Deal +1
As language models become capable of processing increasingly long and complex texts, there has been growing interest in their application within computational literary studies. How…
Oolong: Evaluating Long Context Reasoning and Aggregation Capabilities
Amanda Bertsch, Adithya Pratapa, Teruko Mitamura +2
As model context lengths continue to grow, concerns about whether models effectively use the full context length have persisted. While several carefully designed long-context evalu…
Prompt-MII: Meta-Learning Instruction Induction for LLMs
Emily Xiao, Yixiao Zeng, Ada Chen +3
A popular method to adapt large language models (LLMs) to new tasks is in-context learning (ICL), which is effective but incurs high inference costs as context length grows. In thi…
Efficient Many-Shot In-Context Learning with Dynamic Block-Sparse Attention
Emily Xiao, Chin-Jou Li, Yilin Zhang +2
Many-shot in-context learning has recently shown promise as an alternative to finetuning, with the major advantage that the same model can be served for multiple tasks. However, th…