5 papers
Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute
Nikita Kozodoi, Zainab Afolabi, Jack Butler
Test-time scaling improves LLM accuracy but multiplies inference cost, making the accuracy gained per unit of compute the metric that matters in deployment. Self-consistency is one…
Are we Merging the Right Models? Impact of Expert Training Duration on Model Merging for LLMs
Nikita Kozodoi, Zainab Afolabi, Jack Butler
Multi-task model merging combines separately trained expert models into a single model that handles all tasks without co-training. Standard practice merges experts at their optimal…
SWE-InfraBench: Evaluating Language Models on Cloud Infrastructure Code
Natalia Tarasova, Enrique Balp-Straffon, Aleksei Iancheruk +10
Building infrastructure-as-code (IaC) in cloud computing is a critical task, underpinning the reliability, scalability, and security of modern software systems. Despite the remarka…
Finding the Sweet Spot: Trading Quality, Cost, and Speed During Inference-Time LLM Reflection
Jack Butler, Nikita Kozodoi, Zainab Afolabi +2
As Large Language Models (LLMs) continue to evolve, practitioners face increasing options for enhancing inference-time performance without model retraining, including budget tuning…
ChemPile: A 250GB Diverse and Curated Dataset for Chemical Foundation Models
Adrian Mirza, Nawaf Alampara, Martiño Ríos-García +12
Foundation models have shown remarkable success across scientific domains, yet their impact in chemistry remains limited due to the absence of diverse, large-scale, high-quality da…