3 papers
cs.LG2026
Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute
Nikita Kozodoi, Zainab Afolabi, Jack Butler
Test-time scaling improves LLM accuracy but multiplies inference cost, making the accuracy gained per unit of compute the metric that matters in deployment. Self-consistency is one…
cs.LG2026
Are we Merging the Right Models? Impact of Expert Training Duration on Model Merging for LLMs
Nikita Kozodoi, Zainab Afolabi, Jack Butler
Multi-task model merging combines separately trained expert models into a single model that handles all tasks without co-training. Standard practice merges experts at their optimal…
stat.ML2025
Finding the Sweet Spot: Trading Quality, Cost, and Speed During Inference-Time LLM Reflection
Jack Butler, Nikita Kozodoi, Zainab Afolabi +2
As Large Language Models (LLMs) continue to evolve, practitioners face increasing options for enhancing inference-time performance without model retraining, including budget tuning…