From the 1 of 4 linked papers with an AI index.
4 papers
Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute
Nikita Kozodoi, Zainab Afolabi, Jack Butler
Test-time scaling improves LLM accuracy but multiplies inference cost, making the accuracy gained per unit of compute the metric that matters in deployment. Self-consistency is one…
Are we Merging the Right Models? Impact of Expert Training Duration on Model Merging for LLMs
Nikita Kozodoi, Zainab Afolabi, Jack Butler
The paper investigates how the training duration of domain-specific expert models influences the performance of merged large language models, finding that optimal merging strategie…
Faithful or Just Plausible? Evaluating the Faithfulness of Closed-Source LLMs in Medical Reasoning
Halimat Afolabi, Zainab Afolabi, Elizabeth Friel +13
Closed-source large language models (LLMs), such as ChatGPT and Gemini, are increasingly consulted for medical advice, yet their explanations may appear plausible while failing to…
Finding the Sweet Spot: Trading Quality, Cost, and Speed During Inference-Time LLM Reflection
Jack Butler, Nikita Kozodoi, Zainab Afolabi +2
As Large Language Models (LLMs) continue to evolve, practitioners face increasing options for enhancing inference-time performance without model retraining, including budget tuning…