works on

From the 1 of 5 linked papers with an AI index.

activity
20242026
collaborators

5 papers

cs.LG2026

Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute

Nikita Kozodoi, Zainab Afolabi, Jack Butler

Test-time scaling improves LLM accuracy but multiplies inference cost, making the accuracy gained per unit of compute the metric that matters in deployment. Self-consistency is one…

cs.LG2026

Are we Merging the Right Models? Impact of Expert Training Duration on Model Merging for LLMs

Nikita Kozodoi, Zainab Afolabi, Jack Butler

The paper investigates how the training duration of domain-specific expert models influences the performance of merged large language models, finding that optimal merging strategie…

cs.SE2026

SWE-InfraBench: Evaluating Language Models on Cloud Infrastructure Code

Natalia Tarasova, Enrique Balp-Straffon, Aleksei Iancheruk +10

Building infrastructure-as-code (IaC) in cloud computing is a critical task, underpinning the reliability, scalability, and security of modern software systems. Despite the remarka…

stat.ML2025

Finding the Sweet Spot: Trading Quality, Cost, and Speed During Inference-Time LLM Reflection

Jack Butler, Nikita Kozodoi, Zainab Afolabi +2

As Large Language Models (LLMs) continue to evolve, practitioners face increasing options for enhancing inference-time performance without model retraining, including budget tuning…

stat.ML2024

Fighting Sampling Bias: A Framework for Training and Evaluating Credit Scoring Models

Nikita Kozodoi, Stefan Lessmann, Morteza Alamgir +2

Scoring models support decision-making in financial institutions. Their estimation and evaluation are based on the data of previously accepted applicants with known repayment behav…