computer-use benchmarking 1interactive agents 1long-horizon tasks 1safety auditing 1tool-use evaluation 1
From the 1 of 12 linked papers with an AI index.
Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Test-Time Scaling Makes Overtraining Compute-Optimal
Nicholas Roberts, Sungjun Cho, Zhiqi Gao +7
Modern LLMs scale at test-time, e.g. via repeated sampling, where inference cost grows with model size and the number of samples. This creates a trade-off that pretraining scaling…
cs.LG2025
COSMOS: Predictable and Cost-Effective Adaptation of LLMs
Jiayu Wang, Aws Albarghouthi, Frederic Sala
Large language models (LLMs) achieve remarkable performance across numerous tasks by using a diverse array of adaptation strategies. However, optimally selecting a model and adapta…