API security 1distillation attacks 1large language models 1output perturbation defenses 1threat modeling 1
From the 1 of 4 linked papers with an AI index.
Showing cs.LGShow all
2 papers · 1 filter
cs.LG2025
Pitfalls in Evaluating Language Model Forecasters
Daniel Paleka, Shashwat Goel, Jonas Geiping +1
Large language models (LLMs) have recently been applied to forecasting tasks, with some works claiming these systems match or exceed human performance. In this paper, we argue that…
cs.LG2025
Consistency Checks for Language Model Forecasters
Daniel Paleka, Abhimanyu Pallavi Sudhir, Alejandro Alvarez +4
Forecasting is a task that is difficult to evaluate: the ground truth can only be known in the future. Recent work showing LLM forecasters rapidly approaching human-level performan…