2 citations · 2 across the 2 of their papers we have counts for
5 papers
How Much Does a Reasoning Summary Reveal? An Observability Ladder for Large Language Models
Andres Algaba, Francesca Carlon, Lynn Delcon +3
Large language models often show users a final response and a short reasoning summary while the full reasoning trace stays hidden. We introduce an observability ladder that holds e…
The relationship between reasoning and performance in large language models--o3 (mini) thinks harder, not longer
Marthe Ballon, Andres Algaba, Vincent Ginis
Large language models have demonstrated remarkable progress in mathematical reasoning, leveraging chain-of-thought and reinforcement learning. However, many open questions remain r…
Probing the Trajectories of Reasoning Traces in Large Language Models
Marthe Ballon, Brecht Verbeken, Vincent Ginis +1
Large language models (LLMs) increasingly solve difficult problems by producing "reasoning traces" before emitting a final response. However, it remains unclear how accuracy and de…
Benchmarks Saturate When The Model Gets Smarter Than The Judge
Marthe Ballon, Andres Algaba, Brecht Verbeken +1
Benchmarks are important tools to track progress in the development of Large Language Models (LLMs), yet inaccuracies in datasets and evaluation methods consistently undermine thei…
Estimating problem difficulty without ground truth using Large Language Model comparisons
Marthe Ballon, Andres Algaba, Brecht Verbeken +1
Recent advances in the finetuning of large language models (LLMs) have significantly improved their performance on established benchmarks, emphasizing the need for increasingly dif…