most citedThe relationship between reasoning and performance in large language models--o3 (mini) thinks harder, not longer

2 citations · 2 across the 2 of their papers we have counts for

collaborators

5 papers

cs.LG2026

How Much Does a Reasoning Summary Reveal? An Observability Ladder for Large Language Models

Andres Algaba, Francesca Carlon, Lynn Delcon +3

Large language models often show users a final response and a short reasoning summary while the full reasoning trace stays hidden. We introduce an observability ladder that holds e…

cs.LG20262 cited

The relationship between reasoning and performance in large language models--o3 (mini) thinks harder, not longer

Marthe Ballon, Andres Algaba, Vincent Ginis

Large language models have demonstrated remarkable progress in mathematical reasoning, leveraging chain-of-thought and reinforcement learning. However, many open questions remain r…

cs.LG2026

Probing the Trajectories of Reasoning Traces in Large Language Models

Marthe Ballon, Brecht Verbeken, Vincent Ginis +1

Large language models (LLMs) increasingly solve difficult problems by producing "reasoning traces" before emitting a final response. However, it remains unclear how accuracy and de…

cs.AI2026

Benchmarks Saturate When The Model Gets Smarter Than The Judge

Marthe Ballon, Andres Algaba, Brecht Verbeken +1

Benchmarks are important tools to track progress in the development of Large Language Models (LLMs), yet inaccuracies in datasets and evaluation methods consistently undermine thei…

cs.LG2025

Estimating problem difficulty without ground truth using Large Language Model comparisons

Marthe Ballon, Andres Algaba, Brecht Verbeken +1

Recent advances in the finetuning of large language models (LLMs) have significantly improved their performance on established benchmarks, emphasizing the need for increasingly dif…