most citedThe relationship between reasoning and performance in large language models--o3 (mini) thinks harder, not longer

2 citations · 2 across the 7 of their papers we have counts for

collaborators
Showing cs.LGShow all

10 papers · 1 filter

cs.LG2026

Shorter Reasoning, Earlier Answers? An Evaluation of Reasoning Interfaces

Francesca Carlon, Vincent Ginis, Andres Algaba

Large language models often reason at length before answering, increasing cost and latency. Prompts and trained settings can shorten this reasoning, but a shorter trace may only sh…

cs.LG2026

How Much Does a Reasoning Summary Reveal? An Observability Ladder for Large Language Models

Andres Algaba, Francesca Carlon, Lynn Delcon +3

Large language models often show users a final response and a short reasoning summary while the full reasoning trace stays hidden. We introduce an observability ladder that holds e…

cs.LG20262 cited

The relationship between reasoning and performance in large language models--o3 (mini) thinks harder, not longer

Marthe Ballon, Andres Algaba, Vincent Ginis

Large language models have demonstrated remarkable progress in mathematical reasoning, leveraging chain-of-thought and reinforcement learning. However, many open questions remain r…

cs.LG2026

Scalable Classification of Course Information Sheets Using Large Language Models: A Reusable Institutional Method for Academic Quality Assurance

Brecht Verbeken, Joke Van den Broeck, Inge De Cleyn +4

Purpose: Higher education institutions face increasing pressure to audit course designs for generative AI (GenAI) integration. This paper presents an end-to-end method for using la…

cs.LG2026

Humanity's Last Exam

Long Phan, Alice Gatti, Ziwen Han +1144

Benchmarks are important tools for tracking the rapid advancements in large language model (LLM) capabilities. However, benchmarks are not keeping pace in difficulty: LLMs now achi…

cs.LG2026

Probing the Trajectories of Reasoning Traces in Large Language Models

Marthe Ballon, Brecht Verbeken, Vincent Ginis +1

Large language models (LLMs) increasingly solve difficult problems by producing "reasoning traces" before emitting a final response. However, it remains unclear how accuracy and de…