2 citations · 2 across the 7 of their papers we have counts for
10 papers · 1 filter
Shorter Reasoning, Earlier Answers? An Evaluation of Reasoning Interfaces
Francesca Carlon, Vincent Ginis, Andres Algaba
Large language models often reason at length before answering, increasing cost and latency. Prompts and trained settings can shorten this reasoning, but a shorter trace may only sh…
How Much Does a Reasoning Summary Reveal? An Observability Ladder for Large Language Models
Andres Algaba, Francesca Carlon, Lynn Delcon +3
Large language models often show users a final response and a short reasoning summary while the full reasoning trace stays hidden. We introduce an observability ladder that holds e…
The relationship between reasoning and performance in large language models--o3 (mini) thinks harder, not longer
Marthe Ballon, Andres Algaba, Vincent Ginis
Large language models have demonstrated remarkable progress in mathematical reasoning, leveraging chain-of-thought and reinforcement learning. However, many open questions remain r…
Scalable Classification of Course Information Sheets Using Large Language Models: A Reusable Institutional Method for Academic Quality Assurance
Brecht Verbeken, Joke Van den Broeck, Inge De Cleyn +4
Purpose: Higher education institutions face increasing pressure to audit course designs for generative AI (GenAI) integration. This paper presents an end-to-end method for using la…
Humanity's Last Exam
Long Phan, Alice Gatti, Ziwen Han +1144
Benchmarks are important tools for tracking the rapid advancements in large language model (LLM) capabilities. However, benchmarks are not keeping pace in difficulty: LLMs now achi…
Probing the Trajectories of Reasoning Traces in Large Language Models
Marthe Ballon, Brecht Verbeken, Vincent Ginis +1
Large language models (LLMs) increasingly solve difficult problems by producing "reasoning traces" before emitting a final response. However, it remains unclear how accuracy and de…