most citedThe relationship between reasoning and performance in large language models--o3 (mini) thinks harder, not longer

2 citations · 2 across the 6 of their papers we have counts for

collaborators

19 papers

cs.LG2026

Shorter Reasoning, Earlier Answers? An Evaluation of Reasoning Interfaces

Francesca Carlon, Vincent Ginis, Andres Algaba

Large language models often reason at length before answering, increasing cost and latency. Prompts and trained settings can shorten this reasoning, but a shorter trace may only sh…

cs.LG2026

How Much Does a Reasoning Summary Reveal? An Observability Ladder for Large Language Models

Andres Algaba, Francesca Carlon, Lynn Delcon +3

Large language models often show users a final response and a short reasoning summary while the full reasoning trace stays hidden. We introduce an observability ladder that holds e…

cs.CR2026

Geometric Configurations of Perturbed Jailbreak Prompts

Lynn Delcon, Andres Algaba, Vincent Ginis

Perturbation techniques that turn unsuccessful jailbreak prompts into successful ones are continuously evolving, constituting a major security threat to LLM safety. In this paper,…

cs.LG20262 cited

The relationship between reasoning and performance in large language models--o3 (mini) thinks harder, not longer

Marthe Ballon, Andres Algaba, Vincent Ginis

Large language models have demonstrated remarkable progress in mathematical reasoning, leveraging chain-of-thought and reinforcement learning. However, many open questions remain r…

cs.DL2026

Rebuttals Move Peer-Review Scores, but Initial-Review Structure Bounds the Movement

Mathieu Louis, Tibo Vanleke, Vincent Ginis +1

Author rebuttals are the main post-submission window in peer review, but their effect on reviewer scores remains hard to measure because score updates mix rebuttal content with ini…

cs.CL2026

Thinking Like a Scientist? A Structural Study of LLM-Generated Research Methods

Francesca Carlon, Brecht Verbeken, Vincent Ginis +1

Large Language Models (LLMs) are increasingly used to guide research methodology, yet their default methodological tendencies under minimal prompting remain unclear. Here, we promp…