1 citations · 2 across the 3 of their papers we have counts for
4 papers
Learning to Discover at Test Time
Mert Yuksekgonul, Daniel Koceja, Xinhao Li +8
How can we use AI to discover a new state of the art for a scientific problem? Prior work in test-time scaling, such as AlphaEvolve, performs search by prompting a frozen LLM. We p…
To Err Is Human: Systematic Quantification of Errors in Published AI Papers via LLM Analysis
Federico Bianchi, Yongchan Kwon, Zachary Izzo +2
How many mistakes do published AI papers contain? Peer-reviewed publications form the foundation upon which new research and knowledge are built. Errors that persist in the literat…
ReasonIF: Large Reasoning Models Fail to Follow Instructions During Reasoning
Yongchan Kwon, Shang Zhu, Federico Bianchi +2
The ability of large language models (LLMs) to follow user instructions is central to their reliability, safety, and usefulness. While prior studies assess instruction adherence in…
Dynamic Cheatsheet: Test-Time Learning with Adaptive Memory
Mirac Suzgun, Mert Yuksekgonul, Federico Bianchi +2
Despite their impressive performance on complex tasks, current language models (LMs) typically operate in a vacuum: Each input query is processed separately, without retaining insi…