3 papers
cs.AI2025
Reasoning Models Know When They're Right: Probing Hidden States for Self-Verification
Anqi Zhang, Yulin Chen, Jane Pan +4
Reasoning models have achieved remarkable performance on tasks like math and logical reasoning thanks to their ability to search during reasoning. However, they still suffer from o…
cs.CL2025
Measuring LLM Novelty As The Frontier Of Original And High-Quality Output
Vishakh Padmakumar, Chen Yueh-Han, Jane Pan +2
As large language models (LLMs) are increasingly used for ideation and scientific discovery, it is important to evaluate their ability to generate novel output. Prior work evaluate…
cs.HC2025
When Benchmarks Talk: Re-Evaluating Code LLMs with Interactive Feedback
Jane Pan, Ryan Shar, Jacob Pfau +3
Programming is a fundamentally interactive process, yet coding assistants are often evaluated using static benchmarks that fail to measure how well models collaborate with users. W…