10 papers
Thinking Hard, Not Smart: Reasoning Models Fail to Ration Test-Time Compute Across Questions
Chenrui Fan, Yize Cheng, Ming Li +3
Reasoning language models increasingly use test-time compute to improve performance, but existing evaluations typically study this compute one question at a time. Yet when multiple…
When is Your LLM Steerable?
Chenrui Fan, Yize Cheng, Ming Li +2
Activation steering offers a lightweight approach to control language models' behavior at inference time, but whether it succeeds or fails heavily depends on the prompt, concept, m…
Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy
Kazem Faghih, Yize Cheng, Shoumik Saha +3
Large language models (LLMs) often achieve strong accuracy on benchmarks, yet it remains unclear how reliably they apply this knowledge when the same question is phrased in differe…
Model-Adaptive Tool Necessity Reveals the Knowing-Doing Gap in LLM Tool Use
Yize Cheng, Chenrui Fan, Mahdi JafariRaviz +2
Large language models (LLMs) increasingly act as autonomous agents that must decide when to answer directly vs. when to invoke external tools. Prior work studying adaptive tool use…
Schoenfeld's Anatomy of Mathematical Reasoning by Language Models
Ming Li, Chenrui Fan, Yize Cheng +2
Large language models increasingly expose reasoning traces, yet their underlying cognitive structure and steps remain difficult to identify and analyze beyond surface-level statist…
Your LLM Agents are Temporally Blind: The Misalignment Between Tool Use Decisions and Human Time Perception
Yize Cheng, Arshia Soltani Moakhar, Chenrui Fan +5
Large language model (LLM) agents are increasingly used to interact with and execute tasks in dynamic environments. However, a critical yet overlooked limitation of these agents is…