From the 2 of 50 linked papers with an AI index.
19 papers · 1 filter
Thinking Hard, Not Smart: Reasoning Models Fail to Ration Test-Time Compute Across Questions
Chenrui Fan, Yize Cheng, Ming Li +3
Reasoning language models increasingly use test-time compute to improve performance, but existing evaluations typically study this compute one question at a time. Yet when multiple…
When is Your LLM Steerable?
Chenrui Fan, Yize Cheng, Ming Li +2
Activation steering offers a lightweight approach to control language models' behavior at inference time, but whether it succeeds or fails heavily depends on the prompt, concept, m…
Off-the-Shelf LLMs as Process Scorers: Training-Free Alternative to PRMs for Mathematical Reasoning
Atoosa Chegini, Soheil Feizi
Selecting the best response from multiple small-model samples using a stronger scorer is a simple inference-time strategy, but fails when the small model has already committed to i…
SpecHop: Continuous Speculation for Accelerating Multi-Hop Retrieval Agents
Mehrdad Saberi, Keivan Rezaei, Soheil Feizi
Large language models increasingly use external tools such as web search and document retrieval to solve information-intensive tasks. However, multi-hop tool use in complex tasks i…
Schoenfeld's Anatomy of Mathematical Reasoning by Language Models
Ming Li, Chenrui Fan, Yize Cheng +2
Large language models increasingly expose reasoning traces, yet their underlying cognitive structure and steps remain difficult to identify and analyze beyond surface-level statist…
Your LLM Agents are Temporally Blind: The Misalignment Between Tool Use Decisions and Human Time Perception
Yize Cheng, Arshia Soltani Moakhar, Chenrui Fan +5
Large language model (LLM) agents are increasingly used to interact with and execute tasks in dynamic environments. However, a critical yet overlooked limitation of these agents is…