From the 2 of 41 linked papers with an AI index.
41 papers
Thinking Hard, Not Smart: Reasoning Models Fail to Ration Test-Time Compute Across Questions
Chenrui Fan, Yize Cheng, Ming Li +3
Reasoning language models increasingly use test-time compute to improve performance, but existing evaluations typically study this compute one question at a time. Yet when multiple…
Retrieval is Enough: Training-Free Interpretability with a Tool-Using Agent
Sriram Balasubramanian, Soheil Feizi
Interpretability methods for neural network activations span a wide cost spectrum, from cheap, training-free techniques (such as linear probes, PCA, SVD) to more expensive training…
Do Agent Optimizers Compound? A Continual-Learning Evaluation on Terminal-Bench 2.0
Wenxiao Wang, Priyatham Kattakinda, Soheil Feizi
The paper evaluates whether gains from agent-optimization methods compound over successive optimization phases in a continual‑learning setting, using hard tasks from Terminal‑Bench…
SpurLens: Automatic Detection of Spurious Cues in Multimodal LLMs
Parsa Hosseini, Sumit Nawathe, Mazda Moayeri +2
The paper introduces SpurLens, an automated pipeline that uses GPT-4 and open-set object detectors to find spurious visual cues in multimodal large language models, showing that th…
When is Your LLM Steerable?
Chenrui Fan, Yize Cheng, Ming Li +2
Activation steering offers a lightweight approach to control language models' behavior at inference time, but whether it succeeds or fails heavily depends on the prompt, concept, m…
Off-the-Shelf LLMs as Process Scorers: Training-Free Alternative to PRMs for Mathematical Reasoning
Atoosa Chegini, Soheil Feizi
Selecting the best response from multiple small-model samples using a stronger scorer is a simple inference-time strategy, but fails when the small model has already committed to i…