activity
20242026
most citedPosition: Stop Anthropomorphizing Intermediate Tokens as Reasoning/Thinking Traces!

4 citations · 4 across the 1 of their papers we have counts for

collaborators
Showing cs.AIShow all

6 papers · 1 filter

cs.AI20264 cited

Position: Stop Anthropomorphizing Intermediate Tokens as Reasoning/Thinking Traces!

Subbarao Kambhampati, Karthik Valmeekam, Siddhant Bhambri +6

Intermediate token generation (ITG), where a model produces output before the solution, has become a standard method to improve the performance of language models on reasoning task…

cs.AI2025

Local Coherence or Global Validity? Investigating RLVR Traces in Math Domains

Soumya Rani Samineni, Durgesh Kalwar, Vardaan Gangal +2

Reinforcement Learning with Verifiable Rewards (RLVR)-based post-training of Large Language Models (LLMs) has been shown to improve accuracy on reasoning tasks and continues to att…

cs.AI2024

Incorporating Human Flexibility through Reward Preferences in Human-AI Teaming

Siddhant Bhambri, Mudit Verma, Upasana Biswas +2

Preference-based Reinforcement Learning (PbRL) has made significant strides in single-agent settings, but has not been studied for multi-agent frameworks. On the other hand, modeli…

cs.AI2024

LLMs Can't Plan, But Can Help Planning in LLM-Modulo Frameworks

Subbarao Kambhampati, Karthik Valmeekam, Lin Guan +5

There is considerable confusion about the role of Large Language Models (LLMs) in planning and reasoning tasks. On one side are over-optimistic claims that LLMs can indeed do these…

cs.AI2024

Robust Planning with LLM-Modulo Framework: Case Study in Travel Planning

Atharva Gundawar, Mudit Verma, Lin Guan +3

As the applicability of Large Language Models (LLMs) extends beyond traditional text processing tasks, there is a burgeoning interest in their potential to excel in planning and re…

cs.AI2024

On the Brittle Foundations of ReAct Prompting for Agentic Large Language Models

Mudit Verma, Siddhant Bhambri, Subbarao Kambhampati

The reasoning abilities of Large Language Models (LLMs) remain a topic of debate. Some methods such as ReAct-based prompting, have gained popularity for claiming to enhance sequent…