Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
BODHI: Do LLMs Branch Out and Discover Heterogeneous Inferences?
Soumadeep Saha, Krish Sharma, Akshay Chaturvedi +1
Although reinforcement learning with verifiable rewards (RLVR) has improved the performance of large language models (LLMs) across a variety of reasoning tasks, there is significan…
cs.CL2025
DIMSUM: Discourse in Mathematical Reasoning as a Supervision Module
Krish Sharma, Niyar R Barman, Akshay Chaturvedi +1
We look at reasoning on GSM8k, a dataset of short texts presenting primary school, math problems. We find, with Mirzadeh et al. (2024), that current LLM progress on the data set ma…