3 papers
cs.LG2026
Curriculum Reinforcement Learning from Easy to Hard Tasks Improves LLM Reasoning
Shubham Parashar, Shurui Gui, Xiner Li +8
We aim to improve the reasoning capabilities of language models via reinforcement learning (RL). Recent RL post-trained models like DeepSeek-R1 have demonstrated reasoning abilitie…
cs.AI2025
Complex LLM Planning via Automated Heuristics Discovery
Hongyi Ling, Shubham Parashar, Sambhav Khurana +6
We consider enhancing large language models (LLMs) for complex planning tasks. While existing methods allow LLMs to explore intermediate steps to make plans, they either depend on…
cs.AI2025
Inference-Time Computations for LLM Reasoning and Planning: A Benchmark and Insights
Shubham Parashar, Blake Olson, Sambhav Khurana +4
We examine the reasoning and planning capabilities of large language models (LLMs) in solving complex tasks. Recent advances in inference-time techniques demonstrate the potential…